Abstract
Background. In recent years, advancements in artificial intelligence (AI) have led to the creation of numerous large language models, including Chat Generative Pre-trained Transformer (ChatGPT), a natural language processing model developed by OpenAI. There has been increasing interest in exploring the potential of this tool, as evidenced by several studies in the healthcare domain that have investigated the performance of ChatGPT in responding to specific questions on relevant topics. However, its application in forensic contexts remains largely unexplored.
Objectives. This study aims to evaluate the performance of ChatGPT in estimating the post-mortem interval (PMI) by considering thanatochronological changes as well as through the application of the Henssge nomogram.
Materials and methods. Fifteen questions concerning PMI estimation were presented to ChatGPT. These questions were categorized into 3 domains, each addressing different aspects: 1) theoretical notions, 2) practical application of theoretical notions, and 3) utilization of the Henssge nomogram. The evaluation of responses included 3 sub-criteria: focus, accuracy, and completeness.
Results. The results showed a high level of accuracy and completeness in addressing theoretical issues. However, the AI failed when responding to practical case scenarios and when calculating PMI using the Henssge nomogram.
Conclusions. Although AI may be attractive for application to forensic science questions, uncertain source documents and incomplete access to the scientific literature may affect accuracy. The study concluded that AI should be investigated for use in forensic science; however, it is currently not suitable for practical PMI estimation.
Key words: artificial intelligence, post-mortem interval, performance evaluation, ChatGPT, post-mortem interval estimation
Background
Artificial intelligence (AI) is a discipline that explores the construction of computer systems capable of simulating human thought, problem solving, and decision-making.1, 2, 3 In recent years, significant advances in AI have led to the emergence of many large language models (LLMs). These include Chat Generative Pre-trained Transformer (ChatGPT; OpenAI, San Francisco, USA), a natural language processing model developed by OpenAI, which is currently one of the largest publicly available language models.4, 5
When ChatGPT was launched in November 2022, it represented an unprecedented technological revolution that captured global attention and sparked debate.6, 7 ChatGPT uses deep learning techniques to generate human-like responses to natural language input. The model uses massive datasets, powerful computational resources, and efficient algorithms to build an intelligent system capable of extracting information from vast amounts of textual data and generating complex responses. ChatGPT facilitates multi-touch human–computer interaction by providing responses and feedback in text form.4, 7, 8, 9 The most recent free version currently available and most widely used is ChatGPT-3.5.4, 8, 10 In the past year, interest in the potential of LLMs has grown, with several studies published in the healthcare field evaluating ChatGPT’s ability to respond to specific questions on topics of major interest in various medical specialties.11, 12, 13
AI technology has been applied to sex and age estimation in forensic anthropology, evaluation of tooth development in forensic odontology, and taxonomic identification of diatoms in forensic pathology.14 However, although some papers have been published in the forensic literature on the molecular evaluation of postmortem microbiomes and eye opacity,15, 16 to the best of our knowledge, the performance of chatbots such as ChatGPT in responding to questions about postmortem changes and the interpretation of postmortem interval estimation has not been evaluated.
Indeed, ChatGPT does not yet have specific forensic applications but has been used in cybercrime investigations to analyze complex textual evidence and assist in identifying suspects or missing persons.17, 18, 19 For forensic pathologists involved in post-mortem examinations, the estimation of post-mortem interval (PMI) is a crucial and challenging issue. The estimation of PMI is a complex assessment involving various factors, and it remains a topic of academic debate because it has a considerable margin of error. Numerous articles have been published in the scientific literature regarding methods for its evaluation,20, 21 including the mathematical method based on the construction of the Henssge nomogram. The Henssge nomogram is a commonly used tool for estimating the PMI using body temperature, ambient temperature, and body weight. The Henssge nomogram is most accurate for PMIs of less than 48 h and has some limitations, even though its mathematical model has been widely studied.22, 23, 24, 25, 26, 27, 28 This topic remains cumbersome, especially in real-world forensic casework, and the determination of a reliable and precise time of death often remains impossible. Post-mortem interval estimation is a well-studied topic within forensic science, with an abundance of published studies. It was hypothesized that ChatGPT may be able to generate responses to questions about the PMI, assuming that the relevant scientific literature is accessible to the AI.
Objectives
Considering the potential and limitations of LLMs in forensic practice, this exploratory study aims to assess the performance of ChatGPT in responding to forensic questions and case vignettes related to PMI estimation. The evaluation focuses on 3 distinct domains – general theoretical knowledge, applied knowledge, and application of the Henssge nomogram – and relies on expert ratings of relevance, accuracy, and completeness to appraise the AI-generated responses.
Materials and methods
The study was designed by adapting the methodology of previous studies conducted in clinical settings to the selected forensic topic.11, 12, 13 Questions of increasing difficulty were formulated by experts in the scientific field of interest. The questions were then used to interact with ChatGPT by a single user, and the responses were recorded for further analysis. The responses had to be assessed by 1 or more experts, who assigned a score to each response to indicate whether it was appropriate or not. Having multiple evaluators assess a response can help reduce biases based on personal experiences, perspectives, or preferences, and increase the accuracy and reliability of the evaluation.11, 12, 13
Three Italian forensic pathologists with >5 years of experience were recruited from a single institution to generate a total of 15 questions of increasing difficulty for ChatGPT. The questions were collaboratively developed in English. The experts formulated the questions clearly and avoided ambiguous wording. Three different domains with 5 questions each were created. Domain 1 contained 5 questions concerning theoretical notions related to thanatological changes in PMI estimation. Domain 2 contained 5 questions on the application of theoretical notions to practical cases. Domain 3 contained 5 questions on practical cases and the application of the Henssge nomogram. For illustrative purposes, the questions posed to ChatGPT are presented in quotation marks. The 5 questions within each domain are listed as follows (see Supplementary data).
Domain 1: Theoretical notions on the parameters used in the evaluation of the PMI
1. “Please describe the timing of hypostasis formation in thanatochronology.”
2. “How quickly does rigor mortis develop and dissipate?”
3. “Please describe the post-mortem changes in body temperature known as algor mortis.”
4. “What are the phenomena of post-mortem eye dehydration, and at what point in time do they typically occur?”
5. “Describe the stages of corpse putrefaction and their timing.”
Domain 2: Application of theoretical notions to practical cases
1. “What is the post-mortem interval with fixed hypostasis and rigor mortis resolved in all joints?”
2. “What is the post-mortem interval with fixed hypostasis, even if imprints are still visible, and present, persistent rigor mortis in all joints?”
3. “What is the post-mortem interval with evident rigor mortis in the neck and arm muscles but absent in the leg muscles?”
4. “What is the post-mortem interval with evident rigor mortis in the leg muscles but absent in the arm and neck muscles?”
5. “What is the post-mortem interval with present rigor mortis in the neck, arm, and leg muscles?”
Domain 3: Application of the Henssge nomogram
1. “According to the Henssge nomogram, calculate the time since death for a person weighing 120 kg, with a rectal temperature of 35°C, found in an environment with an ambient temperature of 15°C.”
2. “According to the Henssge nomogram, calculate the time since death for a person weighing 80 kg, with a rectal temperature of 35°C, found in a location with an ambient temperature of 15°C.”
3. “According to the Henssge nomogram, calculate the time since death for a person weighing 80 kg, with a rectal temperature of 35°C, found completely naked in an outdoor location with an ambient temperature of 15°C.”
4. “According to the Henssge nomogram, calculate the time since death for a person weighing 80 kg, with a rectal temperature of 35°C, found indoors wearing a T-shirt, sweatshirt, and winter jacket in an ambient temperature of 15°C.”
5. “According to the Henssge nomogram, calculate the time since death for a person weighing 80 kg, with a rectal temperature of 35°C, found in a river with running water and an ambient temperature of 15°C.”
The 3 experts wrote the scripted questions, and one of them entered the questions into the professional version of ChatGPT-3.5, available by subscription (date of queries: July 10, 2024), in the same order in which they are shown in the list above. The generation of responses was restricted to a single day to avoid the possibility of the addition of sources available to the program, which could potentially have provided more detailed responses if the questions had been asked in the following days or weeks. For each question, a new chatbot session was opened, and a response time within 10 s was observed. The response provided by ChatGPT was copied and pasted into a plain-text document (Microsoft Word 2016; Microsoft Corp., Armonk, USA) without any modification.
According to previous literature,11 we chose 3 experts for response evaluation to reduce bias. In particular, the responses were shared with 3 different Italian forensic pathologists with >5 years of experience, working in the same institution as the question authors. This distinct group scored the responses in terms of focus, accuracy, and completeness. The experts individually scored the responses using a Likert scale. Following the initial scoring, the 3 experts met and collaboratively developed a single evaluation and comment regarding the ChatGPT responses; they then collaboratively provided an overall score for each response based on each assessment. In cases of conflicting opinions, agreement among the specialists was reached by comparing literature sources and discussing the issue. The specialists were always able to reach agreement.
Using a semi-quantitative 5-point Likert-like scale (rating 1: insufficient, rating 2: barely sufficient, rating 3: sufficient, rating 4: good, rating 5: excellent), the evaluation considered 3 sub-criteria for each response:
– focus (rating 1: topic of the question not identified; rating 5: topic of the question properly identified);
– accuracy (rating 1: response completely incorrect; rating 5: response correct);
– completeness (ChatGPT failed to respond to part or all of a question; rating 1: response incomplete; rating 5: response complete).
The evaluation of the sub-criteria was carried out by comparing the responses with the literature20, 21, 22, 23, 24, 25, 29, 30, 31 and by manually applying the Henssge nomogram for calculations. For each response, a commentary by the team of experts was included to explain the rationale for the ratings assigned to the various sub-criteria (Supplementary Tables 1–3).
Results
Due to the length of the experts’ responses and comments, full results are provided in Supplementary Tables 1–3, and a summary of the results for the sub-criteria (focus, accuracy, and completeness) is shown in Table 1. Regarding focus, in all responses, the 3 pathologists who performed the evaluation assigned the maximum score (5); therefore, the focus score reached the maximum value for each question. Accuracy showed variable values in Domain 1 (Theoretical notions), where 3 out of 5 responses were completely correct. In contrast, this variability was not observed in the other 2 domains. In fact, accuracy received the lowest score for all responses in the other 2 domains (Application of theoretical notions and Application of the Henssge nomogram). Completeness was the sub-criterion that showed the greatest variation among the domains, although the values were generally lower in Domain 3 than in the other 2 domains. ChatGPT occasionally provided exhaustive information on theoretical aspects related to the specific case but failed to extract the pertinent details necessary to answer the actual question.
Discussion
The estimation of PMI represents a major challenge for forensic experts. Its assessment requires the integration of theoretical insights regarding essential thanatochronological factors (hypostasis and rigor mortis), along with the results of mathematical calculations using the Henssge nomogram concerning algor mortis.21, 22, 23, 24, 25, 26 Evidence gathered in recent studies regarding the use of AI in the forensic field14, 15, 16, 17 has shown promising results; therefore, this is a topic that deserves further exploration to understand the potential this program offers. Among the various AI systems available, ChatGPT has been one of the most widely used in the evaluation of medical issues in recent years, although, to the best of our knowledge, it has never been applied to answer specific forensic questions. Given that estimation of the time since death through application of the Henssge nomogram is a fundamental task for the forensic pathologist, a better understanding of ChatGPT’s capabilities in this regard is necessary. In the present study, ChatGPT’s performance was evaluated through questions regarding both theoretical concepts and practical applications of PMI estimation (domains), and scores were assigned for 3 distinct sub-criteria: focus, accuracy, and completeness. A more detailed analysis of the individual scores revealed different strengths and weaknesses in ChatGPT’s performance in this forensic context.
At first, a notable finding is that the “focus” sub-criterion received an excellent rating, with the highest score across all responses. This could be due to the careful formulation of questions by forensic pathologists, avoiding ambiguous or difficult-to-understand phrases or terms. This precaution was implemented to prevent the so-called “hallucinations” in ChatGPT’s responses. In fact, numerous studies have reported that ChatGPT may be overly sensitive to variations in question phrasing and may struggle to clarify ambiguous prompts. This sensitivity might lead to the phenomenon of “hallucination”, manifested as logical inconsistencies or nonsensical responses relative to the posed question.32, 33, 34, 35
The results for “completeness” are interesting, with responses to theoretical questions being quite complete, while responses to practical application tasks were incomplete. The responses were generally well organized by the AI, often being subdivided into relevant paragraphs, aiding in the review and comprehension of the response. It is noteworthy that the lower “completeness” scores in the practical application domain should be considered a major limitation to the use of ChatGPT in this forensic context. In particular, the field of PMI calculation requires solutions based on the application of concepts and formulas, which ChatGPT does not seem able to provide.
The high scores found for the “accuracy” sub-criterion in Domain 1 (Theoretical notions) are noteworthy, considering that ChatGPT’s access to scientific literature is restricted to freely available sources or public datasets, whereas a significant portion of the scientific literature requires specific permissions and agreements to access articles or textbooks. In evaluating the accuracy of responses, it is crucial to consider the quality of the cited scientific sources. Indeed, as is well known, scientific articles can be classified as either open-access journals or subscription-based journals. It is not possible to determine whether ChatGPT has access to open-access journals and/or subscription-based journals. This implies that when the authors evaluated the accuracy of the responses in the present study, if a response was found to be inaccurate, it was not possible to establish whether this was due to a selection bias in the sources consulted.
On the other hand, the lowest accuracy values were reported in the responses to Domain 2 (Application of theoretical notions) and Domain 3 (Application of the Henssge nomogram), where the AI was required to evaluate a particular real case and calculate the PMI. This finding appears to be the most significant in the present study and deserves particular attention due to its multiple potential implications in practice.
Indeed, it is notable that ChatGPT provides the least accurate information when requested to apply theoretical notions to practical cases, thereby reporting incorrect answers concerning the specific calculation of the PMI. Moreover, each question in Domain 3 contained all the necessary information (rectal temperature, ambient temperature, and body weight), as well as additional data required for the correction factor (clothing and environmental factors) used in the Henssge nomogram calculation. In all cases where ChatGPT cited misleading calculations and provided a numerical PMI value in its response, these values were totally inaccurate, differing considerably from those derived from the nomogram calculation. It is notable that these responses sometimes included an initially correct theoretical introduction explaining the Henssge nomogram and outlining the factors influencing it, along with partial instructions on how to use it.
It is also worth noting that in Domains 2 and 3, all responses received the lowest possible score for accuracy (i.e., 1 on the Likert scale), indicating a complete failure of the model to provide responses in the application of theoretical notions to practical cases (Domain 2) and PMI estimation using the Henssge nomogram (Domain 3). While this could support a binary (accurate/inaccurate) interpretation, we chose to retain the 5-point scale across all domains for consistency. In Domain 1, accuracy was often partial and heterogeneous, and a Likert-based assessment allowed a more nuanced evaluation. In Domains 2 and 3, however, the uniformity of the ratings means that the practical implications of a binary classification would be identical to those of the current scale. For these reasons, we opted for a unified evaluative framework, while recognizing that in this specific case, the accuracy dimension already reflects a binary outcome in practice.
Previous authors have already addressed the issue of the non-transparent nature of ChatGPT’s source of information.36, 37, 38 The lack of a reference list, which makes it impossible to assess the accuracy of the information presented by ChatGPT, has been reported. On these grounds, it is only possible to speculate about ChatGPT’s inability to respond to nomogram questions. Several considerations may provide insight into this observation. First, in contrast to the more general theoretical and practical questions in Domains 1 (Theoretical notions) and 2 (Practical application of theoretical notions), the questions in Domain 3 (Utilization of the Henssge nomogram) required the application of a mathematical approach employing, e.g., open-source online tools for calculations using the Henssge nomogram based on specific mathematical algorithms.34, 35
It is evident that ChatGPT, as a text-based AI model capable of performing only explicit symbolic or numerical computations, lacks the ability to interact with the drop-down menus of freely available Henssge nomogram calculators; moreover, it cannot directly employ the underlying mathematical formulas or account for the algorithmic limitations on which these calculators are based. Indeed, despite the availability of all the necessary information for the calculation, ChatGPT could not use the online calculator template. Recent research has identified the fundamental architectural reasons underlying these calculation failures. Most fundamentally, ChatGPT operates through pattern matching rather than mathematical computation, explaining why it can provide theoretically correct explanations while failing at specific numerical applications. These limitations are architectural rather than training-related, suggesting that simply providing more forensic literature would not resolve the calculation failures observed in our study. Technical solutions exist for mathematical calculation limitations through API integration, as demonstrated by the Wolfram plugin for ChatGPT, which provides access to dedicated computational engines.36
Furthermore, the complexity of the Henssge nomogram and its application to PMI estimation may be beyond the current capabilities of the chatbot model, requiring a deeper level of understanding and reasoning to utilize dedicated diagrams or mathematical models.
In addition, the lack of clarity about the quality and accuracy of the dataset used to train ChatGPT has a potential impact on the reliability of its responses.13, 37, 38, 39 As previously discussed, it is not possible to distinguish the type of journal (open-access and/or subscription-based) from which the information in the response is derived. This issue holds significant relevance in the forensic field, as it is crucial that the sources cited in court cases are rooted in scientific data. For this reason, even responses in Domains 1 (Theoretical notions) and 2 (Practical application of theoretical notions) that demonstrate the highest level of accuracy could not be introduced in a court trial due to the absence of specific references.
As previously specified, our article analyzes the application of ChatGPT in a forensic context and, in addition to other works that generally analyze the application of AI in the forensic setting, reinforces the findings of Galante et al.14 in their recent review of the application of AI in forensic medicine.
Our study shows the potential of using ChatGPT only as a complementary tool for forensic pathologists. The inability to elucidate the intricate processes by which computational methods analyze and convert input into reliable results precludes insight into why ChatGPT answers certain types of questions incorrectly. It could be interesting in the future to observe the behavior of various updated AI tools in responding to topics of growing interest in forensic pathology40, 41, 42 or other fields of forensic science.43, 44
Limitations of the study
The present study has several limitations. First, it was designed as a small-scale, exploratory analysis aimed at identifying recurring reasoning patterns and qualitative shortcomings in ChatGPT’s forensic answers, rather than at establishing generalizable estimates of performance. Second, the selection of questions and assessment of responses were based on the subjective judgment of the forensic experts who participated in the study. As stated in the Introduction, PMI estimation is a complex field of forensic pathology and requires a specific approach for each individual case. This complexity, together with the fact that the evaluation of specific cases was based on the judgment of a single group of pathologists, even if supported by the international scientific literature, might explain the reduction in the assessment of “accuracy”. As different experts may work from different premises, despite efforts to standardize the responses, the evaluations are derived from the experience of several individuals, making it challenging to grade ChatGPT responses on this particular topic. This issue was particularly relevant for theoretical notions, whereas the authors expected greater accuracy for responses involving the application of mathematical calculations using the nomogram. However, this limitation was mitigated by using 2 groups of 3 forensic pathology experts each to assess both the responses and the questions. In addition, the number of questions developed was relatively small. Nevertheless, the responses to the sub-criteria within each domain showed coherence, indicating a linear trend within each domain and reducing the need for additional questions. Moreover, as previously specified, ChatGPT’s access to scientific literature is restricted to freely available sources, which represents a significant limitation in the formulation of exhaustive responses in a scientific context. Recently, the deep-thought option, which specifies the rationale behind the provided response, has been introduced in some AI tools, including ChatGPT. Future studies could examine this functionality. Indeed, it is important to recognize that ChatGPT is a dynamic AI system that continuously learns and evolves over time, which may limit the reproducibility of the study even if the same questions are asked in the future. This experiment was conducted using ChatGPT v. 3.5 in 2024, but the well-known rapid evolution of AI may change the results over time, and ChatGPT itself may acquire new, currently unexpected skills.
This is noteworthy because it is difficult to predict the outcome of such technological progress and whether LLMs will be able to provide more consistent and reliable results in responding to forensic questions in the future. On this topic, it should also be highlighted that even the same version of the program, given the rapid pace of LLM learning, might provide different responses to the same question if asked at different times, leading to a lack of reproducibility and replicability. Finally, the exclusive use of ChatGPT-3.5 represents a significant limitation, as substantial performance differences exist between model versions. Studies demonstrate that GPT-4 achieves 87.2% accuracy compared to GPT-3.5’s 68.4% in medical examinations, representing an 18.8% absolute improvement.45 Future studies should incorporate comparative analyses of newer model versions, domain-specific AI systems, and hybrid approaches combining language models with specialized computational tools to provide a more comprehensive assessment of AI capabilities in forensic pathology.
Conclusions
This analysis of ChatGPT 3.5’s performance in the field of PMI evaluation in forensic pathology, including the application of the Henssge nomogram, revealed accuracy and completeness in responding to theoretical questions. However, the system still exhibits significant limitations when faced with practical case scenarios and mathematical calculations involving the Henssge nomogram. Therefore, although AI may be popular and appealing, this study demonstrated ChatGPT’s limitations in post-mortem interpretation. It is recommended that it not be utilized for forensic casework, as it does not replace forensic pathologist interpretation, and further research is needed to better understand the potential of such technology for future utilization in actual cases.
Supplementary data
The supplementary materials are available at https://doi.org/10.5281/zenodo.17207038. The package contains the following files:
Supplementary Table 1. Domain 1: Theoretical notions on the parameters used in the evaluation of the PMI: question, responses, and comments.
Supplementary Table 2. Domain 2: Application of the Henssge nomogram: question, response, and comments.
Supplementary Table 3. Domain 3: Application of the Henssge nomogram: question, response, and comments.
Data Availability Statement
Data sharing does not apply to this article, as all data are already included in the manuscript and in the supplementary data.
Consent for publication of personal information
Not applicable.
Use of AI and AI-assisted technologies
Not applicable.



