Korean J Anesthesiol Search

CLOSE


Korean J Anesthesiol > Volume 79(3); 2026 > Article
Zarantonello, De Cassai, Pettenuzzo, Sella, Mormando, Bolzon, Fulvio, Bertoncello, and Boscolo: Artificial intelligence in intensive care units: a scoping review addressing the translational gap to clinical practice

Abstract

Background

Critically ill patients generate large volumes of complex data, creating challenges for timely clinical decision making in intensive care units (ICUs). Artificial intelligence (AI) has emerged as a promising tool for supporting diagnosis, monitoring, prognostication, and workflow optimization in this setting. This scoping review aimed to map current AI applications in critical care and identify practical clinical applications.

Methods

A systematic search of the MEDLINE, Scopus, and EMBASE databases was conducted for studies published between January 2015 and June 2025. Eligible studies evaluated practical AI applications in ICU settings involving patients, relatives, or healthcare professionals. Data pertaining to study design, AI techniques, clinical domains, outcomes, model characteristics, and implementation features were extracted.

Results

In total, 112 studies were included. Most were retrospective observational studies (59.8%) focusing on adult populations. Machine learning was the predominant technology used (76.8%), and the main clinical applications were outcome and mortality predictions, early warning systems, and monitoring, particularly in neurological and respiratory domains. Notably, 24.1% of included studies relied on North American public databases, raising concerns about geographic data monoculture, and only 27.7% of the systems provided real-time bedside applications. Most systems remained at the experimental stage, with limited real-world implementation, heterogeneous performance reporting, and a frequent lack of external validation.

Conclusions

AI applications in ICUs have expanded rapidly and show substantial promise for improving patient care and workflow efficiency. Future research should prioritize prospective multicenter validation, explainability, and implementation science to ensure the safe and effective integration of AI into critical care.

Introduction

Critically ill patients admitted to intensive care units (ICU) generate continuous and vast amounts of data from multiple sources. It is estimated that tens of thousands of data points originating from multiparametric monitors, ventilators, infusion pumps, and electronic health records (EHRs) are recorded daily for each ICU patient [1]. The volume and complexity of this data can overwhelm physicians, particularly because therapeutic decisions in the ICU often depend on timely integration of this information. Although the importance of data for guiding ICU management has been recognized [2], studies have shown that a large portion of the data available in EHRs remain underutilized. Over 25% of the data are never used, and only approximately 33% are accessed more than half the time [3].
In this context, the use of artificial intelligence (AI) tools to process this data and provide useful information to clinicians is of interest, and many areas of intensive care medicine have already incorporated these technologies in recent years. For example, in a highly noise-polluted ICU environment where false alarm rates pose a serious challenge to staff attention, AI-driven systems may be able to filter distracting alerts and improve working conditions [4]. Moreover, to reduce medication errors, automatic systems such as barcode scanning, intelligent infusion pumps, and automated drug dispensers have been implemented [4]. AI may also be able to enhance the management of mechanical ventilation, advancing beyond the closed-loop ventilation systems that represent the most sophisticated technological solutions for ventilatory support to date [5]. Beyond patient care, AI applications may also benefit relatives by identifying factors associated with the development of post-traumatic stress disorder and supporting timely interventions [6]. Unfortunately, due to ethical and legal concerns, the vast majority of ICU data remain stored on hard drives without unlocking their full potential [1].
Given the opportunities and limitations of a topic that is relatively novel and continuously developing, we aimed to perform a scoping review of the literature related to possible applications of AI systems in critical care and, more importantly, to identify their practical applications.  

Materials and Methods

Protocol and registration

This study used the framework for scoping reviews proposed by Peters et al. [7,8]. The protocol for this scoping review was registered in the Open Science Framework (OSF) on June 24, 2025 (reference: fdbp4). This review was prepared and reported in accordance with the Preferred Reporting Items for Systematic Reviews (PRISMA) [9] and the Meta-Analyses Extension for Scoping Reviews (PRISMA-ScR) [10] guidelines, and this checklist is reported in Supplementary Material 1.

Search strategy

We conducted a systematic literature review using the MEDLINE, Scopus, and EMBASE databases to identify studies published from January 1, 2015 to June 30, 2025. The search strategy for each database is available in Supplementary Material 2. The reference lists of the included articles and relevant reviews were also manually searched to identify other potentially eligible studies.

Inclusion criteria

Inclusion criteria for this scoping review are summarized below using the PCC-T (participants, concept, context, type of studies) framework:
Participants: This review included studies involving patients admitted to the ICU, their relatives, and physicians working in the ICU. Studies were included regardless of patient age, diagnosis, or ICU type (medical, surgical, or mixed).
Concept: The concept of interest for this scoping review was the application of AI in ICU settings, including AI-based tools or systems used for diagnosis, monitoring, treatment, prognostics, resource allocation, and the improvement of ICU quality and safety.
Context: The context of this review was the ICU setting. Any study that investigated the implementation of AI tools within ICUs or assessed AI’s impact on patient care, physician workflow, or relatives’ outcomes in the ICU was eligible for inclusion.
Types of studies: This scoping review considered both quantitative and qualitative studies, including randomized controlled trials (RCTs), non-randomized interventional studies, observational studies, cohort studies, and case series that reported practical AI applications in the ICU. Only studies published in peer-reviewed journals and in English were included. Studies focusing exclusively on theoretical AI models, simulations without clinical applications, and non-English-language publications were excluded.

Extracted variables

The following variables were extracted: first author, publication year, study design, study population, AI technology used, primary use of AI, primary outcome of the study, target users, explainable AI (xAI) use, performance metrics, integration with EHRs, and practical applications of the model.
Two investigators (FZ, ADC) independently performed citation screening, and three other authors performed data extraction (AB, GAF, CB). Disagreements were resolved by discussion and, whenever necessary, by a third author (ABB). Data were entered into a custom Microsoft Excel (Microsoft Corporation) database.

Critical appraisal of studies

In line with current methodological guidance [10], scoping reviews are not intended to critically evaluate the overall quality or strength of the evidence. Rather, they are designed to map and describe the existing literature. Accordingly, the PRISMA guidelines do not recommend routine reporting of the risk of bias in scoping reviews but consider a critical appraisal of included studies to be optional. Therefore, two researchers (FZ and ADC) independently assessed the methodological quality of the included studies, focusing on the study design, sample characteristics, validation strategies, and reporting transparency.

Handling and summarizing data

Data were exclusively extracted from included manuscripts. All variables were summarized as reported by the original authors and expressed using numbers and percentages for categorical data and means ± standard deviations or medians with interquartile ranges for continuous data, as appropriate. No data transformation was performed, and no additional statistical analyses were conducted as part of this scoping review.

Results

Study selection

The PRISMA flow diagram in Fig. 1 illustrates the study selection process. In brief, 3456 records were retrieved from the databases, which were then reduced to 2326 after removal of 1130 duplicates. After title and abstract screening, 2107 records were excluded due to lack of relevance to the study topic. Of the remaining articles, 107 full-text papers were excluded, resulting in 112 studies being included in the analysis.

Critical appraisal of included studies

Analysis of the study characteristics revealed several frequently observed methodological strengths that enhanced the credibility and potential applicability of AI tools in intensive care.
• Validation approaches: studies frequently employed a variety of validation strategies, including internal validation (48 studies) and temporal or geographical external validation (18 studies). This represents a positive trend toward more rigorous model assessments.
• Large sample sizes (44 studies): utilization of large databases, such as Medical Information Mart for Intensive Care (MIMIC-III), MIMIC-IV, and eICU, provided substantial data for model development and testing.
• xAI methods (38 studies): reporting of xAI methods can help the final user (i.e., clinicians, nurses, or hospital administrators) better understand the most relevant factors for the model in use.
Common methodological limitations included:
• Limited generalizability (47 studies): often stemming from single-center designs, specific patient populations, narrow inclusion criteria, or particular healthcare system characteristics, raising questions about model performance in diverse clinical settings.
• Lack of external validation (34 studies): the absence of testing on independent cohorts from different institutions or over different time periods limits confidence in model performance across different settings and patient populations.
• Single-center design (41 studies): studies conducted at a single institution may not capture variations in clinical practice patterns, patient case mixes, data quality, or healthcare system factors that could affect the model performance.
• Limited sample size (29 studies): small datasets can lead to model overfitting, unstable predictions, or insufficient statistical power.
The strengths and limitations of the included studies are summarized in Supplementary Material 3. Some studies had both methodological strengths and limitations, which reflects the presence of partial or context-specific strengths (e.g., availability of external validation) alongside relevant methodological limitations (e.g., small or underpowered validation cohorts, single-center design, or limited generalizability). This dual classification is intentional and consistent with the aims of a scoping review, which seeks to map the breadth and characteristics of the available evidence rather than to exclude studies based on quality assessments.

Publication timeline

Publication frequency increased substantially over the study period, particularly from 2020 onwards. Specifically, lesser than 10 studies were published annually before 2020 [1122]. The period from 2020 to 2022 showed exponential growth, with 11 studies in 2020, 23 studies in 2021, and 31 studies in 2022. Subsequently, 11 studies were published in 2023, 13 in 2024, and eight in 2025 (partial year data). The temporal evolution of AI during these two time periods (2015–2019 vs. 2020–2025) is summarized in Table 1.

Study design, study population, and geographic distribution

Retrospective observational studies were the predominant study design (n = 67, 59.8%), followed by prospective observational studies (n = 37, 33.0%), and prospective interventional studies (n = 7, 6.3%). Most studies focused exclusively on adult populations (n = 95, 84.8%), with 15 studies (13.4%) addressing pediatric populations and one study (0.9%) including mixed populations.
The studies were conducted in 25 countries across five continents. The United States had the largest number of studies (n = 41, 36.6%), followed by China (n = 17, 15.2%), Germany (n = 8, 7.1%), Canada (n = 5, 4.5%), and France (n = 4, 3.6%). The United Kingdom contributed three studies (2.7%), while Belgium, Israel, South Korea, and Spain each contributed two studies (1.8% each). Other contributing countries included Australia, Saudi Arabia, Austria, Finland, Denmark, India, Iran, Taiwan, the Netherlands, Switzerland, and Turkey (n = 1 each). Nine studies involved multi-country collaborations, including a large European consortium spanning Scotland, Sweden, Lithuania, Germany, Italy, and Spain.

Data sources and databases

Studies utilized diverse data sources classified into both public and private databases. Private institutional databases predominated, with 41 monocentric (36.6%) and 41 multicentric (36.6%) studies. Public databases were utilized in 27 studies (24.1%), including the MIMIC-III (n = 8; [2228]), MIMIC-IV (n = 3; [2931]), eICU (n = 7; [3238]), MIMIC-II (n = 2; [15,39]), combined MIMIC-III and eICU databases (n = 2; [34,40]), and combined MIMIC-IV and eICU databases (n = 5; [4143]).
Three studies combined public databases with private institutional data, with one study using MIMIC-III with monocentric data [44]; one combining MIMIC-III, eICU, and monocentric data [45]; and one integrating MIMIC-III with another private database [46].

Artificial intelligence technologies and algorithms

Machine learning was the predominant AI technology employed (n = 86, 76.8%) followed by deep learning (n = 12, 10.7%). One study used large language models [32], another combined machine learning with deep learning [47], and another integrated deep learning with linear regression [48]. Twelve studies did not specify the type of AI technology used.
Among the specific algorithms, Random Forest and XGBoost were the most frequently employed (n = 19 each), followed by neural networks (n = 11), logistic regression (n = 11), Support Vector Machines (n = 10), and gradient-boosting methods (n = 8). The less commonly used algorithms included LASSO regression (n = 5); ensemble methods (n = 3); Long Short-Term Memory networks (n = 1); Gated Recurrent Units [49]; Graph Neural Networks [50]; CatBoost [42]; LightGBM [51,52]; reinforcement learning approaches, including Q-learning and SARSA [53]; and natural language processing methods [54]. Many studies employed multiple algorithms for comparative evaluations.

Model characteristics

Model explainability was explicitly addressed in 38 studies (33.9%), while 36 studies (32.1%) did not incorporate explainability features and 37 studies (33.0%) provided unclear information regarding model interpretability. Common explainability approaches included SHapley Additive exPlanations (SHAP) values [53], Local Interpretable Model-agnostic Explanations (LIME), and feature importance analysis. Several studies specifically developed interpretable models, such as risk scores [40], or utilized inherently interpretable algorithms.
Regarding operational deployment, most systems were designed for offline analysis (n = 71, 63.4%), whereas 31 studies (27.7%) developed real-time applications for clinical decision-making support at the bedside. All systems that specified their level of autonomy functioned as decision-support tools rather than autonomous systems, maintaining physician oversight in clinical decision making.

Input data types

The predominant data source was physiological monitoring (n = 71, 63.4%), including vital signs, waveform data, and continuous monitoring parameters. Laboratory values served as the primary input in 17 studies (15.2%), while imaging data (including ultrasound, MRI, and fundoscopy) were utilized in four studies (3.6%; [46,5557]). Five studies combined physiological data with laboratory values, and three studies employed clinical notes or text data [23,44,58]. One study used RNA transcriptome data. Integration with the EHR was reported in 46 studies (41.1%), whereas 22 (19.6%) did not explicitly integrate with EHR systems.

Performance metrics

Performance metrics were reported heterogeneously across studies. Sensitivity (recall), specificity, and accuracy were reported in 62 (55.4 %), 52 (46.4%), and 50 studies (44.6%), respectively. An area under the receiver operating characteristic curve (AUC-ROC) was explicitly reported in a subset of studies, although many studies presented ROC curves graphically. F1 scores were reported in 20 studies (17.9%), and precision was reported in a subset of studies. Variability in metric reporting limited direct comparisons across studies.

Clinical outcomes

The primary application purpose varied across studies, including prognosis and outcome prediction (n = 37, 33.0%), monitoring and early warning (n = 26, 23.2%), diagnosis and detection (n = 13, 11.6%), and treatment response assessments (n = 8, 7.1%). Several studies addressed multiple purposes, including combining prognosis and monitoring (n = 6), monitoring with treatment response (n = 1), and diagnosis with prognosis (n = 2). Mortality prediction was addressed in 44 studies (39.3%), whereas 27 (24.1%) examined ICU or hospital lengths of stay. Clinical outcome assessments were performed in 64 studies (57.1%).

Target users and implementation status

Physicians were the primary intended users in 92 studies (82.1%), and nursing staff were the primary users or co-users in 12 studies (10.7%). Nine studies specifically targeted both physicians and nurses as collaborators. Most systems remained at the experimental stage (n = 89, 79.5%), whereas only six studies (5.4%) evaluated commercially available systems, including DeltaScan for delirium detection [59], ExoAI for cardiac assessment [56], SGC/SGCplus systems for glycemic control [19,60], a transcriptomic classifier for sepsis, and a capillary refill assessment system [61]. Practical clinical application were reported as feasible in 83 studies (74.1%).

Implementation considerations

Privacy and data security measures were explicitly documented in 26 studies (23.2%), and informed consent processes were reported in 31 studies (27.7%). Barriers to implementation, including data quality issues, integration challenges with existing clinical workflows, regulatory considerations, and the need for prospective validation, were discussed in 20 studies (17.9%). Only 11 studies (9.8%) assessed user satisfaction. The majority of studies acknowledged limitations regarding external validation, with most performing validation using internal test sets or single-center data.

Clinical application domains

The clinical applications of AI in ICU settings were diverse, spanning multiple organ systems and clinical domains. Studies were categorized by their primary clinical focus, although several studies addressed multiple conditions simultaneously. Table 2 summarizes the studies in each clinical domain and their limitations to widespread adoption (e.g., missing xAI, offline use, commercial availability).

Neurological applications

Neurological applications represented the largest clinical domain (n=29; 25.9%). Studies addressed traumatic brain injury management and outcome prediction [6265], intracranial pressure prediction and monitoring [45,66,67], delirium detection and prediction [14,59,6871], encephalopathy identification [23], post-cardiac arrest neurological outcomes [7274], anoxic-ischemic brain injury assessment [75], spinal cord injury outcomes [51], neonatal neurological monitoring [44,7678], agitation prediction in mechanically ventilated patients [79], and responsiveness phenotyping.

Pulmonary and respiratory applications

Pulmonary applications constituted the second largest domain (n=30; 26.8%). Studies addressed COVID-19 mortality and prediction of ICU lengths of stay [32,8085], ARDS detection and prediction [15,54,86], ventilator-associated pneumonia prediction [29], mechanical ventilation management including weaning prediction [12,20,37,74,87], extubation outcome prediction [38], pneumothorax detection via lung ultrasound [55], pediatric pneumonia diagnosis [57], difficult airway prediction [88], oxygen saturation monitoring [47], tidal volume estimation [48], hypoxemia assessment [36], and ECMO timing optimization [89]. Another study addressed pulmonary decompensation and hemodynamic instability [49].

Sepsis management

Sepsis-related applications were addressed in 18 (16.1%) studies. Studies focused on septic shock prediction [90], sepsis monitoring and diagnosis [18,31,50,91], sepsis-induced cardiomyopathy detection [92], antibiotic therapy optimization [13,93], treatment response prediction including with corticosteroid therapy [94,95], fluid balance management [33,53], sepsis-induced coagulopathy prediction [42], and bloodstream infection detection in pediatric patients [96]. One study addressed multiple domains including sepsis and pulmonary and metabolic applications [86].

Hemodynamic management

Hemodynamic applications were investigated in 16 (14.3%) studies. Studies addressed hypotension prediction and prevention [17,39,97], cardiogenic shock risk scoring [40], heart failure outcomes [25], cardiac function assessment using AI-enhanced ultrasound [56], hemorrhagic shock prediction in trauma [44,98], hemodynamic optimization in patients [27], vasoactive medication weaning [99], capillary refill time assessment [61], time-to-death prediction after withdrawal of life-sustaining measures [100], clinical deterioration prediction [101], and pediatric post-cardiac surgery complications [102]. One study addressed combined hemodynamic and pulmonary decompensation [49].

Renal applications

Renal applications were examined in ten studies (8.9%). Studies focused on acute kidney injury detection and prediction [11,22,30,46,103], creatinine clearance prediction [104], renal replacement therapy requirement prediction in rhabdomyolysis [34], dialysis requirement in COVID-19 patients [105], and mortality prediction in sepsis-associated acute kidney injury with continuous renal replacement therapy timing evaluations [43]. Another study addressed renal outcomes along with pulmonary and neurological outcomes in trauma patients [106].

ICU admission, discharge, and readmission

ICU admission, discharge, and readmission predictions were examined in eight studies (7.1%). Studies addressed postoperative ICU admission decisions [16], ICU readmission and mortality risk predictions [107], clinical deterioration predictions [101], ICU length of stay predictions [24,28], early warning systems for elevated-risk inpatients [108], and ICU patient phenotyping using EHR data [109].

Other clinical applications

Several additional clinical domains were identified. Metabolic applications (n = 4, 3.6%) included glycemic control [19,60], hypoglycemia prediction [35], and hypokalemia prediction in traumatic brain injury patients [110]. Nutritional applications (n = 2, 1.8%) addressed enteral nutrition optimization [21] and weight gain prediction in neonates [111]. Additional studies addressed pressure injury prediction [58], false alarm management [112], and multidomain applications spanning pulmonary, hemodynamic, metabolic, neurological, and septic conditions [41].

Discussion

This scoping review provides a comprehensive overview of the rapidly expanding landscape of AI applications in intensive care medicine. The findings demonstrate marked growth in research activity since 2020, reflecting an increasing recognition of the potential for AI to support clinical decision making in data-rich and time-critical ICU environments. Most identified studies focused on prognosis, early warning, and monitoring, with neurological and respiratory conditions emerging as the most frequently explored domains. This finding aligns with the clinical priorities of critical care, in which early detection of deterioration and outcome predictions are central to improving patient safety and resource utilization.
When comparing the two time periods (2015–2019 vs. 2020–2025), we see that not only has the number of published studies increased, the methodological sophistication has also advanced. Recent years have seen an emergence of more complex AI approaches, including deep learning and large language models, which reflects a shift toward higher-capacity architectures capable of modeling nonlinear and multimodal data. Nevertheless, most studies continue to rely on conventional machine-learning techniques, likely owing to their interpretability, lower computational demands, and established clinical applicability.
Despite this encouraging growth, important gaps remain between research development and real-world clinical implementation. The predominance of retrospective, single-center, and offline models highlights the fact that most AI tools are still in the developmental or validation phases rather than being fully integrated into routine practice.
Compared to the earlier period, the past 5 years have seen a substantial increase in the absolute number of published studies, while the proportion of AI models providing real-time data analysis and output availability has decreased (from 40% to 26%). This trend likely reflects a growing interest in AI among researchers; however, the integration of real-time AI-generated outputs into complex clinical workflows remains a substantial challenge.
Only a small proportion of studies have evaluated commercially available systems or prospectively tested AI tools at the bedside, specifically in the neurological, sepsis, hemodynamic, and metabolic domains. This disconnect suggests that even though technical performance is often promising, evidence regarding clinical effectiveness, workflow integration, and patient-centered outcomes remains limited.
Another key finding is the heterogeneity of data sources, algorithms, and performance reporting. Although large public databases, such as MIMIC and eICU, have facilitated reproducible research, a heavy reliance on single-institution datasets limits model generalizability.
Furthermore, many models developed by non-American researchers, including those from China [25,30,31,34,37,41,43,44], have been trained on US-based datasets (MIMIC-II, MIMIC-III, MIMIC-IV, and eICU), with nearly 25% of all included studies relying on these databases, and their use has increased from 13% to 21% in recent years (2020–2025). This trend reveals a concerning "data monoculture" in ICU-AI research. Models are disproportionately shaped by the clinical practices, demographic profiles, and treatment patterns of a limited number of North American institutions. Although these datasets have been instrumental in advancing the field, their overrepresentation introduces a systematic bias that may compromise model safety and accuracy when deployed in non-American healthcare systems with different patient populations, disease epidemiology, staffing models, and resource availability. This finding constitutes an urgent call to action for the global critical care community. Researchers from Asia, Europe, Africa, and Latin America should prioritize the development and public sharing of regional ICU databases that capture local disease patterns, treatment protocols, and patient demographics. Such efforts are essential not only for validating existing models across diverse settings but also for developing AI tools that are inherently designed for equitable global applicability. Privacy concerns and regulatory differences represent major barriers to cross-national data sharing, but, in the coming years, the widespread adoption of federated data access, allowing for the analysis and extraction of information from health data without physical data transfer, and federated learning, enabling AI models to be trained and updated within secure local infrastructures, may help overcome these limitations [113]. The ICU-AI field must actively transition from a paradigm of convenience-driven dataset selection to one that treats geographic and ethnic representativeness as prerequisites for clinical deployment.
In addition, inconsistent reporting of performance metrics and the limited use of calibration and clinical impact measures hinder meaningful comparisons across studies. The limited adoption of explainable AI approaches further represents a barrier to clinician trust and safe implementation, particularly in high-risk environments such as the ICU.
The “black-box” nature of AI models, in which inputs are processed into outputs without transparent reasoning, has raised concerns among users, particularly in healthcare settings, where erroneous predictions may have serious consequences. By clarifying the variables or inputs that most strongly influence a model’s output, xAI may enhance transparency and improve user understanding. However, only approximately one-third of the included studies reported the implementation of xAI approaches, highlighting an important gap in the current literature. Previous work has shown that xAI can increase clinicians’ trust in AI systems in nearly half of cases; however, its effects are not uniformly positive. Clear, concise, and clinically relevant explanations tend to enhance trust. In contrast, complex, redundant, or conflicting explanations may reduce confidence or even lead to inappropriate clinical decisions due to an overreliance on model outputs [114]. Importantly, xAI provides insight into how a model arrives at a given output, but not whether that output is correct. Although xAI remains valuable for improving transparency, future research should prioritize rigorous external validation across diverse patient populations, analogous to the evaluation processes required before introducing new pharmacological therapies into clinical practice [115].
Although 41.1% of the studies reported EHR integration, only 10.7% achieved a combination of EHR connectivity and real-time deployment [14,18,21,29,50,52,54,58,76,78,86,108], confirming that data access alone does not ensure operational availability. Future pipelines should adopt common data formats and communication protocols as architectural prerequisites to enable real-time EHR querying without institution-specific customization. A phased validation pathway — retrospective internal, prospective single-center, multicenter external — should be mandated before any clinical deployment, with each stage reported according to TRIPOD-AI [116] standards. Only 6% of the studies used an interventional design [18,54,56,59,60,97,108], underscoring the urgency of this progression.
Beyond technical integration, the human dimension of AI implementation remains underexplored. User satisfaction was formally assessed in only 9.8% of studies [16,18,30,50,54,60,63,80,93,88,93], implementation barriers in 17.9% [17,18,23,28,41,47,54,57,69,71,74,76,77,84,86,87,91,93,103,106], and workflow efficiency in 9.8% [11,38,39,41,55,60,63,80,83,91,104], indicating that algorithmic performance has been consistently prioritized over sociotechnical feasibility. Simple tools, such as the System Usability Scale, have been available for a long time and can be integrated as validation instruments in future studies. Given that these models operate as decision support tools, outcomes, such as clinician trust, cognitive load, and alert fatigue, should act as important safety endpoints in prospective trials.
Ethical, legal, and organizational challenges also emerged as significant barriers. Only a minority of studies explicitly reported on data governance, privacy safeguards, or informed consent processes, raising concerns about transparency and regulatory readiness. Moreover, few studies assessed the user experience, clinician acceptance, or long-term safety, all of which are essential for the sustainable adoption of AI in critical care.
To mitigate the poor reporting of xAI use, lack of data security measures, and slow progression from research to commercial deployment, a strict workflow should be implemented with the following characteristics: (i) alignment with local or international laws (e.g., EU AI Act); (ii) reporting of explainability methods for any high-acuity clinical applications; (iii) explicit disclosure of anonymization procedures and data-sharing agreements in all publications; and (iv) widespread adoption of commonly accepted frameworks (such as DECIDE-AI [117]) for staged early-stage clinical evaluations and safety monitoring prior to large-scale deployment.
To provide a practical synthesis of these implementation requirements, we developed a 10-point "ICU-AI Implementation Checklist" (Fig. 2), which consolidates the key recommendations from TRIPOD-AI [116], DECIDE-AI [117], and other relevant regulatory frameworks into a prescriptive roadmap organized into three phases: design & governance, development & validation, and deployment & monitoring. This tool is intended to guide researchers and clinicians in planning, conducting, and evaluating ICU-AI studies and their clinical implementation.
This scoping review had several limitations. First, only studies published in English and peer-reviewed journals were included, which may have led to language and publication biases. Secondly, although a structured methodological appraisal was performed, the inherently heterogeneous nature of the study designs, AI methods, and reported outcomes limited direct comparisons across studies. Third, as a scoping review, this study was not designed to formally assess the pooled risk of bias or to quantitatively synthesize the results. Finally, the rapid evolution of AI technologies means that some recent developments may not have been captured, potentially limiting the completeness of the current evidence map [118].
The ICU AI literature has demonstrated technical feasibility across multiple clinical domains. Demonstrating clinical safety, implementability, and governance maturity will be a defining challenge over the next decade. Future research should prioritize prospective multicenter and implementation-focused studies with a stronger emphasis on external validation, explainability, standardized reporting, and ethical governance. Bridging the gap between algorithm development and clinical impact is essential to ensure that AI applications translate into meaningful improvements in patient outcomes and safety in intensive care settings.

Funding

None.

Conflicts of Interest

No potential conflict of interest relevant to this article was reported.

Data Availability

Data sharing not applicable to this article as no datasets were generated or analyzed during the current study.

Author Contributions

Francesco Zarantonello (Conceptualization; Formal analysis; Methodology; Project administration; Writing – original draft; Writing – review & editing)

Alessandro De Cassai (Conceptualization; Data curation; Formal analysis; Investigation; Methodology; Project administration; Validation; Visualization; Writing – original draft; Writing – review & editing)

Tommaso Pettenuzzo (Writing – original draft; Writing – review & editing)

Nicolò Sella (Writing – original draft; Writing – review & editing)

Giulia Mormando (Writing – original draft; Writing – review & editing)

Annalisa Bolzon (Conceptualization; Investigation; Writing – original draft; Writing – review & editing)

Giulia Aviani Fulvio (Investigation; Writing – original draft; Writing – review & editing)

Carlo Alberto Bertoncello (Investigation; Writing – original draft; Writing – review & editing)

Annalisa Boscolo (Supervision; Writing – original draft; Writing – review & editing)

Supplementary Materials

Supplementary Material 1.
Prisma checklist.
kja-251142-Supplementary-Marterial-1.pdf
Supplementary Material 2.
Search strings.
kja-251142-Supplementary-Marterial-2.pdf
Supplementary Material 3.
Studies strengths and limitations.
kja-251142-Supplementary-Marterial-3.pdf

Fig. 1.
The PRISMA flow diagram illustrating the study selection process.
kja-251142f1.jpg
Fig. 2.
ICU-AI implementation checklist. A 10-point strategic roadmap for researchers and clinicians, organized across three phases (design & governance, development & validation, deployment & monitoring). The checklist synthesizes recommendations from current guidelines and author suggestions. The dashed arrow indicates a continuous improvement cycle linking post-deployment monitoring back to the validation phase.
kja-251142f2.jpg
Table 1.
Temporal Evolution of Artificial Intelligence Trends
Characteristic Total (n = 112) 2015–2019 (n = 15) 2020–2025 (n = 97)
Number of studies 112 (100) 15 (13.4) 97 (86.6)
AI technology
 Machine learning 86 (76.8) 9 (60.0) 77 (79.4)
 Deep learning* 13 (11.6) 1 (6.7) 12 (12.4)
 Large language model 1 (0.9) 0 (0.0) 1 (1.0)
 Unclear/Others 11 (9.8) 5 (33.3) 6 (6.2)
Data sources
 MIMIC (all versions) 23 (20.5) 2 (13.3) 21 (21.6)
 eICU 15 (13.4) 0 (0.0) 15 (15.5)
 Private–multicentric 41 (36.6) 3 (20.0) 38 (39.2)
 Private–monocentric 41 (36.6) 10 (66.7) 31 (32.0)
 Unclear/others 1 (0.9) 0 (0.0) 1 (1.0)
Data availability
 Real-time 31 (27.7) 6 (40.0) 25 (25.8)
 Offline§ 72 (64.3) 8 (53.3) 64 (66.0)
 Unclear/missing 9 (8.0) 1 (6.7) 8 (8.2)

*Deep learning includes studies reporting convolutional neural networks, recurrent neural networks, and transformer-based architectures but excluding large language models. Studies using more than one database are counted in each applicable category; percentages may therefore exceed 100% when summed. MIMIC: Medical Information Mart for Intensive Care (versions II, III, IV), eICU: eICU Collaborative Research Database. Real-time: model integrated into clinical decision-support or monitoring systems providing immediate outputs. §Offline: model applied to retrospective data without real-time clinical integration.

Table 2.
Clinical Domain Readiness and Barriers
Clinical domain* Studies (n) Commercial implementation Real-time application Explainable AI use§
Neurological 29 1 (3.4) 9 (31.0) 12 (41.4)
Respiratory 30 0 (0.0) 7 (23.3) 10 (33.3)
Sepsis 18 1 (5.6) 6 (33.3) 5 (27.8)
Hemodynamics 16 2 (12.5) 5 (31.3) 7 (43.8)
Renal 10 0 (0.0) 0 (0.0) 5 (50.0)
General ICU management 10 0 (0.0) 2 (20.0) 2 (20.0)
Metabolism/nutrition 8 2 (25.0) 4 (50.0) 2 (25.0)

*Studies pertaining to more than one domain were counted in each applicable domain. Commercial implementation: defined as studies reporting a commercially licensed AI system. Real-time application: defined as studies reporting a model integrated into a real-time clinical workflow. §Explainable AI use: defined as studies reporting use of explainability methods. AI: artificial intelligence, ICU: intensive care unit.

References

1. Lijović L, Elbers P. Leveraging the power of routinely collected ICU data. Intensive Care Med 2025; 51: 163-6.
crossref pmid pdf
2. Vincent JL. Critical care--where have we been and where are we going? Crit Care 2013; 17 Suppl 1(Suppl 1): S2.
crossref pmid pmc pdf
3. Pickering BW, Gajic O, Ahmed A, Herasevich V, Keegan MT. Data utilization for medical decision making at the time of patient admission to ICU. Crit Care Med 2013; 41: 1502-10.
crossref
4. Flint AR, Schaller SJ, Balzer F. How AI can help in error detection and prevention in the ICU? Intensive Care Med 2025; 51: 590-2.
crossref pmid pmc pdf
5. Fritsch SJ, Cecconi M. Setting the ventilator with AI support: challenges and perspectives. Intensive Care Med 2025; 51: 593-5.
crossref pdf
6. Dupont T, Kentish-Barnes N, Pochard F, Duchesnay E, Azoulay E. Prediction of post-traumatic stress disorder in family members of ICU patients: a machine learning approach. Intensive Care Med 2024; 50: 114-24.
crossref pmid pdf
7. Peters MD, Marnie C, Tricco AC, Pollock D, Munn Z, Alexander L, et al. Updated methodological guidance for the conduct of scoping reviews. JBI Evid Synth 2020; 18: 2119-26.
crossref pmid pmc
8. Peters MD, Marnie C, Tricco AC, Pollock D, Munn Z, Alexander L, et al. Updated methodological guidance for the conduct of scoping reviews. JBI Evid Implement 2021; 19: 3-10.
crossref pmid
9. Page MJ, McKenzie JE, Bossuyt PM, Boutron I, Hoffmann TC, Mulrow CD, et al. The PRISMA 2020 statement: an updated guideline for reporting systematic reviews. BMJ 2021; 372: n71.
crossref pmid pmc
10. Tricco AC, Lillie E, Zarin W, O’Brien KK, Colquhoun H, Levac D, et al. PRISMA extension for scoping reviews (PRISMA-ScR): checklist and explanation. Ann Intern Med 2018; 169: 467-73.
crossref pmid pmc
11. Ahmed A, Vairavan S, Akhoundi A, Wilson G, Chiofolo C, Chbat N, et al. Development and validation of electronic surveillance tool for acute kidney injury: a retrospective analysis. J Crit Care 2015; 30: 988-93.
crossref pmid
12. Liu Y, Mu YU, Li GQ, Yu X, Li PJ, Shen ZQ, et al. Extubation outcome after a successful spontaneous breathing trial: a multicenter validation of a 3-factor prediction model. Exp Ther Med 2015; 10: 1591-601.
crossref pmid pmc
13. Downes K, Schriver E, Russo M, Weiss S, Fitzgerald J, Balamuth F, et al. Implementation of a pragmatic biomarker-driven algorithm to guide antibiotic use in the pediatric intensive care unit: the optimizing antibiotic strategies in sepsis (OASIS) II study. Open Forum Infect Dis 2017; 4(Suppl 1): S504.
crossref pmc
14. Moon KJ, Jin Y, Jin T, Lee SM. Development and validation of an automated delirium risk assessment system (Auto-DelRAS) implemented in the electronic health record system. Int J Nurs Stud 2018; 77: 46-53.
crossref pmid
15. Taoum A, Mourad-Chehade F, Amoud H. Early-warning of ARDS using novelty detection and data fusion. Comput Biol Med 2018; 102: 191-9.
crossref pmid
16. Carrano FM, Wang B, Sherman SE, Makarov DV, Berman RS, Newman E, et al. Artificial intelligence outperforms clinical judgment in triage for postoperative ICU care: prospective preliminary results. J Am Coll Surg 2019; 229 (Suppl 1): S141-2.
crossref
17. Donald R, Howells T, Piper I, Enblad P, Nilsson P, Chambers I, et al. Forewarning of hypotensive events using a Bayesian artificial neural network in neurocritical care. J Clin Monit Comput 2019; 33: 39-51.
crossref pmid pdf
18. Giannini HM, Ginestra JC, Chivers C, Draugelis M, Hanish A, Schweickert WD, et al. A machine learning algorithm to predict severe sepsis and septic shock: development, implementation, and impact on clinical practice. Crit Care Med 2019; 47: 1485-92.
crossref pmid pmc
19. Mader JK, Motschnig M, Theiler-Schwetz V, Eibel-Reisz K, Reisinger AC, Lackner B, et al. Feasibility of blood glucose management using intra-arterial glucose monitoring in combination with an automated insulin titration algorithm in critically ill patients. Diabetes Technol Ther 2019; 21: 581-8.
crossref pmid
20. Tsai TL, Huang MH, Lee CY, Lai WW. Data science for extubation prediction and value of information in surgical intensive care unit. J Clin Med 2019; 8: 1709.
crossref pmid pmc
21. Ziemba KJ, Kumar R, Nuss K, Estrada M, Lin A, Ayad O. Clinical decision support tools and a standardized order set enhances early enteral nutrition in critically ill children. Nutr Clin Pract 2019; 34: 916-21.
crossref pmid pdf
22. Zimmerman LP, Reyfman PA, Smith AD, Zeng Z, Kho A, Sanchez-Pinto LN, et al. Early prediction of acute kidney injury following ICU admission using a multivariate panel of physiological measurements. BMC Med Inform Decis Mak 2019; 19(Suppl 1): 16.
crossref pdf
23. Ariño H, Bae SK, Chaturvedi J, Wang T, Roberts A. Identifying encephalopathy in patients admitted to an intensive care unit: Going beyond structured information using natural language processing. Front Digit Health 2023; 5: 1085602.
crossref pmid pmc
24. Khope SR, Elias S. Critical correlation of predictors for an efficient risk prediction framework of ICU patient using correlation and transformation of MIMIC-III dataset. Data Sci Eng 2022; 7: 71-86.
crossref pdf
25. Li F, Xin H, Zhang J, Fu M, Zhou J, Lian Z. Prediction model of in-hospital mortality in intensive care unit patients with heart failure: machine learning-based, retrospective analysis of the MIMIC-III database. BMJ Open 2021; 11: e044779.
crossref pmid pmc
26. Persson I, Östling A, Arlbrandt M, Söderberg J, Becedas D. A machine learning sepsis prediction algorithm for intended intensive care unit use (NAVOY Sepsis): proof-of-concept study. JMIR Form Res 2021; 5: e28000.
crossref pmid pmc
27. Roggeveen L, El Hassouni A, Ahrendt J, Guo T, Fleuren L, Thoral P, et al. Transatlantic transferability of a new reinforcement learning model for optimizing haemodynamic treatment for critically ill patients with sepsis. Artif Intell Med 2021; 112: 102003.
crossref pmid
28. Williams Batista R, Sanchez-Arias R. A methodology for estimating hospital intensive care unit length of stay using novel machine learning tools. In: 2020 19th IEEE International Conference on Machine Learning and Applications (ICMLA). Miami, FL: IEEE; 2020:827-32. Available from https://ieeexplore.ieee.org/document/9356186/. Accessed December 11, 2025.

29. Cao H, Wei J, Hua P, Yang S. Interpretable machine learning for early predicting the risk of ventilator-associated pneumonia in ischemic stroke patients in the intensive care unit. Front Neurol 2025; 16: 1513732.
crossref pmid pmc
30. Hu C, Tan Q, Zhang Q, Li Y, Wang F, Zou X, et al. Application of interpretable machine learning for early prediction of prognosis in acute kidney injury. Comput Struct Biotechnol J 2022; 20: 2861-70.
crossref pmid pmc
31. Wang Y, Gao Z, Zhang Y, Lu Z, Sun F. Early sepsis mortality prediction model based on interpretable machine learning approach: development and validation study. Intern Emerg Med 2025; 20: 909-18.
crossref pmid pmc pdf
32. Bartoszko J, Dranitsaris G, Wilcox ME, Del Sorbo L, Mehta S, Peer M, et al. Development of a repeated-measures predictive model and clinical risk score for mortality in ventilated COVID-19 patients. Can J Anaesth 2022; 69: 343-52.
crossref pdf
33. Lee B, Puri N, Patel S. 1427: Net fluid balance effect on mortality in septic ICU patients with diabetes. Crit Care Med 2022; 50: 716.
crossref
34. Liu C, Yuan Q, Mao Z, Hu P, Wu R, Liu X, et al. Development and validation of a model for the early prediction of the RRT requirement in patients with rhabdomyolysis. Am J Emerg Med 2021; 46: 38-44.
crossref pmid
35. Mantena S, Arévalo AR, Maley JH, Da Silva Vieira SM, Mateo-Collado R, Da Costa Sousa JM, et al. Predicting hypoglycemia in critically Ill patients using machine learning and electronic health records. J Clin Monit Comput 2022; 36: 1297-303.
crossref pmid pmc pdf
36. Patel S, Singh G, Zarbiv S, Ghiassi K, Rachoin JS. Mortality prediction using SaO2/FiO2 ratio based on eICU database analysis. Crit Care Res Pract 2021; 2021: 6672603.
crossref pmid pmc pdf
37. Wang H, Zhao QY, Luo JC, Liu K, Yu SJ, Ma JF, et al. Early prediction of noninvasive ventilation failure after extubation: development and validation of a machine-learning model. BMC Pulm Med 2022; 22: 304.
crossref pmid pmc pdf
38. Wong AI, Kamaleswaran R, Tabaie A, Reyna MA, Josef C, Robichaux C, et al. Prediction of acute respiratory failure requiring advanced respiratory support in advance of interventions and treatment: a multivariable prediction model from electronic medical record data. Crit Care Explor 2021; 3: e0402.
crossref pmid pmc
39. Cherifa M, Blet A, Chambaz A, Gayat E, Resche-Rigon M, Pirracchio R. Prediction of an acute hypotensive episode during an ICU hospitalization with a super learner machine-learning algorithm. Anesth Analg 2020; 130: 1157-66.
crossref pmid
40. Yamga E, Mantena S, Rosen D, Bucholz EM, Yeh RW, Celi LA, et al. Optimized risk score to predict mortality in patients with cardiogenic shock in the cardiac intensive care unit. J Am Heart Assoc 2023; 12: e029232.
crossref pmid pmc
41. Zeng Z, Liu Y, Yao S, Liu J, Xiao B, Liu C, et al. Neural networks based on attention architecture are robust to data missingness for early predicting hospital mortality in intensive care unit patients. Digit Health 2023; 9: 20552076231171482.
crossref pmid pmc pdf
42. Zhao QY, Liu LP, Luo JC, Luo YW, Wang H, Zhang YJ, et al. A machine-learning approach for dynamic prediction of sepsis-induced coagulopathy in critically ill patients with sepsis. Front Med (Lausanne) 2021; 7: 637434.
crossref pmid pmc
43. Zhuang C, Hu R, Li K, Liu Z, Bai S, Zhang S, et al. Machine learning prediction models for mortality risk in sepsis-associated acute kidney injury: evaluating early versus late CRRT initiation. Front Med (Lausanne) 2025; 11: 1483710.
crossref pmid pmc
44. Zhao T, Griffith T, Zhang Y, Li H, Hussain N, Lester B, et al. Early-life factors associated with neurobehavioral outcomes in preterm infants during NICU hospitalization. Pediatr Res 2022; 92: 1695-704.
crossref pmid pmc pdf
45. Schweingruber N, Mader MM, Wiehe A, Röder F, Göttsche J, Kluge S, et al. A recurrent machine learning model predicts intracranial hypertension in neurointensive care patients. Brain 2022; 145: 2910-9.
crossref pmid pmc pdf
46. Shi S. A novel hybrid deep learning architecture for predicting acute kidney injury using patient record data and ultrasound kidney images. Appl Artif Intell 2021; 35: 1329-45.
crossref
47. Radhakrishnan S, Sreedhar GJ, VA B, Akbar R, PB D, Nair SG. Maintaining blood oxygen saturation level with a deep learning neural network model using real-patient data. In: 2024 1st International Conference on Innovative Engineering Sciences and Technological Research (ICIESTR). Muscat, Oman: IEEE; 2024:1-6. Available from https://ieeexplore.ieee.org/document/10798175/. Accessed December 11, 2025.

48. Yang HL, Park SA, Lee HY, Lee H, Ryu HG. Feasibility of estimating tidal volume from electrocardiograph-derived respiration signal and respiration waveform. J Crit Care 2025; 85: 154920.
crossref pmid
49. Mandel C, Stich K, Autexier S, Luth C, Ziehn A, Hochbaum K, et al. Using gated recurrent unit networks for the prediction of hemodynamic and pulmonary decompensation. In: 2022 44th Annual International Conference of the IEEE Engineering in Medicine & Biology Society (EMBC). Glasgow, Scotland: IEEE; 2022:4584-4589. Available from https://ieeexplore.ieee.org/document/9871500/. Accessed December 11, 2025.

50. Teredesai A, Huang S, Stewart T, Hu J, Thakker A, Stern K, et al. Sub-sequence graph representation learning on high variability data for dynamic risk prediction in critical care. In: 2022 IEEE International Conference on Big Data (Big Data). Osaka, Japan: IEEE; 2022:2082-2092. Available from https://ieeexplore.ieee.org/document/10020963/. Accessed December 11, 2025.

51. Karabacak M, Margetis K. Precision medicine for traumatic cervical spinal cord injuries: accessible and interpretable machine learning models to predict individualized in-hospital outcomes. Spine J 2023; 23: 1750-63.
crossref pmid
52. Lin YH, Chang TC, Liu CF, Lai CC, Chen CM, Chou W. The intervention of artificial intelligence to improve the weaning outcomes of patients with mechanical ventilation: Practical applications in the medical intensive care unit and the COVID-19 intensive care unit: A retrospective study. Medicine (Baltimore) 2024; 103: e37500.
crossref pmid pmc
53. Su L, Li Y, Liu S, Zhang S, Zhou X, Weng L, et al. Establishment and implementation of potential fluid therapy balance strategies for ICU sepsis patients based on reinforcement learning. Front Med (Lausanne) 2022; 9: 766447.
crossref pmid pmc
54. Knighton AJ, Kuttler KG, Ranade-Kharkar P, Allen L, Throne T, Jacobs JR, et al. An alert tool to promote lung protective ventilation for possible acute respiratory distress syndrome. JAMIA Open 2022; 5: ooac050.
crossref pmid pmc pdf
55. Clausdorff Fiedler H, Prager R, Smith D, Wu D, Dave C, Tschirhart J, et al. Automated real-time detection of lung sliding using artificial intelligence: a prospective diagnostic accuracy study. CHEST 2024; 166: 362-70.
crossref pmid
56. Gallant C, Bernard L, Kwok C, Wichuk S, Noga M, Punithakumar K, et al. AI-augmented point of care ultrasound in intensive care unit patients: can novices perform a “basic echo” to estimate left ventricular ejection fraction in this acute-care setting? J Clin Med 2025; 14: 2899.
crossref
57. Kessler D, Zhu M, Gregory CR, Mehanian C, Avila J, Avitable N, et al. Development and testing of a deep learning algorithm to detect lung consolidation among children with pneumonia using hand-held ultrasound. PLoS One 2024; 19: e0309109.
crossref
58. Lee JH, Yu JY, Shim SY, Yeom KM, Ha HA, Jekal SY, et al. Development of a pressure injury machine learning prediction model and integration into clinical practice: a prediction model development and validation study. Korean J Adult Nurs 2024; 36: 191-202.
crossref pdf
59. Ditzel FL, Hut SC, Van Den Boogaard M, Boonstra M, Leijten FS, Wils EJ, et al. DeltaScan for the assessment of acute encephalopathy and delirium in ICU and non-ICU patients, a prospective cross-sectional multicenter validation study. Am J Geriatr Psychiatry 2024; 32: 1093-104.
crossref pmid
60. González-Caro MD, Fernández-Castillo RJ, Carmona-Pastor M, Arroyo-Muñoz FJ, González-Fernández FJ, Garnacho-Montero J. Effectiveness and safety of the Space GlucoseControl system for glycaemia control in caring for postoperative cardiac surgical patients. Aust Crit Care 2022; 35: 136-42.
crossref pmid
61. Hunter RB, Jiang S, Nishisaki A, Nickel AJ, Napolitano N, Shinozaki K, et al. Supervised machine learning applied to automate flash and prolonged capillary refill detection by pulse oximetry. Front Physiol 2020; 11: 564589.
crossref pmid pmc
62. Bhattacharyay S, Caruso PF, Åkerlund C, Wilson L, Stevens RD, Menon DK, et al. Mining the contribution of intensive care clinical course to outcome after traumatic brain injury. NPJ Digit Med 2023; 6: 154.
crossref
63. Carra G, Güiza F, Depreitere B, Meyfroidt G. Prediction model for intracranial hypertension demonstrates robust performance during external validation on the CENTER-TBI dataset. Intensive Care Med 2021; 47: 124-6.
crossref pmid pdf
64. Moffet EW, Subramaniam T, Hirsch LJ, Gilmore EJ, Lee JW, Rodriguez-Ruiz AA, et al. Validation of the 2HELPS2B seizure risk score in acute brain injury patients. Neurocrit Care 2020; 33: 701-7.
crossref pmid pdf
65. Palepu A, Murali A, Ballard J, Li R, Ramesh S, Sarma S, et al. 6: A machine learning model to predict brain trauma outcome in the intensive care unit. Crit Care Med 2021; 49: 3.
crossref
66. Fong N, Feng J, Hubbard A, Dang LE, Pirracchio R. IntraCranial pressure prediction algorithm using machine learning (I-CARE): training and validation study. Crit Care Explor 2023; 6: e1024.
crossref pmid pmc
67. Nortvig MJ, Andersen MC, Eriksen NL, Aunan-Diop JS, Pedersen CB, Poulsen FR. Utilizing retinal arteriole/venule ratio to estimate intracranial pressure. Acta Neurochir (Wien) 2024; 166: 445.
crossref pmid pmc pdf
68. Gong K, Lu R, Bergamaschi T, Sanyal A, Guo J, Kim H, et al. 27: Machine learning prediction of intensive care unit delirium. Crit Care Med 2021; 49: 14.
crossref
69. Jeffcock J, Hansen M, Ruiz Garate V. Transformers and human-robot interaction for delirium detection. In: Proceedings of the 2023 ACM/IEEE International Conference on Human-Robot Interaction. Stockholm Sweden: ACM, 2023:466-74. Available from https://dl.acm.org/doi/10.1145/3568162.3576971. Accessed December 11, 2025.
crossref
70. Lu Y, Li Y, Chi S, Feng Y, Li G, Lin X, et al. Comparison of machine learning and logistic regression models for predicting emergence delirium in elderly patients: a prospective study. Int J Med Inform 2025; 199: 105888.
crossref pmid
71. Sun H, Kimchi E, Akeju O, Nagaraj SB, McClain LM, Zhou DW, et al. Automated tracking of level of consciousness and delirium in critical illness using deep learning. Npj Digit Med 2019; 2: 89.
crossref pmid pmc pdf
72. Hunfeld M, Verboom M, Josemans S, Van Ravensberg A, Straver D, Lückerath F, et al. Prediction of survival after pediatric cardiac arrest using quantitative EEG and machine learning techniques. Neurology 2024; 103: e210043.
crossref pmid
73. Sun B, Lei M, Wang L, Wang X, Li X, Mao Z, et al. Prediction of sepsis among patients with major trauma using artificial intelligence: a multicenter validated cohort study. Int J Surg 2025; 111: 467-80.
crossref pmid pmc
74. Sung CW, Shieh JS, Chang WT, Lee YW, Lyu JH, Ong HN, et al. Machine learning analysis of heart rate variability for the detection of seizures in comatose cardiac arrest survivors. IEEE Access 2020; 8: 160515-25.
crossref
75. Mattia GM, Sarton B, Villain E, Vinour H, Ferre F, Buffieres W, et al. Multimodal MRI-based whole-brain assessment in patients in anoxoischemic coma by using 3D convolutional neural networks. Neurocrit Care 2022; 37(Suppl 2): 303-12.
crossref pmid pmc pdf
76. Duncan HP, Fule B, Rice I, Sitch AJ, Lowe D. Wireless monitoring and real-time adaptive predictive indicator of deterioration. Sci Rep 2020; 10: 11366.
crossref pmid pmc pdf
77. Huang D, Yu D, Zeng Y, Song X, Pan L, He J, et al. Generalized camera-based infant sleep-wake monitoring in NICUs: a multi-center clinical trial. IEEE J Biomed Health Inform 2024; 28: 3015-28.
crossref pmid
78. Moghadam SM, Airaksinen M, Nevalainen P, Marchi V, Hellström-Westas L, Stevenson NJ, et al. An automated bedside measure for monitoring neonatal cortical activity: a supervised deep learning-based electroencephalogram classifier with external cohort validation. Lancet Digit Health 2022; 4: e884-92.
crossref pmid
79. Zhang Z, Liu J, Xi J, Gong Y, Zeng L, Ma P. Derivation and validation of an ensemble model for the prediction of agitation in mechanically ventilated patients maintained under light sedation. Crit Care Med 2021; 49: e279-90.
crossref pmid
80. Alabbad DA, Almuhaideb AM, Alsunaidi SJ, Alqudaihi KS, Alamoudi FA, Alhobaishi MK, et al. Machine learning model for predicting the length of stay in the intensive care unit for Covid-19 patients in the eastern province of Saudi Arabia. Inform Med Unlocked 2022; 30: 100937.
crossref pmid pmc
81. Churpek MM, Gupta S, Spicer AB, Hayek SS, Srivastava A, Chan L, et al. Machine learning prediction of death in critically ill patients with Coronavirus Disease 2019. Crit Care Explor 2021; 3: e0515.
crossref
82. Lerner RK, Baharav N, Chen O, Levi A, Lipsky A, Celniker G, et al. 171: Using artificial intelligence models to predict deterioration in critically ill COVID-19 patients. Crit Care Med 2022; 50: 69.
crossref
83. Shanbehzadeh M, Nopour R, Kazemi-Arpanahi H. Using decision tree algorithms for estimating ICU admission of COVID-19 patients. Inform Med Unlocked 2022; 30: 100919.
crossref pmid pmc
84. Sievering AW, Wohlmuth P, Geßler N, Gunawardene MA, Herrlinger K, Bein B, et al. Comparison of machine learning methods with logistic regression analysis in creating predictive models for risk of critical in-hospital events in COVID-19 patients on hospital admission. BMC Med Inform Decis Mak 2022; 22: 309.
crossref pmid pmc pdf
85. Magunia H, Lederer S, Verbuecheln R, Gilot BJ, Koeppen M, Haeberle HA, et al. Machine learning identifies ICU outcome predictors in a multicenter COVID-19 cohort. Crit Care 2021; 25: 295.
crossref pmid pmc pdf
86. Chen JT, Mehrizi R, Aasman B, Gong MN, Mirhaji P. Long short-term memory model identifies ARDS and in-hospital mortality in both non-COVID-19 and COVID-19 cohort. BMJ Health Care Inform 2023; 30: e100782.
crossref pmid pmc
87. Su L, Zhang Z, Zheng F, Pan P, Hong N, Liu C, et al. Five novel clinical phenotypes for critically ill patients with mechanical ventilation in intensive care units: a retrospective and multi database study. Respir Res 2020; 21: 325.
crossref pdf
88. Rodiera C, Fortuny H, Valls A, Borras R, Ramírez C, Ros B, et al. Voice analysis as a method for preoperatively predicting a difficult airway based on machine learning algorithms: original research report. Health Sci Rep 2024; 7: e70246.
crossref pmid pmc
89. Giraud R, Legouis D, Assouline B, De Charriere A, Decosterd D, Brunner ME, et al. Timing of VV-ECMO therapy implementation influences prognosis of COVID-19 patients. Physiol Rep 2021; 9: e14715.
crossref pmid pmc pdf
90. Agor JK, Li R, Özaltın OY. Septic shock prediction and knowledge discovery through temporal pattern mining. Artif Intell Med 2022; 132: 102406.
crossref pmid
91. Benov A, Brand A, Rozenblat T, Antebi B, Ben-Ari A, Amir-Keret R, et al. Evaluation of sepsis using compensatory reserve measurement: a prospective clinical trial. J Trauma Acute Care Surg 2020; 89(2S Suppl 2): S153-60.
crossref pmid
92. Bollen Pinto B, Ribas Ripoll V, Subías-Beltrán P, Herpain A, Barlassina C, Oliveira E, et al. Application of an exploratory knowledge-discovery pipeline based on machine learning to multi-scale OMICS data to characterise myocardial injury in a cohort of patients with septic shock: an observational study. J Clin Med 2021; 10: 4354.
crossref pmid pmc
93. Düvel JA, Lampe D, Kirchner M, Elkenkamp S, Cimiano P, Düsing C, et al. An AI-based clinical decision support system for antibiotic therapy in sepsis (KINBIOTICS): use case analysis. JMIR Hum Factors 2025; 12: e66699.
crossref
94. König R, Kolte A, Ahlers O, Oswald M, Krauss V, Roell D, et al. Use of IFNγ/IL10 ratio for stratification of hydrocortisone therapy in patients with septic shock. Front Immunol 2021; 12: 607217.
crossref pmid pmc
95. Pirracchio R, Hubbard A, Sprung CL, Chevret S, Annane D. Assessment of machine learning to estimate the individual treatment effect of corticosteroids in septic shock. JAMA Netw Open 2020; 3: e2029050.
crossref pmid pmc
96. Han HJ, Kim K, Park JD. Early detection of bloodstream infection in critically ill children using artificial intelligence. Acute Crit Care 2024; 39: 611-20.
crossref pmid pmc pdf
97. Schneck E, Schulte D, Habig L, Ruhrmann S, Edinger F, Markmann M, et al. Hypotension Prediction Index based protocolized haemodynamic management reduces the incidence and duration of intraoperative hypotension in primary total hip arthroplasty: a single centre feasibility randomised blinded prospective interventional trial. J Clin Monit Comput 2020; 34: 1149-58.
crossref pmid pdf
98. Lang E, Neuschwander A, Favé G, Abback PS, Esnault P, Geeraerts T, et al. Clinical decision support for severe trauma patients: Machine learning based definition of a bundle of care for hemorrhagic shock and traumatic brain injury. J Trauma Acute Care Surg 2022; 92: 135-43.
crossref pmid
99. Goldsmith MP, Nadkarni VM, Futterman C, Gazit AZ, Baronov D, Tomczak A, et al. Use of a risk analytic algorithm to inform weaning from vasoactive medication in patients following pediatric cardiac surgery. Crit Care Explor 2021; 3: e0563.
crossref pmid pmc
100. Scales NB, Herry CL, Van Beinum A, Hogue ML, Hornby L, Shahin J, et al. Predicting time to death after withdrawal of life-sustaining measures using vital sign variability: derivation and validation. Crit Care Explor 2022; 4: e0675.
crossref pmid pmc
101. Verma AA, Pou-Prom C, McCoy LG, Murray J, Nestor B, Bell S, et al. Developing and validating a prediction model for death or critical illness in hospitalized adults, an opportunity for human-computer collaboration. Crit Care Explor 2023; 5: e0897.
crossref pmid pmc
102. Zeng X, An J, Lin R, Dong C, Zheng A, Li J, et al. Prediction of complications after paediatric cardiac surgery. Eur J Cardiothorac Surg 2020; 57: 350-8.
crossref pmid pdf
103. Neyra JA, Ortiz-Soriano V, Liu LJ, Smith TD, Li X, Xie D, et al. Prediction of mortality and major adverse kidney events in critically ill patients with acute kidney injury. Am J Kidney Dis 2023; 81: 36-47.
crossref pmid pmc
104. De Vlieger G, Huang CY, Pörteners B, Güiza F, Meyfroidt G. Creatinine clearance in critically ill adults: prospective comparison of prediction by intensive care unit physicians and machine learning models. Intensive Care Med 2024; 50: 1532-4.
crossref pmid pdf
105. Vaid A, Chan L, Chaudhary K, Jaladanki SK, Paranjpe I, Russak A, et al. Predictive approaches for acute dialysis requirement and death in COVID-19. Clin J Am Soc Nephrol 2021; 16: 1158-68.
crossref pmid pmc
106. Moris D, Henao R, Hensman H, Stempora L, Chasse S, Schobel S, et al. Multidimensional machine learning models predicting outcomes after trauma. Surgery 2022; 172: 1851-9.
crossref pmid
107. Dam TA, De Bruin D, Cinà G, Thoral PJ, Elbers PW, Den Uil CA, et al. ICU readmission and mortality risk prediction: generalizability of a multi-hospital model. J Intensive Med 2025; 5: 377-84.
crossref pmid pmc
108. Winslow CJ, Edelson DP, Churpek MM, Taneja M, Shah NS, Datta A, et al. The impact of a machine learning early warning score on hospital mortality: a multicenter clinical intervention trial. Crit Care Med 2022; 50: 1339-47.
crossref pmid
109. Wang Y, Stroh JN, Hripcsak G, Low Wang CC, Bennett TD, Wrobel J, et al. A methodology of phenotyping ICU patients from EHR data: High-fidelity, personalized, and interpretable phenotypes estimation. J Biomed Inform 2023; 148: 104547.
crossref pmid pmc
110. Zhou Z, Huang C, Fu P, Huang H, Zhang Q, Wu X, et al. Prediction of in-hospital hypokalemia using machine learning and first hospitalization day records in patients with traumatic brain injury. CNS Neurosci Ther 2023; 29: 181-91.
crossref pmid pmc pdf
111. Yalcin N, Çelik HT, Demirkan K, Yiğit S. Machine learning algorithms to predict weight gain at discharge in neonatal intensive care unit: state of the art. Clin Nutr ESPEN 2021; 46: S724-5.
crossref pmc
112. Liu Z, Khojandi A, Mohammed A, Li X, Chinthala LK, Davis RL, et al. HeMA: A hierarchically enriched machine learning approach for managing false alarms in real time: a sepsis prediction case study. Comput Biol Med 2021; 131: 104255.
crossref
113. Van Genderen ME, Cecconi M, Jung C. Federated data access and federated learning: improved data sharing, AI model development, and learning in intensive care. Intensive Care Med 2024; 50: 974-7.
crossref pmid pmc pdf
114. Rosenbacke R, Melhus Å, McKee M, Stuckler D. How explainable artificial intelligence can increase or decrease clinicians’ trust in AI applications in health care: systematic review. JMIR AI 2024; 3: e53207.
crossref
115. Ghassemi M, Oakden-Rayner L, Beam AL. The false hope of current approaches to explainable artificial intelligence in health care. Lancet Digit Health 2021; 3: e745-50.
crossref pmid
116. Collins GS, Moons KG, Dhiman P, Riley RD, Beam AL, Van Calster B, et al. TRIPOD+AI statement: updated guidance for reporting clinical prediction models that use regression or machine learning methods. BMJ 2024; 385: e078378.
crossref pmid pmc
117. Vasey B, Nagendran M, Campbell B, Clifton DA, Collins GS, Denaxas S, et al. Reporting guideline for the early stage clinical evaluation of decision support systems driven by artificial intelligence: DECIDE-AI. BMJ 2022; 377: e070904.
crossref pmid pmc
118. Dost B, Turan Eİ, Aydın ME, Ahıskalıoğlu A, Narayanan M, Yılmaz R, et al. Artificial intelligence in anaesthesiology: current applications, challenges, and future directions. Turk J Anaesthesiol Reanim 2025; 53: 282-92.
crossref pmid pmc


ABOUT
ARTICLE CATEGORY

Browse all articles >

BROWSE ARTICLES
AUTHOR INFORMATION
Editorial Office
101-3503, Lotte Castle President, 109 Mapo-daero, Mapo-gu, Seoul 04146, Korea
Tel: +82-2-792-5128    Fax: +82-2-792-4089    E-mail: journal@anesthesia.or.kr                
Business Name: Korean Society of Anesthesiologists
Business Registration: 106-82-07194
Representative: Young-Tae Jeon

Copyright © 2026 by Korean Society of Anesthesiologists.

Developed in M2PI

Close layer
prev next