Korean J Anesthesiol Search

CLOSE


Korean J Anesthesiol > Volume 79(3); 2026 > Article
Kang: Correcting what cannot be corrected: rethinking publication bias analysis methods in clinical meta-analyses

Abstract

Methods to assess and adjust for publication bias are often presented as tools to correct distorted evidence in meta-analyses. However, statistical adjustment cannot recover information that was selectively generated, reported, or disseminated. Clinical evidence syntheses frequently rely on small or selective sets of trials and are characterized by substantial heterogeneity, multiple outcomes and time points, and complex dissemination pathways. Publication bias analysis methods are thus prone to over-interpretation and may yield conflicting conclusions. Therefore, they should be understood as an inferential process that links detection, model-based adjustment, and interpretation under explicit and unverifiable assumptions. We review classical methods to detect publication bias, including funnel plots, tests of small-study effects, and P value-based approaches, and demonstrate their essential role as stress tests of model adequacy rather than as definitive detectors of publication bias. We then examine widely used methods to adjust for publication bias, such as trim-and-fill, selection models, regression-based approaches relating the effect size to study precision, and the Bayesian approach, clarifying their key assumptions and typical failure modes. Using a worked example, we illustrate how applying different publication bias adjustment methods to the same evidence base can yield divergent adjusted effects, emphasizing their assumption dependence. We additionally identify common misuses, propose a framework for evaluations, and discuss emerging challenges related to preprints, umbrella reviews, and AI-assisted evidence synthesis. This review thus aims to help align the strength of clinical conclusions with the robustness or fragility of the underlying data, with direct implications for authors, reviewers, and editors.

Introduction

Publication bias is widely recognized as a major threat to the validity of meta-analyses in clinical medicine [13]. It is often described as a tendency to publish only statistically significant or favorable results; however, in practice, it reflects a much broader set of selective processes across the entire research ecosystem. Selectivity can arise during topic prioritization, protocol design, choice of outcomes and time points, analytic decisions, reporting, journal submission, peer review, and citation. By the time evidence is synthesized, the available studies may thus represent a systematically filtered subset of all research that was conducted [2,4].
Meta-analyses in clinical medicine are particularly vulnerable to such distortions. Many clinical questions are addressed by a small number of trials with variable outcome definitions, inconsistent follow-up windows, and heterogeneous patient populations and interventions. Moreover, research results now appear not only in journals but also in trial registries, preprints, conference abstracts, press releases, and regulatory documents, each with their own selection pressures and timelines [5,6]. Under these conditions, making the assumption that studies are missing solely because their results are nonsignificant is rarely realistic. Standard methods for detecting publication bias such as funnel plots and small-study effect tests are often applied in situations that differ substantially from the ideal settings for which they were originally developed.
Recent methodological work shows that publication bias can distort not only the apparent magnitude of pooled effects, but also the consistency and trustworthiness of the evidence, particularly when data are sparse and heterogeneity is substantial [711]. At the same time, the growing list of methods available to assess and adjust for publication bias, from trim-and-fill and regression-based corrections to selection models, P value-based approaches, and the Bayesian publication bias model, has made it more difficult for analysts to interpret conflicting signals and multiple “corrected” estimates [7,1215]. The proliferation of these methods has sharpened the central tension that analyses of publication bias are often treated as automatic correction tools that recover a hidden “true effect,” even though all available methods to adjust for publication bias depend on strong, untestable assumptions about the selection process [8,9,16,17].
In this review, publication bias analysis is therefore reframed as a single, coherent inferential process that integrates the methods used to detect and adjust for publication bias and their interpretations rather than as a checklist of disconnected techniques or a means to obtain one definitive bias-corrected effect [8,16,17]. The focus is on how each method encodes specific assumptions about which studies are more likely to appear in the literature, how these assumptions interact with between-study heterogeneity and selective outcome reporting, and the extent that clinical conclusions depend on these assumptions [7,14,15]. Drawing on empirical and methodological work on sensitivity analysis, selection models, and model-averaging, the methods used to detect and adjust for publication bias and interpretive guidance are combined into a single framework in this review, with practical recommendations for authors, reviewers, editors, and umbrella review teams [7,14,15]. Rather than aiming to recover a single “true” effect size, this perspective treats methods to assess and adjust for publication bias as tools for disciplined inference that help to align the strength and wording of clinical conclusions with the robustness and stability of the underlying evidence [46,9]. In this sense, publication bias methods are best viewed not as mechanical correction tools, but as structured sensitivity analyses that clarify how clinical conclusions may depend on assumptions about missing evidence.

Publication bias in the research ecosystem

Viewing publication bias only as the end result of selectively publishing significant studies obscures the larger system that produces it. Modern research environments reward novelty, clear positive messages, and statistically significant findings, so researchers are encouraged to choose questions that are likely to “work,” select outcomes and time points that make effects easier to detect, and emphasize favorable analyses when reporting results. Journals, funders, and media outlets pay more attention to studies that appear impactful, and citation patterns amplify these studies further, making them much more visible to later reviewers and meta-analysts [46,18].
“Missingness” is rarely a simple yes/no issue. Not only are some entire studies not published, but bias often arises from selective reporting within published studies, for example, presenting primary but not secondary outcomes, highlighting favorable subgroups, choosing particular analytic models, or focusing on follow-up times that yield statistically significant or directionally desirable results [8,19]. However, many standard methods to detect publication bias were developed under the assumption that entire studies are missing, rather than that certain outcomes or analyses are selectively reported, so they may respond only weakly or in misleading ways when the main problems are outcome switching, multiplicity, or data-driven analysis choices [8,19]. Consequently, meta-analyses in clinical medicine can inherit an evidence base that is distorted both by the studies that are available and by the information reported in those studies, and no statistical adjustment at the synthesis stage can fully reconstruct information that was never transparently generated, recorded, or reported [6,8,9,17,18].
Publication bias is therefore best understood as an ecosystem-level phenomenon driven by the incentives and routines that shape the questions that are asked, how studies are designed and analyzed, the framing of results, and how they circulate through scientific and clinical communities [5,6,8,16]. Within this ecosystem, methods designed to detect and adjust for publication bias are better viewed as tools for testing the robustness of findings under plausible selection and reporting mechanisms, rather than as algorithms that can recover a single underlying truth [8,9,16]. This ecosystem perspective naturally motivates a distinction between detection, adjustment, and interpretation, and underpins the integrated inferential framework developed in the following sections [8,9,16].
Beyond classic publication bias, a wide range of dissemination- and reporting-related mechanisms can distort the body of evidence available for meta-analyses. These include the delayed or selective publication of entire studies, selective reporting of outcomes and analyses within studies, preferential citation of positive results, and differential visibility of findings across journals, preprints, registries, conference abstracts, and regulatory documents [20,21]. Because these mechanisms can generate similar statistical patterns, such as small-study effects or funnel plot asymmetry, while calling for different interpretive responses, distinguishing them conceptually before introducing specific methods to detect and adjust for publication bias is a helpful strategy. Table 1 summarizes a working taxonomy of dissemination-related and reporting biases that are most relevant to meta-analyses in clinical medicine, extending the notion of “publication bias” to a broader ecosystem of selective dissemination.
A well-known example is the Vioxx (rofecoxib) scandal [2224]. Rofecoxib, a selective COX-2 inhibitor widely used for arthritis and pain, was withdrawn from the market after randomized trials and observational studies revealed an increased risk of myocardial infarction and stroke [2224]. Subsequent investigations of regulatory submissions and company documents showed that some cardiovascular events had been incompletely reported, analyzed in non-prespecified ways, or published with substantial delay, and that unfavorable findings were less visible than positive efficacy results [2224]. Conceptually, this episode maps directly onto the taxonomy shown in Table 1, combining publication and time-lag bias, outcome-reporting bias, selective analysis, and location/citation bias into a single ecosystem-level distortion of the evidence base.

Statistical detection of publication bias

Funnel plots as stress tests, not detectors

Funnel plots are often used as the initial check for publication bias in meta-analyses in clinical medicine because they are simple to make and easy to read [12,16,25,26]. They plot study effect estimates against a measure of precision and assume an ideal situation: if studies are an unbiased sample and true effects are broadly similar, less precise studies should be spread roughly symmetrically around the pooled effect. Fig. 1A illustrates this approach using the tropisetron dataset, with each circle representing an individual study positioned according to its log risk ratio (x-axis) and standard error (y-axis), with more precise studies toward the bottom. The contour lines demarcate regions of statistical significance (P < 0.05 and P < 0.01), providing additional context for interpreting asymmetry. Visual inspection reveals an apparent deficit of small, unfavorable studies in the lower-right quadrant and a cluster of small studies reporting large protective effects on the left side, creating an asymmetric funnel shape that could initially suggest publication bias.
However, when the same data are examined with additional structures such as stratification by methodological quality (Fig. 1B), the apparent asymmetry becomes less consistent with selective non-publication. In this example, several small, low-quality trials are clustered on the left side of the funnel with relatively large protective effects, whereas larger trials with higher methodological quality tend to lie closer to the null. Such a pattern can arise even when no studies are missing: studies with a higher risk of bias are more likely to overestimate treatment effects, particularly when sample sizes are small and outcome assessment is less rigorous. In this setting, the funnel plot may appear asymmetric despite being complete, reflecting an exaggeration of the risk-of-bias-related effect rather than publication bias.
In practice, true effects can differ according to baseline risk, disease severity, co-interventions, adherence, or outcome measurement methods, which are factors that often correlate with study size or design. Therefore, funnel plot asymmetry can be caused by real effect modifications, systematic differences in methods, or differences in outcome definitions even when no studies are missing [2,18]. Funnel plots are also difficult to interpret when only a few studies are available [2729]. Random variation or a single influential trial can create patterns that resemble genuine asymmetry [2729]. Conversely, a funnel plot may appear roughly symmetric even in the presence of substantial selective reporting, especially when the main source of bias is within-study selection (e.g., outcome switching or selective choice of follow-up time) rather than non-publication of entire studies [8,19].
For these reasons, funnel plots are best viewed as stress tests of how well a simple model with homogeneous effects and no selection fits the data and not as tools that identify a specific mechanism. They indicate that something is wrong with the simple model and that further investigation is needed; however, they cannot clarify whether publication bias, heterogeneity, or other factors are responsible [6,18].

Egger’s and Begg’s tests: small-study effects are not publication bias

Regression- and rank-based tests, such as Egger’s regression and Begg’s rank-correlation tests, turn funnel plot asymmetry into a statistical test by examining whether effect sizes are related to study precision [30,31]. These tests are often interpreted as direct methods for detecting publication bias. However, in reality, they detect small-study effects, a broader pattern in which smaller studies tend to report different (often larger) effects than larger studies.
Small-study effects do not indicate a single cause; publication bias is only one possible explanation. Other common reasons include real differences in effects across study settings or patient populations; systematic differences in study designs or methodological quality; selective reporting of outcomes; exploratory analyses or the early termination of small trials; and timing effects, in which small studies are conducted early, followed by later, larger, and more definitive trials.
Because of these alternative explanations, a statistically significant Egger’s or Begg’s test should be interpreted as evidence of an unexplained relationship between effect size and study precision and not as proof that publication bias is present. These tests are also sensitive to data limitations: when heterogeneity is high, false-positive results may be produced, and when only a few studies are available, they may lack power and fail to detect real problems. In some cases, the result can be driven largely by a single influential study [2,18].
Conversely, a non-significant asymmetry test does not rule out publication bias. This is particularly true when the number of studies is small, heterogeneity is substantial, or selective reporting occurs within studies rather than through non-publication of whole trials [7,8,19]. To use Egger’s and Begg’s tests appropriately, analysts should consider alternative explanations for small study effects and interpret test results alongside other information such as risk-of-bias assessments, study design features, and evidence from trial registries [2,9,18].

P value‑based methods: p‑curve, p‑uniform, and evidential value under selective reporting

P value-based methods assess publication bias by examining the distribution of statistically significant P values rather than by focusing on missing studies or study size alone. The two most commonly used approaches are the p-curve and p-uniform analyses, which are closely related but serve distinct inferential purposes [19,32].
The p-curve analysis is based on the idea that the distribution of statistically significant P values contains information on whether the reported findings reflect a genuine underlying effect or are mainly the result of selective reporting. When a true non-zero effect exists and statistical models are correctly specified, very small P values should occur more often than P values just below conventional significance thresholds, producing a right-skewed p-curve (Fig. 2A). In contrast, when the true effect is null and statistically significant results arise mainly through selective reporting, analytical flexibility, or repeated testing, significant P values tend to cluster just below the significance cutoff, yielding a flat or left-skewed distribution (Figs. 2B and C). From this perspective, the p-curve functions primarily as a diagnostic tool, evaluating whether a set of statistically significant results contains evidential value rather than estimating the magnitude of the treatment effect [19,33].
The p-uniform analysis builds on similar logic but additionally attempts to estimate a bias-adjusted effect size. It assumes that, conditional on statistical significance, P values should follow a uniform distribution under the true effect. By modeling this distribution, p-uniform analyses seek to infer the underlying effect size that best explains the observed set of significant P values, thus operating more explicitly as a model-based adjustment method, whereas the p-curve focuses on detecting the presence or absence of an evidential value.
Both approaches are appealing because they directly target the selective reporting of statistically significant results, rather than assuming that entire studies are missing. However, this strength is also their main limitation, as P value-based methods rely on several strong assumptions, namely, that statistical tests are correctly specified, analyzed P values are independent, and P values represent a well-defined and unbiased set of all relevant significant results [19,33].
From an editorial and methodological perspective, p-curve results require careful interpretation and may be unreliable in several common clinical research settings. The p-curve should not be relied upon when P values come from extensive outcome multiplicity, flexible analyses, or undisclosed outcome switching, because the set of “significant” results may itself have been selectively assembled. The method is also unreliable when statistical models are specified incorrectly, multiple P values from the same study are treated as independent, or study selection is driven by effect direction or novelty rather than statistical significance. Under these conditions, a right-skewed p-curve may reflect selective reporting or analytical flexibility rather than a true evidential value. Therefore, p-curve results should not be interpreted in isolation, but rather be evaluated alongside protocols, trial registries, risk-of-bias assessments, and other publication bias analyses.
However, these assumptions are often violated in clinical studies. Trials frequently evaluate multiple outcomes, time points, subgroups, and analytical specifications, reporting only a subset of favorable results. When such multiplicity, outcome switching, or selective analysis occurs, the observed P value distribution itself may be selectively constructed. The p-curve may thus suggest an evidential value, even when the apparent signal is driven by selective reporting, and p-uniform estimates may become unstable or misleading [8,19].
Therefore, p-curve and p-uniform analysis results should not be interpreted in isolation but are best used alongside information from study protocols, trial registries, selective outcome reporting assessments, and other publication-bias diagnostics. Rather than providing definitive proof for or against publication bias, P value-based methods should be viewed as a component of a broader sensitivity analysis framework, helping to clarify the extent to which conclusions depend on assumptions regarding reporting practices and selection mechanisms [8,19].

Detection as a transitional step

Methods to detect publication bias, whether graphical or statistical, do not explain why a pattern appears; they only show that the observed relationship between effect size and precision does not fit a simple reference model [7,1416]. Treating detection as the final step (e.g., treating any significant Egger’s test as definitive proof of publication bias and immediately reporting a single “corrected” estimate) encourages premature adjustment and overconfident conclusions [7,1416]. Detection is rather a transitional step in an integrated inferential framework. It should prompt a careful search for plausible explanations, such as heterogeneity, risk-of-bias gradients, outcome reporting bias, or selective analysis, and aid in the selection of sensitivity analyses and methods to adjust for publication bias that are conceptually aligned with those explanations [7,1416]. Thus, detection acts as a bridge between descriptive diagnostics and the counterfactual modeling and interpretive work developed in the following sections [8,9,15].

Statistical adjustment for publication bias

Adjustment as counterfactual modeling

Methods to adjust for publication bias are often described as solutions to publication bias, but logically they do not “remove” the bias in a causal sense [1215]. Instead, they encode explicit or implicit assumptions about the selection process that might be generating the asymmetry and then compute how the pooled effect would appear if those assumptions were true [38,1215]. The resulting adjusted estimates are therefore hypothetical and model-dependent. This dependence on untestable assumptions is thus a central feature that should be made explicit rather than concealed [1215]. From a sensitivity-analysis perspective, the key question is not which adjusted estimate is “correct,” but how much the substantive conclusions change or remain stable across a range of plausible selection and reporting mechanisms [1215].

Trim‑and‑fill: symmetry by construction

Trim‑and‑fill works by editing the funnel plot so that it looks symmetric. In practice, the algorithm first “trims” small, extreme studies from one side of the funnel and estimates where the center would be under symmetry. It then “fills” in mirror‑image studies on the other side and recomputes the pooled effect. This produces a clear visual story: “here are the missing negative studies, and here is the corrected effect,” but the filled‑in studies are not real data; they are values created by the algorithm under specific assumptions [7,8,12,13,16]. As a result, the apparent symmetry can be visually compelling yet methodologically misleading because graphical persuasion may outpace evidential justification (Fig. 3).
This method assumes that (1) the asymmetry is mainly caused by missing small, unfavorable studies, and (2) these missing studies would be arranged symmetrically around the pooled effect if they could be observed. These assumptions are often unrealistic when substantial heterogeneity or systematic differences in the study designs, outcome definitions, follow-up, or risk of bias exist. If the funnel is asymmetric because true effects differ across subgroups, outcomes are defined differently, or high‑risk‑of‑bias studies tend to show larger effects, trim‑and‑fill can be misleading as it will “create” missing studies that reflect the geometry of the algorithm rather than any plausible selection process in the real research ecosystem [7,8,12,13,16].
In such settings, the adjusted pooled effect can reflect the enforced symmetry of the method more than the clinical reality. Empirical work and case studies show that trim‑and‑fill can add several moderately sized, near‑null studies that do not resemble any observed trial and can either over‑correct or under‑correct depending on how heterogeneity and small‑study effects interact. For example, when small high‑quality trials show larger effects than larger low‑quality trials, trim‑and‑fill may “correct” toward the larger, biased studies rather than toward the truth.
For these reasons, trim‑and‑fill is best used as a simple, transparent sensitivity analysis, not as a definitive fix for publication bias. It is most informative when framed explicitly as “If we assume that asymmetry is entirely due to missing small, unfavorable studies arranged symmetrically around the pooled effect, here is how our conclusions would change.” In an integrated framework, its role is to illustrate how sensitive the main conclusions are to this specific symmetry‑based “missing‑study” scenario, alongside other methods to adjust for publication bias that encode different assumptions about selection and reporting [7,8,12,13,16,34].

Selection models: explicit selection mechanisms with strong sensitivity to specification

Selection models attempt to describe the likelihood that a study will be “seen” (e.g., published and then included in a meta-analysis) as a function of its P value, effect direction, or other study characteristics. In a common simple version, the analyst specifies a selection function that assigns various publication probabilities to different P value ranges (e.g., one probability for P < 0.01, another for P = 0.01–0.05, and another for P ≥ 0.05) and then jointly estimates the underlying effect and these selection parameters from the observed studies [8,9,15,16]. This is attractive because it directly encodes the idea that “significant” or favorable results are more likely to appear in the literature, but it also relies on strong assumptions about the exact shape (cut-points, patterns) and strength (likelihood) of selection, which are difficult to identify reliably from a single meta-analysis with limited data [8,9,15,16].
Crucially, the true selection process is unknown [6,1215]. Different plausible selection functions (i.e., using alternative P value cut-off points, different up-weighting of significant results, or asymmetric treatment of positive and negative findings) can often fit the observed data almost equally well while producing considerably different adjusted effect estimates [6,1215]. Simulation studies and applied examples show that even modest changes in how strongly significant studies are up-weighted, or in how many “non-significant” studies are assumed to be suppressed, can move the adjusted pooled effect from clearly beneficial to nearly null, even when standard fit statistics and residual checks look similar [6,1215]. In other words, selection models can be highly sensitive to specification as small changes in the assumed selection mechanism can lead to large changes in the “corrected” effect [6,1215].
Therefore, selection models are best used as part of a structured sensitivity analysis rather than as a one-shot correction that yields a single preferred estimate. In practice, this means fitting several plausible selection models (for example, different P value cut-off points or different assumptions about the extent that positive results are favored), comparing the extent to which the main clinical conclusions change across these models, and, where possible, anchoring the choice of selection functions in external information such as trial registries, protocols, reporting guidelines, or prior empirical evidence about editorial and reporting practices. Within an integrated inferential framework, the main value of selection models is not that they can recover the “true” effect, but that they make the assumed selection process explicit and allow readers to see the extent to which the conclusions depend on those assumptions [8,9,15,16].

PET–PEESE and related regression adjustments

Regression-based approaches such as the precision-effect test (PET) and precision-effect estimate with standard errors (PET–PEESE) are designed to address small-study effects by modeling how estimated treatment effects change with study precision, usually measured using the standard error or its variance. These methods are based on the simple idea that if smaller, less precise studies are more likely to report extreme results and larger, more precise studies are less affected by selective dissemination, then the relationship between effect size and precision may reveal the influence of small-study effects.
With the PET, effect sizes are regressed on the standard error, and the regression line is extrapolated to a standard error of zero. The intercept represents the expected effect in an infinitely precise study, and is interpreted as a bias-adjusted estimate. Importantly, an infinitely precise study does not correspond to any real clinical trial, rather, it is a conceptual reference point, a model-based projection that relies on strong assumptions about how effect size changes as precision increases (Fig. 4A). From an interpretive standpoint, the PET estimate is best understood as a test of whether a meaningful treatment effect remains after accounting for the association between the effect size and imprecision.
These methods are appealing because they are conceptually simple, easy to implement, and naturally fit within standard meta-regression frameworks. However, their validity depends on the key assumption that the relationship between effect size and precision is approximately linear and monotonic. This assumption is unlikely to be valid in several clinical meta-analyses. When heterogeneity is substantial, the precision range is dominated by a few influential studies, or study precisions vary only within a narrow window, PET estimates can become unstable [35].
The PEESE approach is often applied after the PET and is interpreted as a refinement that estimates the likely magnitude of the treatment effect under the assumption that a true effect exists but is distorted by small-study effects. The PEESE models the effect size as a function of variance (squared standard error) rather than the standard error itself (Fig. 4B). This alternative specification can lead to different adjusted estimates, as the divergence between the PET and PEESE regression lines illustrates (Fig. 4C). The PET and PET–PEESE should be interpreted together, as the PET primarily addresses whether the observed effect survives adjustment for imprecision, whereas the PET–PEESE suggests the possible size of that effect under the assumed precision–effect relationship.
In meta-analyses that include a few large pragmatic trials along with several smaller explanatory trials, the PET–PEESE may effectively shift the synthesis toward the largest and most precise studies. Rather than “correcting” bias in a general sense, this method may mainly reflect the results of these large trials. Fig. 4D illustrates this pattern by plotting risk ratios against study precision: smaller, less precise studies tend to report larger benefits, whereas the largest and most precise trials report effects closer to the null. This pattern highlights both the motivation for the PET–PEESE and its limitations, as it performs well only if the observed precision–effect relationship truly reflects small-study bias rather than genuine differences between trials.
Simulation studies show that the PET–PEESE can perform reasonably well under certain selection mechanisms and low heterogeneity but may perform poorly when heterogeneity is moderate to high or when selection operates through mechanisms other than simple significance filtering. For these reasons, the PET–PEESE should be viewed as one element of a broader sensitivity-analysis toolkit and not as a standalone method to adjust for publication bias [14,36].

Bayesian approaches to publication bias adjustment

Bayesian approaches to publication bias provide a flexible framework for modeling selective reporting by treating both the data-generating process and selection mechanism as probabilistic entities. Rather than producing a single bias-corrected estimate through algorithmic adjustment, Bayesian models explicitly encode assumptions about study selection, heterogeneity, and prior beliefs about effect sizes within a unified inferential structure. This allows uncertainty arising from missingness, small-study effects, and between-study variability to propagate through the posterior distribution rather than being handled in a piecemeal fashion [3739].
In practice, Bayesian publication bias models often specify a likelihood for observed study results together with a selection function that governs the probability of a study being observed or published, conditional on its statistical significance or effect size. Priors are then placed on both the treatment effect and the parameters of the selection process. This formulation makes explicit what is implicit in many frequentist adjustments, as conclusions depend critically on assumptions about how studies enter the literature and how strongly selection depends on statistical results. Importantly, these assumptions are not resolved by data alone and must be justified substantively rather than statistically [3739].
Transparency about assumption dependence is one conceptual advantage of Bayesian approaches. By varying priors or selection models, analysts can directly examine changes in posterior estimates under alternative, but plausible, mechanisms of selective reporting. Disagreement between posterior summaries across models is, therefore, informative, signaling the fragility of the evidence rather than the failure of the method. In this sense, Bayesian approaches align naturally with the view of publication bias adjustment as a structured sensitivity analysis rather than as a causal correction [15,37,38].
However, Bayesian models are equally subject to the fundamental limitations of publication bias adjustment. When the number of studies is small, heterogeneity is substantial, or selection mechanisms are complex and poorly understood, posterior inferences can be largely driven by prior choices. Highly informative priors may stabilize estimates but risk masking uncertainty, whereas weakly informative priors may yield diffuse posteriors that offer limited practical guidance. Moreover, different selection models can fit the observed data equally well while implying markedly different conclusions, underscoring the non-identifiability of publication bias mechanisms from the published literature alone [3739].
Accordingly, Bayesian publication bias methods are best used to make uncertainty explicit rather than to eliminate it. Their primary contribution lies in clarifying how conclusions depend on assumptions about selection, heterogeneity, and prior beliefs and in facilitating coherent comparisons across competing inferential scenarios. When reported transparently, along with frequentist adjustments, risk-of-bias assessments, and exploration of heterogeneity, Bayesian approaches can strengthen interpretive discipline without creating an illusion of definitive bias correction [3739].
In practice, Bayesian publication bias models are most commonly formulated as selection, weight function, or hierarchical models with explicit selection components, differing primarily in how the probability of study inclusion is linked to statistical significance, effect size, or both [3739]. Importantly, these Bayesian formulations do not obviate the need for a sensitivity analysis; rather, they formalize it by embedding selection assumptions directly into the inferential model.

Adjustment as sensitivity analysis: what to report

Across all adjustment methods, including both the frequentist and Bayesian approaches, disagreements between results should be viewed as informative rather than indications of failure. Such divergence shows that the main conclusions depend on assumptions about study selection and reporting, which cannot be verified directly [4,8,15,16]. When different methods such as trim-and-fill, selection models, PET–PEESE, and P value-based approaches produce clearly different adjusted effects, this signals assumption dependence, not analytical error.
Therefore, the most appropriate response is transparency. Analysts should clearly state the assumptions each method relies on, the sensitivity of the key clinical conclusions to those assumptions, and how this fragility may affect the certainty of the evidence and strength of recommendations [4,6,8,15,16]. Instead of presenting a single “corrected” estimate, meta-analyses should report a structured set of sensitivity analyses. These may include worst-case bias scenarios and ranges of estimates derived under alternative assumptions, for example through different priors or selection models in a Bayesian framework [4,8,15,16]. The goal is to demonstrate which effect sizes remain compatible with the data under plausible selection mechanisms.
From this perspective, adjustment methods including Bayesian selection models are not automatic correction tools. Rather, they are tools for disciplined inference that help define the limits of what can reasonably be concluded [1215]. When reported transparently and interpreted alongside risk-of-bias assessments, exploration of heterogeneity, and consideration of information size, these methods help align clinical conclusions with the true robustness of the evidence [4,6,8,15,16].
In settings where heterogeneity is low, outcomes and time points are clearly prespecified, and external information (e.g., trial registries) constrains plausible selection processes, publication bias adjustment methods can be particularly informative as sensitivity analyses. Under these conditions, the Bayesian and frequentist approaches may show greater convergence, supporting more confident but conditional interpretations. However, their role remains inferential rather than corrective.

An integrated inferential framework

Publication bias analysis is best viewed as a repeating cycle of reasoning rather than a checklist of separate techniques [8,9,15,16]. In this cycle, the detection methods highlight where the data deviate from a simple reference model. Adjustment methods then offer structured explanations for these deviations and generate alternative effect estimates under explicit assumptions regarding selection. Both frequentist and Bayesian approaches can be used at this stage, with interpretations demonstrating how clinical conclusions should change in light of these competing scenarios [2,79,1416,40]. The goal is not to identify a single “best” method, but to understand how strongly the direction, size, and certainty of conclusions depend on assumptions about which studies and results are more likely to appear [8,9,15,16].
In this framework, the first step is detection. Graphical and statistical diagnostics (e.g., funnel plots, Egger’s and Begg’s tests, P value-based approaches, and other measures of small-study effects) are used to stress test a simple model with homogeneous effects and no selection. Detection results are interpreted as signals that the simple model may be strained and not as proof of publication bias or any specific mechanism. Their main role is to prompt consideration of alternative explanations, including between-study heterogeneity, risk-of-bias gradients, outcome reporting bias, selective analysis, and design-related differences [68,16,19].
The second step is model-based adjustment, in which methods such as trim-and-fill, selection models, PET–PEESE, and Bayesian selection or hierarchical models are applied as sensitivity analyses under different narratives of how selection operates [7,8,1215,32,35]. Each method encodes specific assumptions: trim-and-fill assumes symmetry around a pooled effect; selection models assume explicit, often P value-dependent publication probabilities; regression-based approaches assume a defined relationship between effect size and precision; and Bayesian models embed these assumptions probabilistically through priors on effects, heterogeneity, and selection mechanisms [7,8,1215,32,35]. Rather than concealing these assumptions, the framework treats them as central objects of inference and encourages analysts to explore a range of plausible models, documenting how pooled estimates and uncertainty change across this range [8,9,16].
The third step is interpretation, in which the results from detection and adjustment are combined into judgments regarding robustness and confidence. The key questions include the following: How wide is the range of effect sizes compatible with the data under plausible selection mechanisms? Do conclusions regarding benefits, harm, or equivalence remain stable across methods, including Bayesian and frequentist models, or do they hinge on specific assumptions? When these findings are considered alongside risk of bias assessments, heterogeneity, and limitations in information size, do they justify downgrading the certainty of the evidence or softening clinical recommendations? [6,8,9,16,18]. By centering these questions, the framework shifts attention away from selecting a single “corrected” estimate and toward calibrating claims in proportion to evidential fragility or stability.
In practical terms, the integrated framework operates as a cycle: (1) specify a baseline meta-analytic model and apply detection tools; (2) define plausible selection and reporting mechanisms and apply corresponding adjustment methods as sensitivity analyses, including alternative Bayesian priors or selection structures when appropriate; and (3) compare the resulting estimates and intervals, explicitly mapping how conclusions depend on assumptions and data limitations [6,8,9,16,18]. The worked example, together with later discussions of common misuses, editorial evaluations, and umbrella reviews, illustrates how this inferential loop can guide both authors’ analytical choices and editors’ judgments on whether publication bias has been handled with appropriate rigor and restraint [6,8,9,16,18].

Applied example in clinical meta-analysis

Consider a hypothetical random‑effects meta-analysis of ten randomized controlled trials evaluating tropisetron versus a control for prevention of postoperative nausea and vomiting (PONV) in adult patients undergoing surgery under general anesthesia. The primary outcome is the incidence of PONV across the postoperative period and in the event that data were reported at multiple time points, the first available postoperative time point is selected as the outcome of interest.
A conventional random-effects model yields a pooled risk ratio (RR) of 0.678 (95% CI [0.561–0.821]), which would typically be interpreted as evidence of a clinically meaningful benefit. However, the evidence base is modest and heterogeneous (I2 = 64.5%, prediction interval: 0.405–1.137) (Fig. 5). The total sample size is limited (n = 1502), event rates vary widely (from 12.0% to 81.6%), outcome timing is inconsistent (from 0 to 18 h postoperatively), and study designs differ substantially between small explanatory trials and larger pragmatic trials. These features create a setting in which small-study effects are plausible, even in the absence of selective publications.
A funnel plot shows an apparent lack of small, unfavorable trials and a cluster of small studies with large beneficial effects (Fig. 1A). Egger’s regression test yields a P value of 0.001, and Begg’s rank correlation test is borderline (P = 0.089), which might be interpreted as evidence of publication bias. However, closer inspection suggests that small trials differ systematically from large trials in terms of methodological quality (Fig. 1B). In this context, funnel plot asymmetry may reflect effect modification or methodological gradients rather than selective non-publication alone [7,12,13,41].
To explore this uncertainty, publication bias analyses were treated as structured sensitivity analyses rather than corrective procedures [12,13]. Trim-and-fill attenuated the pooled effect, with the adjusted CI crossing the null (from RR = 0.678, 95% CI [0.561–0.821] to RR = 0.827, 95% CI [0.679–1.007]); however, the imputed studies did not resemble any clearly identifiable group of real trials (Fig. 3). This highlights the fact that the adjustment reflects the symmetry assumption of the method rather than a validated model of the publication process [7,12,13,41].
Selection models provide a complementary perspective on potential publication bias. Three P value-based selection scenarios were examined to assess the sensitivity to increasingly strong and granular assumptions regarding selective publication. In Scenario 1, studies with P < 0.05 were assumed to be twice as likely to be observed as non-significant studies, yielding a modestly attenuated pooled estimate (RR = 0.681, 95% CI [0.545–0.851]). Scenario 2 imposed a stronger selection assumption, a three-fold preferential observation of statistically significant studies, yet produced an almost identical estimate (RR = 0.681, 95% CI [0.545–0.851]). Scenario 3 applied a gradient with greater differentiation (P < 0.01: 2×; 0.01 ≤ P < 0.05: 1.5×; P ≥ 0.05: reference), again resulting in virtually unchanged pooled effects (RR = 0.680, 95% CI [0.544–0.850]) (Table 2).
Importantly, all selection-adjusted estimates were very close to the original random-effects result (RR 0.678, 95% CI: 0.561–0.821), indicating that P value-based selection assumptions (weak or strong) had a limited influence on the pooled effect in this evidence base [8,9,15,16].
However, regression-based approaches (PET–PEESE) yielded markedly different conclusions. The PET produced an estimate close to the null (RR = 1.018, 95% CI [0.945–1.096]), while the PEESE suggested an intermediate effect (RR = 0.889, 95% CI [0.817–0.966]) (Figs. 4AC). These differences arose because each method places implicit trust in different regions of the evidence base (i.e., large pragmatic trials, small explanatory trials, or a compromise between them) (Fig. 4D). The divergence is, therefore, informative, as it shows that conclusions depend on judgments regarding which studies are most credible [35].
Whereas the selection models explored above adopt a frequentist framework, Bayesian approaches offer an alternative inferential structure that makes explicit assumptions about selection, heterogeneity, and prior beliefs within a unified probabilistic model [3739]. To illustrate how Bayesian methods clarify assumption dependence, we applied a series of Bayesian models to a tropisetron meta-analysis with varying prior specifications and selection mechanisms. A baseline Bayesian model using a weakly informative prior (normal [0, 1] on the log risk ratio scale) yielded a posterior mean RR of 0.681 (95% credible interval [CrI]: 0.563–0.823), nearly identical to the frequentist random-effects estimate (RR = 0.678, 95% CI [0.561–0.821]), demonstrating that Bayesian and frequentist approaches converge when selection bias is not explicitly modeled and priors are diffuse (Fig. 6, Table 2).
After selection mechanisms were incorporated into the Bayesian framework, the posterior estimates shifted in a systematic and interpretable manner. Under Scenario 1, which assumed a two-step selection process in which studies with P < 0.05 were twice as likely to be published, the adjusted posterior estimate was RR 0.631 (95% CrI [0.537–0.740]). Imposing a stronger selection gradient in Scenario 2 (i.e., a threefold preferential observation of statistically significant studies) shifted the posterior toward smaller relative risk values, yielding an RR of 0.604 (95% CrI [0.525–0.696]). Scenario 3, which applied a moderate and more differentiated gradient (P < 0.01: 2×; 0.01 ≤ P < 0.05: 1.5×), produced an intermediate estimate (RR = 0.637, 95% CrI [0.541–0.750]) (Fig. 6, Table 2).
In these scenarios, stronger assumptions regarding selective publication were associated with numerically larger estimated treatment effects. Under these assumptions, studies with null or modest effects contribute less to the posterior inferences, while those reporting larger benefits carry greater weight. Consequently, the estimated effect increases. This behavior reflects how the model distributes evidence under the assumed selection mechanisms and does not prove the presence of publication bias.
Prior specifications influenced the posterior inferences, although their impact was smaller than that of the selection mechanism. Table 2 summarizes a prior sensitivity analysis across four specifications that reflect progressively stronger prior skepticism about treatment benefits: weakly informative (normal [0, 3.16]), moderately informative (normal [0, 1]), strongly informative (normal [0, 0.447]), and very skeptical (normal [0, 0.316]).
Under the very weak and moderate priors, both of which place minimal constraints on the effect size and allow the data to dominate, the posterior estimates closely matched the original random-effects meta-analysis (RR = 0.678, 95% CI [0.561–0.821]), yielding an RR of 0.679 (95% CrI [0.561–0.821]) and 0.681 (95% CrI [0.563–0.823]). This indicates that the Bayesian analysis largely reproduces the frequentist result in the absence of strong prior beliefs.
As prior skepticism increased, the posterior estimates gradually shifted toward the null. The skeptical prior produced an intermediate estimate (RR = 0.690, 95% CrI [0.573–0.832]), reflecting a balanced position between prior belief and observed data. Under the very skeptical prior, which encodes a strong belief that the true effect is likely to be small or absent, the posterior moved further toward the null (RR = 0.702, 95% CrI [0.585–0.842]).
Critically, the Bayesian framework makes this assumption dependence explicit and quantifiable, whereas it can be implicit in frequentist adjustments. By reporting posterior estimates under multiple priors and selection models side by side, analysts acknowledge that inferential conclusions are conditional on model choice and invite readers to assess which assumptions seem most plausible given substantive knowledge of the field. The resulting synthesis systematically and transparently documents uncertainty rather than eliminating it.
When analyses are restricted to statistically significant trial-level results (P < 0.05), p-curve methods indicate a right-skewed distribution of P values, which is consistent with the evidential value under standard assumptions. As seen in Fig. 7A, among the five studies with statistically significant results, the P values cluster toward smaller values (left side of the histogram, 0.00–0.01 range), producing the characteristic right-skewed distribution that the p-curve interprets as evidence against null findings being driven purely by selective reporting.
However, conducting a comparison with trial registry information reveals outcome switching and multiplicity in several studies, indicating that the observed set of significant P values is itself selectively constructed. Fig. 7B illustrates this critical limitation by displaying each study’s significance level along with evidence of reporting irregularities. Four of the five significant studies (Studies B, F, J, and D) show registry discrepancies: two with outcome switching (orange bars) and two with multiplicity (darker orange bars). In contrast, the five non-significant studies (P ≥ 0.05) show no evidence of registry discrepancies, all classified as “None” (blue bars).
This pattern demonstrates that even a superficially reassuring p-curve can reflect a selectively constructed evidence base, rather than robust evidential value. The concentration of reporting irregularities, specifically among significant studies (80% of significant studies vs. 0% of non-significant studies), suggests that the right-skewed p-curve may itself be an artifact of selective outcome reporting rather than evidence of true treatment effects. This limits the reassurance that P value-based methods can provide in this context and underscores the fact that apparent evidential value does not necessarily imply robustness when selective outcome reporting is present.
Fig. 6 presents a comprehensive comparison of frequentist (blue) and Bayesian (red) approaches to publication bias adjustment, illustrating the assumption dependence of meta-analytic inference when small-study effects are present. Frequentist methods yield divergent estimates depending on their underlying assumptions: trim-and-fill approaches the null (RR = 0.827, 95% CI [0.679–1.007]) by imputing five missing studies, selection models produce estimates of RR = 0.680–0.681 under varying selection severity assumptions, and regression-based methods range from no effect (PET: RR = 1.018, 95% CI [0.945–1.096]) to moderate benefit (PEESE: RR = 0.889, 95% CI [0.817–0.966]). Bayesian approaches show similar point estimates (RR = 0.604–0.702), with tighter credible intervals but meaningful sensitivity to prior specification, shifting from RR = 0.679 (very weak prior) to RR = 0.702 (very skeptical prior). Therefore, substantial divergence across methods or prior specifications should be interpreted as a signal that clinical conclusions are sensitive to modeling assumptions and warrant cautious interpretation.
The central inferential message is not that one adjustment method “wins,” but rather that the claim of benefit is fragile across plausible assumptions about selection mechanisms and prior beliefs [4,8,16]. The divergence of adjusted estimates, spanning from substantial benefit (RR = 0.604) to no effect (RR = 1.02), combined with evidence of selective outcome reporting in 80% of significant studies (Fig. 7B), indicates that clinical recommendations should be tentative with moderate or low certainty ratings. Rather than selecting a single “best” estimate, decision-makers should consider the entire range of plausible effect sizes when evaluating treatment efficacy, treating the divergence across methods as a result that reflects epistemic uncertainty rather than a methodological problem to be solved [4,6,8,16].

Common misuses in clinical meta-analyses

Empirical assessments of meta-analyses in clinical medicine consistently show that publication bias procedures are often applied in a mechanical, checklist-driven way, with little attention to their assumptions or limits of interpretation [25,42,43]. One recurring error is the tendency to treat funnel plot asymmetry or a statistically significant Egger’s test as direct proof of publication bias, even when substantial heterogeneity, inconsistent outcome definitions, or clear design-related trends across studies exist. Under such conditions, the observed asymmetry can reasonably be explained by real-effect modification, differences in study quality, baseline risk, or follow-up duration, especially when these factors are related to the study size. Calling all such patterns “publication bias” oversimplifies the problem and ignores other plausible explanations, collapsing multiple mechanisms into a single narrow label [2,6,18,19].
A second common misuse is the presentation of one bias-adjusted estimate, most often from a trim-and-fill or single selection model, as the main or final result of the meta-analysis. In many studies, the unadjusted analysis or results from alternative models are pushed into supplementary files or not shown [8,9,12,13,15,16,44]. This turns a sensitivity analysis into a headline conclusion, although different and equally reasonable models can fit the same data and lead to clearly different pooled effects [8,9,1416]. A related problem is that researchers rarely carefully examine how and why the estimate changes after adjustment. Adjusted results are often assumed to be more accurate simply because they are smaller or closer to no effect without considering whether the studies that drive this change make clinical or methodological sense.
Third, measures of heterogeneity such as I2, τ2, τ, or prediction interval are frequently reported but not meaningfully used in interpretation [3]. Authors may acknowledge large between-study variability yet still apply asymmetry tests and bias-adjustment methods as if all studies were estimating the same underlying effect. This internal inconsistency weakens both the detection of bias and any subsequent adjustments because most publication bias methods rely implicitly on relatively homogeneous effects [3,16].
Finally, and perhaps most consequentially, overconfident conclusions remain commonplace. Even when publication bias analyses show instability, such as widely varying adjusted estimates, broad sensitivity ranges, or inadequate information size, authors may still use strong causal language or recommend clinical adoption without appropriately lowering certainty [5,6,8,9,1416]. Collectively, these practices reflect a persistent misunderstanding of publication bias adjustment methods as corrective tools that “fix” biased evidence, rather than as exploratory tools that test the extent to which conclusions depend on untestable assumptions about study selection and reporting [8,9,1416].

Editorial and authorial evaluation: what constitutes adequate handling of publication bias?

From an editorial and methodological perspective, handling publication bias in meta-analyses in clinical medicine cannot be judged simply by whether a funnel plot or Egger’s test has been reported [5,6,16,18,25]. The same limitation applies from an authorial perspective: adequate handling is not demonstrated by performing these analyses, but by their framing, justification, and interpretation. The key question should be framed as “Were these tools used thoughtfully?” Critically, the manuscript should show inferential discipline, with authors recognizing that small-study effects are non-specific, discussing other plausible explanations besides publication bias, and clearly explaining why the chosen sensitivity analyses are appropriate for this particular body of evidence [79,15,17,25]. For authors, reviewers, and editors alike, this means checklist-style evaluations are avoided, and both the internal logic and coherence of the analysis are judged. In particular, attention should be paid to how publication bias assessments are integrated with risk-of-bias evaluations, explorations of heterogeneity, checks for selective outcome reporting, and judgments about the certainty of the evidence, rather than being treated as isolated statistical steps [5,6,8,9,15,18].
From this shared perspective, a role-informed checklist (Table 3) can help clarify expectations across authorship, peer review, and editorial decision making. Rather than functioning as a procedural requirement, the checklist is intended for structural interpretation. The focus is on the framing, justification, and integration of publication bias analyses into the overall inferential narrative and not simply on whether particular statistical techniques were applied [4,8,12,13,15,17]. In this sense, the checklist serves as a guide for judging studies rather than as a substitute. By foregrounding questions about assumptions, coherence, and evidential implications, it aims to promote consistency in evaluations, while discouraging the mechanical use of bias diagnostics as standalone indicators of rigor.
The interpretive role of publication bias analyses aligns closely with the logic of evidence certainty frameworks, such as GRADE [45]. Rather than serving as a mechanism to correct effect estimates, publication bias assessments should inform judgments regarding the credibility and stability of the evidence [46]. When adjusted estimates vary substantially across plausible models or when conclusions depend strongly on unverifiable selection assumptions, this sensitivity should be reflected as a downgrade in certainty rather than resolved through the selection of a single preferred estimate. Therefore, from both authorial and editorial perspectives, the appropriate response to suspected publication bias is not statistical correction, but transparent acknowledgment of uncertainty and proportionate calibration of evidentiary claims.
Editors and reviewers can promote better practice by valuing manuscripts that use publication bias methods to explore the robustness of conclusions, rather than to produce a single “corrected” effect estimate [4,8,15,17]. Authors should frame bias adjustment explicitly as a sensitivity exercise and not as a substitute for the primary analysis. Well-handled analyses treat bias adjustment as a sensitivity exercise, and not as a replacement for the primary analysis. This includes encouraging the regular use of structured sensitivity analyses (e.g., worst-case bias scenarios, model-averaging approaches, and clear demonstrations of the change in the conclusions under different selection assumptions) and explicitly downgrading the certainty of evidence when conclusions depend strongly on unverifiable assumptions [46,8,15,17]. Ultimately, treating publication bias analyses as part of the interpretation rather than as a final diagnostic checkbox can help journals ensure that synthesized evidence is both methodologically rigorous and appropriately cautious in its claims [46,8,15,17].

Preprints and the reshaping of the publication bias ecosystem

The availability of preprints is currently changing the channels of evidence sharing in clinical medicine that impact publication bias and meta-analyses. Preprints are often described as a solution to publication bias because they allow for the rapid public posting of study results without the need for journal gatekeeping. This can reduce outright non-publication and time-lag bias because negative or inconclusive studies can, in principle, be posted as easily as positive ones. However, preprints do not remove selective processes; instead, they change when, how, and which results become visible, thereby introducing new forms of selection while mitigating others [1,47,48].
Traditional publication bias analysis methods were developed for a journal-centered world, which implicitly assumes a relatively static evidence base in which studies are either published or not published. Preprints add an additional dissemination layer that is both dynamic and versioned. Decisions about whether to post a preprint, which analyses to present, and whether or how to update later versions remain under the control of the investigators. As a result, preprints change what is missing and when rather than eliminating missingness altogether [1,48].
This shift has important implications for statistical signals of publication bias. Funnel plot asymmetry and small-study effects may partly reflect the temporal ordering of evidence. Early preprints may come mainly from small, exploratory studies with imprecise and sometimes exaggerated effects, with larger, more definitive trials appearing only later. In such situations, asymmetry does not necessarily indicate selective suppression, but may instead arise from the timing and evolution of dissemination [1,47].
For evidence synthesis, neither automatic inclusion nor systematic exclusion of preprints is methodologically neutral. Including preprints may increase the heterogeneity and instability if many are small, preliminary, or later retracted or heavily revised, whereas excluding them may reintroduce a directionally biased synthesis by preferentially omitting null or unpublished results. A more defensible strategy is to identify preprints transparently, track version histories, and perform sensitivity analyses that compare meta-analysis results with and without preprint data [1,47].
Preprints also interact with higher-order evidence synthesis in distinct ways. Umbrella reviews tend to amplify bias vertically by stacking one synthesis on top of another, whereas preprints redistribute bias horizontally across dissemination stages and over time. Recognizing this distinction helps in understanding how publication and reporting biases propagate in modern evidence ecosystems [1,49].
Consistent with the broader argument of this review, preprints should not be treated as a cure for publication bias. Instead, they represent a structural change in the research ecosystem that alters the form, timing, and visibility of selective processes. In this setting, publication bias analysis remains an exercise in mapping assumption dependence and evidential fragility, rather than in recovering a single bias‑corrected effect [1,50].

Amplification of publication bias in umbrella reviews

Umbrella reviews synthesize evidence from multiple systematic reviews and meta-analyses, and therefore operate one level above primary-study selection processes [5,6,9]. This structure creates opportunities for publication bias to be amplified across synthesis layers, as bias that originates in the primary literature (selective publication and reporting) is propagated into individual meta-analyses (selective inclusion, analytic flexibility), and then further reinforced in umbrella reviews that preferentially include prominent or statistically significant meta-analyses [5,6,9]. The overlap of primary studies across meta-analyses can cause the same influential positive trials to be counted repeatedly, creating a false impression of replication when the same selectively disseminated evidence is reused across syntheses. This recycling effect is particularly problematic when a small number of early or high-impact trials dominate multiple systematic reviews and meta-analyses.
Furthermore, umbrella reviews often classify evidence strength using simple decision rules such as significance thresholds, effect size cut-offs, or criteria based on the number and size of contributing studies [1618]. These rules can be misleading when the pooled estimates themselves are unstable. When evidence grading relies on results that are fragile, highly sensitive to publication bias adjustments, strongly heterogeneous, or based on limited information size, umbrella reviews may turn weak or uncertain signals into categories that appear authoritative and definitive [46,8,15,17]. In this way, statistical uncertainty can be hidden behind formal labels of “strong” or “convincing” evidence.
Publication bias assessment in umbrella reviews should therefore focus on patterns of fragility, study overlap, and robustness across competing assumptions, rather than relying only on single corrected estimates from individual systematic reviews and meta-analyses or uncritical evidence grading systems [46,8,15,17]. Approaches that explicitly address the overlap of primary studies, compare results across alternative inclusion criteria and modeling choices, and summarize publication bias signals across constituent meta-analyses are particularly important at this higher level of synthesis [5,6,9]. Without such safeguards, umbrella reviews risk amplifying bias while providing an illusion of increased certainty.

Future directions: automation, artificial intelligence, and the new face of bias

Automation and artificial intelligence (AI)-assisted evidence synthesis introduce new avenues for bias, beyond classical statistical selection mechanisms. Algorithms trained mainly on published and highly cited studies tend to inherit the publication and citation biases already present in these sources. Consequently, automated screening, ranking, or prioritization systems may systematically favor large, positive, or high-profile studies, while smaller or null studies receive less attention. This process can quietly reinforce the same biases that systematic review and meta-analysis methods are meant to examine. Large language models and other generative tools add an additional layer of risk. They can turn uncertain, assumption-dependent meta-analysis results into fluent, convincing narratives, often downplaying uncertainty, masking fragility, or omitting key caveats regarding publication bias, heterogeneity, or limited information size. When results are presented in a clear and polished manner, readers may confuse how convincing they sound with how strong the underlying evidence is [51].
Addressing these challenges requires methodological and governance safeguards in addition to improved algorithms. Traditional publication bias analyses must be actively integrated into the design and oversight of automated systems [4,8,15,17]. Bias-aware retrieval strategies, such as systematic searches of trial registries, forward and backward citation chasing, and the inclusion of preprints and regulatory documents, are essential to counter the default tendency of automated tools to focus on highly cited journal articles [5,6,18]. Without these countermeasures, automation risks narrowing rather than broadening the evidence base.
In addition, the routine comparison of registry records with published reports, transparent audit trails for automated screening and inclusion decisions, and human-in-the-loop interpretations that explicitly consider publication and citation bias should become standard components of AI-assisted evidence synthesis workflows [5,6,18]. Human oversight is particularly important at interpretive stages, during which judgments about certainty and clinical relevance are made.
In this emerging landscape, publication bias is no longer only a statistical problem, but also a socio-technical problem shaped by how evidence is retrieved, summarized, and communicated. Aligning tool development, reporting standards, and editorial expectations is therefore essential to ensure both that automation supports methodological rigor rather than amplifying bias and that conclusions remain appropriately cautious and well calibrated to the underlying evidence [4-6,8,15,17].

Conclusion

Publication bias analysis methods do not causally correct for bias. Rather than recovering a hidden “true” effect, they reveal the extent to which clinical conclusions depend on unverifiable assumptions about study selection, reporting, and dissemination and the potential fragility of an evidence base under plausible mechanisms of missingness and selective reporting. As such, their primary value lies not in numerical adjustment, but in clarifying the limits of inference. Viewed through a certainty-of-evidence framework such as GRADE, the primary contribution of publication bias analysis is to guide the appropriate downgrading of confidence in the evidence. Divergent adjusted estimates, wide sensitivity ranges, or strong dependence on unverifiable selection assumptions signal fragility that should temper causal language and clinical recommendations and should not be resolved by choosing a single preferred model.
In this sense, publication bias analysis functions as a bridge between statistical modeling and responsible interpretation. When used to characterize uncertainty through structured sensitivity analyses and transparent reporting, it supports interpretive discipline rather than false precision. Treated in this way, publication bias analysis methods help ensure that conclusions drawn from systematic reviews and meta-analyses, and the clinical recommendations that follow, remain proportional to the strength and stability of the underlying evidence rather than being overstated by apparent numerical certainty.

Funding

None.

Conflicts of Interest

Hyun Kang has been a member of the Statistical Rounds of the Korean Journal of Anesthesiology. However, he was not involved in any process of review for this article, including peer reviewer selection, evaluation, or decision-making. There were no other potential conflicts of interest relevant to this article.

Data Availability

All data generated or analyzed during this study are included in this published article and its supplementary information files. The analyses presented in this paper can be performed using the Supplementary R code.

Supplementary Materials

Supplementary Material 1.
R code.
kja-26020-Supplementary-Marterial-1.R
Supplementary Material 2.
Annotated code.
kja-26020-Supplementary-Marterial-2.docx

Fig. 1.
Funnel plots of trials evaluating tropisetron for prevention of postoperative nausea and vomiting. (A) Conventional funnel plot. (B) Quality-stratified funnel plot. Each circle represents an individual study, plotted by log risk ratio and standard error. Shaded contours indicate regions of statistical significance (P < 0.05 and P < 0.01). Blue circles indicate higher-quality studies, whereas red circles indicate lower-quality studies.
kja-26020f1.jpg
Fig. 2.
Distributions of P values under different selective reporting scenarios. (A) Distribution showing an excess of small P values, consistent with a right-skewed pattern often interpreted as evidential value under selective reporting. (B) Distribution compared with an expected uniform frequency, illustrating how deviations from uniformity may inform assessment of selective significance filtering. (C) Distribution characterized by clustering just below conventional significance thresholds, consistent with selective reporting or analytical flexibility rather than a coherent underlying effect. Each panel displays the frequency distribution of P values from a set of studies. Vertical dashed lines indicate conventional significance thresholds (P = 0.05 and P = 0.01).
kja-26020f2.jpg
Fig. 3.
Trim-and-fill-adjusted funnel. Each circle represents an individual study, plotted by log risk ratio and standard error. Shaded contours indicate regions of statistical significance (P < 0.05 and P < 0.01). Open circles denote observed studies, and filled circles denote imputed studies. Diamonds indicate the pooled effect estimates before (gray) and after (white) trim-and-fill adjustment, with corresponding 95% CI.
kja-26020f3.jpg
Fig. 4.
PET–PEESE regression adjustments for publication bias. (A) Precision-effect test (PET), showing regression of log risk ratio on standard error. (B) Precision-effect estimate with standard errors (PEESE), showing regression of log risk ratio on variance. (C) Comparison of pooled effect estimates from the conventional random-effects meta-analysis, PET, and PEESE, displayed as risk ratios with 95% CI. (D) Scatter plot of treatment effect by study precision (1/standard error), illustrating the association between study size and estimated effect. In panels A and B, each circle represents an individual study, with point size proportional to study precision. Dashed horizontal lines indicate the conventional random-effects estimate and the PET- or PEESE-adjusted estimates. Blue dashed lines denote the conventional random-effects estimate, red dashed lines denote the PET-adjusted estimate, and green dashed lines denote the PEESE-adjusted estimate.
kja-26020f4.jpg
Fig. 5.
Forest plot for preventive effect of postoperative nausea and vomiting. Each trial is shown as a filled square proportional to sample size, with the 95% CI depicted as a horizontal line. The blue diamond indicates the pooled effect and 95% CI and the red bar indicates 95% prediction interval. N number, RR risk ratio.
kja-26020f5.jpg
Fig. 6.
Comparison of frequentist and Bayesian approaches to publication bias adjustment. Forest plot comparing publication bias adjustment methods. Frequentist (blue) and Bayesian (red) approaches show varying estimates based on different assumptions. Green line indicates null effect (RR = 1.0). Error bars: 95% CI (frequentist) or CrI (Bayesian). Substantial divergence across methods or prior specifications indicates that the pooled estimate is sensitive to modeling assumptions and should be interpreted with caution. CrI: credible interval.
kja-26020f6.jpg
Fig. 7.
P value distribution and registry-based evidence of selective reporting. (A) Right-skewed p-curve among significant studies suggests evidential value. (B) Registry comparison reveals outcome switching or multiplicity. Red dashed lines indicate P = 0.025 (A) and P = 0.05 (B) thresholds.
kja-26020f7.jpg
Table 1.
Taxonomy of Dissemination‑related and Reporting Biases Relevant to Meta-analyses in Clinical Medicine
Publication bias
 Selective non-publication of entire studies, usually related to statistical significance, effect direction, or the perceived importance of results. This mechanism directly reduces the observable evidence base and is the primary target of funnel-plot-based methods to detect publication bias and selection models as methods to adjust for publication bias.
Time‑lag bias
 Delayed publication of studies with null or unfavorable results. Meta-analyses conducted early in the lifecycle of an intervention may therefore overestimate treatment effects, even if all studies are eventually published, and repeated updates often show attenuation of effects over time.
Outcome reporting bias (within‑study reporting bias)
 Selective reporting of outcomes, time points, analyses, or adverse events within published studies. Pre-specified outcomes may be omitted, downgraded to secondary status, or replaced by more favorable endpoints, producing biased effect estimates even when all trials are nominally “published.” Study-level methods to detect publication bias often have low sensitivity to this mechanism because no entire trials are missing.
Selective analysis and specification bias
 Use of analytic flexibility—for example, alternative covariate sets, transformations, subgroups, or statistical models—in ways that favor statistically significant or desirable findings. This contributes to small-study effects and distorted p value distributions without requiring selective publication or non-publication of whole studies.
Duplicate, location, and citation bias
 Multiple publication of the same or overlapping data, preferential publication of positive studies in high-impact or easily accessible journals, and preferential citation of positive or novel results. These mechanisms can over-represent favorable findings in literature searches and narrative syntheses and can lead to overestimation of treatment effects if duplicates are inadvertently counted more than once in a meta-analysis.
Dissemination channel bias
 Differential visibility of results across journals, trial registries, preprints, conference abstracts, regulatory documents, and grey literature. Positive findings are more likely to appear in prominent, indexed outlets, whereas null or unfavorable results may remain in hard-to-access sources, making study identification and selection more difficult.
Language bias
 Preferential publication or indexing of studies in particular languages, sometimes linked to the nature or direction of the results. Restricting a systematic review to English-language publications can therefore miss relevant evidence and, in some settings, alter pooled estimates or conclusions.
Table 2.
The Results according to Selection Scenario
Method and Scenario Selection probability
Frequentist Conventional random-effects model Same in all cases RR 0.678 [95% CI 0.561–0.821]
Trim-and-Fill RR 0.827 [95% CI 0.679–1.007]
Selection Scenarios Scenario 1: Moderate selection P < 0.05: 2 times, P ≥ 0.05: reference RR 0.681 [95% CI 0.545–0.851]
Scenario 2: Strong selection P < 0.05: 3 times, P ≥ 0.05: reference RR 0.681 [95% CI 0.545–0.851]
Scenario 3: Custom gradient (P < 0.01: 2×, 0.01–0.05: 1.5×) P < 0.01: 2 times, 0.01 ≤ P < 0.05: 1.5 times, P ≥ 0.05: reference RR 0.680 [95% CI 0.544–0.850]
PET- PEESE PET RR 1.018 [95% CI 0.945–1.096]
PEESE RR 0.889 [95% CI 0.817–0.966]
Bayesian Bayesian (Weakly Informative) RR 0.681 [95% CrI 0.563–0.823]
Selection Scenarios Scenario 1 P < 0.05: 2 times, P ≥ 0.05: reference RR 0.631 [95% CrI 0.537–0.740]
Scenario 2 P < 0.05: 3 times, P ≥ 0.05: reference RR 0.604 [95% CrI 0.525–0.696]
Scenario 3 P < 0.01: 2 times, 0.01 ≤ P < 0.05: 1.5 times, P ≥ 0.05: reference RR 0.637 [95% CrI 0.541–0.750]
Prior Sensitivity Very weak prior Normal (0, 3.16) RR 0.679 [95% CrI 0.561–0.821]
Moderate prior Normal (0, 1) RR 0.681 [95% CrI 0.563–0.823]
Skeptical prior Normal (0, 0.447) RR 0.690 [95% CrI 0.573–0.832]
Very skeptical prior Normal (0, 0.316) RR 0.702 [95% CrI 0.585–0.842]
Table 3.
A Role-specific Checklist for Authors, Reviewers, and Editors in Evaluating Publication Bias
Domain Authors Reviewers Editors
Conceptual framing State clearly that funnel plot asymmetry and small-study effects are non-specific signals, not proof of publication bias. Check whether asymmetry or Egger’s test is treated as direct evidence of publication bias. Assess whether the manuscript shows inferential discipline rather than box-ticking use of tests.
Choice of methods Justify detection or adjustment methods in relation to plausible selection processes. Evaluate whether chosen methods fit the clinical and methodological context. Judge alignment between bias methods and real-world publication mechanisms.
Transparency of results Report both unadjusted and bias-adjusted results clearly in the main text. Confirm that key results are not hidden only in supplements. Ensure balanced presentation without overemphasis on a single adjusted estimate.
Multiple approaches Apply more than one complementary method when bias is suspected. Check consistency and differences across methods. Avoid reliance on a single “corrected” estimate.
Interpretation of adjustment Explain direction and magnitude of changes after adjustment. Assess plausibility of added or down-weighted studies. Confirm adjustment is treated as sensitivity analysis, not truth correction.
Heterogeneity integration Interpret bias analyses alongside I², τ², and clinical heterogeneity. Check for inconsistency between heterogeneity discussion and bias analysis. Ensure coherence between heterogeneity and bias handling.
Selective reporting Consider within-study selection (outcome switching, timing). Evaluate discussion of selective reporting beyond nonpublication. Check recognition of limits of funnel plots for detecting selective reporting.
Impact on conclusions Temper conclusions when results are sensitive to assumptions. Match strength of language to robustness of findings. Ensure abstracts and conclusions reflect uncertainty.
Higher-level syntheses Consider accumulation of bias across evidence layers. Check handling of overlap and preferential inclusion. Assess recognition of bias propagation in umbrella reviews.
Overall judgment Demonstrate how conclusions depend on assumptions. Assess analytic coherence. Decide whether synthesis is cautious and methodologically sound.

References

1. Kang H. Beyond the paywall: the role of preprints in overcoming publication bias. J Evid-Based Pract 2025; 1: 7-11.
crossref pdf
2. Almalik O. Adjusting for publication bias in meta-analysis with continuous outcomes: a comparative study. Mathematics 2025; 13: 3487.
crossref
3. Choi GJ, Kang H. Heterogeneity in meta-analyses: an unavoidable challenge worth exploring. Korean J Anesthesiol 2025; 78: 301-14.
crossref pmid pmc pdf
4. Bartoš F, Maier M, Wagenmakers EJ, Nippold F, Doucouliagos H, Ioannidis JP, et al. Footprint of publication selection bias on meta-analyses in medicine, environmental sciences, psychology, and economics. Res Synth Methods 2024; 15: 500-11.
crossref pmid
5. Hennessy EA, Johnson BT, Keenan C. Best practice guidelines and essential methodological steps to conduct rigorous and systematic meta-reviews. Appl Psychol Health Well Being 2019; 11: 353-81.
crossref pmid pmc pdf
6. Johnson BT, Hennessy EA. Systematic reviews and meta-analyses in the health sciences: best practice methods for research syntheses. Soc Sci Med 2019; 233: 237-51.
crossref pmid pmc
7. Moreno SG, Sutton AJ, Ades AE, Stanley TD, Abrams KR, Peters JL, et al. Assessment of regression-based methods to adjust for publication bias through a comprehensive simulation study. BMC Med Res Methodol 2009; 9: 2.
crossref pmid pmc pdf
8. Mathur MB, VanderWeele TJ. Sensitivity analysis for publication bias in meta-analyses. J R Stat Soc Ser C Appl Stat 2020; 69: 1091-119.
crossref pmid pmc pdf
9. Bartoš F, Maier M, Wagenmakers EJ, Doucouliagos H, Stanley TD. Robust Bayesian meta-analysis: model-averaging across complementary publication bias adjustment methods. Res Synth Methods 2023; 14: 99-116.
crossref pmid pmc pdf
10. Nakagawa S, Yang Y, Macartney EL, Spake R, Lagisz M. Quantitative evidence synthesis: a practical guide on meta-analysis, meta-regression, and publication bias tests for environmental sciences. Environ Evid 2023; 12: 8.
crossref pmid pmc pdf
11. Lee S. Heterogeneity in meta-analysis: a path toward more meaningful clinical evidence. Korean J Anesthesiol 2025; 78: 297-8.
crossref pmid pmc pdf
12. Duval S, Tweedie R. Trim and fill: a simple funnel-plot-based method of testing and adjusting for publication bias in meta-analysis. Biometrics 2000; 56: 455-63.
crossref pmid pmc
13. Duval S, Tweedie R. A nonparametric “trim and fill” method of accounting for publication bias in meta-analysis. J Am Stat Assoc 2000; 95: 89-98.
crossref
14. Stanley TD. Limitations of PET-PEESE and other meta-analysis methods. Soc Psychol Personal Sci 2017; 8: 581-91.
crossref pdf
15. Maier M, VanderWeele TJ, Mathur MB. Using selection models to assess sensitivity to publication bias: a tutorial and call for more routine use. Campbell Syst Rev 2022; 18: e1256.
crossref pmid pmc
16. Lin L, Chu H. Quantifying publication bias in meta-analysis. Biometrics 2018; 74: 785-94.
crossref pmid pmc pdf
17. Lin L, Chu H. Rejoinder to “quantifying publication bias in meta-analysis”. Biometrics 2018; 74: 801-2.
crossref pdf
18. Johnson CE, Weerasuria MP, Keating JL. Effect of face-to-face verbal feedback compared with no or alternative feedback on the objective workplace task performance of health professionals: a systematic review and meta-analysis. BMJ Open 2020; 10: e030672.
crossref pmid pmc
19. van Aert RC, Wicherts JM, van Assen MA. Conducting meta-analyses based on p values: reservations and recommendations for applying p-uniform and p-curve. Perspect Psychol Sci 2016; 11: 713-29.
crossref pmid pmc pdf
20. Page MJ, Higgins JP. Rethinking the assessment of risk of bias due to selective reporting: a cross-sectional study. Syst Rev 2016; 5: 108.
crossref pmid pmc pdf
21. Jackson JL, Balk EM, Hyun N, Kuriyama A. Approaches to assessing and adjusting for selective outcome reporting in meta-analysis. J Gen Intern Med 2022; 37: 1247-53.
crossref pdf
22. Jüni P, Reichenbach S, Egger M. COX 2 inhibitors, traditional NSAIDs, and the heart. BMJ 2005; 330: 1342-3.
crossref pmid pmc
23. Garner S, Fidan D, Frankish R, Judd M, Towheed T, Wells G, et al. Rofecoxib for the treatment of rheumatoid arthritis. Cochrane Database Syst Rev 2002; (3): CD003685. Update in: Cochrane Database Syst Rev 2005; (1): CD003685.
crossref pmid pmc
24. Hazlewood GS. Risk/benefit trade-offs in rheumatology: rofecoxib revisited in the era of JAK inhibitors. Rheumatology (Oxford) 2023; 62: 3773-5.
crossref pdf
25. Sutton AJ, Duval SJ, Tweedie RL, Abrams KR, Jones DR. Empirical assessment of effect of publication bias on meta-analyses. BMJ 2000; 320: 1574-7.
crossref pmid pmc
26. Kang H. Statistical considerations in meta-analysis. Hanyang Med Rev 2015; 35: 23-32.
crossref
27. Simmonds M. Quantifying the risk of error when interpreting funnel plots. Syst Rev 2015; 4: 24.
crossref pmid pmc pdf
28. Sterne JA, Egger M, Smith GD. Systematic reviews in health care: investigating and dealing with publication and other biases in meta-analysis. BMJ 2001; 323: 101-5.
crossref pmid pmc
29. Rafla C, Zitek T, Shalaby M. A systematic review of the respiratory effects of ultrasound-guided phrenic nerve block. Anesth Pain Med (Seoul) 2025; 20: 371-83.
crossref pmid pmc pdf
30. Egger M, Davey Smith G, Schneider M, Minder C. Bias in meta-analysis detected by a simple, graphical test. BMJ 1997; 315: 629-34.
crossref pmid pmc
31. Begg CB, Mazumdar M. Operating characteristics of a rank correlation test for publication bias. Biometrics 1994; 50: 1088-101.
crossref pmid
32. Simonsohn U, Nelson LD, Simmons JP. p-curve and effect size: correcting for publication bias using only significant results. Perspect Psychol Sci 2014; 9: 666-81.
crossref pmid pdf
33. Bergman H, Henschke N, Arevalo-Rodriguez I, Buckley BS, Crosbie EJ, Davies JC, et al. Human papillomavirus (HPV) vaccination for the prevention of cervical cancer and other HPV-related diseases: a network meta-analysis. Cochrane Database Syst Rev 2025; 11: CD015364.
crossref pmid pmc
34. Ximenes GF, Costa ÁL, Leite LL, Costa LL, Ribeiro MO, Baima Colares PG, et al. Are artificial intelligence models reliable for clinical application in pediatric fracture detection on radiographs? A systematic review and meta-analysis. Clin Orthop Relat Res 2026; 484: 371-85.
crossref pmid
35. Stanley TD, Doucouliagos H. Meta-regression approximations to reduce publication selection bias. Res Synth Methods 2014; 5: 60-78.
crossref pmid
36. Alinaghi N, Reed WR. Meta-analysis and publication bias: how well does the FAT-PET-PEESE procedure work? Res Synth Methods 2018; 9: 285-311.
crossref pmid pdf
37. Larose DT, Dey DK. Modeling publication bias using weighted distributions in a Bayesian framework. Comput Stat Data Anal 1998; 26: 279-302.
crossref
38. Maier M, Bartoš F, Wagenmakers EJ. Robust Bayesian meta-analysis: addressing publication bias with model-averaging. Psychol Methods 2023; 28: 107-22.
crossref pmid
39. Jung J, Aloe AM. Bayesian workflow for bias-adjustment model in meta-analysis. Res Synth Methods 2026; 17: 293-313.
crossref pmid pmc
40. Cansian JM, Bracht VS, Biolo LV, Schmidt AP. Comparative effectiveness of bupivacaine and lidocaine-bupivacaine mixtures in brachial plexus block: a systematic review and meta-analysis. Anesth Pain Med (Seoul) 2025; 20: 252-65.
crossref pmid pmc pdf
41. Spitschan M, Schmidt MH, Blume C. Principles of open, transparent and reproducible science in author guidelines of sleep research and chronobiology journals. Wellcome Open Res 2021; 5: 172.
crossref pmid pmc pdf
42. Atakpo P, Vassar M. Publication bias in dermatology systematic reviews and meta-analyses. J Dermatol Sci 2016; 82: 69-74.
crossref pmid
43. Mathur MB. Sensitivity analysis for the interactive effects of internal bias and publication bias in meta-analyses. Res Synth Methods 2024; 15: 21-43.
crossref pmid pmc
44. Mahendru K, Kumar A, Pandey K, Sarma R. Comparison of the effects of remimazolam and inhalational anesthesia on postoperative recovery in patients undergoing general anesthesia: a systematic review and meta-analysis of randomized controlled trials. Anesth Pain Med (Seoul) 2025; 20: 393-405.
crossref pmid pmc pdf
45. Malmivaara A. Methodological considerations of the GRADE method. Ann Med 2015; 47: 1-5.
crossref pmid pmc
46. Schünemann HJ, Mustafa RA, Brozek J, Steingart KR, Leeflang M, Murad MH, et al. GRADE guidelines: 21 part 2. Test accuracy: inconsistency, imprecision, publication bias, and other domains for rating the certainty of evidence and presenting it in evidence profiles and summary of findings tables. J Clin Epidemiol 2020; 122: 142-52.
crossref pmid
47. Llanaj E, Muka T. Misleading meta-analyses during COVID-19 pandemic: examples of methodological biases in evidence synthesis. J Clin Med 2022; 11: 4084.
crossref pmid pmc
48. Kang H, Oh HC. Current concerns on journal article with preprint: anesthesia and pain medicine perspectives. Anesth Pain Med (Seoul) 2023; 18: 97-103.
crossref pmid pmc pdf
49. Munn Z, Pollock D, Barker TH, Stone J, Stern C, Aromataris E, et al. The Pandora’s box of evidence synthesis and the case for a living evidence synthesis taxonomy. BMJ Evid Based Med 2023; 28: 148-50.
crossref pmid pmc
50. Turner RM, Spiegelhalter DJ, Smith GC, Thompson SG. Bias modelling in evidence synthesis. J R Stat Soc Ser A Stat Soc 2009; 172: 21-47.
crossref pmid pmc pdf
51. Xing X, Lin L, Murad MH, Tong J. Leveraging AI for meta-analysis: evaluating LLMs in detecting publication bias for next-generation evidence synthesis. Cochrane Evid Synth Methods 2025; 3: e70047.
crossref pmid pmc pdf
TOOLS
Share :
Facebook Twitter Linked In Line it
METRICS Graph View
  • 1 Web of Science
  • 1 Crossref
  • 1 Scopus
  • 2,219 View
  • 68 Download


ABOUT
ARTICLE CATEGORY

Browse all articles >

BROWSE ARTICLES
AUTHOR INFORMATION
Editorial Office
101-3503, Lotte Castle President, 109 Mapo-daero, Mapo-gu, Seoul 04146, Korea
Tel: +82-2-792-5128    Fax: +82-2-792-4089    E-mail: journal@anesthesia.or.kr                
Business Name: Korean Society of Anesthesiologists
Business Registration: 106-82-07194
Representative: Young-Tae Jeon

Copyright © 2026 by Korean Society of Anesthesiologists.

Developed in M2PI

Close layer
prev next