Introduction
Decompressive craniectomy increases intracranial compliance when cerebral swelling or mass effect threatens perfusion, tissue viability, or survival. Its use in severe traumatic brain injury (TBI) and malignant cerebral infarction is supported by guidelines, randomized trials, pooled analyses, and reviews.1-6 International practice also includes alternative decompressive strategies,7 and recent reviews continue to summarize the evidence for decompressive craniectomy in TBI.8 Selective use has also been described for cerebral venous sinus thrombosis and ischemic stroke after thrombolysis,9,10 while broader neurocritical-care literature includes selected intracranial hemorrhage and other causes of intracranial hypertension.11 Severe-TBI management reviews provide additional context.12 Once the bone flap has been removed and the dura opened, the surgeon faces a second decision: reconstruct the dura formally with the intention of obtaining a watertight barrier, or use a faster non-watertight strategy that preserves decompression while reducing additional operative steps.
This technical decision may affect both wound separation from the subarachnoid space and operative burden. A sutured expansile graft intended to be watertight may reduce direct incisional cerebrospinal fluid (CSF) egress, but it requires graft preparation, suturing, hemostatic reassessment, and additional work during a physiologically demanding operation. General reviews and clinical reports describe a range of dural-repair methods and graft or substitute materials.13-16 Related literature has examined incisional CSF leakage and dural reconstruction or sealing strategies.17,18 Collagen-based matrices and suture, patch, or sealant constructs have also been evaluated.19-21 Additional duraplasty and collagen-based closure techniques have been reported.22-24 Published comparisons may change several components simultaneously—material, suture or attachment method, sealant, incision geometry, drain practice, and operative workflow—and therefore compare multicomponent surgical techniques rather than the isolated effect of watertight intent.
Outcome definitions add a second layer of heterogeneity. Direct leakage through the incision is clinically different from a radiographic subdural hygroma, a subgaleal fluid collection, hydrocephalus, meningitis, or a study-defined composite complication. Mortality and function are even less closure-specific and are strongly influenced by the original brain injury, the indication for decompression, and treatment limitation. When studies differ simultaneously in exposure, comparator, etiology, anatomy, outcome definition, and surveillance window, a statistically pooled result may be more precise numerically while becoming less interpretable clinically. When meta-analysis is not appropriate, alternative synthesis methods should be reported transparently; the Synthesis Without Meta-analysis reporting guideline provides a framework for such reporting.25 Certainty of evidence in this review was assessed using the Grading of Recommendations Assessment, Development and Evaluation (GRADE) framework.26
The previous meta-analysis used a broad open-versus-closed framework and included four studies involving 368 participants overall.27 Related systematic reviews and meta-analyses have examined watertight closure after supratentorial craniotomy28; decompressive craniectomy in TBI29,30; hydrocephalus or cranioplasty timing after decompression31-33; seizures after cranioplasty34; and craniotomy versus decompressive craniectomy for acute subdural hematoma.35 However, uncertainty remains about whether formal reconstruction is justified by its additional operative burden. Direct incisional CSF leakage was selected as the primary outcome because it is anatomically closest to the dural intervention and should not be conflated with radiographic hygroma, subgaleal collection, or hydrocephalus. We used an all-etiology decompressive-craniectomy scope and retained etiology and anatomy for interpretation.
Accordingly, this systematic review and meta-analysis aimed to evaluate whether formal watertight-intent dural reconstruction, compared with non-watertight management, was associated with direct postoperative incisional CSF leakage after decompressive craniectomy, and to assess the accompanying operative burden and other clinically relevant outcomes. The primary meta-analysis combined all eligible direct-leak comparisons, with randomized and observational studies examined as design subgroups and with a formal test for subgroup differences. Secondary outcomes included total operation time, wound or surgical-site infection and other complications, mortality, and functional outcomes. Outcomes with irreconcilable definitions, follow-up windows, or denominators were not pooled. The pooled estimates were interpreted as associations across the included studies rather than evidence of a common causal effect.
Materials and methods
Protocol and reporting
The review followed the Preferred Reporting Items for Systematic Reviews and Meta-Analyses (PRISMA) 2020 statement and PRISMA for Searches.36,37 The completed PRISMA 2020 checklist is provided as Supplementary File 1. This systematic review was not prospectively registered. After completion of the review, the protocol was retrospectively deposited on protocols.io (record key: j9kfcr4tp) and is publicly available at https://www.protocols.io/view/j9kfcr4tp for public documentation of the review methods.
Information sources and study selection
PubMed, Scopus, Embase, CENTRAL, and Web of Science were searched from inception through August 1, 2026, without language restriction. Database-specific strategies and execution details are reported in Supplementary Table. 1. Records were deduplicated before title-and-abstract screening. The two authors (ZM and JF) independently assessed titles and abstracts and then full texts; disagreements were resolved through discussion and consensus. The final flow comprised 9,244 records: PubMed 1,945; Scopus 3,071; Embase 2,048; CENTRAL 107; and Web of Science 2,073. After removal of 4,591 duplicates, 4,653 records underwent title-and-abstract screening and 189 reports underwent full-text assessment (Fig. 1).
Eligible reports compared dural-management strategies during decompressive craniectomy, defined as removal of a cranial bone flap with dural opening intended to increase intracranial compliance in the setting of pathological swelling or mass effect. We imposed no age restriction; eligible reports included adult or mixed-age cohorts, and no pediatric-only effect was pooled with an adult effect. Routine elective craniotomy, nondecompressive operations, single-arm technical series without an eligible comparator, and reports without an explicitly stated watertight or sealing intention for the reconstruction arm were excluded. Material-only comparisons were excluded when both arms had the same or unclassifiable closure intention, or when a watertight-intent versus non-watertight contrast could not be identified. Studies remained eligible when an arm with explicit watertight or sealing intent could be distinguished from an explicitly non-watertight comparator, even when material or fixation also differed; these were interpreted as comparisons of multicomponent surgical techniques rather than estimates of the isolated effect of watertightness.
Study identity and clinical classification
Multiple reports and overlapping cohorts were tracked separately to avoid double counting within the same outcome and time point. Etiology, supratentorial versus posterior-fossa anatomy, primary versus secondary decompression, index surgery versus cranioplasty, study design, outcome definition, and surveillance window were preserved throughout.
For analysis and reporting, formal reconstruction required an explicitly stated watertight or sealing intention confirmed in the report. Classification was based on the intended surgical endpoint, not an objective measurement of seal integrity. No included study systematically reported standardized intraoperative leak testing or postoperative verification of achieved watertightness. Ladisich described a watertight-or-near-watertight intended endpoint and was therefore treated as a broader intent-based reconstruction category. It was retained only for outcomes with compatible patient-level denominators; this classification does not imply objectively verified watertightness. Named non-watertight comparators included fixed covers, covers with unclear fixation, free or unsutured onlays, open dura/no repair/rapid approximation, multidural slits, and stellate incisions. These terms describe the overall surgical technique and do not imply that watertightness was the only difference between groups.
Data extraction and verification
The two authors (ZM and JF) independently extracted population, etiology, anatomy, design, randomization and participant flow, dural intent, comparator technique, material, numerator, denominator, mean, dispersion, adjusted estimates, outcome definitions, time points, and source locations. Discrepancies were resolved through discussion and direct rechecking of the full report. For non-English reports, two independent translation-software outputs were compared against the original tables and text. This process assisted interpretation but was not represented as certified professional translation or independent back-translation. Abstracts and previous reviews were not substituted for an available full report. Percentages were not converted into counts without a recoverable numerator, and missing dispersion was not imputed without a defensible source.
The primary outcome was direct postoperative incisional or wound CSF leakage. Subgaleal collection, subdural hygroma, hydrocephalus, wound or surgical-site infection, central nervous system (CNS) infection, and study-defined overall complications were analyzed as separate constructs. Operative outcomes included index operation time and intraoperative blood loss. Patient-important outcomes included mortality and function, retaining the reported time window, scale, and dichotomization. Later cranioplasty literature was used only as contextual evidence in the Discussion and was not included in analyses of index-operation outcomes.
Jeong reported subgaleal fluid collection but did not provide a separable arm-level count for the defined direct incisional CSF-leak outcome; it was therefore excluded from that estimand regardless of effect direction. Mitmuang reported arm-level direct CSF leakage, operation time, mortality, functional categories, reoperation, seizure, central nervous system infection, and subgaleal collection. Its surgical-site infection description did not provide separable arm-level event counts; no zero event was inferred.
Risk of bias and certainty of evidence
Risk of bias was assessed at the result level for the effect of assignment using the revised Cochrane risk-of-bias tool for randomized trials (RoB 2; 22 August 2019) and the 2016 Risk Of Bias In Non-randomized Studies—of Interventions (ROBINS-I) tool.38,39 Domain judgments were supported by evidence from the original reports and were not based on effect direction or statistical significance. GRADE certainty was assessed separately by comparison, design, outcome, and time point.26 Randomized evidence started at high certainty and nonrandomized evidence at low certainty; downgrading considered risk of bias, inconsistency, indirectness, imprecision, and publication bias. Evidence that could not be pooled was synthesized according to Synthesis Without Meta-analysis principles.25
Analysis hierarchy
The primary framework compared reconstruction undertaken with explicit watertight or sealing intent with clearly non-watertight management, separating randomized and observational studies and retaining the named comparator technique. Comparisons in which material and fixation also differed were treated as comparisons of multicomponent surgical techniques rather than isolated watertightness effects. No pooled estimate was produced where outcome definitions, time windows, denominators, or cohort independence could not be reconciled.
The primary direct-leak analysis combined all nine eligible studies. Randomized and observational studies were examined as design subgroups, and subgroup differences were formally tested. Because the overall analysis included different non-watertight techniques, clinical populations, and study designs, its pooled estimate was interpreted as an overall association across the included studies rather than a single causal treatment effect.
Statistical analysis
Binary outcomes were expressed as risk ratios (RRs) with 95% confidence intervals (CIs). Continuous outcomes reported in a common unit were expressed as mean differences (MDs), oriented as formal reconstruction minus non-watertight management. Double-zero comparisons were retained descriptively and were not forced into a relative-effect model by a mechanical continuity correction.40 Adjusted and unadjusted estimates were not combined in the same causal analysis.
Between-study variance was estimated using restricted maximum likelihood (REML).41,42 Wald-type 95% CIs were used for the primary presentation because several design-specific syntheses contained only two studies, for which Hartung–Knapp intervals were excessively unstable. Hartung–Knapp intervals were retained as sensitivity analyses.43 A single study was reported only as an individual effect; two-study pools were considered exploratory. Heterogeneity was summarized with I², tau-squared, and Cochran Q; a confidence interval for tau-squared was not calculated. Prediction intervals were displayed only for syntheses with at least five studies and were interpreted cautiously because between-study variance remains imprecisely estimated at low k. Design differences were evaluated with mixed-effects meta-regression using a shared REML variance and a Knapp–Hartung F test. Leave-one-out models refitted the primary nine-study overall analysis after removing each study. Figure 2 displays the range of refitted point estimates and the envelope of their Wald CIs; the Hartung–Knapp envelope and all individual refits are retained in the supplement. Neither envelope is a new pooled CI because its endpoints arise from different deletion models. Funnel plots, Egger tests, and meta-regression beyond the design contrast were not performed because no eligible synthesis contained at least ten studies.
Analyses used R 4.4.3 and metafor 5.0-1. Binary and continuous random-effects models were fitted with rma.uni and method=“REML”; test=“z” generated the primary Wald intervals, and the same models were refitted with test=“knha” for Hartung–Knapp sensitivity inference. Analytic code and session information were retained, and all figures were generated directly from the analytic outputs.
Results
Study selection and characteristics
Of 4,653 deduplicated records screened, 4,464 were excluded at title-and-abstract review and 189 full texts were assessed. The 177 full-text exclusions comprised no eligible comparator or a noncomparative study (n = 68), unclear watertight intent or technique (n = 38), material-only or material/closure comparisons without an identifiable eligible intent contrast (n = 23), no usable or separable data (n = 22), nondecompressive surgery or the wrong surgical stage (n = 14), and duplicate reports or ineligible publication types (n = 12). After these exclusions, 12 reports representing 12 studies were included (Fig. 1). Nine studies contributed calculable direct incisional CSF-leak effects and all 12 contributed at least one quantitative outcome. The included evidence comprised five randomized or provisionally randomized and seven nonrandomized studies, with traumatic and mixed populations and predominantly supratentorial decompression.44-55 Ladisich contributed only outcomes with compatible patient denominators54; Vychopen was excluded because the reconstruction report did not explicitly state watertight or sealing intent.56 Study characteristics and the clinical technique descriptions are summarized in Table 1.44-55 Age, sex, admission or preoperative Glasgow Coma Scale (GCS), and other reported severity measures were variably available and are summarized without imputation in Supplementary Table. 2. Notable reported imbalances included GCS in Mitmuang, sex in Sun, and GCS and etiology in Ladisich.
| Study | Country / site / recruitment | Design and participant flow | Etiology / anatomy | Formal reconstruction | Non-watertight comparator | Main contribution |
|---|
| Vieira 201845 | Brazil; Hospital da Restauração Neurotrauma Service Recife; 2012-01 to 2013-12 | Randomized; single-center prospective randomized controlled trial; allocated/treated/analyzed by reported arms: 28/27 | Mixed TBI, cerebrovascular, other; supratentorial | Watertight expansile duraplasty; n = 28 | Rapid closure without watertight repair; n = 27 | Direct CSF leak; infection; mortality; complications |
| Jeong 202046 | Korea; Gachon University Gil Medical Center single trauma center Incheon; 2017-01 to 2018-12 | Observational; retrospective single-center comparative cohort; dural technique selected by surgeon preference; allocated/treated/analyzed by reported arms: 69/37 | TBI; supratentorial | Watertight sutured reconstruction; n = 69 | Free artificial-dura onlay; n = 37 | Operation time; infection; fluid collections; cranioplasty |
| Kumar V 202647 | India; Pandit Bhagwat Dayal Sharma University of Health Sciences Postgraduate Institute of Medical Sciences Rohtak Haryana; 2020-01-01 to 2022-12-31 | Randomized; prospective randomized controlled study using lottery allocation; allocated/treated/analyzed by reported arms: 100/100 | TBI; supratentorial | Watertight duraplasty; n = 100 | Free or unsutured onlay; n = 100 | Direct CSF leak; infection; hydrocephalus; complications |
| Ammar 202648 | Egypt; Neurosurgery Department, Menoufia University, Shibin Elkom, Menoufia; 2022-01 to 2025-01 | Observational; mixed retrospective-prospective comparative cohort; retrospective classic-closure control versus prospectively collected fast-closure intervention; allocated/treated/analyzed by reported arms: 38/21 | Mixed TBI, cerebrovascular, other; supratentorial | Watertight classic duraplasty; n = 38 | Stellate dural incisions; n = 21 | Direct CSF leak; operation time; infection; mortality/function |
| Kumar A 202650 | India; Department of Neurosurgery, ABVIMS/Dr Ram Manohar Lohia Hospital, New Delhi; 2021-05-15 to 2022-04-15 | Randomized; prospective randomized comparative study using sealed-envelope block randomization; allocated/treated/analyzed by reported arms: 28/28 | Mixed TBI, cerebral infarction, SAH, DVST; supratentorial | Watertight dural reconstruction; n = 28 | Non-watertight cover, attachment unclear; n = 28 | Direct CSF leak; operation time; infection |
| Kumar P 202349 | India; Department of Neurosurgery neurotrauma unit King Georges Medical University Lucknow; 2020-01 to 2021-09 | Randomized; parallel prospective randomized controlled trial; allocated/treated/analyzed by reported arms: 60/60 | TBI; supratentorial | Watertight autologous duraplasty; n = 60 | Open dura/no repair; n = 60 | Direct CSF leak; hydrocephalus |
| Mitmuang 201655 | Thailand; Chumphon Khet Udomsakdi Hospital; 2008-01 to 2014-12 | Observational; single-center retrospective comparative cohort; allocated/treated/analyzed by reported arms: 116/91 | TBI, acute subdural hematoma; supratentorial, unilateral | Watertight sutured duraplasty; n = 116 | Unsutured layered cover; n = 91 | Direct CSF leak; operation time; complications; mortality/function |
| Sun 201852 | China; Affiliated Hospital of Logistics University of People's Armed Police Force; Sixth Department of Neurosurgery; 2011-01 to 2013-12 | Observational; adjusted analyses also reported; retrospective single-center historical comparative cohort; technique changed in July 2012; raw and adjusted analyses reported; allocated/treated/analyzed by reported arms: 192/195 | Severe TBI; supratentorial | Watertight artificial-dura repair; n = 195 | No dural repair; n = 192 | Direct CSF leak; sensitivity analysis |
| Trnka 202353 | Czech Republic; Olomouc University Hospital Neurosurgical Clinic versus Liberec Hospital Neurosurgical Department; 2019 to 2020 | Observational; retrospective bicentric comparative study; center-specific technique; allocated/treated/analyzed by reported arms: 82/45 | TBI; supratentorial, unilateral | Watertight duraplasty; n = 82 | Open dura/no duraplasty; n = 45 | Direct CSF leak; operation time; sensitivity analysis |
| Barooah 202044 | India; Guwahati Medical College; recruitment period not reported | Randomized; comparative study; allocation method not reported in sufficient detail; allocated/treated/analyzed by reported arms: 15/15 | TBI; anatomical site not reported | Complete watertight sutured temporalis-fascia closure; n = 15 | Sutureless collagen-matrix onlay; n = 15 | Operation time; material and fixation both differed |
| Zhang 202251 | China; Second Affiliated Hospital of Wenzhou Medical University and Yueqing Affiliated Hospital; 2017-01 to 2020-12 | Observational; retrospective consecutive two-center comparative cohort; allocated/treated/analyzed by reported arms: 39/64 | Severe TBI; supratentorial | Tightly sutured sealing-intent NormalGEN reconstruction; n = 64 | Sutureless DuraMax onlay; n = 39 | Direct CSF leak; hydrocephalus; material and fixation both differed |
| Ladisich 202654 | Austria; two tertiary neurosurgical departments; exact site-to-technique mapping not reported; recruitment began in March 2022; end date unclear because the Methods section reports January 2012, which is chronologically inconsistent with the stated start | Observational; retrospective bicentric consecutive comparative cohort; each center used a standard duraplasty technique; allocated/treated/analyzed by reported arms: 50/50 | Mixed etiologies; supratentorial | Watertight-or-near-watertight sutured expansile closure; n = 50 | Unsutured expansile cover; n = 50 | Hydrocephalus; complications; mortality/function; compatible patient denominators only |
Direct postoperative incisional CSF leakage
Nine studies with 1,314 participants reported separable arm-level direct-leak events (Fig. 3). Every reconstruction arm had an explicitly stated watertight or sealing intention; comparisons included fixed covers, covers with unclear fixation, free onlays, open/no repair, multidural slits, or stellate incisions. In the primary analysis, watertight-intent reconstruction was associated with a lower reported risk of incisional CSF leakage: RR 0.54 (Wald REML 95% CI 0.34–0.86), corresponding to a 46% lower relative risk. The Hartung–Knapp sensitivity interval also excluded 1 (0.31–0.93). Heterogeneity statistics were I² = 23.3%, tau² = 0.111, and Q = 10.57 on 8 degrees of freedom (P = 0.23); the prediction interval was 0.24–1.20 and therefore included the null value.
In the randomized subgroup, leakage occurred in 9/216 (4.2%) reconstruction and 12/215 (5.6%) comparator participants; the subgroup summary was RR 0.75 (Wald REML 95% CI 0.32–1.73; I² = 0%). The five observational studies (883 participants) yielded RR 0.43 (Wald REML 95% CI 0.20–0.90; I² = 60.5%). Its Hartung–Knapp sensitivity interval was 0.14–1.32 and crossed 1, unlike the Wald interval; the observational inference was therefore sensitive to interval method. The ratio of subgroup RRs was 0.62 (Wald REML 95% CI 0.20–1.89; P for subgroup difference = 0.41 by Knapp–Hartung test). This low-powered interaction test does not establish equivalence of the subgroup estimates.
Mitmuang contributed the largest explicitly watertight-intent observational comparison: direct CSF leakage occurred in 2/116 reconstruction patients and 12/91 patients managed with an unsutured layered cover, RR 0.13 (95% CI 0.03–0.57).55 The estimate is large in magnitude but unadjusted. Technique selection was not randomized, postoperative surveillance time was not explicit, and baseline injury severity differed between groups.
Robustness analyses
Figure 2 places the primary overall analysis first, followed by randomized-only, observational-only, and TBI-only estimates. Every summary point estimate favored reconstruction. The first row reproduces the overall Wald interval from Figure 3, and the randomized-only and observational-only results were as reported above. Restriction to TBI-only studies produced RR 0.46 (Wald REML 95% CI 0.23–0.89; k = 6; N = 1,144), whereas its Hartung–Knapp sensitivity interval was 0.18–1.13 and crossed 1. Thus, the direction of association was consistent, although whether the observational-only and TBI-only estimates excluded the null depended on the interval method used.
Across nine leave-one-out refits of the primary overall analysis, the association consistently favored reconstruction, and the pooled RRs ranged from 0.45 to 0.62. Figure 2 displays the Wald REML 95% CI envelope of 0.29–1.01; the Hartung–Knapp envelope was 0.25–1.15 and is reported in Supplementary Table. 3. Excluding Zhang 2022, the study that combined changes in material and fixation with closure intent, yielded RR 0.59 (Wald REML 95% CI 0.37–0.92; Hartung–Knapp 0.35–0.98). Thus, Zhang did not independently determine the direction of the primary result or whether the confidence interval excluded the null value. Removal of Sun 2018 produced the widest upper limits, and both interval methods crossed 1. After removal of Kumar A 2026, the Wald interval was 0.32–0.89 whereas the Hartung–Knapp interval was 0.28–1.02. The envelope endpoints came from different deletion models and do not form a pooled confidence interval. The direction of the point estimate was stable, but interval-based statistical significance was not completely robust to study deletion or interval method.
Index operation time and blood loss
Six studies involving 585 participants provided usable operation-time data, including two randomized or provisionally randomized studies and four observational studies. All point estimates indicated longer operations with formal reconstruction, but their magnitudes varied substantially. The two-study randomized summary was MD 51.40 minutes (Wald REML 95% CI 34.71–68.08; I² = 69.7%) and was treated as exploratory. Its Hartung–Knapp sensitivity interval was −56.76 to 159.56 minutes and crossed 0, unlike the Wald interval, showing the instability of low-k inference. Among the four observational comparisons, the pooled MD was 33.18 minutes (Wald REML 95% CI 14.79–51.58; I² = 93.7%). The across-design estimate was 39.85 minutes (Wald REML 95% CI 25.14–54.56; I² = 97.4%), with a prediction interval from 3.46 to 76.24 minutes. Complete Hartung–Knapp results are provided in Supplementary Table. 4.
All six study-level estimates indicated longer reported operation time with formal reconstruction. Mitmuang reported 92.9 ± 12.1 versus 53.4 ± 5.5 minutes, MD 39.50 (95% CI 37.03–41.98). Blood-loss evidence was sparse and was retained as unpooled outcome-specific evidence.
Wound or surgical-site infection and hydrocephalus
Five studies (476 participants) reported wound or surgical-site infection using sufficiently similar labels for a pooled display (Fig. 4). The across-design estimate was RR 0.99 (Wald REML 95% CI 0.50–1.97; I² = 0%). Design-specific estimates remained imprecise. These intervals allow clinically important benefit and harm and cannot establish equivalent infection risk. CNS infection, meningitis, abscess, and cranioplasty infection were kept separate. Mitmuang did not provide separable arm-level surgical-site infection counts, so it did not contribute to this pool and was not assigned a zero event.
Four studies (523 participants) reported hydrocephalus. The randomized and observational two-study summaries were both exploratory and highly imprecise. The across-design estimate was RR 0.91 (Wald REML 95% CI 0.57–1.46; I² = 0%). Because k = 4, no prediction interval was displayed. This synthesis included different comparator techniques and cannot establish benefit or harm. Study-defined overall complications, subdural hygroma, subgaleal fluid collection, and blood loss remained separate outcome constructs and were not combined.
Mortality and function
Mortality estimates were ordered by the reported follow-up window, and functional outcomes were separated by scale and threshold (Fig. 5). No common mortality or functional-outcome diamond was drawn. Mitmuang’s full cohort showed lower postoperative mortality with reconstruction, RR 0.40 (95% CI 0.24–0.68), and more source-defined favorable outcomes, RR 1.45 (1.10–1.91). However, the reconstruction group contained no patients with admission Glasgow Coma Scale scores of 4–5, whereas the non-watertight group contained 33. In the overlapping Glasgow Coma Scale 6–8 subgroup, mortality was RR 1.06 (0.49–2.32) and favorable function RR 0.92 (0.73–1.16). These subgroup estimates are not adjusted effects and cannot be combined with the full cohort; they illustrate the potential magnitude of baseline-severity confounding.
Other studies reported in-hospital, discharge, 6-month, or unspecified mortality and used the Glasgow Outcome Scale, the extended Glasgow Outcome Scale, or study-defined favorable and unfavorable thresholds. The intervals were generally wide and did not support a common patient-important effect.
Risk of bias and certainty of evidence
Among included studies, 25 randomized result-level RoB 2 judgments were available: 22 had some concerns and three were at high risk; none was low risk. The 42 nonrandomized result-level ROBINS-I judgments comprised 26 serious, 12 critical, and four with insufficient information; none was low or moderate risk. These totals include the final risk-of-bias assessments for Barooah’s index-operation-time result and Zhang’s reported outcome domains (Supplementary Table. 5). Concerns most often involved allocation or analysis flow, deviations from intended intervention, missing data, selective reporting, confounding by indication, center or era effects, participant selection, and exposure classification.
Certainty was very low for every key outcome. For direct leakage, the pooled Wald and Hartung–Knapp intervals excluded 1, but risk of bias, indirectness, and imprecision warranted downgrading. Operation-time, wound-infection, hydrocephalus, mortality, and functional evidence was also downgraded for risk of bias together with indirectness, inconsistency, or imprecision. Integrating the final Barooah and Zhang result-level assessments did not change any certainty rating. Supplementary Table. 6 reports the domain-level reasons and design-specific starting levels; Table 2 summarizes the key clinical outcomes.
| Outcome | Evidence | Effect | 95% CI | Certainty | Interpretation |
|---|
| Direct incisional CSF leakage | 9 studies; 1,314 participants; randomized events 9/216 vs 12/215 | RR 0.54 | Wald REML 0.34 to 0.86; HK sensitivity 0.31 to 0.93 | Very low | Formal reconstruction was associated with lower reported direct leakage in the primary overall analysis; causal certainty remains very low |
| Index operation time | 6 studies; 585 participants | MD + 39.85 min | Wald REML + 25.14 to +54.56; HK sensitivity + 20.64 to +59.06 | Very low | Formal reconstruction was associated with longer reported operation time; the magnitude is highly heterogeneous and not directly transferable |
| Wound or surgical-site infection | 5 studies; 476 participants | RR 0.99 | Wald REML 0.50 to 1.97; HK sensitivity 0.49 to 2.00 | Very low | The evidence is very uncertain about comparative wound-infection risk |
| Hydrocephalus | 4 studies; 523 participants | RR 0.91 | Wald REML 0.57 to 1.46; HK sensitivity 0.49 to 1.69 | Very low | The evidence is very uncertain about comparative hydrocephalus risk |
| Mortality | Individual effects at incompatible time windows | No pooled estimate | Not pooled | Very low | No common survival effect can be inferred |
| Functional outcome | Source-defined dichotomies and mean GOS at incompatible times | No pooled estimate | Not pooled | Very low | No common functional effect can be inferred |
Discussion
Principal findings
The primary overall analysis identified a consistent association: across nine studies, watertight-intent reconstruction was associated with fewer reported direct incisional CSF leaks (RR 0.54, Wald 95% CI 0.34–0.86), and the Hartung–Knapp sensitivity interval also excluded 1. Both design-subgroup estimates and every leave-one-out point estimate favored reconstruction. The randomized subgroup was imprecise, the prediction interval included the null value, and some restricted and leave-one-out intervals crossed 1. The direction of the association was generally consistent, but the findings do not establish that watertight intent alone caused the reduction.
Operation time showed the opposite pattern. All six study-level estimates indicated longer procedures with formal reconstruction, and the overall Wald interval excluded 0. However, I² exceeded 97%, the prediction interval ranged from approximately 3 to 76 additional minutes, operative-time definitions were not uniform, and the two-study randomized Wald and Hartung–Knapp intervals differed in whether they included the null value. Thus, non-watertight management was associated with shorter operations in the available studies, but the magnitude of this difference may not apply to other centers. Infection and hydrocephalus estimates were imprecise, and mortality and function could not be synthesized into common effects.
Why technique and material matter
A watertight intention describes a planned endpoint, not a measured physiological state. Few reports documented leak testing, suture-line pressure, or postoperative seal integrity. More importantly, most studies compared multicomponent surgical techniques. The broader dural-repair literature and technical reports demonstrate substantial variation in graft or substitute type, suturing or attachment, sealant use, sutureless or onlay constructs, dural incision geometry, and operative workflow.13-24,57-63 The observed contrast therefore cannot be attributed to watertightness alone, even when allocation was randomized.
Material and technique are also frequently entangled. Autologous fascia, pericranium, artificial dura, collagen matrix, nonabsorbable sheets, and sealants may influence tissue reaction, infection, adhesion, fluid passage, and later dissection.44,51,64 Some included reports, particularly Barooah and Zhang, changed material and fixation together with the intended seal. Their estimates therefore compare multicomponent surgical techniques and cannot isolate the effect of watertightness. Future comparisons should either hold material constant or factorially separate material from attachment and watertight intent.
Causal interpretation and clinical heterogeneity
The observational evidence is vulnerable to confounding by indication. Surgeons may choose a rapid strategy in unstable patients, severe swelling, contaminated wounds, friable dura, or technically difficult anatomy, while selecting formal reconstruction when physiology and tissue quality are more favorable. Most observational effects were unadjusted, and all were at serious, critical, or insufficient-information risk of bias. Differences in center experience, operative era, drain use, antibiotic practice, craniectomy size, and thresholds for diagnosing or treating CSF leakage could not be accounted for using arm-level data.
Mitmuang illustrates this problem. Its large reduction in reported leakage and full-cohort mortality favored reconstruction, but the pronounced baseline Glasgow Coma Scale imbalance means that mortality and functional differences cannot be assigned to dural management. The overlapping severity-restricted subgroup moved those patient-important estimates toward no association. The direct-leak endpoint is closer to the surgical exposure than mortality or function, but it remains vulnerable to selection, co-intervention, and ascertainment bias. Outcome assessment was generally unblinded, and centers may have applied different thresholds to wound dampness, intermittent drainage, clinically obvious fistula, or leakage requiring intervention.
Randomization reduces confounding but does not erase indirectness. The four randomized direct-leak studies compared different bundled techniques, populations, and reporting windows. Their Wald interval ranged from a substantial reduction to a possible increase in leakage. The absence of a design interaction is not evidence that randomized and observational studies estimated the same effect; with only nine studies the interaction test had little power.
Clinical heterogeneity further limits exchangeability. Studies included TBI, acute subdural hematoma, infarction, subarachnoid hemorrhage, venous thrombosis, and mixed etiologies; supratentorial and posterior-fossa operations; unilateral and other decompression patterns; and multiple non-watertight techniques. These settings differ in swelling mechanism, wound geometry, CSF dynamics, drain use, infection risk, and operative goal. A posterior-fossa estimate has particularly limited applicability to supratentorial TBI practice. Low statistical heterogeneity in the broad leak model cannot establish clinical homogeneity.
Relation to previous evidence
The previous meta-analysis included four studies and 368 participants overall; its CSF-leak analysis included 281 patients and reported an odds ratio of 1.04 (95% CI 0.33–3.25).27 The present review included nine studies and 1,314 participants in the direct-leak analysis. The two pooled estimates should not be compared as direct updates of the same model because the reviews differed in eligible studies, outcome definitions, and effect measures.
Compared with the previous review, the present analysis included additional and more recent studies, distinguished direct incisional CSF leakage from other fluid collections, and reported randomized and observational evidence separately. It also added prediction intervals, leave-one-out analyses, operation-time results, and outcome-specific GRADE assessments.
Infection, fluid collections, and patient-important outcomes
An incisional leak plausibly increases wound maceration, need for lumbar drainage or revision, and exposure of deeper tissues, but the present infection data were too sparse to verify that mechanistic sequence. Pooling only wound or surgical-site infection avoided conflating it with meningitis, abscess, or cranioplasty infection. The near-null summary remains compatible with meaningful benefit or harm and should not be described as evidence of equal safety.
Similarly, hygroma, subgaleal collection, and hydrocephalus should not be collapsed into a generic “CSF-related complication.” They arise from different pressure gradients, anatomical spaces, absorption pathways, and treatment thresholds. A technique could reduce incisional leakage while increasing a contained collection, or alter radiographic fluid without affecting symptoms. Keeping these constructs separate precludes a larger pooled sample but preserves clinical meaning.
Mortality and function are essential but remote from the closure decision. The original neurological injury, decompression timing, systemic insults, treatment limitation, and rehabilitation dominate these outcomes. Closure trials require adequate sample size and balanced baseline severity before they can detect a plausible mediated effect. Until such evidence exists, non-significant mortality or function comparisons must not be described as safety, equivalence, non-inferiority, or interchangeability.
Later cranioplasty was not a directly comparable outcome of the included index-procedure studies. The contextual reports were conditioned on survival and receipt of cranioplasty and therefore provide only indirect evidence.65-71
Implications for practice
The observed association with fewer reported incisional leaks suggests a clinically relevant advantage of formal watertight-intent reconstruction when prevention of wound CSF leakage is a high priority. That potential benefit must be weighed against the consistently longer operations observed with reconstruction and against the patient’s physiology, anatomy, tissue quality, and local expertise. The evidence does not mandate a uniform strategy, identify an optimal non-watertight technique, or establish comparative infection, survival, or functional effects; nor should these uncertainties be misread as evidence that the strategies are equivalent.
Operative records and future registries should document the intended endpoint, graft material, whether the graft was sutured or free, attachment pattern, sealant, leak testing, drains, and deviations from the planned technique. Recording only “duraplasty,” “open,” or “closed” prevents clinically useful comparison and promotes material–technique misclassification.
Future research
Future randomized trials should specify watertight intent before allocation and standardize the comparator technique or stratify randomization by it. Material should be held constant where possible. Allocation concealment, complete participant flow, blinded outcome adjudication, and prespecified intention-to-treat analysis are needed. Trials should report both closure-specific time and total operation time, direct incisional CSF leakage using a fixed surveillance window and severity grading, wound or surgical-site infection separately from CNS infection, individual fluid-collection outcomes, mortality and the extended Glasgow Outcome Scale at common time points, and the proportion surviving to cranioplasty.
Cranioplasty follow-up should use both patient and procedure denominators and report dissection time, dural injury, blood loss with variance, infection, hematoma, hydrocephalus, and reoperation. A prospective core outcome set would reduce selective reporting and allow perioperative benefits and harms to be assessed more consistently across studies.
Strengths and limitations
Strengths include full-report verification, careful tracking of multiple reports and overlapping cohorts, clinical technique descriptions, separation of direct leak from other fluid outcomes, design-specific synthesis, prediction intervals, leave-one-out refitting, and figures generated directly from the analytic outputs.
Several limitations should be considered. First, only nine studies and 116 leakage events informed the primary synthesis, limiting precision in design-specific and sensitivity analyses. Second, most comparisons evaluated multicomponent surgical techniques across different etiologies, anatomical settings, and comparator techniques, so the effect of watertight intent could not be separated from material, fixation, and other procedural differences. Third, most observational estimates were unadjusted and susceptible to confounding, while the randomized evidence remained limited in size and imprecise. Fourth, outcome definitions, surveillance windows, event ascertainment, and operation-time reporting were not fully standardized; publication bias could not be assessed because no synthesis contained at least ten studies. These factors reduced certainty in the exact magnitude and causal interpretation of the pooled effects. Nevertheless, every leave-one-out point estimate favored reconstruction for direct leakage, and all study-level operation-time estimates indicated longer procedures with reconstruction. The retrospective timing of protocol deposition limits assessment of deviations from the written plan.