Skip Navigation
Skip to contents

Epidemiol Health : Epidemiology and Health

OPEN ACCESS
SEARCH
Search

Articles

Page Path
HOME > Epidemiol Health > Volume 48; 2026 > Article
Review Paper
Beyond cohort versus case-control: the primacy of exposure assessment over study design in occupational cancer epidemiology
Yangwoo Kim1,2orcid, Dong-Uk Park3orcid
Epidemiol Health 2026;48:e2026026.
DOI: https://doi.org/10.4178/epih.e2026026
Published online: June 8, 2026

1Department of Social and Preventive Medicine, Inha University College of Medicine, Incheon, Korea

2Department of Occupational and Environmental Medicine, Inha University Hospital, Incheon, Korea

3Department of Environmental Health, Korea National Open University, Seoul, Korea

Correspondence: Yangwoo Kim Department of Social and Preventive Medicine, Inha University College of Medicine, 27 Inhang-ro, Jung-gu, Incheon 22332, Korea E-mail: oem.ywkim@gmail.com
• Received: March 4, 2026   • Revised: April 20, 2026   • Accepted: May 27, 2026

© 2026, Korean Society of Epidemiology

This is an open-access article distributed under the terms of the Creative Commons Attribution License (http://creativecommons.org/licenses/by/4.0/), which permits unrestricted use, distribution, and reproduction in any medium, provided the original work is properly cited.

prev next
  • 1,782 Views
  • 49 Download
  • In occupational cancer epidemiology, prospective cohort studies are conventionally ranked above case-control studies. However, exposure assessment quality—central to occupational risk estimation—receives less methodological attention than study design. This review argues that divergent cohort and case-control risk estimates reflect exposure assessment limitations rather than study design. Three findings support this. First, variance decomposition demonstrates that group-level exposure assignment attenuates effect estimates by a reliability ratio approaching zero as within-worker variability becomes dominant. Second, Berkson error in job-exposure matrices, combined with model non-linearity, biases estimates toward the null. Third, the healthy worker effect, immortal time bias, and left truncation compound misclassification and shift cohort estimates further toward the null. Two case studies illustrate this pattern. For night shift work and breast cancer, the case-control odds ratio (OR) of 1.34 contrasted with the cohort relative risk (RR) of 0.98; at 30 years’ cumulative exposure, these estimates were 1.88 and 1.13. For welding fumes and lung cancer, the pooled cohort RR of 1.29 contrasted with a case-control OR of 1.87; Sørensen’s cohort, using individual breathing-zone measurements, recovered a hazard rate ratio of 2.34 at the highest cumulative exposure. Thus, an occupational cancer study’s inferential value depends on its exposure assessment method rather than its design label. Group-level assignment attenuates effect estimates and may compromise etiological inference regardless of nominal design. Individual exposure reconstruction—through occupational interviews, industrial hygiene assessment, or personal monitoring—supports precise dose-response estimation. Quality assessment tools and meta-analyses should stratify evidence by exposure assessment method rather than design label.
For occupational carcinogens with long latency, group-level exposure assignment in prospective cohort studies systematically attenuates risk estimates. Regulatory and quality-assessment frameworks should evaluate studies by the precision of their exposure information—distinguishing individual-level reconstruction from group-level assignment—rather than by their nominal design label. Null findings from cohorts relying on job-title classification warrant caution due to this systematic attenuation.
In the conventional hierarchy of epidemiological evidence, prospective cohort studies are often ranked as the highest-quality observational design. The underlying logic is well established: exposure is ascertained before disease onset, recall bias is minimized, loss to follow-up can be quantified, and temporal precedence is unambiguous [1,2]. Systematic reviews and quality assessment tools routinely weight cohort evidence more heavily than case-control evidence [1]. By contrast, the International Agency for Research on Cancer (IARC) Monographs preamble evaluates studies according to methodological quality and informativeness rather than design label, treating well-conducted cohort and case-control studies as potentially equivalent [3]. Nevertheless, the assumption that cohorts are inherently superior continues to pervade evidence-evaluation practice in occupational epidemiology.
This ranking, however, is challenged by a persistent empirical paradox. Across IARC Group 1 and Group 2A occupational carcinogens, case-control studies consistently yield substantially higher risk estimates than cohort studies of the same exposure–outcome pair. This discrepancy was formally documented in a 4-part series in 1989 [4-7], and subsequent work has addressed bias management in occupational epidemiology [8]. However, no integrated structural explanation for this divergence has emerged.
Part of this difference is arithmetical: the odds ratio (OR) exceeds the risk ratio (RR) when the outcome is common [2]. However, occupational cancers are rare; cumulative incidence rarely exceeds 5% at any single site, and annual rates are typically below 1 per 1,000 person-years (PY). Under these conditions, the rare-disease assumption generally holds, and the OR closely approximates the RR. Yet the observed gaps exceed the likely contribution of non-collapsibility: a pooled OR of 1.34 versus an RR of 0.98 for night shift work and breast cancer [9] and an OR of 1.87 versus an RR of 1.29 for welding fumes and lung cancer [10]. A structural explanation beyond OR arithmetic is therefore required.
We argue that this paradox arises not from study design itself but from the reliance on group-level exposure assignment in prospective occupational cancer cohorts. Most prospective cohorts, as illustrated in the sections that follow, function as surveillance systems: they record job titles or broad occupational categories and assign exposures at the group rather than individual level. For carcinogens with 15–40-year latency, this approach cannot reconstruct the individual lifetime exposure histories required for etiological inference. By contrast, case-control studies conduct retrospective exposure assessment at case identification and reconstruct job-by-job histories through industrial hygiene measurements, expert judgment, and detailed interviews. When studies are evaluated by the quality of their exposure information rather than by their design label, the conventional hierarchy often inverts.
We first explain why group-level exposure assignment is limited, drawing on variance decomposition and classical measurement error theory. We then examine the mechanics of job-exposure matrices (JEMs) and Berkson-error attenuation toward the null. Next, we identify structural biases—the healthy worker effect (HWE), immortal time bias, and left truncation—that compound misclassification and push estimates below the underlying association. Two empirical case studies, night shift work and welding fume exposure, document cohort-versus-case-control divergence across pooled analyses. We conclude by synthesizing these findings, considering counterarguments, and offering recommendations for study design and evidence evaluation.
Previous reviews have addressed related themes, including misclassification [11], the HWE [12], and limitations of quality assessment tools [13]. However, none have linked exposure assessment practice to systematic underestimation in occupational cohorts. This review integrates those themes and reframes how study quality should be evaluated in occupational cancer epidemiology. In classical epidemiological theory, Miettinen argued that cohort and case-control studies are not categorically distinct designs but alternative methods of sampling the study base [14,15]. Rather than centering the debate on design hierarchy, this formulation leads to a simpler empirical question: for a given study, which sampling approach yields the less biased effect estimate? The answer depends primarily on the quality of exposure assessment rather than the nominal design label.
Structural constraints on prospective cohort designs
Occupational cancers are characterized by long latency—the interval between first exposure and clinical detection [1]. Therefore, the etiologically relevant exposure occurs long before a prospective cohort enrolls its first participant, making retrospective reconstruction a biological requirement that prospective cohorts rarely meet. Instead, these studies typically classify participants into exposure groups at baseline using job titles, census occupational codes, or self-reported trade, then carry these classifications forward. Such classifications describe occupational identity, not individual dose history.
Kromhout et al. [16] demonstrated this limitation using approximately 20,000 personal measurements across 165 Dutch job-factory groups. When total variance was decomposed into its between-worker component (σB2, between individuals within a job group) and within-worker component (σW2, day-to-day variability within workers), within-worker variance equaled or exceeded between-worker variance in most settings. Geometric standard deviations were approximately 3.0 for between-worker variability and 3.1 for within-worker variability, indicating that both varied by more than an order of magnitude. Only 25% of job-factory groups satisfied a reasonable homogeneity criterion, defined as all individual long-term means falling within a twofold range. The implication is clear: assigning a single group mean to all members of a job-title category does not approximate individual exposure but instead replaces it with noise.
Categorizing a continuous exposure distribution further reduces statistical power. Greenland [17] demonstrated analytically that dichotomizing at the median discards approximately one-third of the information; even tertiles or quartiles retain less information than continuous analysis. In terms of precision, these losses are equivalent to discarding the same fraction of the study sample. Brenner & Loomis [18] extended this work by showing that categorization can convert non-differential misclassification into differential misclassification, thereby distorting estimates unpredictably. These distortions are compounded when exposure distributions are right-skewed, as is typical in occupational settings, and when thresholds are quantile-based rather than biologically based. Early formal treatments of misclassification arising from the categorization of continuous exposures appeared in the Korean epidemiology literature [11].
The attenuation caused by classical measurement error is quantified by the reliability ratio [19]. The observed regression coefficient is attenuated relative to the true coefficient by λ: βobserved=λ×βtrue, where λ=σ2X/(σ2X+σ2U) [19]. Here, σ2X is the variance of true individual-level exposure, and σ2U is the variance of random measurement error. This equation applies to the classical-error component of exposure assessment; group-mean assignment also has a Berkson-error component, discussed below. When σ2U is large relative to σ2X, λ approaches zero and observed associations approach the null. Kromhout et al. [16]’s data do not estimate λ directly, but their substantial within-group variability shows why group-level proxies can be imprecise. In cardiovascular research, regression dilution from a single baseline blood pressure measurement underestimates true associations by approximately 60% [20]. Similar degradation is expected in occupational settings, although exact magnitudes for cancer outcomes remain unvalidated. Statistical correction does not restore significance because noise also reduces precision. Studies relying on group-level assignment may therefore be doubly penalized: their estimates can be attenuated and their tests underpowered.
These constraints are not artifacts of poor data quality; they are structural features of prospective cohort designs as applied to occupational cancer. Case-control studies that use individual interviews or expert job-history reconstruction can recover individual-level variation that group-level assignment destroys. The relevant distinction is therefore between group-level exposure assignment, such as census codes, job titles, or baseline questionnaires, and individual-level reconstruction, such as occupational interviews, expert evaluation, or personal monitoring. Observational studies of occupational cancer should be classified not only by design—case-control versus cohort—but also by function: surveillance, which is essential for population monitoring, versus etiological research. This functional distinction should be grounded in the exposure assessment method rather than the design label. Prospective occupational cohorts predominantly serve the surveillance function; etiological inference requires individual-level assessment, which they rarely achieve (Table 1) [21-24].
Job-exposure matrices and the Berkson error problem
JEMs assign exposure estimates to workers based on job title or occupational code and are the standard tool when individual interview data are unavailable. A single matrix can classify hundreds of thousands of participants at negligible cost, but this efficiency carries a measurable statistical penalty. Bouyer et al. [25] showed in a Monte Carlo simulation of 500 samples across 36 configurations that the variance of the log-OR from JEM classification was up to 4 times larger than that from expert assessment and was 3 times higher on average. Power to detect an OR of 2.0 with 250 exposed participants decreased from 0.90 under expert classification to 0.62 under JEM classification, requiring approximately 3 times the sample size for equivalent power [25]. Moreover, JEMs are anchored to job-title averages and cannot capture temporal variation in workplace exposures, task-level differences within the same title, or individual work practices that determine the biologically relevant dose.
Bhatti et al. [26] compared expert hygienist ratings with a JEM for cumulative lead exposure in the National Cancer Institute Brain Tumor Study. The interaction between the ALAD G177C polymorphism and lead exposure in meningioma was statistically significant under expert assessment (p=0.04) but entirely absent under JEM classification (p=0.90). Among ALAD2 carriers above the 95th percentile of exposure, the expert-derived OR was 13.2 (95% confidence interval [CI], 2.4 to 72.9), whereas the JEM-derived estimate was 1.1 (95% CI, 0.1 to 12.0). JEM sensitivity relative to expert judgment was only 0.52 (κ=0.47), classifying 25% of participants as exposed compared with 40% under expert assessment [26]. Thus, a biologically plausible interaction was rendered invisible by the choice of exposure assessment tool. Expert assessment is feasible in case-control studies but impractical at the cohort scale.
The statistical mechanism underlying JEM-induced attenuation is Berkson error. Armstrong [19] defined Berkson error as a situation in which “the same approximate exposure (proxy) is used for many subjects; the true exposures vary randomly about this proxy, with mean equal to it”. For example, individual exposures within a single job-title group routinely vary by more than an order of magnitude [16]; thus, a JEM-assigned group mean of 1 mg/m3 may correspond to actual exposures spanning 0.1–10.0 mg/m3. In classical measurement error, random error is added independently to each individual’s true value (Xobserved=Xtrue+ε), and the resulting misclassification attenuates regression coefficients toward the null. Berkson error reverses this structure (Xtrue=Xproxy+δ): the proxy is fixed, and individual true exposures vary around it.
Under the Berkson structure, linear regression produces no point-estimate bias: because the proxy equals the conditional mean of the true value, the slope estimated through proxy values equals the slope estimated through true individual means. The coefficient remains consistent even as the variance increases. This property has been invoked to defend JEM-based cohort analyses.
That defense does not extend to the standard models used in cancer epidemiology. Logistic regression and Cox proportional hazards models are non-linear on the natural scale, and Berkson error can induce attenuation bias in these frameworks [27,28]. Küchenhoff et al. [27] demonstrated correction methods for additive and multiplicative Berkson error in Cox models, but extension to conditional logistic models remains incomplete. Oraby et al. [28] confirmed this limitation in 813 participants in the INTEROCC study: naive surrogate exposures produced log-OR bias of −0.28 to −0.30 at a true log-OR of 0.40, whereas Berkson correction restored approximate unbiasedness in ordinary logistic regression. In conditional logistic models, however, every method examined—naive, corrected, and simulation-based—failed to recover the true parameter, with bias ranging from −0.30 to −0.35. The authors concluded that no current method provides accurate risk estimates in stratified analyses [28].
Refinements to standard JEMs, including task-exposure matrices, company-specific matrices, and population-specific matrices, partially address these limitations. When realistic misclassification is applied to both JEM and expert approaches, the performance gap narrows [25]. Nevertheless, all group-level approaches retain the Berkson structure: a shared proxy is assigned to individuals whose true exposures vary around it. The trade-off between classical error from imprecise recall and Berkson error from group-level assignment therefore persists. Only individual-level exposure measurement can avoid both sources of error.
Structural biases compounding the cohort problem
Attenuation from group-level exposure assessment does not operate in isolation. Prospective cohorts face additional structural biases that also deflect estimates toward the null. As a result, compounded effects can make null cohort findings difficult to interpret as evidence of safety.

Immortal time bias

Misalignment between the exposure-assignment time origin and the start of follow-up creates an “immortal” interval during which no event can occur. This interval is termed “immortal” because, by design, participants in the exposed group are guaranteed to survive it. When this interval is incorrectly attributed to the exposed group, the hazard ratio (HR) is artificially deflated [29]. The resulting bias can be substantial: Lévesque et al. [30] reanalyzed a cohort in which the crude HR suggested protection (HR, 0.74), but after correction for immortal time, the estimate reversed and indicated harm (HR, 1.97). Registry-based cohorts are particularly susceptible because employment records often anchor exposure status to dates that postdate cohort entry. Case-control studies that use risk-set sampling are structurally protected against this bias because each control is selected from those at risk at the exact moment of the matched case’s event, aligning exposure and event time and preventing immortal intervals.

Healthy worker effect and healthy worker survivor effect

Employment itself selects for health. The HWE suppresses standardized mortality ratios (SMRs) because workers must be healthy enough to obtain and retain employment [31]. The healthy worker survivor effect (HWSE) compounds this bias: workers whose health deteriorates selectively leave the workforce, yielding a progressively healthier remaining cohort [32]. The meta-analysis by Hwang et al. [12] of 6 semiconductor worker cohorts illustrates the magnitude of this effect. Without HWE adjustment, the all-cancer summary SMR was 0.70; after reference SMR (rSMR) correction, it increased to 1.38 [12]. For leukemia, an SMR of 0.88 became an rSMR of 1.88, reversing apparent protection into nearly twofold excess risk.

Left truncation

Occupational cohorts often enroll mid-career workers, who have, by definition, survived early working hazards. This left truncation introduces survivor bias: the observed cohort underrepresents workers most susceptible to early occupational harm, including those who left the workforce or died before observation began. Left truncation operates independently of the HWE. Even among workers initially healthy enough for hire, those who experience early harm are excluded if they do not survive to the study’s observation window. Applebaum et al. [33] quantified this problem through Monte Carlo simulation, demonstrating that a true exposure coefficient of 0.05 was attenuated to −0.18 among workers hired 30 or more years before baseline. In the framework of Applebaum et al. [33], the true exposure coefficient was the underlying regression coefficient that would have been observed with complete exposure histories from employment onset, absent measurement or truncation error. The progressive downward gradient with increasing pre-study tenure supports a frailty-selection mechanism. Left truncation is inherently a cohort-design problem; population-based case-control studies do not accumulate this distortion equivalently, although nested case-control studies inherit it from their parent cohort.

Compound bias and the problem of null findings

These 4 sources—group-level misclassification, immortal time bias, the healthy worker and survivor effects, and left truncation—are not independent perturbations that might cancel across studies. They are structurally embedded in cohort designs and directionally consistent, with each pushing observed estimates toward or below the null. Their simultaneous operation increases the probability that a true occupational carcinogen will yield a null or apparently protective cohort finding. Accordingly, a null cohort result should not be interpreted as evidence of safety without careful consideration of distortions arising from these design-inherent features; instead, it may reflect compounded structural bias (Figure 1).

G-methods as partial, data-dependent solutions

G-estimation, inverse probability weighting, and the parametric g-formula offer principled approaches to time-varying confounding and the HWSE [34]. These methods address confounding, including time-varying selection into employment, but they cannot correct the information loss inherent in group-level exposure assignment. When cohorts provide repeatedly measured individual data, such as individual radiation dosimetry records, these methods can yield more valid estimates. However, cohorts that rely on JEM-derived assignments typically lack the individual time-varying exposure and covariate data required for g-methods. In this common circumstance, analytic sophistication cannot compensate for structural data deficiency, and the compound bias described above may be difficult to eliminate regardless of the statistical approach.
Empirical case studies

Case study 1: night shift work and breast cancer

The association between night shift work and breast cancer is the most extensively studied exposure–outcome pair in occupational cancer epidemiology and provides a useful test of the theoretical framework developed above. The design-dependent divergence is among the largest consistently replicated patterns in this literature, and the biological mechanism is well characterized.
The proposed pathway runs from light-at-night exposure through suprachiasmatic nucleus disruption, melatonin suppression, estrogen synthesis upregulation, and mammary cell proliferation. IARC classified shift work involving circadian disruption as a Group 2A probable carcinogen in 2007 and reaffirmed this classification in 2019 [35,36]. The IARC Working Group concluded that animal evidence supports the carcinogenic effects of altered light–dark schedules [37]. When the mechanism is plausible but the signal is absent in one design, the design itself warrants scrutiny.
Van et al. [9] conducted a meta-analysis of 15 cohort studies, the pooled RR was 0.98 (95% CI, 0.93 to 1.03). Night shift exposure was typically captured by a single baseline questionnaire item, usually a binary ever/never variable, with no measurement of duration, intensity, rotation speed, or chronotype. Individual cohort studies with more refined exposure metrics yielded elevated risk estimates at long durations: Nurses’ Health Study I, RR 1.36 (95% CI, 1.04 to 1.78) for ≥30 years; Nurses’ Health Study II, RR 1.79 (95% CI, 1.06 to 3.01) for >20 years; the WOLF cohort, HR 2.02 (95% CI, 1.03 to 3.95) for shifts with night work; and the Swedish Twin Registry, HR 1.68 (95% CI, 0.98 to 2.88) for >20 years and HR 1.77 (95% CI, 1.03 to 3.04) for the subgroup aged <60 years [38,39,42,43]. In two nested case-control analyses within the Norwegian nurses’ cohort, Lie et al. [44] reported an OR of 2.21 (95% CI, 1.10 to 4.45) for 30 or more years of night work, the same group [45] subsequently found an OR of 1.80 (95% CI, 1.10 to 2.80) among women with 5 or more years of work involving 6 or more consecutive night shifts. This gradient within the cohort literature implicates measurement quality, rather than a true absence of effect, as a determinant of the null finding.
Case-control studies with detailed lifetime exposure interviews present a markedly different pattern. In the same meta-analysis, 13 case-control studies yielded a pooled OR of 1.34 (95% CI, 1.17 to 1.53, I2=63.8%) [9]. The pooled estimate from Cordina-Duverger et al. [46] reached an OR of 2.55 (95% CI, 1.03 to 6.30) among premenopausal women with at least 10 years of exposure at 3 or more nights per week. In the CECILE study, Menegaux et al. [47] reported an OR of 1.35 (95% CI, 1.01 to 1.80) for overnight shift work. Studies with the most detailed exposure assessment produced even larger estimates: the GENICA study reported an OR of 4.73 (95% CI, 1.22 to 18.36) for estrogen receptor-negative tumors among women with at least 20 years of night shift work [48], and Hansen & Lassen [49], in a nested case-control study among women in the Danish military cohort questionnaire study, found an OR of 3.9 (95% CI, 1.6 to 9.5) among women with morning chronotype and high cumulative night shift exposure (≥884 shifts). These studies share a feature that most cohorts lack: complete occupational histories with data on shift frequency, consecutive nights, rotation, and chronotype.
This divergence widens systematically with cumulative exposure, as expected under non-differential misclassification, in which measurement error accumulates with longer exposure duration and progressively attenuates the observed estimate. Van et al. [9] reported this contrast directly: case-control OR 1.34 (95% CI, 1.17 to 1.53) versus cohort RR 0.98 (95% CI, 0.93 to 1.03). The two-stage dose-response meta-analysis by Moon et al. [50], which incorporated 10 cohort and 11 case-control studies, extends this contrast across the full exposure continuum. At 30 years of cumulative night shift work, the cohort-derived RR was 1.13 (95% CI, 1.04 to 1.23), whereas the case-control-derived OR was 1.88 (95% CI, 1.38 to 2.57), corresponding to a 13% versus 88% increase in pooled risk, respectively (Table 2 and Figure 2A). This near-sevenfold disparity in excess risk widens with cumulative exposure, consistent with progressive misclassification of long-term, high-intensity workers when only a single early-career questionnaire item is available.
The principal objection to case-control evidence is differential recall. Under this hypothesis, cases may overreport because of illness awareness or underreport because of denial or minimization of occupational risks. Vestergaard et al. [51] linked self-reported night shift histories from 27,438 Danish respondents, including 225 breast cancer cases and 1,800 controls, to employer payroll records. Sensitivity was 86.2% in cases and 80.6% in controls, yielding a differential of 5.6 percentage points with a CI crossing zero. Correcting for this differential reduced the OR from 1.12 to 1.05 (95% CI, 0.95 to 1.16), an approximately 6% attenuation [51]. A 6% attenuation from recall bias cannot explain a difference of 0.36 on the effect-estimate scale, as observed between the case-control OR of 1.34 and the cohort RR of 0.98. Differential recall is therefore a minor, quantifiable source of bias in this comparison, whereas exposure assessment quality in cohort designs appears to be the dominant driver of the observed attenuation.

Case study 2: welding fumes and lung cancer

IARC reclassified welding fumes from Group 2B to Group 1 in 2017 (Monograph 118), based on sufficient evidence from both cohort and case-control studies and supported by mechanisms including chronic pulmonary inflammation and local immunosuppression from metal particulate deposition [52]. As a second empirical test, welding fumes provide a deliberate contrast to shift work: they are a chemically complex aerosol acting through pulmonary toxicity and are defined by a manual trade rather than work scheduling. Despite these differences, the same pattern of design-dependent effect estimates emerges (Table 3) [10,21-23,53-57].
The most comprehensive synthesis to date is the World Health Organization/International Labour Organization Joint Estimates systematic review and meta-analysis by Loomis et al. [53], which pooled 40 studies, including 29 case-control and 11 cohort studies, with more than 1.26 million participants. The pooled lung cancer incidence RR was 1.48 (95% CI, 1.29 to 1.70; I2=24%; 23 studies; 57,931 participants), and the evidence was graded as moderate quality and classified as “sufficient evidence of harmfulness.” The mortality estimate was 1.27 (95% CI, 1.04 to 1.56). Three features of this synthesis bear directly on the design-divergence question. First, exposure was classified dichotomously (any/high vs. no/low), typically by job title or self-report; this group-level approach approximates the cohort-style assessment that this review identifies as a principal source of attenuation. Second, case-control ORs were converted to RRs under the rare-disease assumption and pooled with cohort estimates rather than stratified by design, leaving design-related differences unexamined. Third, several included studies were judged to be at high risk of misclassification bias, independently corroborating that exposure assessment quality is a dominant methodological concern in this literature. The pooled RR of 1.48, derived largely from group-level exposure assessment, therefore serves as an authoritative lower-bound estimate against which design-stratified analyses with detailed exposure reconstruction can be benchmarked.
Three design-stratified meta-analyses across 3 decades reveal a striking pattern within this evidence base. Moulin [54] found a case-referent RR of 1.72 versus a cohort RR of 1.27 across all welding categories. Ambroise et al. [55], restricting the analysis to 60 studies without positive reporting bias, reported nearly equal pooled estimates: 1.27 (95% CI, 1.11 to 1.46) for case-control studies and 1.29 (95% CI, 1.19 to 1.40) for cohort studies. This finding suggests that part of the earlier case-control excess reflected selective reporting. However, in the most recent and largest design-stratified analysis, Honaryar et al. [10] reported a pooled RR of 1.29 for cohort studies versus a pooled OR of 1.87 for case-control studies. The pooled cohort estimate is remarkably stable across all 3 meta-analyses (1.27 to 1.29), whereas the case-control estimate varies with the inclusion of studies using detailed exposure reconstruction. This pattern—a ceiling on cohort effects alongside upward heterogeneity in case-control effects when methods are rigorous—is consistent with the argument that group-level exposure assignment through occupational codes and generic JEMs caps detectable effects in cohorts, whereas individual-level assessment in case-control studies through detailed interviews and expert-defined occupation categories can recover larger effects when exposure is precisely reconstructed, as in the SYNERGY pooled analysis described below.
The SYNERGY pooled analysis provides the most comprehensive quantification of the association. Kendzia et al. [21] pooled 15,483 men lung cancer cases and 18,388 controls from 16 studies conducted between 1985 and 2010; 81% of participants completed face-to-face interviews, and exposure was classified using ISCO-68 expert-defined welding categories. The overall OR for ever-regular welding was 1.44 (95% CI, 1.25 to 1.67), adjusted for age, study center, smoking, and other occupational lung cancer risk factors. Among never-smokers, the OR was 2.34 (95% CI, 1.31 to 4.17); among never-smokers and light smokers combined (<10 pack-years), it was 1.96 (95% CI, 1.37 to 2.79). Welding duration >25 years yielded an OR of 1.77 (95% CI, 1.31 to 2.39; p for trend <0.001). Squamous cell and small cell histologies had ORs of 1.58 and 1.41, respectively, and a relative excess risk due to interaction of 3.72 indicated positive additive interaction with smoking. In the subgroup analysis by Honaryar et al. [10], 8 case-control studies adjusted for both smoking and asbestos yielded a pooled OR of 1.17 (95% CI, 1.04 to 1.38), compared with 1.87 in the full 15-study pooled estimate. A parallel pattern appeared in cohort studies: among 6 cohorts adjusting for smoking, the pooled meta-RR decreased from 1.29 to 1.10. These subgroup comparisons indicate that residual confounding contributes to elevated estimates but does not fully explain the design-dependent divergence, because confounding affects both designs similarly.
Additional case-control studies with detailed exposure assessment support this pattern: Vallières et al. [56] reported elevated lung cancer risks among light smokers exposed to gas-welding fumes (OR, 2.9; 95% CI, 1.7 to 4.8) or arc-welding fumes (OR, 2.3; 95% CI, 1.3 to 3.8), and Pesch et al. [57] found excess risks at high exposure to welding fumes (OR, 1.55; 95% CI, 1.17 to 2.05), hexavalent chromium (OR, 1.85; 95% CI, 1.35 to 2.54), and nickel (OR, 1.60; 95% CI, 1.21 to 2.12).
The contrast between 2 welding cohorts with radically different exposure assessment methods illustrates how measurement precision drives observed estimates. Sørensen et al. [22] followed 4,539 Danish metal workers for 125,762 PY from 1968 to 2003, reconstructing lifetime cumulative welding-fume exposure from more than 1,000 process-specific breathing-zone measurements combined with detailed occupational questionnaires. Among stainless steel welders, the adjusted hazard rate ratio (HRR) increased monotonically across cumulative-exposure categories of 0–5, 6–10, and ≥11 mg/m3·yr, reaching 2.34 (95% CI, 1.03 to 5.28) in the highest category after adjustment for age, smoking, and asbestos. The overall standardized incidence ratio was 1.35 (95% CI, 1.06 to 1.70); for stainless steel welders with ≥21 years of employment, it was 3.69 (95% CI, 1.77 to 6.79). Mild steel welding showed no significant dose-response, consistent with the lower toxicity of iron oxide relative to hexavalent chromium and nickel from stainless steel welding. In contrast, Siew et al. [23] applied a generic Finnish JEM to the 1970 census cohort of economically active Finnish men and found a modest RR of 1.15 (95% CI, 0.90 to 1.46) in the highest welding-fume exposure category. The divergence between these 2 cohorts reflects exposure assessment precision: individual breathing-zone quantification reveals dose-response gradients that JEM-based census cohorts cannot detect (Table 3, Figure 2B).
Because the HRR of 2.34 in the study of Sørensen et al. [22] was obtained after individual-level adjustment for smoking and asbestos, the design-dependent attenuation observed in most cohorts cannot be attributed primarily to differential confounding. The operative factor appears to be exposure assessment precision: cohorts with individual measurement detect dose-response gradients that proxy-based cohorts render invisible.
The theoretical framework developed above identifies 3 sources of attenuation toward the null: group-level misclassification, Berkson error, and structural biases. The 2 case studies support this pattern empirically. These lines of evidence converge on a central conclusion: most occupational cohort studies that rely on group-level assignment function as surveillance systems rather than etiological instruments. The Korean Workers’ Compensation-National Health Insurance Service (KoWorC-NHIS) cohort, which links 858,793 compensated workers to NHIS records through occupational codes and injury/disease registries, exemplifies this surveillance model [11,24,58]. Surveillance evidence—group-level monitoring based on job titles, census codes, or registry linkage—is well suited to describing patterns of disease occurrence and informing research priorities, but its coarse exposure resolution is inadequate for dose-response estimation. Recognizing this boundary is essential for methodological clarity.
This framing does not deny the value of cohort studies. When cohorts achieve individual-level exposure quantification, as demonstrated in the welding cohort by Sørensen et al. [22], they recover dose-response gradients comparable in magnitude to high-quality case-control estimates [21].
Two objections deserve attention. Recall bias is a genuine concern in case-control studies because cases may selectively recall or misattribute occupational exposures after a cancer diagnosis. The payroll-record validation by Vestergaard et al. [51], however, documented a differential accuracy of only 5.6 percentage points, translating to an approximately 6% attenuation in the OR, which is far too small to explain a 0.36 gap on the effect-estimate scale between case-control and cohort studies [9,51]. Confounding is a more substantial concern. In the meta-analysis by Honaryar et al. [10], the subgroup of 8 case-control studies adjusted for both smoking and asbestos yielded a pooled OR of 1.17, compared with 1.87 in the full 15-study pooled estimate. A similar pattern appeared in cohort studies, in which smoking-adjusted estimates (pooled meta-RR, 1.10) were lower than unadjusted estimates (pooled meta-RR, 1.29), indicating that confounding affects both designs. The cohort of Sørensen et al. [22], with individual-level adjustment for smoking and asbestos, still revealed a significant cumulative-exposure gradient reaching an HRR of 2.34 at the highest exposure level, indicating that the association persists when exposure is individually measured.
On the cohort side, Papantoniou & Hansen [59] proposed a 5-point taxonomy of cohort limitations: crude exposure definitions, absence of intensity metrics, incomplete occupational histories, insufficient follow-up, and inclusion of long-tenured survivor participants. This taxonomy maps directly onto the structural mechanisms identified above.
A further bias operates at the level of evidence synthesis. In the review by Bae [13] of nutritional-epidemiology meta-analyses, the Newcastle-Ottawa Scale rated 91% of cohort studies as high quality versus 34% of case-control studies. This pattern reflects the scale’s design-level weighting and generalizes beyond nutrition, as the scoring criteria equate temporal structure with methodological quality while disregarding exposure assessment [13]. Cohort studies consequently dominate systematic reviews, contributing attenuated pooled estimates that may delay or weaken regulatory action.
This review is narrative rather than systematic, and its conclusions apply specifically to occupational carcinogens with long latency. Systematic extension to all IARC Group 1 and Group 2A carcinogens, as well as evaluation of whether the same divergence holds for acute outcomes, remain areas for future research.
We propose 4 changes to current practice. First, nested case-control studies with individual exposure assessment should be the default design for effect estimation. Second, cohort studies using group-level JEM assignment should be explicitly characterized as monitoring instruments: valuable for tracking cancer burden but unsuitable for dose-response estimation. Third, quality assessment tools should incorporate exposure assessment quality as a scored domain [13]. Fourth, meta-analyses should stratify pooled estimates by exposure assessment method rather than design label. The contrast between the dose-response gradient reported by Sørensen et al. [22], with an HRR of 2.34 at ≥11 mg/m3·yr cumulative exposure using individual breathing-zone measurement, and Siew et al. [23]’s flat JEM-based census cohort estimate, with an RR of 1.15, demonstrates that “cohort study” encompasses a range of inferential quality that a single pooled estimate cannot represent.
The conventional evidence hierarchy cannot be applied uncritically in occupational cancer epidemiology. Prospective cohort studies that rely on job-title or JEM-based classification are structurally constrained to surveillance-level exposure assessment. Kromhout et al. [16]’s variance decomposition, the attenuation coefficient λ from classical measurement error theory, directional biases from the HWE, immortal time bias, and left truncation, and the empirical evidence from the case studies all converge on the same conclusion: cohort studies of occupational carcinogens with long latency routinely underestimate underlying associations, sometimes to the point of obscuring them entirely.
Case-control studies with individual-level exposure reconstruction through retrospective industrial hygiene assessment, expert review, and detailed occupational history should be recognized as the methodologically appropriate default for etiological research on occupational carcinogens. Cohort studies retain value as monitoring instruments for cancer incidence trends across broad occupational groups. The few prospective cohorts that achieve individual-level measurement can function as true etiological studies and should be distinguished from the group-assignment majority. The critical distinction for both research design and evidence synthesis is not the nominal study label but the quality of exposure information achieved.

Conflict of interest

The authors have no conflicts of interest to declare for this study.

Funding

None.

Acknowledgements

None.

Author contributions

Both authors contributed equally to conceiving the study, analyzing the data, and writing this paper.

Figure 1.
Compound bias mechanisms attenuating effect estimates in prospective occupational cohort studies using group-level exposure assessment. The left pathway depicts 4 directionally consistent bias sources that independently push estimates toward the null: exposure misclassification through λ attenuation, Berkson error in non-linear models, the healthy worker effect and HWSE, and immortal time bias compounded by left truncation. The right pathway shows that case-control studies with individual-level exposure reconstruction and risk-set sampling can avoid these mechanisms. Each mechanism independently attenuates; compound operation renders null findings ambiguous. OR, odds ratio; RR, relative risk; JEM, job-exposure matrix; λ, attenuation factor/reliability ratio; σB2, between-group variance; σW2, within-group variance; HWSE, healthy worker survivor effect; SMR, standardized mortality ratio; HR, hazard ratio.
epih-48-e2026026f1.jpg
Figure 2.
Design-specific effect estimates for night shift work and breast cancer (A) and welding fumes and lung cancer (B). Squares denote cohort studies with group-level exposure assessment, triangles denote cohort studies with individual-level exposure measurement, and circles denote case-control studies. Filled symbols represent individual studies, and open symbols indicate summary or meta-analysis estimates. All data points include 95% confidence intervals (CIs), shown as horizontal whiskers. The dashed line indicates the null value of 1.0, and the x-axis is plotted on a logarithmic scale. In Panel B, the bracket highlights the cohort study by Sørensen et al. [22] using individual air sampling (hazard rate ratio, 2.34; 95% CI, 1.03 to 5.28), which aligns with case-control estimates rather than group-assignment cohort estimates. JEM, job-exposure matrix.
epih-48-e2026026f2.jpg
epih-48-e2026026f3.jpg
Table 1.
Exposure assessment methods in occupational cancer epidemiology by study design
Characteristics Census-based codes JEM Expert assessment Individual measurement
Precision relative to expert assessment Lower than JEM ~1/3 of expert1 Reference Variable
Bias direction (logistic/Cox) Toward null Toward null Unpredictable2 Minimal
Assignment level Group Group Semi-individual Individual
Typical study design Registry-linked cohort Prospective cohort Case-control Specialized cohort
Berkson error Yes Yes Minimal No
Classical error Minimal Minimal Moderate Minimal
Temporal coverage Baseline only Baseline or cross-sectional Lifetime reconstruction Duration of monitoring
Feasibility at cohort scale High High Low Low
Example studies Population registries; Min et al. 2024 (KoWorC–NHIS) [24] Siew et al. 2008 (FINJEM) [23] Kendzia et al. 2013 (SYNERGY) [21] Sørensen et al. 2007 (welding fumes) [22]

JEM, job-exposure matrix; KoWorC–NHIS, Korean Workers’ Cohort–National Health Insurance Service; FINJEM, Finnish job-exposure matrix; OR, odds ratio.

1 Bouyer et al. [25]: the JEM variance of log-OR was 2–3 times larger than expert assessment; the power to detect an OR of 2.0 dropped from 0.90 to 0.62.

2 Expert assessment introduces classical measurement error, which can bias estimates away from the null when misclassification is differential.

Table 2.
Night shift work and breast cancer: key studies by design
Studies Study design Country n (cases/total or controls) Exposure metric Effect estimate (95% CI)1 Exposure assessment
Meta-analyses
 Van et al., 2021 [9] Meta (15 cohorts) Multiple Pooled Ever/never+duration RR 0.98 (0.93, 1.03) Baseline questionnaire
 Van et al., 2021 [9] Meta (13 case-controls) Multiple Pooled Ever/never+duration OR 1.34 (1.17, 1.53) Lifetime occupational interview
 Moon et al., 2024 [50] Dose-response meta (10 cohorts) Multiple Pooled Cumulative years RR 1.13 (1.04, 1.23) at 30 yr Baseline questionnaire
 Moon et al., 2024 [50] Dose-response meta (11 case-controls) Multiple Pooled Cumulative years OR 1.88 (1.38, 2.57) at 30 yr Lifetime occupational interview
Cohort studies
 Travis et al., 2016 [40] Prospective cohort (Million Women Study) UK 4,809/522,246 Ever night shift RR 1.00 (0.92, 1.08) Baseline questionnaire
 Akerstedt et al., 2015 [42] Prospective cohort (Swedish Twin Registry) Sweden 463/13,656 Years of night work HR 1.68 (0.98, 2.88) for >20 yr Computer-assisted telephone interview
Case-control studies
 Lie et al., 2006 [44] Nested case-control Norway 537/2,143 Years of night work OR 2.21 (1.10, 4.45) for ≥30 yr Nurse-registry and census-based work-history reconstruction
 Lie et al., 2011 [45] Nested case-control Norway 699/895 Consecutive night shifts and years OR 1.80 (1.10, 2.80) for ≥5 yr with ≥6 consecutive night shifts Computer-assisted telephone interview of lifetime occupational history
 Cordina-Duverger et al., 2018 [46] Pooled population-based case-control (5 studies) Australia, Canada, France, Germany, Spain 6,093/6,933 Duration×frequency OR 2.55 (1.03, 6.30) for ≥10 yr at ≥3 night/wk among premenopausal women Lifetime occupational interview
 Menegaux et al., 2013 [47] Population-based case-control (CECILE) France 1,232/1,317 Duration and frequency OR 1.35 (1.01, 1.80) for overnight shifts Face-to-face interview of lifetime occupational history
 Rabstein et al., 2013 [48] Case-control (GENICA) Germany 857/892 Duration, ER status OR 4.73 (1.22, 18.36) for ER-negative Lifetime interview
 Hansen et al., 2012 [49] Nested case-control Denmark 141/551 Cumulative no. of night shifts+chronotype OR 3.9 (1.6, 9.5) for morning chronotype with high cumulative exposure (≥884 shifts) Postal structured questionnaire, supplemented by telephone interview
Validation study
 Vestergaard et al., 2024 [51] Cross-sectional validation Denmark 225 cases/1,800 controls Self-report vs. payroll Cases: sensitivity 86.2%, specificity 82.6% Payroll record linkage

CI, confidence interval; ER, estrogen receptor; HR, hazard ratio; OR, odds ratio; RR, relative risk.

1 Effect estimates are presented as reported in the original publications; Studies are grouped by design to illustrate the systematic divergence in effect magnitude; The validation study by Vestergaard et al. [51] found a differential recall of 5.6 percentage points (CI crossing zero), corresponding to approximately 6% attenuation of the OR.

Table 3.
Welding fumes and lung cancer: key studies by design
Studies Study design Country n (cases/controls orPY) Exposure assessment Effect estimate (95% CI) Smoking adjustment
Meta-analyses
 Honaryar et al., 2019 [10] Meta (cohort studies) Multiple 22 cohort studies JEM/occupational codes Meta-RR 1.29 (1.20, 1.39) Varied
 Honaryar et al., 2019 [10] Meta (case-control studies) Multiple 15 case-control studies Expert/interview Meta-OR 1.87 (1.53, 2.29) Varied
 Honaryar et al., 2019 [10] Meta (adjusted for smoking+asbestos) Multiple 8 case-control studies Mixed Meta-OR 1.17 (1.04, 1.38) Yes
 Moulin, 1997 [54] Meta-analysis Multiple Pooled Mixed Excess lung cancer in welders2 Varied
 Ambroise et al., 2006 [55] Updated meta-analysis Multiple Pooled Mixed Excess lung cancer confirmed2 Varied
 Loomis et al., 2022 [53] WHO/ILO systematic review and meta-analysis Multiple 40 studies; >1,265,512 participants Indirect methods including job title, self-report, and expert assessment Incidence RR 1.48 (1.29, 1.70); Mortality RR 1.27 (1.04, 1.56) Varied
Cohort studies1
 Sørensen et al., 2007 [22] Prospective cohort Denmark 75 cases/125,762 PY Individual (1,000+breathing-zone measurements; mg/m3×yr) HRR 2.34 (1.03, 5.28) at ≥11 mg/m3×yr Yes
 Sørensen et al., 2007 [22] Prospective cohort Denmark Same cohort Individual SIR 3.69 (1.77, 6.79) for SS employment ≥21 yr -
 Siew et al., 2008 [23] Prospective cohort Finland 30,137 lung cancer cases/1.2 million men JEM (FINJEM) linked to 1970 census occupation RR 1.15 (0.90, 1.46) for highest cumulative welding-fume exposure Yes
Case-control studies
 Kendzia et al., 2013 [21] Pooled case-control (SYNERGY) Europe, Canada, China, New Zealand 15,483/18,388 Expert-defined categories, face-to-face interviews (81%) OR 1.44 (1.25, 1.67) overall Yes
 Kendzia et al., 2013 [21] Same (never-smokers) Same Never-smokers subgroup Same OR 2.34 (1.31, 4.17) N/A
 Kendzia et al., 2013 [21] Same (never+light smokers, <10 pack-yr) Same Never+light smokers subgroup Same OR 1.96 (1.37, 2.79) N/A
 Kendzia et al., 2013 [21] Same (>25 yr duration) Same >25 yr duration subgroup Same OR 1.77 (1.31, 2.39), p-trend <0.001 Yes
 Vallières et al., 2012 [56] Pooled population-based case-control studies Canada 1,593/1,960 Lifetime occupational interview+expert chemist-hygienist assessment Light smokers: gas-fume OR 2.9 (1.7, 4.8); arc-fume OR 2.3 (1.3, 3.8) Yes
 Pesch et al., 2019 [57] Pooled population-based case-control studies Germany 3,418/3,488 Welding-process exposure matrix linked to job-specific questionnaire High exposure: fumes OR 1.55 (1.17, 2.05); Cr(VI) OR 1.85 (1.35, 2.54); Ni OR 1.60 (1.21, 2.12) Yes

PY, person-years; CI, confidence interval; JEM, job-exposure matrix; WHO, World Health Organization; ILO, International Labour Organization; OR, odds ratio; RR, relative risk; HRR, hazard rate ratio; SIR, standardized incidence ratio; SS, stainless steel welding; Cr(VI), chromium 6; Ni, nickel; N/A, not available.

1 Sørensen et al. [22] reported the only prospective study with individual cumulative exposure measurement (breathing-zone air sampling), providing a direct comparison with the JEM-based Siew et al. [23] cohort of similar design.

2 Specific pooled estimates not reported here as they predate the Honaryar et al. [10] comprehensive meta-analysis.

Figure & Data

References

    Citations

    Citations to this article as recorded by  

      Figure
      • 0
      • 1
      • 2
      Beyond cohort versus case-control: the primacy of exposure assessment over study design in occupational cancer epidemiology
      Image Image Image
      Figure 1. Compound bias mechanisms attenuating effect estimates in prospective occupational cohort studies using group-level exposure assessment. The left pathway depicts 4 directionally consistent bias sources that independently push estimates toward the null: exposure misclassification through λ attenuation, Berkson error in non-linear models, the healthy worker effect and HWSE, and immortal time bias compounded by left truncation. The right pathway shows that case-control studies with individual-level exposure reconstruction and risk-set sampling can avoid these mechanisms. Each mechanism independently attenuates; compound operation renders null findings ambiguous. OR, odds ratio; RR, relative risk; JEM, job-exposure matrix; λ, attenuation factor/reliability ratio; σB2, between-group variance; σW2, within-group variance; HWSE, healthy worker survivor effect; SMR, standardized mortality ratio; HR, hazard ratio.
      Figure 2. Design-specific effect estimates for night shift work and breast cancer (A) and welding fumes and lung cancer (B). Squares denote cohort studies with group-level exposure assessment, triangles denote cohort studies with individual-level exposure measurement, and circles denote case-control studies. Filled symbols represent individual studies, and open symbols indicate summary or meta-analysis estimates. All data points include 95% confidence intervals (CIs), shown as horizontal whiskers. The dashed line indicates the null value of 1.0, and the x-axis is plotted on a logarithmic scale. In Panel B, the bracket highlights the cohort study by Sørensen et al. [22] using individual air sampling (hazard rate ratio, 2.34; 95% CI, 1.03 to 5.28), which aligns with case-control estimates rather than group-assignment cohort estimates. JEM, job-exposure matrix.
      Graphical abstract
      Beyond cohort versus case-control: the primacy of exposure assessment over study design in occupational cancer epidemiology
      Characteristics Census-based codes JEM Expert assessment Individual measurement
      Precision relative to expert assessment Lower than JEM ~1/3 of expert1 Reference Variable
      Bias direction (logistic/Cox) Toward null Toward null Unpredictable2 Minimal
      Assignment level Group Group Semi-individual Individual
      Typical study design Registry-linked cohort Prospective cohort Case-control Specialized cohort
      Berkson error Yes Yes Minimal No
      Classical error Minimal Minimal Moderate Minimal
      Temporal coverage Baseline only Baseline or cross-sectional Lifetime reconstruction Duration of monitoring
      Feasibility at cohort scale High High Low Low
      Example studies Population registries; Min et al. 2024 (KoWorC–NHIS) [24] Siew et al. 2008 (FINJEM) [23] Kendzia et al. 2013 (SYNERGY) [21] Sørensen et al. 2007 (welding fumes) [22]
      Studies Study design Country n (cases/total or controls) Exposure metric Effect estimate (95% CI)1 Exposure assessment
      Meta-analyses
       Van et al., 2021 [9] Meta (15 cohorts) Multiple Pooled Ever/never+duration RR 0.98 (0.93, 1.03) Baseline questionnaire
       Van et al., 2021 [9] Meta (13 case-controls) Multiple Pooled Ever/never+duration OR 1.34 (1.17, 1.53) Lifetime occupational interview
       Moon et al., 2024 [50] Dose-response meta (10 cohorts) Multiple Pooled Cumulative years RR 1.13 (1.04, 1.23) at 30 yr Baseline questionnaire
       Moon et al., 2024 [50] Dose-response meta (11 case-controls) Multiple Pooled Cumulative years OR 1.88 (1.38, 2.57) at 30 yr Lifetime occupational interview
      Cohort studies
       Travis et al., 2016 [40] Prospective cohort (Million Women Study) UK 4,809/522,246 Ever night shift RR 1.00 (0.92, 1.08) Baseline questionnaire
       Akerstedt et al., 2015 [42] Prospective cohort (Swedish Twin Registry) Sweden 463/13,656 Years of night work HR 1.68 (0.98, 2.88) for >20 yr Computer-assisted telephone interview
      Case-control studies
       Lie et al., 2006 [44] Nested case-control Norway 537/2,143 Years of night work OR 2.21 (1.10, 4.45) for ≥30 yr Nurse-registry and census-based work-history reconstruction
       Lie et al., 2011 [45] Nested case-control Norway 699/895 Consecutive night shifts and years OR 1.80 (1.10, 2.80) for ≥5 yr with ≥6 consecutive night shifts Computer-assisted telephone interview of lifetime occupational history
       Cordina-Duverger et al., 2018 [46] Pooled population-based case-control (5 studies) Australia, Canada, France, Germany, Spain 6,093/6,933 Duration×frequency OR 2.55 (1.03, 6.30) for ≥10 yr at ≥3 night/wk among premenopausal women Lifetime occupational interview
       Menegaux et al., 2013 [47] Population-based case-control (CECILE) France 1,232/1,317 Duration and frequency OR 1.35 (1.01, 1.80) for overnight shifts Face-to-face interview of lifetime occupational history
       Rabstein et al., 2013 [48] Case-control (GENICA) Germany 857/892 Duration, ER status OR 4.73 (1.22, 18.36) for ER-negative Lifetime interview
       Hansen et al., 2012 [49] Nested case-control Denmark 141/551 Cumulative no. of night shifts+chronotype OR 3.9 (1.6, 9.5) for morning chronotype with high cumulative exposure (≥884 shifts) Postal structured questionnaire, supplemented by telephone interview
      Validation study
       Vestergaard et al., 2024 [51] Cross-sectional validation Denmark 225 cases/1,800 controls Self-report vs. payroll Cases: sensitivity 86.2%, specificity 82.6% Payroll record linkage
      Studies Study design Country n (cases/controls orPY) Exposure assessment Effect estimate (95% CI) Smoking adjustment
      Meta-analyses
       Honaryar et al., 2019 [10] Meta (cohort studies) Multiple 22 cohort studies JEM/occupational codes Meta-RR 1.29 (1.20, 1.39) Varied
       Honaryar et al., 2019 [10] Meta (case-control studies) Multiple 15 case-control studies Expert/interview Meta-OR 1.87 (1.53, 2.29) Varied
       Honaryar et al., 2019 [10] Meta (adjusted for smoking+asbestos) Multiple 8 case-control studies Mixed Meta-OR 1.17 (1.04, 1.38) Yes
       Moulin, 1997 [54] Meta-analysis Multiple Pooled Mixed Excess lung cancer in welders2 Varied
       Ambroise et al., 2006 [55] Updated meta-analysis Multiple Pooled Mixed Excess lung cancer confirmed2 Varied
       Loomis et al., 2022 [53] WHO/ILO systematic review and meta-analysis Multiple 40 studies; >1,265,512 participants Indirect methods including job title, self-report, and expert assessment Incidence RR 1.48 (1.29, 1.70); Mortality RR 1.27 (1.04, 1.56) Varied
      Cohort studies1
       Sørensen et al., 2007 [22] Prospective cohort Denmark 75 cases/125,762 PY Individual (1,000+breathing-zone measurements; mg/m3×yr) HRR 2.34 (1.03, 5.28) at ≥11 mg/m3×yr Yes
       Sørensen et al., 2007 [22] Prospective cohort Denmark Same cohort Individual SIR 3.69 (1.77, 6.79) for SS employment ≥21 yr -
       Siew et al., 2008 [23] Prospective cohort Finland 30,137 lung cancer cases/1.2 million men JEM (FINJEM) linked to 1970 census occupation RR 1.15 (0.90, 1.46) for highest cumulative welding-fume exposure Yes
      Case-control studies
       Kendzia et al., 2013 [21] Pooled case-control (SYNERGY) Europe, Canada, China, New Zealand 15,483/18,388 Expert-defined categories, face-to-face interviews (81%) OR 1.44 (1.25, 1.67) overall Yes
       Kendzia et al., 2013 [21] Same (never-smokers) Same Never-smokers subgroup Same OR 2.34 (1.31, 4.17) N/A
       Kendzia et al., 2013 [21] Same (never+light smokers, <10 pack-yr) Same Never+light smokers subgroup Same OR 1.96 (1.37, 2.79) N/A
       Kendzia et al., 2013 [21] Same (>25 yr duration) Same >25 yr duration subgroup Same OR 1.77 (1.31, 2.39), p-trend <0.001 Yes
       Vallières et al., 2012 [56] Pooled population-based case-control studies Canada 1,593/1,960 Lifetime occupational interview+expert chemist-hygienist assessment Light smokers: gas-fume OR 2.9 (1.7, 4.8); arc-fume OR 2.3 (1.3, 3.8) Yes
       Pesch et al., 2019 [57] Pooled population-based case-control studies Germany 3,418/3,488 Welding-process exposure matrix linked to job-specific questionnaire High exposure: fumes OR 1.55 (1.17, 2.05); Cr(VI) OR 1.85 (1.35, 2.54); Ni OR 1.60 (1.21, 2.12) Yes
      Table 1. Exposure assessment methods in occupational cancer epidemiology by study design

      JEM, job-exposure matrix; KoWorC–NHIS, Korean Workers’ Cohort–National Health Insurance Service; FINJEM, Finnish job-exposure matrix; OR, odds ratio.

      Bouyer et al. [25]: the JEM variance of log-OR was 2–3 times larger than expert assessment; the power to detect an OR of 2.0 dropped from 0.90 to 0.62.

      Expert assessment introduces classical measurement error, which can bias estimates away from the null when misclassification is differential.

      Table 2. Night shift work and breast cancer: key studies by design

      CI, confidence interval; ER, estrogen receptor; HR, hazard ratio; OR, odds ratio; RR, relative risk.

      Effect estimates are presented as reported in the original publications; Studies are grouped by design to illustrate the systematic divergence in effect magnitude; The validation study by Vestergaard et al. [51] found a differential recall of 5.6 percentage points (CI crossing zero), corresponding to approximately 6% attenuation of the OR.

      Table 3. Welding fumes and lung cancer: key studies by design

      PY, person-years; CI, confidence interval; JEM, job-exposure matrix; WHO, World Health Organization; ILO, International Labour Organization; OR, odds ratio; RR, relative risk; HRR, hazard rate ratio; SIR, standardized incidence ratio; SS, stainless steel welding; Cr(VI), chromium 6; Ni, nickel; N/A, not available.

      Sørensen et al. [22] reported the only prospective study with individual cumulative exposure measurement (breathing-zone air sampling), providing a direct comparison with the JEM-based Siew et al. [23] cohort of similar design.

      Specific pooled estimates not reported here as they predate the Honaryar et al. [10] comprehensive meta-analysis.


      Epidemiol Health : Epidemiology and Health
      TOP