Skip Navigation
Skip to contents

Perspect Integr Med : Perspectives on Integrative Medicine

OPEN ACCESS
SEARCH
Search

Articles

Page Path
HOME > Perspect Integr Med > Volume 5(2); 2026 > Article
Review Article
Operational Definitions of Population, Intervention, Comparison, Outcome for Generating Real-World Evidence Using Korean Health Insurance Claims Sample Data in Integrative Medicine Research: A Methodological Guide
Haein Kimorcid, Seungwon Shin*orcid
Perspectives on Integrative Medicine 2026;5(2):73-84.
DOI: https://doi.org/10.56986/pim.2026.06.001
Published online: June 19, 2026

College of Korean Medicine, Sangji University, Wonju, Republic of Korea

*Corresponding author: Seungwon Shin, College of Korean Medicine, Sangji University, 83 Sangjidae-gil, Wonju 26339, Republic of Korea, Email: ssw.kmd@gmail.com
• Received: December 18, 2025   • Revised: March 8, 2026   • Accepted: April 29, 2026

©2026 Jaseng Medical Foundation

This is an open access article under the CC BY-NC license (http://creativecommons.org/licenses/by-nc/4.0/).

next
  • 914 Views
  • 18 Download
  • Real-world evidence studies using health insurance claims data are utilized in clinical research. This review provides operational definitions of the study population, intervention/exposure, comparison, and outcomes (PICO) in integrative medicine research (conventional and Korean medicine approaches), with case-based examples. Illustrative claims-based studies using Korean national health insurance data were selected, and how PICO elements had been operationally defined in the research were reviewed. Key variables were categorized into general information, diagnosis information, and medical service information, and mapped to Health Insurance Review and Assessment Service-National Patient Sample and National Health Insurance Service-National Sample Cohort tables to describe their roles in constructing each PICO component. Population was primarily defined using diagnosis information, with variations depending on the breadth of diagnosis code inclusion, while medical service utilization information enabled more refined patient selection. Intervention was mainly operationalized using treatment and prescription codes, service types, and treatment intensity. Comparison groups were constructed by confirming no exposure and improving clinical comparability using confounding-control strategies. Outcomes were defined using combinations of diagnosis records, healthcare utilization events, and mortality data. This review provides a practical framework for operationally defining PICO elements in claims-based real-world evidence studies and may improve rigor and reproducibility in future research.
Real-world evidence (RWE) refers to medical evidence generated through the analysis of real-world data (RWD) collected from actual clinical practice rather than from controlled experimental settings. RWD includes electronic health records, patient registries, data from wearable devices or mobile applications, and health insurance claims data [1]. In Korea, claims data provided by the Health Insurance Review and Assessment Service (HIRA) and the National Health Insurance Service (NHIS) cover more than 98% of the total population; accordingly, they have been extensively used in clinical and health services research [2,3]. Claims data enable the evaluation of the effectiveness, safety, and cost-effectiveness of interventions that are difficult to assess through randomized controlled trials and thus serve as a key data source for generating RWE [4].
A key challenge in using RWD is establishing clear operational definitions which specify how concepts are measured and classified [5]. This enhances reproducibility and reduces ambiguity across studies [68]. This is particularly critical in health insurance claims-based research [9] because administrative claims data are generated primarily for reimbursement [10,11], often complicating the direct definition of the study population and variables. Nevertheless, by combining claims-based data elements (e.g., diagnosis codes, procedure/prescription codes, and demographics) researchers can construct meaningful definitions. In this context, the Population, Intervention/Exposure, Comparison, and Outcomes (PICO) framework provides a useful structure for translating a clinical question into claims-based operational definitions [1214].
Numerous studies have investigated the validity of operational definitions for specific diseases in claims-based research [1517]. For example, a study on osteoporotic hip fracture compared nine different algorithms combining diagnosis and procedure codes, and identified the operational definition with the highest accuracy [15]. In a colorectal cancer study, age-standardized incidence rates derived from different definitions were compared to determine the definition most consistent with the actual incidence [16]. Similarly, a review on hepatocellular carcinoma proposed sex-specific criteria for valid case identification [17].
Despite these discussions on disease-specific operational definitions in claims data, there remains a lack of research on comprehensive operationalizing frameworks that encompass the entire study PICO framework. Therefore, this review aimed to provide a methodological guide for interpreting administrative claims data for clinical research purposes and a translation of operational definitions using the PICO framework. Through worked examples, the level of specificity for each PICO element and how to construct these definitions will be illustrated by combining variables from relevant data.
1. Data source and variable categorization
The Health Insurance Review and Assessment Service-National Patient Sample (HIRA-NPS) and National Health Insurance Service-National Sample Cohort (NHIS-NSC) are nationally representative sample databases based on the Korean health insurance system, designed to reflect the entire population of Korea [18]. The HIRA-NPS is a repeated cross-sectional dataset constructed through annual resampling based on sex and 5-year age groups, and includes approximately 2%–3% of all patients who used medical services in Korea each year (about 1.0–1.4 million individuals between 2009 and 2020) [19]. In contrast, the NHIS-NSC is a longitudinal cohort database established by sampling 2% (approximately 1 million individuals) of the population who maintained National Health Insurance or Medical Aid eligibility in 2006, and follows them over a long-term period (2009–2019), with newborn samples added annually to maintain sample size and representativeness [20,21]. Studies using the HIRA-NPS and NHIS-NSC are subject to the data access and approval procedures from the respective data providers, and typically undergo ethics committee review to ensure compliance with relevant ethical and data security requirements [22,23].
Both datasets are structured with multiple tables linked through a common identifier [11] (Table 1). Both datasets include the tables for general specification, medical service specification, diagnosis statement, drug prescription, and healthcare institution [24]. The general specification table contains sociodemographic variables such as sex, age, and insurance qualification, as well as core information including primary and secondary diagnoses, and medical costs. The medical service specification table provides detailed information on medical services delivered during inpatient and outpatient care. The diagnosis statement table records all diagnoses as individual entries, and the drug prescription table contains information on medications prescribed [25]. The healthcare institution table provides institutional characteristics, including facility type, bed capacity, and medical workforce [26]. In contrast to the HIRA-NPS, the NHIS-NSC also provides the health insurance qualification table, the birth and death (BND) and the health screening tables [27], enabling the analysis of patients’ socioeconomic status, mortality, and overall health status.
In this review, variables were selected from each table of the NHIS-NSC and HIRA-NPS based on the official codebooks and manuals provided by the HIRA and NHIS, and were classified into 3 categories: (1) general information, (2) diagnosis information, and (3) medical service information. Variables from the 2 databases were compared to identify common and dataset-specific variables, and consistency in meaning, coding, and data structure was verified even when variable names differed. The extracted variables were classified based on semantic and functional similarity.
2. Literature identification and case selection
This article is a methodological guide and is not intended to be a systematic review of all claims-based studies. Examples were selected for their illustrative usefulness in demonstrating specific methodological elements. Articles published between April 2012 (when the HIRA-NPS became available) and November 2025 were retrieved using major domestic and international academic databases (PubMed, KoreaMed, and the Korean Medical Database). Search terms combined: (1) claims-data sources (NHIS, HIRA); (2) claims-related terms (claims, insurance claim, administrative data); and (3) topic terms (integrative, traditional, herbal, acupuncture, Korean medicine). When needed, manual searches were conducted using condition- and intervention-specific terms.
Titles and abstracts were manually screened to remove duplicates and assess relevance to the study objective of the review. Studies were included if they (1) were observational studies using the NHIS and/or HIRA claims data, (2) provided explicit claims-based operational definitions for at least 1 PICO element, and (3) addressed a medical topic. From these eligible studies, 15 were judged to have high illustrative usefulness as case-based examples and were selected for inclusion. When multiple studies were similarly illustrative, studies on integrative medicine (combining Western conventional and Korean medicine) were prioritized over those limited to either conventional Western medicine or complementary medicine, with preference given to more recently published articles, while seeking to cover diverse conditions and research contexts.
3. PICO elements in previous claims-based research
To clarify how and to what extent each PICO element can be operationally defined from claims data for clinical research, PICO definitions from previous claims-based studies were reviewed and mapped. We extracted and categorized the operational definitions of the study population, the intervention/exposure, the comparison, and the outcomes. For each component, the underlying data sources (e.g., diagnosis codes, procedure/prescription codes, demographics) were identified and mapped to the corresponding tables and variables in the HIRA-NPS and NHIS-NSC. If specific tables and variables were not explicitly reported, based on the described code types and data sources, an inference was made.
1. Operational definitions of study population
The study population is the element defined by the most comprehensive combination of the 3 domains: general information, diagnosis information, and medical service information. Each domain reflects different aspects of the patient: (1) general information represents sociodemographic characteristics; (2) diagnosis information reflects disease characteristics; and (3) medical service information represents healthcare utilization patterns. Depending on the research objective, these domains can be combined to operationally define the study population.
Sociodemographic information serves as the fundamental criteria for defining baseline patient characteristics. In addition to basic demographic variables (such as age, sex, and residential area), the insurance type and premium level reflect patients’ socioeconomic status and healthcare accessibility [28]. General information, when combined with diagnosis and medical service information, allows a more precise definition of the study population. For example, a study analyzing the long-term effects of acupuncture in patients with idiopathic Parkinson’s disease (IPD) included adults aged 19 years or older who were newly diagnosed with IPD between 2012 and 2016, and had no registered disability at the time of initial diagnosis [29]. General information is available in T200 for the HIRA-NPS and in the health insurance qualification table, and the BND table for the NHIS-NSC; information on residential area, disability status, and death is provided only in the NHIS-NSC (Table 2).
Diagnosis information is a core component in defining the study population. Patients are typically included based on the primary diagnosis. For example, a study analyzing claims patterns and healthcare utilization among patients with allergic rhinitis using conventional Western medicine and Korean medicine (KM) defined the study population as patients diagnosed with allergic rhinitis as the primary diagnosis [30]. However, relying solely on the primary diagnosis may substantially limit the number of eligible patients for certain conditions. Therefore, target populations can be defined by additionally including selected secondary diagnoses. For example, a study evaluating the effectiveness and safety of herbal medicine treatment for facial nerve palsy identified patients based on both the primary and the 1st secondary diagnosis codes [31]. An asthma study showed that the healthcare utilization rate in the total population was only 2.9% based on the primary diagnosis alone, but increased to 5.7% when the primary and 1st secondary diagnoses were included, and this rate further increased to 11.4% when all secondary diagnoses were included [32].
For some conditions, the disease of interest may be recorded more frequently as a secondary diagnosis rather than as the primary one. For example, a study reported that 65.7% of hospitalized patients with COVID-19 had recorded COVID-19 codes as a secondary diagnosis [33]. It has been reported that more than 15 secondary diagnosis fields are required to achieve sufficient capture of comorbidities and complications, and that such a range is necessary for adequate characterization of clinical outcomes [34]. Therefore, researchers should define diagnostic criteria by appropriately combining primary and secondary diagnoses according to the research objective and analytic scope. While T200 (T20) provides the primary diagnosis and a limited number of secondary diagnoses, T400 (T40) comprehensively contains all diagnosis codes assigned to each patient, enabling a complete ascertainment of diagnostic information (Table 2).
Depending on the research objective, specific disease groups may need to be excluded from the study population, which can also be implemented using diagnosis codes. For example, a study evaluating the risk of chronic kidney disease and diabetes in patients with colorectal cancer minimized confounding by excluding patients diagnosed with inflammatory bowel disease and familial or hereditary polyposis based on diagnosis codes [35].
Medical service information also plays a complementary role when disease identification based solely on diagnosis information is insufficient. The most fundamental and comprehensive variable within medical service information is the service type. It allows identification of patient use of inpatient or outpatient services of a conventional medical institution, dental clinic, KM institution, psychiatric institution, or public health center. At a more specific level, the detailed information for the procedures or prescriptions can also be used to define the study population. For example, a study analyzing prescription patterns of “insured” herbal preparations from 2010 to 2019 defined the study population as individuals who had been prescribed herbal medicine at least once at KM institutions during the study period [36]. Another study evaluating the association between acupuncture use and prognosis in patients with ischemic stroke ensured diagnostic validity by incorporating hospitalization status after onset, the performance of neuroimaging (computed tomography / magnetic resonance imaging) at onset, and history of antithrombotic therapy before and after the event [37]. In a gout study, the study population was identified using not only the gout diagnosis code but also prescriptions for allopurinol or febuxostat, and healthcare utilization records within 1 year after diagnosis [38]. Furthermore, medical service information can be used to exclude patients with a specific healthcare utilization history. For example, in a study evaluating conventional synthetic disease-modifying antirheumatic drug treatment prior to biological therapy in patients with rheumatoid arthritis, dialysis patients were excluded in advance to minimize potential confounding [39].
Previous studies have reported that the positive predictive value of disease identification improves when medical utilization information is used in addition to diagnosis codes [40,41]. General medical service information (e.g., service type, service dates) can be identified from T200 (T20), whereas detailed information on the type, dosage, and frequency of specific procedures and prescriptions must be obtained from T300 (T30). While T200 (T20) contains summary claim information, T300 (T30) contains detailed treatment records, and the 2 tables are linked in a one-to-many relationship which must be considered in the analysis. In addition, in-hospital prescriptions are recorded in T300 (T30), whereas out-of-hospital prescriptions must be separately identified from T530 (T60) (Table 2).
Taken together, the study population can be defined through the combined use of general information, diagnosis information, and medical service information. For example, a study evaluating the impact of continuity of care and medication adherence in patients with angina pectoris identified newly diagnosed patients aged 30 years or older. Claims data were used from 2002 to 2019 to sequentially exclude patients with insufficient repeated prescription records, prior diagnoses of angina or related complications before baseline, a history of relevant procedures or surgeries, and Emergency Department visits or hospitalizations within 2 years before the index date [42]. This process represents a case in which a stable and valid study population was established by simultaneously integrating general information (age), diagnosis information (angina and complication codes), and medical service information (treatment records and service types).
2. Operational definitions of intervention/exposure
In KM research, the simplest approach to identify whether a patient received the KM service is to use the type of the claim (KM inpatient or outpatient care) in T20 (T200). For example, a study examining the association between KM treatment and the risk of Parkinson’s disease in patients with inflammatory bowel disease defined patients with a history of KM treatment as the integrative treatment group using classification codes [43].
In studies requiring a more specific definition of the intervention, the intervention group is defined by selecting specific procedures or prescriptions. KM treatment types identifiable in the T40 (T400) table can be categorized into acupuncture, moxibustion, cupping, psychotherapy, and “insured” herbal preparations [43]. Each treatment category corresponds to a predefined code range. For example, the codes corresponding to Meridian Acupuncture are 40011 and 40012, while those for Microsystem Acupuncture include 40120–0129 and 40131–40134 (Supplement). Code ranges were verified as of January 2025 and may change with updates to reimbursement schedules. In a study analyzing the long-term effects of acupuncture in patients with IPD, acupuncture treatment was defined using codes corresponding to general acupuncture and electroacupuncture [29].
Additional criteria based on treatment frequency or dosage are also applied. For instance, a study defined the acupuncture group as patients who received acupuncture at least 6 times within 1 year, while excluding those who received acupuncture 1 to 5 times [29]. Information on the dosage and frequency of procedures and prescriptions is available in T300 (T30), whereas out-of-hospital prescriptions must be identified from T530 (T60; Table 2).
3. Operational definitions of comparison
The comparison group is established to evaluate differences relative to the intervention group and can be identified to classify patients who did not receive the specific intervention(s). The basic approach to constructing a comparison group is to simply assign patients according to whether the intervention was performed or not. This is appropriate for studies that aim to describe patient characteristics or healthcare utilization patterns. For example, a study analyzing determinants of KM utilization after spinal surgery compared patients with and without KM use in a regression model [44]. However, in most observational studies, evaluating intervention effects such a simple comparison is insufficient because the intervention and comparison groups are highly likely to differ in clinical characteristics, socioeconomic background, and comorbidities. These differences introduce confounding, making it difficult to interpret outcome differences between groups as true intervention effects [45]. Therefore, it is essential to construct a clinically comparable comparison group by considering general information (e.g., age, sex, and insurance status) and diagnosis information [Korean Standard Classification of Diseases (KCD) codes]. This method of defining the comparison group is a key determinant of the internal validity and real-world relevance of the study [46].
Statistical methods used to minimize confounding include regression adjustment, propensity score matching (PSM), propensity score stratification or adjustment, and inverse probability of treatment weighting (IPTW) [47]. PSM is a statistical procedure that constructs pairs of participants with similar propensity scores between the intervention and comparison groups, allowing outcome comparisons in a manner analogous to randomized controlled trials; matching is typically performed at a 1:1 or 1:N ratio [48]. For example, in a study analyzing the association between acupuncture therapy and the incidence of lumbar surgery in patients with low back pain, 1:1 PSM was performed based on age, sex, income level, and the Charlson comorbidity index was used to adjust for disease severity between the study groups [49]. IPTW assigns each participant a weight equal to the inverse of the probability of receiving the actual treatment (1/propensity score), thereby balancing baseline characteristics between the intervention and comparison groups. While PSM reduces the sample size by selecting only participants most similar to the intervention group, IPTW retains the full sample and achieves balance through weighting [48]. For instance, in a study evaluating whether immunocompromised patients with COVID-19 had an increased mortality risk, propensity scores were calculated using multivariable logistic regression, and IPTW was then applied to balance baseline characteristics between groups [50]. Because confounding adjustment requires the simultaneous consideration of general information, diagnosis information, and medical service information, constructing comparison groups through the integrated use of T200 (T20), T300 (T30), T400 (T40), and T530 (T60) is essential (Table 2).
To ensure the internal validity of claims-based comparisons, a structured approach covering design, assumptions, limitations, and diagnostics is needed beyond the simple selection of a weighting or matching technique. Firstly, at the design stage, consistent index date alignment is important to establish the correct temporality between covariate assessment, exposure, and outcomes. In addition, designs such as an active-comparator and a new-user design can improve comparability by aligning treatment indication and baseline disease severity between groups [51]. Secondly, propensity score-based methods rely on key causal identification assumptions, including conditional exchangeability (i.e., no unmeasured confounding), consistency, and positivity [52]. Because these methods can only address measured confounders, they remain vulnerable to residual confounding from unmeasured factors. Careful covariate specification is therefore essential, although unmeasured confounding cannot be eliminated completely [48,52]. Finally, after matching or weighting, covariate balance should be assessed using standardized mean difference (SMD), with an absolute SMD < 0.1 often used as a practical benchmark for negligible imbalance between groups [48]. Sensitivity analyses can further strengthen causal interpretation by evaluating the robustness of findings to unmeasured confounding; examples include negative control analyses [53] and E-value-based assessments [54]. Moreover, if exposure status can change over time, the intervention and comparison groups do not necessarily have to be defined at a fixed time point. In studies where patients’ clinical status or treatment exposure may change over the study period, time-dependent analysis should be considered [55]. Time-dependent analysis incorporates changes in individual exposure status over time into the statistical model and helps prevent immortal time bias arising from differences in intervention timing [56,57]. In more complex settings, involving time-varying exposure and time-dependent confounding, approaches such as marginal structural models estimated using inverse probability weighting may also be considered [58]. For example, in a study evaluating the effects of acupuncture in patients with ischemic stroke, the period before acupuncture initiation was treated as unexposed person-time to reflect changes in exposure status over time, and to minimize immortal time bias [37]. Such time-dependent analyses are feasible because daily service dates are recorded in T200 (T20; Table 2).
4. Operational definitions of outcomes
Health insurance claims data generally do not include clinical outcomes commonly used in clinical trials such as symptom scores, vital signs, laboratory results, or physical examination findings. Due to these limitations, claims-based studies inevitably rely on indirect or proxy outcome measures such as post-intervention diagnosis records, patterns of medical service, and/or mortality [59]. Accordingly, in claims-based research, outcomes are operationally defined by combining diagnosis information, medical service information, and mortality data, each of which captures different types of clinical events.
Diagnosis information is used to identify the occurrence of new diagnoses including disease recurrence, complications, and adverse events. The criteria for defining recurrence or adverse events must be clearly specified according to the research objective. For example, in a study analyzing the association between acupuncture treatment and prognosis in patients with ischemic stroke, major complications were defined to include pneumonia, urinary tract infection, pressure ulcers, gastrointestinal bleeding, and femoral fractures [37]. Diagnosis-based outcomes can be defined using different ranges of diagnostic fields such as primary diagnosis, 1st secondary diagnosis, or additional secondary diagnoses, with the choice between T200 (T20) and T400 (T40) depending on the breadth of diagnostic information required (Table 2).
Medical service-based outcomes can also be used to define the outcomes: readmission, emergency visits, procedures or prescriptions, medical costs, or healthcare utilization frequency. Readmission and emergency visits may serve as indicators of clinical deterioration or changes in patient status. Rather than defining readmission based solely on hospitalization records, more precise definitions can be constructed by combining them with specific medical procedures. For example, in a study of patients with ischemic stroke, readmission was defined as hospitalization for at least 1 day with the same primary diagnosis after the index event, which was accompanied by computed tomography or magnetic resonance imaging examinations [37]. Records of specific procedures or prescriptions can also be used to indirectly assess disease progression. For instance, in a study evaluating the long-term effects of acupuncture in patients with IPD, the occurrence of a 1st deep brain stimulation surgery was included as a secondary outcome to assess symptom worsening [29]. In addition, medical costs and the number of healthcare visits can be used to estimate patients’ healthcare burden and overall health expenditure. For example, changes in medical costs and the frequency of healthcare visits during a defined follow-up period can be analyzed to evaluate the burden-reducing effects of a given treatment. Medical service information can be obtained from T200 (T20), T300 (T30), and T530 (T60; Table 2).
Mortality outcomes can be classified as either all-cause mortality or cause-specific mortality. For example, in a study of patients with prediabetes and diabetes, both all-cause mortality and cancer-specific mortality were defined as outcomes to investigate the disease-specific risk of cancer-related death [60]. Information on the date and cause of death is available from the BND table in the NHIS-NSC, while death information is not provided in the HIRA-NPS (Table 2). The summarized approaches for the operational definitions of the PICO elements are described in Table 3.
5. Worked example
To demonstrate end-to-end PICO operationalization using Korean claims data, a worked example using the NHIS-NSC is presented. Patients with low back pain were identified using the KCD codes, and a 12-month look-back period was applied to exclude prior lumbar surgery. Acupuncture exposure is then defined as receiving ≥ 6 acupuncture sessions within 8 weeks after the index date. An 8-week landmark design is applied to classify exposure, and follow-up begins at the landmark among individuals who remain event-free up to that point. As an active comparator, new initiators of physical therapy/rehabilitation are defined using the same look-back period. Confounding is addressed using PSM based on baseline covariates, measured during the look-back period, with covariate balance assessed using SMDs. The outcome is lumbar surgery within 365 days after the landmark (time zero). Follow-up was censored at the earliest of lumbar surgery, end of data availability, or Day 365. Detailed operational definitions and data domains used for each PICO element are summarized in Table 4.
In order to provide practical strategies for operational definitions of PICO in integrative medicine (conventional Western medicine and Korean medicine) research, this review examined how the PICO elements are operationally defined in actual research, and examples were reviewed. In this way, how each PICO element can be implemented in actual study designs through the integrated use of diverse claims-based information can be demonstrated.
The definition of the study population is closely linked to the process of identifying disease groups primarily using diagnosis information, which can vary substantially depending on how the scope of primary and secondary diagnoses is used. In addition, integrating medical service information enables a more precise definition of the study population. The intervention can be specifically defined based on medical service information, including procedure and prescription codes, types of medical services, and the dose of the procedures. The comparison group requires a strategy in which the absence of intervention exposure is 1st confirmed, followed by the construction of a clinically comparable control group using general information, diagnosis information, and medical service information. Outcomes can be defined by identifying clinical events such as recurrence, complications, and adverse events through diagnosis information. By assessing prognosis using medical utilization records and confirming mortality using death-related information, the investigators can evaluate more comprehensive interventional results.
A particularly important methodological issue arises when the same healthcare utilization variables are used to define both population and intervention. This overlap blurs the distinction between eligibility and exposure, potentially introducing selection bias and immortal time bias. Immortal time bias typically occurs when a time-dependent exposure is inappropriately incorporated into the population definition, such as applying cumulative exposure criteria (e.g., ≥ 6 acupuncture sessions) to identify the target population. As a general principle to prevent this, the study population should be defined as the eligible disease group at study entry using information available before the index date, whereas the intervention/exposure should be defined using exposure measured at or after the index date. Furthermore, to mitigate immortal time bias when exposure varies over time, exposure can be modeled as a time-dependent variable, or landmark analysis may be considered [56]. Furthermore, confounding by indication represents another critical challenge in comparative claims-based research. This bias arises when disease severity or clinical status influencing treatment selection also affects outcomes, resulting in different baseline risks between comparison groups. This problem may be mitigated by restricting the study population to patients with the specific treatment indication; when the exact indication cannot be measured accurately due to data limitations, an active comparator design may be an effective practical alternative [61].
Separate from these design-related issues, claims databases also have intrinsic limitations. Firstly, miscoding may occur in diagnosis codes, procedure/prescription codes, and service dates; this can be partly mitigated by excluding implausible values and using more specific operational definitions, such as repeated diagnoses (e.g., ≥ 2 outpatient visits) or definitions that combine diagnosis and medical service information. Second, non-covered services are not fully captured in claims databases, which is particularly relevant in KM, where the proportion of non-covered expenses is relatively high [62]. Thirdly, the HIRA-NPS and NHIS-NSC differ in sampling structure, which has implications for representativeness and study design. The HIRA-NPS is an annually resampled database more appropriate for cross-sectional analyses or short-term follow-up studies (less than 1 year), whereas the NHIS-NSC is a longitudinal cohort database suitable for analyses requiring long-term follow-up [63]. These limitations cannot be fully corrected analytically and should therefore be addressed through cautious interpretation, appropriate dataset selection, and the use of validated and specific operational definitions whenever possible.
The PICO framework is more appropriate because it enables researchers to systematically map core elements of clinical research onto claims-based data variables. Alternative frameworks may be more intuitive for specific research questions, although they can be regarded as variants of PICO [6468]. Therefore, PICO was adopted as the primary framework because it provides a broadly applicable and conceptually consistent basis for operationalizing claims-based clinical research.
A key strength of this review lies in its restructuring of the complex claims data architecture into core information domains from a researcher’s perspective and in clearly demonstrating the roles of these domains in defining each PICO element. For researchers who are not familiar with health insurance claims data, it is often difficult to determine which variables should be selected and how they should be applied among the enormous number of available variables. By intuitively organizing how general information, diagnosis information, and medical service information function in the definitions of population, intervention/exposure, comparison, and outcomes, this study enhances the understanding of claims data structure. In addition, by presenting real-world research examples, this study demonstrates how PICO operational definitions are constructed and applied in research, thereby providing concrete and actionable guidance that researchers can directly apply to their own studies.
Nevertheless, several limitations of this study should be acknowledged. Firstly, although the HIRA and NHIS provide customized datasets for research purposes [11], this study focused exclusively on national sample datasets (HIRA-NPS and NHIS-NSC). Secondly, although operational definitions in claims-based research can vary substantially according to disease characteristics, study design, and analytic objectives, this study addressed only selected examples for illustrative purposes, and therefore does not encompass all possible approaches to operational definition. Thirdly, although some studies apply modified frameworks rather than the conventional PICO structure depending on their design, this study adopted the PICO perspective for conceptual consistency.
In recent years, effort to promote RWE generation through pseudonymized linkage between claims data and external datasets has been expanding. The HIRA and NHIS provide services that enable the safe integration and provision of datasets from different organizations. For example, studies linking claims data with mortality data of Statistics Korea, the Korean Labor and Income Panel Study data of the Korea Labor Institute, and clinical data of specific KM hospitals have been reported [6971]. In integrative medicine research, data linkage can compensate for the limited availability of KM-related information in claims data and enable more refined analyses of the effects of integrative interventions as well as the relationships among diverse social and clinical factors. Although this study focused on the claims data, future studies are warranted to explore how operational definitions can be expanded in linked-data environments, where a broader range of variables becomes available.
This review aimed to provide a practical approach for the operational definition of PICO elements in RWE research using Korean health insurance claims sample data. Rather than conducting a systematic review, illustrative claims-based studies were selected and reviewed to determine how individual PICO elements have been defined in prior research. Based on this review, the operational definitions of population, intervention/exposure, comparison, and outcomes were mapped onto the core information domains available in the HIRA-NPS and NHIS-NSC, including general information, diagnosis information, and medical service information. By providing a flexible and broadly applicable framework for constructing PICO definitions, this review may contribute to the enhancement of methodological rigor and the level of evidence in future claims-based RWE studies in integrative medicine.
Supplementary materials are available at doi: https://doi.org/10.56986/pim.2026.06.001.

Author Contributions

Conceptualization: SS. Methodology: SS and HK. Formal investigation: HK. Data analysis: HK. Writing original draft: HK. Writing - review and editing: SS and HK.

Conflicts of Interest

The authors have no conflicts of interest to declare.

Author Use of AI Tools Statement

During the preparation of this manuscript, the authors used ChatGPT 5.1 (OpenAI, USA) for improving language clarity and grammar. All content was subsequently reviewed and revised by the authors, who accept full responsibility for the final version of the work.

Funding

This research was supported by a grant of the Korea Health Technology R&D Project through the Korea Health Industry Development Institute (KHIDI), funded by the Ministry of Health & Welfare, Republic of Korea (grant no.: RS-2024-00441486).

Ethics Statement

This research did not involve any human or animal experiments.

All relevant data are included in this manuscript.
pim-2026-06-001f1.jpg
Table 1
Structure of the HIRA-NPS and NHIS-NSC Databases
Category HIRA-NPS NHIS-NSC
Health insurance qualification - BNC
Birth & death - BND
Medical service claims General specification T200 T20
Medical service specification T300 T30
Diagnosis statements T400 T40
Drug prescriptions T530 T60
Healthcare institutions YKIHO INST
National health screening program - G1E

The table labels of HIRA-NPS & NHIS-NSC are indicated for each category.

HIRA = Health Insurance Review and Assessment Service (Republic of Korea); NHIS = National Health Insurance Service (Republic of Korea); NPS = National Patient Sample; NSC = National Sample Cohort.

Table 2
Key Variables of the HIRA-NPS and the NHIS-NSC Databases
Category Key variable HIRA-NPS NHIS-NSC
General information Gender T200 BNC
Age (y) T200 (5-y group) BND (year of birth)
Residential area - BNC
Insurance qualification T200 (Type) BNC (Type, premium decile)
Disability (type & severity) - BNC
Death (cause) - BND
Diagnosis information Primary diagnosis (KCD codes) T200, T400 T20, T40
1st secondary diagnosis (KCD codes) T200, T400 T20, T40
All additional secondary diagnoses (KCD codes) T400 T20 (up to 4 secondary diagnoses), T40
Medical service information Service type T200 T20
Medical department T200, T400 T20, T40
Service dates T200 T20
Codes (Fee schedule, insurance, main ingredients, medical device) T300 T30
Dosage & frequency T300, T530 T30, T60
Surgery T200 T20
Medical cost T200 T20
Main ingredients T300, T530 T60
ATC codes T300, T530 -
Drug classification - T30, T60

ATC = anatomical therapeutic chemical; BMI = body mass index; DDM = doctor of dental medicine; DKM = doctor of Korean medicine; DM = doctor of medicine; DRG = diagnosis-related group; HIRA = Health Insurance Review and Assessment Service (Republic of Korea); KCD = Korean Standard Classification of Diseases; NHIS = National Health Insurance Service (Republic of Korea); NPS = National Patient Sample; NSC = National Sample Cohort.

Table 3
Approach to Define PICO Elements Operationally in Claims-Based Research
PICO element Information domain Examples Table: key variables
P (Population) General information Defining eligibility criteria with sociodemographic characteristics BNC: sex, residential area, insurance qualification, disability
BND: age, death
Diagnosis information Defining the target diseases with diagnosis codes
Refining eligibility criteria with prior/comorbid history
T200 (T20): main diagnosis, 1st secondary diagnosis
T400 (T40): all additional secondary diagnosis
Medical service information Defining the study population with procedure/prescription history
Defining exclusion criteria with prior healthcare utilization
T200 (T20): service dates, service type
T300 (T30): procedure & prescription code
I (Intervention or exposure) General information Defining the intervention with healthcare service type T200 (T20): service type
Medical service information Defining the intervention with procedure/prescription records T200 (T20): service dates
T300 (T30): procedure & prescription code, dosage & frequency
T530 (T60): dosage & frequency
C (Comparison) General information Defining sociodemographic covariates for constructing comparable groups BNC: sex, residential area, insurance qualification
BND: age
Diagnosis information Defining clinical covariates for constructing comparable groups T200 (T20): main diagnosis, 1st secondary diagnosis
T400 (T40): all additional secondary diagnosis
Medical service information Defining non-exposed group with the absence of intervention of interest
Defining healthcare-utilization covariates for comparison group construction
T200 (T20): service dates
T300 (T30): procedure & prescription code, dosage & frequency
T530 (T60): dosage & frequency
O (Outcome) General information Identifying mortality status & cause of death BND: death
Diagnosis information Identifying recurrence, complications, or adverse events T200 (T20): main diagnosis, 1st secondary diagnosis
T400 (T40): all additional secondary diagnosis
Medical service information Identifying rehospitalization, emergency visits, procedures, prescriptions, or healthcare utilization measures T200 (T20): service dates, medical cost
T300 (T30): procedure & prescription code, dosage & frequency
T530 (T60): dosage & frequency
Table 4
Worked Example: Claims-Based Operationalization of PICO Elements in a Low Back Pain Cohort Using the NHIS-NSC Database
PICO element Operational definition Table: key variables Notes
P (population) Adults (age at cohort entry, ≥ 18 y) with low back pain identified using KCD codes (main or 1st secondary diagnosis); ≥ 2 outpatient claims or ≥ 1 inpatient claim T20: service dates
T40: main diagnosis, 1st secondary diagnosis
BND: age
A 12-mo look-back period was used to exclude prior lumbar surgery
I (Intervention or exposure) Acupuncture treatment defined using procedure/treatment codes; exposure defined as ≥ 6 acupuncture sessions within 8 weeks after index date T20: service dates
T30: procedure & prescription code
Index date was defined as 1st observed acupuncture service after a 12-mo acupuncture-free look-back
C (comparison) New initiators of physical therapy/rehabilitation among patients with low back pain; confounding control via PSM using baseline covariates (age, sex, insurance qualification, CCI) T20: service dates, main diagnosis, 1st secondary diagnosis
T30: procedure & prescription code
T40: all additional secondary diagnosis
BNC: sex, residential area, insurance qualification
BND: age
Active-comparator new-user design with same look-back period; Covariate balance assessed using SMDs
O (outcome) Lumbar surgery within 365 days after landmark time T20: service dates
T30: procedure & prescription code
Follow-up began at 8-week landmark among event-free individuals (up to that point)

CCI = Charlson comorbidity index; KCD = Korean Standard Classification of Diseases; NHIS = National Health Insurance Service (Republic of Korea); NSC = National Sample Cohort; PSM = propensity score matching; SMD = standardized mean difference.

  • [1] Park S. [Thesis] Status of awareness and use of Real-World Data (RWD) and Real-World Evidence (RWE) in the domestic drug approval process: Domestic survey study. Seoul (Korea), Sungkyunkwan University. 2024.
  • [2] Kwon S. Thirty years of national health insurance in South Korea: lessons for achieving universal health care coverage. Health Policy Plan 2009;24(1):63−71.ArticlePubMedPMC
  • [3] Sungkyunkwan University Research & Business Foundation. Research for approval system development using real world data (RWD). Seoul (Korea), Sungkyunkwan University Research & Business Foundation, 2019.
  • [4] Kim DS, Byun JH, Kim JH, Kim SY, Lee EJ, Kim SH, et al. Development of an RWE-based platform for reimbursement management: a claims data analysis. Wonju (Korea), Health Insurance Review & Assessment Service, 2020, Report no.: G00F8Q-2020-108 [in Korean].
  • [5] Muhl C, Mulligan K, Bayoumi I, Ashcroft R, Godfrey C. Establishing internationally accepted conceptual and operational definitions of social prescribing through expert consensus: a Delphi study. BMJ Open 2023;13(7):e070184. ArticlePubMedPMC
  • [6] Hunter J, Harnett JE, Chan WJJ, Pirotta M. What is integrative medicine? Establishing the decision criteria for an operational definition of integrative medicine for general practice health services research in Australia. Integr Med Res 2023;12(4):100995. ArticlePubMedPMC
  • [7] Dzieciatko-Szendrei B, Pantić N, Joksimović S, Gašević D, Viry G. Systematic review of the operational definitions and indicators of teacher communities. Educ Res Rev 2024;45:100640. Article
  • [8] Xiong Y, Yang J, Wong A, Chan SSW, Li HHS, Tam LHP, et al. Operational definitions improve reliability of the age-related white matter changes scale. Eur J Neurol 2011;18(5):744−9.ArticlePubMed
  • [9] U.S. Food and Drug Administration [Internet]. Real-world data: assessing electronic health records and medical claims data to support regulatory decision-making for drug and biological products: 2024 [cited 2025 Nov 14]. Available from: https://www.fda.gov/regulatory-information/search-fda-guidance-documents/real-world-data-assessing-electronic-health-records-and-medical-claims-data-support-regulatory.
  • [10] Lawson EH, Louie R, Zingmond DS, Brook RH, Hall BL, Han L, et al. A comparison of clinical registry versus administrative claims data for reporting of 30-day surgical complications. Ann Surg 2012;256(6):973−81.ArticlePubMed
  • [11] Park I. How to use health insurance data effectively for healthcare research. J Health Inform Stat 2022;47(Suppl 2):S31−9.ArticlePDF
  • [12] Schlosser RW, Koul R, Costello J. Asking well-built questions for evidence-based practice in augmentative and alternative communication. J Commun Disord 2007;40(3):225−38.ArticlePubMed
  • [13] Richardson WS, Wilson MC, Nishikawa J, Hayward RS. The well-built clinical question: a key to evidence-based decisions. ACP J Club 1995;123(3):A12−3.Article
  • [14] Hosseini M-S, Jahanshahlou F, Akbarzadeh MA, Zarei M, Vaez-Gharamaleki Y. Formulating research questions for evidence-based studies. J Med Surg Pub Health 2024;2:100046. Article
  • [15] Lee Y, Yoo J, Kim T, Ha Y, Koo K, Choi H, et al. Validation of operational definition to identify patients with osteoporotic hip fractures in administrative claims data. Healthcare (Basel) 2022;10(9):1724. ArticlePubMedPMC
  • [16] Park H, Kim YR, Pyun Y, Joo H, Shin A. Operational definitions of colorectal cancer in the Korean national health insurance database. J Prev Med Public Health 2023;56(4):312−8.ArticlePubMedPMCPDF
  • [17] Kim Y, Back J, Seo S, Park H, Cho S, Shin A, et al. Operational definition of liver cancer in studies using data from the national health insurance service: a systematic review. J Cancer Prev 2023;28(2):47−52.ArticlePubMedPMC
  • [18] Kim HK, Song SO, Noh J, Jeong I, Lee B. Data configuration and publication trends for the Korean national health insurance and health insurance review & assessment database. Diabetes Metab J 2020;44(5):671−8.ArticlePubMedPMCPDF
  • [19] Health Insurance Review & Assessment Service. Guideline for using patient sample dataset. Wonju (Korea), Health Insurance Review & Assessment Service, 2022.
  • [20] Chung H, Kim SY, Kim HS. Clinical research from a health insurance database: practice and perspective. Korean J Med 2019;94(6):463−70. [in Korean].ArticlePDF
  • [21] National Health Insurance Service [Internet]. Sample cohort DB 2.2 user manual: Wonju (Korea); National Health Insurance Service: 2022 [cited 2025 Mar 9]. Available from: https://nhiss.nhis.or.kr/lp/z/z/999/lpzzcms.do?cntsKeyVl=D [in Korean].
  • [22] Lim SJ, Jang SI. Leveraging national health insurance service data for public health research in Korea: structure, applications, and future directions. J Korean Med Sci 2025;40(8):e111. ArticlePubMedPMCPDF
  • [23] Lee H, Lim Y, Kim D, Park K, Lee YJ, Ha I, et al. Comparative analysis of infertility healthcare utilization before and after insurance coverage of assisted reproductive technology: a cross-sectional study using National Patient Sample data. PLoS One 2023;18(11):e0294903. ArticlePubMedPMC
  • [24] Yoon CY, Ahn JJ, Lee G, Choi Y, Kim L, Ha DY, et al. Developing the new national patient sample and evaluating representations. Health Insurance Rev Assess Serv Res 2021;1(2):166−78. [in Korean].Article
  • [25] Kim L, Kim J-A, Kim S. A guide for the utilization of health insurance review and assessment service national patient samples. Epidemiol Health 2014;36:e2014008. ArticlePubMedPMC
  • [26] Health Insurance Review & Assessment Service. Patient sample dataset variable description manual. Wonju (Korea), Health Insurance Review & Assessment Service, 2021.
  • [27] National Health Insurance Service. Sample cohort DB 2.2 layout. Wonju (Korea), National Health Insurance Service, 2022.
  • [28] Park D, Lee SY, Jeong E, Hong D, Kim M, Choi JH, et al. The effects of socioeconomic and geographic factors on chronic phase long-term survival after stroke in South Korea. Sci Rep 2022;12(1):4327. ArticlePubMedPMCPDF
  • [29] Hwang YC, Lee J, Kang D, Lee H, Kwon S, Choi S, et al. A nationwide retrospective cohort study of the association between acupuncture exposure and clinical outcomes of idiopathic Parkinson’s disease using health insurance claim data in South Korea. Integr Med Res 2025;14(2):101146. ArticlePubMedPMC
  • [30] Kang C, Kim H, Kim J, Hwang J, Lee D. Outcomes analysis for western medicine and Korean medicine using the propensity score matching in allergic rhinitis: data from the health insurance review and assessment service. J Korean Med Ophthalmol Otolaryngol Dermatol 2021;34(2):53−69. [in Korean]. https://doi.org/10.6114/jkood.2021.34.2.053.
  • [31] Kim SD, Park MY, Cho E, Cha J, Yang C, Kim S. Herbal medicine evaluation for reimbursement-based facial palsy (HERB-FP): a retrospective analysis using Korean health insurance claim data 2020–2024. BMC Complement Med Ther 2025;25(1):307. ArticlePubMedPMCPDF
  • [32] Yoo S, Kim D, Kim Y, Park J, Kim Y, Cho K, et al. Data resource profile: the allergic disease database of the Korean National Health Insurance Service. Epidemiol Health 2021;43:e2021010. ArticlePubMedPMC
  • [33] Milićević MŠ, Rosić N, Vujetić M, Pavlović N, Stevanović A, Jovanović V, et al. Which burden of COVID-19 was greater: main diagnosis or comorbidity? Eur J Public Health 2023;33(Suppl 2):ckad160.978. PMC
  • [34] Drösler SE, Romano PS, Sundararajan V, Burna B, Colin C, Pincus H, et al. How many diagnosis fields are needed to capture safety events in administrative data? Findings and recommendations from the WHO ICD-11 Topic Advisory Group on Quality and Safety. Int J Qual Health Care 2013;26(1):16−25.ArticlePubMedPMC
  • [35] Oh HJ, Lee HA, Moon CM, Ryu DR. The combined impact of chronic kidney disease and diabetes on the risk of colorectal cancer depends on sex: a nationwide population-based study. Yonsei Med J 2020;61(6):506−14.ArticlePubMedPMCPDF
  • [36] Kim H, Lee M, Kim J, Cho J. General prescription pattern of insured herbal preparation in South Korea: a nationwide cohort study. J Korean Med 2024;45(3):14−30. [in Korean].Article
  • [37] Choi SR, Kim ES, Jang BH, Jung B, Ha I. A time-dependent analysis of association between acupuncture utilization and the prognosis of ischemic stroke. Healthcare (Basel) 2022;10(5):756. ArticlePubMedPMC
  • [38] Kim HJ, Ghang B, Kim J, Ahn HS. Regional variations of cardiovascular risk in gout patients: a nationwide cohort study in Korea. J Rheum Dis 2023;30(3):185−97.ArticlePubMedPMC
  • [39] Kim MJ, Park EH, Shin A, Ha Y, Lee YJ, Lee EB, et al. Assessment on treatments with conventional synthetic disease-modifying drugs before initiating biologics in patients with rheumatoid arthritis in Korea: a population-based study. J Rheum Dis 2022;29(2):79−88.ArticlePubMedPMC
  • [40] Papani R, Sharma G, Agarwal A, Callahan SJ, Chan WJ, Kuo Y, et al. Validation of claims-based algorithms for pulmonary arterial hypertension. Pulm Circ 2018;8(2):2045894018759246. ArticlePubMedPMCPDF
  • [41] Shima D, Ii Y, Higa S, Kohro T, Hoshide S, Kono K, et al. Validation of novel identification algorithms for major adverse cardiovascular events in a Japanese claims database. J Clin Hypertens (Greenwich) 2021;23(3):646−55.ArticlePubMedPMCPDF
  • [42] Kim D, Cha J. Continuity of care and medication adherence in patients with angina: a retrospective cohort study using Korea’s National Health Insurance data. BMJ Open 2025;15(6):e098903. ArticlePubMedPMC
  • [43] Noh H, Jang J, Kwon S, Cho S, Jung W, Moon S, et al. The impact of Korean medicine treatment on the incidence of Parkinson’s disease in patients with inflammatory bowel disease: a nationwide population-based cohort study in South Korea. J Clin Med 2020;9(8):2422. ArticlePubMedPMC
  • [44] Kim D, Lee YJ, Jang BH, Park J, Park S, D’Adamo CR, et al. Analysis of factors associated with the use of Korean medicine after spinal surgery using a nationwide database in Korea. Sci Rep 2023;13(1):20177. ArticlePubMedPMCPDF
  • [45] Grobbee DE, Hoes AW. Confounding and indication for treatment in evaluation of drug treatment for hypertension. BMJ 1997;315(7116):1151−4.ArticlePubMedPMC
  • [46] D’Arcy M, Stürmer T, Lund JL. The importance and implications of comparator selection in pharmacoepidemiologic research. Curr Epidemiol Rep 2018;5(3):272−83.ArticlePubMedPMCPDF
  • [47] Kim DW. Statistical methods for baseline adjustment and cohort analysis in Korean national health insurance claims data: a review of PSM, IPTW, and survival analysis with future directions. J Korean Med Sci 2025;40(8):e110. ArticlePubMedPMCPDF
  • [48] Austin PC. An introduction to propensity score methods for reducing the effects of confounding in observational studies. Multivariate Behav Res 2011;46(3):399−424.ArticlePubMedPMC
  • [49] Koh W, Kang K, Lee YJ, Kim M, Shin J, Lee J, et al. Impact of acupuncture treatment on the lumbar surgery rate for low back pain in Korea: a nationwide matched retrospective cohort study. PLoS One 2018;13(6):e0199042. ArticlePubMedPMC
  • [50] Baek MS, Lee MT, Kim WY, Choi JC, Jung S. COVID-19-related outcomes in immunocompromised patients: A nationwide study in Korea. PLoS One 2021;16(10):e0257641. ArticlePubMedPMC
  • [51] Lund JL, Richardson DB, Stürmer T. The active comparator, new user study design in pharmacoepidemiology: historical foundations and contemporary application. Curr Epidemiol Rep 2015;2(4):221−8.ArticlePubMedPMCPDF
  • [52] Shiba K, Kawahara T. Using propensity scores for causal inference: pitfalls and tips. J Epidemiol 2021;31(8):457−63.ArticlePubMedPMC
  • [53] Lipsitch M, Tchetgen Tchetgen E, Cohen T. Negative controls: a tool for detecting confounding and bias in observational studies. Epidemiology 2010;21(3):383−8.PubMedPMC
  • [54] VanDerWeele TJ, Ding P. Sensitivity analysis in observational research: introducing the e-value. Ann Intern Med 2017;167(4):268−74.ArticlePubMedPMCPDF
  • [55] Shintani AK, Girard TD, Eden SK, Arbogast PG, Moons KGM, Ely EW. Immortal time bias in critical care research: application of time-varying Cox regression for observational cohort studies. Crit Care Med 2009;37(11):2939−45.ArticlePubMedPMC
  • [56] Jones M, Fowler R. Immortal time bias in observational studies of time-to-event outcomes. J Crit Care 2016;36:195−9.ArticlePubMed
  • [57] Zabor EC, Assel M. On the need for landmark analysis or time dependent covariates. J Urol 2023;209(6):1060−2.ArticlePubMedPMC
  • [58] Shinozaki T, Suzuki E. Understanding marginal structural models for time-varying exposures: pitfalls and tips. J Epidemiol 2020;30(9):377−89.ArticlePubMedPMC
  • [59] Kyoung DS, Kim HS. Understanding and utilizing claim data from the Korean National Health Insurance Service (NHIS) and Health Insurance Review & Assessment (HIRA) database for research. J Lipid Atheroscler 2021;11(2):103−10.ArticlePubMedPMCPDF
  • [60] Tran TXM, Kim S, Song H, Park B. Increased risk of cancer and cancer-related mortality in middle-aged Korean women with prediabetes and diabetes: a population-based study. Epidemiol Health 2023;45:e2023080. ArticlePubMedPMCPDF
  • [61] Sendor R, Stürmer T. Core concepts in pharmacoepidemiology: confounding by indication and the role of active comparators. Pharmacoepidemiol Drug Saf 2022;31(3):261−9.ArticlePubMedPMCPDF
  • [62] National Health Insurance Service. 2023 medical expenditure survey for national health insurance patients. Wonju (Korea), National Health Insurance Service, 2025.
  • [63] Kim S, Kim M, You S, Jung S. Conducting and reporting a clinical research using Korean healthcare claims database. Korean J Fam Med 2020;41(3):146−52.ArticlePubMedPMCPDF
  • [64] Morgan RL, Whaley P, Thayer KA, Schünemann HJ. Identifying the PECO: a framework for formulating good questions to explore the association of environmental and other exposures with health outcomes. Environ Int 2018;121(Pt 1):1027−31.ArticlePubMedPMC
  • [65] Riva JJ, Malik KMP, Burnie SJ, Endicott AR, Busse JW. What is your research question? An introduction to the PICOT format for clinicians. J Can Chiropr Assoc 2012;56(3):167−71.PubMedPMC
  • [66] Wildridge V, Bell L. How CLIP became ECLIPSE: a mnemonic to assist in searching for health policy/management information. Health Info Libr J 2002;19(2):113−5.ArticlePubMedPMC
  • [67] Cooke A, Smith D, Booth A. Beyond PICO: the SPIDER tool for qualitative evidence synthesis. Qual Health Res 2012;22(10):1435−43.ArticlePubMedPMCPDF
  • [68] Bramer WM, de Jonge GB, Rethlefsen ML, Mast F, Kleijnen J. A systematic approach to searching: an efficient and complete method to develop literature searches. J Med Libr Assoc 2018;106(4):531−41.ArticlePubMedPMCPDF
  • [69] Kwon CY, Park IS. National health insurance data as a research tool in Korean medicine: a guide to database utilization and methodological approaches. J Pharmacopuncture 2025;28(1):1−10.ArticlePubMedPMC
  • [70] Bahk J, Kim YY, Kang HY, Lee J, Kim I, Lee J, et al. Using the National Health Information Database of the National Health Insurance Service in Korea for monitoring mortality and life expectancy at national and local levels. J Korean Med Sci 2017;32(11):1764−70.ArticlePubMedPMCPDF
  • [71] Health Insurance Review & Assessment Service. A study on healthcare utilization patterns and equity by life cycle. Wonju (Korea), Health Insurance Review & Assessment Service, 2022.

Figure & Data

References

    Citations

    Citations to this article as recorded by  

      • PubReader PubReader
      • ePub LinkePub Link
      • Cite
        Download Citation
        Download a citation file in RIS format that can be imported by all major citation management software, including EndNote, ProCite, RefWorks, and Reference Manager.

        Format:
        • RIS — For EndNote, ProCite, RefWorks, and most other reference management software
        • BibTeX — For JabRef, BibDesk, and other BibTeX-specific software
        Include:
        • Citation for the content below
        Operational Definitions of Population, Intervention, Comparison, Outcome for Generating Real-World Evidence Using Korean Health Insurance Claims Sample Data in Integrative Medicine Research: A Methodological Guide
        Perspect Integr Med. 2026;5(2):73-84.   Published online June 19, 2026
        Close
      • XML DownloadXML Download
      Figure
      • 0
      Operational Definitions of Population, Intervention, Comparison, Outcome for Generating Real-World Evidence Using Korean Health Insurance Claims Sample Data in Integrative Medicine Research: A Methodological Guide
      Image
      Graphical abstract
      Operational Definitions of Population, Intervention, Comparison, Outcome for Generating Real-World Evidence Using Korean Health Insurance Claims Sample Data in Integrative Medicine Research: A Methodological Guide
      Category HIRA-NPS NHIS-NSC
      Health insurance qualification - BNC
      Birth & death - BND
      Medical service claims General specification T200 T20
      Medical service specification T300 T30
      Diagnosis statements T400 T40
      Drug prescriptions T530 T60
      Healthcare institutions YKIHO INST
      National health screening program - G1E
      Category Key variable HIRA-NPS NHIS-NSC
      General information Gender T200 BNC
      Age (y) T200 (5-y group) BND (year of birth)
      Residential area - BNC
      Insurance qualification T200 (Type) BNC (Type, premium decile)
      Disability (type & severity) - BNC
      Death (cause) - BND
      Diagnosis information Primary diagnosis (KCD codes) T200, T400 T20, T40
      1st secondary diagnosis (KCD codes) T200, T400 T20, T40
      All additional secondary diagnoses (KCD codes) T400 T20 (up to 4 secondary diagnoses), T40
      Medical service information Service type T200 T20
      Medical department T200, T400 T20, T40
      Service dates T200 T20
      Codes (Fee schedule, insurance, main ingredients, medical device) T300 T30
      Dosage & frequency T300, T530 T30, T60
      Surgery T200 T20
      Medical cost T200 T20
      Main ingredients T300, T530 T60
      ATC codes T300, T530 -
      Drug classification - T30, T60
      PICO element Information domain Examples Table: key variables
      P (Population) General information Defining eligibility criteria with sociodemographic characteristics BNC: sex, residential area, insurance qualification, disability
      BND: age, death
      Diagnosis information Defining the target diseases with diagnosis codes
      Refining eligibility criteria with prior/comorbid history
      T200 (T20): main diagnosis, 1st secondary diagnosis
      T400 (T40): all additional secondary diagnosis
      Medical service information Defining the study population with procedure/prescription history
      Defining exclusion criteria with prior healthcare utilization
      T200 (T20): service dates, service type
      T300 (T30): procedure & prescription code
      I (Intervention or exposure) General information Defining the intervention with healthcare service type T200 (T20): service type
      Medical service information Defining the intervention with procedure/prescription records T200 (T20): service dates
      T300 (T30): procedure & prescription code, dosage & frequency
      T530 (T60): dosage & frequency
      C (Comparison) General information Defining sociodemographic covariates for constructing comparable groups BNC: sex, residential area, insurance qualification
      BND: age
      Diagnosis information Defining clinical covariates for constructing comparable groups T200 (T20): main diagnosis, 1st secondary diagnosis
      T400 (T40): all additional secondary diagnosis
      Medical service information Defining non-exposed group with the absence of intervention of interest
      Defining healthcare-utilization covariates for comparison group construction
      T200 (T20): service dates
      T300 (T30): procedure & prescription code, dosage & frequency
      T530 (T60): dosage & frequency
      O (Outcome) General information Identifying mortality status & cause of death BND: death
      Diagnosis information Identifying recurrence, complications, or adverse events T200 (T20): main diagnosis, 1st secondary diagnosis
      T400 (T40): all additional secondary diagnosis
      Medical service information Identifying rehospitalization, emergency visits, procedures, prescriptions, or healthcare utilization measures T200 (T20): service dates, medical cost
      T300 (T30): procedure & prescription code, dosage & frequency
      T530 (T60): dosage & frequency
      PICO element Operational definition Table: key variables Notes
      P (population) Adults (age at cohort entry, ≥ 18 y) with low back pain identified using KCD codes (main or 1st secondary diagnosis); ≥ 2 outpatient claims or ≥ 1 inpatient claim T20: service dates
      T40: main diagnosis, 1st secondary diagnosis
      BND: age
      A 12-mo look-back period was used to exclude prior lumbar surgery
      I (Intervention or exposure) Acupuncture treatment defined using procedure/treatment codes; exposure defined as ≥ 6 acupuncture sessions within 8 weeks after index date T20: service dates
      T30: procedure & prescription code
      Index date was defined as 1st observed acupuncture service after a 12-mo acupuncture-free look-back
      C (comparison) New initiators of physical therapy/rehabilitation among patients with low back pain; confounding control via PSM using baseline covariates (age, sex, insurance qualification, CCI) T20: service dates, main diagnosis, 1st secondary diagnosis
      T30: procedure & prescription code
      T40: all additional secondary diagnosis
      BNC: sex, residential area, insurance qualification
      BND: age
      Active-comparator new-user design with same look-back period; Covariate balance assessed using SMDs
      O (outcome) Lumbar surgery within 365 days after landmark time T20: service dates
      T30: procedure & prescription code
      Follow-up began at 8-week landmark among event-free individuals (up to that point)
      Table 1 Structure of the HIRA-NPS and NHIS-NSC Databases

      The table labels of HIRA-NPS & NHIS-NSC are indicated for each category.

      HIRA = Health Insurance Review and Assessment Service (Republic of Korea); NHIS = National Health Insurance Service (Republic of Korea); NPS = National Patient Sample; NSC = National Sample Cohort.

      Table 2 Key Variables of the HIRA-NPS and the NHIS-NSC Databases

      ATC = anatomical therapeutic chemical; BMI = body mass index; DDM = doctor of dental medicine; DKM = doctor of Korean medicine; DM = doctor of medicine; DRG = diagnosis-related group; HIRA = Health Insurance Review and Assessment Service (Republic of Korea); KCD = Korean Standard Classification of Diseases; NHIS = National Health Insurance Service (Republic of Korea); NPS = National Patient Sample; NSC = National Sample Cohort.

      Table 3 Approach to Define PICO Elements Operationally in Claims-Based Research

      Table 4 Worked Example: Claims-Based Operationalization of PICO Elements in a Low Back Pain Cohort Using the NHIS-NSC Database

      CCI = Charlson comorbidity index; KCD = Korean Standard Classification of Diseases; NHIS = National Health Insurance Service (Republic of Korea); NSC = National Sample Cohort; PSM = propensity score matching; SMD = standardized mean difference.


      Perspect Integr Med : Perspectives on Integrative Medicine
      TOP