Chinese population chronic obstructive pulmonary disease risk assessment modeling method and system
Through whole-genome sequencing and environmental data integration among Chinese populations, a COPD risk assessment model was established, which solved the adaptability of genetic characteristics and environmental factors in COPD risk assessment, achieved efficient personalized prevention and standardized processes, and improved the prediction accuracy and clinical application value of the model.
Patent Information
- Application Number
- CN202510563133.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-30
- Publication Date
- 2025-08-15
AI Technical Summary
The existing COPD risk assessment model has insufficient standardization and quantification of genetic characteristics and environmental factors in the Chinese population, resulting in limited application value of the model in clinical practice.
Sample collection is adopted using unified standard operating procedures, whole genome sequencing is performed using the Illumina NovaSeq 6000 platform, feature selection is performed in combination with Plink v1.9 and elastic network regression, predictive models are established through multi-factor logistic regression, intervenable environmental factors are integrated, and user-friendly application tools are developed.
It significantly improves the prediction efficiency of genetic markers, innovatively integrates environmental factors, provides a basis for personalized prevention, establishes a standardized implementation process, lowers the threshold for clinical use, and improves the scientificity and practicality of the model.
Smart Images

Figure CN120496828A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of bioinformatics and precision medicine, and in particular to a method for modeling risk assessment of chronic obstructive pulmonary disease (COPD) in the Chinese population, and a system for modeling risk assessment of COPD in the Chinese population based on this method. Background Art
[0002] Chronic obstructive pulmonary disease (COPD), a common chronic respiratory disease, poses an increasing global burden. According to the World Health Organization, COPD has become the third leading cause of death worldwide, causing approximately 3.18 million deaths in 2019. In China, the burden of COPD is particularly severe. Epidemiological surveys show that the prevalence of COPD among people over 40 years old is as high as 13.7%, with an estimated total number of patients approaching 100 million. More worryingly, the death toll from COPD has continued to rise over the past decade, becoming the third leading cause of death after ischemic heart disease and stroke. In 2013, approximately 900,000 people died from COPD nationwide, surpassing only hypertension and diabetes, and the incidence rate has been increasing annually. In fact, COPD is affected by genetics, personal health and the environment. The currently known environmental factors with the greatest impact on COPD include smoking, age and PM2.5: smokers are more than twice as likely to develop COPD as non-smokers; in China, the proportion of people over 20 years old suffering from COPD is about 8.6%, while for those over 40 years old, the proportion is 14%; PM2.5 in the Pearl River Delta is better than in the Yangtze River Delta, and in the Yangtze River Delta is better than in North China and the Northeast Plain. The proportions of COPD in the three regions are 3.9%, 6.1% and 6.9% respectively.
[0003] The pathogenesis of COPD is complex and is the result of long-term interactions between genetic and environmental factors. In terms of genetics, genome-wide association studies (GWAS) have identified hundreds of single nucleotide polymorphism (SNP) sites associated with COPD. However, existing studies are mainly based on European populations. Due to differences in genetic background between populations, the predictive efficacy of these genetic markers in the Chinese population is significantly limited. Studies have shown that the predictive accuracy (AUC) of the polygenic risk score constructed based on the European population for the Chinese population is usually only slightly higher than 0.5, which seriously restricts its application value in clinical practice.
[0004] Regarding environmental factors, smoking is widely recognized as the most significant risk factor for COPD. China, the world's largest tobacco consumer, has a smoking rate of over 50% among adult men, directly contributing to the high incidence of COPD. Furthermore, long-term exposure to air pollution, particularly PM2.5, is also closely linked to COPD. In North my country, due to factors such as winter coal heating, PM2.5 concentrations are significantly higher than in other regions, correspondingly leading to a significantly higher incidence of COPD in the region. Other environmental risk factors include occupational dust exposure and indoor air pollution.
[0005] The limitations of existing technologies are as follows: First, most PRS models are developed based on European population data and fail to fully consider the genetic characteristics of the Chinese population; second, existing models often ignore the contribution of environmental factors, or simply incorporate them without standardized quantification methods; these defects make it difficult for existing technologies to meet the actual needs of COPD prevention and treatment in my country. Summary of the Invention
[0006] In order to overcome the defects of the existing technology, the technical problem to be solved by the present invention is to provide a risk assessment modeling method for chronic obstructive pulmonary disease in the Chinese population, which can solve the adaptability problems of genetic characteristics of the Chinese population, the standardized quantification problems of environmental factors, and the clinical practicality of the model.
[0007] The technical solution of the present invention is: This modeling method for risk assessment of chronic obstructive pulmonary disease in the Chinese population includes the following steps:
[0008] (1) Sample collection and processing: Samples were collected using a unified standard operating procedure (SOP). All samples were collected from 5 mL of peripheral venous blood and stored in EDTA anticoagulant tubes. DNA extraction was completed within 48 hours.
[0009] (2) Whole-genome sequencing and quality control: Sequencing was performed on the Illumina NovaSeq 6000 platform using the PE150 sequencing strategy; sequence alignment was performed using BWA-MEM v0.7.17, with the reference genome being GRCh38 / hg38;
[0010] (3) Environmental data collection, statistical analysis, and model construction: Environmental data were collected using standardized questionnaires. Genome-wide association analysis was performed using Plink v1.9. Elastic net regression was used for feature selection. A prediction model was then established using multivariate logistic regression, adjusting for age, sex, and principal components. The significance threshold was set at P < 5 × 10 -8 ; Environmental factors were integrated using a multivariate logistic regression model.
[0011] The beneficial technical effects of the present invention are:
[0012] 1) Designed specifically for the Chinese population, the predictive efficacy of genetic markers is significantly improved;
[0013] 2) Innovatively integrates modifiable environmental factors to provide a basis for personalized prevention;
[0014] 3) Established a standardized implementation process to ensure the reproducibility of results;
[0015] 4) Developed user-friendly application tools to lower the threshold for clinical use.
[0016] These innovations give the present invention significant advantages in terms of scientificity, practicality and economy, and provide new technical support for the precise prevention and treatment of COPD.
[0017] A risk assessment modeling system for chronic obstructive pulmonary disease in the Chinese population is also provided, which includes:
[0018] The sample collection and processing module uses a unified standard operating procedure (SOP) for sample collection. All samples are collected from 5 mL of peripheral venous blood, stored in EDTA anticoagulant tubes, and DNA extraction is completed within 48 hours.
[0019] Whole-genome sequencing and quality control module: Sequencing was performed on the Illumina NovaSeq 6000 platform using the PE150 sequencing strategy; sequence alignment was performed using BWA-MEM v0.7.17, with the reference genome being GRCh38 / hg38.
[0020] Environmental data collection, statistical analysis, and model building modules were used. Environmental data were collected using standardized questionnaires. Genome-wide association analysis was performed using Plink v1.9. Elastic net regression was used for feature selection. A prediction model was then established using multivariate logistic regression, adjusting for age, sex, and principal components. The significance threshold was set at P < 5 × 10 -8 ; Environmental factors were integrated using a multivariate logistic regression model. BRIEF DESCRIPTION OF THE DRAWINGS
[0021] Figure 1 This is a flowchart of the COPD risk assessment modeling method for the Chinese population based on the present invention. This modular design intuitively presents the entire process from sample collection to model application. The flowchart is divided into two sections: the upper section displays the experimental process, including DNA extraction, library construction, and sequencing analysis; the lower section presents the data analysis process, covering quality control, alignment, variant detection, association analysis, and model construction. Key parameters and quality control points are annotated for each step, such as sequencing depth ≥30× and Q30 >85%.
[0022] Figure 2 This is a Manhattan plot of the genome-wide association analysis, showing the strength of association between each chromosome region and the risk of COPD. The figure uses a double-panel design: the upper figure is the analysis result of the entire sample, and the lower figure is the analysis result of the subgroup of smokers. The horizontal axis marks the chromosome position, and the vertical axis shows -log10 (P value). The whole genome threshold line (P = 5×10 -7 ) and the suggestive threshold line (P = 1 × 10 -5 ).
[0023] Figure 3The number distribution of significant SNP sites at different P value thresholds is shown. The figure is in the form of a bar chart, with the horizontal axis representing the progressive P value threshold (from 5×10 -4 to 5×10 -8 ), the vertical axis represents the number of SNPs screened at the corresponding threshold. The color depth of the column indicates the source literature of the locus.
[0024] Figure 4 This is the result of gene function enrichment analysis, generated using Metascape software. The top figure shows pathway enrichment for all significant sites, highlighting pathways related to lung development and smoking. The bottom figure shows the specific enrichment pattern of sites in smoking-related samples, highlighting pathways such as nicotine metabolism and airway remodeling.
[0025] Figure 5 The receiver operating characteristic (ROC) curves compare the prediction performance of different models. The left figure shows the ROC curve of the simple PRS model (AUC = 0.68), and the right figure shows the comprehensive model after integrating environmental factors (AUC = 0.77).
[0026] Figure 6 A multi-panel design was used to demonstrate the dose-effect relationship of environmental factors.
[0027] Figure 7 Schematic diagram of the overall process of the modeling method for risk assessment of chronic obstructive pulmonary disease in the Chinese population according to the present invention. DETAILED DESCRIPTION
[0028] In order to make the purpose, technical solutions and advantages of the present invention more clearly understood, the present invention will be further described in detail below with reference to the accompanying drawings and specific embodiments. It should be understood that the specific embodiments described herein are only used to explain the present invention and are not intended to limit the present invention.
[0029] To provide a more detailed and complete description of the present invention, the following illustrative descriptions of embodiments and examples of the present invention are provided; however, these descriptions are not intended to be the only ways to implement or use the embodiments of the present invention. The embodiments cover features of various embodiments, as well as the method steps and sequences for constructing and operating these embodiments. However, other embodiments may be used to achieve the same or equivalent functionality and sequence of steps.
[0030] like Figure 7 As shown in Figure 2, this modeling method for COPD risk assessment in the Chinese population includes the following steps:
[0031] (1) Sample collection and processing: Samples were collected using a unified standard operating procedure (SOP). All samples were collected from 5 mL of peripheral venous blood and stored in EDTA anticoagulant tubes. DNA extraction was completed within 48 hours.
[0032] (2) Whole-genome sequencing and quality control: Sequencing was performed on the Illumina NovaSeq 6000 platform using the PE150 sequencing strategy; sequence alignment was performed using BWA-MEM v0.7.17, with the reference genome being GRCh38 / hg38;
[0033] (3) Environmental data collection, statistical analysis, and model construction: Environmental data were collected using standardized questionnaires. Genome-wide association analysis was performed using Plink v1.9. Elastic net regression was used for feature selection. A prediction model was then established using multivariate logistic regression, adjusting for age, sex, and principal components. The significance threshold was set at P < 5 × 10 -8 ; Environmental factors were integrated using a multivariate logistic regression model.
[0034] The beneficial technical effects of the present invention are:
[0035] 1) Designed specifically for the Chinese population, the predictive efficacy of genetic markers is significantly improved;
[0036] 2) Innovatively integrates modifiable environmental factors to provide a basis for personalized prevention;
[0037] 3) Established a standardized implementation process to ensure the reproducibility of results;
[0038] 4) Developed user-friendly application tools to lower the threshold for clinical use.
[0039] These innovations give the present invention significant advantages in terms of scientificity, practicality and economy, and provide new technical support for the precise prevention and treatment of COPD.
[0040] Preferably, in step (3), a two-stage strategy is adopted for the genetic score: first, significant loci are screened through genome-wide association analysis, and then the pruning + threshold method of the LDpred algorithm is applied to construct a polygenic risk score, multiple progressive P-value thresholds are set, and the optimal model is selected through cross-validation.
[0041] Preferably, in step (3), in terms of environmental factor integration, smoking intensity is quantified by the number of packs smoked per day × the number of years of smoking; PM2.5 exposure is calculated based on the annual average concentration based on historical monitoring data of the place of residence; and occupational exposure adopts the internationally accepted grading standard.
[0042] Preferably, in step (3), elastic network regression is used for feature selection, and then a final prediction model is established by multi-factor logistic regression. The model formula is as follows:
[0043] logit(P)=α+β1×PRS+β2×Smoking features+β4×Age+β5×Gender+ε
[0044] The PRS was standardized by Z-score; smoking intensity was grouped into quintiles; PM2.5 used the WHO-recommended grading standard; age and gender were included as covariates; α, β1, β2, β4, and β5 represent unknown coefficients required by the model, ε is the error term, PRS represents the polygenic risk score output by the LDpred software package, Age represents age, Gender represents gender, Smokingfeatures represents daily smoking volume, and P represents the risk of COPD.
[0045] Preferably, in step (3), the model validation randomly divides the original cohort into a training set and a validation set in a ratio of 7:3, and collects samples from an independent clinical center as an external validation cohort; validation indicators include AUC, sensitivity, specificity, decision curve analysis and reclassification improvement index.
[0046] Preferably, in step (1), DNA is extracted using a Qiagen Blood DNA Kit, and the extracted DNA is tested by Nanodrop, requiring the A260 / A280 ratio to be between 1.8 and 2.0 and the concentration to be ≥50 ng / μL; qualified samples are fragmented using a Covaris ultrasonic crusher, with a target fragment size of 350 bp; library construction is performed using an Illumina TruSeq DNA PCR-Free Library Prep Kit, and the library quality is tested by an Agilent 2100 Bioanalyzer, requiring the fragment distribution peak to be within the range of 350±50 bp.
[0047] Preferably, in step (2), sequencing is performed on an Illumina NovaSeq 6000 platform using a PE150 sequencing strategy, with a 5% PhiX control set in each lane to monitor sequencing quality. The raw data undergoes the following quality control steps:
[0048] (2.1) Trimmomatic v0.39 was used to remove adapter sequences and bases with poor quality (Q < 20);
[0049] (2.2) Data quality was checked by FastQC v0.11.9, with Q30 > 85% and GC content between 40% and 60%;
[0050] (2.3) Use bbduk.sh of BBTools to remove contaminating sequences, and qualified data will enter subsequent analysis.
[0051] Preferably, in step (2), after sequence alignment, SAMtools v1.11 is used for sorting, and Picard v2.23.5 is used to mark repeated sequences; the alignment quality requirements are: average depth ≥30×, coverage ≥95%, and alignment rate ≥98%; variation detection uses the HaplotypeCaller module of GATK v4.2.0, and the GVCF mode is used for joint calling; variation quality control standards include: QD ≥2.0, FS ≤60.0, MQ ≥40.0, MQRankSum ≥-12.5, and ReadPosRankSum ≥-8.0; and finally 9,112,201 high-quality variation sites are obtained.
[0052] A risk assessment modeling system for chronic obstructive pulmonary disease in the Chinese population is also provided, which includes:
[0053] The sample collection and processing module is configured to use a unified standard operating procedure (SOP) for sample collection. All samples are collected from 5 mL of peripheral venous blood, stored in EDTA anticoagulant tubes, and DNA extraction is completed within 48 hours.
[0054] Whole-genome sequencing and quality control modules were configured for sequencing on the Illumina NovaSeq 6000 platform using the PE150 sequencing strategy. Sequence alignment was performed using BWA-MEM v0.7.17, using the GRCh38 / hg38 reference genome.
[0055] Environmental data collection, statistical analysis, and model building modules were used. Environmental data were collected using standardized questionnaires. Genome-wide association analysis was performed using Plink v1.9. Elastic net regression was used for feature selection. A prediction model was then established using multivariate logistic regression, adjusting for age, sex, and principal components. The significance threshold was set at P < 5 × 10 -8 ; Environmental factors were integrated using a multivariate logistic regression model.
[0056] Preferably, the environmental data collection, statistical analysis, and model building modules use elastic network regression for feature selection, and then establish a final prediction model through multi-factor logistic regression. The model formula is as follows:
[0057] logit(P)=α+β1×PRS+β2×Smoking features+β4×Age+β5×Gender+ε
[0058] The PRS was standardized by Z-score; smoking intensity was grouped into quintiles; PM2.5 used the WHO-recommended grading standard; and age and gender were included as covariates.
[0059] The advantages of the present invention are mainly reflected in the following aspects:
[0060] At the data level, this invention has established the largest COPD research cohort in the Chinese population. Through cooperation with clinical institutions such as the China-Japan Friendship Hospital, a total of 5,943 samples that strictly met the GOLD diagnostic criteria were included (2,991 patients / 2,952 controls). Detailed epidemiological surveys were completed for all samples, and environmental data such as smoking history, occupational exposure history, and air pollution levels in the place of residence were systematically collected. To ensure data quality, a unified standard operating procedure was used for sample collection, processing, and storage to minimize batch effects and technical variations.
[0061] In terms of experimental technology, the present invention uses the Illumina NovaSeq high-throughput sequencing platform for whole-genome sequencing, achieving an average depth of 30× and a Q30 ratio exceeding 85%. Compared to conventional genotyping arrays, whole-genome sequencing can uncover more population-specific genetic variants, providing more comprehensive genetic information for model construction. Sequencing data undergoes a rigorous three-level quality control process: raw data is quality-filtered using Trimmomatic to remove low-quality reads; the BWA-MEM algorithm is used for alignment, using the hg38 reference genome; and variant detection follows the GATK best practice process to ensure the accuracy and reliability of the results.
[0062] Model construction is the core innovation of this invention. In terms of genetic scoring, a two-stage strategy is adopted: first, significant loci are screened through genome-wide association analysis, and then the "pruning + thresholding" (P+T) method of the LDpred algorithm is applied to construct a polygenic risk score. Compared with traditional methods, this invention sets multiple progressive P value thresholds (from 3×10 -1 to 1×10 -8 ) and select the optimal model through cross-validation. This strategy significantly improves the robustness and accuracy of the model. To integrate environmental factors, a standardized exposure assessment method was developed: smoking intensity was quantified using "pack-years" (number of packs smoked per day x number of years smoked); PM2.5 exposure was calculated based on historical residential monitoring data for annual average concentrations; and occupational exposure was graded using an internationally recognized grading standard. These standardized measures ensured data consistency and comparability.
[0063] This invention innovatively combines machine learning with traditional statistical methods. First, feature selection is performed using Elastic Net Regression, which can prevent model overfitting while performing feature selection. Then, a final prediction model is established through multi-factor logistic regression. This hybrid modeling strategy retains the interpretability of the statistical model while taking into account the ability of machine learning to process high-dimensional data. The logistic regression model formula is as follows:
[0064] logit(P)=α+β1×PRS+β2×Smoking features+β4×Age+β5×Gender+ε
[0065] The PRS was Z-score standardized; smoking intensity was categorized into quintiles; PM2.5 was categorized using the WHO-recommended grading system; and age and sex were included as covariates to ensure the model's applicability across diverse populations.
[0066] The model validation adopted a strict design scheme. The original cohort was randomly divided into a training set and a validation set in a ratio of 7:3. At the same time, 750 samples (267 patients / 483 controls) were collected from an independent clinical center as an external validation cohort. The validation indicators included not only traditional AUC, sensitivity, and specificity, but also introduced clinical practical indicators such as decision curve analysis and net reclassification improvement index. The results showed that the predictive performance of the integrated model was significantly better than that of the single-dimensional model. The AUC in the external validation set reached 0.77, which was 9% higher than the model containing only genetic factors.
[0067] To translate this into clinical applications, the present invention has developed a complementary risk assessment tool. This user-friendly tool supports multiple data input methods and outputs information including risk level, prevention recommendations, and follow-up plans. To accommodate diverse medical scenarios, the system is available in both a full and simplified version. The simplified version significantly reduces testing costs while maintaining predictive accuracy.
[0068] The above description is merely a preferred embodiment of the present invention and does not constitute any form of limitation to the present invention. Any simple modifications, equivalent changes and modifications made to the above embodiments based on the technical essence of the present invention are still within the scope of protection of the technical solution of the present invention.
Claims
1. A modeling method for risk assessment of chronic obstructive pulmonary disease in the Chinese population, characterized by: It includes the following steps: (1) Sample collection and processing: Samples were collected using a unified standard operating procedure (SOP). All samples were collected from 5 mL of peripheral venous blood and stored in EDTA anticoagulant tubes. DNA extraction was completed within 48 hours. (2) Whole-genome sequencing and quality control: Sequencing was performed on the Illumina NovaSeq 6000 platform using the PE150 sequencing strategy; sequence alignment was performed using BWA-MEM v0.7.17, with the reference genome being GRCh38 / hg38; (3) Environmental data collection, statistical analysis, and model construction: Environmental data were collected using standardized questionnaires. Genome-wide association analysis was performed using Plink v1.
9. Elastic net regression was used for feature selection. A prediction model was then established using multivariate logistic regression, adjusting for age, sex, and principal components. The significance threshold was set at P < 5 × 10 -8 ; Environmental factors were integrated using a multivariate logistic regression model.
2. The modeling method for risk assessment of chronic obstructive pulmonary disease in the Chinese population according to claim 1, characterized in that: In step (3), a two-stage strategy is adopted for genetic scoring: first, significant loci are screened through genome-wide association analysis, and then the pruning + threshold method of the LDpred algorithm is applied to construct a polygenic risk score, multiple progressive P-value thresholds are set, and the optimal model is selected through cross-validation.
3. The modeling method for risk assessment of chronic obstructive pulmonary disease in the Chinese population according to claim 2, characterized in that: In step (3), in terms of environmental factor integration, smoking intensity is quantified by the number of packs smoked per day × the number of years of smoking; PM2.5 exposure is calculated based on the annual average concentration based on historical monitoring data of the place of residence; and occupational exposure adopts the internationally accepted grading standard.
4. The modeling method for risk assessment of chronic obstructive pulmonary disease in the Chinese population according to claim 3, characterized in that: In step (3), elastic network regression is used for feature selection, and then the final prediction model is established by multi-factor logistic regression. The model formula is as follows: logit(P)=α+β1×PRS+β2×Smoking features+β4×Age+β5×Gender+ε The PRS was standardized by Z-score; smoking intensity was grouped into quintiles; PM2.5 used the WHO-recommended grading standard; age and gender were included as covariates; α, β1, β2, β4, and β5 represent unknown coefficients required by the model, ε is the error term, PRS represents the polygenic risk score output by the LDpred software package, Age represents age, Gender represents gender, Smokingfeatures represents daily smoking volume, and P represents the risk of COPD.
5. The modeling method for risk assessment of chronic obstructive pulmonary disease in the Chinese population according to claim 4, characterized in that: In step (3), the model validation randomly divides the original cohort into a training set and a validation set in a ratio of 7:3, and collects samples from an independent clinical center as an external validation cohort; validation indicators include AUC, sensitivity, specificity, decision curve analysis and reclassification improvement index.
6. The modeling method for risk assessment of chronic obstructive pulmonary disease in the Chinese population according to claim 1, characterized in that: In step (1), DNA was extracted using the Qiagen Blood DNA Kit, and the extracted DNA was tested by Nanodrop, requiring the A260 / A280 ratio to be between 1.8 and 2.0 and the concentration to be ≥50 ng / μL; qualified samples were fragmented using a Covaris ultrasonic crusher, with a target fragment size of 350 bp; library construction was performed using the Illumina TruSeq DNA PCR-Free Library Prep Kit, and the library quality was tested by Agilent 2100 Bioanalyzer, requiring the fragment distribution peak to be within the range of 350±50 bp.
7. The modeling method for risk assessment of chronic obstructive pulmonary disease in the Chinese population according to claim 6, characterized in that: In step (2), sequencing was performed on the Illumina NovaSeq 6000 platform using a PE150 sequencing strategy. A 5% PhiX control was set in each lane to monitor sequencing quality. The raw data passed the following quality control steps: (2.1) Trimmomatic v0.39 was used to remove adapter sequences and bases with poor quality (Q < 20); (2.2) Data quality was checked by FastQC v0.11.9, with Q30 > 85% and GC content between 40% and 60%; (2.3) Use bbduk.sh of BBTools to remove contaminating sequences, and qualified data will enter subsequent analysis.
8. The modeling method for risk assessment of chronic obstructive pulmonary disease in the Chinese population according to claim 7, characterized in that: In step (2), after sequence alignment, SAMtools v1.11 was used for sorting, and Picard v2.23.5 was used to mark repeated sequences; the alignment quality requirements were: average depth ≥30×, coverage ≥95%, and alignment rate ≥98%; variation detection used the HaplotypeCaller module of GATK v4.2.0, and joint calling was performed using the GVCF mode; variation quality control standards included: QD ≥2.0, FS ≤60.0, MQ ≥40.0, MQRankSum ≥-12.5, and ReadPosRankSum ≥-8.0; and finally 9,112,201 high-quality variation sites were obtained.
9. The Chinese population COPD risk assessment modeling system is characterized by: It includes: The sample collection and processing module uses a unified standard operating procedure (SOP) for sample collection. All samples are collected from 5 mL of peripheral venous blood, stored in EDTA anticoagulant tubes, and DNA extraction is completed within 48 hours. Whole-genome sequencing and quality control module: Sequencing was performed on the Illumina NovaSeq 6000 platform using the PE150 sequencing strategy; sequence alignment was performed using BWA-MEM v0.7.17, with the reference genome being GRCh38 / hg38. Environmental data collection, statistical analysis, and model building modules were used. Environmental data were collected using standardized questionnaires. Genome-wide association analysis was performed using Plink v1.
9. Elastic net regression was used for feature selection. A prediction model was then established using multivariate logistic regression, adjusting for age, sex, and principal components. The significance threshold was set at P < 5 × 10 -8 ; Environmental factors were integrated using a multivariate logistic regression model.
10. The Chinese population COPD risk assessment modeling system according to claim 9, characterized in that: The environmental data collection, statistical analysis, and model building modules use elastic network regression for feature selection, and then establish the final prediction model through multi-factor logistic regression. The model formula is as follows: logit(P)=α+β1×PRS+β2×Smoking features+β4×Age+β5×Gender+ε The PRS was standardized by Z-score; smoking intensity was grouped into quintiles; PM2.5 used the WHO-recommended grading standard; and age and gender were included as covariates.