Method and system for early warning and analyzing high-risk groups with high altitude polycythemia
By screening, structuring, and model generation, an early warning system for high-risk groups of polycythemia vera in high-altitude areas was established. This system addresses the limitations of missing data dimensions and static assessment, achieving efficient and accurate early warning for polycythemia vera in high-altitude areas, and is applicable to remote high-altitude regions.
Patent Information
- Application Number
- CN202511143249.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-08-15
- Publication Date
- 2025-11-21
AI Technical Summary
The lack of a systematic early warning analysis model for high-risk groups of polycythemia vera in existing technologies leads to missing data dimensions, limitations in static assessment, and crude privacy mechanisms, resulting in one-sided and unreliable early warning results.
By selecting subject datasets from extremely high-altitude areas, performing structured processing, extracting core predictive factors, generating candidate early warning models, and selecting the optimal early warning model, an early warning system for high-altitude polycythemia was established, including logistic regression, XGBoost, and random forest models, and the effectiveness was verified.
It enables multi-dimensional analysis of high-quality data, accurately identifies high-risk groups, dynamically couples environmental exposure and behavioral factors, generates personalized risk coefficients, supports targeted prevention, improves the sensitivity and specificity of the early warning system, and is suitable for remote plateau areas.
Smart Images

Figure CN120998531A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to a method and system for early warning analysis of high-risk groups for polycythemia vera at high altitudes, belonging to the field of biomedical engineering technology. Background Technology
[0002] High Altitude Polycythemia (HAPC) is a common chronic disease in extremely high-altitude areas above 4,500 meters. It has a high incidence rate and low patient awareness. In recent years, although a large number of studies have focused on the pathogenesis and risk factors of the disease, revealing the complex relationship between the hypoxic environment of high altitude and excessive erythrocyte proliferation, there is still a lack of clear guidelines and relevant research support on how to prevent the disease through lifestyle interventions.
[0003] Lifestyle interventions have been proven to play an important role in the prevention and management of chronic diseases. Studies have shown that unhealthy lifestyles, such as smoking, high-salt diets, and lack of exercise, are the main factors leading to the occurrence and development of chronic diseases. Through health education and lifestyle interventions, the incidence of chronic diseases and the incidence of complications can be significantly reduced. For example, health education can help patients better understand disease-related knowledge and improve their health behaviors, thereby effectively controlling the disease. In addition, lifestyle interventions also have significant economic benefits, with prevention costs far lower than treatment costs.
[0004] Lifestyle interventions are also of great significance in the prevention and treatment of high altitude polycythemia. Studies have shown that lifestyle changes such as reducing smoking, maintaining good work and rest habits, and controlling diet can effectively alleviate the symptoms of high altitude polycythemia. However, there are relatively few studies on lifestyle interventions for this disease, and there is a lack of systematic early warning and intervention models.
[0005] The aforementioned three defects—missing data dimensions, limitations of static assessment, and crude privacy mechanisms—jointly lead to the core problem of one-sided and unreliable early warning results. Summary of the Invention
[0006] This invention provides a method and system for early warning analysis of high-risk groups for polycythemia vera at high altitudes. Its main purpose is to solve the core problem that the early warning results are one-sided and unreliable due to the triple defects of missing data dimensions, limitations of static evaluation, and crude privacy mechanisms.
[0007] To achieve the above objectives, this invention provides a method for early warning analysis of high-risk groups for high-altitude polycythemia, comprising: The subject dataset was selected from extremely high altitude areas, including HAPC group datasets and non-HAPC group datasets. The subject dataset is processed in a structured manner to obtain a structured feature matrix, and core predictive factors are extracted from the structured feature matrix. Generate candidate early warning models for the core predictive factors, wherein the candidate early warning models include logistic regression models, XGBoost models, and random forest models; The optimal early warning model is selected from the candidate early warning models, and an early warning system for high-altitude polycythemia of the subject dataset is established using the optimal early warning model. The effectiveness of the high-altitude polycythemia early warning system was verified to achieve early warning analysis and processing of high-risk groups for high-altitude polycythemia in the subject dataset.
[0008] Optionally, the process of screening the subject dataset from extremely high altitude areas includes: Initial respondents were identified in health improvement programs at extremely high altitudes. The initial respondents were categorized into different social classes based on their place of residence, age, and gender. Random sampling was performed on the initial respondents corresponding to each stratum to obtain the random sample population; The subject inclusion operation was performed on the randomly sampled population to obtain the subject dataset; The extremely high altitude areas refer to areas with extreme low oxygen levels at an altitude of ≥4500 meters.
[0009] Optionally, the process of performing a subject inclusion operation on the randomly sampled population to obtain a subject dataset includes: Establish diagnostic criteria for HAPC in the randomly sampled population; The diagnostic criteria for HAPC include a hemoglobin level ≥210 g / L for men and a hemoglobin level ≥190 g / L for women. The randomly sampled population was grouped according to the HAPC diagnostic criteria to obtain the HAPC group and the non-HAPC group. Physiological indicators and questionnaire data were collected from the HAPC group and the non-HAPC group. The subject dataset was determined using the HAPC group population, the non-HAPC group population, the physiological indicators, and the questionnaire data.
[0010] Optionally, the step of performing structured processing on the subject dataset to obtain a structured feature matrix includes: Extract key indicators from the subject dataset; The key indicators include gender, age, blood oxygen saturation, BMI, smoking amount, tea consumption, blood pressure, and laboratory indicators. The key indicators are combined into a structured feature matrix.
[0011] Optionally, extracting core predictive factors from the structured feature matrix includes: The structured feature matrix is reduced in dimensionality using the LASSO regression method to obtain the reduced-dimensional features; The importance of the dimensionality-reduced features was analyzed by combining the random forest method and the XGBoost method, and the importance analysis results were obtained. Candidate factors are selected from the importance analysis results; The candidate factors were statistically significant using univariate and multivariate analysis methods, and the results of the significance verification were obtained. Based on the significance verification results, core predictive factors directly related to lifestyle interventions are selected from the candidate factors; The core predictive factors include tea consumption, smoking volume, and BMI. Among them, tea consumption is verified as a protective factor when tea consumption is ≥3 cups / day, smoking volume is verified as a risk factor when smoking volume is ≥10 cigarettes / day, and BMI is set to ≥25.7 based on the body fat distribution characteristics in extremely high altitude areas.
[0012] Optionally, the candidate early warning model for generating the core predictive factor includes: Extract the training set from the core predictive factors; The preset logistic regression model, preset XGBoost model, and preset random forest model are trained using the training set to obtain the logistic regression model, XGBoost model, and random forest model, respectively. The logistic regression model, the XGBoost model, and the random forest model are used as candidate early warning models.
[0013] Optionally, selecting the optimal early warning model from the candidate early warning models includes: Calculate the performance metrics of the candidate early warning model; The performance metrics include AUC, accuracy, sensitivity, and specificity; ROC curves and decision curves for the performance metrics are plotted. Based on the ROC curve and the decision curve, the optimal early warning model is selected from the candidate early warning models.
[0014] Optionally, the step of establishing a high-altitude polycythemia early warning system for the subject dataset using the optimal early warning model includes: Obtain the logistic regression coefficients of the optimal early warning model; The maximum absolute value among the logistic regression coefficients is used as the benchmark value; Based on the benchmark value, the logistic regression coefficients are converted into integer scores using the following formula:
[0015] in, This represents the integer score corresponding to the i-th logistic regression coefficient. This represents the i-th logistic regression coefficient. Indicates the baseline value. Represents the floor function; Set the risk threshold for the integer score; Based on the risk threshold, a risk level corresponding to the integer score is established to create an early warning system for high-altitude polycythemia in the subject dataset.
[0016] Optionally, the effectiveness verification of the high-altitude polycythemia early warning system includes: Extract a test set from the core predictive factors; The sensitivity and specificity of the high-altitude polycythemia early warning system were verified using the test set, and the sensitivity and specificity verification results were obtained. The effectiveness verification process of the high-altitude polycythemia early warning system is achieved by using the sensitivity verification results and the specificity verification results. The sensitivity verification results meet the clinical detection rate criteria, which means a clinical detection rate ≥ 85%, and the specificity verification results are determined by a finger-clip pulse oximeter.
[0017] To address the aforementioned problems, the present invention also provides a system for early warning and analysis of high-risk groups for polycythemia vera at high altitudes, the system comprising: A data filtering module is used to filter subject datasets from extremely high altitude areas, wherein the subject datasets include HAPC group datasets and non-HAPC group datasets; The factor extraction module is used to perform structured processing on the subject dataset to obtain a structured feature matrix, and extract core predictive factors from the structured feature matrix. The model generation module is used to generate candidate early warning models for the core predictive factors, wherein the candidate early warning models include logistic regression models, XGBoost models, and random forest models. The system establishment module is used to select the optimal early warning model from the candidate early warning models, and establish an early warning system for high-altitude polycythemia of the subject dataset through the optimal early warning model; The effectiveness verification module is used to verify the effectiveness of the high-altitude polycythemia early warning system, so as to realize the early warning analysis and processing of the high-risk population of high-altitude polycythemia in the subject dataset.
[0018] Compared to the problems described in the background section, this invention addresses the issue of missing data dimensions by selecting a dataset of subjects from extremely high-altitude regions. Stratified sampling covers residence (3 counties), age (6 age groups), and gender (36 social classes), ensuring the diversity of the sample representing areas at altitudes ≥4500 meters. The dataset includes 1089 subjects (109 in the HAPC group and 980 in the non-HAPC group), encompassing multidimensional data on physiology, biochemistry, and lifestyle, avoiding the limitations of a single data source. It also allows for precise identification of high-risk groups, employing strict diagnostic criteria (HGB ≥210g / L for males and ≥190g / L for females), significantly differentiating clinical characteristics between the HAPC and non-HAPC groups (e.g., median SpO2 of 79.0% in the HAPC group). The difference between the non-HAPC group and the non-HAPC group was 83.0%, p<0.001, providing a high-quality data foundation for model training. This invention, through screening key lifestyle factors and using LASSO regression to reduce the dimensionality of 82 indicators, combined with feature importance analysis using random forest and XGBoost, identified core factors such as tea consumption, smoking, and BMI. Tea consumption ≥3 cups / day was identified as a protective factor (OR=0.47, 95%CI: 0.26-0.85), and smoking ≥10 cigarettes / day was identified as a risk factor (OR=2.72, 95%CI: 1.21-5.72), directly correlated with lifestyle interventions. This also overcomes the limitations of static assessments by dynamically coupling environmental exposure. ) and behavioral factors (smoking / tea drinking) to generate personalized risk coefficients (e.g. (The OR of 4.35 (<83% OR = 4.35) is the strongest predictor), supporting targeted prevention. This invention optimizes performance through multi-model comparison. The AUCs of logistic regression, XGBoost, and random forest are 0.848, 0.828, and 0.735, respectively. The logistic regression model combines high interpretability and clinical applicability, outputting an interpretable risk coefficient, facilitating adjustments to lifestyle intervention strategies by medical personnel and overcoming the limitations of static assessment. This invention selects the optimal early warning model from the candidate models and compares the performance of logistic regression, XGBoost, and random forest. The results... The results show that the logistic regression model has the highest AUC value, indicating its high sensitivity and specificity in identifying high-risk groups for high-altitude polycythemia vera (HAPC). Although the random forest and XGBoost models perform similarly on some indicators, the logistic regression model has stronger interpretability, making it easier for clinical application and promotion. Furthermore, decision curve analysis (DCA) further confirms the application value of the logistic regression model in clinical decision-making. This invention's embodiments validate sensitivity ≥85% using a test set, addressing the issue of inadequate privacy mechanisms. Specificity validation utilizes portable devices such as finger-clip pulse oximeters, making it suitable for remote high-altitude areas and meeting the needs of public health scenarios. Therefore, the method and system for early warning analysis of high-risk groups for high-altitude polycythemia vera provided by this invention can solve the core problem of one-sided and unreliable early warning results caused by the triple defects of missing data dimensions, limitations of static evaluation, and inadequate privacy mechanisms. Attached Figure Description
[0019] Figure 1 This is a flowchart illustrating a method for early warning analysis of high-risk groups for polycythemia vera at high altitudes, provided in an embodiment of the present invention. Figure 2 This is a flowchart illustrating the process of screening subject datasets for a method to predict and analyze high-risk populations for polycythemia vera in high-altitude areas, as provided in an embodiment of the present invention. Figure 3a This is a schematic diagram of the feature importance analysis using random forest method to implement an early warning analysis method for high-risk groups of polycythemia vera in high altitude areas, according to an embodiment of the present invention. Figure 3b This is a schematic diagram of the feature importance analysis using the XGBoost method to implement an early warning analysis method for high-risk groups of polycythemia vera in high altitude regions, according to an embodiment of the present invention. Figure 4 This is a schematic diagram of a logistic regression forest plot used in an embodiment of the present invention to implement an early warning analysis method for high-risk groups of polycythemia vera at high altitudes. Figure 5a This is a schematic diagram of the decision curve for implementing an early warning analysis method for high-risk groups of polycythemia vera in high-altitude areas, according to an embodiment of the present invention. Figure 5bThis is a schematic diagram of the decision curve for implementing an early warning analysis method for high-risk groups of polycythemia vera in high-altitude areas, according to an embodiment of the present invention. Figure 5c This is a schematic diagram of the decision curve for implementing an early warning analysis method for high-risk groups of polycythemia vera in high-altitude areas, according to an embodiment of the present invention. Figure 6a This is a schematic diagram illustrating the performance indicators of a method for early warning analysis of high-risk groups for polycythemia vera in high-altitude areas, provided in an embodiment of the present invention. Figure 6b This is a schematic diagram illustrating the performance indicators of a method for early warning analysis of high-risk groups for polycythemia vera in high-altitude areas, provided in an embodiment of the present invention. Figure 7 This is a schematic diagram of a module for implementing the early warning and analysis system for high-risk groups of plateau polycythemia, provided in an embodiment of the present invention.
[0020] The objectives, features, and advantages of this invention will be further explained in conjunction with the embodiments and with reference to the accompanying drawings. Detailed Implementation
[0021] It should be understood that the specific embodiments described herein are merely illustrative of the invention and are not intended to limit the invention.
[0022] This application provides a method for early warning analysis of high-risk groups for high-altitude polycythemia. The executing entity of this method includes at least one electronic device, such as a server or a terminal, that can be configured to execute the method provided in this application. In other words, the method for early warning analysis of high-risk groups for high-altitude polycythemia can be executed by software or hardware installed on a terminal device or a server device. The server includes a single server, a server cluster, a cloud server, or a cloud server cluster. Example 1:
[0023] Reference Figure 1 The diagram shown is a flowchart illustrating a method for early warning analysis of high-risk groups for high-altitude polycythemia according to an embodiment of the present invention. In this embodiment, the method for early warning analysis of high-risk groups for high-altitude polycythemia includes: S1. Select subject datasets from extremely high altitude areas, wherein the subject datasets include HAPC group datasets and non-HAPC group datasets.
[0024] This invention addresses the issue of missing data dimensions by selecting a subject dataset from extremely high-altitude regions. Stratified sampling covers residence (3 counties), age (6 age groups), and gender (36 social strata), ensuring the diversity of the sample representing areas at altitudes ≥4500 meters. The dataset includes 1089 subjects (109 in the HAPC group and 980 in the non-HAPC group), encompassing multidimensional data on physiology, biochemistry, and lifestyle, avoiding the limitations of a single data source. It also allows for precise identification of high-risk groups, employing strict diagnostic criteria (HGB ≥210g / L for males and ≥190g / L for females) to significantly differentiate the clinical characteristics between the HAPC and non-HAPC groups (e.g., HAPC group). The median was 79.0% (v vs 83.0% in the non-HAPC group, p < 0.001), providing a high-quality data foundation for model training.
[0025] In one embodiment of the present invention, the step of screening the subject dataset from extremely high altitude areas includes: querying the initial respondents of the health improvement project in extremely high altitude areas; classifying the initial respondents into different social strata according to their place of residence, age, and gender; randomly sampling the initial respondents corresponding to each social stratum to obtain a randomly sampled population; and performing a subject inclusion operation on the randomly sampled population to obtain a subject dataset; wherein, the extremely high altitude areas refer to extremely low oxygen areas with an altitude ≥ 4500 meters.
[0026] It should be noted that when launching the High Altitude Health Improvement (HI-VHA) project (i.e., the health improvement project), 3,564 people were recruited as initial respondents from three counties: Shuanghu County (4,700 meters above sea level), Nima County (4,530 meters above sea level), and Ando County (4,570 meters above sea level). These counties include 8 townships and 51 villages at altitudes above 4,500 meters. To ensure the representativeness of the sampling, the population was divided into 36 strata according to place of residence (three counties), age group (7-20 years, 20-30 years, 30-40 years, 40-50 years, 50-60 years, and ≥60 years), and sex (male and female). Researchers randomly sampled from each stratum. After implementing strict inclusion criteria, a total of 1,089 individuals from the HI-VHA group who had HAPC and had undergone pulmonary function tests were included in this study, which is the included subject dataset.
[0027] See Figure 2 The diagram shown is a flowchart illustrating the process of screening a subject dataset for a method to predict and analyze high-risk populations for polycythemia vera at high altitudes, according to an embodiment of the present invention. Figure 2In this context, "Very High Altitude Population Sampling" indicates sampling of people at extremely high altitudes; "Identifying the study population" indicates identifying the study population; "Random sampling" indicates random sampling; "Divided by residence," "Divided by ages," and "Divided by gender" indicate stratification by residence, age, and gender, respectively; "Collection of 3654 samples" indicates collecting 3654 samples; "Samples are taken from each sampling group to ensure that the sample size for each stratum is not less than 10%" indicates that samples are drawn from each sampling group to ensure that the sample size for each stratum is not less than 10%; "Excluded 1344 unhealthy individuals" and "Excluded 653 individuals ≤ 18 years old" indicate excluding 1344 unhealthy individuals and 653 individuals aged ≤ 18 years; and "Identify 1656 study healthy populations" indicates identifying 1656 healthy study populations.
[0028] In another embodiment of the present invention, the step of performing subject inclusion operations on the randomly sampled population to obtain a subject dataset includes: setting HAPC diagnostic criteria for the randomly sampled population; wherein, the HAPC diagnostic criteria include hemoglobin levels ≥210 g / L for males and ≥190 g / L for females; grouping the randomly sampled population according to the HAPC diagnostic criteria to obtain an HAPC group and a non-HAPC group; collecting physiological indicators and questionnaire data from the HAPC group and the non-HAPC group; and determining the subject dataset using the HAPC group, the non-HAPC group, the physiological indicators, and the questionnaire data.
[0029] Please refer to Table 1 below, which is a schematic table of the subject dataset. In Table 1, the first column on the left includes physiological indicators and data such as tea consumption and smoking volume extracted from the questionnaire data:
[0030] Table 1 above lists a total of 1089 participants in this study, including 980 in the non-HAPC group and 109 in the HAPC group. Table 1 shows that, in terms of epidemiological and physiological indicators, the hemoglobin (Hgb) level in the HAPC group was significantly higher than that in the Non-HAPC group (medians were 218 g / L and 168 g / L, respectively). The median blood pressure in the HAPC group was significantly higher than that in the Non-HAPC group (64% vs 44%, p<0.001), and the median age in the HAPC group was also higher than that in the Non-HAPC group (46 years vs 41 years, p=0.002). In addition, the blood oxygen saturation (SpO2) in the HAPC group was significantly lower than that in the Non-HAPC group (medians were 79.0% and 83.0%, respectively, p<0.001), while the smoking rate and tea drinking habits did not differ significantly between the two groups (p>0.05). The body mass index (BMI) in the HAPC group was significantly higher than that in the Non-HAPC group (medians were 25.7 and 22.7, respectively, p<0.001), and the systolic blood pressure (SBP), diastolic blood pressure (DBP), and mean arterial pressure (MBP) in the HAPC group were all significantly higher than those in the Non-HAPC group (p<0.001).
[0031] In terms of laboratory indicators, the HAPC group showed significant abnormalities in several biochemical indicators, including higher levels of total bilirubin (TBIL), direct bilirubin (DBIL), indirect bilirubin (IBIL), urea (UREA), creatinine (CREA), uric acid (UA), triglycerides (TG), low-density lipoprotein cholesterol (LDLC), and C-reactive protein (CRP), while high-density lipoprotein cholesterol (HDLC) levels were lower (p<0.001). In addition, the serum potassium (K) level in the HAPC group was significantly higher than that in the Non-HAPC group (p<0.001), while the levels of chloride (CL), calcium (CA), magnesium (MG), and phosphorus (Pi) did not differ significantly (p>0.05). The homocysteine (HCY) level in the HAPC group was significantly higher than that in the Non-HAPC group (p<0.001), and the estimated glomerular filtration rate (eGFR) was slightly lower in the HAPC group than in the Non-HAPC group (p=0.002).
[0032] Regarding pulmonary function indicators, the HAPC group showed no significant differences compared to the Non-HAPC group in terms of maximum voluntary ventilation (VCMAX), residual volume (ERV), inspiratory volume (IC), minute ventilation (MV), tidal volume (VT), forced expiratory volume (FVCEX, FEV1, FEV1 / FVCEX), peak expiratory flow (PEF), and maximum mid-expiratory flow (MEF75, MEF50, MEF25, MEF25-75) (p>0.05). However, some indicators, such as FEV1 / FVCEX and PEF, were slightly lower in the HAPC group (p>0.05). These results indicate that patients with high-altitude polycythemia have significant differences from non-patients in multiple clinical characteristics, biochemical indicators, and pulmonary function indicators, and these differences may be closely related to the occurrence and development of the disease.
[0033] S2. Perform structured processing on the subject dataset to obtain a structured feature matrix, and extract core predictive factors from the structured feature matrix.
[0034] This invention, through screening key lifestyle factors and using LASSO regression to reduce the dimensionality of 82 indicators, combined with feature importance analysis using random forest and XGBoost, identifies core factors such as tea consumption, smoking, and BMI. Tea consumption ≥3 cups / day is considered a protective factor (OR=0.47, 95%CI: 0.26-0.85), while smoking ≥10 cigarettes / day is considered a risk factor (OR=2.72, 95%CI: 1.21-5.72), directly linking to lifestyle interventions. It also overcomes the limitations of static assessment by dynamically coupling environmental exposure (SpO2) with behavioral factors (smoking / tea consumption) to generate personalized risk coefficients (e.g., OR=4.35 for SpO2 <83% is the strongest predictor), supporting targeted prevention.
[0035] In one embodiment of the present invention, the step of performing structured processing on the subject dataset to obtain a structured feature matrix includes: extracting key indicators from the subject dataset; wherein the key indicators include gender, age, blood oxygen saturation, BMI, smoking amount, tea consumption, blood pressure, and laboratory indicators; and combining the key indicators into a structured feature matrix.
[0036] The structured feature matrix is mainly composed of the key indicators, for example, the key indicators are combined into a matrix with one row and multiple columns.
[0037] In one embodiment of the present invention, the step of extracting core predictive factors from the structured feature matrix includes: performing feature dimensionality reduction on the structured feature matrix using LASSO regression to obtain dimensionality-reduced features; performing feature importance analysis on the dimensionality-reduced features using random forest and XGBoost methods to obtain importance analysis results; screening candidate factors from the importance analysis results; performing statistical significance verification on the candidate factors using univariate analysis and multivariate analysis to obtain significance verification results; and selecting core predictive factors directly related to lifestyle intervention from the candidate factors based on the significance verification results. The core predictive factors include tea consumption, smoking volume, and BMI, wherein tea consumption is verified as a protective factor when tea consumption is ≥3 cups / day, smoking volume is verified as a risk factor when smoking volume is ≥10 cigarettes / day, and BMI is set to ≥25.7 based on the body fat distribution characteristics in extremely high-altitude areas.
[0038] The BMI ≥ 25.7 was set based on the body fat distribution characteristics in extremely high altitude areas because: the body fat distribution of people in extremely high altitude (≥ 4500 meters) is different from that in plain areas due to the low oxygen environment. The median BMI of the HAPC group was significantly higher than that of the non-HAPC group (25.7 vs 22.7, p < 0.001).
[0039] Optionally, the LASSO regression method is used to reduce the dimensionality of the structured feature matrix, resulting in the following dimensionality-reduced features: Feature selection is performed using the LASSO model to screen key predictive variables from 82 epidemiological factors and physiological and biochemical indicators, including tea consumption, smoking, gender, blood pressure, age, blood oxygen saturation (SpO2), and body mass index (BMI). Further, the candidate factors are statistically significant using univariate and multivariate analysis methods. The results show that the independent association between each candidate factor (e.g., smoking amount, tea consumption) and high-altitude polycythemia vera (HAPC) is tested individually. Factors with a p-value <0.1 are preliminarily screened, and factors with significant univariate results are included in the logistic regression model to correct for confounding factors. Multivariate analysis was conducted to identify confounding factors (such as age and SpO2) and verify their independent effects (OR value and 95% CI not exceeding 1.0). Factors that were statistically significant (P<0.05) and had consistent effects in the multivariate analysis (e.g., smoking OR=2.72 is a risk factor) were output. Furthermore, based on the significance verification results, core predictive factors directly related to lifestyle intervention were selected from the candidate factors as follows: only factors significantly related to HAPC in the multivariate analysis (e.g., smoking, tea drinking, BMI) were retained, non-interventional factors (e.g., age) were removed, and factors that could be modified through lifestyle (e.g., drinking ≥3 cups of tea / day is a protective factor) were retained. The intervention threshold was determined based on the effect size (e.g., OR=1.06 for BMI≥25.7) and clinical consensus.
[0040] Please refer to Table 2 below, which illustrates the methods of univariate analysis and multivariate analysis:
[0041] See Figure 3a The diagram shown illustrates a random forest feature importance analysis method for implementing an early warning analysis method for high-risk groups of plateau polycythemia, according to an embodiment of the present invention. Figure 3a In the table, the importance analysis results obtained by performing feature importance analysis on the dimensionality-reduced features using the random forest method are shown. The first column on the left represents each dimensionality-reduced feature, and the length of the rectangle in each row represents the importance of each dimensionality-reduced feature.
[0042] See Figure 3b The diagram shown illustrates the feature importance analysis using the XGBoost method for early warning analysis of high-risk groups for polycythemia vera in high-altitude areas, as provided in an embodiment of the present invention. Figure 3b In the table, the importance analysis results obtained by performing feature importance analysis on the dimensionality-reduced features using the XGBoost method are shown. The first column on the left represents each dimensionality-reduced feature, and the length of the rectangle in each row represents the importance of each dimensionality-reduced feature.
[0043] See Figure 4 The image shown is a schematic diagram of a logistic regression forest plot used in an embodiment of the present invention to implement an early warning analysis method for high-risk groups of polycythemia vera at high altitudes. Figure 4 In this context, p-value represents the P-value, odds ratio represents the odds ratio (or advantage ratio), and the logistic regression forest plot is a visualization of the multifactor analysis data mentioned above.
[0044] It should be noted that, through comprehensive analysis of the feature importance of both Random Forest and XGBoost machine learning algorithms, and single / multivariate analysis, the top ten variables contributing most to the model's predictive ability were selected. These variables include spo2 (blood oxygen saturation), UA (uric acid), TBIL (total bilirubin), IBIL (indirect bilirubin), CL (chloride), left (left brain oxygen saturation), BMI (body mass index), DBIL (direct bilirubin), K (potassium), and CA (calcium). Among these variables, factors closely related to lifestyle, such as tea consumption, smoking, gender, blood pressure, age, blood oxygen saturation, and BMI, were particularly emphasized. Multivariate logistic regression revealed that spo2 (blood oxygen saturation) showed extremely high importance in both algorithms (Random Forest and XGBoost), indicating that blood oxygen levels have a significant impact on the predictive target variable. Furthermore, BMI (body mass index), as an indicator reflecting the ratio of an individual's weight to height, also ranked highly in both algorithms, further confirming the importance of body oxygen saturation in predicting the target variable. The importance of management in health prediction was demonstrated. Male sex was significant in logistic regression analysis (p<0.001), indicating it as a risk factor. Tea consumption and smoking, as representative lifestyle factors, also showed significant predictive power in the logistic regression forest plot. Tea consumption had a significant negative correlation with the prediction results (odds ratio = 0.350, 95% confidence interval: 0.160–0.880, p = 0.018), while smoking had a significant positive correlation (odds ratio = 2.720, 95% confidence interval: 1.210–5.720, p = 0.011). In summary, through comprehensive analysis of feature importance, random forest feature importance, and logistic regression forest plot, a set of lifestyle-based predictive factors was selected. These factors included tea consumption, smoking, sex, blood pressure, age, blood oxygen saturation, and BMI. They demonstrated significant predictive power in the model, providing an important basis for subsequent model building.
[0045] S3. Generate candidate early warning models for the core predictive factors, wherein the candidate early warning models include logistic regression models, XGBoost models, and random forest models.
[0046] The embodiments of this invention optimize performance through multi-model comparison. The AUCs of logistic regression, XGBoost, and random forest are 0.848, 0.828, and 0.735, respectively. Among them, the logistic regression model has both high interpretability and clinical applicability. The model outputs an interpretable risk coefficient, which makes it easier for medical staff to adjust lifestyle intervention strategies and overcome the limitations of static assessment.
[0047] In one embodiment of the present invention, generating a candidate early warning model for the core predictive factor includes: extracting a training set from the core predictive factor; training a preset logistic regression model, a preset XGBoost model, and a preset random forest model using the training set to obtain the logistic regression model, the XGBoost model, and the random forest model; and using the logistic regression model, the XGBoost model, and the random forest model as candidate early warning models.
[0048] The dataset for the core predictors was divided into a training set and a test set, with a ratio of 80:20.
[0049] S4. Select the optimal early warning model from the candidate early warning models, and establish an early warning system for high-altitude polycythemia of the subject dataset using the optimal early warning model.
[0050] This invention employs a method to select the optimal early warning model from the candidate early warning models and compare the performance of three models: logistic regression, XGBoost, and random forest. The results show that the logistic regression model has the highest AUC value, indicating that it has high sensitivity and specificity in identifying high-risk individuals for HAPC. Although the random forest and XGBoost models perform similarly on some indicators, the logistic regression model has stronger interpretability and is easier to apply and promote in clinical practice. Furthermore, decision curve analysis (DCA) further confirms the application value of the logistic regression model in clinical decision-making.
[0051] In one embodiment of the present invention, the step of selecting the optimal early warning model from the candidate early warning models includes: calculating the performance indicators of the candidate early warning models; wherein the performance indicators include AUC, accuracy, sensitivity, and specificity; plotting the ROC curve and decision curve of the performance indicators; and selecting the optimal early warning model from the candidate early warning models based on the ROC curve and the decision curve.
[0052] It should be noted that the study used decision curve analysis (DCA) to evaluate the clinical utility of three different prediction models—random forest (rf_model), XGBoost (xgb_model), and logistic regression (logistics_model)—at different threshold probabilities.
[0053] See Figure 5a The image shown is a schematic diagram of the decision curve for an early warning analysis method for high-risk groups of polycythemia vera at high altitudes, provided by an embodiment of the present invention. Figure 5aThe decision curve of the logistic regression model (lr_model) is shown in the figure. Compared with the other two models, the net return of the logistic regression model is slightly lower at certain threshold probabilities. Nevertheless, in the lower threshold probability range, the net return of the model is still higher than the "do nothing" strategy, showing certain clinical application value.
[0054] See Figure 5b The image shown is a schematic diagram of the decision curve for an early warning analysis method for high-risk groups of polycythemia vera at high altitudes, provided by an embodiment of the present invention. Figure 5b The decision curve of the XGBoost model (xgb_model) is shown in the figure. Similar to the random forest model, the XGBoost model also shows high net returns at lower threshold probabilities. In most threshold ranges, the net returns of the XGBoost model are also higher than the "do all" and "do none" strategies, indicating that the model has high application value in clinical decision-making.
[0055] See Figure 5c The image shown is a schematic diagram of the decision curve for an early warning analysis method for high-risk groups of polycythemia vera at high altitudes, provided by an embodiment of the present invention. Figure 5c In the study, the random forest model (rf_model) showed high net returns at lower threshold probabilities (0.0 to 0.2). As the threshold probability increased, the net returns gradually decreased. However, within most threshold ranges, the net returns of this model were higher than the "do all" (red curve) and "do none" (green curve) strategies, showing good potential for clinical application.
[0056] Furthermore, a comprehensive comparison of the decision curves of the three models reveals that, within a lower threshold probability range, the net returns of the Random Forest and XGBoost models are higher than those of the Logistic Regression model, demonstrating better clinical utility. However, within a higher threshold probability range, the net returns of the three models are not significantly different and are all lower than the "do all" strategy.
[0057] Furthermore, in the model evaluation, the performance of three models—logistic regression, XGBoost, and random forest—was compared, and the results are presented in... Figure 6a and Figure 6b middle.
[0058] See Figure 6a The diagram shown illustrates the performance indicators of a method for early warning analysis of high-risk groups for polycythemia vera at high altitudes, provided by an embodiment of the present invention. Figure 6a The diagram shows the ROC curves of three models. The model performance is evaluated by plotting the sensitivity and specificity at different thresholds. The closer the ROC curve is to the upper left corner, the better the model's performance. The ROC curves of XGBoost and Random Forest are quite similar, and both are significantly better than the Logistic Regression model.
[0059] See Figure 6b The diagram shown illustrates the performance indicators of a method for early warning analysis of high-risk groups for polycythemia vera at high altitudes, provided by an embodiment of the present invention. Figure 6b The table shows the scores of the three models on four metrics: accuracy, area under the curve (AUC), sensitivity, and specificity. It can be seen that the three models perform similarly on these metrics, but there are subtle differences in some aspects. The Random Forest model is slightly higher than the other two models in accuracy and AUC, indicating its better ability to distinguish between positive and negative samples. In terms of sensitivity, the three models perform almost identically, showing their comparable ability to identify positive samples. In terms of specificity, the XGBoost model is slightly higher than the other two models, showing its advantage in identifying negative samples. Random Forest and XGBoost have higher net returns, but ultimately logistic regression is chosen.
[0060] In one embodiment of the present invention, establishing a high-altitude polycythemia early warning system for the subject dataset using the optimal early warning model includes: obtaining the logistic regression coefficients of the optimal early warning model; using the maximum absolute value of the logistic regression coefficients as a benchmark value; and converting the logistic regression coefficients into integer scores using the following formula based on the benchmark value:
[0061] in, This represents the integer score corresponding to the i-th logistic regression coefficient. This represents the i-th logistic regression coefficient. Indicates the baseline value. Represents the floor function; Set a risk threshold for the integer score; based on the risk threshold, establish a risk level corresponding to the integer score to establish a high-altitude polycythemia early warning system for the subject dataset.
[0062] The high-altitude polycythemia early warning system is based on a quantitative scoring model constructed using core predictive factors (tea consumption, smoking, BMI, etc.). It uses integer scores to classify risk levels and enables dynamic monitoring and lifestyle intervention guidance for high-risk groups.
[0063] Optionally, the risk level corresponding to the integer score is established based on the risk threshold. For example, if the sum of the integer scores is S, S belonging to the A~B range (i.e., the risk threshold) is low risk, S belonging to the C~D range (i.e., the risk threshold) is low risk, and S belonging to the E~F range (i.e., the risk threshold) is low risk.
[0064] S5. Verify the effectiveness of the high-altitude polycythemia early warning system to achieve early warning analysis and processing of the high-risk population for high-altitude polycythemia in the subject dataset.
[0065] The embodiments of this invention verify a sensitivity of ≥85% through a test set, thus addressing the issue of crude privacy mechanisms. The specificity verification employs portable devices such as finger-clip pulse oximeters, making it suitable for remote high-altitude areas and meeting the needs of public health scenarios.
[0066] In one embodiment of the present invention, the effectiveness verification of the high-altitude polycythemia early warning system includes: extracting a test set from the core predictive factors; using the test set to perform sensitivity verification and specificity verification on the high-altitude polycythemia early warning system, respectively, to obtain sensitivity verification results and specificity verification results; and realizing the effectiveness verification processing of the high-altitude polycythemia early warning system through the sensitivity verification results and the specificity verification results; wherein, the sensitivity verification results meet the clinical detection rate condition, the clinical detection rate condition refers to a clinical detection rate ≥85%, and the specificity verification results are determined by a finger-clip pulse oximeter.
[0067] Optionally, the sensitivity and specificity of the high-altitude polycythemia vera early warning system are verified using the test set. The sensitivity and specificity verification results are as follows: the detection rate (true positive rate) of the model for HAPC positive samples is calculated in the test set, which must be ≥85%; SpO2 is measured using a finger-clip pulse oximeter (gold standard); the exclusion rate (true negative rate) of the model for negative samples is calculated; a confusion matrix is generated; and the sensitivity and specificity values are quantified (e.g., sensitivity 92%, specificity 89%). Further, the high-altitude polycythemia vera early warning system is implemented based on the sensitivity and specificity verification results. The effectiveness verification process is as follows: the sensitivity verification result must meet the requirement of a clinical detection rate ≥85% (i.e., a missed diagnosis rate ≤15%), and the specificity verification result must be confirmed by a finger-clip pulse oximeter (portable device) to ensure practical feasibility in high-altitude areas. If both indicators meet the standards, the early warning system passes the verification (e.g., sensitivity 92%>85%, specificity 89%>80%). Among them, the finger-clip pulse oximeter, as an SpO2 detection tool, has been listed as a core diagnostic item of HAPC. The specificity verification result is determined by the finger-clip pulse oximeter as follows: the specificity verification uses a finger-clip pulse oximeter to measure SpO2 and compares it with the gold standard (hemoglobin detection) to determine the specificity.
[0068] Compared to the problems described in the background section, this invention addresses the issue of missing data dimensions by selecting a dataset of subjects from extremely high-altitude regions. Stratified sampling covers residence (3 counties), age (6 age groups), and gender (36 social classes), ensuring the diversity of the sample representing areas at altitudes ≥4500 meters. The dataset includes 1089 subjects (109 in the HAPC group and 980 in the non-HAPC group), encompassing multidimensional data on physiology, biochemistry, and lifestyle, avoiding the limitations of a single data source. It also allows for precise identification of high-risk groups, employing strict diagnostic criteria (HGB ≥210g / L for males and ≥190g / L for females), significantly differentiating clinical characteristics between the HAPC and non-HAPC groups (e.g., median SpO2 of 79.0% in the HAPC group). The difference between the non-HAPC group and the non-HAPC group was 83.0%, p<0.001, providing a high-quality data foundation for model training. This embodiment of the invention screens key lifestyle factors, reduces the dimensionality of 82 indicators using LASSO regression, and combines random forest and XGBoost feature importance analysis to identify core factors such as tea consumption, smoking, and BMI. Tea consumption ≥3 cups / day is considered a protective factor (OR=0.47, 95%CI: 0.26-0.85), while smoking ≥10 cigarettes / day is considered a risk factor (OR=2.72). (95% CI: 1.21-5.72), directly related to lifestyle interventions, and can also overcome the limitations of static assessment. It dynamically couples environmental exposure (SpO2) and behavioral factors (smoking / tea drinking) to generate personalized risk coefficients (e.g., OR=4.35 for SpO2<83% is the strongest predictor), supporting targeted prevention. In this embodiment of the invention, the performance is optimized through multi-model comparison. The AUCs of logistic regression, XGBoost, and random forest are 0.848, 0.828, and 0.735, respectively. Among them, the logistic regression model has both high interpretability and clinical applicability. The model outputs interpretable risk coefficients, which facilitates medical staff to adjust lifestyle intervention strategies and overcomes the limitations of static assessment. In this embodiment of the invention, the candidate early warning model... The optimal early warning model was selected from the three models (logistic regression, XGBoost, and random forest) to compare their performance. The results showed that the logistic regression model had the highest AUC value, indicating that it has high sensitivity and specificity in identifying high-risk groups for HAPC. Although the random forest and XGBoost models performed similarly on some indicators, the logistic regression model had stronger interpretability, making it easier to apply and promote in clinical practice. In addition, decision curve analysis (DCA) further confirmed the application value of the logistic regression model in clinical decision-making. The embodiments of this invention verified sensitivity ≥85% through the test set, solving the problem of crude privacy mechanisms. The specificity verification used portable devices such as finger clip pulse oximeters, which are suitable for remote high-altitude areas and meet the needs of public health scenarios.Therefore, the method and system for early warning analysis of high-risk groups of polycythemia vera in high altitude areas provided by the embodiments of the present invention can solve the core problem that the early warning results are one-sided and unreliable due to the triple defects of missing data dimensions, limitations of static evaluation, and crude privacy mechanisms. Example 2:
[0069] like Figure 7 The diagram shown is a functional module diagram of an early warning and analysis system for high-risk groups of high-altitude polycythemia according to the present invention.
[0070] The high-risk population early warning analysis system 700 for high-altitude polycythemia, as described in this invention, can be installed in an electronic device. Depending on the functions implemented, the high-risk population early warning analysis system for high-altitude polycythemia may include a data filtering module 701, a factor extraction module 702, a model generation module 703, a system establishment module 704, and an effect verification module 705. The modules described in this invention can also be referred to as units, which are a series of computer program segments that can be executed by the processor of an electronic device and perform a fixed function, stored in the memory of the electronic device.
[0071] In this embodiment of the invention, the functions of each module / unit are as follows: The data filtering module 701 is used to filter subject datasets from extremely high altitude areas, wherein the subject datasets include HAPC group datasets and non-HAPC group datasets. The factor extraction module 702 is used to perform structured processing on the subject dataset to obtain a structured feature matrix, and extract core predictive factors from the structured feature matrix. The model generation module 703 is used to generate candidate early warning models for the core predictive factors, wherein the candidate early warning models include logistic regression models, XGBoost models, and random forest models. The system establishment module 704 is used to select the optimal early warning model from the candidate early warning models, and establish a high-altitude polycythemia early warning system for the subject dataset through the optimal early warning model. The effect verification module 705 is used to verify the effect of the high-altitude polycythemia early warning system, so as to realize the early warning analysis and processing of the high-risk population of high-altitude polycythemia in the subject dataset.
[0072] In detail, the modules in the high-risk population early warning and analysis system 700 for polycythemia vera described in this embodiment of the invention employ the same methods as described above. Figure 1 The same technical means are used to achieve the same early warning analysis method for high-risk groups of polycythemia vera in high altitude areas, and can produce the same technical effect, so it will not be elaborated here.
[0073] It will be apparent to those skilled in the art that the present invention is not limited to the details of the exemplary embodiments described above, and that the present invention can be implemented in other specific forms without departing from the spirit or essential characteristics of the present invention.
[0074] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention and are not intended to limit it. Although the present invention has been described in detail with reference to preferred embodiments, those skilled in the art should understand that modifications or equivalent substitutions can be made to the technical solutions of the present invention without departing from the spirit and scope of the technical solutions of the present invention.
Claims
1. A method for early warning analysis of high-risk groups for polycythemia vera at high altitudes, characterized in that, The method includes: The subject dataset was selected from extremely high altitude areas, including HAPC group datasets and non-HAPC group datasets. The subject dataset is processed in a structured manner to obtain a structured feature matrix, and core predictive factors are extracted from the structured feature matrix. Generate candidate early warning models for the core predictive factors, wherein the candidate early warning models include logistic regression models, XGBoost models, and random forest models; The optimal early warning model is selected from the candidate early warning models, and an early warning system for high-altitude polycythemia of the subject dataset is established using the optimal early warning model. The effectiveness of the high-altitude polycythemia early warning system was verified to achieve early warning analysis and processing of high-risk groups for high-altitude polycythemia in the subject dataset.
2. The method for early warning analysis of high-risk groups for plateau polycythemia as described in claim 1, characterized in that, The selection of subject datasets from extremely high altitude areas includes: Initial respondents were identified in health improvement programs at extremely high altitudes. The initial respondents were categorized into different social classes based on their place of residence, age, and gender. Random sampling was performed on the initial respondents corresponding to each stratum to obtain the random sample population; The subject inclusion operation was performed on the randomly sampled population to obtain the subject dataset; The extremely high altitude areas refer to areas with extreme low oxygen levels at an altitude of ≥4500 meters.
3. The method for early warning analysis of high-risk groups for plateau polycythemia as described in claim 2, characterized in that, The process of enrolling participants in the randomly sampled population to obtain a participant dataset includes: Establish diagnostic criteria for HAPC in the randomly sampled population; The diagnostic criteria for HAPC include a hemoglobin level ≥210 g / L for men and a hemoglobin level ≥190 g / L for women. The randomly sampled population was grouped according to the HAPC diagnostic criteria to obtain the HAPC group and the non-HAPC group. Physiological indicators and questionnaire data were collected from the HAPC group and the non-HAPC group. The subject dataset was determined using the HAPC group population, the non-HAPC group population, the physiological indicators, and the questionnaire data.
4. The method for early warning analysis of high-risk groups for plateau polycythemia as described in claim 1, characterized in that, The process of structuring the subject dataset to obtain a structured feature matrix includes: Extract key indicators from the subject dataset; The key indicators include gender, age, blood oxygen saturation, BMI, smoking amount, tea consumption, blood pressure, and laboratory indicators. The key indicators are combined into a structured feature matrix.
5. The method for early warning analysis of high-risk groups for high-altitude polycythemia as described in claim 1, characterized in that, The extraction of core prediction factors from the structured feature matrix includes: The structured feature matrix is reduced in dimensionality using the LASSO regression method to obtain the reduced-dimensional features; The importance of the dimensionality-reduced features was analyzed by combining the random forest method and the XGBoost method, and the importance analysis results were obtained. Candidate factors are selected from the importance analysis results; The candidate factors were statistically significant using univariate and multivariate analysis methods, and the results of the significance verification were obtained. Based on the significance verification results, core predictive factors directly related to lifestyle interventions are selected from the candidate factors; The core predictive factors include tea consumption, smoking volume, and BMI. Among them, tea consumption is verified as a protective factor when tea consumption is ≥3 cups / day, smoking volume is verified as a risk factor when smoking volume is ≥10 cigarettes / day, and BMI is set to ≥25.7 based on the body fat distribution characteristics in extremely high altitude areas.
6. The method for early warning analysis of high-risk groups for plateau polycythemia as described in claim 1, characterized in that, The candidate early warning model for generating the core predictive factors includes: Extract the training set from the core predictive factors; The preset logistic regression model, preset XGBoost model, and preset random forest model are trained using the training set to obtain the logistic regression model, XGBoost model, and random forest model, respectively. The logistic regression model, the XGBoost model, and the random forest model are used as candidate early warning models.
7. The method for early warning analysis of high-risk groups for high-altitude polycythemia as described in claim 1, characterized in that, The step of selecting the optimal early warning model from the candidate early warning models includes: Calculate the performance metrics of the candidate early warning model; The performance metrics include AUC, accuracy, sensitivity, and specificity; ROC curves and decision curves for the performance metrics are plotted. Based on the ROC curve and the decision curve, the optimal early warning model is selected from the candidate early warning models.
8. The method for early warning analysis of high-risk groups for plateau polycythemia as described in claim 1, characterized in that, The high-altitude polycythemia early warning system established using the optimal early warning model for the subject dataset includes: Obtain the logistic regression coefficients of the optimal early warning model; The maximum absolute value among the logistic regression coefficients is used as the benchmark value; Based on the benchmark value, the logistic regression coefficients are converted into integer scores using the following formula: in, This represents the integer score corresponding to the i-th logistic regression coefficient. This represents the i-th logistic regression coefficient. Indicates the baseline value. Represents the floor function; Set the risk threshold for the integer score; Based on the risk threshold, a risk level corresponding to the integer score is established to create an early warning system for high-altitude polycythemia in the subject dataset.
9. The method for early warning analysis of high-risk groups for plateau polycythemia as described in claim 1, characterized in that, The effectiveness verification of the high-altitude polycythemia early warning system includes: Extract a test set from the core predictive factors; The sensitivity and specificity of the high-altitude polycythemia early warning system were verified using the test set, and the sensitivity and specificity verification results were obtained. The effectiveness verification process of the high-altitude polycythemia early warning system is achieved by using the sensitivity verification results and the specificity verification results. The sensitivity verification results meet the clinical detection rate criteria, which means a clinical detection rate ≥ 85%, and the specificity verification results are determined by a finger-clip pulse oximeter.
10. A system for early warning and analysis of high-risk groups for polycythemia vera at high altitudes, characterized in that, The system includes: A data filtering module is used to filter subject datasets from extremely high altitude areas, wherein the subject datasets include HAPC group datasets and non-HAPC group datasets; The factor extraction module is used to perform structured processing on the subject dataset to obtain a structured feature matrix, and extract core predictive factors from the structured feature matrix. The model generation module is used to generate candidate early warning models for the core predictive factors, wherein the candidate early warning models include logistic regression models, XGBoost models, and random forest models. The system establishment module is used to select the optimal early warning model from the candidate early warning models, and establish an early warning system for high-altitude polycythemia of the subject dataset through the optimal early warning model; The effectiveness verification module is used to verify the effectiveness of the high-altitude polycythemia early warning system, so as to realize the early warning analysis and processing of the high-risk population of high-altitude polycythemia in the subject dataset.
Citation Information
Cited By
Plateau newborn inherited metabolic disease screening system, computer readable storage medium and application of plateau newborn inherited metabolic disease screening system
CN121885165A
A plateau newborn genetic metabolic disease screening system, a computer readable storage medium and an application thereof
CN121885165B