Pre-diabetic exhaled gas biomarker composition and identification method, screening method and device thereof, medium, program product and terminal
Through gas chromatography-mass spectrometry combined technology and multi-stage screening strategy, significant combinations of exhaled gas markers were screened, and combined with the XGBoost model, the cumbersome and specificity of prediagnostic methods in the prior art were solved, and a non-invasive diagnosis with high accuracy and high sensitivity was achieved.
Patent Information
- Application Number
- CN202510666673.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-22
- Publication Date
- 2025-08-08
AI Technical Summary
The existing prediagnostic methods for diabetes are cumbersome and invasive, and have limited sensitivity and specificity in early screening. There is a lack of unified standards for exhaled gas research, limited sample size, and insufficient exploration of multiple marker combinations, resulting in detection signals being easily disturbed by background noise and insufficient specificity.
A gas chromatography-mass spectrometer combined with an adsorption concentration device and a thermal desorption instrument was used to screen out marker gas parameters significantly higher than ambient air. Through a multi-stage screening strategy, including peak area comparison, outlier filtration, correlation analysis and principal component analysis, marker combinations with contribution rates higher than the threshold were extracted, and the XGBoost model was used for diagnosis.
It improves the accuracy and reliability of the detection of exhaled gas markers in prediabetes, achieves a non-invasive diagnosis with high specificity and high sensitivity, and is suitable for large-scale population screening.
Smart Images

Figure CN120446356A_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the technical field of early diagnosis of prediabetes, and in particular to a prediabetes biomarker composition and its identification method, screening method, device, medium, program product and terminal. Background Art
[0002] Prediabetes, a critical transitional stage in the development of diabetes, is reversible and has important clinical intervention value. Approximately 541 million people worldwide are in this stage. Without timely intervention, 5%-18% of patients will progress to diabetes each year. However, existing diagnostic methods face major challenges: traditional oral glucose tolerance tests, fasting blood glucose, and glycated hemoglobin tests are not only cumbersome and invasive, but also have limited sensitivity and specificity in early screening. These methods are difficult to meet the needs of large-scale population screening, especially in areas with scarce medical resources.
[0003] In recent years, exhaled breath analysis has attracted much attention as an emerging non-invasive detection technology. This method reflects the metabolic state of the body by analyzing volatile organic compounds (VOCs) in exhaled breath, and has the advantages of simple operation and strong reproducibility. However, research on exhaled breath in prediabetes has been slow, mainly due to the following technical bottlenecks: First, the metabolic changes in prediabetes are relatively mild, resulting in small changes in the concentration of related VOCs, and the detection signal is easily interfered by background noise; second, the composition of exhaled gas is easily affected by multiple factors such as smoking, diet, and oral flora, resulting in insufficient specificity; third, existing studies lack unified standards for data processing, and the results of different research teams are difficult to verify each other. In addition, the sample size of most studies is limited, and there is a lack of systematic exploration of multi-marker combinations, which seriously restricts the practical application of this technology in prediabetes screening. Summary of the Invention
[0004] In view of the shortcomings of the prior art described above, the purpose of this application is to provide a prediabetes biomarker composition and its identification method, screening method, device, medium, program product, and terminal for large-scale prediabetes screening. A gas chromatography-mass spectrometer with an adsorption concentration device and a thermal desorber is used to address the issues of low concentrations and small fluctuations of relevant VOCs, and the susceptibility of detection signals to background noise. Clear inclusion and exclusion criteria and strict sampling conditions are used to avoid the susceptibility of exhaled gas to influence and the lack of specificity in detection results. Data processing methods are clarified to ensure universality of the results.
[0005] To achieve the above-mentioned objectives and other related objectives, the first aspect of the present application provides a biomarker composition for diagnosing or assisting in the diagnosis of prediabetes, wherein the biomarker composition is selected from any one, multiple or all of the following: undecane, 4-ethyloctane, 2,5-bis(1,1-dimethylethyl)-phenol, 1-methoxy-2-propanone, acetone, and tetradecane.
[0006] To achieve the above-mentioned objectives and other related objectives, the second aspect of the present application provides the use of biomarkers and / or substances for detecting biomarkers in the preparation of products for diagnosing or assisting in the diagnosis of prediabetes, wherein the biomarkers are the biomarker composition as described in the first aspect.
[0007] To achieve the above-mentioned and other related objectives, the present application provides, in a third aspect, a method for diagnosing or assisting in the diagnosis of prediabetes, comprising:
[0008] The concentration of each single item in the biomarker composition as described in the first aspect of the test sample is obtained, and a prediabetes diagnostic model is used to determine whether the patient has prediabetes.
[0009] To achieve the above-mentioned objectives and other related objectives, the fourth aspect of the present application provides a diagnostic or auxiliary diagnostic device for prediabetes, including: a data module: used to obtain the concentration data of each item in the biomarker composition as described in the first aspect of the test sample; a judgment module: used to judge whether it is prediabetes based on the concentration data of each item in the biomarker composition as described in the first aspect of the sample using a prediabetes diagnostic model.
[0010] To achieve the above-mentioned purpose and other related purposes, the fifth aspect of the present application provides a method for screening markers of exhaled gas in prediabetes, comprising: obtaining indoor air parameters and exhaled gas parameters of a subject; wherein the exhaled gas parameters include a plurality of marker gas parameters; performing a peak area comparison operation on the exhaled gas parameters and the indoor air parameters, screening marker gas parameters with a peak area higher than that of indoor air, and performing an outlier filtering operation to generate a first data set; performing a correlation analysis operation on the first data set, and obtaining a second data set by screening through a preset correlation number; performing a principal component analysis operation on the second data set, calculating the contribution rate of each marker feature to all principal components; and extracting markers with a contribution rate higher than a preset threshold.
[0011] To achieve the above-mentioned purpose and other related purposes, the sixth aspect of the present application provides a marker screening device for exhaled gas in prediabetes, including: a data acquisition module: used to obtain indoor air parameters and exhaled gas parameters of a subject; wherein the exhaled gas parameters include multiple marker gas parameters; a marker screening module: used to perform a peak area comparison operation on the exhaled gas parameters and the indoor air parameters, screen marker gas parameters with a peak area higher than that of indoor air, and perform an outlier filtering operation to generate a first data set; perform a correlation analysis operation on the first data set, and obtain a second data set by screening through a preset correlation number; perform a principal component analysis operation on the second data set, calculate the contribution rate of each marker feature to all principal components; and extract markers with a contribution rate higher than a preset threshold; the markers with a contribution rate higher than the preset threshold include the above-mentioned marker combination.
[0012] To achieve the above-mentioned purpose and other related purposes, the seventh aspect of the present application provides a computer-readable storage medium having a computer program stored thereon. When the computer program is executed by a processor, it implements the method for diagnosing or assisting in diagnosing prediabetes, or the method for screening markers of exhaled gas for prediabetes.
[0013] To achieve the above-mentioned objectives and other related objectives, the eighth aspect of the present application provides a computer program product, which includes computer program code. When the computer program code is run on a computer, the computer implements the diagnosis or auxiliary diagnosis method for prediabetes, or the marker screening method for exhaled gas of prediabetes.
[0014] To achieve the above-mentioned objectives and other related objectives, the ninth aspect of the present application provides an electronic terminal comprising a memory, a processor and a computer program stored on the memory; the processor executes the computer program to implement the diagnosis or auxiliary diagnosis method for prediabetes, or the marker screening method for exhaled gas of prediabetes.
[0015] As described above, the present application's exhaled gas biomarker composition for prediabetes and its identification method, screening method, device, medium, program product, and terminal have the following beneficial effects: obtaining the subject's exhaled gas parameters and indoor air parameters by gas chromatography-mass spectrometry; comparing the peak areas of the two to screen out marker gas parameters that are significantly higher than those of ambient air, and performing an outlier filtering operation to generate a first data set; performing correlation analysis and screening to obtain a second data set; calculating the contribution rate of each marker by principal component analysis, and extracting marker combinations with a contribution rate higher than a preset threshold. This method innovatively adopts a multi-level screening strategy, effectively solving key problems in the prior art such as poor marker specificity and insufficient detection standardization, and significantly improving the accuracy and reliability of prediabetes exhaled gas marker detection. Through strict environmental parameter correction and statistical screening, a marker combination with high specificity and sensitivity can be obtained, providing a new technical solution for the non-invasive diagnosis of prediabetes; the exhaled gas sample involved in this application has a non-invasive, simple and rapid collection process. By detecting the concentration of each item in the biomarker composition of the sample, it is possible to accurately determine whether the test subject has prediabetes, with high accuracy, high sensitivity, high specificity, high precision and high F1 Score. BRIEF DESCRIPTION OF THE DRAWINGS
[0016] Figure 1 Shown is a flowchart of a method for diagnosing or assisting in diagnosing prediabetes in one embodiment of the present application.
[0017] Figure 2 Shown is a schematic diagram of a diagnostic or auxiliary diagnostic device for prediabetes in one embodiment of the present application.
[0018] Figure 3 A flow chart showing an embodiment of a method for screening exhaled gas markers for prediabetes according to the present application is shown.
[0019] Figure 4 A schematic structural diagram of an embodiment of a device for screening exhaled gas markers for prediabetes according to the present application is shown.
[0020] Figure 5 A schematic structural diagram of an embodiment of a terminal for screening markers of prediabetic exhaled gas of the present application is shown. DETAILED DESCRIPTION
[0021] The following describes the embodiments of the present application through specific examples. Those skilled in the art can easily understand the other advantages and effects of the present application from the content disclosed in this specification. The present application can also be implemented or applied through other different specific embodiments. The details in this specification can also be modified or changed based on different viewpoints and applications without departing from the spirit of the present application. It should be noted that the following embodiments and features in the embodiments can be combined with each other unless they conflict.
[0022] Before further explaining the present invention in detail, the nouns and terms involved in the embodiments of the present invention are explained. The nouns and terms involved in the embodiments of the present invention are subject to the following interpretations:
[0023] <1> "Diagnosis" generally refers to the process of determining the existence, nature and type of a disease from the patient's symptoms, signs, laboratory test results and other information through various methods and means.
[0024] <2> "Assisted diagnosis" generally refers to the use of various auxiliary tools and methods in conjunction with a physician's clinical judgment during the medical diagnostic process to improve diagnostic accuracy and reliability. The primary purpose of auxiliary diagnosis is to support and supplement a physician's clinical diagnosis, helping to determine the nature, extent, and stage of a disease, thereby guiding subsequent treatment plans.
[0025] <3> "Prediabetes" generally refers to the transitional stage before the onset of diabetes, including impaired fasting glucose (IFG), impaired glucose tolerance (IGT), and a mixed state of the two (IFG + IGT). It is an intermediate hyperglycemic state between normal blood sugar and diabetes. In prediabetes, blood sugar levels are higher than normal but have not yet reached the diagnostic criteria for diabetes.
[0026] Based on the discoveries and research of the present application, the principle of the present application is that prediabetes is a metabolic disorder characterized by insulin resistance or insufficient insulin secretion, which leads to elevated blood sugar levels but does not yet meet the diagnostic criteria for diabetes. Due to the limited effect of insulin, glucose metabolism and fat metabolism in the body become abnormal, resulting in a series of changes in metabolites. These metabolites enter the lungs through the blood and interact with volatile organic compounds in exhaled breath, causing changes in the concentration of certain specific compounds in the exhaled breath. By detecting changes in the concentration of specific volatile organic compounds in the human body's exhaled breath, it is possible to determine whether or not prediabetes exists.
[0027] Therefore, the present application first provides a biomarker composition for diagnosing or assisting in the diagnosis of prediabetes, wherein the biomarker composition is selected from any one or more of the following: undecane, 4-ethyloctane, 2,5-bis(1,1-dimethylethyl)-phenol, 1-methoxy-2-propanone, acetone, and tetradecane, which solves the problems of cumbersome and invasive processes in the early screening of prediabetes in the prior art, and can be used for diagnosing or assisting in the diagnosis of prediabetes, with high accuracy, high sensitivity, high specificity, high precision, and high F1 Score.
[0028] In one embodiment of the present application, the biomarker composition for diagnosing or assisting in the diagnosis of prediabetes includes: undecane, 4-ethyloctane, 2,5-bis(1,1-dimethylethyl)-phenol, 1-methoxy-2-propanone, acetone, and tetradecane.
[0029] Among them, undecane, the molecular formula is C 11 H 24 , CAS number is 1120-21-4; 4-ethyloctane, molecular formula is C 10 H 22 , CAS number is 15869-86-0; 2,5-bis(1,1-dimethylethyl)-phenol, molecular formula is C 14 H 22 O, CAS number is 5875-45-6; 1-methoxy-2-propanone, molecular formula is C4H8O2, CAS number is 5878-19-3; acetone, molecular formula is C3H6O, CAS number is 67-64-1; tetradecane, molecular formula is C 14 H 30 , CAS number is 629-59-4.
[0030] In one embodiment of the present application, the biomarker is derived from the exhaled gas of the test subject, specifically alveolar exhaled gas.
[0031] The present application also provides the use of biomarkers and / or substances for detecting biomarkers in the preparation of products for diagnosing or assisting in the diagnosis of prediabetes, wherein the biomarker is the above-mentioned biomarker composition.
[0032] In one embodiment of the present application, the biomarker can be used as a standard with a known concentration for quantitative detection.
[0033] In one embodiment of the present application, the substance for detecting biomarkers is a substance for detecting the concentration of biomarkers, such as an instrument and / or reagent for detecting the concentration of biomarkers; more specifically, for example, it can be the instrument and / or reagent required for detecting the concentration of biomarkers using gas chromatography-mass spectrometry, which refers to the instruments or reagents used in the process of completing the detection of biomarker concentration using gas chromatography-mass spectrometry, for example, it can include a gas chromatography-mass spectrometer (including a gas chromatography part and a mass spectrometry part), an automatic sampler, a high-purity carrier gas system, a data workstation, etc., as well as a chromatographic column, high-purity carrier gas, standard gas, cleaning reagents, etc.
[0034] In one embodiment of the present application, the biomarker detection apparatus may further include instruments and / or reagents required for pre-processing. For example, the pre-processing may include steps such as collection, adsorption concentration, and desorption. The biomarker detection apparatus may further include the following: an exhaled gas collection device, a pre-concentration device, an adsorption tube, a thermal desorber, and the required reagents for performing pre-processing steps such as gas collection, adsorption concentration, and desorption.
[0035] In one embodiment of the present application, the test sample of the product is the exhaled gas of the test subject, specifically the alveolar exhaled gas.
[0036] For example, in one embodiment of the present application, a method for detecting the concentration of exhaled gas in prediabetes is specifically as follows: collecting an exhaled gas sample from a test subject through a CO2 control-alveolar exhaled gas collection device, transferring the exhaled gas in the air bag to an adsorption tube filled with Tenax TA adsorption material for pre-concentration, placing the adsorption tube with concentrated exhaled gas components in a thermal desorber for heating to desorb the exhaled gas components, and delivering the desorbed exhaled gas components to a gas chromatograph-mass spectrometer with a DB624 UI chromatographic column to detect the concentration of each marker (compound) in a combination of exhaled gas biomarkers that can diagnose prediabetes, wherein the compounds are undecane, 4-ethyloctane, 2,5-bis(1,1-dimethylethyl)-phenol, 1-methoxy-2-propanone, acetone, and tetradecane. Then, the substances for detecting the above-mentioned biomarkers can be the instruments or reagents used in this method.
[0037] In certain embodiments of the present application, the product may be a system, a kit, an instrument, a chip, or the like.
[0038] Based on the use of the above-mentioned biomarkers and / or substances for detecting biomarkers in the preparation of products for diagnosing or assisting in the diagnosis of prediabetes, the present application can also provide a product for diagnosing or assisting in the diagnosis of prediabetes, which includes the biomarkers and / or substances for detecting biomarkers in the above-mentioned uses.
[0039] Based on the use of the above-mentioned biomarkers and / or substances for detecting biomarkers in the preparation of products for diagnosing or assisting in the diagnosis of prediabetes, the present application also provides a method for diagnosing or assisting in the diagnosis of prediabetes. Figure 1 Detailed description. Figure 1 A flowchart of a method for diagnosing or assisting in diagnosing prediabetes according to an embodiment of the present invention is shown, including:
[0040] The concentration data of each single item in the above-mentioned biomarker composition of the test sample is obtained, and a prediabetes diagnostic model is used to determine whether the patient has prediabetes.
[0041] The concentration data is the concentration of each single compound in the biomarker composition in the exhaled breath.
[0042] In one embodiment of the present application, in the method of obtaining concentration data of each individual item in the above-mentioned biomarker composition of the test sample, the test sample is an exhaled breath sample. Specifically, the exhaled breath sample of the test subject is collected, pre-treated by adsorption concentration, desorption, etc., and then gas chromatography-mass spectrometry is used to obtain the concentration of each individual item of the above-mentioned compound in the exhaled breath sample of the test subject. For example, the above-mentioned method for detecting the concentration of exhaled breath in prediabetes is used to obtain the concentration data of each individual biomarker in the above-mentioned biomarker composition.
[0043] In one embodiment of the present application, strict sampling conditions as described in the embodiment are used to exclude interference from smoking, diet, oral bacteria, chemicals, detergents, drugs, etc., to avoid the problem that exhaled gas is easily affected and the test results are not specific enough.
[0044] In one embodiment of the present application, a gas chromatography-mass spectrometer equipped with an adsorption concentration device and a thermal desorber is used to detect the concentration of each individual item in the biomarker combination in the present application, thereby solving the problems of low concentration of relevant VOCs, small variation range, and detection signal susceptibility to background noise interference.
[0045] In one embodiment of the present application, the prediabetes diagnostic model is obtained by training with the concentration data of each individual item in the biomarker composition of known samples, wherein the known samples used for training adopt the inclusion and exclusion criteria as clearly defined in the embodiment to avoid the problem that exhaled gas is easily affected and the test results are not specific enough.
[0046] In one embodiment of the present application, the obtained prediabetes diagnostic model may also be subjected to a model performance test, and the samples used for the test are in accordance with the inclusion and exclusion criteria specified in the embodiment.
[0047] The prediabetes diagnosis model is an XGBoost model, which is trained using the following command line to obtain the prediabetes diagnosis model.
[0048] The use of the prediabetes diagnostic model involves substituting concentration data into the prediabetes diagnostic model to determine whether the test subject has prediabetes. The XGBoost model can accurately determine whether the test subject has prediabetes with high accuracy, high sensitivity, high specificity, high precision, and high F1 Score, solving the problems of cumbersome, invasive, and limited specificity in the early screening of prediabetes in the prior art.
[0049] xgb_modelXGBClassifier(
[0050] colsample_bytree=0.8,
[0051] gamma=0.2,
[0052] learning_rate=0.01,
[0053] max_depth=5,
[0054] n_estimators=200,
[0055] subsample=0.8,
[0056] eval_metric = 'logloss',
[0057] random_state=72 )
[0059] xgb_model.fit(X_train,y_train), where xgb_model represents the XGBoost model, X_train represents the training set data, and y_train represents the training set labels. The dataset is divided into 20 test sets and 70 training sets. Use train_test_split to split the dataset and ensure that the class distribution is consistent: X_train,X_test,y_train,y_test=train_test_split(X,y,test_size=20,train_size=70,stratify=y,random_state=72)
[0060] In one embodiment of the present application, all steps of the method for diagnosing or assisting in diagnosing prediabetes are implemented by a computer.
[0061] In one embodiment of the present application, the diagnosis or auxiliary diagnosis method of prediabetes is for non-disease diagnosis and treatment purposes, for example, it can be used for early screening, risk assessment, health management, etc.
[0062] This application also provides a diagnostic or auxiliary diagnostic device for prediabetes.
[0063] Figure 2 This is a schematic block diagram of a pre-diabetes diagnosis or auxiliary diagnosis device provided in an embodiment of the present application. Figure 2 As shown, the device includes a data module 201 and a judgment module 202.
[0064] The data module 201 is used to obtain the concentration data of each individual item in the above-mentioned biomarker composition of the test sample; the judgment module 202 is used to determine whether the sample is prediabetes based on the concentration data of each individual item in the above-mentioned biomarker composition of the sample using the prediabetes diagnostic model.
[0065] It should be understood that the specific process of each module executing the above corresponding steps has been described in detail in the above method embodiment, and for the sake of brevity, it will not be repeated here.
[0066] In one embodiment of the present application, the detection objects of the above-mentioned products, methods, and devices can be mammals, such as but not limited to humans, primates, livestock (such as sheep, cattle, horses, donkeys, pigs), pets (such as dogs, cats), laboratory test animals (such as mice, rabbits, rats, guinea pigs, hamsters) or captured wild animals (such as foxes, deer); preferably, the detection object is a primate; more preferably, the detection object is a human.
[0067] To facilitate understanding of the embodiments of this application, first Figure 3 Detailed description. Figure 3 The flowchart of a method for screening exhaled breath markers for prediabetes according to an embodiment of the present invention is shown. The method for screening exhaled breath markers for prediabetes according to this embodiment mainly includes the following steps:
[0068] Step S31: Acquire indoor air parameters and exhaled gas parameters of the subject; wherein the exhaled gas parameters include parameters of multiple marker gases;
[0069] In one embodiment of the present application, a gas chromatography-mass spectrometry system (GC-MS) equipped with a DB624 UI chromatographic column is used to detect and analyze gas samples. The detection process includes: comparing the collected gas sample mass spectrometry data with the mass spectrometry database of the National Institute of Standards and Technology (NIST14), and determining that the compound is effectively identified when the matching degree is ≥80%. After detection and analysis, 67 volatile organic compounds (VOCs) were successfully identified in more than 80% of the samples. The system automatically classifies and stores the identification results, and generates structured exhaled gas parameter data sets and indoor air parameter data sets, respectively, to provide a data basis for subsequent analysis.
[0070] Step S32: performing a peak area comparison operation on the exhaled gas parameter and the indoor air parameter, screening marker gas parameters having a peak area higher than that of the indoor air, and performing an outlier filtering operation to generate a first data set;
[0071] Preferably, marker gas parameters with peak areas greater than twice that of indoor air were screened. This screening process yielded a total of 63 VOCs. To further enhance data reliability, the screening results were corrected for environmental background by subtracting the peak area responses of the corresponding compounds in the ambient air sample. This generated a primary dataset that effectively eliminated environmental interference factors and contained multiple specific marker signatures.
[0072] In this example, the outlier filtering process involves filling missing values in the dataset using a second-order polynomial interpolation method to ensure data integrity and avoid analytical bias caused by missing data. The outlier removal process involves calculating the coefficient of variation and outlier ratio for each compound, and removing compounds with a coefficient of variation greater than 0.3 or an outlier ratio exceeding 10% to eliminate anomalous data caused by experimental error and individual differences, thereby improving analytical reliability. By way of example, this example screened 58 compounds through outlier removal. The natural logarithm transformation process involves applying an ln(x) transformation to the peak area data of all retained compounds to reduce the order of magnitude difference between high- and low-concentration compounds, thereby ensuring that the data distribution is more consistent with statistical analysis requirements. This example establishes a baseline data collection reference system by collecting parameters from indoor air and exhaled breath, reducing environmental interference. Peak area comparison and outlier filtering then form an objective preliminary screening mechanism, making the entire data processing process clear and universal.
[0073] Step S33: performing a correlation analysis operation on the first data set, and obtaining a second data set containing multiple marker features by screening using a preset correlation number;
[0074] In one embodiment of the present application, the process of performing a correlation analysis operation on the first data set includes: performing a correlation analysis on the first data set using the Pearson correlation number method, and calculating the degree of correlation between each category of markers and their corresponding P values. By setting a significance level threshold (P<0.05), characteristic compounds with statistically significant correlation are screened out. For example, as shown in Table 1, a total of 9 compounds that meet the significance requirements are screened out in this embodiment, constituting a second data set containing multiple marker characteristics. This embodiment provides a clear operating procedure for the feature selection process by performing correlation analysis and setting a clear correlation threshold standard, thereby enhancing the standardization and repeatability of the method for extracting a prediabetes biomarker composition, thereby improving the feasibility and practical value of its application in large-scale population screening.
[0075]
[0076]
[0077] Table 1: P values for chromatographic peak areas of exhaled breath compounds between prediabetic patients and healthy subjects
[0078] Step S34: performing principal component analysis on the second data set, calculating the contribution rate of each marker feature to all principal components; and extracting markers whose contribution rate is higher than a preset threshold.
[0079] In one embodiment of the present application, principal component analysis was used to reduce the dimensionality of the second dataset to reduce redundant information and extract compounds with a contribution rate of 95% as the final diagnostic markers. In this embodiment, the diagnostic markers included the following six combinations: undecane, 4-ethyloctane, 2,5-bis(1,1-dimethylethyl)phenol, 1-methoxy-2-propanone, acetone, and tetradecane.
[0080] In one embodiment of the present application, the random forest algorithm is used to calculate the contribution rate of each of the 9 compounds obtained by the above screening, and the calculation process also includes a consistency verification operation: first, the contribution rate of each marker in the second data set is calculated based on the random forest algorithm, and the contribution rate is sorted from high to low. Subsequently, the step-by-step addition method is used to input the sorted markers into the exhaled gas detection model in sequence, and the model parameters are updated after each input and the updated classification accuracy is calculated. When the classification accuracy reaches a stable state and no longer increases, the key marker combination at this time is recorded. Finally, the combination is compared with the marker category screened out by the principal component analysis. If the category is consistent, the validity verification is determined to be passed. This verification operation ensures that the screened marker combination has reliable classification performance. This embodiment uses the contribution rate ranking in the principal component analysis to evaluate the importance of the marker by mathematical methods, so that the data processing results are objective and universal.
[0081] Compounds can be added to the random forest algorithm in the following order to verify compound consistency: undecane, 4-ethyloctane, 2,5-bis(1,1-dimethylethyl)phenol, 1-methoxy-2-propanone, acetone, tetradecane, 3-methylheptane, 4,6-dimethylundecane, and 3-ethyl-3-methylheptane. The consistency verification result is consistent with the composition type with a contribution rate greater than 95% obtained after principal component analysis, indicating that the current consistency verification has passed.
[0082] The method for screening exhaled breath markers for prediabetes in this example provides a standardized data processing workflow, helping to address the issue of inconsistent data processing standards in prediabetes biomarker composition research. This clear processing workflow reduces subjective factors in the research process, improves the reproducibility of data processing, and provides a standardized method for reference by different research teams, making research results more comparable and facilitating large-scale population screening for prediabetes biomarkers.
[0083] In the embodiments of this application, terms such as "first" and "second" are used to distinguish between identical or similar items with substantially the same function or effect. For example, the first and second data sets are merely used to distinguish different data sets and do not define their order. Those skilled in the art will understand that terms such as "first" and "second" do not define the quantity or execution order, and do not necessarily imply differences.
[0084] It should be noted that in the embodiments of this application, words such as "exemplary" or "for example" represent examples, illustrations, or descriptions. Any embodiment or design described in this application as "exemplary" or "for example" should not be interpreted as being preferred or advantageous over other embodiments or designs. Rather, the use of words such as "exemplary" or "for example" is intended to present the relevant concepts in a concrete manner.
[0085] In the embodiments of the present application, "at least one" refers to one or more, and "more" refers to two or more. "And / or" describes the association relationship of associated objects, indicating that three relationships may exist. For example, A and / or B can represent: the existence of A alone, the existence of A and B at the same time, and the existence of B alone, where A and B can be singular or plural. The character " / " generally indicates that the previous and next associated objects are in an "or" relationship. "At least one of the following items" or similar expressions refers to any combination of these items, including any combination of single items or plural items. For example, at least one of a, b or c can represent: a, b, c, ab, ac, bc or abc, where a, b, c can be single or multiple.
[0086] Figure 4 Schematic diagram of the marker screening device 400 for prediabetes exhaled gas provided in the embodiment of the present application. Figure 4 As shown, the device includes a data acquisition module 401 and a marker screening module 402.
[0087] Data acquisition module 401: used to acquire indoor air parameters and exhaled gas parameters of the subject; wherein the exhaled gas parameters include multiple marker gas parameters.
[0088] Marker screening module 402: used to perform a peak area comparison operation on the exhaled gas parameters and the indoor air parameters, screen marker gas parameters with a peak area higher than that of indoor air, and perform an outlier filtering operation to generate a first data set; perform a correlation analysis operation on the first data set, and obtain a second data set by screening through a preset correlation number; perform a principal component analysis operation on the second data set, calculate the contribution rate of each marker feature to all principal components; and extract markers with a contribution rate higher than a preset threshold; the markers with a contribution rate higher than the preset threshold include the biomarker combination as described in claim 1.
[0089] It should be understood that the specific process of each module executing the above corresponding steps has been described in detail in the above method embodiment, and for the sake of brevity, it will not be repeated here.
[0090] It should also be understood that the division of modules in the embodiments of the present application is illustrative and is merely a logical functional division. In actual implementation, other division methods may be used. Furthermore, the functional modules in the various embodiments of the present application may be integrated into a single processor, or may exist physically separately, or two or more modules may be integrated into a single module. The aforementioned integrated modules may be implemented in the form of hardware or software functional modules.
[0091] Figure 5: is a schematic block diagram of an electronic terminal provided in an embodiment of the present application. Figure 5 As shown, the electronic terminal includes: at least one processor 501, a memory 502, at least one network interface 503 and a user interface 505. The various components in the device are coupled together via a bus system 504. It is understood that the bus system 504 is used to achieve connection and communication between these components. In addition to including a data bus, the bus system 504 also includes a power bus, a control bus and a status signal bus. However, for the sake of clarity, Figure 5 Various buses are labeled as bus systems.
[0092] The user interface 505 may include a display, a keyboard, a mouse, a trackball, a click gun, keys, buttons, a touch pad or a touch screen.
[0093] It will be appreciated that the memory 502 may be a volatile memory or a non-volatile memory, or may include both volatile and non-volatile memories. Among them, the non-volatile memory may be a read-only memory (ROM) or a programmable read-only memory (PROM), which is used as an external cache. By way of example but not limitation, many forms of RAM are available, such as static random access memory (SRAM) and synchronous static random access memory (SSRAM). The memory described in the embodiments of the present invention is intended to include, but is not limited to, these and any other suitable types of memory.
[0094] The memory 502 in the embodiment of the present invention is used to store various categories of data to support the operation of the electronic terminal 500. Examples of these data include: any executable program for operating on the electronic terminal 500, such as an operating system 5021 and an application 5022; the operating system 5021 includes various system programs, such as a framework layer, a core library layer, a driver layer, etc., for implementing various basic services and processing hardware-based tasks. The application 5022 can include various applications, such as a media player (Media Player), a browser (Browser), etc., for implementing various application services. The method for diagnosing or assisting diagnosis of prediabetes or the method for screening markers of exhaled gas in prediabetes provided in the embodiment of the present invention can be included in the application 5022.
[0095] The methods disclosed in the above embodiments of the present invention can be applied to or implemented by processor 501. Processor 501 may be an integrated circuit chip with signal processing capabilities. During implementation, each step of the above method can be completed by hardware integrated logic circuits in processor 501 or by software instructions. The above processor 501 may be a general-purpose processor, a digital signal processor (DSP), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. Processor 501 can implement or execute the various methods, steps, and logic block diagrams disclosed in the embodiments of the present invention. The general-purpose processor 501 may be a microprocessor or any conventional processor. The steps of the accessory optimization method provided in the embodiments of the present invention can be directly implemented and executed by a hardware decoding processor, or by a combination of hardware and software modules in the decoding processor. The software module can be located in a storage medium located in a memory. The processor reads the information in the memory and, in conjunction with its hardware, completes the steps of the above method.
[0096] In an exemplary embodiment, the electronic terminal 500 may be configured to execute the aforementioned method using one or more application specific integrated circuits (ASICs), DSPs, programmable logic devices (PLDs), or complex programmable logic devices (CPLDs).
[0097] According to the method provided in the embodiments of the present application, the present application also provides a computer program product, which includes: computer program code, which, when the computer program code is run on a computer, enables the computer to execute the diagnosis or auxiliary diagnosis method for prediabetes, or the marker screening method for exhaled gas of prediabetes, as described in any of the embodiments shown above.
[0098] According to the method provided in the embodiments of the present application, the present application also provides a computer-readable storage medium, which stores program code. When the program code is run on a computer, the computer executes the diagnosis or auxiliary diagnosis method for prediabetes, or the marker screening method for exhaled gas of prediabetes, as described in any of the embodiments shown above.
[0099] As used in this specification, the terms "component," "module," "system," and the like are used to represent computer-related entities, hardware, firmware, a combination of hardware and software, software, or software in execution. For example, a component can be, but is not limited to, a process running on a processor, a processor, an object, an executable file, an execution thread, a program, and / or a computer. By way of illustration, both an application running on a computing device and a computing device can be a component. One or more components can reside in a process and / or an execution thread, and a component can be located on a computer and / or distributed between two or more computers. In addition, these components can be executed from various computer-readable media having various data structures stored thereon. Components can communicate, for example, via local and / or remote processes based on signals having one or more data packets (e.g., data from two components interacting with another component on a local system, a distributed system, and / or a network, such as the Internet interacting with other systems via signals).
[0100] Those skilled in the art will appreciate that the various illustrative logical blocks and steps described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, or a combination of computer software and electronic hardware. Whether these functions are performed in hardware or software depends on the specific application and design constraints of the technical solution. Professional and technical personnel can use different methods to implement the described functions for each specific application, but such implementation should not be considered beyond the scope of this application.
[0101] Those skilled in the art will clearly understand that, for the convenience and brevity of description, the specific working processes of the systems, devices and units described above can refer to the corresponding processes in the aforementioned method embodiments and will not be repeated here.
[0102] In the several embodiments provided in this application, it should be understood that the disclosed systems, devices and methods can be implemented in other ways. For example, the device embodiments described above are merely schematic. For example, the division of units is only a logical function division. In actual implementation, there may be other division methods, such as multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the mutual coupling or direct coupling or communication connection shown or discussed can be through some interfaces, indirect coupling or communication connection of devices or units, which can be electrical, mechanical or other forms.
[0103] Units described as separate components may or may not be physically separate, and components shown as units may or may not be physical units, that is, they may be located in one place or distributed across multiple network units. Some or all of these units may be selected to achieve the purpose of this embodiment according to actual needs.
[0104] In addition, each functional unit in each embodiment of the present application may be integrated into one processing unit, or each unit may exist physically separately, or two or more units may be integrated into one unit.
[0105] In the above embodiments, the functions of each functional unit can be implemented in whole or in part by software, hardware, firmware or any combination thereof. When software is used for implementation, it can be implemented in whole or in part in the form of a computer program product. The computer program product includes one or more computer instructions (programs). When the computer program instructions (program) are loaded and executed on a computer, the process or function according to the embodiment of the present application is generated in whole or in part. The computer can be a general-purpose computer, a special-purpose computer, a computer network or other programmable device. The computer instructions can be stored in a computer-readable storage medium or transmitted from one computer-readable storage medium to another computer-readable storage medium. For example, the computer instructions can be transmitted from a website, computer, server or data center to another website, computer, server or data center by wired (such as coaxial cable, optical fiber, digital subscriber line (DSL)) or wireless (such as infrared, wireless, microwave, etc.) mode. The computer-readable storage medium can be any available medium that a computer can access or a data storage device such as a server or data center that includes one or more available media integrations. Available media can be magnetic media (e.g., floppy disks, hard disks, tapes), optical media (e.g., high-density digital video discs (DVDs), or semiconductor media (e.g., solid-state drives (SSDs)).
[0106] If the function is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present application, or the part that contributes to the prior art, or the part of the technical solution, can be embodied in the form of a software product. The computer software product is stored in a storage medium and includes several instructions for enabling a computer device (which can be a personal computer, server, or network device, etc.) to execute all or part of the steps of the various embodiments of the present application. The aforementioned storage medium includes: various media that can store program codes, such as a USB flash drive, a mobile hard disk, a read-only memory (ROM), a random access memory (RAM), a magnetic disk or an optical disk.
[0107] The present application is described below through specific embodiments.
[0108] Inclusion and exclusion criteria for samples used in the examples:
[0109] 1. Inclusion criteria for patients with prediabetes
[0110] (1) After diagnosis by a doctor, the patient meets the diagnostic criteria for prediabetes published by the American Diabetes Association in 2023: impaired fasting glucose (IFG): 5.6mmol / L≤IFG<7mmol / L; or impaired two-hour postprandial glucose (IGT): 7.8mmol / L≤IGT≤11.0mmol / L; or both IFG and ITG exist.
[0111] (2) age range was 18 to 80 years old, regardless of gender;
[0112] (3) Must have certain cognitive abilities, be in good physiological condition, and be able to independently collect breath samples;
[0113] (4) Voluntarily participated in this study and signed the informed consent form.
[0114] 2. Inclusion criteria for healthy subjects
[0115] (1) Normal fasting blood glucose level (<5.6 mmol / L) and no symptoms related to diabetes or prediabetes;
[0116] (2) age range was 18 to 80 years old, regardless of gender;
[0117] (3) Have good cognitive ability and physiological status, and be able to independently collect breath samples;
[0118] (4) Voluntarily participated in this study and signed the informed consent form.
[0119] 3. Exclusion criteria
[0120] (1) Those who do not meet the diagnostic criteria for prediabetes or healthy control group mentioned above;
[0121] (2) Those who engage in behaviors that affect exhaled air, such as long-term smoking, alcoholism, or taking drugs;
[0122] (3) Pregnant or breastfeeding women;
[0123] (4) Suffering from endocrine and metabolic diseases such as hyperthyroidism that may affect blood sugar;
[0124] (5) Suffering from serious heart, liver, kidney, tumor and other diseases;
[0125] (6) Those who suffer from mental illness, lack self-awareness, and cannot express themselves effectively;
[0126] (7) Those who are unwilling or unable to cooperate with the experiment.
[0127] Example 1 Exhaled Gas Collection and Detection
[0128] 1. Exhaled gas collection
[0129] The alveolar exhaled gas of 120 subjects (60 healthy people and 60 pre-diabetic patients) was collected using a CO2 control-alveolar exhaled gas collection device. The total amount of air exhaled by the human body each time is about 500 ml, which is divided into two parts. The first 150 ml of exhaled air comes from the upper respiratory tract, also known as "dead space air". The remaining 350 ml of exhaled air comes from the alveoli, called "alveolar gas", which is produced after exchange with the blood through the pulmonary circulation and is considered to be the headspace air in the blood. Due to the special properties of alveolar gas, it is regarded as a target component for disease diagnosis.
[0130] All subjects completed the collection process under guidance, and all breath samples were collected in a dedicated clinic. The clinic was free of any chemicals, cleaning agents, or medications that could affect the samples, minimizing environmental interference. To ensure quality, all samples were collected between 8:00 and 10:30 a.m. Prior to collection, subjects were required to fast for at least 8 hours and avoid smoking, drinking, and consuming spicy or irritating foods for 24 hours; strenuous exercise should be avoided within two hours of sampling. Furthermore, participants were asked to avoid using perfume before sampling to minimize potential interference. Subjects rinsed their mouths with purified water before exhaled breath was collected. These preparations effectively ensured the accuracy and reliability of the breath samples. Subjects exhaled calmly, and the gas was collected into a Tedlar bag using a CO2-controlled alveolar exhaled gas collection device. Specifically, this application utilizes the Fowler model, using a CO2 concentration in the exhaled gas that is higher than 50% (C50) of the CO2 concentration in the alveolar gas to distinguish between anatomical dead space exhaled gas and alveolar exhaled gas. When the exhaled CO2 concentration is below C50, the exhaled breath collection device does not collect this portion of exhaled breath. When the exhaled CO2 concentration is above C50, the exhaled breath collection device collects this portion of exhaled breath into a Teflon bag with a collection volume of 3 L. All sample bags are stored at room temperature (22-26°C), and sample pretreatment is completed within 6 hours.
[0131] 2. Exhaled gas detection
[0132] Exhaled breath was transferred to a Tenax TA adsorbent tube for preconcentration and then detected by gas chromatography-mass spectrometry using thermal desorption injection. Chromatographic conditions included a DB624 UI column; a carrier gas flow rate of 1.0 mL / min in constant flow mode; and an oven temperature of 40°C for 5 minutes, then increased to 200°C at a rate of 10°C min-1 and held for 10 minutes. The mass spectrometer operated in scan mode over the m / z range of 35 to 390 amu. Exhaled breath compounds were identified based on retention times from the NIST14 database and reference standards, and quantification was performed using peak area.
[0133] Example 2 Construction and verification of prediction model
[0134] 1. Screening of markers
[0135] (1) We screened all volatile organic compounds (VOCs) whose median peak areas in breath samples were more than twice the corresponding peak areas in indoor air, and selected 63 compounds. We then subtracted the peak areas of the corresponding compounds in ambient air samples to reduce the impact of environmental factors.
[0136] (2) Second-order polynomial interpolation method is used to fill missing values to ensure data integrity.
[0137] (3) The coefficient of variation and outlier ratio of each compound were calculated. Compounds with a coefficient of variation greater than 0.3 or an outlier ratio greater than 10% were screened and deleted to reduce the impact of individual differences on the experimental results. A total of 58 compounds were screened.
[0138] (4) The peak area data of all compounds were transformed into natural logarithms to reduce the large differences between different volatile organic compounds and optimize the data distribution.
[0139] (5) Pearson correlation coefficients were used to calculate P values, and characteristic compounds with P values less than 0.05 were screened for significant correlation. A total of 9 compounds were screened. The P values of the chromatographic peak areas of the exhaled breath compounds between prediabetic patients and healthy subjects are shown in Table 1 above.
[0140] (6) The “feature importance scores” of these nine compounds were calculated using the random forest algorithm and ranked from high to low in importance: undecane, 4-ethyloctane, 2,5-bis(1,1-dimethylethyl)-phenol, 1-methoxy-2-propanone, acetone, tetradecane, 3-methylheptane, 4,6-dimethylundecane, and 3-ethyl-3-methylheptane.
[0141] (7) Principal component analysis was used to reduce data dimensionality, reduce redundant information, and select the number of principal components with a cumulative contribution rate of 95%. The six compounds finally screened out happened to be the top six substances in the feature importance ranking: undecane, 4-ethyloctane, 2,5-bis(1,1-dimethylethyl)-phenol, 1-methoxy-2-propanone, acetone, and tetradecane.
[0142] (8) The classification effect of the data after principal component analysis dimensionality reduction was evaluated using the random forest algorithm and the support vector machine algorithm. Features were added gradually in descending order of importance. When the number of features increased to 6, the classification accuracy of the model was the highest.
[0143] 2. Model training and validation
[0144] An XGBoost model was constructed based on a combination of six compounds: undecane, 4-ethyloctane, 2,5-bis(1,1-dimethylethyl)phenol, 1-methoxy-2-propanone, acetone, and tetradecane for evaluation:
[0145] The 120 samples collected were divided into a training set (90 samples) and an independent validation set (30 samples) at a ratio of 3:1. The training set was further divided into a training set of 70 samples and a test set of 20 samples at a ratio of 7:2 for model training and preliminary validation.
[0146] Stratified sampling was used to partition the data to ensure that the distribution of categories in the training set and the test set was consistent. The machine learning model was built using XGBoost:
[0147] (1) Using the XGBoost model to learn the training set, we obtained a prediabetes diagnosis model:
[0148] Use the following command line to train:
[0149] xgb_modelXGBClassifier(
[0150] colsample_bytree=0.8,
[0151] gamma=0.2,
[0152] learning_rate=0.01,
[0153] max_depth=5,
[0154] n_estimators=200,
[0155] subsample=0.8,
[0156] eval_metric = 'logloss',
[0157] random_state=72 )
[0159] xgb_model.fit(X_train,y_train), where xgb_model represents the XGBoost model, X_train represents the training set data, and y_train represents the training set labels. The dataset is divided into 20 test sets and 70 training sets. Use train_test_split to split the dataset and ensure that the class distribution is consistent: X_train,X_test,y_train,y_test=train_test_split(X,y,test_size=20,train_size=70,stratify=y,random_state=72)
[0160] Subsequently, the obtained model was evaluated, and its evaluation indicators were accuracy, sensitivity, specificity, and AUC value. Accuracy: The proportion of correctly identified prediabetic subjects and healthy people among all subjects, indicating the proportion of correct predictions of the model as a whole. Sensitivity: The proportion of correctly identified prediabetic subjects among all prediabetic subjects, which reflects the ability to identify true positives. Specificity: The proportion of correctly identified healthy people among all healthy people, which indicates the ability to identify true negatives. AUC: reflects the ability of the model to distinguish between positive and negative samples under various thresholds. The AUC value represents the area under the ROC curve. The larger the area under the ROC curve, the stronger the classification ability of the model. Specifically, the closer the AUC value is to 1, the more accurately the model can distinguish between prediabetic subjects and healthy people.
[0161] (2) Substitute the test set into the prediabetes diagnosis model for evaluation.
[0162] Use the trained model to predict the test set and calculate evaluation indicators, including accuracy, sensitivity, specificity, precision, F1 Score, ROC, AUC value, etc. The specific commands are as follows:
[0163] y_test_pred=xgb_model.predict(X_test)
[0164] conf_matrix=confusion_matrix(y_test,y_test_pred)
[0165] Accuracy=accuracy_score(y_test,y_test_pred)
[0166] Precision=precision_score(y_test,y_test_pred)
[0167] f1=f1_score(y_test,y_test_pred)
[0168] roc_auc=roc_auc_score(y_test,xgb_model.predict_proba(X_test)[:,1])
[0169] sensitivity,specificity = compute_sensitivity_specificity(conf_matrix), where conf_matrix represents the confusion matrix, y_test_pred represents the predicted result, and y_test represents the true label. Calculate sensitivity and specificity based on the confusion matrix.
[0170] (3) The external validation set was substituted into the prediabetes diagnosis model for evaluation.
[0171] Substitute the XGBoost model parameters into a new dataset to predict and evaluate the model's performance, and calculate the accuracy, sensitivity, specificity, precision, F1 Score, ROC, AUC value, etc. The specific commands are as follows:
[0172] y_pred_new=best_xgb_model.predict(X_new)
[0173] conf_matrix_new=confusion_matrix(y_new,y_pred_new)
[0174] accuracy_new=accuracy_score(y_new,y_pred_new)
[0175] precision_new=precision_score(y_new,y_pred_new)
[0176] f1_new=f1_score(y_new,y_pred_new)
[0177] roc_auc_new=roc_auc_score(y_new,best_xgb_model.predict_proba(X_new)[:,1])
[0178] sensitivity_new=conf_matrix_new[1,1] / (conf_matrix_new[1,1]+conf_matrix_new[1,0])
[0179] specificity_new=conf_matrix_new[0,0] / (conf_matrix_new[0,0]+conf_matrix_new[0,1])
[0180] The evaluation results of the above-mentioned prediabetes diagnostic model are shown in Table 2, which shows that the diagnostic model constructed by the biomarkers screened in this application has high accuracy, sensitivity, specificity and precision in identifying prediabetes.
[0181] Table 2 Results of XGBoost model classification of healthy people and prediabetic patients
[0182]
[0183] Accuracy represents the proportion of prediabetic and healthy subjects correctly identified among all subjects, indicating the proportion of correct predictions made by the model overall. Sensitivity represents the proportion of prediabetic subjects correctly identified among all prediabetic subjects, reflecting its ability to identify true positives. Specificity represents the proportion of healthy subjects correctly identified among all healthy subjects, reflecting its ability to identify true negatives. AUC reflects the model's ability to distinguish between positive and negative samples at various thresholds. The AUC value represents the area under the receiver operating characteristic (ROC) curve. A larger area under the ROC curve indicates a stronger classification ability. Specifically, an AUC value closer to 1 indicates that the model can more accurately distinguish between prediabetic and healthy subjects. Precision represents the proportion of subjects identified as prediabetic by the model who actually have prediabetes. The F1 score is the harmonic mean of precision and sensitivity, calculated as shown in Formula 1. A higher AUC value indicates a better overall model performance in accurately identifying positive samples.
[0184]
[0185] The above description is merely a specific embodiment of the present application, but the scope of protection of the present application is not limited thereto. Any changes or substitutions that can be easily conceived by a person skilled in the art within the technical scope disclosed in this application should be included in the scope of protection of the present application. Therefore, the scope of protection of the present application should be based on the scope of protection of the claims.
[0186] In summary, the present application provides a composition of exhaled breath biomarkers for prediabetes and its identification method, screening method, device, medium, program product and terminal, which obtains indoor air parameters and exhaled gas parameters of subjects by gas chromatography-mass spectrometry; compares the peak areas of the two, screens out marker gas parameters that are significantly higher than those of ambient air, and performs outlier filtering operations to generate a first data set; performs correlation analysis and screening to obtain a second data set; calculates the contribution rate of each marker by principal component analysis, and extracts marker combinations with a contribution rate higher than a preset threshold. This method innovatively adopts a multi-level screening strategy, which effectively solves the key problems of poor marker specificity and insufficient detection standardization in the prior art, and significantly improves the accuracy and reliability of exhaled breath marker detection for prediabetes. Through strict environmental parameter correction and statistical screening, a biomarker combination with high specificity and sensitivity can be obtained, which provides a new technical solution for non-invasive screening of prediabetes. By detecting the concentration of each item in the biomarker combination of the sample, it is possible to accurately identify whether the detection object is prediabetes. The identification method has high accuracy, high sensitivity, high specificity, high precision and high F1Score. Therefore, the present invention effectively overcomes various shortcomings of the prior art and has high industrial utilization value.
[0187] The above embodiments are merely illustrative of the principles and effects of this application and are not intended to limit this application. Anyone skilled in the art may modify or alter the above embodiments without departing from the spirit and scope of this application. Therefore, all equivalent modifications or alterations made by one of ordinary skill in the art without departing from the spirit and technical concepts disclosed in this application shall be covered by the claims of this application.
Claims
1. A biomarker composition for diagnosing or assisting in the diagnosis of prediabetes, characterized in that: The biomarker composition is selected from any one, multiple or all of the following: undecane, 4-ethyloctane, 2,5-bis(1,1-dimethylethyl)-phenol, 1-methoxy-2-propanone, acetone, and tetradecane.
2. Use of biomarkers and / or substances for detecting biomarkers in the preparation of products for diagnosing or assisting in the diagnosis of prediabetes, characterized in that: The biomarker is the biomarker composition according to claim 1; Preferably, the substance for detecting biomarkers is an instrument and / or reagent for detecting biomarker concentrations.
3. A method for diagnosing or assisting diagnosis of prediabetes, characterized in that: include: The concentration data of each single item in the biomarker composition according to claim 1 is obtained for the test sample, and a prediabetes diagnostic model is used to determine whether the patient is in the prediabetes stage.
4. The method according to claim 3, wherein The prediabetes diagnostic model is obtained by training with concentration data of each single item in the biomarker composition of known samples; preferably, the prediabetes diagnostic model is an XGBoost model.
5. A diagnostic or auxiliary diagnostic device for prediabetes, characterized in that: include: Data module: used to obtain concentration data of each single item in the biomarker composition according to claim 1 for the test sample; A judgment module is used to judge whether a patient has prediabetes based on the concentration data of each item in the biomarker composition according to claim 1 of the sample and using a prediabetes diagnostic model.
6. The device according to claim 5, characterized in that The prediabetes diagnostic model is obtained by training with concentration data of each single item in the biomarker composition of known samples; preferably, the prediabetes diagnostic model is an XGBoost model.
7. A method for screening exhaled breath markers for prediabetes, characterized in that: include: Acquiring indoor air parameters and exhaled gas parameters of the subject; wherein the exhaled gas parameters include multiple marker gas parameters; performing a peak area comparison operation on the exhaled gas parameter and the indoor air parameter, screening marker gas parameters having a peak area greater than that of the indoor air, and performing an outlier filtering operation to generate a first data set; Performing a correlation analysis operation on the first data set, and filtering the data using a preset correlation number to obtain a second data set; Performing a principal component analysis operation on the second data set to calculate the contribution rate of each marker feature to all principal components; sorting the markers of each category from high to low according to the contribution rate, and extracting markers with a contribution rate higher than a preset threshold.
8. A device for screening markers of prediabetic exhaled breath, characterized in that: include: Data acquisition module: used to obtain indoor air parameters and exhaled gas parameters of the subject; wherein the exhaled gas parameters include multiple marker gas parameters; A marker screening module: used to perform a peak area comparison operation on the exhaled gas parameters and the indoor air parameters, screen marker gas parameters with a peak area higher than that of indoor air, and perform an outlier filtering operation to generate a first data set; perform a correlation analysis operation on the first data set, and obtain a second data set by screening through a preset correlation number; perform a principal component analysis operation on the second data set, calculate the contribution rate of each marker feature to all principal components; and extract markers with a contribution rate higher than a preset threshold; the markers with a contribution rate higher than the preset threshold include the marker combination as described in claim 1.
9. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the method for diagnosing or assisting in diagnosing prediabetes according to claim 3 or 4, or the method for screening exhaled breath markers for prediabetes according to any one of claim 8 is implemented.
10. A computer program product, characterized in that The computer program product includes computer program code, and when the computer program code is run on a computer, the computer implements the method for diagnosing or assisting diagnosis of prediabetes as described in claim 3 or 4, or the method for screening markers in exhaled gas of prediabetes as described in any one of claim 8.
11. An electronic terminal comprising a memory, a processor, and a computer program stored in the memory, characterized in that: The processor executes the computer program to implement the method for diagnosing or assisting diagnosis of prediabetes according to claim 3 or 4, or the method for screening exhaled breath markers for prediabetes according to any one of claim 8.