Screening method, device and terminal equipment for disease breath markers

By constructing a sample and VOCs similarity network and combining exhaled gas characteristics and expression profiles, respiratory biomarkers related to diseases were screened out, which solved the problem of low reliability in existing technologies and achieved stronger disease biomarker discovery and metabolic pathway elucidation.

CN115575523BActive Publication Date: 2026-03-27CHANGSHA UNIVERSITY
View PDF 2 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-09-14
Publication Date
2026-03-27

AI Technical Summary

Technical Problem

Existing technologies, when using statistical analysis to identify VOCs with significant differences in exhaled gases as disease biomarkers, do not consider the impact of detection methods on respiratory expression profiles and the physiological and pathological significance of VOCs, resulting in low reliability of screening results and a lack of clinical trial value.

Method used

We constructed a sample similarity network and a VOCs similarity network, combined the characteristic sequences of exhaled gases and the expression profiles of VOCs, and screened out target respiratory biomarkers related to diseases through clustering and label propagation algorithms, taking into account the relationship between different pathogenic mechanisms of diseases and biomarker combinations.

Benefits of technology

This improved the reliability of respiratory biomarker screening, revealed the metabolic pathways of disease development, discovered multiple combinations of respiratory biomarkers, and enhanced the reliability and clinical application value of the biomarkers.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115575523B_ABST
    Figure CN115575523B_ABST
Patent Text Reader

Abstract

The application is suitable for the technical field of computer application, and provides a screening method for disease respiratory markers, comprising: determining a sample similarity network corresponding to a sample data set according to the similarity between the collection and detection parameters of each exhaled gas and the information of the corresponding sampling object, then determining a VOCs expression profile corresponding to VOCs contained in each exhaled gas according to the respiratory fingerprint of each exhaled gas, then determining a candidate respiratory marker included in the VOCs and a VOCs similarity network corresponding to the candidate respiratory marker according to each VOCs expression profile, and finally screening a target respiratory marker related to a disease according to the sample similarity network and the VOCs similarity network. Thus, by constructing the sample similarity network and the VOCs similarity network, the experimental conditions and the difference information of the sampling objects between samples are introduced, and the physiological and pathological correlation information between VOCs is introduced, so that the accuracy of the screening of the disease respiratory markers is improved.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application belongs to the technical field of computer application, and particularly relates to a disease respiratory marker screening method and device and a terminal device. BACKGROUND

[0002] Respiratory diagnosis technology reflects the physiological and pathological state of a subject by collecting and detecting exhaled air, is a non-invasive, safe and simple method for observing in-vivo biological metabolism and physiological processes, and is considered as a promising early disease screening technology. In the past few decades, it has been a research hotspot in the medical field worldwide and is increasingly widely applied in medical practice. Volatile organic compounds (VOCs) in human exhaled air carry a large amount of information about the health status of the human body and can be used as ideal markers for malignant tumors and respiratory diseases. However, there is currently no certain respiratory marker for most diseases.

[0003] Currently, the related art generally performs a "case-control" test, that is, exhaled air of a disease patient and a healthy subject is collected respectively, and the types and contents of VOCs contained in the exhaled air samples of the two groups of people are detected, and then statistical analysis is performed to find VOCs that have significant differences between the two groups of people as respiratory markers of the disease. However, pure statistical analysis does not consider the influence of the detection method on the respiratory expression profile and the physiological and pathological significance of VOCs in exhaled air, thereby resulting in low reliability of the screening result and often lacking clinical test value. SUMMARY

[0004] The disease respiratory marker screening method and device and the terminal device provided by the embodiments of the present application can solve the problem that pure statistical analysis does not consider the influence of the detection method on the respiratory expression profile and cannot give the physiological and pathological significance of VOCs in exhaled air, thereby resulting in low reliability of the screening result and lacking clinical test value.

[0005] In a first aspect, an embodiment of the present application provides a disease respiratory marker screening method, comprising: obtaining a sample data set corresponding to exhaled gas, wherein the sample data set comprises a plurality of respiratory fingerprints corresponding to exhaled gas, sampling parameters, and sampling object information corresponding to exhaled gas; determining a feature sequence corresponding to each exhaled gas according to the sampling parameters corresponding to each exhaled gas and the sampling object information corresponding to each exhaled gas; determining a similarity network corresponding to the sample data set according to the similarity between the feature sequences corresponding to each exhaled gas; determining an expression profile corresponding to volatile organic compounds (VOCs) contained in each exhaled gas according to the respiratory fingerprints corresponding to each exhaled gas; determining a candidate respiratory marker included in the VOCs and a VOCs similarity network corresponding to the candidate respiratory marker according to the expression profile corresponding to each VOC; and screening a target respiratory marker related to a disease from the candidate respiratory marker according to the respiratory fingerprints corresponding to each exhaled gas, each candidate respiratory marker, the sample similarity network, and the VOCs similarity network.

[0006] In a possible implementation form of the first aspect, the disease comprises a plurality of disease subtypes, and the determination of the candidate respiratory marker included in the VOCs and the VOCs similarity network corresponding to the candidate respiratory marker according to the expression profile corresponding to each VOC comprises:

[0007] determining an enzyme network corresponding to each VOC according to the metabolic pathways related to each VOC and the enzyme list related to each metabolic pathway;

[0008] performing clustering on each VOC according to the expression profile corresponding to each VOC and the enzyme network to determine a candidate respiratory marker included in the VOCs and a clustering center corresponding to each candidate respiratory marker and a disease subtype, wherein each clustering center corresponds to a disease subtype;

[0009] determining a VOCs similarity network corresponding to the candidate respiratory marker according to the clustering center corresponding to each candidate respiratory marker;

[0010] Accordingly, the screening of the target respiratory marker related to the disease from the candidate respiratory marker according to the respiratory fingerprints corresponding to each exhaled gas, each candidate respiratory marker, the sample similarity network, and the VOCs similarity network comprises:

[0011] screening a target respiratory marker related to each disease subtype from the candidate respiratory marker according to the respiratory fingerprints corresponding to each exhaled gas, each candidate respiratory marker, the sample similarity network, and the VOCs similarity network.

[0012] Optionally, in a further possible implementation manner of the first aspect, the sampling object information comprises a real disease type, and the screening of the target respiratory biomarker related to each disease type from the candidate respiratory biomarkers according to the respiratory fingerprint corresponding to each exhaled gas, the sample similarity network, and the VOCs similarity network comprises:

[0013] determining a relationship matrix between the sample data set and the candidate respiratory biomarkers according to the candidate respiratory biomarkers contained in each exhaled gas;

[0014] dividing each exhaled gas in the sample data set into a training sample and a prediction sample;

[0015] respectively determining a first effectiveness vector corresponding to each training sample according to the real disease type corresponding to each training sample, the disease type corresponding to each candidate respiratory biomarker, and the serial number of each training sample in the sample data set, wherein the first effectiveness vector is used to represent each candidate respiratory biomarker related to the real disease type corresponding to the training sample and the serial number of the training sample in the sample data set;

[0016] respectively determining an initial effectiveness vector corresponding to each prediction sample according to the serial number of each prediction sample in the sample data set;

[0017] processing the sample similarity network, the VOCs similarity network, the relationship matrix, each first effectiveness vector, and each initial effectiveness vector by using a preset label propagation algorithm, to determine a prediction effectiveness vector corresponding to each prediction sample, wherein the prediction effectiveness vector is used to represent each candidate respiratory biomarker related to the prediction disease type of the prediction sample and the serial number of the training sample in the sample data set;

[0018] determining a prediction disease type corresponding to each prediction sample according to the prediction effectiveness vector corresponding to each prediction sample;

[0019] determining the target respiratory biomarker related to each disease type from the candidate respiratory biomarkers according to the matching degree between the prediction disease type corresponding to each prediction sample and the real disease type.

[0020] Optionally, in a further possible implementation manner of the first aspect, the clustering of each VOCs according to the expression profile corresponding to each VOCs and the enzyme network, to determine the candidate respiratory biomarker included in the VOCs, and the cluster center corresponding to each candidate respiratory biomarker and the disease type, comprises:

[0021] inputting the expression profile corresponding to each VOCs into the preset self-organizing neural network and the enzyme network to determine the expression profile similarity between each VOCs and each cluster center in the preset self-organizing neural network, and the metabolic similarity between each VOCs and each cluster center;

[0022] determining the distance between each VOCs and each cluster center according to the expression profile similarity between each VOCs and each cluster center in the preset self-organizing neural network, the metabolic similarity between each VOCs and each cluster center, and a preset weight;

[0023] determining the candidate breath marker included in the VOCs, and the cluster center corresponding to each candidate breath marker and the disease classification according to the distance between each VOCs and each cluster center.

[0024] Optionally, in a further possible implementation manner of the first aspect, the VOCs include M VOCs, and the preset self-organizing neural network includes N cluster centers, where M and N are positive integers; accordingly, the inputting the expression profile corresponding to each VOCs into the preset self-organizing neural network and the enzyme network to determine the expression profile similarity between each VOCs and each cluster center in the preset self-organizing neural network, and the metabolic similarity between each VOCs and each cluster center includes:

[0025] determining the expression profile similarity between the ith VOCs and the jth cluster center according to the distance between the expression profile corresponding to the ith VOCs and the expression profile corresponding to the jth cluster center, where i is a positive integer greater than or equal to 1 and less than or equal to M, and j is a positive integer greater than or equal to 1 and less than or equal to N;

[0026] fusing the enzyme network corresponding to the ith VOCs and the enzyme network corresponding to the jth cluster center to generate a fused enzyme network;

[0027] determining the metabolic similarity between the ith VOCs and the jth cluster center according to the difference between the enzyme network corresponding to the jth cluster center and the fused enzyme network.

[0028] Optionally, in a further possible implementation manner of the first aspect, the sampling parameters include at least one of the following parameters: an experimental control condition, a gas sampling instrument parameter, a thermal desorption parameter, a chromatographic parameter, and a mass spectrometric parameter; and the sampling object information includes at least one of the following information: personal basic information, living environment and habit information, and medical information.

[0029] Optionally, in a further possible implementation form of the first aspect, the feature sequence corresponding to each exhaled gas comprises a state sequence and a value sequence, and the feature sequence corresponding to each exhaled gas is determined according to the sampling parameter corresponding to each exhaled gas and the sampling object information corresponding to each exhaled gas, comprising:

[0030] determining the numerical data and the non-numerical data in the sampling parameter and the sampling object information corresponding to each exhaled gas;

[0031] normalizing the numerical data corresponding to each exhaled gas to determine the value sequence corresponding to each exhaled gas;

[0032] mapping the state of the non-numerical data corresponding to each exhaled gas to determine the state sequence corresponding to each exhaled gas.

[0033] Optionally, in a further possible implementation form of the first aspect, before determining the sample similarity network corresponding to the sample data set according to the similarity between the feature sequences corresponding to the exhaled gases, the method further comprises:

[0034] determining the similarity between the value sequences corresponding to the exhaled gases and the similarity between the state sequences corresponding to the exhaled gases;

[0035] determining the similarity between the feature sequences corresponding to the exhaled gases according to the similarity between the value sequences corresponding to the exhaled gases and the similarity between the state sequences corresponding to the exhaled gases.

[0036] In a second aspect, an embodiment of the present application provides a disease respiratory marker screening device, comprising: a first obtaining module configured to obtain a sample data set corresponding to exhaled gas, wherein the sample data set comprises a plurality of respiratory fingerprints corresponding to exhaled gas, sampling parameters, and sampling object information corresponding to exhaled gas; a first determining module configured to determine a feature sequence corresponding to each exhaled gas according to the sampling parameter corresponding to each exhaled gas and the sampling object information corresponding to each exhaled gas; a second determining module configured to determine a similarity network corresponding to the sample data set according to the similarity between the feature sequences corresponding to the exhaled gases; a third determining module configured to determine an expression profile of VOCs contained in each exhaled gas according to the respiratory fingerprint corresponding to each exhaled gas; a fourth determining module configured to determine a candidate respiratory marker included in the VOCs and a VOC similarity network corresponding to the candidate respiratory marker according to the expression profile corresponding to each VOC; and a first screening module configured to screen a target respiratory marker related to a disease from the candidate respiratory marker according to the respiratory fingerprint corresponding to each exhaled gas, each candidate respiratory marker, the sample similarity network, and the VOC similarity network.

[0037] Optionally, in a possible implementation manner of the second aspect, the diseases include a plurality of disease types, and the fourth determining module includes:

[0038] The first determining unit is configured to determine an enzyme network corresponding to each VOC according to each VOC-related metabolic pathway and each metabolic pathway-related enzyme list;

[0039] The second determining unit is configured to cluster each VOC according to the expression profile and the enzyme network corresponding to each VOC, to determine candidate breath markers included in the VOCs, and a clustering center corresponding to each candidate breath marker and a disease type, wherein each clustering center corresponds to one disease type.

[0040] The third determining unit is configured to determine a VOC similarity network corresponding to each candidate breath marker according to the clustering center corresponding to each candidate breath marker.

[0041] Correspondingly, the first screening module includes:

[0042] The first screening unit is configured to screen target breath markers related to each disease type from the candidate breath markers according to the breath fingerprint corresponding to each exhaled gas, each candidate breath marker, the sample similarity network, and the VOC similarity network.

[0043] Optionally, in another possible implementation manner of the second aspect, the first screening unit is specifically configured to:

[0044] determine a relationship matrix between the sample data set and the candidate breath markers according to the candidate breath markers included in each exhaled gas;

[0045] divide each exhaled gas in the sample data set into a training sample and a prediction sample;

[0046] determine a first effectiveness vector corresponding to each training sample according to the real disease type corresponding to each training sample, the disease type corresponding to each candidate breath marker, and the serial number of each training sample in the sample data set, wherein the first effectiveness vector is used to represent each candidate breath marker related to the real disease type corresponding to the training sample and the serial number of the training sample in the sample data set.

[0047] determine an initial effectiveness vector corresponding to each prediction sample according to the serial number of each prediction sample in the sample data set;

[0048] The preset label propagation algorithm is used to process the sample similarity network, the VOCs similarity network, the relationship matrix, each first effectiveness vector and each initial effectiveness vector, so as to determine a predicted effectiveness vector corresponding to each prediction sample, wherein the predicted effectiveness vector is used to represent each candidate respiratory marker and a serial number of the training sample in the sample data and the VOCs which are related to a predicted disease classification of the prediction sample;

[0049] According to the predicted effectiveness vector corresponding to each prediction sample, a predicted disease classification corresponding to each prediction sample is determined.

[0050] According to a matching degree between the predicted disease classification corresponding to each prediction sample and the true disease classification, a target respiratory marker related to each disease classification is determined from the candidate respiratory markers.

[0051] Optionally, in another possible implementation manner of the second aspect, the second determining unit is specifically configured to:

[0052] The expression profile corresponding to each VOCs is input into the preset self-organizing neural network, so as to determine an expression profile similarity between each VOCs and each cluster center in the preset self-organizing neural network, and a metabolic similarity between each VOCs and each cluster center;

[0053] According to the expression profile similarity between each VOCs and each cluster center in the preset self-organizing neural network, the metabolic similarity between each VOCs and each cluster center and a preset weight, a distance between each VOCs and each cluster center is determined.

[0054] According to the distance between each VOCs and each cluster center, a candidate respiratory marker included in the VOCs and a cluster center corresponding to each candidate respiratory marker and a disease classification are determined.

[0055] Optionally, in another possible implementation manner of the second aspect, the VOCs include M VOCs, the preset self-organizing neural network includes N cluster centers, and M and N are positive integers; correspondingly, the second determining unit is further configured to:

[0056] According to the distance between the expression profile corresponding to the i th VOCs and the expression profile corresponding to the j th cluster center, an expression profile similarity between the i th VOCs and the j th cluster center is determined, wherein i is a positive integer greater than or equal to 1 and less than or equal to M, and j is a positive integer greater than or equal to 1 and less than or equal to N;

[0057] The enzyme network corresponding to the i th VOCs and the enzyme network corresponding to the j th cluster center are fused to generate a fused enzyme network;

[0058] According to the difference between the enzyme network corresponding to the jth cluster center and the fused enzyme network, metabolic similarity between the ith VOC and the jth cluster center is determined.

[0059] Optionally, in a possible implementation of the second aspect, the sampling parameters include at least one of the following: experimental control conditions, gas sampling instrument parameters, thermal desorption parameters, chromatographic parameters, and mass spectrometry parameters; and the sampling object information includes at least one of the following: personal basic information, living environment and habit information, and medical information.

[0060] Optionally, in a possible implementation of the second aspect, the first determining module includes:

[0061] The fourth determining unit is configured to determine numerical data and non-numerical data in the sampling parameters and the sampling object information corresponding to each exhaled gas;

[0062] The fifth determining unit is configured to perform normalization processing on the numerical data corresponding to each exhaled gas to determine a value sequence corresponding to each exhaled gas;

[0063] The sixth determining unit is configured to perform state mapping on the non-numerical data corresponding to each exhaled gas to determine a state sequence corresponding to each exhaled gas.

[0064] Optionally, in a possible implementation of the second aspect, the apparatus further includes:

[0065] The fifth determining module is configured to determine similarity between the value sequences corresponding to the exhaled gases and similarity between the state sequences corresponding to the exhaled gases;

[0066] The sixth determining module is configured to determine similarity between the feature sequences corresponding to the exhaled gases according to the similarity between the value sequences corresponding to the exhaled gases and the similarity between the state sequences corresponding to the exhaled gases.

[0067] In a third aspect, an embodiment of the present application provides a terminal device, including a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor implements the disease respiratory marker screening method as described above when executing the computer program.

[0068] In a fourth aspect, an embodiment of the present application provides a computer readable storage medium having a computer program stored thereon, wherein the computer program is executable by a processor to implement the disease respiratory marker screening method as described above.

[0069] In a fifth aspect, an embodiment of the present application provides a computer program product, which, when running on a terminal device, causes the terminal device to perform the disease respiratory marker screening method according to any one of the first aspect.

[0070] Compared with the prior art, the embodiment of the present application has the beneficial effect that: by combining the sample similarity network and the VOCs similarity network to construct a heterogeneous network system conforming to the actual clinical process, the relationship between different pathogenic mechanisms of diseases and marker combinations is found, thereby overcoming the limitations and sample dependence of a pure statistical method, and it is beneficial to find various respiratory marker combinations of diseases and reveal metabolic pathways closely related to the occurrence and development of diseases, and the discovered markers are more reliable. BRIEF DESCRIPTION OF DRAWINGS

[0071] In order to more clearly illustrate the technical solutions in the embodiments of the present application, the drawings needed to be used in the embodiments or prior art description will be briefly introduced. Obviously, the drawings in the following description are only some embodiments of the present application, and other drawings can be obtained by those skilled in the art without creative labor.

[0072] Figure 1 is a flowchart of the disease respiratory marker screening method provided by an embodiment of the present application;

[0073] Figure 2 is a flowchart of the disease respiratory marker screening method provided by another embodiment of the present application;

[0074] Figure 3 is a structural schematic diagram of the disease respiratory marker screening device provided by an embodiment of the present application;

[0075] Figure 4 is a structural schematic diagram of the terminal device provided by an embodiment of the present application. DETAILED DESCRIPTION

[0076] In the following description, specific details are set forth in order to provide a thorough understanding of the embodiments of the present application. However, persons skilled in the art should understand that the present application can also be implemented in other embodiments without these specific details. In other cases, detailed descriptions of well-known systems, devices, circuits, and methods are omitted so as not to obscure the description of the present application with unnecessary details.

[0077] It will be understood that the term “includes,” “comprises, ” “comprising,” “has,” “contains” and / or “including,” when used in the specification and / or the claims of this application, specifies the presence of stated features, integers, steps, operations, elements, and / or components but do not preclude the presence or addition of one or more other features, integers, steps, operations, elements, components, and / or groups thereof.

[0078] It should also be understood that the term “and / or” as used herein refers to and encompasses any and all possible combinations of one or more of the associated listed items, and that the term “includes” or “comprising” means “including, but not limited to.”

[0079] As used in the description of the application and the appended claims, the term “if’ can be interpreted to mean “when” or “upon” or “in response to determining” or “in response to detecting” depending on the context. Similarly, the phrase “if it is determined” or “if [a described condition or event] is detected” can be interpreted to mean “upon determining” or “in response to determining” or “upon [the described condition or event] being detected” or “in response to [the described condition or event] being detected,” depending on the context.

[0080] In addition, the terms “first,” “second,” “third,” etc. as used in the description of embodiments herein and throughout the claims (if any) are not used to connote any relative importance but are used differently from their meanings in the background art, which are used merely to distinguish one element from another.

[0081] Reference to “one embodiment” or “some embodiments” or “one implementation” or “some implementations” etc. in the description herein means that a particular feature, structure, or characteristic described in connection with the embodiment is included in at least one embodiment of the application. The appearances of the phrases “in one embodiment” or “in some embodiments” or “in other embodiments” or “in still other embodiments” or other similar phrases in various places in the specification are not necessarily all referring to the same embodiment, nor are they necessarily referring to some, but not all, embodiments. The terms “including,” “comprising,” “having” and variations thereof herein are meant to encompass the item listed thereafter but do not exclude additional, unrecited items. Unless otherwise indicated herein, the use of relational terms and / or adjectives such as “by way of illustration,” “exemplification,” “example,” and / or “embodiment” are used herein to help illustrate the application and are not intended to be interpreted as limiting.

[0082] A method for screening a disease respiratory marker, an apparatus, a terminal device, a storage medium and a computer program are provided in the application.

[0083] Figure 1 A flowchart of a method for screening a disease respiratory marker is shown.

[0084] At step 101, a sample data set corresponding to exhaled gas is obtained, wherein the sample data set includes a plurality of respiratory fingerprints corresponding to exhaled gas, sampling parameters, and sampling object information corresponding to the exhaled gas.

[0085] It should be noted that the disease respiratory marker screening method of the embodiments of the present application can be executed by the disease respiratory marker screening device of the embodiments of the present application. The disease respiratory marker screening device of the embodiments of the present application can be configured in any terminal device to execute the disease respiratory marker screening method of the embodiments of the present application. For example, the device of the embodiments of the present application can be configured in any terminal device that can realize the screening of disease respiratory markers, such as a respiratory detection device, and the like, which is not limited in the embodiments of the present application.

[0086] The respiratory fingerprint can be a spectrum indicating the types of VOCs contained in the exhaled gas and the corresponding concentrations of each type of VOCs.

[0087] The sampling parameters can include experimental condition parameters, environmental condition parameters, and the like.

[0088] The sampling object information can be personal information of the sampling object. In actual use, the sampling object can include a diseased population and a healthy population, so as to take the difference between the exhaled gas of the diseased population and the healthy population as one of the factors for screening the disease respiratory marker.

[0089] In the embodiments of the present application, the exhaled gas of the real sampling object can be collected, and each collected exhaled gas can be detected to obtain the respiratory fingerprint corresponding to each exhaled gas, and the sampling parameters and the sampling object information when each exhaled gas is collected can be recorded to construct each sample data in the sample data set; or, relevant literatures can be obtained from the related patent and paper database, and the respiratory fingerprint of the exhaled gas, the sampling parameters, and the sampling object information corresponding to the exhaled gas involved in the literatures can be extracted as the sample data in the sample data set, so as to improve the scale and data richness of the sample data set, and make the subsequent disease respiratory marker screening more accurate.

[0090] It should be noted that when the sample data set is obtained, the sampling object can be selected according to the type of disease to be studied. For example, when the disease respiratory marker screening method of the embodiments of the present application is applied to the screening of lung cancer respiratory markers, a certain amount of lung cancer patients and healthy people can be selected as the sampling object.

[0091] In step 102, the characteristic sequence corresponding to each exhaled gas is determined according to the sampling parameters corresponding to each exhaled gas and the sampling object information corresponding to each exhaled gas.

[0092] The characteristic sequence corresponding to the exhaled gas can be used to represent the characteristics of the sampling parameters and the sampling object information corresponding to the exhaled gas.

[0093] In the embodiments of the present application, the sampling parameters corresponding to the exhaled gas and the sampling object information can be combined to generate the feature sequence corresponding to the exhaled gas; or the sampling parameters corresponding to the exhaled gas and the sampling object information can also be normalized, and then the normalized sampling parameters and the sampling object information are combined to generate the feature sequence of the exhaled gas.

[0094] In step 103, the sample similarity network corresponding to the sample data set is determined according to the similarity between the feature sequences corresponding to each exhaled gas.

[0095] The sample similarity network can include the similarity between the feature sequences corresponding to any two exhaled gases, that is, the sample similarity network can be a symmetric matrix, and each element in the sample similarity network can represent the similarity between the feature sequences corresponding to any two exhaled gases.

[0096] In the embodiments of the present application, the feature sequence corresponding to the exhaled gas can be divided into two parts, that is, the feature sequence of the sampling parameters corresponding to the exhaled gas and the feature sequence of the sampling object information corresponding to the exhaled gas; then for any two exhaled gases, the first similarity between the feature sequences of the sampling parameters corresponding to the two exhaled gases and the second similarity between the feature sequences of the sampling object information corresponding to the two exhaled gases can be determined first, and then the weighted sum of the first similarity and the second similarity can be determined as the similarity between the feature sequences corresponding to the two exhaled gases. After determining the similarity between all exhaled gases in the sample data set, the similarity between any two exhaled gases can be filled into the corresponding position in the sample similarity network according to the serial numbers of the exhaled gases to generate the sample similarity network corresponding to the sample data set.

[0097] For example, there are m samples in the sample data set, and the sample similarity network is an m x m symmetric matrix, and the element in the i-th row and j-th column of the sample similarity network is the similarity between the feature sequences corresponding to the i-th exhaled gas and the j-th exhaled gas in the sample data set, where i and j are positive integers greater than or equal to 1 and less than or equal to m.

[0098] It should be noted that in actual use, the weight used for weighting the first similarity and the second similarity can be determined according to actual needs and specific application scenarios, and the embodiments of the present application do not limit this.

[0099] In step 104, the expression profile of VOCs contained in each exhaled gas is determined according to the breath fingerprint corresponding to each exhaled gas.

[0100] The VOCs can refer to the VOCs involved in the breath fingerprints of all exhaled gases in the sample data set.

[0101] The expression profile corresponding to VOCs can be an expression profile constituted by the content of the VOCs contained in each breathprint. For example, the sample data set contains 1000 breathprints corresponding to exhaled gases, and 3000 VOCs are involved in the 1000 breathprints. For any one of the 3000 VOCs, the content of the VOCs in the 1000 breathprints can be extracted to constitute the expression profile corresponding to the VOCs, that is, the expression profile corresponding to each VOC can be a 1000-dimensional vector exhaled gas.

[0102] As a possible implementation, the breathprints corresponding to each exhaled gas can be processed by using a spectrum recognition software to generate the expression profile corresponding to each VOC.

[0103] In step 105, the candidate breath markers included in the VOCs and the VOCs similarity network corresponding to the candidate breath markers are determined according to the expression profile corresponding to each VOC.

[0104] The VOCs similarity network can include the similarity between any two VOCs, that is, the VOCs similarity network can be a symmetric matrix, and each element in the VOCs similarity network can represent the similarity between any two VOCs.

[0105] For example, n VOCs are involved in each breathprint in the sample data set, the VOCs similarity network is an n*n symmetric matrix, and the element in the ith row and jth column of the VOCs similarity network is the similarity between the ith VOC and the jth VOC, where i and j are positive integers greater than or equal to 1 and less than or equal to n.

[0106] In the embodiments of the present application, the similarity between each VOC can be determined according to the expression profile corresponding to each VOC, and the VOCs similarity network corresponding to the candidate breath markers can be constructed according to the similarity between each VOC.

[0107] In the embodiments of the present application, all VOCs can be used as candidate breath markers, or part of the VOCs can be selected as candidate breath markers according to the expression profile corresponding to the VOCs.

[0108] Further, the disease includes multiple disease types, and the VOCs can be clustered and screened by clustering to determine the disease types related to the VOCs and the VOCs related to each disease type. That is, in a possible implementation of the embodiments of the present application, the above step 105 can include:

[0109] determine an enzyme network corresponding to each VOC according to each VOC-related metabolic pathway and each metabolic pathway-related enzyme list;

[0110] cluster each VOC according to its corresponding expression profile and enzyme network to determine candidate breath markers included in the VOCs and a cluster center corresponding to each candidate breath marker and a disease type, wherein each cluster center corresponds to a disease type;

[0111] determine a VOC similarity network corresponding to each candidate breath marker according to the cluster center corresponding to each candidate breath marker.

[0112] The VOC-related metabolic pathway can be a human metabolic pathway related to the VOC in a metabolic process.

[0113] The enzyme network corresponding to the VOC can be generated according to various enzymes related to the metabolic process of the VOC and biochemical connections between the various enzymes.

[0114] In the embodiments of the present application, the metabolic pathways related to each VOC can be obtained according to the compound data in the Kyoto Encyclopedia of Genes and Genomes (KEGG) database, and then the metabolic pathway set corresponding to each VOC can be constructed using the metabolic pathways related to each VOC. However, in some metabolic pathways, the metabolic process of the VOC in the human body, such as the conversion and generation of the VOC, can not be involved, so these metabolic pathways cannot represent the correlation between the VOC and the human metabolism, and thus these metabolic pathways can be removed, and only the metabolic pathways related to the metabolic process of the VOC in the human body can be retained. Therefore, for a VOC, the enzyme-related metabolic pathways can be screened from the metabolic pathway set corresponding to the VOC according to the enzyme list related to each metabolic pathway corresponding to the VOC, and then the biochemical connections between any two enzymes can be generated according to the number of each enzyme-related metabolic pathway, and then the enzyme network corresponding to the VOC can be generated according to the biochemical connections between the various enzymes related to the VOC, so that the metabolic characteristics corresponding to the VOC can be represented by the enzyme network corresponding to the VOC.

[0115] In the embodiments of the present application, the artificial neural network can be used to cluster VOCs according to the enzyme network corresponding to the VOCs and the expression profile corresponding to the VOCs, and in the clustering process, according to the enzyme network corresponding to each VOC and the expression profile, a plurality of cluster centers that can represent different disease subtypes can be automatically generated, and the cluster center corresponding to each VOC is determined, and then the disease subtype related to the cluster center corresponding to the VOC is determined as the disease subtype corresponding to the VOC.

[0116] As a possible implementation, if a certain VOC fails to be divided into any cluster center in the clustering process, the VOC can be removed, and all VOCs with corresponding cluster centers are determined as candidate respiratory markers.

[0117] As another possible implementation, after clustering each VOC and determining the disease subtype corresponding to each VOC, the gene transcription pathway related to the target gene of each disease subtype can be obtained from the KEGG database, and enrichment analysis can be performed on each gene transcription pathway to determine whether there is an association between each gene transcription pathway and the metabolic pathway corresponding to each VOC. If the metabolic pathway corresponding to a VOC is associated with the gene transcription pathway, it can be determined that the VOC is indeed associated with the related disease subtype, and the VOC can be retained and the weight of the cluster center corresponding to the VOC can be increased for the next clustering. If the metabolic pathway corresponding to a VOC is not associated with all gene transcription pathways, it can be determined that the VOC may not be related to the disease, and thus the VOC can be removed or the weight of the cluster center corresponding to the VOC can be reduced. Then, the parameters of the model are adjusted according to the adjusted weights of each cluster center, and the adjusted model is used to continue clustering the remaining VOCs until the VOCs corresponding to all cluster centers match the pathway subtype result of the target gene, then the clustering process of the VOCs can be determined to be completed, and the finally screened VOCs are determined as the respiratory markers to be screened.

[0118] It should be noted that when obtaining the gene transcription pathway related to the target gene of each disease subtype from the KEGG database, the target gene of each disease subtype can be determined according to the specific application scenario. For example, when screening respiratory markers for lung cancer, the target gene related to the gene transcription pathway of each lung cancer disease subtype can be obtained from the KEGG database according to each disease subtype of lung cancer.

[0119] Further, when clustering VOCs, the expression profile similarity and metabolic similarity between VOCs can be considered simultaneously to improve the accuracy of the clustering results. That is, in a possible implementation manner of the embodiment of the present application, the above-mentioned clustering each VOC according to the expression profile and enzyme network corresponding to each VOC to determine the candidate breath marker included in the VOCs, and the clustering center corresponding to each candidate breath marker and the disease classification can include:

[0120] inputting the expression profile and enzyme network corresponding to each VOC into a preset self-organizing neural network to determine the expression profile similarity between each VOC and each clustering center in the preset self-organizing neural network, and the metabolic similarity between each VOC and each clustering center;

[0121] determining the distance between each VOC and each clustering center according to the expression profile similarity between each VOC and each clustering center in the preset self-organizing neural network, the metabolic similarity between each VOC and each clustering center, and a preset weight;

[0122] determining the candidate breath marker included in the VOCs, and the clustering center corresponding to each candidate breath marker and the disease classification according to the distance between each VOC and each clustering center.

[0123] The self-organizing neural network can be an artificial neural network using an unsupervised competitive learning mechanism to discover the internal laws of input data by self-organizing adjustment of network parameters and structure.

[0124] In the embodiment of the present application, the preset self-organizing neural network can generate clustering centers autonomously in the process of clustering according to the input expression profile and enzyme network corresponding to each VOC. That is, in the initial stage of clustering, the preset self-organizing neural network can determine the distance between two VOCs according to the expression profile and enzyme network corresponding to the two VOCs, and when the distance between the two VOCs is less than a distance threshold, it is determined that the two VOCs belong to one clustering center, and a clustering center is generated according to the two VOCs; if the distance between the two VOCs is greater than the distance threshold, it can be determined that the two VOCs belong to two clustering centers respectively, and the two VOCs are determined as clustering centers respectively. After one or more clustering centers are determined initially, the input VOCs can be clustered according to the generated clustering centers.

[0125] Correspondingly, the similarity between the expression profile corresponding to the input VOCs and the expression profile corresponding to each cluster center, and the similarity between the enzyme network corresponding to the VOCs and the enzyme network corresponding to each cluster center, can be determined first, and then the similarity between the expression profile corresponding to the VOCs and the expression profile corresponding to each cluster center and the similarity between the enzyme network corresponding to the VOCs and the enzyme network corresponding to each cluster center are weighted and summed by using a preset weight, so as to determine the distance between the VOCs and each cluster center. Then, the cluster center with a distance less than a distance threshold to the VOCs can be determined as the cluster center corresponding to the VOCs. If the distance between the VOCs and each cluster center is greater than the distance threshold, the VOCs can be determined as a new cluster center. The above process is sequentially performed on each VOCs until the cluster result corresponding to each VOCs is determined, and then it is determined that the clustering is completed.

[0126] As a possible implementation manner, the expression profile similarity and the metabolic similarity between the VOCs and the cluster center can be determined by the following manner:

[0127] The expression profile similarity between the ith VOCs and the jth cluster center is determined according to the distance between the expression profile corresponding to the ith VOCs and the expression profile corresponding to the jth cluster center, wherein i is a positive integer greater than or equal to 1 and less than or equal to M, and j is a positive integer greater than or equal to 1 and less than or equal to N.

[0128] The enzyme network corresponding to the ith VOCs and the enzyme network corresponding to the jth cluster center are fused to generate a fused enzyme network.

[0129] The metabolic similarity between the ith VOCs and the jth cluster center is determined according to the difference between the enzyme network corresponding to the jth cluster center and the fused enzyme network.

[0130] As a possible implementation manner, the distance between the expression profile corresponding to the VOCs and the expression profile corresponding to the cluster center can be determined as the expression profile similarity between the VOCs and the cluster center. For example, the Euclidean distance between the expression profile corresponding to the VOCs and the expression profile corresponding to the cluster center can be determined as the expression profile similarity between the VOCs and the cluster center.

[0131] As a possible implementation manner, when the metabolic similarity between the VOCs and the cluster center is determined, the enzyme networks of all VOCs corresponding to the cluster center can be fused to generate an enzyme network corresponding to the cluster center, and then the enzyme network corresponding to the VOCs and the enzyme network corresponding to the cluster center can be fused to generate a fused enzyme network, and then the average metabolic similarity of the enzyme network corresponding to the cluster center and the average metabolic similarity of the fused enzyme network can be determined respectively. Specifically, the average metabolic similarity of the enzyme network corresponding to the jth cluster center can be represented by the following formula:

[0132]

[0133] wherein, is the average metabolic similarity of the enzyme network corresponding to the jth cluster center, and mn ρ i represents the biochemical connection between the enzyme m and the enzyme n in the enzyme network corresponding to the jth cluster center, and J represents the number of enzymes involved in the enzyme network corresponding to the jth cluster center.

[0134] After the average metabolic similarity of the enzyme network corresponding to the cluster center and the average metabolic similarity of the fused enzyme network are determined, the difference between the average metabolic similarity of the fused enzyme network and the average metabolic similarity of the enzyme network corresponding to the cluster center is determined as the difference between the enzyme network corresponding to the cluster center and the fused enzyme network, and the lower the metabolic similarity between the enzyme network corresponding to the cluster center and the fused enzyme network is, the lower the difference between the enzyme network corresponding to the cluster center and the fused enzyme network is.

[0135] Correspondingly, after the expression profile similarity and the metabolic similarity between the ith VOCs and the jth cluster center are determined by the above-mentioned manner, the distance between the ith VOCs and the jth cluster center can be determined by the following formula:

[0136]

[0137] wherein, p i is the expression profile of the ith VOCs, and j ω i is the expression profile of the jth cluster center, and j ||2 represents the expression profile similarity between the ith VOCs and the jth cluster center, is the metabolic similarity of the enzyme network corresponding to the jth cluster center when the ith VOCs does not belong to the jth cluster center, is the metabolic similarity of the fused enzyme network corresponding to the jth cluster center when the ith VOCs belongs to the jth cluster center, represents the metabolic similarity between the ith VOCs and the jth cluster center, and a is a preset weight.

[0138] As a possible implementation, after the clustering of each VOCs is completed, i.e., the cluster centers to which each VOCs belongs are determined, the similarity between two VOCs can be determined according to the distance between the cluster centers to which the two VOCs belong, and then the similarity network of VOCs can be constructed by using the similarity between each pair of VOCs.

[0139] As an example, the similarity between the ith VOCs and the jth VOCs can be determined by the following formula:

[0140]

[0141] wherein sim(vi, vj) is the similarity between the ith VOCs and the jth VOCs, vi is the ith VOCs, vj is the jth VOCs, ωi is the cluster center to which the ith VOCs belongs, and ωj is the cluster center to which the jth VOCs belongs. That is, if the cluster centers to which the two VOCs belong are the same, the similarity between the two VOCs can be determined as 1, and if the cluster centers to which the two VOCs belong are different, the Euclidean distance between the cluster centers to which the two VOCs belong can be determined as the similarity between the two VOCs. i j

[0142] In step 106, the target respiratory marker related to the disease is screened from the candidate respiratory markers according to the respiratory fingerprint corresponding to each exhaled gas, each candidate respiratory marker, the sample similarity network, and the VOCs similarity network.

[0143] In the embodiments of the present application, since the sample similarity network contains the similarity between the sampling conditions and the sampling objects of each exhaled gas, and the VOCs similarity network contains the similarity information between each candidate respiratory marker, the respiratory fingerprint related to each exhaled gas and each candidate respiratory marker can be determined according to the respiratory fingerprint corresponding to each exhaled gas and each candidate respiratory marker, and then the respiratory fingerprint related to each exhaled gas and each candidate respiratory marker is used to determine the disease information contained in the sampling objects corresponding to each exhaled gas, and the sample similarity network and the VOCs similarity network are introduced to comprehensively consider the differences in the sampling environment, the sampling objects, and the characteristics of VOCs, so as to screen the target respiratory marker related to the disease from each candidate respiratory marker.

[0144] ​​Further, since the same disease can include multiple disease subtypes, in the case where the disease subtypes of each candidate respiratory marker are determined by clustering in the above steps, the embodiments of the present application can further select target respiratory markers related to each disease subtype from each candidate respiratory marker to further improve the accuracy of disease respiratory marker screening. That is, in one possible implementation of the embodiments of the present application, the above step 106 can include:

[0145] According to the respiratory fingerprint corresponding to each exhaled gas, each candidate respiratory marker, the sample similarity network and the VOCs similarity network, target respiratory markers related to each disease subtype are selected from the candidate respiratory markers.

[0146] As one possible implementation, the target respiratory markers related to each disease subtype can be selected from the candidate respiratory markers by the following method:

[0147] According to the candidate respiratory markers contained in each exhaled gas, a relationship matrix between the sample data set and the candidate respiratory markers is determined;

[0148] Each exhaled gas in the sample data set is divided into training samples and prediction samples;

[0149] According to the true disease subtype corresponding to each training sample, the disease subtype corresponding to each candidate respiratory marker, and the serial number of each training sample in the sample data set, a first effectiveness vector corresponding to each training sample is determined respectively, wherein the first effectiveness vector is used to represent the candidate respiratory markers related to the true disease subtype corresponding to the training sample and the serial number of the training sample in the sample data set;

[0150] According to the serial number of each prediction sample in the sample data set, an initial effectiveness vector corresponding to each prediction sample is determined respectively;

[0151] Using a preset label propagation algorithm, the sample similarity network, the VOCs similarity network, the relationship matrix, each first effectiveness vector and each initial effectiveness vector are processed to determine a prediction effectiveness vector corresponding to each prediction sample, wherein the prediction effectiveness vector is used to represent the candidate respiratory markers related to the predicted disease subtype of the prediction sample and the serial number of the training sample in the sample data set;

[0152] According to the prediction effectiveness vector corresponding to each prediction sample, a predicted disease subtype corresponding to each prediction sample is determined;

[0153] According to the matching degree between the predicted disease subtype and the true disease subtype corresponding to each prediction sample, target respiratory markers related to each disease subtype are determined from the candidate respiratory markers.

[0154] The relation matrix can be a matrix generated based on the expression data of each candidate respiratory biomarker in the respiratory fingerprint corresponding to each exhaled gas, which can be used to characterize the correlation between the sample dataset and the candidate respiratory biomarkers.

[0155] For example, if there are n candidate respiratory biomarkers and the sample dataset contains m exhaled gas samples, for a single exhaled gas in the sample dataset, the content of each candidate respiratory biomarker can be determined from the respiratory fingerprint corresponding to that exhaled gas, and an n-dimensional column vector can be generated. Then, the m n-dimensional column vectors corresponding to the m exhaled gas samples can be combined to generate an n×m-dimensional matrix, which serves as the relationship matrix R between the sample dataset and the exhaled biomarkers to be screened.

[0156] As one possible implementation, after determining the relation matrix R, the sample similarity matrix, the VOCs similarity matrix, and the relation matrix R can be combined to generate a matrix A that simultaneously contains similarity information between samples, similarity information between each candidate respiratory biomarker, and similarity information between each sample and each candidate respiratory biomarker. That is, when the sample dataset contains m samples and the number of candidate respiratory biomarkers is n, the sample similarity network is an m×m dimensional matrix Sim. S The similarity network of VOCs is represented by an n×n matrix Sim. V Then matrix A can be a (n+m)×(n+m) dimensional matrix:

[0157]

[0158] As one possible implementation, for each training sample, the first validity vector y can be determined based on the actual disease type, the disease type corresponding to each candidate respiratory biomarker, and the sequence number of each training sample in the sample dataset, all included in the sampling object information corresponding to each training sample. For a training sample s, the constructed first validity vector y = [y1, y2, ..., y n ,y n+1 ,…,y n+m ] T Where, when k = 1, 2, ..., n, if the true disease type corresponding to the training sample s is the same as the disease type related to the kth VOCs, then y k =1 otherwise 0; when k = n+1, n+2, ..., n+m, if the training sample s is the kn-th sample in the sample dataset, then y k =1.

[0159] For example, if the training sample s is the second sample in the sample data set, and the real disease type corresponding to the training sample s is type A, and the disease type related to the candidate respiratory marker VOC1, VOC2, VOC3 is also type A, that is, the real disease type corresponding to the training sample s is the same as the disease type related to the candidate respiratory marker VOC1, VOC2, VOC3, it can be determined that in the first effectiveness vector corresponding to the training sample s, the values of y1, y2, y3 are all 1, and the values of y4 to yn are all 0; the value of yn+2 is 2, and the values of yn+1, yn+3 to yn+m are 0.

[0160] As a possible implementation, for each prediction sample in the sample data set, an initial effectiveness vector corresponding to each prediction sample can be constructed only according to the serial number of each prediction sample in the sample data set, to predict the VOCs related to the prediction sample, and then to screen reliable respiratory markers from the candidate respiratory markers as target respiratory markers according to the prediction result.

[0161] For example, if the prediction sample a is the 100th sample in the sample data set, it can be determined that the value of the element yn+100 in the initial effectiveness vector y corresponding to the prediction sample is 1, and the values of the other elements are all 0.

[0162] In the embodiments of the present application, after the first effectiveness vector corresponding to each training sample and the initial effectiveness vector corresponding to each prediction sample are determined, the first effectiveness vector and the initial effectiveness vector can be combined to generate a (n+m) x m matrix Y. Then, the matrix A and the matrix Y can be input into a preset label propagation algorithm, the label information of the training sample is propagated to the whole heterogeneous network, the known disease type is propagated to other samples, until the nodes reach a stable state, to predict the values of the first n elements of the initial effectiveness vector corresponding to each prediction sample, and generate a prediction effectiveness vector corresponding to each prediction sample.

[0163] In the embodiments of the present application, the values of the first n elements in the prediction validity vector corresponding to each prediction sample can represent the predicted VOCs related to the prediction sample, so that for a prediction sample, the candidate respiratory markers predicted to be related to the prediction sample can be determined according to the values of the first n elements in the prediction validity vector corresponding to the prediction sample, and then the disease type corresponding to the candidate respiratory markers predicted to be related to the prediction sample is determined as the predicted disease type corresponding to the prediction sample, and then whether the candidate respiratory markers predicted to be related to the prediction sample can be used as the target respiratory marker and the disease type corresponding to the target respiratory marker is determined according to the matching degree between the predicted disease type corresponding to the prediction sample and the true disease type. The above processing is performed on each prediction sample in the same way, and then all target respiratory markers included in the candidate respiratory markers and the disease type related to each target respiratory marker can be determined.

[0164] For example, if the values of y2-y4 of the prediction validity vector corresponding to a prediction sample are 1, it can be determined that the predicted VOCs related to the prediction sample are VOC2, VOC3 and VOC4, and the disease type related to VOC2, VOC3 and VOC4 is type A, and the true disease type corresponding to the prediction sample is also type A, so it can be determined that the candidate respiratory markers VOC2, VOC3 and VOC4 have strong correlation with the disease type A and can be used to accurately predict the presence of the disease type A, so VOC2, VOC3 and VOC4 can be determined as the target respiratory marker related to the disease type A. In the above example, if the disease type related to VOC2, VOC3 and VOC4 is type A, and the true disease type corresponding to the prediction sample is type B, it can be determined that the correlation between VOC2, VOC3 and VOC4 and the disease type A is unreliable, so VOC2, VOC3 and VOC4 can be removed and not used as target respiratory markers.

[0165] It should be noted that the above examples are only exemplary and should not be regarded as limiting the present application. In actual use, the specific way of screening the target respiratory marker according to the matching degree between the predicted disease type and the true disease type can be determined according to actual needs and specific application scenarios, and the embodiments of the present application do not limit this step. For example, a matching degree threshold can also be preset, and the target respiratory marker can be determined according to the relationship between the matching degree and the matching degree threshold.

[0166] The disease respiratory marker screening method provided in the application is to obtain a sample data set of exhaled gas, then determine a feature sequence corresponding to each exhaled gas according to the sampling parameters and sampling object information of each exhaled gas, then determine a sample similarity network of the sample data set according to the similarity between the feature sequences of the exhaled gas, then determine an expression profile of VOCs contained in each exhaled gas according to the respiratory fingerprint of each exhaled gas, then determine a candidate respiratory marker included in each VOC and a VOC similarity network corresponding to the candidate respiratory marker according to the expression profile of each VOC, and finally screen a target respiratory marker related to a disease according to the respiratory fingerprint corresponding to each exhaled gas, each candidate respiratory marker, the sample similarity network and the VOC similarity network. In this way, by constructing the sample similarity network and the VOC similarity network, the statistical differences between the exhaled gas of the diseased population and the exhaled gas of the healthy population are considered in the disease respiratory marker screening, and the experimental condition difference information between samples, the sampling object difference information, and the characteristic difference information between VOCs and the physiological and pathological correlation data of VOCs are introduced, so that the accuracy of the disease respiratory marker screening is improved, the screening result is closer to the demand of precision medicine, and the possibility of success in subsequent biological principle research, identification performance verification and clinical experiment of the marker is greater.

[0167] In a possible implementation form of the application, since the sampling parameters of the exhaled gas and the data types included in the sampling object information of the exhaled gas are relatively complex, different data processing methods can be used according to the data types of the parameters to process the sampling parameters and the sampling object information of the exhaled gas to generate the feature sequence corresponding to the exhaled gas.

[0168] The disease respiratory marker screening method provided in the application will be further described below. Figure 2 The disease respiratory marker screening method provided in the application will be further described below.

[0169] Figure 2 Another flowchart of the disease respiratory marker screening method provided in the application is shown.

[0170] In step 201, a sample data set corresponding to the exhaled gas is obtained, wherein the sample data set includes the respiratory fingerprint, the sampling parameters and the sampling object information corresponding to the exhaled gas.

[0171] The specific implementation process and principles of the above step 201 can refer to the detailed description of the above embodiments, which will not be described here.

[0172] It should be noted that the sampling parameters can include at least one of the following parameters: experimental control conditions, gas sampling instrument parameters, thermal desorption parameters, chromatographic parameters, mass spectrometry parameters; the sampling object information can include at least one of the following information: personal basic information, living environment and habit information, medical information.

[0173] As a possible implementation, the experimental control conditions can include parameters such as pre-gas state, breathing mode, whether fasting, sampling mode, etc.; the gas sampling instrument parameters can include parameters such as sample volume, gas sampling flow rate, gas sampling temperature, etc.; the thermal desorption parameters can include parameters such as purge flow rate, purge temperature, primary desorption parameters, secondary desorption parameters, inlet split and outlet split, etc.; the chromatographic parameters can include parameters such as carrier gas flow rate, sampling time, sampling port temperature, temperature program, etc.; the mass spectrometry parameters can include parameters such as interface temperature, ion source temperature, scan mode, scan range, etc.; the personal basic information can include information such as age, gender, weight, etc.; the living environment and habit information can include information such as related occupation, air quality in living place, smoking history, passive smoking history, etc.; the medical information can include information such as lung infection history, gastrointestinal infection history, tumor history, medication history.

[0174] It should be noted that the medical information can be related to the type of disease that needs to be screened for respiratory markers, and the embodiments of the present application do not limit this. In actual use, the disease-related medical information can be obtained according to the actual disease research scene.

[0175] Step 202, determining the numerical data and non-numerical data in the sampling object information and the sampling parameters corresponding to each exhaled gas;

[0176] The numerical data can be data whose value can be directly represented by a specific numerical value. For example, the gas sampling instrument parameters, the thermal desorption parameters, the carrier gas flow rate, the sampling time, the sampling port temperature, etc. in the chromatographic parameters, the mass spectrometry parameters, the personal basic information, etc.

[0177] The non-numerical data can be data whose value cannot be directly represented by a specific numerical value. For example, the experimental control conditions, the temperature program in the chromatographic parameters, the living environment and habit information, the medical information, etc.

[0178] In the embodiments of the present application, since the sampling parameters and the sampling object information corresponding to the exhaled gas are directly combined to form the feature sequence corresponding to the exhaled gas, it is not only inconvenient for data processing, but also affects the accuracy of respiratory marker screening. Therefore, the embodiments of the present application can further process the parameters to generate a feature sequence containing effective information, so that the numerical data and non-numerical data in the sampling parameters and the sampling object information can be first determined for data processing.

[0179] Step 203, normalize the numerical data corresponding to each exhaled gas to determine the value sequence corresponding to each exhaled gas.

[0180] In the embodiments of the present application, the above-mentioned sampling parameters and sampling object information are classified, and for the numerical data among them: gas sampling instrument parameters, thermal desorption parameters, carrier gas flow rate, sample injection time, sample injection port temperature and other parameters in the chromatographic parameter, mass spectrometry parameters, and personal basic information, the above-mentioned various parameters can be normalized between 0 and 1 to determine the normalized numerical value corresponding to each parameter, and the numerical value of the numerical data corresponding to each exhaled gas after normalization is combined to generate the value sequence corresponding to each exhaled gas.

[0181] Step 204, state mapping is performed on the non-numerical data corresponding to each exhaled gas to determine the state sequence corresponding to each exhaled gas.

[0182] In the embodiments of the present application, the above-mentioned sampling parameters and sampling object information are classified, and for the non-numerical data among them: experimental control conditions, temperature program in the chromatographic parameter, living environment and habit information, medical information, the above-mentioned non-numerical data can be abstracted to determine the state corresponding to each non-data type value, and then the state corresponding to the non-numerical data corresponding to each exhaled gas is combined to generate the state sequence corresponding to each exhaled gas.

[0183] For example, each non-numerical data can be abstracted according to whether a certain way is used or whether it is in a certain state. For example, for the breathing mode in the experimental control condition, the breathing mode can be abstracted according to whether the breathing mode is a certain breathing mode (such as tidal breathing), if yes, the value of the breathing mode can be determined as YES, if not, the value of the breathing mode can be determined as NO.

[0184] Step 205, determine the sample similarity network corresponding to the sample data set according to the similarity between the feature sequences corresponding to each exhaled gas.

[0185] The specific implementation process and principles of the above-mentioned step 205 can refer to the detailed description of the above-mentioned embodiments, which will not be repeated here.

[0186] Further, since the feature sequence can include the value sequence and the state information, the similarity between the feature sequences corresponding to each exhaled gas can be determined according to the similarity between the value sequences corresponding to the exhaled gas and the similarity between the state sequences. That is, in one possible implementation manner of the embodiments of the present application, the above-mentioned step 205 can further include:

[0187] determine similarity between the value sequence corresponding to each exhaled gas, and similarity between the state sequence corresponding to each exhaled gas;

[0188] determine similarity between the value sequence corresponding to each exhaled gas, and similarity between the state sequence corresponding to each exhaled gas;

[0189] In the embodiment of the present application, since the sampling parameters corresponding to the exhaled gas and the sampling object information both contain numerical data and non-numerical data, the value sequence corresponding to the exhaled gas can include the value sequence corresponding to the sampling parameters and the value sequence corresponding to the sampling object information; the state sequence corresponding to the exhaled gas can include the state sequence corresponding to the sampling parameters and the state sequence corresponding to the sampling object information. Therefore, in the embodiment of the present application, the similarity between the value sequence corresponding to the sampling parameters, the similarity between the value sequence corresponding to the sampling object information, the similarity between the state sequence corresponding to the sampling parameters, and the similarity between the state sequence corresponding to the sampling object information can be weighted and summed to determine the similarity between the feature sequences corresponding to two exhaled gases. For example, the similarity between the feature sequences corresponding to the exhaled gas S1 and the exhaled gas S2 is:

[0190] sim(s1,s2)=A*TC mtd (s1,s2)+B*Ω mtd (s1,s2)+C*TC cln (s1,s2)+D*Ω cln (s1,s2)

[0191] wherein, TC mtd (s1,s2) can refer to the similarity between the state sequence corresponding to the sampling parameters of the exhaled gas S1 and the exhaled gas S2, Ω mtd (s1,s2) can refer to the similarity between the value sequence corresponding to the sampling parameters of the exhaled gas S1 and the exhaled gas S2; TC cln (s1,s2) can refer to the similarity between the state sequence corresponding to the sampling object information of the exhaled gas S1 and the exhaled gas S2, Ω cln (s1,s2) can refer to the similarity between the value sequence corresponding to the sampling object information of the exhaled gas S1 and the exhaled gas S2, wherein A, B, C, and D are weight values, and the value of A+B+C+D can be 1.

[0192] It should be noted that in actual use, the values of the weights A, B, C, and D can be determined according to actual needs and specific application scenarios, and the embodiment of the present application does not limit this.

[0193] Step 206, determining the expression profile of VOCs contained in each exhaled gas according to the breathprint of each exhaled gas.

[0194] Step 207, determining the candidate breath marker included in the VOCs and the VOCs similarity network corresponding to the candidate breath marker according to the expression profile of each VOC.

[0195] Step 208, screening the target breath marker related to the disease from the candidate breath marker according to the breathprint of each exhaled gas, each candidate breath marker, the sample similarity network and the VOCs similarity network.

[0196] The specific implementation process and principle of steps 206-208 described above can refer to the detailed description of the above embodiments, which will not be repeated here.

[0197] The screening method of disease breath marker provided in the present application, by obtaining the sample data set of exhaled gas, then determining the value sequence in the feature sequence corresponding to each exhaled gas according to the numerical data in the sampling parameters and sampling object information of each exhaled gas, and determining the state sequence in the feature sequence corresponding to each exhaled gas according to the non-numerical data in the sampling parameters and sampling object information of each exhaled gas, then determining the sample similarity network of the sample data set according to the similarity between the value sequences and the similarity between the state sequences in the feature sequence of each exhaled gas, then determining the expression profile of VOCs contained in each exhaled gas according to the breathprint of each exhaled gas, then determining the candidate breath marker included in the VOCs and the VOCs similarity network corresponding to the candidate breath marker according to the expression profile of each VOC, and finally screening the target breath marker related to the disease from the candidate breath marker according to the breathprint of each exhaled gas, each candidate breath marker, the sample similarity network and the VOCs similarity network. Thus, by constructing the sample similarity network and the VOCs similarity network, not only the statistical differences between the exhaled gases of the diseased and healthy populations are considered in the screening of disease breath markers, but also the experimental condition difference information between samples, the sampling object difference information, and the characteristic difference information between VOCs, and the physiological and pathological correlation data of VOCs are introduced, thereby improving the accuracy of disease breath marker screening, making the screening results closer to the needs of precision medicine, and making the subsequent biological principle research of the marker, identification performance verification and clinical experiment more likely to be successful.

[0198] It should be understood that the size of the serial number of each step in the above embodiments does not mean the order of execution, and the execution order of each process should be determined by its function and internal logic, and should not constitute any limitation on the implementation process of the embodiments of the present application.

[0199] The screening method of the disease breath marker corresponding to the above embodiment, Figure 3 A structural block diagram of a screening device of a disease breath marker provided by an embodiment of the present application is shown. For ease of illustration, only parts related to the embodiments of the present application are shown.

[0200] With reference to Figure 3 The device 30 comprises:

[0201] A first obtaining module 31 is configured to obtain a sample data set corresponding to exhaled gas, wherein the sample data set comprises a plurality of breathprints corresponding to exhaled gas, sampling parameters, and sampling object information corresponding to exhaled gas.

[0202] A first determining module 32 is configured to determine a feature sequence corresponding to each exhaled gas according to the sampling parameters corresponding to each exhaled gas and the sampling object information corresponding to each exhaled gas.

[0203] A second determining module 33 is configured to determine a sample similarity network corresponding to the sample data set according to the similarity between the feature sequences corresponding to the exhaled gas.

[0204] A third determining module 34 is configured to determine an expression profile corresponding to VOCs contained in each exhaled gas according to the breathprint corresponding to each exhaled gas.

[0205] A fourth determining module 35 is configured to determine a candidate breath marker included in the VOCs and a VOCs similarity network corresponding to the candidate breath marker according to the expression profile corresponding to each VOC.

[0206] A first screening module 36 is configured to screen a target breath marker related to a disease from the candidate breath marker according to the breathprint corresponding to each exhaled gas, each candidate breath marker, the sample similarity network, and the VOCs similarity network.

[0207] In actual use, the screening device of a disease breath marker provided by the embodiments of the present application can be configured in any terminal device to perform the aforementioned screening method of a disease breath marker.

[0208] The disease respiratory marker screening device provided by the application, by obtaining a sample data set of exhaled gas, then determining the feature sequence corresponding to each exhaled gas according to the sampling parameters and sampling object information of each exhaled gas, then determining the sample similarity network of the sample data set according to the similarity between the feature sequences of each exhaled gas, then determining the expression profile corresponding to the VOCs contained in each exhaled gas according to the respiratory fingerprint of each exhaled gas, then determining the candidate respiratory marker included in each VOCs and the VOCs similarity network corresponding to the candidate respiratory marker according to the expression profile of each VOCs, and finally screening the target respiratory marker related to the disease from the candidate respiratory marker according to the respiratory fingerprint corresponding to each exhaled gas, each candidate respiratory marker, the sample similarity network and the VOCs similarity network. Therefore, by constructing the sample similarity network and the VOCs similarity network, not only the statistical differences between the exhaled gas of the diseased and healthy population are considered in the disease respiratory marker screening, but also the experimental condition difference information between the samples, the sampling object difference information, and the characteristic difference information between the VOCs and the physiological and pathological correlation data of the VOCs are introduced, thereby improving the accuracy of the disease respiratory marker screening, making the screening result closer to the demand of precision medicine, and the possibility of success in subsequent biological principle research, identification performance verification and clinical experiment of the marker is greater.

[0209] In a possible implementation manner of the embodiment of the application, the disease includes multiple disease types; correspondingly, the fourth determining module 35 includes:

[0210] The first determining unit is configured to determine the enzyme network corresponding to each VOC according to the metabolic pathways related to each VOC and the enzyme list related to each metabolic pathway;

[0211] The second determining unit is configured to cluster each VOC according to the expression profile and the enzyme network corresponding to each VOC, to determine the candidate respiratory marker included in the VOCs, and the cluster center corresponding to each candidate respiratory marker and the disease type, wherein each cluster center corresponds to one disease type.

[0212] The third determining unit is configured to determine the VOCs similarity network corresponding to the candidate respiratory marker according to the cluster center corresponding to each candidate respiratory marker.

[0213] Correspondingly, the first screening module 36 includes:

[0214] The first screening unit is configured to screen the target respiratory marker related to each disease type from the candidate respiratory marker according to the respiratory fingerprint corresponding to each exhaled gas, each candidate respiratory marker, the sample similarity network and the VOCs similarity network.

[0215] Further, in a possible implementation of the embodiments of the present application, the first screening unit is specifically configured to:

[0216] determine a relationship matrix between the sample data set and the candidate respiratory markers according to each exhaled gas containing a candidate respiratory marker;

[0217] divide each exhaled gas in the sample data set into a training sample and a prediction sample;

[0218] determine a first effectiveness vector corresponding to each training sample according to a real disease type corresponding to the training sample, disease types corresponding to each candidate respiratory marker, and a serial number of the training sample in the sample data set, wherein the first effectiveness vector is used to represent each candidate respiratory marker related to the real disease type corresponding to the training sample and the serial number of the training sample in the sample data set;

[0219] determine an initial effectiveness vector corresponding to each prediction sample according to a serial number of the prediction sample in the sample data set;

[0220] process the sample similarity network, the VOCs similarity network, the relationship matrix, each first effectiveness vector, and each initial effectiveness vector by using a preset label propagation algorithm to determine a prediction effectiveness vector corresponding to each prediction sample, wherein the prediction effectiveness vector is used to represent each candidate respiratory marker related to a prediction disease type of the prediction sample and the serial number of the training sample in the sample data set;

[0221] determine a prediction disease type corresponding to each training sample according to the prediction effectiveness vector corresponding to the training sample;

[0222] determine a target respiratory marker related to each disease type from the candidate respiratory markers according to a matching degree between the prediction disease type corresponding to each training sample and a real disease type.

[0223] Further, in a possible implementation of the embodiments of the present application, the second determining unit is specifically configured to:

[0224] input the expression profile corresponding to each VOCs and the enzyme network into a preset self-organizing neural network to determine an expression profile similarity between each VOCs and each cluster center in the preset self-organizing neural network, and a metabolic similarity between each VOCs and each cluster center;

[0225] determine a distance between each VOCs and each cluster center according to the expression profile similarity between each VOCs and each cluster center in the preset self-organizing neural network, the metabolic similarity between each VOCs and each cluster center, and a preset weight.

[0226] According to the distance between each VOCs and each cluster center, the candidate respiratory markers included in the VOCs, and the cluster center corresponding to each candidate respiratory marker are determined and the disease classification.

[0227] Further, in another possible implementation manner of the embodiment of the present application, the VOCs include M VOCs, the preset self-organizing neural network includes N cluster centers, and M and N are positive integers; correspondingly, the second determining unit is further configured to:

[0228] According to the distance between the expression profile corresponding to the ith VOCs and the expression profile corresponding to the jth cluster center, the expression profile similarity between the ith VOCs and the jth cluster center is determined, wherein i is a positive integer greater than or equal to 1 and less than or equal to M, and j is a positive integer greater than or equal to 1 and less than or equal to N;

[0229] The enzyme network corresponding to the ith VOCs and the enzyme network corresponding to the jth cluster center are fused to generate a fused enzyme network;

[0230] According to the difference between the enzyme network corresponding to the jth cluster center and the fused enzyme network, the metabolic similarity between the ith VOCs and the jth cluster center is determined.

[0231] Further, in another possible implementation manner of the embodiment of the present application, the sampling parameters include at least one of the following parameters: experimental control conditions, gas sampling instrument parameters, thermal desorption parameters, chromatographic parameters, and mass spectrometry parameters; and the sampling object information includes at least one of the following information: personal basic information, living environment and habit information, and medical information.

[0232] Further, in another possible implementation manner of the embodiment of the present application, the first determining module 32 includes:

[0233] The fourth determining unit is configured to determine the numerical data and the non-numerical data in the sampling parameters and the sampling object information corresponding to each exhaled gas;

[0234] The fifth determining unit is configured to perform normalization processing on the numerical data corresponding to each exhaled gas to determine a value sequence corresponding to each exhaled gas;

[0235] The sixth determining unit is configured to perform state mapping on the non-numerical data corresponding to each exhaled gas to determine a state sequence corresponding to each exhaled gas.

[0236] Further, in another possible implementation manner of the embodiment of the present application, the apparatus 30 further includes:

[0237] The fifth determining module is configured to determine the similarity between the value sequences corresponding to the respective exhaled gases and the similarity between the state sequences corresponding to the respective exhaled gases.

[0238] The sixth determining module is configured to determine the similarity between the feature sequences corresponding to the respective exhaled gases according to the similarity between the value sequences corresponding to the respective exhaled gases and the similarity between the state sequences corresponding to the respective exhaled gases.

[0239] It should be noted that the information interaction and execution process between the above apparatuses / units are based on the same concept as the method embodiments of the present application, and the specific functions and technical effects brought by the above apparatuses / units can be referred to the method embodiments, which will not be described here.

[0240] Those skilled in the art can clearly understand that, for the convenience and brevity of description, only the division of the above functional units / modules is taken as an example for illustration, and in actual application, the above functions can be completed by different functional units / modules according to needs, that is, the internal structure of the apparatus is divided into different functional units or modules to complete all or part of the functions described above. Each functional unit / module in the embodiments can be integrated in one processing unit, or each unit can be physically present separately, or two or more units can be integrated in one unit. The integrated unit can be realized in the form of hardware or in the form of software functional unit. In addition, the specific names of each functional unit / module are only for easy distinction, and do not limit the protection scope of the present application. The specific working process of the units / modules in the above system can refer to the corresponding process in the method embodiments, which will not be described here.

[0241] In order to realize the above-mentioned embodiments, the present application further provides a terminal device.

[0242] Figure 4 The structural schematic diagram of the terminal device of one embodiment of the present application.

[0243] As shown in Figure 4 The terminal device 200 includes:

[0244] The memory 210 and the at least one processor 220, the bus 230 connecting different components including the memory 210 and the processor 220, the memory 210 storing a computer program, and the processor 220 executing the program to realize the disease respiratory marker screening method of the embodiments of the present application.

[0245] Bus 230 generally represents what in the art will be broadly understood as a bus structure for a communication architecture. It can represent one or more of several types of bus structures including memory buses or memory controllers, peripheral buses, graphics acceleration buses (e.g., AGP bus) and a local bus using many of the bus structures above.

[0246] Terminal device 200 typically includes a variety of computer system readable media. These media can be any available media that is accessible by terminal device 200 and includes both volatile and nonvolatile media, removable and non-removable media.

[0247] Memory 210 also can include computer system readable media in the form of volatile memory, such as random access memory (RAM) 240 and / or cache memory 250. Terminal device 200 can further include other removable / non-removable, volatile / non-volatile computer system storage media. By way of example only, storage system 260 can be provided for reading from and writing to a non-removable, non-volatile magnetic media (e.g., a "hard drive"). Figure 4 Although not shown, a magnetic disk drive can also be utilized in some embodiments for reading from and writing to a removable, non-volatile magnetic disk (e.g., a "floppy disk"), and an optical disk drive can be utilized in some embodiments for reading from and writing to a removable, non-volatile optical disk (e.g., a CD-ROM, DVD-ROM or other optical media). In these instances, each can be connected to bus 230 by one or more data media interfaces. As will be further depicted and described below, memory 210 can include at least one program product having a set (e.g., at least one) of program modules that are configured to carry out the functions of embodiments of the application. Figure 4 As will be appreciated, various items can be stored in one or more of the aforementioned memory or storage locations, including computer-readable instructions, program modules, program data, and the like. In this regard, the program modules, which are implemented in software, are given the control and processing advantages of both hardware and software. Note also that the computer system of this illustrative embodiment, as well as the computers typically employed in connection with computer systems, are preferably in electronic communication with one or more computer networks, such as an intranet, the Internet, a local area network or a wide area network. Such networked environments will in many cases be commonplace in the commercial and home computer markets, but will not necessarily be so in less sophisticated implementations. Accordingly, the various network-based implementations of this illustrative embodiment can have additional elements that are not discussed in detail herein, such as firewalls, gateways, load balancers, load distributors, and the like.

[0248] Program / utility 280 having a set (at least one) of program modules 270 can be stored in memory 210 by way of example, and not limitation, as well as an operating system, one or more application programs, other program modules, and program data, and can include an implementation of a network environment, each or a combination thereof. Generally, these elements can carry out the functions and / or methodologies described in this specification.

[0249] Terminal device 200 can also be in communication with one or more external devices 290 such as a keyboard, a pointing device, a display 291, etc.; one or more devices that enable a user to interact with terminal device 200; and / or any devices (e.g., network card, modem, etc.) that enable terminal device 200 to communicate with one or more other computing devices. Such communication can be facilitated by an Input / Output (I / O) interface 292. Still yet, terminal device 200 can be in communication with one or more networks (e.g., a local area network (LAN), a wide area network (WAN), and / or the Internet) through a network adapter 293. As depicted, network adapter 293 communicates with the other components of terminal device 200 through bus 230. It should be appreciated that although not shown, other hardware and / or software modules could be used in conjunction with terminal device 200. Such modules include, but are not limited to, microcode, device drivers, redundant processing units, external disk drive arrays, RAID systems, tape drives, and data archival storage systems, etc.

[0250] Processor 220 performs various function applications and data processing by running programs stored in memory 210.

[0251] It should be noted that the implementation process and technical principles of the terminal device of the present embodiment are described above in the explanation of the disease respiratory marker screening method of the present embodiment, which will not be repeated here.

[0252] The present embodiment also provides a computer readable storage medium, which stores a computer program, and the computer program is executed by a processor to implement the steps in each of the above method embodiments.

[0253] The present embodiment provides a computer program product, which, when executed on a terminal device, enables the terminal device to implement the steps in each of the above method embodiments.

[0254] The integrated unit, if implemented in the form of a software function unit and sold or used as an independent product, can be stored in a computer readable storage medium. Based on such understanding, the present application can implement all or part of the processes in the above-mentioned embodiment methods through a computer program to instruct relevant hardware to complete, and the computer program can be stored in a computer readable storage medium. When the computer program is executed by a processor, the steps of each method embodiment described above can be implemented. The computer program includes computer program code, which can be in the form of source code, object code, executable files or some intermediate forms. The computer readable medium can at least include any entity or device capable of carrying the computer program code to the photographing device / terminal equipment, recording medium, computer memory, read-only memory (ROM), random access memory (RAM), electrical carrier signal, telecommunication signal and software distribution medium. For example, U disk, mobile hard disk, magnetic disk or optical disk, etc. In some jurisdictions, according to legislation and patent practice, the computer readable medium can not be an electrical carrier signal and a telecommunication signal.

[0255] In the above embodiments, the description of each embodiment has its own focus, and the parts not described or recorded in detail in a certain embodiment can be referred to the relevant description of other embodiments.

[0256] Those of ordinary skill in the art can realize that the units and algorithm steps of each example described in connection with the embodiments disclosed herein can be realized by electronic hardware, or a combination of computer software and electronic hardware. Whether the functions are performed in hardware or software depends on the specific application and design constraints of the technical solution. Those skilled in the art can use different methods to implement the described functions for each specific application, but such implementation should not be considered beyond the scope of the present application.

[0257] In the embodiments provided by the present application, it should be understood that the disclosed apparatus / terminal equipment and method can be implemented in other ways. For example, the above-described apparatus / terminal equipment embodiments are merely schematic, and the division of the modules or units is merely a logical function division, and actual implementation can have another division manner, for example, a plurality of units or components can be combined or integrated into another system, or some features can be ignored or not executed. In addition, the coupling or direct coupling or communication connection between the displayed or discussed each other can be indirect coupling or communication connection through some interfaces, devices or units, which can be electrical, mechanical or other forms.

[0258] The units described as separate components may or may not be physically separate, and the components displayed as units may or may not be physical units, that is, may be located in one place, or may also be distributed to multiple network units. Part or all of the units can be selected to achieve the purpose of the embodiment scheme according to actual needs.

[0259] The above embodiments are only used to illustrate the technical solutions of the present application, but not limit them; although the present application has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that the technical solutions recorded in the foregoing embodiments can still be modified, or some technical features can be replaced by equivalents; and these modifications or replacements do not make the essence of the corresponding technical solutions deviate from the spirit and scope of the technical solutions of the embodiments of the present application, and should be included in the protection scope of the present application.

Claims

1. A method for screening respiratory biomarkers for diseases, characterized in that, include: Obtain a sample dataset corresponding to exhaled gases, wherein the sample dataset includes multiple respiratory fingerprints, sampling parameters, and sampling object information corresponding to the exhaled gases, and the sampling object information includes real disease classification; Based on the sampling parameters corresponding to each exhaled gas and the sampling object information corresponding to each exhaled gas, a feature sequence corresponding to each exhaled gas is determined. The feature sequence corresponding to each exhaled gas is used to characterize the features of the sampling parameters and sampling object information corresponding to the exhaled gas. The feature sequence corresponding to each exhaled gas is generated by combining the sampling parameters and sampling object information corresponding to the exhaled gas. Based on the similarity between the feature sequences corresponding to each of the exhaled gases, a sample similarity network corresponding to the sample dataset is determined; Based on the respiratory fingerprint corresponding to each exhaled gas, the expression profile of the volatile organic compounds (VOCs) contained in each exhaled gas is determined. The expression profile of the VOCs is generated based on the VOC content in all the respiratory fingerprints in the sample dataset. If there are n VOCs involved in all the respiratory fingerprints, the expression profile corresponding to the i-th VOC contains the content of the i-th VOC in each respiratory fingerprint. n is an integer greater than or equal to 1, and i is an integer greater than or equal to 1 and less than or equal to n. Based on the metabolic pathways associated with each VOC and the list of enzymes associated with each metabolic pathway, determine the enzyme network corresponding to each VOC; An artificial neural network is used to cluster each VOCs based on its expression profile and enzyme network to determine the cluster center for each VOCs. Gene transcription pathways associated with the target genes for each disease subtype of the disease were obtained from the Kyoto Encyclopedia of Genes and Genomes (KEGG) database. Enrichment analysis was performed on the gene transcription pathways corresponding to each of the disease subtypes to determine whether there is a correlation between the gene transcription pathways corresponding to each of the disease subtypes and the metabolic pathways corresponding to each of the VOCs. If the metabolic pathway corresponding to the VOCs is associated with the gene transcription pathway corresponding to any disease subtype, then the association between the VOCs and the disease subtype is determined, and the VOCs are retained and the weight of the cluster center corresponding to the VOCs is increased for the next clustering. If the metabolic pathways corresponding to the VOCs are not associated with the gene transcription pathways corresponding to each of the disease subtypes, then the VOCs are removed. The parameters of the clustering model are adjusted according to the adjusted weights of each cluster center, and the adjusted clustering model is used to continue clustering the remaining VOCs until the VOCs corresponding to each cluster center match the gene transcription pathway corresponding to any of the disease subtypes. Then the clustering process of the VOCs is determined to be complete, and each of the finally screened VOCs is determined as a candidate respiratory biomarker. The disease subtypes associated with the cluster centers corresponding to each candidate respiratory biomarker are determined as the disease subtypes corresponding to each candidate respiratory biomarker, wherein each cluster center corresponds to a disease subtype. Based on the distance between the cluster centers corresponding to each candidate respiratory biomarker, a VOCs similarity network is determined for each candidate respiratory biomarker, wherein the VOCs similarity network is constructed using the pairwise similarity between all the candidate respiratory biomarkers. Based on the candidate respiratory biomarkers contained in each of the exhaled gases, a relationship matrix between the sample dataset and the candidate respiratory biomarkers is determined, wherein the relationship matrix is ​​used to characterize the correlation between the sample dataset and the candidate respiratory biomarkers; The exhaled gases in the sample dataset are divided into training samples and prediction samples; Based on the actual disease type corresponding to each training sample, the disease type corresponding to each candidate respiratory biomarker, and the sequence number of each training sample in the sample dataset, a first validity vector is determined for each training sample, wherein the first validity vector is used to represent each candidate respiratory biomarker related to the actual disease type corresponding to the training sample and the sequence number of the training sample in the sample dataset; Based on the index of each predicted sample in the sample dataset, determine the initial validity vector corresponding to each predicted sample; Using a preset label propagation algorithm, the sample similarity network, the VOCs similarity network, the relation matrix, each of the first validity vectors, and each of the initial validity vectors are processed to determine the prediction validity vector corresponding to each prediction sample. The prediction validity vector is used to represent each of the candidate respiratory biomarkers related to the predicted disease classification of the prediction sample and the sequence number of the prediction sample in the sample dataset. Based on the prediction validity vector corresponding to each prediction sample, the predicted disease type corresponding to each prediction sample is determined; Based on the matching degree between the predicted disease subtype and the actual disease subtype corresponding to each predicted sample, target respiratory biomarkers related to each disease subtype are determined from the candidate respiratory biomarkers.

2. The method as described in claim 1, characterized in that, The artificial neural network is a pre-defined self-organizing neural network. The step of using the artificial neural network to cluster each VOC based on its expression profile and enzyme network to determine the cluster center for each VOC includes: The expression profile and enzyme network corresponding to each VOCs are input into the preset self-organizing neural network to determine the expression profile similarity between each VOCs and each cluster center in the preset self-organizing neural network, and the metabolic similarity between each VOCs and each cluster center. Based on the expression spectrum similarity between each VOCs and each cluster center in the preset self-organizing neural network, the metabolic similarity between each VOCs and each cluster center, and the preset weights, the distance between each VOCs and each cluster center is determined. The cluster center corresponding to each VOC is determined based on the distance between each VOC and each cluster center.

3. The method as described in claim 2, characterized in that, The VOCs include M types of VOCs, and the preset self-organizing neural network includes N cluster centers, where M and N are positive integers. The step of inputting the expression profile and enzyme network corresponding to each VOC into the preset self-organizing neural network to determine the expression profile similarity between each VOC and each cluster center in the preset self-organizing neural network, and the metabolic similarity between each VOC and each cluster center, includes: The expression spectrum similarity between the i-th VOC and the j-th cluster center is determined based on the distance between the expression spectrum corresponding to the i-th VOC and the expression spectrum corresponding to the j-th cluster center, where i is a positive integer greater than or equal to 1 and less than or equal to M, and j is a positive integer greater than or equal to 1 and less than or equal to N. The enzyme network corresponding to the i-th VOCs is fused with the enzyme network corresponding to the j-th cluster center to generate a fused enzyme network. The metabolic similarity between the i-th VOC and the j-th cluster center is determined based on the difference between the enzyme network corresponding to the j-th cluster center and the fused enzyme network.

4. The method according to any one of claims 1-3, characterized in that, The sampling parameters include at least one of the following: experimental control conditions, gas sampling instrument parameters, thermal desorption parameters, chromatographic parameters, and mass spectrometry parameters; the sampling object information includes at least one of the following: basic personal information, living environment and habit information, and medical information.

5. The method as described in claim 4, characterized in that, The feature sequence corresponding to the exhaled gas includes a state sequence and a value sequence. Determining the feature sequence corresponding to each exhaled gas based on the sampling parameters and sampling object information corresponding to each exhaled gas includes: Determine the numerical and non-numerical data in the sampling parameters and sampling object information corresponding to each of the exhaled gases; The numerical data corresponding to each exhaled gas is normalized to determine the value sequence corresponding to each exhaled gas; A state mapping is performed on the non-numerical data corresponding to each exhaled gas to determine the state sequence corresponding to each exhaled gas.

6. The method as described in claim 5, characterized in that, Before determining the sample similarity network corresponding to the sample dataset based on the similarity between the feature sequences corresponding to each of the exhaled gases, the method further includes: Determine the similarity between the value sequences corresponding to each of the exhaled gases, and the similarity between the state sequences corresponding to each of the exhaled gases; The similarity between the feature sequences corresponding to each exhaled gas is determined based on the similarity between the value sequences corresponding to each exhaled gas and the similarity between the state sequences corresponding to each exhaled gas.

7. A device for screening respiratory biomarkers for diseases, characterized in that, include: The first acquisition module is used to acquire a sample dataset corresponding to exhaled gas, wherein the sample dataset includes multiple respiratory fingerprints, sampling parameters, and sampling object information corresponding to the exhaled gas, and the sampling object information includes real disease subtyping; The first determining module is used to determine a feature sequence corresponding to each exhaled gas based on the sampling parameters corresponding to each exhaled gas and the sampling object information corresponding to each exhaled gas. The feature sequence corresponding to the exhaled gas is used to characterize the features of the sampling parameters and sampling object information corresponding to the exhaled gas. The feature sequence corresponding to the exhaled gas is generated by combining the sampling parameters and sampling object information corresponding to the exhaled gas. The second determining module is used to determine the sample similarity network corresponding to the sample dataset based on the similarity between the feature sequences corresponding to each of the exhaled gases. The third determining module is used to determine the expression spectrum of volatile organic compounds (VOCs) contained in each exhaled gas according to the respiratory fingerprint corresponding to each exhaled gas. The expression spectrum of VOCs is generated based on the VOC content in all the respiratory fingerprints in the sample dataset. There are n VOCs involved in each respiratory fingerprint. The expression spectrum corresponding to the i-th VOC contains the content of the i-th VOC in each respiratory fingerprint. n is an integer greater than or equal to 1, and i is an integer greater than or equal to 1 and less than or equal to n. The fourth determination module is used to determine the enzyme network corresponding to each VOC based on the metabolic pathways associated with each VOC and the enzyme list associated with each metabolic pathway; to use an artificial neural network to perform clustering on each VOC based on the expression profile and enzyme network corresponding to each VOC to determine the cluster center corresponding to each VOC; to obtain the gene transcription pathways related to the target genes of each disease subtype from the Kyoto Encyclopedia of Genes and Genomes (KEGG) database; to perform enrichment analysis on the gene transcription pathways corresponding to each disease subtype to determine whether there is an association between the gene transcription pathways corresponding to each disease subtype and the metabolic pathways corresponding to each VOC; if there is an association between the metabolic pathways corresponding to the VOCs and the gene transcription pathways corresponding to any disease subtype, then it is determined that the VOCs are associated with any disease subtype, and the VOCs are retained and the weight of the cluster centers corresponding to the VOCs is increased for the next clustering; if the metabolic pathways corresponding to the VOCs and the gene transcription pathways corresponding to each disease subtype are associated ... If no correlation exists between the gene transcription pathways corresponding to the disease subtypes, the VOCs are removed, or the weights of the cluster centers corresponding to the VOCs are reduced. The parameters of the clustering model are adjusted based on the adjusted weights of each cluster center, and the adjusted clustering model is used to continue clustering the remaining VOCs until each VOC corresponding to a cluster center matches a gene transcription pathway corresponding to any disease subtype. This completes the VOC clustering process, and the final selected VOCs are identified as candidate respiratory biomarkers. The disease subtypes associated with the cluster centers corresponding to each candidate respiratory biomarker are identified as the disease subtypes corresponding to each candidate respiratory biomarker, where each cluster center corresponds to one disease subtype. A VOCs similarity network is determined based on the distance between the cluster centers corresponding to each candidate respiratory biomarker, where the VOCs similarity network is constructed using the pairwise similarities between all candidate respiratory biomarkers. The first screening module determines a relationship matrix between the sample dataset and the candidate respiratory biomarkers based on the candidate respiratory biomarkers contained in each exhaled gas, wherein the relationship matrix is ​​used to characterize the correlation between the sample dataset and the candidate respiratory biomarkers; divides each exhaled gas in the sample dataset into training samples and prediction samples; determines a first validity vector for each training sample based on the actual disease type corresponding to each training sample, the disease type corresponding to each candidate respiratory biomarker, and the sequence number of each training sample in the sample dataset, wherein the first validity vector is used to represent each candidate respiratory biomarker related to the actual disease type corresponding to the training sample and the sequence number of the training sample in the sample dataset; and determines a first validity vector for each prediction sample based on the actual disease type corresponding to the training sample and the sequence number of the training sample in the sample dataset. The sequence number in the dataset is used to determine the initial validity vector corresponding to each predicted sample. A preset label propagation algorithm is used to process the sample similarity network, the VOCs similarity network, the relation matrix, each of the first validity vectors, and each of the initial validity vectors to determine the prediction validity vector corresponding to each predicted sample. The prediction validity vector represents each candidate respiratory biomarker related to the predicted disease type of the predicted sample and the sequence number of the predicted sample in the sample dataset. Based on the prediction validity vector corresponding to each predicted sample, the predicted disease type corresponding to each predicted sample is determined. Based on the matching degree between the predicted disease type corresponding to each predicted sample and the actual disease type, the target respiratory biomarker related to each disease type is determined from the candidate respiratory biomarkers.

8. A terminal device, comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that, When the processor executes the computer program, it implements the method as described in any one of claims 1 to 6.

Citation Information

Patent Citations

  • Prediction method for correlation between circular RNA and disease based on gradient enhancement decision-making tree

    CN110459264A

  • Early lung cancer screening device based on detection of volatile organic compounds in exhaled gas

    CN114137158A