Feature selection method and system for expiration analysis
By dividing feature groups and selecting representative features in expiratory analysis, combined with the twin support vector machine recursive feature elimination algorithm, the problem of reducing importance scores caused by highly correlated features in expiratory analysis is solved, and more accurate feature selection and model performance optimization is achieved.
Patent Information
- Application Number
- CN202510353062.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-25
- Publication Date
- 2025-07-11
AI Technical Summary
In the existing exhalation analysis, due to the cross-response characteristics of the gas sensor, the generated feature set has high dimensionality and high correlation, resulting in a reduced distinction between features, affecting the performance of the machine learning model and the accuracy of disease identification.
By obtaining the correlation of the feature set of exhaled samples, dividing it into multiple feature groups, and selecting representative features and inherited features. The twin support vector machine recursive feature elimination algorithm and feature elimination algorithm are used to select feature, and combining membership and correlation matrix optimization, grouping and sorting are automatically completed.
It effectively avoids mutual interference between highly correlated features, improves the accuracy of feature selection and model performance, optimizes the dimensionality reduction of feature data, and enhances the interpretability and robustness of the analysis results.
Smart Images

Figure CN120296375A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of electronic nose signal and information processing, and particularly relates to a feature selection method and system for exhaled breath analysis. Background Art
[0002] With the continuous progress of electronic nose technology, exhaled breath analysis has shown remarkable effects in clinical diagnosis, especially in the fields of diabetes, kidney diseases, and respiratory diseases. The electronic nose system captures the characteristics of exhaled breath samples through a gas sensor array and generates a feature set. However, due to the cross-response characteristics of gas sensors, the generated feature set often has the characteristics of high dimensionality and high correlation. This not only increases the complexity of data processing but also reduces the discrimination between features, making it difficult to accurately identify the features that are truly meaningful for diagnosis.
[0003] Current wrapper-based feature selection methods generally adopt an iterative strategy. First, a machine learning model is used to calculate the importance scores of each feature. Subsequently, the features are sorted according to the importance scores, and the features with the lowest importance scores, that is, the features considered to be the least important, are gradually removed until the predetermined number of features is reached or the model performance no longer improves significantly.
[0004] However, there is a potential problem with current feature selection methods, namely correlation bias. Specifically, when features are highly correlated, these methods tend to give relatively low importance scores to highly correlated feature groups, even if some of these features may be crucial for disease diagnosis. This bias directly affects the performance of the machine learning model, resulting in a decrease in the accuracy of disease identification. Summary of the Invention
[0005] To solve the above problems, the present invention discloses a feature selection method and system for exhaled breath analysis.
[0006] The present invention discloses a feature selection method for exhaled breath analysis, including the following steps:
[0007] Obtain a feature set of exhaled breath samples, and determine the correlation between every two exhaled breath features in the exhaled breath sample feature set;
[0008] Divide all exhaled breath features into multiple exhaled breath feature groups according to the correlation, and determine the representative feature of each exhaled breath feature group according to a first preset rule;
[0009] Determine the membership degree of each exhaled breath feature to each representative feature, and determine the inherited features corresponding to each representative feature according to the membership degree;
[0010] Sort the inherited features according to a feature elimination algorithm and a second preset rule, and use the inherited features and their sorting as the feature selection result.
[0011] Preferably, the membership degrees of each exhalation feature and each representative feature are determined, and the inheritance features corresponding to each representative feature are determined according to the membership degrees, specifically as follows:
[0012] Determine the first membership degree of each exhalation feature and the representative features of this group;
[0013] Determine the second membership degrees of each exhalation feature and each of the other representative features one by one;
[0014] Using an optimization algorithm, select the exhalation feature with the largest first membership degree and the smallest sum of second membership degrees within each exhalation feature group as the inheritance feature corresponding to each group of representative features.
[0015] Preferably, all exhalation features are divided into multiple exhalation feature groups according to the correlation, specifically as follows:
[0016] Determine the grouping order of each exhalation feature according to the exhalation sample feature set;
[0017] Determine the current grouping feature and the feature to be grouped according to the grouping order;
[0018] Divide the current grouping feature and the features to be grouped whose correlation with the current grouping feature is greater than the first threshold into one exhalation feature group.
[0019] Preferably, the correlation is the correlation between every two exhalation features under the influence of other exhalation features;
[0020] Correspondingly, determine the correlation between every two exhalation features in the exhalation sample feature set, specifically as follows:
[0021] Determine the adjacency matrix according to a plurality of preset correlation evaluation indexes;
[0022] Optimize the adjacency matrix according to power iteration and geometric series attributes to obtain a feature correlation matrix;
[0023] Determine the correlation between every two exhalation features in the exhalation sample feature set according to the feature correlation matrix.
[0024] Preferably, determine the representative features of each exhalation feature group according to the first preset rule, specifically as follows:
[0025] Take the mean value of all exhalation features in each exhalation feature group as the representative feature of this exhalation feature group.
[0026] Preferably, sort the inheritance features according to the feature elimination algorithm and the second preset rule, specifically as follows:
[0027] According to the feature elimination algorithm and the second preset rule, iteratively eliminate all representative features to obtain the order in which each representative feature is eliminated;
[0028] Determine the importance ranking of the corresponding inherited features according to the order or reverse order in which each representative feature is eliminated.
[0029] Preferably, according to the feature elimination algorithm and the second preset rule, iteratively eliminate all representative features to obtain the order in which each representative feature is eliminated. Specifically:
[0030] Obtain the representative features that have not been eliminated before this iteration, denoted as surviving representative features;
[0031] Determine the ranking of multiple surviving representative features according to the feature elimination algorithm;
[0032] According to the ranking of the multiple surviving representative features and the second preset rule, iteratively eliminate all representative features to obtain the order in which each representative feature is eliminated.
[0033] Preferably, the second preset rule is specifically:
[0034] When the surviving number of the surviving representative features is greater than or equal to the second preset threshold, the number of representative features eliminated in this iteration is the preset ratio of the surviving number;
[0035] When the surviving number is less than the second preset threshold, eliminate a preset number of representative features in this iteration.
[0036] The present invention also discloses a feature selection system for exhaled breath analysis, including:
[0037] An acquisition module for acquiring an exhaled breath sample feature set and determining the correlation between every two exhaled breath features in the exhaled breath sample feature set;
[0038] A grouping module for dividing all exhaled breath features into multiple exhaled breath feature groups according to the correlation and determining the representative feature of each exhaled breath feature group according to the first preset rule;
[0039] An inheritance module for determining the membership degree of each exhaled breath feature to each representative feature and determining the inherited feature corresponding to each representative feature according to the membership degree;
[0040] A selection module for sorting the inherited features according to the feature elimination algorithm and the second preset rule and taking the inherited features and their sorting as the feature selection result.
[0041] Compared with the prior art, the present invention has the following beneficial effects:
[0042] (1) The present invention groups based on the correlation between exhalation features and selects representative features to participate in the subsequent feature selection process, which can effectively avoid the mutual interference of redundant features that are highly similar to other features during feature selection, and effectively solves the problem of reduced importance scores caused by the mutual influence between highly correlated features.
[0043] (2) The present invention selects the feature with the highest membership degree to this representative feature and the lowest sum of membership degrees to other representative features to participate in the subsequent exhalation analysis process. Therefore, the features selected by the present invention can well represent the overall characteristics of exhalation features and avoid mutual interference between different groups of features. Thus, the features selected by the method of the present invention are more accurate.
[0044] (3) The present invention simulates the correlation between exhalation features based on multiple correlation evaluation indicators and can automatically complete grouping based on the correlation. The grouping is more accurate and the operation is simple, without the need to artificially preset the number of groups and the number of features within each group.
[0045] (4) Based on the above feature selection method and feature grouping method, the present invention can be effectively used for dimensionality reduction of feature data in exhalation analysis, optimizing the model performance, and improving the interpretability of analysis results. Description of the Drawings
[0046] Figure 1 is the feature importance of the artificial dataset calculated using formula (3) in TWSVM, and the results have been normalized;
[0047] Figure 2 is the feature importance of the artificial dataset calculated using formula (4) in TWSVM;
[0048] Figure 3 is the flowchart of the operation of TWSVM - RFE - ICC;
[0049] Figure 4 is the schematic diagram of the electronic nose system;
[0050] Figure 5 is the flowchart of the method of the present invention. Detailed Embodiments
[0051] In the following description, specific details such as specific system structures and technologies are presented for the purpose of illustration rather than limitation, so as to thoroughly understand the embodiments of the present invention. However, those skilled in the art should clearly understand that the present invention can also be implemented in other embodiments without these specific details. In other cases, the detailed descriptions of well-known systems, devices, circuits, and methods are omitted to avoid unnecessary details from interfering with the description of the present invention.
[0052] The feature selection method mentioned in the background art of the present invention is an improvement based on the Recursive Feature Elimination (RFE) method. For the convenience of subsequent description, the Twin Support Vector Machine Recursive Feature Elimination (TWSVM-RFE) of the prior art is briefly introduced as follows:
[0053] The Twin Support Vector Machine (TWSVM) is a fast and accurate advanced classifier that uses a pair of non-parallel approximate hyperplanes for classification as follows:
[0054]
[0055] where w 1,2 and b 1,2 are the weights and bias terms of the hyperplane. By defining this pair of hyperplanes, the samples of one class are placed near its corresponding hyperplane and away from the other hyperplane at the same time. These two hyperplanes can be defined by solving the following two Quadratic Programming Problems (QPPs):
[0056]
[0057] and
[0058]
[0059] where A and B represent the positive and negative class samples respectively, c1, c2, c3, c4 > 0 are penalty parameters, ξ1 and ξ2 are slack variables, and e1, e2 are all-ones vectors.
[0060] In the Twin Support Vector Machine Recursive Feature Elimination (TWSVM-RFE), the absolute value of w 1,2 reflects the importance or contribution of the feature to the classification result. For the importance of feature k, it can be represented by the square value of the k-th element in W as follows:
[0061]
[0062] and
[0063]
[0064] where ||·||2 is the 2-norm and W is the sum of the weights of the two hyperplanes after normalization.
[0065] That is to say, the present invention first obtains the parameters W1, W2, b1, b2 of the TWSVM model by the methods described in formulas (3) and (4), and the TWSVM model is as shown in formulas (1) and (2);
[0066] Then all the exhaled features in the exhaled sample feature set are input into the model to obtain the importance of each feature as shown in formula (5), and then the feature with the lowest score is proposed;
[0067] Retrain a new TWSVM model based on the surviving features, and according to the feature ranking criteria described in formula (5), eliminate the features with the lowest scores. The remaining features participate in the next iteration. Therefore, the features eliminated in later iterations are considered more important. Obviously, eliminating multiple features in one iteration can effectively reduce the computational cost, but may further exacerbate the impact of the correlation bias, thereby reducing the performance of the algorithm.
[0068] The feature selection method of the present invention is an improvement based on the infinite feature selection method of the prior art.
[0069] The infinite feature selection method regards features as nodes in an undirected fully connected weighted graph, and the custom relationship between features is represented as an edge. Therefore, any feature subset of any size can be represented by a path in it. Using the power series of the matrix, the scores of paths of any length can be conveniently calculated.
[0070] The graph can be represented by its adjacency matrix, where the element represents the relationship between feature and, and can be defined as:
[0071]
[0072] where s i and s j represent the scores of f i and f j respectively, and can be calculated using a custom metric, such as the Fisher criterion or mutual information.
[0073] The correlation in formula (7) represents the independent correlation between any two exhalation features, that is, the correlation in the absence of the influence of other exhalation features.
[0074] Make represent a path of length l between node and node , and the total score W P of this path should be:
[0075]
[0076] Because there are multiple paths of length l between node and , assuming there is a set that contains all paths that meet the constraints, based on standard matrix algebra, the total score of should be:
[0077]
[0078] where Al is the power iteration of the adjacency matrix. Therefore, for a certain feature f in a path of length l i the score should be:
[0079] c l (i) = ∑ j∈V A l (i,j) (10)
[0080] where V is the set of all nodes in graph G. Therefore, when l is equal to the total number of nodes, i.e., the total number of features, the importance of all features can be obtained by arranging c l in descending order. The larger the value, the more important the feature is considered. By expanding the length l to infinity and using the geometric series property of the adjacency matrix, the calculation of c l can be greatly simplified as follows:
[0081]
[0082] where r is the regularization parameter used to ensure the convergence of A l .
[0083] Based on the convergence property of the adjacency matrix, Equation 11 can be rewritten as follows:
[0084] c(i) = [((I - rA) -1 - I)e] i (12)
[0085] where c(i) is the importance score of feature f i , I is the identity matrix, e is the all-ones vector, and [·] i represents extracting the i-th element in the matrix.
[0086] Preferably, before step S1, the present invention studies the phenomenon of correlation deviation. Specifically, the present invention constructs an artificial dataset for experiments. The artificial dataset consists of 500 samples and 140 features. These features are artificially divided into 24 highly correlated groups, each group having different predefined importance weights and group sizes, as shown in Table 1. (For example, the column corresponding to feature indices 21 - 40 can be understood as that among the 21st to 40th features, there are 2 highly correlated groups, each group containing 10 highly correlated features with each feature having a predefined importance of 1.)
[0087] Table 1 Feature indices, importance weights, and group sizes of highly correlated feature groups in the artificial dataset
[0088]
[0089] Specifically, the process of artificially dividing 24 highly correlated groups is as follows:
[0090] First, 24 different prototype vectors, i.e., 1, …, 24, were sampled from the normal distribution N(0, 1). By adding a perturbation term sampled from the normal distribution N(0, 0.01) to a prototype vector, a set of highly correlated features can be generated, thus forming a highly correlated group as follows:
[0091] f ij = t i + ∈, j = 1, …, k i , i = 1, …, 24 (13)
[0092] where k i is the group size of the i-th group, i.e., how many mutually correlated features the group contains. Second, the sample label y is determined based on the prototype vectors and their importance weights as follows:
[0093]
[0094] where η is a perturbation term randomly sampled from the normal distribution N(0, 0.01). According to the predefined importance weights, when ω = 1, the feature is considered strongly correlated with the label; when ω = 0.5, the feature is considered weakly correlated with the label; and when ω = 0, the feature is considered noise.
[0095] Finally, the features of all the generated highly correlated groups and their corresponding labels are combined to form a complete artificial dataset containing 24 highly correlated groups.
[0096] The present invention evaluates the importance of features in the artificial dataset using the twin support vector machine (TWSVM), but does not apply the RFE process.
[0097] As Figure 1 , Figure 2 shown, the feature importance of the artificial dataset calculated using formulas (3) and (4) of TWSVM respectively is presented, where the feature importance is normalized for comparison. The line chart illustrates the importance of each feature, and the bar chart illustrates the average importance of the highly correlated feature groups.
[0098] The results clearly show that the importance of features is significantly affected by the number of their correlated features. As the size of the correlated group increases, the average importance of features within the group decreases, clearly indicating the negative impact of CB on TWSVM-RFE. For example, when the group size is 20, even if the predefined importance weight is 1, the features in the group are almost regarded as noise. Therefore, when the number of features eliminated in each iteration is large enough, it may lead to the deletion of the entire correlated group in a single iteration of the RFE process. Therefore, when dealing with highly correlated features, the presence of CB may lead to the incorrect elimination of important features in TWSVM-RFE, resulting in a deviation in the feature results.
[0099] As Figure 5 shown, the present invention discloses a feature selection method for breath analysis, comprising the following steps:
[0100] S1. Obtain a breath sample feature set and determine the correlation between every two breath features in the breath sample feature set;
[0101] The breath sample feature set is collected by a sensor array and the features are extracted, and the feature extraction includes at least one of transient feature extraction, frequency domain feature extraction, and decomposed signal feature extraction.
[0102] In one embodiment, the correlation refers to the correlation between every two breath features under the influence of other breath features, and for the sake of distinction, it is denoted as indirect correlation.
[0103] To determine the correlation between every two breath features in the breath sample feature set, the following is specifically as follows:
[0104] Determine an adjacency matrix according to a plurality of preset correlation evaluation indexes;
[0105] To calculate the adjacency matrix, the present invention constructs an undirected fully connected weighted graph, where the nodes represent features. The relationship (variable) between nodes is defined as a weighted linear combination of two correlation evaluation indexes, which is defined as follows:
[0106]
[0107] where Sp is the Spearman rank correlation coefficient and Pe is the Pearson correlation coefficient. The value range of the coefficient α is [0, 1], which specifically depends on whether the features are linearly correlated or non-linearly correlated, as well as the content of noise and outliers, and can also be optimized and selected through cross-validation. In the present invention, the coefficient α is used to balance the contribution ratio of various customizable correlation evaluation methods.
[0108] It should be noted that the correlation coefficient in formula (15) is customizable and can be replaced with other metrics according to the specific requirements of the application, such as the Kendall rank correlation coefficient or mutual information.
[0109] Optimize the adjacency matrix according to the power iteration and geometric series properties to obtain the feature correlation matrix m c ;
[0110] According to the feature correlation matrix m c Determine the correlation between every two exhalation features in the exhalation sample feature set.
[0111] After defining the calculation method of the relationship between features, each indirect correlation can be calculated through the geometric series of the adjacency matrix, based on formula (16).
[0112] c = ((I - rA) -1 - I)e (16)
[0113] Where c is the feature - to - feature correlation matrix, I is the identity matrix, and e is the all - ones vector.
[0114] Formula (16) in the present invention is an approximation of formula (12), but in the present invention, c represents the indirect correlation between every two exhalation features, while in formula (12), c represents the importance score, and their meanings are different.
[0115] S2. Divide all exhalation features into multiple exhalation feature groups according to the correlation, and determine the representative feature of each exhalation feature group according to the first preset rule;
[0116] Preferably, dividing all exhalation features into multiple exhalation feature groups according to the correlation is specifically as follows:
[0117] S21. Determine the grouping order of each exhalation feature according to the exhalation sample feature set; when extracting the exhalation sample feature set, an order has been automatically formed for each exhalation feature.
[0118] S22. Determine the current grouped feature and the feature to be grouped according to the grouping order;
[0119] S23. Divide the current grouped feature and the features to be grouped whose correlation with the current grouped feature is greater than the first threshold into one exhalation feature group.
[0120] In the present invention, each exhalation feature is sorted in turn according to the grouping order. If an exhalation feature ranked later has already been grouped with a certain feature ranked earlier, it will be excluded during subsequent grouping, that is, each feature appears only once in the grouping and will not be repeatedly grouped.
[0121] T cThe first threshold for determining whether a pair of features is relevant during the grouping process can be specified by those skilled in the art. If it is too large, it may damage the ability of the present invention to process CB, while if it is too small, it may lead to the loss of key information.
[0122] Preferably, the representative feature of each exhalation feature group is determined according to the first preset rule, specifically:
[0123] The mean value of all exhalation features in each exhalation feature group is used as the representative feature of the exhalation feature group.
[0124] In other embodiments, those skilled in the art can also specify the representative feature by themselves, such as using the feature ranked first in importance as the representative feature.
[0125] S3. Determine the membership degree of each exhalation feature to each representative feature, and determine the inherited feature corresponding to each representative feature according to the membership degree. After obtaining the representative feature, the present invention determines the importance of the exhalation features within the relevant group according to the membership degree of each exhalation feature to the representative feature of the group to which it belongs. The higher the membership degree value, the more important it is.
[0126] Preferably, S3 is specifically:
[0127] S31. Determine the first membership degree of each exhalation feature to the representative feature of this group respectively;
[0128] S32. Determine the second membership degree of each exhalation feature to each other representative feature one by one;
[0129] S33. Use an optimization algorithm to select the exhalation feature with the largest first membership degree and the smallest sum of second membership degrees within each exhalation feature group as the inherited feature corresponding to each group's representative feature.
[0130] The second membership degree corresponding to each exhalation feature is multiple.
[0131] The membership degree between a certain exhalation feature and a certain representative feature is calculated according to the following formula:
[0132]
[0133] Where i = 1,..., N and j = 1,..., M, N is the total number of features, and M is the total number of relevant groups; f i is a certain exhalation feature; c j is any representative feature; c m is other representative features.
[0134] The present invention selects the feature with the highest membership degree to the representative feature of each relevant group and the smallest sum of membership degrees to other representative features to inherit the ranking of its corresponding representative. The feature selection is optimized to ensure representativeness while reducing redundancy, improving the accuracy and efficiency of exhaled gas analysis of the electronic nose system.
[0135] In one embodiment, the feature selection method of the present invention shown in S1 - S3 is named infinite relevant clustering, and the running process is as follows:
[0136] Algorithm 1 Infinite Relevant Clustering.
[0137] Input: Feature list F; Weight coefficient α; Correlation threshold T c
[0138] Output: Representative feature list C; Inherited feature list F r
[0139] 1: Initialize the group list G and the membership degree matrix Mm
[0140] 2: M c ← Calculate the correlation matrix between features in F using α according to Formula 12 and Formula 15
[0141] 3: for f i belonging to F do
[0142] 4:
[0143] 5:
[0144] 6: G ← G ∪ G i
[0145] 7: end if
[0146] 8: end for
[0147] 9: for G i belonging to G do
[0148] 10: N ← the number of features included in group G i
[0149] 11: For f j ∈ G i
[0150] 12: C ← C ∪ C i
[0151] 13: end for
[0152] 14: for f i belonging to F do
[0153] 15: for G j belonging to G do
[0154] 16: M m [i, j] ← Calculate the membership degree between the feature and the representative feature according to formula 17
[0155] 17: end for
[0156] 18: end for
[0157] 19: for G i belonging to G do
[0158] 20: f i ← the feature with the highest membership degree to the representative feature of its own group in M m and the lowest total membership degree to other representative features
[0159] 21: F r ← F r ∪ f i
[0160] 22: end for
[0161] 23: Output C, F r
[0162] S4. According to the feature elimination algorithm and the second preset rule, sort the inherited features, and use the inherited features and their sorting as the feature selection result;
[0163] Preferably, sort the inherited features according to the feature elimination algorithm and the second preset rule, specifically:
[0164] S41. According to the feature elimination algorithm and the second preset rule, iteratively eliminate all representative features to obtain the order in which each representative feature is eliminated;
[0165] S41 is specifically:
[0166] S411. Obtain the representative features that have not been eliminated before this iteration, denoted as surviving representative features;
[0167] S412. Determine the sorting of multiple surviving representative features according to the feature elimination algorithm; The iterative elimination algorithm includes: Recursive Feature Elimination (RFE), Twin Support Vector Machine Recursive Feature Elimination (TWSVM-RFE) or SVM-RFE, etc.
[0168] S413. According to the sorting of multiple surviving representative features and the second preset rule, iteratively eliminate all representative features to obtain the order in which each representative feature is eliminated.
[0169] Preferably, the second preset rule is specifically as follows:
[0170] When the surviving number of surviving representative features is greater than or equal to the second preset threshold, the number of representative features eliminated in this iteration is a preset ratio of the surviving number; the preset ratio can be specified by those skilled in the art themselves.
[0171] When the surviving number is less than the second preset threshold, a preset number of representative features are eliminated in this iteration.
[0172] S42. Determine the importance ranking of the corresponding inherited features according to the order or reverse order of elimination of each representative feature.
[0173] In the present invention, the reverse order of elimination of each representative feature represents a gradual decrease in its importance degree, and the inherited features are sorted in sequence according to the importance degree of their corresponding representative features.
[0174] In one embodiment, the method of the present invention shown in S4 is named TWSVM-RFE-ICC, and the running process is as follows:
[0175] Algorithm 2 TWSVM-RFE-ICC
[0176] Input: Feature list F; Sample label L
[0177] Output: Feature ranking R f
[0178] 1: C, F r ← Calculate the representative feature list and the inherited feature list according to Algorithm 1
[0179] 2: Initialize the surviving representative feature index S f ← {1,..., N} and the eliminated representative feature list where N is the total number of related groups
[0180] 3:
[0181] 4: ω1, ω2 ← Train a new TWSVM model through the surviving representative features C[S f and the sample label L
[0182] 5: And sort the results in descending order
[0183] 6: K ← If |S f | is greater than the threshold then Otherwise K = 1
[0184] 7: E ← Select the K representative features with the lowest ranking according to W
[0185] 8: S f ←S f -E
[0186] 9: E f ←E f +E
[0187] 10: end while
[0188] 11: R C ← Reverse E f
[0189] 12: R f ← According to R C Reorder F r by having the inherited features take over the ranking of their respective representative features
[0190] 13: Output R f where the top features represent the most important features
[0191] As Figure 3 shown, it is the flowchart of the operation of TWSVM - RFE - ICC. In the present invention, the final feature ranking is determined by weighting the importance ranking of representative features and membership degrees. By comprehensively considering the importance ranking of representative features and membership degrees, the importance of each exhalation feature in the overall exhalation analysis can be evaluated more comprehensively. This comprehensive consideration method reduces the one - sidedness of single - index evaluation, making the final feature ranking more accurate and reliable. Furthermore, it contributes to the transparency of the model decision - making process. In exhalation analysis, it is possible to clearly understand which features have an important impact on the prediction results, thereby enhancing the interpretability of the model and facilitating subsequent analysis and optimization.
[0192] By comprehensively considering feature information in multiple dimensions such as representative features and membership degrees, the present invention improves the robustness of exhalation analysis against noise and abnormal data. Even in the face of complex and variable exhalation data, it can maintain high analysis accuracy and stability.
[0193] After step S4, the method of the present invention was compared with the prior art:
[0194] During the comparison process, the present invention uses an independently designed electronic nose system based on multi - channel gas sensors to collect the required respiration analysis data set, and its structure is as Figure 3As shown, it includes five main components: a sensor array, a signal processing circuit, a DAQ data acquisition card, an air pump, and a host computer. The sensor array is a key component of the electronic nose system, including a carbon dioxide sensor, six MOS sensors, and a humidity-temperature sensor, which can detect a wide range of target gases, including acetone, ammonia, hydrogen sulfide, carbon monoxide, and ethanol, so as to diagnose various diseases. The specific information is shown in Table 2. The MOS sensors are used to detect biomarkers in exhaled breath, the carbon dioxide sensor (TGS2161) is used to compensate for the change of alveolar air, and the humidity-temperature sensor (HTG3515CH) compensates for the influence of humidity on the sensor response.
[0195] For each sample, it takes 144 seconds (s) to complete the sampling, and the sampling is divided into four stages: baseline, injection, reaction, and purge. In the baseline stage (0 - 15 s), the system records the initial value of each sensor as the baseline, which is used to eliminate the influence of baseline drift during preprocessing. In the injection stage (1 - 8 s), the exhaled gas is pumped into the gas chamber, and the response signal of the sensor begins to change. In the reaction stage (8 - 64 s), the sensor interacts fully with the exhaled gas, and the response signal reaches the maximum value. In the purge stage (64 - 144 s), external air is pumped into the gas chamber for cleaning, and the response signal gradually returns to the baseline.
[0196] Referring to the above process, a total of 320 healthy samples, 192 diabetic samples, 268 nephropathy samples, 218 lung disease samples, and 211 liver disease samples were collected in the breath analysis dataset. Before feature extraction, the samples must be preprocessed. First, subtract the baseline from the samples, usually the average response in the baseline stage, to eliminate the influence of baseline drift. Second, establish a linear humidity response model to compensate for the humidity of the samples. Finally, use median filtering to remove high-frequency noise.
[0197] Table 2 Sensor models used in the electronic nose system
[0198]
[0199]
[0200] Feature extraction is an important step in the field of pattern recognition. Medical data, especially breath analysis data, usually contains a large amount of redundant information and noise. Feature extraction methods reduce the data dimension and improve the model performance by identifying meaningful information. In this paper, we used a feature set containing three types of features: transient features, frequency domain features, and signal decomposition features. A total of 17 feature extraction methods were used to comprehensively analyze the internal relationship of the data. Table 3 provides the detailed information of these methods.
[0201] Details of the transient features, frequency domain features, and signal decomposition features used in Table 3
[0202]
[0203]
[0204] Through these feature extraction methods, compared with previous studies, the present invention constructs a larger and more diverse feature set, enabling a wider exploration of features effective for respiratory analysis. Additionally, it helps to accurately evaluate the importance of sensors, providing key insights for the development of future systems. Finally, it can perform performance evaluation of the proposed method on a large and highly relevant feature set.
[0205] Since disease diagnosis through respiratory analysis belongs to clinical applications, this paper adopts the performance metrics of four widely recognized feature selection methods: the size of the optimal feature subset, accuracy, sensitivity, and specificity. Sensitivity represents the proportion of true disease cases accurately identified as positive cases by the model, as follows:
[0206]
[0207] And Specificity represents the proportion of true healthy samples accurately identified as negative cases, as follows:
[0208]
[0209] This part compares the actual performance of various Correlation Bias (CB) solutions on the artificial dataset. The establishment process of the artificial dataset is as described above. For the model parameters, the penalty parameters of TWSVM in the feature selection and classification processes are both set to c1 = c3 = 2 4 and c2 = c4 = 2 6 , to ensure a fair performance comparison. Through ten-fold cross-validation, in TWSVM-RFE-ICC, the weight α of the correlation evaluation method is searched within the range of [0, 1], and the correlation threshold T c is searched within the range of [0.9, 1]. For the comparison methods, in TWSVM-RFE-CBR, the correlation threshold is searched within [0.65, 0.95], and the group size threshold is searched from [1, 3]. In TWSVM-RFE-HC, the number of clusters is set to 42, measured according to the silhouette coefficient. In TWSVM-RFE-IRFS, the correlation threshold is searched within [0.9, 1], and the retained feature threshold is searched within [1, 30]. During the TWSVM-RFE process, according to the number of predefined highly correlated groups in the artificial dataset, the elimination threshold is set to 24.
[0210] Hereinafter, to avoid redundancy, TWSVM-RFE-CBR, TWSVM-RFE-HC, TWSVM-RFE, and TWSVM-RFE-IRFS are respectively abbreviated as CBR, HC, RFE, and IRFS.
[0211] Table 4 summarizes the distribution of feature importance obtained by five different RFE methods. In the synthetic dataset, there are 60 strongly correlated features with a weight of 1, 60 weakly correlated features with a weight of 0.5, and 20 noise features with a weight of 0. Obviously, in an ideal situation, the feature selection method should assign the highest importance to the features with a weight of 1, followed by the features with a weight of 0.5, and the lowest importance to the noise features with a weight of 0. However, the results show that due to the negative impact of CB, for TWSVM-RFE, 25 features with a weight of 1 are wrongly assigned importance, and 4 of them are wrongly classified as noise. In addition, 18 noise features are given a lot of attention, which may significantly affect the feature selection results and the rules summarized from them.
[0212] Table 4 shows the distribution of the number of features with different predefined weights in three intervals.
[0213]
[0214]
[0215] Interval 1 represents the importance from the 1st to the 60th, interval 2 represents the importance from the 61st to the 120th, and interval 3 represents the importance from the 121st to the 140th. Among various CB solutions, the ICC method we proposed successfully assigns the correct importance to all features and achieves the highest accuracy in subsequent classification experiments, as shown in Table 5. This shows that ICC can accurately group highly correlated features, effectively mitigate the impact of CB on TWSVM, and the feature ranking it provides can improve the performance of subsequent tasks.
[0216] For other CB solutions, similar to ICC, IRFS also successfully assigned the correct importance to all features. However, in the classification experiment, although the retained feature threshold reduced the computational cost, it also limited the algorithm's ability to obtain the optimal result. In the case of HC, due to the error in determining the number of relevant groups, a small number of features were misassigned. However, compared with the original TWSVM-RFE, it still effectively alleviated the impact of CB. CBR is a special case because its workflow is that when the entire relevant group is deleted, the most important feature in the group is put back into the retained list without addressing the importance bias of other features. Therefore, CBR did not provide any improvement to the distribution of feature importance. However, since the classifier only needs to extract a small number of features from each relevant group for classification, CBR performed satisfactorily in the classification task and ranked second only to our method.
[0217] Table 5 Comparison of classification performance of different TWSVM-RFE improvement strategies
[0218]
[0219]
[0220] For the respiratory analysis dataset, the experimental configuration is the same as before. During the RFE process, to balance performance and computational cost, the elimination threshold was set to 100. In ICC, the weight α of the correlation evaluation method was searched within the range of [0, 1], and the correlation threshold T c was searched within the range of [0.7, 1]. All experiments used ten-fold cross-validation.
[0221] Tables 6 to 9 show the performance of our proposed ICC on the real gas sensor dataset and compare it with three existing CB solutions. It can be observed that when using all features, although the computational cost is high, the accuracy is still sub-optimal. This indicates that although the feature set contains valuable classification information, it also includes some redundant and even irrelevant information. This also emphasizes the importance of the feature selection step in high-dimensional data analysis.
[0222] For respiratory analysis data, the CB generated by highly correlated features not only seriously affects the knowledge that can be summarized from the feature selection results, but also reduces the classification accuracy. In all four disease diagnosis tasks, ICC effectively improves the accuracy of TWSVM-RFE and is always superior to other CB solutions. Compared with the original TWSVM-RFE, ICC increases the size of the optimal feature subset. This is because in the original algorithm, some important features were wrongly assigned low ranks, and noise took their places. Therefore, when the optimal subset tries to include all important features, it has to cover some of the top-ranked noise, which reduces the accuracy rate, causing the search for the optimal subset to fall into a local optimum and thus reducing the size of the optimal subset. As expected, the classification accuracy rate of CBR is superior to that of HC and IRFS in most cases. This is because CBR focuses on identifying non-redundant features that contribute to classification rather than correcting feature ranking. Among the solutions based on related groups, the accuracy of HC ranks second only to ICC. However, for HC, how to determine the optimal number of clusters in advance remains a challenging task. ICC and IRFS avoid this challenge because they can both automatically group similar features together. However, compared with ICC, the accuracy of IRFS is lower than the baseline because the retained feature threshold forces the algorithm to always sacrifice accuracy to improve the calculation speed.
[0223] Table 6 Performance comparison of algorithms in distinguishing healthy samples and diabetic samples
[0224] Algorithm Optimal Dimension Sensitivity (%) Specificity (%) Accuracy Rate (%) TWSVM 1536 94.76±4.71 99.06±2.00 97.45±1.77 TWSVM-RFE 38 95.82±4.58 99.38±1.25 98.04±1.75 TWSVM-RFE-CBR 32 95.82±3.15 100.00±0.00 98.44±1.18 TWSVM-RFE-HC 40 95.29±4.37 99.69±0.94 98.04±1.52 TWSVM-RFE-IRFS 80 93.18±5.30 99.69±0.94 97.26±1.80 TWSVM-RFE-ICC 56 97.39±1.30 99.69±0.94 98.83±1.30
[0225] Table 7 Performance comparison of algorithms in distinguishing healthy samples and nephropathy samples
[0226] Algorithm Optimal Dimension Sensitivity (%) Specificity (%) Accuracy Rate (%) TWSVM 1536 76.45±1.10 88.44±5.05 82.98±5.59 TWSVM-RFE 43 84.34±7.37 90.31±5.84 87.59±3.56 TWSVM-RFE-CBR 80 83.23±5.28 90.63±4.84 87.25±3.75 TWSVM-RFE-HC 51 86.92±6.78 90.00±5.19 88.61±4.24 TWSVM-RFE-IRFS 60 76.07±8.44 86.56±5.42 81.76±4.92 TWSVM-RFE-ICC 82 87.66±5.64 89.69±4.43 88.77±4.19
[0227] Table 8 Performance comparison of algorithms in distinguishing healthy samples and lung disease samples
[0228]
[0229]
[0230] Table 9 Performance comparison of algorithms in distinguishing healthy samples and liver disease samples
[0231] Algorithm Optimal Dimension Sensitivity (%) Specificity (%) Accuracy Rate (%) TWSVM 1536 66.92±9.44 82.19±6.41 75.95±5.35 TWSVM-RFE 50 78.72±5.82 84.38±8.62 82.06±4.50 TWSVM-RFE-CBR 40 79.17±6.21 84.69±6.16 82.43±4.27 TWSVM-RFE-HC 57 81.90±7.88 84.38±5.04 83.36±4.63 TWSVM-RFE-IRFS 40 66.01±8.08 82.19±6.26 75.58±5.15 TWSVM-RFE-ICC 37 83.24±8.90 85.63±8.97 84.64±4.33
[0232] The present invention also discloses a feature selection system for exhaled breath analysis, comprising:
[0233] An acquisition module, configured to acquire an exhaled breath sample feature set and determine the correlation between every two exhaled breath features in the exhaled breath sample feature set;
[0234] A grouping module, configured to divide all exhalation features into multiple exhalation feature groups according to relevance, and determine a representative feature for each exhalation feature group according to a first preset rule;
[0235] An inheritance module, configured to determine the membership degree of each exhalation feature to each representative feature, and determine an inheritance feature corresponding to each representative feature according to the membership degree;
[0236] A selection module, configured to sort the inheritance features according to a feature elimination algorithm and a second preset rule, and use the inheritance features and their sorting as a feature selection result.
[0237] Non-invasive disease diagnosis through breath analysis is an important application of an electronic nose system. Due to the cross-response characteristics of gas sensors, a breath analysis data set usually contains highly correlated features. Therefore, the method proposed by the present invention will be tested on a breath analysis data set collected by us.
[0238] Compared with the prior art, the present invention has the following beneficial effects:
[0239] (1) The present invention groups based on the relevance between exhalation features and selects representative features to participate in the subsequent feature selection process, which can effectively avoid the mutual interference of redundant features that are highly similar to other features in feature selection, and effectively solve the problem of reduced importance scores caused by the mutual influence between highly correlated features;
[0240] (2) The present invention selects the feature with the highest membership degree to the representative feature and the lowest sum of membership degrees to other representative features to participate in the subsequent exhalation analysis process. Therefore, the features selected by the present invention can well represent the overall characteristics of exhalation features and avoid mutual interference between different groups of features. Therefore, the features selected by the method of the present invention are more accurate;
[0241] (3) The present invention simulates the relevance between exhalation features based on multiple relevance evaluation indicators, and can automatically complete grouping based on relevance. The grouping is more accurate and the operation is simple, without the need to artificially preset the number of groups and the number of features within the group;
[0242] (4) Based on the above feature selection method and feature grouping method, the present invention can be effectively used for feature data dimensionality reduction in exhalation analysis, optimize model performance, and improve the interpretability of analysis results.
[0243] The above are only several embodiments of the present application and do not impose any form of limitation on the present application. Although the present application is disclosed above with preferred embodiments, it is not intended to limit the present application. Any person skilled in the relevant art can make some changes or modifications within the scope of the technical solution of the present application by using the disclosed technical content, which are all equivalent to equivalent embodiments and fall within the scope of the technical solution.
Claims
1. A feature selection method for exhaled breath analysis, characterized in that, It includes the following steps: Obtain an exhaled breath sample feature set, and determine the correlation between every two exhaled breath features in the exhaled breath sample feature set; Divide all exhaled breath features into multiple exhaled breath feature groups according to the correlation, and determine the representative feature of each exhaled breath feature group according to a first preset rule; Determine the membership degree of each exhaled breath feature to each representative feature, and determine the inherited feature corresponding to each representative feature according to the membership degree; Sort the inherited features according to a feature elimination algorithm and a second preset rule, and use the inherited features and their sorting as the feature selection result.
2. The feature selection method for breath analysis according to claim 1, wherein, Determine the membership degree of each exhaled breath feature to each representative feature, and determine the inherited feature corresponding to each representative feature according to the membership degree. Specifically: Determine the first membership degree of each exhaled breath feature to the representative feature of this group; Determine the second membership degree of each exhaled breath feature to each other representative feature one by one; Use an optimization algorithm to select, from within each exhaled breath feature group, the exhaled breath feature with the largest first membership degree and the smallest sum of second membership degrees as the inherited feature corresponding to the representative feature of each group.
3. The feature selection method for breath analysis according to claim 1, wherein Divide all exhaled breath features into multiple exhaled breath feature groups according to the correlation. Specifically: Determine the grouping order of each exhaled breath feature according to the exhaled breath sample feature set; Determine the current grouped feature and the to-be-grouped feature according to the grouping order; Divide the current grouped feature and the to-be-grouped features whose correlation with the current grouped feature is greater than a first threshold into one exhaled breath feature group.
4. The feature selection method for breath analysis according to claim 1, characterized in that, The correlation is the correlation between every two exhaled breath features under the influence of other exhaled breath features; Correspondingly, determine the correlation between every two exhaled breath features in the exhaled breath sample feature set. Specifically: Determine an adjacency matrix according to a plurality of preset correlation evaluation indexes; Optimize the adjacency matrix according to power iteration and geometric series properties to obtain a feature correlation matrix; Determine the correlation between every two exhaled breath features in the exhaled breath sample feature set according to the feature correlation matrix.
5. The feature selection method for exhaled gas analysis according to claim 1, characterized in that, Determine the representative feature of each exhaled breath feature group according to a first preset rule. Specifically: Use the mean value of all exhaled breath features in each exhaled breath feature group as the representative feature of this exhaled breath feature group.
6. The feature selection method for breath analysis according to claim 1 or 2, characterized in that, Sort the inherited features according to a feature elimination algorithm and a second preset rule. Specifically: Iteratively eliminate all representative features according to the feature elimination algorithm and the second preset rule to obtain the elimination order of each representative feature; Determine the importance sorting of the corresponding inherited features according to the elimination order or reverse order of each representative feature.
7. The feature selection method for exhaled gas analysis according to claim 6, characterized in that, Iteratively eliminate all representative features according to the feature elimination algorithm and the second preset rule to obtain the elimination order of each representative feature. Specifically: Obtain the representative features that have not been eliminated before this iteration, denoted as surviving representative features; Determine the sorting of multiple surviving representative features according to the feature elimination algorithm; Iteratively eliminate all representative features according to the sorting of the multiple surviving representative features and the second preset rule to obtain the elimination order of each representative feature.
8. The feature selection method for breath analysis according to claim 7, wherein The second preset rule is specifically: When the surviving number of the surviving representative features is greater than or equal to a second preset threshold, the number of representative features eliminated in this iteration is a preset ratio of the surviving number; When the surviving quantity is less than a second preset threshold, a preset quantity of representative features is eliminated in this iteration.
9. A feature selection system for breath analysis, characterized in that, Including: An acquisition module, configured to acquire an exhaled breath sample feature set and determine the correlation between every two exhaled breath features in the exhaled breath sample feature set; A grouping module, configured to divide all exhaled breath features into multiple exhaled breath feature groups according to the correlation, and determine the representative feature of each exhaled breath feature group according to a first preset rule; An inheritance module, configured to determine the membership degree of each exhaled breath feature to each representative feature, and determine the inheritance feature corresponding to each representative feature according to the membership degree; A selection module, configured to sort the inheritance features according to a feature elimination algorithm and a second preset rule, and use the inheritance features and their sorting as a feature selection result.