A Feature Selection Method for Early Screening of Coronary Artery Disease Based on Soft Path Cost Accumulation
By using a feature selection method based on soft path cost accumulation, early coronary heart disease features with high interaction are screened out, which solves the problem that existing technologies fail to fully utilize the interaction between multiple diagnostic indicators, and achieves more efficient and accurate early screening for coronary heart disease.
Patent Information
- Application Number
- CN202410792942.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-06-19
- Publication Date
- 2025-12-02
- Estimated Expiration
- 2044-06-19
AI Technical Summary
Existing early screening and prediction models for coronary heart disease fail to fully utilize the interactions between multiple diagnostic indicators, limiting the predictive power of the models and the effective use of medical resources.
A feature selection method based on soft path cost accumulation is adopted. By using the soft path cost accumulation function and the stopping score function, early features of coronary heart disease with high interaction are screened out. The feature selection process is optimized by utilizing the complex relationship between multiple diagnostic indicators.
It has improved the accuracy and efficiency of early screening and prediction models for coronary heart disease, reduced the waste of medical resources, and ensured timely diagnosis for high-risk patients.
Smart Images

Figure CN118888119B_ABST
Abstract
Description
Technical Field
[0001] This invention belongs to the field of medical data processing technology, and more specifically, relates to a method for selecting early screening features for coronary heart disease based on soft path cost accumulation. Background Technology
[0002] In its early stages, coronary artery disease (CAD) often presents with no significant symptoms, making it crucial to identify and implement prevention strategies for high-risk groups. Traditional medical assessment methods evaluate CAD risk by integrating various clinical indicators. However, because the human body is a highly complex biological system, many diagnostic indicators are interrelated, often leading to redundant diagnoses and unnecessary waste of medical resources. Feature selection techniques, by identifying key information, provide patients with more accurate and efficient diagnostic solutions. Feature selection algorithms used for CAD diagnosis can be broadly categorized into filter methods, wrapper methods, and embedded methods. Filter methods are widely used in large-scale medical datasets because they do not rely on the results of any classifier, and their statistical principles offer faster training speeds and lower computational costs.
[0003] Nevertheless, most existing filter-based early screening and prediction models for coronary artery disease (CAD) only consider the relationship between a single diagnostic indicator and CAD, failing to fully utilize the interactions between multiple diagnostic indicators, which limits the predictive power of the models. Although some methods consider the interactions between diagnostic indicators, most only consider the interaction between a pair of diagnostic indicators, without taking advantage of the complex relationships between multiple diagnostic indicators. Therefore, developing a new feature selection method that can deeply analyze and utilize these interactions is of great value for improving the accuracy of CAD prediction models and reducing medical costs. Summary of the Invention
[0004] To address the shortcomings and improvement needs of existing technologies, this invention provides a feature selection method for early screening of coronary heart disease based on soft path cost accumulation. Its purpose is to improve the accuracy and efficiency of feature selection for early screening of coronary heart disease by fully utilizing the interaction between multiple diagnostic indicators, thereby enhancing the accuracy and efficiency of early screening prediction models for coronary heart disease.
[0005] To achieve the above objectives, according to one aspect of the present invention, a method for selecting early coronary heart disease screening features based on soft path cost accumulation is provided, comprising: S1, selecting a candidate early coronary heart disease feature from the candidate feature set and adding it to the selected feature set with the goal of maximizing the soft path cost accumulation function, and removing the selected candidate early coronary heart disease feature from the candidate feature set; S2, calculating the difference between the information gain ratio of the current selected feature set to the entire dataset and the information gain ratio of the selected feature set before its addition to the entire dataset, and using it as the stopping scoring function; S3, repeating S1-S2 until the stopping scoring function obtained k times consecutively is less than 0, where k is a set hyperparameter threshold, or the candidate feature set is empty; S4, taking the current selected feature set corresponding to the maximum stopping scoring function as the optimal early coronary heart disease feature set.
[0006] Furthermore, the soft path cost accumulation function is:
[0007]
[0008] Where SPA is the soft path cost accumulation function, F i For the selected early features of coronary heart disease in the selected feature set, F j Here, C represents the candidate early features for coronary heart disease, C is the category label, and best represents the selected feature set. MIE(F) j C) is F j The mutual information balance between F and C, Rel(F) i ,F j C) is F i F j The interaction measure between C and C.
[0009] Furthermore, Rel(F i ,F j C) is:
[0010]
[0011] Among them, I(F) i ,F j C) is F i F j Three-way interactive mutual information with C, MIE(F) i ,F j ) is F i and F j The balance of mutual information between them.
[0012] Furthermore, MIE(F i ,F j ), MIE(F j C) are respectively:
[0013]
[0014]
[0015] Among them, I(F) i ,F j ) is F i and F j Mutual information between them, I(F) j C) is F j Mutual information between H and C, H(F) i ) is F i The entropy, H(F) j ) is F j The entropy of C is H(C).
[0016] Furthermore, the stopping scoring function is:
[0017] score = g R (D,best)-g R (D,best_before)
[0018]
[0019]
[0020] Where score is the stopping score function, g R (D,best) represents the information gain ratio of the current selected feature set to the entire dataset, and g R (D,best_before) represents the information gain ratio of the selected feature set before inclusion to the entire dataset, g(D,best) represents the information gain of the current selected feature set to the entire dataset, and H... best (D) represents the information entropy of the current selected feature set relative to the entire dataset, g(D, best_before) represents the information gain of the selected feature set relative to the entire dataset before its inclusion, and H represents the information gain of the selected feature set relative to the entire dataset before its inclusion. best_before (D) represents the information entropy of the selected feature set before inclusion relative to the entire dataset, where D represents the entire dataset, best represents the current selected feature set, and best_before represents the selected feature set before inclusion.
[0021] Furthermore, before S1, the method further includes: acquiring a coronary heart disease dataset, wherein the samples in the coronary heart disease dataset contain multiple early features of coronary heart disease to be screened and corresponding category labels, wherein the category labels indicate whether the patient has coronary heart disease; and sequentially performing data cleaning, removing missing values and data discretization on the coronary heart disease dataset to obtain an initial candidate feature set.
[0022] According to another aspect of the present invention, a feature selection system for early screening of coronary heart disease based on soft path cost accumulation is provided, comprising: a processor; and a memory storing a computer-executable program, wherein when the program is executed by the processor, the processor performs the feature selection method for early screening of coronary heart disease based on soft path cost accumulation as described above.
[0023] According to another aspect of the invention, a computer-readable storage medium is provided having a computer program stored thereon, which, when executed by a processor, implements the method for selecting early screening features for coronary heart disease based on soft path cost accumulation as described above.
[0024] In summary, the above-described technical solutions conceived in this invention can achieve the following beneficial effects:
[0025] (1) A method for early screening feature selection of coronary heart disease based on soft path cost accumulation is provided. The feature selection is based on maximizing the soft path cost accumulation function. This method can evaluate the interaction between multiple causes, which is more consistent with the complex relationship between causes of coronary heart disease and provides higher efficiency and stability for early screening feature selection of coronary heart disease.
[0026] The stopping feature selection method uses a stopping score function to monitor the impact of newly added features, ensuring that each iteration moves toward improving the overall model performance. This iterative process allows the model to adaptively find the optimal solution in the feature space, rather than relying on preset fixed rules. Furthermore, due to the dynamic stopping condition, this method can adjust the iterative process and feature selection strategy according to the characteristics of different datasets, exhibiting high adaptability and flexibility.
[0027] (2) By introducing the interaction metric Rel(F) i ,F j The C) and soft path cost accumulation function SPA achieve the reduction of redundancy while maintaining the correlation between features, effectively reducing unnecessary features in the early screening of coronary heart disease, while retaining features with high interaction value and effectively quantifying the interaction between features.
[0028] (3) The soft path feature selection (SPFS) algorithm selects the most representative set of coronary heart disease features by accurately balancing the correlation and redundancy of each feature and using soft path accumulation technology to comprehensively consider the interrelationship between features. This process not only optimizes the allocation of medical resources, but also ensures that high-risk patients can receive accurate diagnosis in a timely manner, significantly improving the value and efficiency of medical services. Attached Figure Description
[0029] Figure 1 A flowchart of a method for early screening feature selection of coronary heart disease based on soft path cost accumulation provided in an embodiment of the present invention;
[0030] Figure 2 This is a schematic diagram illustrating the operation process of the early screening feature selection method for coronary heart disease based on soft path cost accumulation provided in an embodiment of the present invention. Detailed Implementation
[0031] To make the objectives, technical solutions, and advantages of this invention clearer, the invention will be further described in detail below with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are merely illustrative and not intended to limit the invention. Furthermore, the technical features involved in the various embodiments of this invention described below can be combined with each other as long as they do not conflict with each other.
[0032] In this invention, the terms "first," "second," etc. (if present) in the invention and the accompanying drawings are used to distinguish similar objects and are not necessarily used to describe a specific order or sequence.
[0033] Example 1
[0034] A method for feature selection in early screening of coronary heart disease based on soft path cost accumulation, see [reference]. Figure 1 , combined Figure 2 The method for selecting early screening features for coronary heart disease based on soft path cost accumulation in this embodiment is described in detail. The method includes operations S1-S4.
[0035] In this embodiment, an initial candidate feature set needs to be prepared before performing operation S1. Specifically, this includes: obtaining a coronary heart disease dataset, in which samples contain multiple early coronary heart disease features to be screened and corresponding category labels, the category labels indicating whether the patient has coronary heart disease; and sequentially performing data cleaning, missing value removal, and data discretization on the coronary heart disease dataset to obtain the initial candidate feature set.
[0036] After data cleaning, missing value removal, and data discretization, a dataset D containing n samples, m features, and a class label C is obtained. All features F = {F1, F2, ..., F...} are then processed. m Let} be the initial candidate feature set.
[0037] It should be noted that before executing operation S1, the following operations should also be included: In the initial state, the selected feature set is empty, and the candidate feature set is the set of all features in the entire dataset D. At this time, it is necessary to calculate each candidate early coronary heart disease feature F in the candidate feature set. j Mutual information balance (MIE) between the category label C and the category label C j ,C), will maximize MIE(F) jC) The corresponding early coronary heart disease features to be selected are added to the selected feature set to make the selected feature set not empty, and then the subsequent operations S1-S4 are executed.
[0038] Operation S1 aims to maximize the soft path cost accumulation function, select a candidate early coronary heart disease feature from the candidate feature set and add it to the selected feature set, and remove the selected candidate early coronary heart disease feature from the candidate feature set.
[0039] According to an embodiment of the present invention, the soft path cost accumulation function is:
[0040]
[0041] Where SPA is the soft path cost accumulation function, F i For the selected early features of coronary heart disease in the selected feature set, F j Here, C represents the candidate early features for coronary heart disease, C is the category label, and best represents the selected feature set. MIE(F) j C) is F j The mutual information balance between F and C, Rel(F) i ,F j C) is F i F j The interaction measure between and C. MIE(F) j C) is used to evaluate the direct association between each feature and the category label.
[0042] By using the soft path cost accumulation (SPA) function described above, the algorithm not only considers the interaction effect between features, but also explicitly calculates the direct information contribution between the candidate features and the class label C, thus optimizing the decision criteria in the feature selection process.
[0043] Rel(F i ,F j C) is:
[0044]
[0045] Among them, I(F) i ,F j C) is F i F j Three-way interactive mutual information with C, MIE(F) i ,F j ) is F i and F j The balance of mutual information between them.
[0046] MIE(F i ,F j ), MIE(F j C) are respectively:
[0047]
[0048]
[0049] Among them, I(F) i ,F j ) is F i and F j Mutual information between them, I(F) j C) is F j Mutual information between H and C, H(F) i ) is F i The entropy, H(F) j ) is F j The entropy of C is H(C).
[0050] Interaction metric Rel(F) i ,F j C) By combining three-way interactive mutual information I(F) i ,F j The mutual information balance (MIE) allows the algorithm to comprehensively evaluate the information gain of feature combinations, thereby identifying highly redundant feature combinations that contribute little to the model’s predictive power.
[0051] In this embodiment, the objective is to maximize the soft path cost accumulation function. A candidate feature for early coronary heart disease is selected from the candidate feature set and added to the already selected feature set. This ensures the strongest interaction between the selected candidate feature and the already selected features, thereby improving the overall predictive ability of the feature set. This method, the soft path feature selection scoring criterion, can be expressed as:
[0052]
[0053] Where candidate is the current set of candidate features.
[0054] Operation S2 calculates the difference between the information gain ratio of the current selected feature set to the entire dataset and the information gain ratio of the selected feature set to the entire dataset before its inclusion, and uses this as the stopping scoring function.
[0055] In this embodiment, a stopping score function is introduced to evaluate whether the newly added features significantly improve the model's information gain ratio, thereby optimizing the feature selection process and preventing premature stopping of selection. The stopping score function is:
[0056] score = g R (D,best)-g R (D,best_before)
[0057]
[0058]
[0059] Where score is the stopping scoring function, g R (D,best) represents the information gain ratio of the current selected feature set to the entire dataset, and g R (D,best_before) represents the information gain ratio of the selected feature set before inclusion to the entire dataset, g(D,best) represents the information gain of the current selected feature set to the entire dataset, and H... best (D) represents the information entropy of the current selected feature set relative to the entire dataset, g(D, best_before) represents the information gain of the selected feature set relative to the entire dataset before its inclusion, and H represents the information gain of the selected feature set relative to the entire dataset before its inclusion. best_before (D) represents the information entropy of the selected feature set before inclusion relative to the entire dataset, where D represents the entire dataset, best represents the current selected feature set, and best_before represents the selected feature set before inclusion.
[0060] Operation S3 repeats operations S1-S2 until the stopping score function obtained after k consecutive operations is less than 0, where k is the set hyperparameter threshold, or the candidate feature set is empty.
[0061] In this embodiment, the value of newly added candidate early coronary heart disease features is continuously evaluated using the stopping score function, and new candidate early coronary heart disease features are selected and added to the selected feature set until the feature selection stopping criterion is met.
[0062] Specifically, each time a new candidate early coronary heart disease feature is added, the stopping score function (score) of the selected feature set is calculated. If the score < 0, it is considered that the newly added candidate early coronary heart disease feature does not actually help the model's prediction accuracy, but only increases the model's complexity. To avoid prematurely eliminating potentially effective features, it is still added to the selected feature set. If the score > 0, it is considered that the newly added candidate early coronary heart disease feature brings more useful information and helps the model make correct predictions, and should be added to the selected feature set. When using the score to evaluate the newly added candidate early coronary heart disease feature, if the score is less than 0 for several consecutive iterations, feature selection is stopped. Figure 2 As shown.
[0063] Operation S4 takes the current selected feature set corresponding to the maximum stopping score function as the optimal early feature set for coronary heart disease.
[0064] Preferably, the currently selected feature set and its k-1 previously selected feature sets are extracted. Among these k selected feature sets, the feature set with the largest score is selected as the optimal early feature set for coronary heart disease.
[0065] The feature selection method for early screening of coronary heart disease based on soft path cost accumulation provided in this invention not only focuses on the information gain of individual features and labels, but also analyzes the interaction between features in greater depth. This allows the model to identify feature combinations with higher information gain in complex data structures, providing higher efficiency and stability. In particular, it demonstrates better performance and reliability when dealing with large-scale imbalanced datasets.
[0066] Example 2
[0067] A feature selection system for early screening of coronary artery disease based on soft path cost accumulation includes: a processor; and a memory storing a computer-executable program. When the program is executed by the processor, it causes the processor to perform the aforementioned feature selection method for early screening of coronary artery disease based on soft path cost accumulation. The related technical solution is the same as in Embodiment 1 and will not be repeated here.
[0068] Example 3
[0069] A computer-readable storage medium stores a computer program thereon, which, when executed by a processor, implements the aforementioned method for selecting early screening features for coronary heart disease based on soft path cost accumulation. The related technical solution is the same as in Embodiment 1, and will not be repeated here.
[0070] Those skilled in the art will readily understand that the above description is merely a preferred embodiment of the present invention and is not intended to limit the present invention. Any modifications, equivalent substitutions, and improvements made within the spirit and principles of the present invention should be included within the scope of protection of the present invention.
Claims
1. A method for selecting early screening features for coronary heart disease based on soft path cost accumulation, characterized in that, include: S1, with the goal of maximizing the soft path cost accumulation function, select a candidate early coronary heart disease feature from the candidate feature set and add it to the selected feature set, and remove the selected candidate early coronary heart disease feature from the candidate feature set; S2, calculate the difference between the information gain ratio of the current selected feature set to the whole dataset and the information gain ratio of the selected feature set to the whole dataset before its inclusion, and use it as the stopping scoring function; S3, repeat S1-S2 until the stopping score function obtained k times consecutively is less than 0, where k is the set hyperparameter threshold, or the candidate feature set is empty; S4, take the current selected feature set corresponding to the maximum stopping score function as the best early feature set for coronary heart disease; The soft path cost accumulation function is: in, This is the soft path cost accumulation function. These are the early features of coronary heart disease selected from the already selected feature set. Early characteristics of coronary heart disease to be selected For category labels, For the selected feature set, for and Mutual information balance between them for , and A measure of the interaction between them; for: in, for , and Three-way interactive information, for and Mutual information balance between them; , They are respectively: in, for and Mutual information between them for and Mutual information between them for entropy, for entropy, for The entropy.
2. The method for selecting early screening features for coronary heart disease based on soft path cost accumulation as described in claim 1, characterized in that, The stopping scoring function is: in, For the stopping scoring function, The information gain ratio of the current selected feature set to the entire dataset. The information gain ratio of the selected feature set before inclusion to the entire dataset. The information gain of the current selected feature set over the entire dataset. The information entropy of the current selected feature set relative to the entire dataset. The information gain of the selected feature set before inclusion over the entire dataset. The information entropy of the selected feature set before inclusion relative to the entire dataset. For the entire dataset, For the current selected feature set, This is the set of features selected before inclusion.
3. The method for selecting early screening features for coronary heart disease based on soft path cost accumulation as described in any one of claims 1-2, characterized in that, Before S1, the following also applies: Obtain a coronary heart disease dataset, wherein the samples in the coronary heart disease dataset contain multiple early features of coronary heart disease to be screened and corresponding category labels, wherein the category labels indicate whether the patient has coronary heart disease; The coronary heart disease dataset was sequentially cleaned, missing values were removed, and the data was discretized to obtain an initial candidate feature set.
4. A feature selection system for early screening of coronary heart disease based on soft path cost accumulation, characterized in that, include: processor; A memory storing a computer-executable program, which, when executed by the processor, causes the processor to perform the feature selection method for early screening of coronary heart disease based on soft path cost accumulation as described in any one of claims 1-3.
5. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the program is executed by the processor, it implements the early screening feature selection method for coronary heart disease based on soft path cost accumulation as described in any one of claims 1-3.
Citation Information
Patent Citations
Feature selection method and device for model training and electronic equipment
CN111104572A