Feature selection method for identifying stroke markers of patients with asymptomatic cerebral infarction

By employing a two-stage neighborhood rough set feature selection method that combines static and dynamic features, the problem of not considering time information in the identification of stroke markers in asymptomatic cerebral infarction patients is solved, thereby improving the identification accuracy and achieving more comprehensive feature selection.

CN121117554APending Publication Date: 2025-12-12UNIV OF ELECTRONICS SCI & TECH OF CHINA +1
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202511200855.3
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-08-26
Publication Date
2025-12-12

AI Technical Summary

Technical Problem

Existing technologies fail to effectively combine time-informed follow-up data when identifying stroke markers in asymptomatic stroke patients, neglecting the dynamic characteristics of image features changing over time. This leads to the omission of some important features and affects the accuracy of identification.

Method used

A two-stage neighborhood rough set feature selection method is adopted. The importance and weight of feature subsets are calculated using the overall data sample set and the follow-up data sample set respectively. Combining static and dynamic features, a multi-granularity and time-weighted neighborhood rough set method is used for feature screening, and time information is introduced to adjust feature importance.

Benefits of technology

It improves the accuracy of stroke biomarker identification in asymptomatic cerebral infarction patients, solves the problem of not considering the risk of dynamic feature temporal superposition in existing technologies, and obtains feature selection results with overall sample interpretability.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121117554A_ABST
    Figure CN121117554A_ABST
Patent Text Reader

Abstract

The invention belongs to the technical field of SBI patient stroke biomarker recognition and software, and particularly relates to a feature selection method for recognizing stroke markers of an asymptomatic cerebral infarction patient. The method is a feature selection method of a two-stage neighborhood rough set which introduces change information of multiple follow-up visit data of the SBI patient, and the identification accuracy of the stroke biomarker of the SBI patient can be improved. According to the method, the neighborhood rough set is used, irregular longitudinal data change feature importance is introduced into feature selection of the SBI patient stroke marker, and the problem that the dynamic feature time superposition risk is ignored due to the fact that follow-up visit data with time information is not combined in existing SBI patient stroke biomarker recognition is solved. And meanwhile, by combining the importance of a static feature subset and the weight of a dynamic feature subset, the features are screened, and a feature selection result which has overall sample interpretability and considers the influence of a dynamic feature time superposition risk is obtained.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention belongs to the field of stroke biomarker identification and software technology for SBI patients, specifically a feature selection method for identifying stroke biomarkers in asymptomatic cerebral infarction patients. Background Technology

[0002] Stroke biomarkers in patients with asymptomatic brain infarction (SBI) are a key concern for both patients and healthcare professionals. Asymptomatic brain infarction refers to a stroke lesion detected on imaging, but the patient does not exhibit related neurological deficits or symptoms. Stroke is the leading cause of disability among single diseases. Studies show that SBI patients are at high risk of stroke, with a five times higher probability of developing symptomatic stroke compared to those without SBI. Furthermore, SBI patients constitute a significant proportion of the elderly population; studies have found that over 20% of elderly individuals have SBI, with 30%–40% of SBI patients being over 70 years old. Therefore, identifying stroke biomarkers in a large SBI population is clinically significant, providing pathological analysis evidence for early prevention and treatment of stroke events in SBI patients.

[0003] A common approach to identifying stroke biomarkers in SBI patients is feature selection from a variety of existing characteristics. Common biomarkers typically include genetic factors, physiological indicators, and lifestyle habits. Currently, clustering algorithms are generally used for feature selection. It is worth noting that in some diseases, such as cancer, survival analysis models are used for feature selection. These models are characterized by incorporating regular temporal features as labels.

[0004] Chinese patent "CN114091607B A Semi-Supervised Multi-Label Online Flow Feature Selection Algorithm Based on Neighborhood Rough Sets" proposes a method to predict new features through defined neighborhood relationships, perform online feature importance evaluation, and update and select the optimal feature subset. This patent maximizes the dependency of candidate feature sets on labels and selects features with high relevance and low redundancy.

[0005] Chinese patent CN118553432A, "A Method and Application for Establishing a Predictive Model for Survival Time after Radiotherapy for Unresectable Bile Duct Cancer," proposes a predictive model for survival time after radiotherapy for unresectable bile duct cancer, incorporating time features into a feature selection method. This patent combines a risk function with time features, uses LASSO regression for feature selection, screens candidate predictive indicators, and finally substitutes the selected independent risk factors into a Cox model to obtain a survival time model based on these factors.

[0006] In the mining of biomarkers for stroke in the SBI population, the imaging features of SBI and their impact over time are often overlooked. Since infarcts are irreversible, the type, distribution, number, and size of infarcts in imaging features are all important indicators. Furthermore, existing machine learning methods for biomarker analysis of dynamic data from multiple follow-up studies require the completeness of the longitudinal time series data. However, existing SBI follow-up data is irregular, making it impossible to directly incorporate time-varying covariates for analysis. This leads to the neglect of the importance of changes in some physiological indicators and imaging features in SBI.

[0007] The method described in Chinese patent "CN114091607B A Semi-Supervised Multi-Label Online Flow Feature Selection Algorithm Based on Neighborhood Rough Set" can be used to process dynamic flow features, but it does not take into account the temporal characteristics of the data.

[0008] The method described in Chinese patent "CN118553432A: A Method for Establishing and Applying a Predictive Model for Survival Time after Radiotherapy for Unresectable Bile Duct Cancer" uses data with regular time characteristics for feature selection, which is not suitable for irregular SBI follow-up data. Furthermore, this method does not reflect the dynamic nature of the data; for example, it uses changes in specific examination information over time as a feature, treating time merely as a static feature. Summary of the Invention

[0009] To address the aforementioned problems, this invention proposes a feature selection method for identifying stroke biomarkers in asymptomatic stroke patients. This invention employs a two-stage neighborhood rough set feature selection method that incorporates change information from multiple follow-up data of SBI patients, which can improve the accuracy of stroke biomarker identification in SBI patients.

[0010] The technical solution of this invention is:

[0011] A feature selection method for identifying stroke biomarkers in asymptomatic cerebral infarction patients includes the following steps:

[0012] S1. Obtain physical examination data of asymptomatic stroke patients, and construct an overall data sample set and a follow-up data sample set after preprocessing; the method for constructing the overall data sample set is as follows: record the physical examination data of patients obtained through the physical examination system as an electronic health record set. ,in This represents the record of the i-th patient. N, where N is the total number of electronic health records; defined This includes k medical examination data points, all of which have been converted to digital format. The corresponding follow-up information set is represented as Follow-up information includes time information, lifestyle information, biochemical test data, imaging data, and rating scale information. This refers to the follow-up information in the j-th physical examination data, where t represents time information. For lifestyle information, For biochemical test data, For image data, For rating scale information, Disassembled into ,and , r represents demographic information, and h represents medical history information. This represents the follow-up information in the j-th physical examination data after deleting the time information, where d represents the stroke outcome information; it will be... Disassembled As a sample data point, through analysis of... The overall data sample set is obtained by decomposition. The method for constructing the follow-up data sample set is as follows: extract the follow-up information from E, and take the i-th electronic health record... The extracted follow-up information was subtracted pairwise in chronological order to obtain... New sample record , , k, will be obtained As a sample data point, through analysis of... The follow-up information in each electronic health record is processed to obtain a follow-up data sample set. ;

[0013] S2. Obtain the feature subset importance index and feature subset weight from the overall data sample set and the follow-up data sample set, respectively;

[0014] The specific method for obtaining the feature subset importance index from the overall data sample set is as follows:

[0015] Based on the six categories of attribute information corresponding to the overall data sample set—namely, demographic information, medical history information, lifestyle information, biochemical test data, imaging data, and rating scale information—the overall data sample set is divided into six groups. Each group contains the corresponding data and labels for each attribute. as well as Each set of data is used to construct a quadruple ( , , , =( , ∪ , , , = ∪ , Indicates the number of groups and L = 6 Indicates the first Group of overall data sample sets, Representing attribute characteristics, by and Composition, in which Indicates a conditional attribute. This represents the decision attribute, corresponding to stroke outcome information d. It refers to the attribute value range, which includes the conditional attribute value range and the decision attribute value range for all samples. It is a mapping. It reflects the relationship between objects, attributes, and attribute values, that is... Therefore, there is ;

[0016] For any sample Equivalence class based on decision attribute d Composed of samples with the same class label, d ,Will The equivalence class is defined as:

[0017]

[0018] Given a feature subset P, P⊆ , The neighborhood set on P contains and Similar samples; given neighborhood distance The neighborhood set is defined as:

[0019]

[0020] in This indicates that the sample in the feature subset P and Euclidean distance;

[0021] In distance The importance of the neighborhood dependency of a feature subset P based on the neighborhood dependency method is expressed as follows:

[0022]

[0023] Where | represents the number of samples in the set, The value range is from 0 to 1;

[0024] Introducing neighborhood credibility and neighborhood coverage In the feature subset P, for the equivalence class based on decision attribute d Its neighborhood credibility and neighborhood coverage They are respectively:

[0025]

[0026]

[0027] Using neighborhood credibility and neighborhood coverage The neighborhood dependency importance of the feature subset P is adjusted, denoted as... :

[0028]

[0029] This refers to the importance index of feature subsets;

[0030] The specific method for obtaining feature subset weights from the follow-up data sample set is as follows:

[0031] Follow-up data sample set Treat it as an attribute set and construct a quadruple (U, R, V, f). =(U, T∪C∪D, V, ,in, =(T∪C∪D , For conditional attributes, For decision attributes, For time attributes, V 2 It is the attribute value range of all samples. It is a mapping relationship, and the decision attribute only considers the stroke outcome d;

[0032] Using the neighborhood radius of the attribute Divide into equivalence classes, for , ,in For attributes Standard deviation of the data It is a set parameter;

[0033] For any feature subset P⊆ Equivalence classes are defined by dynamic features with time attributes. / ={ }, through decision attribute d The division of equivalence classes is denoted as / ={ }, then the information entropy of P and the joint information entropy of P and d are expressed as follows:

[0034]

[0035]

[0036] in i = 1, 2, ..., m, representing The probability, The cardinality of equivalence classes. Let i = 1, 2, ..., m, j = 1, 2, ..., n, representing The probability of;

[0037] Due to the existence - The conditional information entropy of d relative to P is defined as:

[0038]

[0039] in , i=1,2,...,m, j=1,2,...,n;

[0040] The mutual information between P and d is defined as:

[0041]

[0042] right The feature subset P is weighted under the decision attribute d. Defined as:

[0043]

[0044] in For normalization processing;

[0045] S3. After obtaining the feature subset importance index and feature subset weights, input them into the feature selection module. Use dynamic feature weights to adjust the feature subset importance index, obtain the conditional feature weighted importance of feature subsets in each group of features, and perform feature selection, specifically:

[0046] For the feature subset P⊆ Its importance is recorded as ,because The larger the value, the more divergent P is; therefore, for a given conditional feature... Conditional feature weighted importance The calculation method is as follows:

[0047]

[0048]

[0049] in This is a weighted importance stacking function for feature subsets. The feature subset weights of feature subset P under decision attribute d are calculated from the follow-up data sample set.

[0050] The feature selection algorithm is a two-stage neighborhood rough set feature selection, specifically including:

[0051] (1) Input sample set ( V f , ( V f , Middle Neighborhood Radius , , Middle Neighborhood Radius Calculation Parameters ;

[0052] (2) Initialize feature subset ;

[0053] (3) To middle Each quadruple is used for feature subset selection. For ( , , , , , = ∪ Calculate each conditional feature Weighted importance, selection Features ,make { }, ;

[0054] (4) Repeat step (3) until or Obtain the feature subset of the quadruple. , where the threshold Obtained from experiments;

[0055] (5) Perform a union of the obtained L feature subsets. This yields the final selected feature subset P.

[0056] The beneficial effects of this invention are as follows:

[0057] Compared with existing technologies, this invention uses neighborhood rough sets to introduce the importance of irregular longitudinal data variation features into feature selection for stroke biomarker identification in SBI patients. This solves the problem that current SBI stroke biomarker identification does not incorporate follow-up data with time information, thus ignoring the risk of dynamic feature temporal superposition. Simultaneously, by combining the importance of static feature subsets and the weights of dynamic feature subsets, features are screened to obtain feature selection results that have overall sample interpretability and consider the impact of dynamic feature temporal superposition risk. Attached Figure Description

[0058] Figure 1 This is the overall flowchart of the present invention.

[0059] Figure 2 This is a schematic diagram of the two-stage neighborhood rough set feature selection algorithm in this invention.

[0060] Figure 3 This is a schematic diagram of the overall architecture of the present invention.

[0061] Figure 4 This is a schematic diagram of the multi-granularity weighted neighborhood rough set in this invention.

[0062] Figure 5 This is a schematic diagram of the single-granularity time-weighted neighborhood rough set in this invention. Detailed Implementation

[0063] The technical principles and solutions of the present invention will now be described in detail with reference to the accompanying drawings:

[0064] The feature selection method of the two-stage neighborhood rough set (DNRS) that incorporates change information from multiple follow-up data of SBI patients described in this invention is as follows: Figure 1 As shown, the specific steps are as follows:

[0065] Step 1: Data Preprocessing

[0066] A two-stage neighborhood rough set was constructed, comprising a global data sample set and a follow-up data sample set. The global data sample set includes all physical examination data of SBI patients, while the follow-up data sample set contains only physical examination data of SBI patients with follow-up data. The objectives are as follows: to predict stroke outcome for any data point in the global data set using markers identified through feature selection methods; and to investigate the impact of changes in sample features on the feature selection results using the follow-up data sample set.

[0067] (1) In the hospital's physical examination system, all physical examination data for each patient is recorded as an electronic health record, and the electronic health record set is recorded as follows: ,in Let N represent the record of the i-th patient, and N be the total number of electronic health records.

[0068] (2) Given that feature selection methods can be used for biomarker identification in SBI, and stroke outcome can be predicted using any single physical examination data point from asymptomatic stroke patients, it is necessary to further refine the feature selection method. The patient-centric data is split so that each patient's physical examination data is treated as a separate sample.

[0069] from The electronic health records were broken down item by item to obtain the overall data sample set. Based on domain knowledge, the features of the overall data sample set were classified into seven categories of attribute information: demographic information, medical history information, lifestyle information, time information, biochemical test data, imaging data, and rating scale information, as well as a stroke outcome label. Since all physical examination data for the same patient share a single set of demographic information, medical history information, and label information, and each physical examination data record contains a follow-up record, which includes time information, lifestyle information, biochemical test data, imaging data, and rating scale information, the specific breakdown method is as follows:

[0070] Record an electronic health record There are k medical examination records, therefore... Disassembled into , Where r represents demographic information and h represents medical history information, express The follow-up information set of k physical examination data. This refers to the follow-up information in the j-th physical examination data, where t represents time information. For lifestyle information, For biochemical test data, For image data, For rating scale information, represents the follow-up information in the j-th physical examination data after deleting the time information t, and d represents the stroke outcome information. Each data point is broken down into its constituent parts. As a single sample data point, it constitutes a complete data sample set containing six types of attribute information and one outcome information, denoted as... .

[0071] (3) Since the features in the follow-up information are all dynamic features, in order to explore the impact of changes in sample features, this paper extracts the follow-up information and outcome information in E, treats all follow-up information of the same patient as the same type of attribute information, and subtracts them from each other in order of time to obtain the feature changes in different time periods. This approach makes the data changes in each group of different time periods associated with the predicted stroke outcome.

[0072] Extract follow-up information from E to construct a follow-up data sample set. That is, for any electronic health record... If follow-up information More than one, that is At that time, generate New sample record = , k. Extracted As a single sample data point, a follow-up data sample set is constructed, denoted as... .

[0073] Step 2: Obtain the feature subset importance index and feature subset weight from the overall data sample set and the follow-up data sample set respectively:

[0074] The importance of classification features is measured based on the overall data sample set:

[0075] Due to the large number of influencing features, in order to reduce computational load and improve efficiency, this method, during the initial screening of overall influencing factors, is based on the six categories of attribute information (excluding time information) identified in step one, and then applies this information to the overall data. Based on attribute information, the data and corresponding labels under each attribute category are used as initial conditions to construct four-tuples. This involves using a computer to process the input... Automatically divided into as well as Six groups. Because the properties and number of features differ among groups, the number of selected features may vary. Therefore, a multi-granularity weighted neighborhood rough set is used, with different neighborhood radii... The neighborhood set is defined above. Meanwhile, considering the dominant role of consistent samples, this invention proposes a weighted feature importance measurement method, which introduces neighborhood confidence and neighborhood coverage when calculating the importance of feature subsets in each feature group.

[0076] (1) Input the overall data sample set Each attribute information category is used to construct a quadruple using the corresponding data and its corresponding label. , , , =( , ∪ , , , = ∪ , Indicates the number of groups and In this paper, L = 6. Representing attribute characteristics, by and Composition, in which Indicates a conditional attribute. Represents decision attributes (i.e., class labels in classification). It refers to the attribute value range, which includes the conditional attribute value range and the decision attribute value range for all samples. It is a mapping. It reflects the relationship between objects, attributes, and attribute values, that is... Therefore, there is Here, the decision attributes only consider stroke outcomes.

[0077] (2) For any sample Equivalence class based on decision attribute d It consists of samples with the same class label, where d .

[0078] The equivalence class can be defined as

[0079] (1)

[0080] (3) Given a feature subset P, P⊆ , The neighborhood set on P contains and Similar samples. Given neighborhood distance This paper defines the neighborhood set as follows:

[0081] (2)

[0082] in This indicates that the sample in the feature subset P and Euclidean distance.

[0083] (4) In summary, in terms of distance The importance of the neighborhood dependency of a feature subset P based on the neighborhood dependency method can be expressed as:

[0084] (3)

[0085] Where | represents the number of samples in the set, The value range is from 0 to 1.

[0086] (5) Considering the dominant role of consistent samples, neighborhood confidence is introduced when calculating the importance of feature subset P. and neighborhood coverage The approximate accuracy of the neighborhood of the equivalence class is evaluated, as well as the proportion of samples of the same class in the neighborhood, to emphasize the discriminative power of the feature subset for consistent samples. In the feature subset P, for the equivalence class based on decision attribute d... Its neighborhood credibility and neighborhood coverage They are respectively

[0087] (4)

[0088] (5)

[0089] (6) Use neighborhood credibility and neighborhood coverage The neighborhood dependency importance of the feature subset P is adjusted, denoted as... :

[0090] (6)

[0091] Therefore, in the first stage, we calculated the importance of different feature subsets in each set of data features using the overall data set.

[0092] Weighting of dynamic features based on follow-up data sample set:

[0093] In longitudinal data on asymptomatic ischemic stroke, some features change with increasing follow-up frequency. These changing features, such as infarct size in imaging features, are associated with the stroke incidence rate in asymptomatic ischemic stroke patients. Furthermore, the rate of feature change is closely related to stroke outcome. Therefore, when selecting features, it is necessary to consider the impact of these changing factors and assign them weights. Since follow-up of SBI patients is not periodic, this approach uses a time-weighted single-granularity neighborhood rough set feature selection method.

[0094] (1) The follow-up data sample set Treat it as an attribute set and construct a quadruple (U, R, V, f). =(U, T∪C∪D, V, .in, =(T∪C∪D , For conditional attributes, For decision attributes, The time attribute is considered, where the influence of time on the characteristics of change is taken into account. Extracted separately. ={t}。 V 2 It is the attribute value range of all samples. This is a mapping relationship. Here, the decision attribute only considers stroke outcome.

[0095] When screening dynamic features in longitudinal data analysis based on follow-up data sample sets, it is not necessary to group the data features. Existing studies on feature significance and correlation only consider the data itself.

[0096] (2) Using the neighborhood radius of the attribute Divide into equivalence classes, for , ,in For attributes Standard deviation of the data It is a set parameter used to adjust the neighborhood size based on the classification accuracy.

[0097] (3) For any feature subset P⊆ Equivalence classes are defined by dynamic features with time attributes. / ={ }, through decision attribute d The division of equivalence classes is denoted as / ={ }, then the information entropy of P and the joint information entropy of P and d are expressed as:

[0098] (7)

[0099] (8)

[0100] in i = 1, 2, ..., m, representing / The probability, The cardinality of the equivalence class. Let i = 1, 2, ..., m, j = 1, 2, ..., n, representing / The probability of.

[0101] (4) Due to the existence - Therefore, the conditional information entropy of d relative to P can be defined as:

[0102] (9)

[0103] in , i=1,2,...,m, j=1,2,...,n.

[0104] The mutual information between P and d is defined as:

[0105] (10)

[0106] (5) In the overall data sample set, this paper classifies the data. Data features belonging to the same category Only exist and Two scenarios. Therefore, for The feature subset weights of feature subset P under decision attribute d It can be defined as:

[0107] (11)

[0108] in This is for normalization purposes.

[0109] Step 3: Feature selection based on conditional feature weighted importance metric:

[0110] After obtaining the dynamic feature weights calculated based on the follow-up data sample set and the importance of feature subsets in each group of features based on the overall data sample set, the data is input into the feature selection module. In this module, the dynamic feature weights are used to adjust the importance of feature subsets, obtain the conditional feature weighted importance of feature subsets in each group of features, and then perform feature selection.

[0111] For the feature subset P⊆ Its importance is recorded as ,because The larger the value, the more divergent P is, and the more... Only exist and Two scenarios. Therefore, for a given conditional feature... Conditional feature weighted importance The calculation method is as follows:

[0112]

[0113]

[0114] in This is a weighted importance stacking function for feature subsets. The feature subset weights of feature subset P under decision attribute d are calculated from the follow-up data sample set.

[0115] The feature selection algorithm is a two-stage neighborhood rough set feature selection, such as... Figure 2 As shown. The algorithm flow is as follows:

[0116] (1) Input sample set ( V f , ( V f , Middle Neighborhood Radius , , Middle Neighborhood Radius Calculation Parameters ;

[0117] (2) Initialize feature subset ;

[0118] (3) To middle Each quadruple is subjected to feature subset selection. , , , , , = ∪ Through formula Calculate each conditional feature Weighted importance, selection Features ,make { }, ;

[0119] (4) Repeat step (3) until or Obtain the feature subset of the quadruple. , where the threshold Based on experiments, the range is generally set between [10-100].

[0120] (5) Perform a union of the L feature subsets obtained in the above steps. This yields the final selected feature subset P.

[0121] The key point of this invention is:

[0122] Most common feature selection methods are static, focusing only on the static characteristics of the data and neglecting time series or dynamic changes. This invention, however, considers feature changes over time by introducing longitudinal data with temporal information, addressing the problem of not considering the temporal superposition risk of risk factors when selecting features for stroke biomarkers in SBI patients. Furthermore, this invention proposes a two-stage neighborhood rough set feature selection method, combining static and dynamic features to obtain more comprehensive data information.

[0123] In the first stage, since different types of features have different characteristics, this invention groups features in electronic health record information using domain knowledge and proposes a weighted multi-granularity neighborhood rough set method that can adaptively determine the neighborhood radius based on the local geometry of the data, helping to better capture the intrinsic structure and distribution characteristics of the data. Considering the dominant role of consistent samples and emphasizing the discriminative power of feature subsets for consistent samples, neighborhood confidence and neighborhood coverage are introduced when calculating the importance of feature subsets.

[0124] Because SBI follow-up data is irregular, survival analysis methods cannot be used to stack the risks of dynamic features. Therefore, in the second stage, a time-weighted single-granularity neighborhood rough set is proposed. A neighborhood rough set combined with time is constructed using the same neighborhood radius. Weights are calculated for dynamic data features in the follow-up data sample set, thereby introducing importance weights for changing features. The feature importance obtained in the first stage is further adjusted to optimize feature selection.

Claims

1. A feature selection method for identifying stroke biomarkers in asymptomatic cerebral infarction patients, characterized in that, Includes the following steps: S1. Obtain physical examination data of asymptomatic stroke patients, and construct an overall data sample set and a follow-up data sample set after preprocessing; the method for constructing the overall data sample set is as follows: record the physical examination data of patients obtained through the physical examination system as an electronic health record set. ,in This represents the record of the i-th patient. N, where N is the total number of electronic health records; defined This includes k medical examination data points, all of which have been converted to digital format. The corresponding follow-up information set is represented as Follow-up information includes time information, lifestyle information, biochemical test data, imaging data, and rating scale information. This refers to the follow-up information in the j-th physical examination data, where t represents time information. For lifestyle information, For biochemical test data, For image data, For rating scale information, Disassembled into ,and , r represents demographic information, and h represents medical history information. This represents the follow-up information in the j-th physical examination data after deleting the time information, where d represents the stroke outcome information; it will be... Disassembled As a sample data point, through analysis of... The overall data sample set is obtained by decomposition. The method for constructing the follow-up data sample set is as follows: extract the follow-up information from E, and take the i-th electronic health record... The extracted follow-up information was subtracted pairwise in chronological order to obtain... New sample record , , k, will be obtained As a sample data point, through analysis of... The follow-up information in each electronic health record is processed to obtain a follow-up data sample set. ; S2. Obtain the feature subset importance index and feature subset weight from the overall data sample set and the follow-up data sample set, respectively; The specific method for obtaining the feature subset importance index from the overall data sample set is as follows: Based on the six categories of attribute information corresponding to the overall data sample set—namely, demographic information, medical history information, lifestyle information, biochemical test data, imaging data, and rating scale information—the overall data sample set is divided into six groups. Each group contains the corresponding data and labels for each attribute. as well as Each set of data is used to construct a quadruple ( , , , =( , ∪ , , , = ∪ , Indicates the number of groups and L = 6 Indicates the first Group of overall data sample sets, Representing attribute characteristics, by and Composition, in which Indicates a conditional attribute. This represents the decision attribute, corresponding to stroke outcome information d. It refers to the attribute value range, which includes the conditional attribute value range and the decision attribute value range for all samples. It is a mapping. It reflects the relationship between objects, attributes, and attribute values, that is... Therefore, there is ; For any sample Equivalence class based on decision attribute d Composed of samples with the same class label, d ,Will The equivalence class is defined as: , Given a feature subset P, P⊆ , The neighborhood set on P contains and Similar samples; given neighborhood distance The neighborhood set is defined as: , in This indicates that the sample in the feature subset P and Euclidean distance; In distance The importance of the neighborhood dependency of a feature subset P based on the neighborhood dependency method is expressed as follows: , Where | represents the number of samples in the set, The value range is from 0 to 1; Introducing neighborhood credibility and neighborhood coverage In the feature subset P, for the equivalence class based on decision attribute d Its neighborhood credibility and neighborhood coverage They are respectively: , , Using neighborhood credibility and neighborhood coverage The neighborhood dependency importance of the feature subset P is adjusted, denoted as... : , This refers to the importance index of feature subsets; The specific method for obtaining feature subset weights from the follow-up data sample set is as follows: Follow-up data sample set Treat it as an attribute set and construct a quadruple (U, R, V, f). =(U, T∪C∪D, V, ,in, =(T∪C∪D , For conditional attributes, For decision attributes, For time attributes, V 2 It is the attribute value range of all samples. It is a mapping relationship, and the decision attribute only considers the stroke outcome d; Using the neighborhood radius of the attribute Divide into equivalence classes, for , ,in For attributes Standard deviation of the data It is a set parameter; For any feature subset P⊆ Equivalence classes are defined by dynamic features with time attributes. / ={ }, through decision attribute d The division of equivalence classes is denoted as / ={ }, then the information entropy of P and the joint information entropy of P and d are expressed as follows: , , in i = 1, 2, ..., m, representing The probability, The cardinality of equivalence classes. Let i = 1, 2, ..., m, j = 1, 2, ..., n, representing / The probability of; Due to the existence - The conditional information entropy of d relative to P is defined as: , in , i=1,2,...,m, j=1,2,...,n; The mutual information between P and d is defined as: , right The feature subset P is weighted under the decision attribute d. Defined as: , in For normalization processing; S3. After obtaining the feature subset importance index and feature subset weights, input them into the feature selection module. Use dynamic feature weights to adjust the feature subset importance index, obtain the conditional feature weighted importance of feature subsets in each group of features, and perform feature selection, specifically: For the feature subset P⊆ Its importance is recorded as ,because The larger the value, the more divergent P is; therefore, for a given conditional feature... Conditional feature weighted importance The calculation method is as follows: , , in This is a weighted importance stacking function for feature subsets. The feature subset weights of feature subset P under decision attribute d are calculated from the follow-up data sample set. The feature selection algorithm is a two-stage neighborhood rough set feature selection, specifically including: (1) Input sample set ( V f , ( V f , Middle Neighborhood Radius , , Middle Neighborhood Radius Calculation Parameters ; (2) Initialize feature subset ; (3) To middle Each quadruple is used for feature subset selection. For ( , , , , , = ∪ Calculate each conditional feature Weighted importance, selection Features ,make { }, ; (4) Repeat step (3) until or Obtain the feature subset of the quadruple. , where the threshold Obtained from experiments; (5) Perform a union of the obtained L feature subsets. This yields the final selected feature subset P.

Citation Information

Patent Citations

  • A semi-supervised multi-label online stream feature selection method based on neighborhood rough sets

    CN114091607B

  • Establishment method and application of prediction model for survival time of non-resectable bile duct cancer after radiotherapy

    CN118553432A