A method and system for preprocessing stroke screening data

Through the oversampling technology of MAHAKIL and isolated forests based on the maximum information coefficient and principal component analysis method, the problem of feature selection and data imbalance in stroke screening data was solved, and the efficiency and accuracy of data analysis were improved.

CN115114995BActive Publication Date: 2025-07-29BEIJING ZHENTAI LIFE TECHNOLOGY CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202210834869.0
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-07-15
Publication Date
2025-07-29
Estimated Expiration
2042-07-15

AI Technical Summary

Technical Problem

When processing stroke screening data, the prior art has the problem of limitations in feature selection algorithms and the single synthesis of samples in traditional oversampling methods, which affects the efficiency of data analysis.

Method used

The correlation feature selection method based on the maximum information coefficient and the principal component analysis method are used to reduce the feature dimensions. Combined with the oversampling technology of MAHAKIL and isolated forests, redundant noise is removed and a few types of samples tend to be equalized.

Benefits of technology

The analysis efficiency and accuracy of stroke screening data are improved, and the data non-balance problem is handled through feature selection and oversampling technology, which enhances the balance and analysis effect of the data set.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115114995B_ABST
    Figure CN115114995B_ABST
Patent Text Reader

Abstract

The present invention relates to a method and system for preprocessing stroke screening data. The present invention uses a feature dimensionality reduction method to remove redundant noise from stroke screening data, performs feature selection on it, and performs feature transformation on the selected features; in view of the data imbalance characteristic of stroke screening data, oversampling is performed to make the samples of the minority class and the majority class tend to be balanced, thereby improving the analysis efficiency of stroke screening data.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of data processing, and particularly to a method and a system for preprocessing stroke screening data. Background Art

[0002] With the accelerating aging process in China, stroke has the characteristics of high disability rate and high fatality rate. Stroke screening data presents the characteristics of high-dimensional and imbalanced data sets, which greatly affects the analysis efficiency of this kind of data. When processing stroke screening data, in terms of feature dimensionality reduction, the existing feature selection algorithms have the advantage of not changing the distribution of the original features, but only selecting the original features, which has certain limitations; the defect of feature extraction is that it changes the distribution characteristics of the original feature data. In terms of imbalanced data, traditional oversampling methods such as the SMOTE method have defects such as single synthetic samples and inability to synthesize minority class samples rich in more information. Summary of the Invention

[0003] To solve the above problems existing in the prior art, the present invention provides a method and a system for preprocessing stroke screening data.

[0004] To achieve the above object, the present invention provides the following solutions:

[0005] A method for preprocessing stroke screening data includes:

[0006] Obtain an original data set; the original data set includes stroke screening data;

[0007] Perform feature selection on the original data set by using a correlation characteristic feature selection method based on the maximum information coefficient to obtain a feature subset;

[0008] Perform feature extraction on the feature subset by using the principal component analysis method to obtain a feature transformation subset;

[0009] Process the feature transformation subset by using an oversampling technique to obtain a balanced data set; use the balanced data set as the preprocessed stroke screening data set.

[0010] Preferably, the performing feature selection on the original data set by using a correlation characteristic feature selection method based on the maximum information coefficient to obtain a feature subset specifically includes:

[0011] Extract the features of the original data set to obtain a feature set;

[0012] Determine the evaluation value of each feature in the feature set, and sort the features in the feature set based on the evaluation value to obtain a feature sequence;

[0013] Construct an empty set;

[0014] Extract the feature with the largest evaluation value in the said feature sequence to obtain a feature subsequence;

[0015] Put the extracted feature with the largest evaluation value into the said empty set to obtain an intermediate feature subset;

[0016] Select the current feature with the largest evaluation value in the said feature subsequence;

[0017] Judge whether the evaluation value of the said current feature is less than the evaluation value of the feature in the said intermediate feature subset to obtain a judgment result;

[0018] When the judgment result is less, delete the said current feature; when the judgment result is greater than or equal to, put the said current feature into the said intermediate feature subset to obtain the said feature subset;

[0019] Return to the step of selecting the current feature with the largest evaluation value in the said feature subsequence until all the features in the said feature subsequence are traversed.

[0020] Preferably, process the said feature transformation subset by using the MRIF oversampling technique based on MAHAKIL and isolation forest to obtain a balanced data set.

[0021] Preferably, the process of using the MRIF oversampling technique based on MAHAKIL and isolation forest to process the said feature transformation subset to obtain a balanced data set specifically includes:

[0022] Extract the class attribute of each feature in the said feature transformation subset;

[0023] Divide the features in the said feature transformation subset into majority class samples and minority class samples according to the said class attribute; use the Mahalanobis random oversampling MARO method to synthesize new samples for the said minority class samples;

[0024] Use the isolation forest algorithm to remove the noise factors in the said new samples to obtain new minority class samples;

[0025] Recombine the said majority class samples, the said minority class samples and the said new minority class samples to obtain the said balanced data set.

[0026] According to the specific embodiments provided by the present invention, the following technical effects are disclosed by the present invention:

[0027] The stroke screening data preprocessing method provided by the present invention uses a feature dimensionality reduction method to remove redundant noise in stroke screening data, performs feature selection on it, and performs feature transformation on the selected features; aiming at the data imbalance characteristic of stroke screening data, oversampling is performed to make the minority class and majority class samples tend to be balanced, thereby improving the analysis efficiency of stroke screening data.

[0028] Corresponding to the above-provided preprocessing method for stroke screening data, the present invention also provides a preprocessing system for stroke screening data, which system includes:

[0029] A dataset acquisition module, configured to acquire an original dataset; the original dataset includes stroke screening data;

[0030] A first feature selection module, configured to perform feature selection on the original dataset by using a correlation feature selection method based on the maximum information coefficient to obtain a feature subset;

[0031] A second feature selection module, configured to perform feature extraction on the feature subset by using the principal component analysis method to obtain a feature transformation subset;

[0032] An oversampling processing module, configured to process the feature transformation subset by using an oversampling technique to obtain a balanced dataset; and use the balanced dataset as the preprocessed stroke screening dataset.

[0033] Preferably, the first feature selection module includes:

[0034] A first feature extraction unit, configured to extract features of the original dataset to obtain a feature set;

[0035] A feature sequence construction unit, configured to determine an evaluation value of each feature in the feature set, and sort the features in the feature set based on the evaluation value to obtain a feature sequence;

[0036] An empty set construction unit, configured to construct an empty set;

[0037] A second feature extraction unit, configured to extract the feature with the largest evaluation value in the feature sequence to obtain a feature subsequence;

[0038] An intermediate feature subset construction unit, configured to put the extracted feature with the largest evaluation value into the empty set to obtain an intermediate feature subset;

[0039] A feature selection unit, configured to select the current feature with the largest evaluation value in the feature subsequence;

[0040] A judgment unit, configured to judge whether the evaluation value of the current feature is less than the evaluation value of the feature in the intermediate feature subset to obtain a judgment result; when the judgment result is less than, delete the current feature; when the judgment result is greater than or equal to, put the current feature into the intermediate feature subset to obtain the feature subset;

[0041] A return traversal unit, configured to return to the feature selection unit until all features in the feature subsequence are traversed.

[0042] Preferably, the oversampling processing module includes:

[0043] An oversampling processing unit for processing the feature transformation subset by using the MRIF oversampling technique based on MAHAKIL and isolation forest to obtain a balanced data set.

[0044] Preferably, the oversampling processing unit includes:

[0045] A feature extraction subunit for extracting the class attribute of each feature in the feature transformation subset;

[0046] A sample classification subunit for classifying the features in the feature transformation subset into majority class samples and minority class samples according to the class attribute;

[0047] A sample synthesis subunit for synthesizing new samples for the minority class samples by using the Mahalanobis random oversampling MARO method;

[0048] A denoising subunit for removing noise factors in the new samples by using the isolation forest algorithm to obtain new minority class samples;

[0049] A data recombination subunit for recombining the majority class samples, the minority class samples and the new minority class samples to obtain the balanced data set.

[0050] Since the technical effects achieved by the stroke screening data preprocessing system provided by the present invention are the same as those achieved by the stroke screening data preprocessing method provided above, they will not be elaborated here. Description of the Drawings

[0051] In order to more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the following will briefly introduce the drawings required for use in the embodiments. Obviously, the drawings in the following description are only some embodiments of the present invention. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.

[0052] Figure 1 It is a flowchart of the stroke screening data preprocessing method provided by the present invention;

[0053] Figure 2 It is a flowchart of the hybrid feature dimensionality reduction method provided by the embodiments of the present invention;

[0054] Figure 3 It is a flowchart of the MARO oversampling method provided by the embodiments of the present invention;

[0055] Figure 4 It is an overall process framework diagram provided by the embodiments of the present invention;

[0056] Figure 5 Schematic structural diagram of the stroke screening data preprocessing system provided by the present invention. Specific embodiments

[0057] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of the present invention.

[0058] The purpose of the present invention is to provide a stroke screening data preprocessing method and system, which can improve the efficiency and accuracy of stroke screening data preprocessing, and further improve the efficiency of stroke screening data analysis.

[0059] To make the above objects, features, and advantages of the present invention more obvious and understandable, the present invention will be further described in detail below in conjunction with the accompanying drawings and specific embodiments.

[0060] As Figure 1 shown, the stroke screening data preprocessing method provided by the present invention includes:

[0061] Step 100: Obtain the original data set. The original data set includes stroke screening data.

[0062] Step 101: Use the correlation feature selection method based on the maximum information coefficient to perform feature selection on the original data set to obtain a feature subset.

[0063] The implementation process of this step is as follows:

[0064] Step 1010: Extract the features of the original data set to obtain a feature set.

[0065] Step 1011: Determine the evaluation value of each feature in the feature set, and sort the features in the feature set based on the evaluation value to obtain a feature sequence.

[0066] Step 1012: Construct an empty set.

[0067] Step 1013: Extract the feature with the largest evaluation value in the feature sequence to obtain a feature subsequence.

[0068] Step 1014: Put the extracted feature with the largest evaluation value into the empty set to obtain an intermediate feature subset.

[0069] Step 1015: Select the current feature with the largest evaluation value in the feature subsequence.

[0070] Step 1016: Determine whether the evaluation value of the current feature is less than the evaluation value of the features in the intermediate feature subset to obtain a judgment result. When the judgment result is less, delete the current feature. When the judgment result is greater than or equal, put the current feature into the intermediate feature subset to obtain a feature subset.

[0071] Return to execute Step 1015 until all the features in the feature subsequence are traversed.

[0072] Step 102: Use the principal component analysis method (Principal Component Analysis, PCA) to extract features from the feature subset to obtain a feature transformation subset.

[0073] Step 103: Use the oversampling technique to process the feature transformation subset to obtain a balanced dataset. Use the balanced dataset as the preprocessed stroke screening dataset. Among them, the oversampling technique of the present invention uses the oversampling MAHAKIL (oversampling based on the theory of inheritance and the Mahalanobis distance) and the Mahalanobis random isolation forest MRIF (MAHAKIL Random and Isolation Forest) oversampling technique based on genetic theory and Mahalanobis distance. The specific process of processing the feature transformation subset by this oversampling technique includes:

[0074] Step 1030: Extract the class attributes of each feature in the feature transformation subset.

[0075] Classify the features in the feature transformation subset into majority class samples and minority class samples according to the class attributes. For example, when processing stroke data with two class attribute information, the data corresponding to the class attribute with fewer samples is the minority class sample, and the data corresponding to the class attribute with more samples is the majority class sample.

[0076] Step 1031: Use the Mahalanobis random oversampling MARO (MAHAKIL Random Oversampling) method to synthesize new samples for the minority class samples.

[0077] Step 1032: Use the isolation forest algorithm to remove the noise factors in the new samples to obtain new minority class samples.

[0078] Step 1033: Recombine the majority class samples, minority class samples and new minority class samples to obtain a balanced dataset.

[0079] The following is based on as Figure 4The implementation framework shown below illustrates the specific implementation process of the above-provided preprocessing method for stroke screening data. The data parameters used in this embodiment are only for illustrative purposes and are not specific limitations of the present invention.

[0080] Based on the descriptions in the above Step 101 and Step 102, the design concept of the hybrid feature dimensionality reduction method that combines feature selection and feature extraction in the present invention is as follows:

[0081] First, the present invention uses the correlation criterion for feature selection. Existing methods use information gain as the correlation criterion, which has a tendency to select feature subsets with more feature attributes. The present invention introduces the Maximal Information Coefficient (MIC) as the selection criterion, abbreviated as the correlation-based feature selection method based on the maximal information coefficient (MCFS).

[0082] Secondly, for the subset after feature selection, the present invention selects the PCA feature extraction algorithm to further reduce the dimensionality of the filtered feature subset, thereby obtaining the optimal feature transformation subset F2. The flowchart of the hybrid feature dimensionality reduction method is as Figure 2 shown, and its implementation process is as follows:

[0083] 1) Use the MCFS method to perform feature selection on the original stroke screening data set F, calculate the feature value according to the evaluation function, and use the best-first search method to select the streamlined feature subset, denoted as F1.

[0084] 2) Select the PCA algorithm to further extract features from the obtained feature subset F1 to obtain the optimal feature transformation subset, denoted as F2.

[0085] Among them, the MCFS method uses the maximal information coefficient to correct the correlation between features. The advantage of this innovative idea is that it not only considers the redundancy between features but also the correlation between features and categories.

[0086] The MCFS algorithm evaluates the classification ability of the feature subset and uses the best-first search method to select the feature subset with the highest value. That is, after calculating the obtained evaluation value Merits F and sorting it in descending order, if the selected feature subset includes M features, then take the first M evaluation values Merits F The corresponding obtained feature combination is the feature subset. Merits F The calculation formula of the evaluation function is shown in the following formula (1):

[0087]

[0088] In the formula, k represents the number of features, F is the selected feature subset, is the average correlation coefficient between the feature subset features and the class, is the average correlation coefficient between the features in the feature subset F.

[0089] The best-first search method is used to obtain the feature subset F1: Specifically, in this embodiment, the search starts from an empty set, calculates the evaluation values of all single features and sorts them. First, the feature with the largest evaluation value is put into the empty set, and then the currently largest feature is selected in turn, and the evaluation value of the feature subset at this time is judged. If the evaluation value decreases, the feature is removed, otherwise the feature is retained, and so on, to find the feature subset that makes the evaluation value the largest, which is the feature subset F1.

[0090] In this embodiment, the correlation coefficient between the features and the class uses the Symmetrical Uncertainty (SU) to solve, in order to compensate for the symmetrical uncertainty of the information gain deviation, where the solution of the symmetrical uncertainty coefficient SU is shown in Equation (2):

[0091]

[0092] In the formula, S and R respectively represent the feature and the class subset, H(S) and H(R) are the information entropies of S and R respectively, and are solved using Equation (3). gain(S,R) is the information gain and is solved using Equation (5).

[0093] The larger the symmetrical uncertainty SU, the higher the correlation between S and R, indicating that the feature has a more significant impact on the class.

[0094]

[0095]

[0096] gain(S,R) = H(S) - H(S|R) (5)

[0097] In the formula, s is the feature in the feature subset S, r is the class feature in the class subset R; p(s) represents the probability mass function of the feature s, p(r) represents the probability mass function of the class feature r; H(S|R) is the conditional entropy between S and R; p(s|r) is the probability function of the feature s under the condition of the r class;

[0098] The existing correlation-based feature selection methods usually use information gain to solve the correlation between features, and its defect is that it is more inclined to select features with more value numbers. In this embodiment, the correlation between features is measured by MIC, and its calculation formula is shown in Equation (6):

[0099]

[0100] In the above formula, I(S1, S2) is the mutual information between features S1 and S2, which is calculated using the joint probability density. The MIC index discretizes the relationship between features S1 and S2 into the two-dimensional scatter plot space, where |x| and |y| are the number of divisions of the grid in the two-dimensional scatter plot space, N is the number of samples, and B(N) is a function of the samples. In this embodiment, B(N) = N 0.6 . The smaller the MIC of two features, the lower the correlation. When selecting features in the present invention, the denominator value should be as large as possible, that is, the correlation between features should be as small as possible.

[0101] The optimal feature transformation subset F2 obtained based on the above process has the characteristics of an imbalanced data set. For this data set, the oversampling technique is used to process the class imbalance problem between classes in the data set. In this embodiment, the MRIF oversampling technique based on MAHAKIL and the isolation forest is proposed. The principle of the MRIF oversampling technique is as follows:

[0102] First, the optimal feature transformation subset F2 is divided into majority class samples and minority class samples according to different labels. In this embodiment, only the stroke data with two category attribute information is processed. Among them, the data corresponding to the category attribute with a small number of samples is the minority class sample, and the data corresponding to the category attribute with a large number of samples is the majority class sample. The Mahalanobis random oversampling MARO method is used to synthesize new samples for the minority class samples. After the synthesized new samples are further processed by the isolation forest algorithm to remove noise factors, the majority class samples, minority class samples, and the new minority class samples after denoising are recombined.

[0103] The defect of the MAHAKIL algorithm is that when synthesizing offspring samples, the average value of the parent samples is taken, ignoring the problem of sample diversity. The Mahalanobis random oversampling MARO oversampling technique of the present invention uses random numbers to synthesize offspring samples, which can make the generated samples have more characteristics and retain the characteristics of the parent, and further uses the isolation forest to remove the noise samples in the newly generated minority class samples. The description of the Mahalanobis random oversampling MARO oversampling technique is shown in Table 1:

[0104] Table 1 Description table of the Mahalanobis random oversampling MARO oversampling technique

[0105]

[0106]

[0107] The newly synthesized samples by oversampling often generate noise, which interferes with the classification performance. To avoid the influence of noise and improve the efficiency of synthesized samples, this embodiment adopts the Isolation Forest (ISF) anomaly detection technology, and the detected outliers are the noise factors.

[0108] The identified noise only comes from the synthesized minority class samples generated during the oversampling process, rather than the original samples. Eliminate the samples identified as noise, and then reorganize the majority class and minority class samples. As Figure 3 shown in the flowchart of the MRIF oversampling technique.

[0109] It can be seen that in this embodiment, first, some minority class samples are synthesized through the MARO algorithm, then the noise samples are removed using the Isolation Forest algorithm, and finally, the synthesized new samples and the original samples are passed through the SMOTE algorithm together to balance the dataset.

[0110] Based on the above description, compared with the prior art, the present invention also has the following advantages:

[0111] 1. Aiming at the defect of redundant features in the stroke screening dataset, the present invention proposes a hybrid feature dimensionality reduction method FS-FE. First, improve the Correlation-based Feature Selection (CFS) algorithm through the Maximal Information Coefficient (MIC) to make up for the defect that CFS tends to select more feature attributes with more values during feature selection. Then, further streamline the selected feature subset using the PCA feature extraction algorithm to obtain the optimal feature subset.

[0112] 2. Due to the imbalance of stroke screening data, the present invention proposes an oversampling technique based on the MAHAKIL oversampling technique and the isolation forest algorithm (MAHAKIL Random and Isolation Forest, MRIF). First, in order to improve the diversity of newly generated samples, the method of simply taking the average value to generate new samples in the MAHAKIL oversampling technique is replaced with random numbers. In addition, considering that noise samples are easily generated when synthesizing new samples, the present invention uses the isolation forest algorithm to detect and remove the noise samples from the newly synthesized samples. The MRIF oversampling technique has certain limitations for highly imbalanced samples. Considering that the SMOTE algorithm is prone to exacerbating the intra-class imbalance of the dataset when synthesizing new samples, but more minority class information will be carried when synthesizing new samples from similar minority classes, a two-stage oversampling technique that combines the MRIF algorithm and the SMOTE algorithm is also proposed. First, a part of minority class samples are synthesized from the original imbalanced dataset through the MRIF algorithm, and the obtained new samples are combined with the original data to form a new imbalanced dataset. Finally, the SMOTE algorithm is used to obtain the final balanced dataset.

[0113] In addition, corresponding to the above-provided stroke screening data preprocessing method, the present invention also provides a stroke screening data preprocessing system, as Figure 5 shown, the system includes:

[0114] A dataset acquisition module 1 for acquiring an original dataset. The original dataset includes stroke screening data.

[0115] A first feature selection module 2 for performing feature selection on the original dataset by using a correlation feature selection method based on the maximum information coefficient to obtain a feature subset.

[0116] A second feature selection module 3 for performing feature extraction on the feature subset by using the principal component analysis method to obtain a feature transformation subset.

[0117] An oversampling processing module 4 for processing the feature transformation subset by using an oversampling technique to obtain a balanced dataset. The balanced dataset is used as the preprocessed stroke screening dataset.

[0118] Preferably, the first feature selection module 2 includes:

[0119] A first feature extraction unit for extracting the features of the original dataset to obtain a feature set.

[0120] A feature sequence construction unit for determining the evaluation value of each feature in the feature set and sorting the features in the feature set based on the evaluation value to obtain a feature sequence.

[0121] An empty set construction unit for constructing an empty set.

[0122] The second feature extraction unit is used to extract the feature with the largest evaluation value in the feature sequence to obtain a feature subsequence.

[0123] The intermediate feature subset construction unit is used to put the extracted feature with the largest evaluation value into an empty set to obtain an intermediate feature subset.

[0124] The feature selection unit is used to select the current feature with the largest evaluation value in the feature subsequence.

[0125] The judgment unit is used to judge whether the evaluation value of the current feature is less than the evaluation value of the feature in the intermediate feature subset to obtain a judgment result. When the judgment result is less than, the current feature is deleted. When the judgment result is greater than or equal to, the current feature is put into the intermediate feature subset to obtain a feature subset.

[0126] The return traversal unit is used to return to the feature selection unit until all the features in the feature subsequence are traversed.

[0127] Preferably, the oversampling processing module 4 includes:

[0128] The oversampling processing unit is used to process the feature transformation subset by using the MRIF oversampling technology based on MAHAKIL and the isolated forest to obtain a balanced data set.

[0129] Among them, the oversampling processing unit includes:

[0130] The feature extraction subunit is used to extract the class attribute of each feature in the feature transformation subset.

[0131] The sample classification subunit is used to classify the features in the feature transformation subset into majority class samples and minority class samples according to the class attribute.

[0132] The sample synthesis subunit is used to synthesize new samples for the minority class samples by using the MARO method.

[0133] The denoising subunit is used to remove the noise factors in the new samples by using the isolated forest algorithm ISF to obtain new minority class samples.

[0134] The data recombination subunit is used to recombine the majority class samples, minority class samples and new minority class samples to obtain a balanced data set.

[0135] In this specification, each embodiment is described in a progressive manner. The key point of each embodiment is to illustrate the differences from other embodiments. The same or similar parts among the embodiments can be referred to each other. For the system disclosed in the embodiment, since it corresponds to the method disclosed in the embodiment, the description is relatively simple, and the relevant parts can be referred to the description of the method part.

[0136] In this article, specific examples are used to illustrate the principles and implementation manners of the present invention. The description of the above embodiments is only used to help understand the method and its core idea of the present invention; at the same time, for those of ordinary skill in the art, according to the idea of the present invention, there will be changes in the specific implementation manners and application scopes. In summary, the content of this specification should not be construed as a limitation to the present invention.

Claims

1. A method for preprocessing stroke screening data, characterized in that, Including: Obtain the original dataset; the original dataset includes stroke screening data; Use the correlation feature selection method based on the maximum information coefficient to perform feature selection on the original dataset to obtain a feature subset; Use the principal component analysis method to perform feature extraction on the feature subset to obtain a feature transformation subset; Use the oversampling technique to process the feature transformation subset to obtain a balanced dataset; use the balanced dataset as the preprocessed stroke screening dataset; Use the correlation feature selection method based on the maximum information coefficient to perform feature selection on the original dataset, calculate the feature value according to the evaluation function, and use the best-first search method to select the feature subset F1; select the PCA algorithm to further perform feature extraction on the obtained feature subset F1 to obtain the optimal feature transformation subset F2; Use the MRIF oversampling technique based on MAHAKIL and the isolation forest to process the feature transformation subset to obtain a balanced dataset, including: 1) Divide the dataset D into a majority-class subset D maj ={d maj1 , d maj2 , …, d Nmaj ,} and a minority-class subset D min , where N maj represents the number of majority-class samples and N min represents the number of minority-class samples. Calculate the number of samples N new to be generated as N maj = N min ; D new : The new sample subset to be generated, initialized to 0; N newnum : Used to count the number of generated samples, initialized to 0; 2) Select the central sample to obtain D min That is, the Mahalanobis distance between each sample in the minority class subset and the central sample, and arrange the minority class samples in descending order according to the obtained Mahalanobis distance and save them to D minMA ; 3) Find D min Find the center point of which is N mid Let N min be the lower approximation of N / 2: Divide the samples in D according to the Mahalanobis distance from the central position, and divide the samples into two parts D minMA and D par1 and D par2 , and assign the same corresponding labels to the two parts of data to achieve pairing; The pairing is done in sequence, that is, the samples at the same position in the two parts of the samples are given the same label and marked as a pair; sequentially for D par1 and D par2 All samples are paired in sequence; the pairing method is as follows: Among them, D par1 and D par2 The data in are all arranged in descending order; Let d par1i ∈ D par1 , d par2i ∈ D par2 , for D par1 and D par2 the samples in the same sorted position are assigned the same label l i , i = 1, 2, ... N mid ; 4) If i = 1, …, N mid ; From D par1 , D par2 Separate samples d a , d b , and d a , d b The two sample labels are the same, which are paired samples: l i (d a ) = l i (d b ); Generate new samples: d newi = d a × random + d b × (1 - random); random is a random number between 0 and 1; Add d newi to D new ; N newnum +1; 5) If N newnum < N new , then recombine the samples in D new with the samples in D par and update D min with them, and repeat steps 2) to 6) to generate new minority class samples; 6) Until D newnum >= N new , D new All the samples in are the generated minority class samples; 7) Use the isolation forest ISF algorithm to remove the noise samples in the newly generated samples to obtain a new unbalanced dataset D'; 8) Use the SMOTE oversampling technique to generate new samples, and merge the new samples with D′ to obtain the final balanced dataset D final .

2. A data preprocessing system for stroke screening, characterized in that, The stroke screening data preprocessing system is used to implement the stroke screening data preprocessing method as described in claim 1; The system includes: A dataset acquisition module for obtaining the original dataset; the original dataset includes stroke screening data; A first feature selection module for using the correlation feature selection method based on the maximum information coefficient to perform feature selection on the original dataset to obtain a feature subset; A second feature selection module for using the principal component analysis method to perform feature extraction on the feature subset to obtain a feature transformation subset; An oversampling processing module for using the oversampling technique to process the feature transformation subset to obtain a balanced dataset; use the balanced dataset as the preprocessed stroke screening dataset.