Cerebral stroke diagnosis system based on improved SMOTE algorithm and ensemble learning

By improving the SMOTE algorithm and integrated learning method, a few samples that are more in line with the actual distribution are generated, and combined with the prediction results of multiple basic learners, the data imbalance problem in early diagnosis of stroke is solved, the robustness and prediction accuracy of the model are improved, and reliable clinical diagnosis auxiliary tools are provided.

CN120280119APending Publication Date: 2025-07-08ZHONGBEI UNIV
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510153054.X
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-02-12
Publication Date
2025-07-08

AI Technical Summary

Technical Problem

When dealing with early diagnosis of stroke, the existing technology has data imbalance problems, resulting in low model learning efficiency and poor prediction accuracy. The samples generated by the traditional SMOTE algorithm are inconsistent with the actual distribution, which affects the generalization ability of the model.

Method used

The improved SMOTE algorithm is used combined with the integrated learning method to improve the robustness and generalization ability of the model by generating a few class samples that are more in line with the actual distribution, and using the Stacking algorithm combined with the prediction results of multiple basis learners.

Benefits of technology

By generating a few samples closer to the actual distribution and integrated learning methods, the prediction accuracy of early stroke diagnosis is improved, the risk of misdiagnosis and misdiagnosis is reduced, and reliable clinical diagnosis assistance tools are provided.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120280119A_ABST
    Figure CN120280119A_ABST
Patent Text Reader

Abstract

The invention discloses a cerebral apoplexy diagnosis system based on an improved SMOTE algorithm and integrated learning. The cerebral apoplexy diagnosis system comprises a data preparation and preprocessing module, an unbalanced data set processing module, a data set generation module, a data set division module, a learner training module and a prediction module. By generating minority class samples closer to actual distribution, the problem that samples generated by a traditional SMOTE algorithm are inconsistent with actual distribution is solved. Meanwhile, the robustness and generalization ability of the model are improved by combining an integrated learning Stacking method with prediction results of a plurality of base learners, so that a better prediction effect is obtained in early diagnosis of cerebral apoplexy.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical fields of artificial intelligence and integrated learning, and particularly relates to a stroke diagnosis system based on an improved SMOTE algorithm and integrated learning. Background Art

[0002] Stroke is a common acute cerebrovascular disease with a high incidence rate, disability rate, and mortality rate. Early diagnosis and prevention are crucial. However, in the early stage of stroke, patients may not have obvious symptoms, resulting in scarce relevant case data, especially a significant shortage in the number of minority class samples (such as stroke patients). This data imbalance problem seriously affects the learning efficiency and prediction accuracy of the model.

[0003] Research on the method for processing imbalanced datasets based on the improved SMOTE algorithm studied the application of the SMOTE algorithm in medical data classification, especially the augmentation of minority class samples in rare disease classification. The article proposed an improved SMOTE algorithm. By adjusting the interpolation method of the generated samples, the diversity of minority class samples was increased, and the problem of model overfitting caused by overly concentrated data points was avoided. The study also compared different sampling methods, indicating that the improved method has significantly improved diagnostic accuracy and robustness in medical diagnosis tasks.

[0004] Research on prediction models for imbalanced medical data proposed an oversampling method applicable to minority class medical data. By generating minority class samples based on information in the adaptive neighborhood, the diversity and reliability of the generated samples were improved. This method dynamically adjusted the positions of the generated samples, making the distribution of the new samples more in line with the actual data and enhancing the generalization ability of the classification model in diagnostic tasks such as rare diseases. The article verified through experiments that this method has better classification performance compared to traditional oversampling techniques and is suitable for imbalanced medical datasets.

[0005] Although traditional oversampling methods such as SMOTE can alleviate the data imbalance problem to a certain extent, since they generate new samples using linear interpolation, this may result in the generated samples being inconsistent with the actual distribution, affecting the generalization ability of the model. In addition, a single model may have limitations when dealing with complex data.

[0006] Therefore, there are defects in the prior art and improvements are needed. Summary of the Invention

[0007] The present invention proposes a stroke diagnosis system based on an improved SMOTE algorithm and integrated learning in view of the deficiencies of the prior art, aiming to generate minority class samples that are more in line with the actual distribution and improve the diagnostic performance and generalization ability of the model.

[0008] The technical solution of the present invention is as follows:

[0009] A stroke diagnosis system based on an improved SMOTE algorithm and ensemble learning, comprising: a data preparation and preprocessing module, an imbalanced dataset processing module, a dataset generation module, a dataset division module, a learner training module, and a prediction module;

[0010] Data preparation and preprocessing module: used to collect case data, divide the samples into normal samples and diseased samples, and respectively form a majority class dataset and a minority class dataset; then, preprocess the dataset, including numerically converting categorical features, filling in missing values, and standardizing all features to unify the scale of the data in each dimension;

[0011] Imbalanced dataset processing module: use the improved SMOTE algorithm to process the imbalanced dataset to obtain a balanced dataset;

[0012] The dataset generation module is used to synthesize a new dataset: that is, merge the data processed by the improved SMOTE algorithm with the original majority class samples to form a new and relatively balanced dataset;

[0013] The dataset division module divides the processed balanced dataset, including the feature matrix X and the label vector Y, into a training set and a test set;

[0014] Learner training module: use the Stacking ensemble learning algorithm to train the base learner and the meta-learner of the Stacking algorithm for final prediction;

[0015] Prediction module: generate the metadata of the test set and make predictions.

[0016] For the stroke diagnosis system described above, the processing flow of the imbalanced dataset processing module is as follows:

[0017] Step 2.1 Select a reference point: randomly select a point as the reference point key from all minority class samples. Assume the minority class sample set is D minority ={x1, x2,..., x n}, then the reference point key can be obtained through a random selection operation: key = x key where x key is a sample point randomly selected from D minority ;

[0018] Step 2.2 Select neighbor points: use the K-nearest neighbor algorithm in the minority class samples to find the nearest neighbor set of the reference point. After excluding the reference point itself, randomly select two points from the remaining neighbors, and name them ref_point and extra_point respectively;

[0019] Step 2.3 Calculate the direction vectors: Calculate the direction vectors from the reference point to these two neighbor points respectively, that is, the vector from the reference point key to ref_point and the vector from the reference point key to extra_point; the formula is as follows:

[0020] V ref = X ref - X key ,V extra = X extra - X key

[0021] Step 2.4 Introduce the weights weights in multiple dimensions: Adjust each feature dimension. The weights are calculated based on the variances of the features of the minority class samples. The larger the weight, the more important the feature is. After calculating the weights, adjust the direction vectors in combination with the weights; the weight calculation formula:

[0022]

[0023] where d is the number of feature dimensions, and Var(X minority,i ) is the variance of the i-th feature. The adjusted direction vectors are V' ref and V' extra ;

[0024] Step 2.5 Calculate the vector lengths: Calculate the lengths of the two direction vectors, and select the smaller value as the "radius" of the sector, called min_radius, to control the range of generating new points; the formula is as follows:

[0025] len_ref = ||V' ref ||, len_ref = ||V' extra ||

[0026] min_radius = min(len_ref, len_extra)

[0027] Step 2.6 Calculate the included angle: Use the dot product formula to calculate the included angle between the two direction vectors; this included angle will be used to determine the angular range of the sector area; the formula is as follows:

[0028]

[0029] where · represents the dot product, and then obtain the included angle through the inverse cosine function:

[0030] angle = arccos(cos(θ))

[0031] Step 2.7 Calculate the arc length: Calculate the arc length arc_length of a sector area based on the minimum radius min_radius and the included angle angle. The formula is as follows:

[0032] arc_length = min_radius × angle

[0033] Step 2.8 Randomly generate the position of the new point: Randomly select a position random_position on the sector arc length arc_length and calculate the angular offset θ. The formula is as follows:

[0034]

[0035] θ = α · angle

[0036] Step 2.9 Generate a new vector through linear interpolation: Use the angle θ to construct a new direction vector new_vector, which is the result of the weighted combination of vector_ref and vector_extra by sine and cosine. Finally, generate a new sample point new_point by adding the new vector new_vector to the reference point key:

[0037]

[0038] X new = X key + V new .

[0039] For the stroke diagnosis system described above, the dataset partitioning module: To generate metadata, the training set needs to be further partitioned into a training part and a validation part. Using the K-fold cross-validation method, the training set is divided into K subsets. Each time, K - 1 subsets are used to train the model, and the remaining one subset is used for validation.

[0040] For the stroke diagnosis system described above, the learner training module performs the following steps:

[0041] Step 5.1 Select the base learner: Use support vector machine, random forest, and logistic regression as the base learners;

[0042] Step 5.2 Train the base learner:

[0043] Training process of support vector machine:

[0044]

[0045] s.t.y i (w T x i + b) ≥ 1 - ζ i , ζi ≥0, i = 1, 2, ..., N

[0046] Prediction process:

[0047]

[0048] Random forest training process: Randomly select samples and features to construct multiple decision trees; each tree uses the Bootstrap sampling method to draw samples from the training set; each node randomly selects a part of the features for splitting;

[0049] Prediction process:

[0050]

[0051] where f(x) is the prediction result of the j-th tree, and T is the number of trees;

[0052] Logistic Regression: Training process:

[0053] where is the sample x i belonging to the positive class probability.

[0054] Prediction process:

[0055] where is the sigmoid function.

[0056] Step 5.3 Generate metadata: Combine the prediction results of all base learners on the validation set to form a new feature matrix, i.e., metadata; the feature of each sample in the metadata set is the prediction result of each base learner;

[0057] Step 5.4 Generate metadata on the validation set:

[0058] where and are the prediction results of SVM, random forest, and logistic regression on the validation set respectively;

[0059] Step 5.5 Generate metadata on the test set:

[0060] Z test = [M SVM (X test ), M RF (X test ), M LR (X test )]

[0061] Step 5.6 Training the meta-learner: The meta-learner selects the Random Forest algorithm, and its training process and prediction process are the same as those in Step 5.2; Use the generated meta-data to train the meta-learner. The input of the meta-learner is the prediction result of the base learner, and the output is the final prediction result;

[0062] For the described stroke diagnosis system, the prediction module uses the trained meta-learner to predict the meta-data of the test set and generate the final classification result;

[0063] Among them, M meat is the trained meta-learner, and Z test is the meta-data of the test set.

[0064] Adopting the above scheme, the present invention solves the problem that the samples generated by the traditional SMOTE algorithm are inconsistent with the actual distribution by generating minority class samples closer to the actual distribution. At the same time, by combining the prediction results of multiple base learners through the ensemble learning Stacking method, the robustness and generalization ability of the model are improved, so as to achieve better prediction results in the early diagnosis of stroke. Brief Description of the Drawings

[0065] Figure 1 : Overall flowchart;

[0066] Figure 2 : Flowchart of benchmark point and neighbor selection operation;

[0067] Figure 3 : Sample distribution diagram;

[0068] Figure 4 : Flowchart of direction vector combined with weight adjustment;

[0069] Figure 5 : Flowchart of limiting the generation range of new samples;

[0070] Figure 6 : Schematic diagram of samples generated by improved SMOTE algorithm;

[0071] Figure 7 : Flowchart of Stacking algorithm training;

[0072] Figure 8 : Specific flowchart of example calculation; Detailed Description of the Invention

[0073] The present invention will be described in detail below in conjunction with specific embodiments.

[0074] The present invention provides a stroke diagnosis system based on an improved SMOTE algorithm and ensemble learning, and the overall process is as Figure 1As shown below. The specific steps are as follows:

[0075] Step 1 Data Preparation and Preprocessing: First, collect case data and divide the samples into normal samples (majority class) and diseased samples (minority class), forming a majority class dataset and a minority class dataset respectively; then, preprocess the datasets, including numerical conversion of categorical features, filling missing values, and standardizing all features to unify the scale of the data in each dimension.

[0076] Step 2 Process the imbalanced dataset using the improved SMOTE algorithm to obtain a balanced dataset.

[0077] Step 2.1 Select a reference point: Randomly select a point as the reference point (key) from all minority class samples. Assume the minority class sample set is D minority ={x1, x2,..., x n}, then the reference point key can be obtained through a random selection operation: key = x key where x key is a sample point randomly selected from D minority ;

[0078] Step 2.2 Select neighbor points: Use the K-Nearest Neighbors algorithm (KNN) in the minority class samples to find the nearest neighbor set of the reference point. After excluding the reference point (key) itself, randomly select two points from the remaining neighbors, named ref_point and extra_point respectively; the selection process of the reference point and neighbors is as Figure 2 shown, and the sample distribution is as Figure 3 shown;

[0079] Step 2.3 Calculate the direction vectors: Calculate the direction vectors from the reference point to these two neighbor points respectively, that is, the vector from the reference point key to ref_point and the vector from the reference point key to extra_point; the formula is as follows:

[0080] V ref = X ref - X key , V extra = X extra - X key

[0081] Step 2.4 Introduce weights weights in multiple dimensions: Adjust each feature dimension. The weights are calculated based on the variances of the features of the minority class samples. The larger the weight, the more important the feature represents. After calculating the weights, adjust the direction vectors in combination with the weights; the weight calculation formula:

[0082]

[0083] where d is the number of dimensions of the feature, and Var(X minority,i ) is the variance of the i-th feature. The adjusted direction vectors are V' ref and V' extra . The process of calculating the direction vector for weight adjustment is as shown in Figure 4 ;

[0084] Step 2.5 Calculate the vector lengths: Calculate the lengths of the two direction vectors, and select the smaller value as the "radius" of the sector (referred to as min_radius) to control the range for generating new points; the formula is as follows:

[0085] len_ref = ||V' ref ||, len_ref = ||V' extra ||

[0086] min_radius = min(len_ref, len_extra)

[0087] Step 2.6 Calculate the included angle: Use the dot product formula to calculate the included angle between the two direction vectors. This included angle will be used to determine the angular range of the sector area; the formula is as follows:

[0088]

[0089] where · represents the dot product, and then obtain the included angle through the inverse cosine function:

[0090] angle = arccos(cos(θ))

[0091] Step 2.7 Calculate the arc length: Calculate the arc length arc_length of a sector area based on the minimum radius min_radius and the included angle angle, and the formula is as follows:

[0092] arc_length = min_radius × angle

[0093] The main functions of Steps 2.5, 2.6, and 2.7 are: to limit the generation range of new samples, and the process is as shown in Figure 5 ;

[0094] Step 2.8 Randomly generate the position of the new point: Randomly select a position random_position on the sector arc length arc_length and calculate the angular offset θ; the formula is as follows:

[0095]

[0096] θ = α · angle

[0097] Step 2.9 Generate a new vector through linear interpolation: Use the angle θ to construct a new direction vector new_vector, which is the result of the weighted combination of vector_ref and vector_extra by sine and cosine; finally, generate a new sample point new_point by adding the new vector new_vector to the reference point key:

[0098]

[0099] X new = X key + V new

[0100] Schematic diagram of the improved algorithm for generating new samples, as shown in Figure 6 shown.

[0101] Step 3 Synthesize a new dataset: Combine the data processed by the improved SMOTE algorithm with the original majority-class samples to form a new and relatively balanced dataset.

[0102] Step 4 Divide the processed balanced dataset (including the feature matrix X and the label vector Y) into a training set and a test set, with the training set accounting for 70% and the test set accounting for 30%.

[0103] Step 4.1 Further division of the training set: To generate metadata, the training set needs to be further divided into a training part and a validation part. The K-fold cross-validation method can be used to divide the training set into K subsets. Each time, K - 1 subsets are used to train the model, and the remaining one subset is used for validation.

[0104] Step 5 Use the Stacking ensemble learning algorithm to train the base learners and the meta-learner of the Stacking algorithm for final prediction. The specific flowchart is as shown in Figure 7 shown.

[0105] Step 5.1 Select base learners: Adopt Support Vector Machine (SVM), Random Forest, and Logistic Regression as base learners.

[0106] Step 5.2 Train the base learners:

[0107] Training process of Support Vector Machine (SVM):

[0108]

[0109] s.t.y i (wT x i + b) ≥ 1 - ζ i , ζ i ≥ 0, i = 1, 2,..., N

[0110] Prediction process:

[0111]

[0112] Random Forest training process: Randomly select samples and features to construct multiple decision trees; each tree uses the Bootstrap sampling method to draw samples from the training set; each node randomly selects a part of the features for splitting.

[0113] Prediction process:

[0114]

[0115] where f(x) is the prediction result of the j-th tree, and T is the number of trees;

[0116] Logistic Regression: Training process:

[0117]

[0118] where, is the probability that the sample x i belongs to the positive class.

[0119] Prediction process:

[0120]

[0121] where, is the sigmoid function.

[0122] Step 5.3 Generate metadata: Combine the prediction results of all base learners on the validation set to form a new feature matrix, i.e., metadata. The feature of each sample in the metadata set is the prediction result of each base learner;

[0123] Step 5.4 Generate metadata on the validation set:

[0124]

[0125] where, and are the prediction results of SVM, Random Forest, and Logistic Regression on the validation set respectively;

[0126] Step 5.5 Generate metadata on the test set:

[0127] Ztest = [M SVM (X test ), M RF (X test ), M LR (X test )]

[0128] Step 5.6 Train the meta - learner: The meta - learner selects the Random Forest algorithm, and its training process and prediction process are the same as those in Step 5.2; Use the generated meta - data to train the meta - learner. The input of the meta - learner is the prediction result of the base learner, and the output is the final prediction result;

[0129] Step 6 Generate the meta - data of the test set and make predictions: Use the trained meta - learner to predict the meta - data of the test set to generate the final classification result.

[0130]

[0131] Among them, M meat is the trained meta - learner, and Z test is the meta - data of the test set.

[0132] Example 1:

[0133] The publicly available dataset used in this example is healthcare - dataset - stroke - data.csv. This dataset contains 5110 records. Among them, the number of samples in the majority class (no stroke) is 4861, accounting for 95.1%; The number of samples in the minority class (stroke) is 249, accounting for 4.9%. The number of minority - class samples is much less than that of the majority class, and the distribution of these data shows a serious imbalance.

[0134] First, select all features (including gender, age, whether having hypertension, whether having heart disease, marital status, work type, average blood glucose level, BMI, smoking status, etc.) and label values (whether having stroke) in the dataset for pre - processing. For categorical features, we used One - Hot Encoding to numericalize them. For the missing values in the BMI feature, we used the mean filling method to complete them. Due to the large difference in the numerical scales of the features, we also standardized all features to ensure that they are in the same range to avoid the impact of feature scale differences on model training. Finally, after pre - processing, the data is divided into two parts: majority - class samples (non - stroke cases) and minority - class samples (stroke cases), which are respectively used for subsequent imbalanced sample processing.

[0135] Next, the minority class samples are augmented by improving the SMOTE algorithm. Specifically, a sample is randomly selected from the minority class samples as the reference point (key). Suppose the set of minority class samples is D minority ={x1, x2,..., x n}, and the reference point can be randomly selected from it, denoted as key = X key . To define the local distribution characteristics of the minority class samples, the K-Nearest Neighbor algorithm (KNN) is used to find the nearest neighbor set of the reference point in D minority . In this embodiment, the number of nearest neighbors K is set to 5. After excluding the reference point itself, two points are randomly selected from its neighbors, named the reference point (ref_point) and the additional point (extra_point) respectively. The existence of these two points provides the direction and range for generating new samples, and defines the local geometric distribution characteristics.

[0136] Based on the selected reference point and neighbor points, two direction vectors are calculated: the vector V ref =X ref -X key from the reference point to the reference point, and the vector V extra =X extra -X key . To ensure that the generated new sample points are more reasonably distributed in the multi-dimensional feature space, feature weights are further introduced to adjust the direction vectors. The calculation of the weights is based on the variance of each feature in the minority class samples. The calculated weight w i reflects the importance of each feature. The larger the weight, the more significant the influence of the feature on the distribution of the minority class samples. The adjusted direction vectors are V' ref and V' extra .

[0137] To control the generation range of the new samples, the lengths and angles of the direction vectors are further restricted. First, calculate the lengths of the two direction vectors len_ref = ||V' ref || and len_ref = ||V' extra ||, and select the smaller value as the radius min_radius of the sector area. Subsequently, use the dot product formula to calculate the angle θ between the two direction vectors, and obtain the specific angle value angle through the arccosine function. This angle θ is used to limit the angle range of the sector area, thereby controlling the direction distribution of the new sample generation. In addition, based on the radius and angle, calculate the arc length of the sector arc_length = min_radius · angle, which is further used to randomly generate the position of the new sample points.

[0138] Randomly select a position on the arc length within the sector area, and determine the angle of the new direction by calculating the angle offset θ = α·angle. Use this angle to generate a new direction vector V in combination with the weighted direction vector. new Finally, add the new direction vector to the reference point position to generate a new sample point X. new =X key +V new 。Generate multiple minority class samples by repeating the above process multiple times until the target quantity is reached.

[0139] After completing the augmentation of the minority class samples, next, merge the newly synthesized dataset with the original majority class samples to form a new balanced dataset. Ensure the correctness of the sample labels during the merging process. Then, re-partition the merged dataset into a training set and a validation set (or test set), usually with a ratio of 70% for the training set and 30% for the validation set (or test set), ensuring the uniformity of the class distribution during the partitioning. The training set is X train and y train ,The validation set is X val and y val 。

[0140] Perform 5-fold cross-validation on the training set X train and y train 。Divide the training set into 5 subsets, and each time use 4 subsets to train the model, and the remaining 1 subset is used for validation. The training set is divided into 5 subsets X1, X2, X3, X4, X5, and the corresponding labels are y1, y2, y3, y4, y5. For each base learner (random forest, support vector machine, and logistic regression), in the first iteration: use (X1, X2, X3, X4) to train the model and use X5 to validate the model, and so on. After each iteration of the base learner, record the performance metrics (such as accuracy, recall, F1 score, etc.) of the model on the validation set, and finally take the average of these metrics as the performance evaluation of the base learner.

[0141] Use the trained base learner to make predictions on the validation set X val to obtain the prediction probabilities of each base learner. Assume that the prediction probabilities of the base learners are P RF 、P SVM and P LP ,Then each column in the meta-data is:

[0142] Meta-Data = [P RF , P SVM , P LR

[0143] Finally, use the generated meta-data Meta-Data as the input and the true class label y val ​As output, train the meta-learner. Select the random forest as the meta-learner. Use the trained meta-learner to predict the meta-data of the test set and generate the final classification result. The experimental results of this model in this embodiment are shown in Table 1:

[0144] The acute ischemic stroke disease prediction method based on the improved SMOTE algorithm and ensemble learning effectively deals with the class imbalance problem and significantly improves the recognition ability of the minority class (stroke cases). The experimental results show that while improving the prediction accuracy of stroke diseases, this method effectively reduces the risks of misdiagnosis and missed diagnosis, providing a reliable auxiliary tool for clinical diagnosis.

[0145] Table 1 Experimental results of the model on the public dataset: healthcare-dataset-stroke-data

[0146]

[0147] It should be understood that those of ordinary skill in the art can make improvements or transformations according to the above description, and all such improvements and transformations shall fall within the protection scope of the appended claims of the present invention.

Claims

1. A stroke diagnosis system based on an improved SMOTE algorithm and ensemble learning, characterized in that, Including: Data preparation and preprocessing module, imbalanced dataset processing module, dataset generation module, dataset division module, learner training module, prediction module; Data preparation and preprocessing module: Used to collect case data, divide samples into normal samples and diseased samples, and respectively form a majority class dataset and a minority class dataset; Then, preprocess the dataset, including numerical conversion of categorical features, filling missing values, and standardizing all features to unify the scale of the data in each dimension; Imbalanced dataset processing module: Use the improved SMOTE algorithm to process the imbalanced dataset to obtain a balanced dataset; Dataset generation module is used to synthesize a new dataset: That is, merge the data processed by the improved SMOTE algorithm with the original majority class samples to form a new and relatively balanced dataset; Dataset division module divides the processed balanced dataset, including the feature matrix X and the label vector Y, into a training set and a test set; Learner training module: Use the Stacking ensemble learning algorithm to train the base learner and the meta-learner of the Stacking algorithm for final prediction; Prediction module: Generate metadata of the test set and make predictions.

2. The stroke diagnosis system according to claim 1, wherein The processing flow of the imbalanced dataset processing module is as follows: Step 2.1 Select a reference point: Randomly select a point from all the minority class samples as the reference point key. Assume the minority class sample set is D minority ={x1, x2,..., x n}, then the reference point key can be obtained through a random selection operation: key = x key where x key is a sample point randomly selected from D minority ; Step 2.2 Select neighbor points: Use the K-nearest neighbor algorithm in the minority class samples to find the nearest neighbor set of the reference point. After excluding the reference point itself, randomly select two points from the remaining neighbors and name them ref_point and extra_point respectively; Step 2.3 Calculate the direction vectors: Calculate the direction vectors from the reference point to these two neighbor points respectively, that is, the vector from the reference point key to ref_point and the vector from the reference point key to extra_point; The formula is as follows: V ref = X ref - X key , V extra = X extra - X key Step 2.4 Introduce weights weights in multiple dimensions: Adjust each feature dimension. The weights are calculated based on the variances of the features of the minority class samples. The larger the weight, the more important the feature; After calculating the weights, adjust the direction vectors in combination with the weights; Weight calculation formula: where d is the number of dimensions of the feature, and Var(X minority,i ) is the variance of the i-th feature; the adjusted direction vectors are V' ref and V' extra ; Step 2.5 Calculate the vector length: Calculate the lengths of the two direction vectors, and select the smaller value as the "radius" of the sector, called min_radius, to control the range of generating new points; The formula is as follows: len_ref = ||V' ref ||, len_ref = ||V' extra || min_radius = min(len_ref, len_extra) Step 2.6 Calculate the included angle: Use the dot product formula to calculate the included angle between the two direction vectors; This included angle will be used to determine the angle range of the sector area; The formula is as follows: Among them, · represents the dot product, and then obtain the included angle through the arccosine function: angle = arccos(cos(θ)) Step 2.7 Calculate the arc length: Calculate the arc length arc_length of a sector area according to the minimum radius min_radius and the included angle angle, and the formula is as follows: arc_length = min_radius × angle Step 2.8 Randomly generate the position of the new point: Randomly select a position random_position on the sector arc length arc_length, and calculate the angular offset θ; the formula is as follows: θ = α · angle Step 2.9 Generate a new vector through linear interpolation: Use the angle θ to construct a new direction vector new_vector, which is the result of the weighted combination of vector_ref and vector_extra by sine and cosine; finally, generate a new sample point new_point by adding the new vector new_vector to the reference point key: X new = X key + V new .

3. The stroke diagnosis system according to claim 1, wherein The described dataset partitioning module: To generate metadata, the training set needs to be further divided into a training part and a validation part; using the K-fold cross-validation method, the training set is divided into K subsets, and each time K-1 subsets are used to train the model, and the remaining one subset is used for validation.

4. The stroke diagnosis system according to claim 1, characterized in that, The learner training module performs the following steps: Step 5.1 Select the base learners: Use support vector machines, random forests, and logistic regression as the base learners; Step 5.2 Train the base learners: Support vector machine training process: s.t.y i (w T x i +b)≥1-ζ i ,ζ i ≥0, i=1,2,...,N Prediction process: Random forest training process: Randomly select samples and features to construct multiple decision trees; each tree uses the Bootstrap sampling method to draw samples from the training set; for each node, randomly select a part of the features for splitting; Prediction process: where f(x) is the prediction result of the j-th tree, and T is the number of trees; Logistic Regression: Training process: Among them, is the probability that the sample x i belongs to the positive class; Prediction process: Among them, is the sigmoid function; Step 5.3 Generate metadata: Combine the prediction results of all base learners on the validation set to form a new feature matrix, that is, metadata; the feature of each sample in the metadata set is the prediction result of each base learner; Step 5.4 Generate metadata on the validation set: Among them, and are the prediction results of SVM, random forest, and logistic regression on the validation set, respectively; Step 5.5 Generate metadata on the test set: Z test = [M SVM (X test ), M RF (X test ), M LR (X test )] Step 5.6 Train the meta-learner: The meta-learner selects the Random Forest algorithm, and its training process and prediction process are the same as those in Step 5.2; use the generated metadata to train the meta-learner, the input of the meta-learner is the prediction result of the base learner, and the output is the final prediction result.

5. The stroke diagnosis system according to claim 1, characterized in that The prediction module uses the trained meta-learner to predict the metadata of the test set to generate the final classification result; Among them, M meat is a trained meta-learner, and Z test is the metadata of the test set.