Breast cancer prediction method based on feature reinforcement learning

Through feature enhancement learning methods, highly correlated breast cancer enhancement features are generated, which solves the problems of unrelated feature interference and uneven category distribution in the existing models, and improves the accuracy and accuracy of breast cancer prediction.

CN119943355APending Publication Date: 2025-05-06THE FIRST AFFILIATED HOSPITAL OF JINZHOU MEDICAL UNIV
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510137424.0
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-02-07
Publication Date
2025-05-06

AI Technical Summary

Technical Problem

Existing breast cancer prediction models ignore irrelevant or irrelevant features when processing patient data, resulting in a decrease in prediction accuracy. Due to the uneven proportion of benign and malignant in the dataset, the model is prone to missed detection of malignant lesions.

Method used

A method based on feature enhancement learning is adopted to generate enhanced features highly correlated with breast cancer detection results through cluster analysis, and a breast cancer prediction model is constructed to solve the problems of unrelated feature interference and uneven category distribution.

Benefits of technology

It improves the accuracy of the breast cancer prediction model, reduces the missed detection rate, and ensures the accuracy of the prediction results.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119943355A_ABST
    Figure CN119943355A_ABST
Patent Text Reader

Abstract

The invention discloses a breast cancer prediction method based on feature reinforcement learning, and the method comprises the steps: obtaining a first data set, carrying out the cleaning and preprocessing of the first data set, obtaining a training set and a test set, and dividing the training set into a first data set and a second data set; clustering the first data set to obtain a first cluster and a first clustering center, and similarly obtaining a second cluster and a second clustering center; calculating the distance from the first data set to the first clustering center to obtain a first distance; similarly, obtaining a second distance; combining the first distance and the second distance as a first enhanced feature; and performing feature reinforcement learning on the features of the data in the training set, calculating the distance from the test set to a third cluster center, obtaining a second reinforcement feature to replace the original features of the test set data, calculating the test set data and the training set data, and finding a preset number of nearest neighbors to obtain a final prediction result. Interference of irrelevant or irrelevant features on the model is solved, model prediction precision is improved, and unbalanced distribution of benign and malignant proportions in a data set is relieved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of data processing, and in particular to a breast cancer prediction method based on feature enhancement learning. Background Art

[0002] Breast cancer is one of the most common cancers in women, with high morbidity and mortality. According to a report by the International Agency for Research on Cancer (IARC) in December 2020, breast cancer is diagnosed more frequently in women than lung cancer. On average, one in five people may develop breast cancer, and the incidence rate will continue to rise in the future, becoming one of the leading causes of death among women worldwide. Therefore, early diagnosis and prediction of breast cancer are particularly important for breast cancer prevention and treatment. In recent years, with the advancement of artificial intelligence and intelligent medical technology, researchers have collected and sorted historical examination data of breast cancer patients, analyzed them with the help of machine learning models, and built an efficient breast cancer prediction model for early screening and diagnosis of cancer.

[0003] Existing machine learning solutions mainly build a breast cancer prediction model based on models such as SVM (Support Vector Machine), LogisticRegression, KNN (K-Nearest Neighbors), multi-layer perceptron, and deep neural network. Although these methods have achieved good prediction results, they also face the following two challenges: First, existing solutions ignore the fact that there may be many factors in patient data records that are irrelevant or unrelated to breast cancer test results. Most existing solutions directly build models based on raw data features, and irrelevant or unrelated feature factors will interfere with the model and reduce the accuracy of prediction.

[0004] Second, existing solutions ignore the problem of uneven distribution of benign and malignant ratios in patient records. The benign ratio in the collected data is usually much higher than the malignant ratio. For example, in the famous BCWD (Breast Cancer Wisconsin Diagnostic) public dataset, the benign ratio is 62.74% and the malignant ratio is 37.26%. Unbalanced category distribution will cause the model to be more inclined to predict benign, thus causing the problem of missed detection. Summary of the invention

[0005] The purpose of the present invention is to overcome the deficiencies of the prior art and provide a breast cancer prediction method based on feature enhancement learning.

[0006] The objective of the present invention is achieved through the following technical solutions: The present invention discloses a breast cancer prediction method based on feature enhancement learning, comprising the following steps: S1, obtaining a first data set and cleaning it, then preprocessing the first data set to obtain a training set and a test set, and dividing the training set into a first data set and a second data set; S2, clustering the first data set to obtain a first cluster and a corresponding first cluster center, and clustering the second data set to obtain a second cluster and a corresponding second cluster center; S3, calculating the distance from all data in the first data set to the first cluster center to obtain a first distance; calculating the distance from all data in the second data set to the second cluster center to obtain a second distance; merging the first distance and the second distance into a vector as a first enhanced feature; S4, setting the dimension of the first enhanced feature to the dimension of the original features of all data in the training set; S5. Perform feature enhancement learning on the features of all data in the training set, that is, use the first enhanced feature to replace the original feature of the training set data, then calculate the distance from the data in the test set to the center of the third cluster to obtain the second enhanced feature; use the second enhanced feature to replace the original feature of the test set data, and based on the second enhanced feature, calculate the Euclidean distance between the test set data and the training set data, find a preset number of nearest neighbors, and vote according to the breast cancer results of the preset number of nearest neighbors to obtain the final prediction result.

[0007] Further, step S1 specifically includes: obtaining a first data set and cleaning it, and then preprocessing the first data set, wherein the preprocessing operation includes digital encoding, obtaining a training set and test set, the training set Divide into the first data set and the second data set ;in Indicates The feature vector of the training data, Indicates the corresponding breast cancer test results, is the data of the training set, is the feature dimension.

[0008] Preferably, step S2 specifically includes: based on K-Means, the first data set Perform cluster analysis using the Euclidean distance formula Clustering is performed, where Represents the first data set No. cluster centers, and obtain The first clusters and the corresponding first cluster centers ; Based on K-Means for the second data set Clustering is performed using the Euclidean distance formula Clustering is performed; Represents the second data set No. cluster centers, and obtain Second clusters and corresponding second cluster centers .

[0009] Preferably, step S3 specifically includes: for the first cluster center , calculate the training data To the cluster center Distance: ,get First distance ; For the second cluster center , calculate the training data To the cluster center Distance: ,get Second distance ; Will First distance and Second distance Combine into one vector As the first enhancement feature.

[0010] Preferably, step S4 specifically includes: setting the dimension of the first enhanced feature to the dimension of the original features of all data in the training set, that is, ; The number of data records in the first data set is , the number of the first cluster and the corresponding first cluster center is updated to ; The number of data records in the second data set is , the number of the second clusters and the corresponding second cluster centers is updated to .

[0011] Preferably, step S5 specifically includes: performing feature enhancement learning on the features of all data in the training set, that is, using the first enhanced feature to replace the original feature of the training set data, and then calculating the distance from the data in the test set to the third cluster center, where the third cluster center is The first cluster center plus The second enhanced features are used to replace the original features of the test set data. Based on the second enhanced features, the Euclidean distance between the test set data and the training set data is calculated to find the K nearest neighbors. The final prediction result is obtained by voting according to the breast cancer results of the K nearest neighbors.

[0012] The beneficial effects of the present invention are: 1) The present invention solves the interference of irrelevant or unrelated features on the model and improves the accuracy of model prediction.

[0013] 2) Alleviate the problem of uneven distribution of benign and malignant ratios in the data set, and reduce the missed detection rate of model prediction, that is, predicting malignant as benign. BRIEF DESCRIPTION OF THE DRAWINGS

[0014] Figure 1 A flowchart of a method for predicting breast cancer based on feature-enhanced learning according to an embodiment of the present invention; Figure 2 It is a schematic diagram comparing the accuracy and AUC value results of the present invention and the existing KNN method according to an embodiment of the present invention. DETAILED DESCRIPTION

[0015] The technical solution of the present invention will be clearly and completely described below in conjunction with the embodiments. Obviously, the described embodiments are only part of the embodiments of the present invention, rather than all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative work are within the scope of protection of the present invention.

[0016] The present invention discloses a breast cancer prediction method based on feature enhancement learning, which enhances the original data features and builds a breast cancer prediction model based on the enhanced features, which can better solve the two problems existing in the prior art. Figure 1 As shown, the method comprises the following steps: S1, obtaining a first data set and cleaning it, then preprocessing the first data set to obtain a training set and a test set, and dividing the training set into a first data set and a second data set; S2, clustering the first data set to obtain a first cluster and a corresponding first cluster center, and clustering the second data set to obtain a second cluster and a corresponding second cluster center; S3, calculating the distance from all data in the first data set to the first cluster center to obtain a first distance; calculating the distance from all data in the second data set to the second cluster center to obtain a second distance; and summing the first distance and the second distance. The data are separated and merged into a vector as the first enhanced feature; S4, the dimension of the first enhanced feature is set to the dimension of the original feature of all the data in the training set; S5, feature enhancement learning is performed on the features of all the data in the training set, that is, the original features of the training set data are replaced by the first enhanced feature, and then the distance from the data in the test set to the third cluster center is calculated to obtain the second enhanced feature; the original features of the test set data are replaced by the second enhanced feature, and based on the second enhanced feature, the Euclidean distance between the test set data and the training set data is calculated to find the preset number of nearest neighbors, and the final prediction result is obtained by voting according to the breast cancer results of the preset number of nearest neighbors.

[0017] For example, step S1 specifically includes: obtaining a first data set (obtaining a BCWD public data set) and cleaning it, removing data records containing missing attribute values, and then preprocessing the features and tags of the first data set, wherein the preprocessing operation includes numerical coding, obtaining a training set and test set, the training set Divide into the first data set (benign set) and the second data set (a vicious set); Indicates The feature vector of the training data, Indicates the corresponding breast cancer test result, with a value of 0 (benign) or 1 (malignant). is the data of the training set, is the feature dimension.

[0018] Specifically, the data set is clustered to obtain the cluster centers of the benign and malignant sets respectively. Step S2 specifically includes: clustering the first data set based on K-Means Perform cluster analysis using the Euclidean distance formula Clustering is performed, where Represents the first data set No. cluster centers, and obtain The first clusters and the corresponding first cluster centers ; Based on K-Means for the second data set Clustering is performed using the Euclidean distance formula Clustering is performed; Represents the second data set No. cluster centers, and obtain Second clusters and corresponding second cluster centers .

[0019] Specifically, cluster center-guided feature enhancement learning, benign collection The cluster center of is obtained by clustering all benign records, so this cluster center encodes the features related to benign records; similarly, the malignant set The cluster center is obtained by clustering all malignant records. This cluster center encodes the features related to malignant records. We calculate the distance from each training data to these cluster centers, and merge these distances into a vector as the learned enhanced features that are strongly related to breast cancer detection results. Based on these two cluster center sets, feature enhancement learning that is strongly related to breast cancer results is performed. Step S3 specifically includes: for the first cluster center , calculate the training data To the cluster center Distance: ,get First distance ; For the second cluster center , calculate the training data To the cluster center Distance: ,get Second distance ;Will First distance and Second distance Combine into one vector As the first enhancement feature.

[0020] Specifically, feature enhancement learning parameter setting, there is usually a class imbalance problem in the training data, that is, the proportion of benign data is much higher than the proportion of malignant data. Next, we use different parameters to perform feature enhancement learning to alleviate the class imbalance problem; step S4 specifically includes: setting the dimension of the first enhanced feature to the dimension of the original feature of all data in the training set, that is ; The number of data records in the first data set is , the number of the first cluster and the corresponding first cluster center is updated to ; The number of data records in the second data set is , the number of the second clusters and the corresponding second cluster centers is updated to When the proportion of benign records is higher than that of malignant records, the present invention will generate more cluster centers related to malignant results, thereby learning more enhanced features associated with malignant results and alleviating the problem of unbalanced category distribution.

[0021] Specifically, a K-nearest neighbor breast cancer prediction model based on enhanced features is constructed; step S5 specifically includes: performing feature enhancement learning on the features of all data in the training set, that is, using the first enhanced feature to replace the original feature of the training set data, and then calculating the distance from the data in the test set to the third cluster center, the third cluster center is The first cluster center plus The second enhanced feature is used to replace the original feature of the test set data, and based on the second enhanced feature, the Euclidean distance between the test set data and the training set data is calculated to find K nearest neighbors, and the breast cancer results of the K nearest neighbors are voted to obtain the final prediction result, wherein the preset number K is a number preset according to the working environment and is not limited to a specific value. Usually, the value of K is an odd number. In this embodiment, K is usually set to 5.

[0022] Most of the existing solutions directly build a machine learning model for breast cancer prediction on the original features. Compared with these solutions, the key point of the present invention is to design a new feature enhancement learning method for breast cancer prediction, and build a breast cancer prediction model based on the enhanced features. In the present invention, the enhanced features are obtained based on the cluster centers of the benign and malignant data sets, so the enhanced features are highly correlated with the breast cancer detection results. The present invention solves the interference of irrelevant or irrelevant features on the prediction model. There may be many features that are irrelevant or irrelevant to the breast cancer detection results in the original data records, which interfere with the construction of the prediction model and reduce the model prediction accuracy. Compared with the original features, the enhanced features are highly correlated with the breast cancer results. The breast cancer prediction model built based on the enhanced features is more robust and has higher classification accuracy. Alleviate the missed detection problem caused by the imbalanced distribution of categories. Usually, there is an imbalanced category problem in the training data, that is, the benign ratio is much higher than the malignant ratio. Most of the existing solutions do not consider the imbalanced distribution of categories, and the constructed models are more inclined to predict the test data as benign, which increases the missed detection rate of the model. Compared with these solutions, the present invention alleviates the problem of imbalanced distribution of categories and reduces the missed detection problem of the model by setting different parameters in the feature enhancement learning process. We compared the prediction results of the present invention and the KNN algorithm on the public BCWD data set, with half of the data used as a training set and the other half as a test set. Figure 2 compares the accuracy and AUC values ​​of the present invention method and the KNN method on the test set, and the present invention method achieves higher accuracy and AUC values. It alleviates the problem of uneven distribution of benign and malignant ratios in the data set and reduces the missed detection rate of model prediction (i.e., predicting malignant as benign). The benign ratio in the BCWD data is 62.74% and the malignant ratio is 37.26%, and the benign ratio is much higher than the malignant ratio. Therefore, the existing scheme is more inclined to predict benign and has a higher missed detection rate. In contrast, the present invention alleviates this problem by enhancing features. Table 1 is the confusion matrix results of the prediction results of the present invention on the BCWD data set, in which only one of the 211 malignant results was missed, greatly reducing the missed detection rate of the model.

[0023] Table 1: Confusion matrix results of the prediction results of the present invention on the BCWD dataset Predicted to be malignant Prediction is benign The truth is evil 210 1 Truth is good 3 354 The above is only a preferred embodiment of the present invention. It should be understood that the present invention is not limited to the form disclosed herein, and should not be regarded as excluding other embodiments, but can be used in various other combinations, modifications and environments, and can be modified within the scope of the concept described herein through the above teachings or the technology or knowledge of the relevant field. The changes and modifications made by those skilled in the art shall not deviate from the spirit and scope of the present invention, and shall be within the scope of protection of the claims attached to the present invention.

Claims

1. A breast cancer prediction method based on feature enhancement learning, characterized in that: The following steps are involved: S1, obtaining a first data set and cleaning it, then preprocessing the first data set to obtain a training set and a test set, and dividing the training set into a first data set and a second data set; S2, clustering the first data set to obtain a first cluster and a corresponding first cluster center, and clustering the second data set to obtain a second cluster and a corresponding second cluster center; S3, calculating the distance from all data in the first data set to the first cluster center to obtain a first distance; Calculate the distance from all data in the second data set to the second cluster center to obtain a second distance; merge the first distance and the second distance into a vector as a first enhanced feature; S4, setting the dimension of the first enhanced feature to the dimension of the original features of all data in the training set; S5, performing feature enhancement learning on the features of all data in the training set, that is, using the first enhanced feature to replace the original feature of the training set data, and then calculating the distance from the data in the test set to the third cluster center to obtain the second enhanced feature; The second enhanced feature is used to replace the original feature of the test set data. Based on the second enhanced feature, the Euclidean distance between the test set data and the training set data is calculated to find a preset number of nearest neighbors. The final prediction result is obtained by voting based on the breast cancer results of the preset number of nearest neighbors.

2. A breast cancer prediction method based on feature enhancement learning according to claim 1, characterized in that: Step S1 specifically includes: obtaining a first data set and cleaning it, and then preprocessing the first data set, wherein the preprocessing operation includes digital coding, obtaining a training set and test set, the training set Divide into the first data set and the second data set ;in Indicates The feature vector of the training data, Indicates the corresponding breast cancer test results, is the data of the training set, is the feature dimension.

3. A breast cancer prediction method based on feature enhancement learning according to claim 2, characterized in that: Step S2 specifically includes: based on K-Means, the first data set Perform cluster analysis using the Euclidean distance formula Clustering is performed, where Represents the first data set No. cluster centers, and obtain The first clusters and the corresponding first cluster centers ; Based on K-Means for the second data set Clustering is performed using the Euclidean distance formula Clustering is performed; Represents the second data set No. cluster centers, and obtain Second clusters and corresponding second cluster centers .

4. The method for predicting breast cancer based on feature-enhanced learning according to claim 3, characterized in that: Step S3 specifically includes: for the first cluster center , calculate the training data To the cluster center Distance: ,get First distance ; For the second cluster center , calculate the training data To the cluster center Distance: ,get Second distance ;Will First distance and Second distance Combine into one vector As the first enhancement feature.

5. The method for predicting breast cancer based on feature-enhanced learning according to claim 4, characterized in that: Step S4 specifically includes: setting the dimension of the first enhanced feature to the dimension of the original features of all data in the training set, that is, ; The number of data records in the first data set is , the number of the first cluster and the corresponding first cluster center is updated to ; The number of data records in the second data set is , the number of the second clusters and the corresponding second cluster centers is updated to .

6. A method for predicting breast cancer based on feature-enhanced learning according to claim 5, characterized in that: Step S5 specifically includes: performing feature enhancement learning on the features of all data in the training set, that is, using the first enhanced feature to replace the original feature of the training set data, and then calculating the distance from the data in the test set to the third cluster center, where the third cluster center is The first cluster center plus The second enhanced features are used to replace the original features of the test set data. Based on the second enhanced features, the Euclidean distance between the test set data and the training set data is calculated to find the K nearest neighbors. The final prediction result is obtained by voting according to the breast cancer results of the K nearest neighbors.