Breast cancer risk prediction method based on machine learning and multi-dimensional data

By integrating multi-dimensional data through machine learning, a breast cancer risk prediction model is constructed, which solves the limitations of the single data source in existing technologies and realizes high-precision, personalized and dynamic breast cancer risk assessment, which is suitable for early prediction and health management of the general population.

CN120674051APending Publication Date: 2025-09-19THE SECOND AFFILIATED HOSPITAL OF GUANGXI UNIV OF SCI & TECH
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
CN202510635048.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-05-16
Publication Date
2025-09-19

AI Technical Summary

Technical Problem

Existing breast cancer risk assessment methods rely on a single data source, cannot fully reflect an individual's health status, have limited prediction accuracy, restricted applicability, high cost, and are difficult to achieve personalized and real-time dynamic risk prediction.

Method used

A multi-dimensional data integration method based on machine learning is used to integrate multi-source data such as diet, lifestyle habits, and physiological indicators through complementary feature selection and support vector machine models to construct a breast cancer risk prediction model for personalized assessment and dynamic prediction.

Benefits of technology

It improves the comprehensiveness and accuracy of breast cancer risk prediction, provides personalized customized solutions, reduces costs, is suitable for early risk prediction in the general population, and supports dynamic health management.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120674051A_ABST
    Figure CN120674051A_ABST
Patent Text Reader

Abstract

The invention discloses a breast cancer risk prediction method based on machine learning and multi-dimensional data, and belongs to the technical field of medical health information.The method comprises the steps that based on an NHANES database, diet, living habits and other information are collected, and a data set is formed; performing pretreatment; screening meaningful data features by using three feature selection methods of LASSO regression, mRMR and forward selection, and obtaining a final feature data set after intersection; dividing a training set and a test set; establishing a risk prediction model by using an SVM machine learning method, and learning the training set; performing model performance analysis on the test set to obtain a risk prediction probability of a final training model; the method has the advantages of multi-source data integration, high-precision prediction, personalized evaluation, dynamic updating and the like, risk factors of the breast cancer can be effectively mined, theoretical support is provided for prevention and treatment of the breast cancer, high-risk group screening is guided, morbidity reduction is assisted, early diagnosis and early treatment are achieved, and development of female health undertaking is promoted.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of medical health information technology, and in particular to a breast cancer risk prediction method based on machine learning and multi-dimensional data. Background Art

[0002] Breast cancer is one of the most common malignant tumors in women worldwide. Early prediction and risk assessment are important for improving cure rates and reducing mortality. Currently, breast cancer risk assessment mainly relies on the following methods:

[0003] 1. Clinical risk assessment models: These include the Gail model and the Tyrer-Cuzick model. These models are typically based on clinical data such as a patient's age, family history, and reproductive history. However, their drawback is that these models rely on a single type of clinical data and fail to fully reflect an individual's health status, particularly risk factors related to diet and lifestyle.

[0004] 2. Genetic risk assessment: This method assesses breast cancer risk based on genetic testing data (such as BRCA1 / BRCA2 gene mutations). However, its drawbacks are that genetic testing is expensive and only applicable to a small number of people with a clear genetic predisposition, not the majority of the general population.

[0005] 3. Imaging tests, such as mammography and ultrasound, are used for early screening of breast cancer. However, these tests are limited in that they are typically used for screening in the middle and late stages of the disease, cannot predict early risk, and carry the risk of radiation exposure and misdiagnosis.

[0006] 4. Traditional statistical analysis methods: Risk assessment is based on statistical methods such as regression analysis and correlation analysis, combined with a small amount of health data (such as body mass index and smoking history). However, their drawbacks are that traditional statistical methods have difficulty processing high-dimensional, multi-source data and have limited predictive accuracy.

[0007] 5. Single-data-source machine learning models: In recent years, some studies have attempted to use machine learning techniques to predict breast cancer risk, but these studies typically rely on a single type of data (such as clinical or genetic data). This limitation is that a single data source cannot fully reflect an individual's health status, resulting in inaccurate and unreliable predictions.

[0008] In general, the defects and shortcomings of the existing technology are mainly reflected in:

[0009] 1. Data uniformity: Existing methods typically rely on a single type of data (such as clinical data or genetic data) and are unable to integrate multi-source data (such as diet, lifestyle habits, physiological indicators, etc.), resulting in insufficient comprehensiveness and accuracy in risk assessment.

[0010] 2. Limited prediction accuracy: Traditional statistical methods and machine learning models based on a single data source perform poorly when processing complex, high-dimensional data and have difficulty capturing nonlinear relationships between multiple factors.

[0011] 3. Limited applicability: Genetic risk assessment and imaging examinations are expensive and only applicable to specific populations. They cannot be widely used for early risk prediction in the general population.

[0012] 4. Lack of personalization: Existing methods are usually based on group data, which makes it difficult to achieve accurate individual risk assessment and cannot fully consider personalized factors such as individual diet and lifestyle habits.

[0013] 5. Insufficient real-time and dynamic performance: Existing methods are usually based on static data and cannot dynamically reflect changes in an individual’s health status, making it difficult to achieve real-time risk prediction. Summary of the Invention

[0014] This paper proposes a breast cancer risk prediction method based on machine learning and multi-dimensional data. By constructing an accurate prediction model, it provides a scientific basis for early diagnosis, intervention, and personalized prevention and treatment strategies, transforms the results of etiology research into a practical risk assessment tool, and provides theoretical support for breast cancer prevention and treatment.

[0015] To achieve the above objectives, the technical solution adopted by the present invention is: a breast cancer risk prediction method based on machine learning and multi-dimensional data, comprising the following steps:

[0016] Obtain multi-dimensional data based on breast cancer influencing factors to form a data set;

[0017] Preprocessing the data set to obtain characteristic variables of features in the data set;

[0018] Using complementary feature selection method to perform feature selection on the feature variables to form multiple target features;

[0019] The target features are integrated by using an intersection method to select common selection features;

[0020] Dividing the data set corresponding to the common selected features into a training set and a test set;

[0021] The SVM machine learning model is trained using the training set, and then the trained SVM machine learning model is used to predict the test set to output a probability or classification result that the sample is a breast cancer patient.

[0022] Preferably, the feature variables are composed of categorical variables and continuous variables; the categorical variables are subjected to hot encoding processing, and the continuous variables are subjected to standardization processing.

[0023] Preferably, the complementary feature selection method includes forward selection method, LASSO regression method and maximum relevance minimum redundancy method.

[0024] Preferably, the forward selection method comprises the following steps:

[0025] S11: take an empty feature set as the current feature set;

[0026] S12: Select an unselected feature from the data set, and add it to the current feature set for training the model;

[0027] S13: Select the feature that improves the model performance the most and add it to the current feature set;

[0028] S14: Repeat steps S12-S13, and use the improvement in model performance and the number of features as the model iteration stopping conditions.

[0029] Preferably, the model iteration stopping condition is specifically:

[0030] Obtain the AUC value of the model and determine whether the AUC value is lower than a preset threshold; if so, stop training and use the features in the current feature set as target features; or

[0031] Determine whether the current number of features reaches the preset number of features. If the current number of features is equal to the preset number of features, stop training and use the features in the current feature set as target features.

[0032] Preferably, the LASSO regression method comprises the following steps:

[0033] S21: Z-score standard is performed on the continuous variables;

[0034] S22: Select the optimal λ value through cross-validation;

[0035] S23: Using the LASSO objective function to perform feature selection on the categorical variables and the continuous variables processed in step S21, the features corresponding to the feature variables with non-zero regression coefficients are used as target features.

[0036] Preferably, in the intersection method, the intersection of target features selected by any two complementary feature selection methods is selected as the common selection feature.

[0037] Preferably, the ratio of the training set to the test set is 7:3.

[0038] Due to the adoption of the above technical solution, the present invention has the following beneficial effects:

[0039] 1. Multi-source data integration to improve prediction comprehensiveness: This invention can comprehensively reflect an individual's health status by integrating multi-source data such as dietary intake, living habits, and physiological indicators. It overcomes the limitation of existing technologies that rely on a single data source and significantly improves the comprehensiveness and accuracy of breast cancer risk prediction.

[0040] 2. High-precision prediction: This invention uses advanced machine learning algorithms that can effectively process high-dimensional, nonlinear data and capture the complex relationships between multiple factors, thereby providing more accurate risk prediction results.

[0041] 3. Personalized risk assessment: This invention fully considers individual factors such as diet and living habits, and can provide customized risk assessment solutions for different groups of people to meet the needs of personalized health management.

[0042] 4. Wide applicability: The present invention does not rely on high-cost genetic testing or imaging examinations, is suitable for early risk prediction in the general population, and has broad application prospects.

[0043] 5. Dynamic real-time prediction: The present invention supports the input and analysis of dynamic data, can reflect the changes in an individual's health status in real time, provide dynamic risk assessment services, and help users adjust their health management strategies in a timely manner.

[0044] 6. Low cost and high efficiency: The present invention is based on existing health data (such as dietary records, lifestyle questionnaires, physiological indicator measurements, etc.), does not require additional expensive testing equipment, reduces costs, and improves evaluation efficiency through automated algorithms.

[0045] 7. Strong interpretability: The machine learning model used in this invention has good interpretability and can output key risk factors and their contribution, helping users understand the prediction results and take targeted preventive measures.

[0046] 8. Promote disease prevention and health management: This invention is not only applicable to breast cancer risk prediction, but can also provide technical support for the early prevention and health management of other chronic diseases, and has broad social benefits and health value. BRIEF DESCRIPTION OF THE DRAWINGS

[0047] Figure 1 This is a flow chart of the prediction method proposed in Example 1 of the present invention;

[0048] Figure 2 This is a flow chart of the inclusion and exclusion criteria for the study population in the prediction method proposed in Example 1 of the present invention;

[0049] Figure 3 This is a LASSO model regression coefficient trajectory diagram proposed in Example 1 of the present invention;

[0050] Figure 4This is a graph showing the change in the LASSO regression coefficient proposed in Example 1 of the present invention;

[0051] Figure 5 This is a schematic diagram of the confusion matrix of the support vector machine proposed in Example 1 of the present invention;

[0052] Figure 6 The area under the curve (AUC) and ROC curve of the SVM model on the validation set and test set proposed in Example 1 of the present invention are shown. DETAILED DESCRIPTION

[0053] The following will clearly and completely describe the technical solutions in the embodiments of the present invention in conjunction with the accompanying drawings. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without making creative efforts are within the scope of protection of the present invention.

[0054] Example 1

[0055] like Figure 1 As shown in the figure, this paper proposes a breast cancer risk prediction method based on machine learning and multidimensional data. Focusing on using machine learning to address challenging issues, the method leverages the NHANES database to deeply analyze multidimensional factors such as personal, dietary, and emotional factors within a large sample, exploring factors influencing breast cancer and identifying high-risk populations. After using three feature selection methods, LASSO regression, mRMR, and forward selection, to reduce data dimensionality and extract key features, the method is trained and tested using a support vector machine learning algorithm for cancer screening.

[0056] (1) Data source

[0057] The data used came from 101,316 respondents living in a certain country during the survey period of 1999-2018, including 49,893 males and 22,815 females under the age of 20, resulting in a total of 28,608 adult females included. Figure 2 Inclusion and exclusion criteria for the study population are shown.

[0058] (2) Data preprocessing

[0059] There are a total of 1,620 feature variables in the dataset, including 644 categorical variables and 976 continuous variables. Categorical variables are hot-coded, and continuous variables are standardized. For quantitative features, the present invention adopts a standardization method of subtracting the mean from the original data and then dividing it by the standard deviation to make the feature values ​​dimensionless, making the data stable and comparable. For feature variables such as whether to smoke, whether to suffer from hypertension and depression, when the feature value is "yes", it is assigned a value of "1"; when the feature value is "no", it is assigned a value of "0", such as "0" indicates not suffering from breast cancer, and "1" indicates suffering from breast cancer.

[0060] (3) Selection of important characteristic variables

[0061] This paper uses three complementary feature selection methods: forward selection, LASSO regression (Least Absolute Shrinkage and Selection Operator), and maximum relevance and minimum redundancy (mRMR). The following describes their principles, implementation steps, and application in breast cancer risk prediction.

[0062] ①Forward Selection

[0063] Forward selection is a greedy search algorithm and a wrapper feature selection method. Its core idea is to start with an empty feature set and gradually add features that contribute most to model performance until a preset stopping condition is reached.

[0064] Implementation steps:

[0065] (1) Initialization: Start with an empty feature set.

[0066] (2) Feature evaluation: For each unselected feature, add it to the current feature set, train the model and evaluate the performance (such as the cross-validation AUC value).

[0067] (3) Feature selection: Select the feature X* that improves the model performance the most and add it.

[0068] (4) Iteration: Repeat steps (2)-(3) until the model performance no longer improves significantly or the preset number of features is reached.

[0069] (5) Stopping condition: The present invention uses the improvement of AUC value (AUC<0.01) as the stopping condition.

[0070] The variables finally screened are shown in Table 1.

[0071] Table 1 Selected variables for forward regression

[0072]

[0073]

[0074]

[0075] ② LASSO regression

[0076] The objective function of LASSO is:

[0077]

[0078] Among them, λ is the regularization strength parameter, which controls the feature sparsity.

[0079] Implementation steps:

[0080] (1) Data standardization: Z-score standardization is performed on continuous variables to ensure consistent feature scales.

[0081] (2) Parameter tuning: Select the optimal λ value through cross-validation (such as 10-fold cross-validation), such as Figure 3 and Figure 4 shown.

[0082] (3) Feature screening: retain features with non-zero regression coefficients and eliminate features with zero regression coefficients.

[0083] The variables finally screened are shown in Table 2.

[0084] Table 2 Lasso regression selection variables

[0085]

[0086]

[0087]

[0088]

[0089] ③mRMR is a spatial search method that determines the contribution of a new variable by adding it to a set of other variables, determining the relevance and non-redundancy of the new variable. It iteratively selects features that maximize mutual information and minimize redundancy with respect to the target class, acquiring information about all features already in the selected subset.

[0090] Mutual information measures how much information one variable has about another variable, and the expression is as follows:

[0091]

[0092] Where n is the total number of instances, x is the y random variable, and x i y i is the ith value in variables x and y, P(x i ) respectively P(y i ) is the probability x i and y i Happens, is x i and y i The P(x i ,y i ) joint probability.

[0093] The redundancy of a variable f with respect to a set of other variables S can be simply defined as the total mutual information from the variable to the attribute set, as follows,

[0094]

[0095] Where f is a random variable, S is a set of random variables, x i is the i-th value in variable x, and I(x,f) is the mutual information between x and f.

[0096] Therefore, at each iteration, the goal of mRMR is to maximize the feature x of the formula j .

[0097]

[0098] Where X is the set of all attributes, S is the set of selected features, I is the mutual information function, and R is the redundancy function. This work uses a fixed number of K-selected features as the stopping criterion for the algorithm. This method generates a weighted ranking of the K-best features. The final selected variables are shown in Table 3.

[0099] Table 3 Variables selected for mRMR

[0100]

[0101]

[0102]

[0103]

[0104]

[0105] (4) Target feature integration

[0106] To enhance the robustness of feature selection, this paper uses an intersection method to integrate the results of the three methods, retaining only features that were selected by at least two methods. As shown in Table 4, the final selected features include: demographic characteristics (age, income level); lifestyle characteristics (smoking history, drinking frequency, exercise habits); clinical indicators (BMI, estrogen level, breast density); and genetic factors (family history of breast cancer).

[0107] Table 4 Number of feature selections

[0108]

[0109] Table 5 shows the intersection features of feature selection using forward selection, lasso regression, and mRMR for dietary data, measurement data, laboratory data, and questionnaire data, respectively.

[0110] Table 5 Detailed intersection features of feature selection

[0111]

[0112]

[0113]

[0114]

[0115]

[0116]

[0117] (5) Divide the training set into the test set

[0118] The training set and test set are divided into two sets at a ratio of 7:3. Stratified random sampling is used to ensure that the distribution of risk labels in the two datasets is consistent with the original data. Preferably, the training set accounts for 70% of the total sample and the test set accounts for 30%. Experimental verification shows that this ratio ensures sufficient model training while effectively evaluating generalization performance.

[0119] (6) SVM machine learning modeling and prediction

[0120] Support Vector Machine (SVM) is one of the most popular machine learning methods. It classifies samples by finding a hyperplane in feature space that maximizes the margin between points of different classes. Each data point in the dataset is plotted in N-dimensional space, where N is the number of features. Then, a hyperplane or a set of hyperplanes are found that create a boundary separating the different data classes. The hyperplane should be formed at the most stable point among the farthest data points of the different classes. It can handle nonlinearly separable data by using various kernels such as linear, polynomial, and radial basis functions, transforming the original feature space into a high-dimensional space. Table 6 shows the optimal combination of support vector machine hyperparameters.

[0121] Support Vector Machine (SVM) is a relatively new classification or prediction method developed by Cortes and Vapnik. It represents a data-driven approach that does not require assumptions about the data distribution. The core purpose of SVM is to identify a separating boundary (called a hyperplane) to help classify cases.

[0122]

[0123] x is the input vector, ω is the weight vector, b is the bias vector, Mapping the input data into a feature space with higher dimensionality than the original data.

[0124] Table 6 Main parameter settings of support vector machine

[0125]

[0126]

[0127] like Figure 5 and Figure 6 As shown in Table 7, in the prediction phase, the preprocessed data is input into the trained SVM model. The model outputs the probability that the sample is a breast cancer patient or directly gives the classification result based on the learned decision rules, as shown in Table 7.

[0128] Table 7 Support Vector Machine-Confusion Matrix

[0129]

[0130] The above description is a detailed description of the preferred embodiments of the present invention, but the embodiments are not intended to limit the scope of the patent application of the present invention. Any equivalent changes or modifications completed under the technical spirit suggested by the present invention should fall within the patent scope covered by the present invention.

Claims

1. A breast cancer risk prediction method based on machine learning and multi-dimensional data, characterized in that: The following steps are involved: Obtain multi-dimensional data based on breast cancer influencing factors to form a data set; Preprocessing the data set to obtain characteristic variables of features in the data set; Using complementary feature selection method to perform feature selection on the feature variables to form multiple target features; The target features are integrated by using an intersection method to select common selection features; Dividing the data set corresponding to the common selected features into a training set and a test set; The SVM machine learning model is trained using the training set, and then the trained SVM machine learning model is used to predict the test set to output a probability or classification result that the sample is a breast cancer patient.

2. The breast cancer risk prediction method based on machine learning and multi-dimensional data according to claim 1, characterized in that: The feature variables consist of categorical variables and continuous variables; the categorical variables are subjected to hot encoding processing, and the continuous variables are subjected to standardization processing.

3. The breast cancer risk prediction method based on machine learning and multi-dimensional data according to claim 2, characterized in that: The complementary feature selection method includes forward selection method, LASSO regression method and maximum relevance minimum redundancy method.

4. The breast cancer risk prediction method based on machine learning and multi-dimensional data according to claim 3, characterized in that: The forward selection method comprises the following steps: S11: take an empty feature set as the current feature set; S12: Select an unselected feature from the data set, and add it to the current feature set for training the model; S13: Select the feature that improves the model performance the most and add it to the current feature set; S14: Repeat steps S12-S13, and use the improvement in model performance and the number of features as the model iteration stopping conditions.

5. The breast cancer risk prediction method based on machine learning and multi-dimensional data according to claim 4, characterized in that: The model iteration stopping condition is specifically: Obtain the AUC value of the model and determine whether the AUC value is lower than a preset threshold; if so, stop training and use the features in the current feature set as target features; or Determine whether the current number of features reaches the preset number of features. If the current number of features is equal to the preset number of features, stop training and use the features in the current feature set as target features.

6. The method for predicting breast cancer risk based on machine learning and multi-dimensional data according to claim 3, characterized in that: The LASSO regression method includes the following steps: S21: Z-score standard is performed on the continuous variables; S22: Select the optimal λ value through cross-validation; S23: Using the LASSO objective function to perform feature selection on the categorical variables and the continuous variables processed in step S21, the features corresponding to the feature variables with non-zero regression coefficients are used as target features.

7. The method for predicting breast cancer risk based on machine learning and multi-dimensional data according to claim 3, characterized in that: In the intersection method, the intersection of target features selected by any two complementary feature selection methods is selected as the common selection feature.

8. The method for predicting breast cancer risk based on machine learning and multi-dimensional data according to claim 1, characterized in that: The ratio of the training set to the test set is 7:3.

Citation Information

Cited By

  • Nuclear accident health effect assessment method and system based on dual-threshold dynamic correction

    CN120913859A