Multi-modal breast cancer risk assessment method and system

Through the multimodal data fusion method, the problem of low accuracy in traditional breast cancer risk assessment is solved, and comprehensive evaluation of imaging, clinical and genetic data is used to improve the evaluation accuracy and accuracy.

CN120376125APending Publication Date: 2025-07-25TONGJI HOSPITAL ATTACHED TO TONGJI MEDICAL COLLEGE HUAZHONG SCI TECH
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510328469.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-19
Publication Date
2025-07-25

AI Technical Summary

Technical Problem

Traditional breast cancer risk assessment methods rely on a single data source and fail to fully consider multiple risk factors, resulting in low evaluation accuracy.

Method used

Multimodal data fusion method is adopted, including image data, clinical data and gene data, and the evaluation accuracy is improved through preprocessing, feature extraction, correlation analysis and comprehensive evaluation of risk assessment models.

Benefits of technology

By integrating multiple data sources, the accuracy of breast cancer risk assessment is improved, especially by incorporating multiple risk factors through the modified Gail model, and the prediction accuracy is improved to 60%-63%.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120376125A_ABST
    Figure CN120376125A_ABST
Patent Text Reader

Abstract

The embodiment of the invention provides a multi-modal breast cancer risk assessment method and system, and relates to the technical field of risk assessment. The method comprises the following steps: acquiring multi-modal data; preprocessing the multi-modal data to obtain first data; performing feature extraction on the first data through a preset first algorithm model to obtain a first feature; performing correlation analysis and screening on the first feature to obtain a second feature; and performing risk assessment on the second feature through a pre-trained assessment model to obtain a risk assessment result. Through the risk assessment method and device, the problem of low risk assessment precision is solved, and the effect of improving the risk assessment precision is achieved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] Embodiments of the present invention relate to the field of risk assessment, and more particularly, to a multi-modal breast cancer risk assessment method and system. Background Art

[0002] Traditional breast cancer risk assessment methods mainly rely on a single data source, such as imaging data (e.g., mammography or ultrasound imaging) or clinical data (e.g., patient medical records, family history, etc.). Although these methods are effective to a certain extent, they have limitations because the risk assessment models are relatively simple and do not fully consider the combined effects of multiple risk factors, resulting in low risk assessment accuracy and affecting the popularization of the risk assessment system. Summary of the Invention

[0003] Embodiments of the present invention provide a multi-modal breast cancer risk assessment method and system to improve the problem of low risk assessment accuracy in related technologies.

[0004] According to an embodiment of the present invention, a multi-modal breast cancer risk assessment method is provided, including: obtaining multi-modal data, where the multi-modal data at least includes imaging data, clinical data, and genetic data; preprocessing the multi-modal data to obtain first data; extracting features from the first data through a preset first algorithm model to obtain first features; performing correlation analysis and screening on the first features to obtain second features; and performing risk assessment on the second features through a pre-trained assessment model to obtain a risk assessment result.

[0005] In an exemplary embodiment, after performing correlation analysis and screening on the first features to obtain second features, the method further includes: constructing a feature matrix based on the second features; calculating a first correlation coefficient of the feature matrix; calculating a first test value and a second test value of the feature matrix based on the first correlation coefficient; and determining that the second features are abnormal when there is an abnormality between the first test value and the second test value.

[0006] In an exemplary embodiment, after preprocessing the multi-modal data to obtain first data, the method further includes: performing variable screening processing on the first data through a regression analysis algorithm to obtain second data; and performing grading processing on the second data through a preset second algorithm model to obtain third data, where the first data includes the third data.

[0007] In an exemplary embodiment, after preprocessing the multimodal data to obtain first data, the method further includes: performing transformation and association processing on the first data to obtain pixel gray values corresponding to the first data; determining a first image corresponding to the first data based on the pixel gray values; calculating a probability value for the first image through a preset third model, and determining that the first data is normal when the probability value is within a preset range.

[0008] According to another embodiment of the present invention, a multimodal breast cancer risk assessment system is provided, including: a data acquisition module for acquiring multimodal data, where the multimodal data at least includes imaging data, clinical data, and genetic data; a preprocessing module for preprocessing the multimodal data to obtain first data; a feature extraction module for extracting features from the first data through a preset first algorithm model to obtain first features; a first screening module for performing correlation analysis and screening on the first features to obtain second features; and a risk assessment module for performing risk assessment on the second features through a pre-trained assessment model to obtain a risk assessment result.

[0009] In an exemplary embodiment, it further includes: a matrix construction module for constructing a feature matrix according to the second features after performing correlation analysis and screening on the first features to obtain the second features; a correlation coefficient module for calculating a first correlation coefficient of the feature matrix; a test value module for respectively calculating a first test value and a second test value of the feature matrix based on the first correlation coefficient; and an abnormality determination module for determining that the second features are abnormal when there is an abnormality between the first test value and the second test value.

[0010] In an exemplary embodiment, it further includes: a variable screening module for performing variable screening processing on the first data through a regression analysis algorithm after preprocessing the multimodal data to obtain first data to obtain second data; and a grading processing module for performing grading processing on the second data through a preset second algorithm model to obtain third data, where the first data includes the third data.

[0011] In an exemplary embodiment, it further includes: an association processing module for performing transformation and association processing on the first data to obtain pixel gray values corresponding to the first data after preprocessing the multimodal data to obtain first data; an image determination module for determining a first image corresponding to the first data based on the pixel gray values; and a probability calculation module for calculating a probability value for the first image through a preset third model, and determining that the first data is normal when the probability value is within a preset range.

[0012] According to another embodiment of the present invention, there is also provided a computer-readable storage medium storing a computer program, wherein the computer program is configured to execute the steps in any one of the above method embodiments when running.

[0013] According to another embodiment of the present invention, there is also provided an electronic device including a memory and a processor, wherein the memory stores a computer program, and the processor is configured to run the computer program to execute the steps in any one of the above method embodiments.

[0014] Through the present invention, by sorting and integrating multi-modal data, comprehensive evaluation can be performed based on relevant data, effectively improving the accuracy of risk assessment. Therefore, the problem of low accuracy of risk assessment can be improved, achieving the effect of improving the accuracy of risk assessment. BRIEF DESCRIPTION OF THE DRAWINGS

[0015] Figure 1 is a flowchart of a multi-modal breast cancer risk assessment method according to an embodiment of the present invention;

[0016] Figure 2 is a structural block diagram of a multi-modal breast cancer risk assessment system according to an embodiment of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0017] Next, the technical solutions in the embodiments of the present application will be described in conjunction with the accompanying drawings in the embodiments of the present application. Obviously, the described embodiments are only a part of the embodiments of the present application, rather than all of the embodiments.

[0018] Hereinafter, terms such as "first" and "second" are only used for descriptive purposes, and cannot be understood as indicating or implying relative importance or implicitly specifying the quantity of the indicated technical features. Thus, features defined with "first", "second", etc. may explicitly or implicitly include one or more of such features. In the description of the present application, unless otherwise specified, the meaning of "a plurality" is two or more.

[0019] In addition, in the present application, orientation terms such as "upper", "lower", "left", "right", etc. may include but are not limited to being defined relative to the schematic placement of components in the drawings. It should be understood that these directional terms may be relative concepts, and they are used for relative description and clarification, and they may change accordingly with the change of the orientation of the components placed in the drawings.

[0020] In this application, unless otherwise clearly specified and defined, the term "connection" shall be understood in a broad sense. For example, "connection" can be a fixed connection, a detachable connection, or an integral one; it can be directly connected or indirectly connected through an intermediate medium. In addition, the term "coupling" can be a way of electrical connection for signal transmission.

[0021] As used herein, "about", "substantially" or "approximately" includes the stated value and the average value within an acceptable deviation range of the specific value, where the acceptable deviation range is determined by those of ordinary skill in the art considering the measurement being discussed and the errors associated with the measurement of a particular quantity (i.e., the limitations of the measurement system).

[0022] In this embodiment, a multimodal breast cancer risk assessment method is provided. Figure 1 It is a flowchart of a multimodal breast cancer risk assessment method according to an embodiment of the present invention, as Figure 1 shown, and this process includes the following steps:

[0023] Step S11, obtain multimodal data, where the multimodal data at least includes image data, clinical data, and gene data;

[0024] In this embodiment, the image data can be high-definition medical image data, such as CT, MRI, etc. Such data is usually stored in formats such as.nii. In particular, the high-definition medical images are usually Whole Slide Imaging (WSI) digital pathology slides, which are digital images converted by high-resolution scanning; the clinical data includes the medical record information, examination results, etc. of the patient, and such data is usually stored in tabular form; the gene data includes gene sequencing data, which is usually stored in formats such as FASTQ, BAM, etc.

[0025] Step S12, preprocess the multimodal data to obtain first data;

[0026] In this embodiment, the preprocessing includes processes such as cleaning and filtering the original data, cropping and splitting, filling, and performing data association conversion on the original data to convert it into a format suitable for algorithm model processing.

[0027] Taking the medical image in.nii format as an example, the preprocessing steps include:

[0028] A1, cropping: Crop out the region of interest (ROI).

[0029] A2, resampling: Resample the image to be isotropic, such as [1, 1, 1].

[0030] A3, Gray-scale region limitation: Limit the gray-scale value within a specific range, such as [-1000, -100].

[0031] A4, Normalization: Normalize the gray-scale value to the range of [0, 1].

[0032] A5, Saving: Save the preprocessed data as a new.nii file.

[0033] Step S13, Extract features from the first data through a preset first algorithm model to obtain first features;

[0034] In this embodiment, a deep learning model (such as a convolutional neural network) is used to automatically extract data features from multi-modal data such as image data, such as the shape, size, and edge features of tumors in image data, the indicators, data integrity, data diversity indicators, data real-time nature, etc. included in the examination results in clinical data, and features such as genes, non-coding regions, and repetitive sequences in gene data. For example, pre-train the BMU-Net model through transfer learning to extract modality-specific features from mammograms and ultrasound images respectively, and so on.

[0035] Step S14, Perform correlation analysis and screening on the first features to obtain second features;

[0036] In this embodiment, after extracting relevant features, it is also necessary to screen the relevant features to ensure that the relevant features can be used for risk assessment, thereby ensuring the accuracy of risk assessment.

[0037] Among them, the correlation analysis and screening can be to screen out the features highly correlated with breast cancer risk from image data, clinical data, and genomic data through methods such as correlation analysis and variance analysis. For example, retain the features with a Pearson correlation coefficient greater than 0.6, and select the most significant features from each data type, etc.

[0038] Specifically, for screening the second features from clinical data, it includes extracting clinical features related to breast cancer risk from the collected data, such as age, BMI, family history, C-reactive protein, D-dimer, homocysteine, etc. Subsequently, a Logistic regression model is used to screen out specific features. Among them, features such as C-reactive protein, D-dimer, low-density lipoprotein cholesterol, homocysteine, etc. are identified as independent risk factors (i.e., the required second features) in the multivariate Logistic regression analysis; Similarly, for the second features of gene data, a Cox regression risk model can be used to explore the key molecular events affecting the prognosis of breast cancer patients. For example, an increased expression of SFRP1 indicates a good prognosis (HR = 0.9, P = 0.015), while when has-miR-342-5p inhibits the expression of SFRP1, it indicates a poor prognosis for breast cancer patients (HR = 1.88, P = 0.016); For extracting the second features from imaging data, a Logistic regression model can be used to screen out relevant features. For example, spiculation or crab-claw-like morphology and microcalcifications are retained as significant features in the Logistic regression model, and so on.

[0039] Step S15, perform a risk assessment on the second features through a pre-trained assessment model to obtain a risk assessment result.

[0040] In this embodiment, the pre-trained model usually undergoes a large amount of data training and parameter optimization, and can more accurately identify the feature patterns related to breast cancer risk; at the same time, by comprehensively analyzing multiple risk factors through multiple models, such as age, family history, gene mutations, etc., the breast cancer risk of an individual can be more comprehensively evaluated; In particular, the improved Gail model adopted in this application finally incorporates 7 risk assessment factors including age, race, age at menarche, age at first birth, personal breast disease history, breast cancer family history, and number of breast biopsies, etc., improving the prediction accuracy to 60% - 63%.

[0041] Specifically, input the screened features into the pre-trained assessment model. For example, input features such as age, family history, BMI, etc. into the Gail model; Subsequently, the model outputs a risk score or risk level, indicating the risk of an individual having breast cancer, and this risk score or risk level is the assessment result. For example, in the Gail model, if the incidence risk within 5 years ≥ 1.67% is considered a high-risk individual, and so on.

[0042] It should be noted that considering that there are many models for evaluating the recurrence risk of breast cancer internationally at present, the most commonly used ones include 21-gene recurrence score (Oncotype DX, which requires collecting tumor tissue specimens of patients for gene testing), 70-gene recurrence score (MammaPrint, which also requires collecting tumor specimens of patients for gene testing), 50-gene recurrence risk score (PAM50), EndoPredict, Breast Cancer Index (BCI), and STEPP. To improve the recognition accuracy of the model of this application, the outputs of one or more of the above-mentioned commonly used models (such as the 21-gene score of Oncotype DX and the 70-gene score of MammaPrint) can also be used as input features of the first algorithm model, and at the same time, integrate the clinical information (such as age, tumor stage, hormone receptor status) relied on by STEPP, etc. as structured inputs, directly incorporate them into the training of the first algorithm model, and then adapt to the new dataset through the Fine-tuning algorithm, or convert gene scores, clinical indicators, imaging features, etc. into a unified feature space (such as through dimensionality reduction or embedding layers), input them into the classifier for joint training, and then provide supplementary information through pathological section image imaging data (such as MRI, ultrasound), thereby improving the recognition accuracy of the entire algorithm model.

[0043] Through the above steps, by sorting and fusing multi-modal data, it is possible to comprehensively evaluate relevant data, effectively improving the risk assessment accuracy, solving the problem of low risk assessment accuracy, and improving the risk assessment accuracy.

[0044] Among them, the execution subject of the above steps can be a base station, a terminal, etc., but is not limited thereto.

[0045] In an optional embodiment, after performing the correlation analysis and screening on the first feature to obtain the second feature, the method further includes:

[0046] Step S141, constructing a feature matrix according to the second feature;

[0047] Step S142, calculating the first correlation coefficient of the feature matrix;

[0048] Step S143, based on the first correlation coefficient, respectively calculating the first test value and the second test value of the feature matrix;

[0049] Step S144, in the case that there is an abnormality between the first test value and the second test value, determining that the second feature is abnormal.

[0050] In this embodiment, when performing risk assessment, if the significance is insufficient, it will lead to a large deviation in the risk assessment result. Therefore, it is necessary to judge the significance of the second feature.

[0051] Among them, the first correlation coefficient r can be calculated based on the Pearson coefficient; the first test value can be the t value used to reflect whether the difference between the sample mean and the population mean is significant. Generally, the larger the absolute value of the t value, the more significant the difference; the second test value can be the p value used to test the significant difference of the data. Generally, the smaller the p value, the more significant the difference between the two groups of data; it is easy to understand that in normal cases, the larger the t value, the smaller the p value. If the situation where the larger the t value, the larger the p value or the change is not obvious (that is, there is an abnormality between the p value and the t value) occurs, it means that the second feature is abnormal, and so on.

[0052] It should be noted that the t value is calculated according to the following formula 1, and the p value is calculated according to the following formula 2:

[0053]

[0054] In the formula, n is the sample size

[0055] p = (1 - tcdf(t, n - 2)) × 2 (Formula 2)

[0056] tcdf(t, n - 2)

[0057] In the formula, it is the cumulative distribution function of the t distribution. Because two-sided detection is required, it needs to be multiplied by 2.

[0058] In an alternative embodiment, after preprocessing the multi-modal data to obtain the first data, the method further includes:

[0059] Step S121, performing variable screening processing on the first data through a regression analysis algorithm to obtain second data;

[0060] Step S122, performing grading processing on the second data through a preset second algorithm model to obtain third data, where the first data includes the third data.

[0061] In this embodiment, during the process of screening data, in addition to screening the first feature, selection can also be made according to the change situation of the data itself.

[0062] Specifically, single-factor Logistic regression analysis or Cox regression analysis is used to preliminarily screen the data related to the target variable (such as the risk of disease occurrence), and at the same time calculate the p-value between the target variable and the related data. If the p-value is less than a certain threshold (usually 0.05), it is determined as abnormal and excluded; otherwise, subsequent grading is performed, and the third data after grading is used for feature extraction. Further grading of the second data is to further extract and optimize features, improve the representativeness of features, reduce the dimension of features, retain the most important information, and thus improve the accuracy of the subsequent model.

[0063] In an optional embodiment, after preprocessing the multimodal data to obtain the first data, the method further includes:

[0064] Step S123: Perform conversion and association processing on the first data to obtain the pixel gray value corresponding to the first data;

[0065] Step S124: Based on the pixel gray value, determine the first image corresponding to the first data;

[0066] Step S125: Calculate the probability value of the first image through a preset third model, and determine that the first data is normal when the probability value is within a preset range.

[0067] In this embodiment, in addition to directly evaluating the data analysis results through an algorithm model, the data can also be converted into pixel points, and whether the relevant data is reasonable can be judged according to the gray image composed of pixel points and the distribution of pixel points.

[0068] Specifically, the processed first image can be transmitted to a neural network model. Input data (such as an image) first passes through a convolutional layer where the convolutional kernel slides over the input data to extract local features and generate multiple feature maps. These feature maps then pass through a pooling layer to reduce the spatial dimension of the feature maps and retain important feature information. Subsequently, the multi-dimensional feature maps after convolution and pooling are unfolded into one-dimensional column vectors. This step converts multi-dimensional data into one-dimensional vectors for input into the fully connected layer. Then, the fully connected layer performs a linear combination through weights and biases and undergoes a non-linear transformation through an activation function (such as ReLU) to generate a new feature representation. After that, the output of the fully connected layer is input into the Softmax function, which converts the output vector into a probability distribution. Each output value represents the probability that the input data belongs to a certain category. At this time, if the probability value is within a certain threshold range, it is judged as normal; otherwise, it is judged as abnormal. Or when the distribution of the gray values of pixel points does not meet the preset requirements. For example, the distribution density of pixel points with a certain gray value in area A is A1, and the minimum and maximum values of the distances between pixel points are D1 and D2 respectively. If the distribution density and the minimum and maximum distance values in this area do not meet the above requirements, it indicates that the relevant data is abnormal, and so on.

[0069] Through the description of the above embodiments, those skilled in the art can clearly understand that the method according to the above embodiments can be implemented by means of software plus a necessary general hardware platform. Of course, it can also be implemented by hardware, but in many cases, the former is a better implementation method. Based on such an understanding, the technical solution of the present invention, in essence, or the part that contributes to the prior art, can be embodied in the form of a software product. This computer software product is stored in a storage medium (such as ROM / RAM, magnetic disk, optical disk) and includes several instructions for causing a terminal device (which can be a mobile phone, computer, server, or network device, etc.) to execute the methods described in various embodiments of the present invention.

[0070] In this embodiment, a multi-modal breast cancer risk assessment system is also provided. This system is used to implement the above embodiments and preferred implementation methods, and those that have been described will not be repeated. As used hereinafter, the term "module" can be a combination of software and / or hardware that can achieve a predetermined function. Although the devices described in the following embodiments are preferably implemented in software, implementation in hardware, or a combination of software and hardware is also possible and contemplated.

[0071] Figure 2 is a structural block diagram of a multi-modal breast cancer risk assessment system according to an embodiment of the present invention. As Figure 2 shown, the system includes:

[0072] The data acquisition module 21 is used to obtain multimodal data, where the multimodal data at least includes image data, clinical data, and gene data;

[0073] The preprocessing module 22 is used to preprocess the multimodal data to obtain first data;

[0074] The feature extraction module 23 is used to extract features from the first data through a preset first algorithm model to obtain first features;

[0075] The first screening module 24 is used to perform correlation analysis and screening on the first features to obtain second features;

[0076] The risk assessment module 25 is used to perform risk assessment on the second features through a pre-trained assessment model to obtain a risk assessment result.

[0077] In an optional embodiment, it further includes:

[0078] The matrix construction module is used to construct a feature matrix according to the second features after performing correlation analysis and screening on the first features to obtain the second features;

[0079] The correlation coefficient module is used to calculate the first correlation coefficient of the feature matrix;

[0080] The test value module is used to calculate the first test value and the second test value of the feature matrix respectively based on the first correlation coefficient;

[0081] The anomaly judgment module is used to determine that the second features are abnormal when there is an anomaly between the first test value and the second test value.

[0082] In an optional embodiment, it further includes:

[0083] The variable screening module is used to perform variable screening processing on the first data through a regression analysis algorithm after preprocessing the multimodal data to obtain the first data to obtain second data;

[0084] The grading processing module is used to perform grading processing on the second data through a preset second algorithm model to obtain third data, and the first data includes the third data.

[0085] In an optional embodiment, it further includes:

[0086] The association processing module is used to perform transformation and association processing on the first data after preprocessing the multimodal data to obtain the first data to obtain the pixel gray value corresponding to the first data;

[0087] An image determination module, configured to determine a first image corresponding to the first data based on the pixel gray values;

[0088] A probability calculation module, configured to calculate a probability value of the first image through a preset third model, and determine that the first data is normal when the probability value is within a preset range.

[0089] It should be noted that the above-mentioned various modules can be implemented by software or hardware. For the latter, it can be implemented in the following ways, but not limited to: the above-mentioned modules are all located in the same processor; or, the above-mentioned various modules are respectively located in different processors in any combination form.

[0090] An embodiment of the present invention further provides a computer-readable storage medium, in which a computer program is stored. Wherein, the computer program is configured to execute the steps in any one of the above method embodiments when running.

[0091] In an exemplary embodiment, the above computer-readable storage medium may include, but is not limited to: USB flash drives, read-only memories (ROM for short), random access memories (RAM for short), mobile hard disks, magnetic disks, or optical disks and other various media that can store computer programs.

[0092] An embodiment of the present invention further provides an electronic device, including a memory and a processor. A computer program is stored in the memory, and the processor is configured to run the computer program to execute the steps in any one of the above method embodiments.

[0093] In an exemplary embodiment, the above electronic device may further include a transmission device and an input / output device. Wherein, the transmission device is connected to the above processor, and the input / output device is connected to the above processor.

[0094] Through the description of the above embodiments, those skilled in the art can clearly understand that for the convenience and brevity of description, only the above division of each functional module is used as an example. In actual applications, the above functions can be allocated to different functional modules according to needs, that is, the internal structure of the device is divided into different functional modules to complete all or part of the functions described above.

[0095] In several embodiments provided by the present application, it should be understood that the disclosed devices and methods can be implemented in other ways. For example, the device embodiments described above are merely illustrative. For example, the division of modules or units is only a logical function division. In actual implementation, there may be other division methods. For example, multiple units or components can be combined or integrated into another device, or some features can be ignored or not executed. Another point is that the displayed or discussed couplings or direct couplings or communication connections between each other can be through some interfaces. The indirect couplings or communication connections of devices or units can be in electrical, mechanical or other forms.

[0096] The units described as separate components may or may not be physically separated. The components displayed as units may be one physical unit or multiple physical units, that is, they can be located in one place, or they can be distributed to multiple different places. Some or all of the units can be selected according to actual needs to achieve the purpose of the solution of this embodiment.

[0097] In addition, in each embodiment of the present application, each functional unit can be integrated in a processing unit, or each unit can exist physically alone, or two or more units can be integrated in one unit. The above-mentioned integrated units can be implemented in the form of hardware or in the form of software functional units.

[0098] If the integrated unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a readable storage medium. Based on such an understanding, the technical solution of the embodiments of the present application, in essence, or the part that contributes to the prior art, or all or part of the technical solution, can be embodied in the form of a software product. The software product is stored in a storage medium and includes several instructions for causing a device (which can be a single-chip microcomputer, a chip, etc.) or a processor to execute all or part of the steps of the methods of the various embodiments of the present application. The aforementioned storage medium includes: various media such as USB flash drives, mobile hard disks, read-only memories (ROM), random access memories (RAM), magnetic disks or optical discs that can store program codes.

[0099] The above content is only the specific implementation manner of the present application, but the protection scope of the present application is not limited thereto. Any changes or substitutions within the technical scope disclosed in the present application should be covered by the protection scope of the present application. Therefore, the protection scope of the present application should be subject to the protection scope of the claims.

Claims

1. A multimodal breast cancer risk assessment method, characterized in that, including: Obtain multimodal data, where the multimodal data at least includes image data, clinical data, and genetic data; Preprocess the multimodal data to obtain first data; Extract features from the first data through a preset first algorithm model to obtain first features; Perform correlation analysis and screening on the first features to obtain second features; Perform risk assessment on the second features through a pre-trained evaluation model to obtain a risk assessment result.

2. The method according to claim 1, wherein After performing correlation analysis and screening on the first features to obtain second features, the method further includes: Construct a feature matrix based on the second features; Calculate the first correlation coefficient of the feature matrix; Based on the first correlation coefficient, calculate the first test value and the second test value of the feature matrix respectively; When there is an abnormality between the first test value and the second test value, determine that the second features are abnormal.

3. The method according to claim 1, wherein After preprocessing the multimodal data to obtain first data, the method further includes: Perform variable screening on the first data through a regression analysis algorithm to obtain second data; Perform grading on the second data through a preset second algorithm model to obtain third data, and the first data includes the third data.

4. The method according to claim 1, characterized in that After preprocessing the multimodal data to obtain first data, the method further includes: Perform transformation and association on the first data to obtain the pixel gray value corresponding to the first data; Based on the pixel gray value, determine the first image corresponding to the first data; Calculate the probability value of the first image through a preset third model, and when the probability value is within a preset range, determine that the first data is normal.

5. A multimodal breast cancer risk assessment system, characterized in that, including: A data acquisition module for obtaining multimodal data, where the multimodal data at least includes image data, clinical data, and genetic data; A preprocessing module for preprocessing the multimodal data to obtain first data; A feature extraction module for extracting features from the first data through a preset first algorithm model to obtain first features; A first screening module for performing correlation analysis and screening on the first features to obtain second features; A risk assessment module for performing risk assessment on the second features through a pre-trained evaluation model to obtain a risk assessment result.

6. The system according to claim 5, wherein It further includes: A matrix construction module for constructing a feature matrix based on the second features after performing correlation analysis and screening on the first features to obtain second features ; A correlation coefficient module for calculating the first correlation coefficient of the feature matrix; A test value module for calculating the first test value and the second test value of the feature matrix respectively based on the first correlation coefficient; An abnormality determination module for determining that the second features are abnormal when there is an abnormality between the first test value and the second test value.

7. The system according to claim 5, characterized in that It further includes: A variable screening module, configured to perform variable screening processing on the first data obtained after preprocessing the multimodal data to obtain second data through a regression analysis algorithm; A grading processing module, configured to perform grading processing on the second data through a preset second algorithm model to obtain third data, where the first data includes the third data.

8. The system according to claim 5, characterized in that, It further includes: An association processing module, configured to perform conversion and association processing on the first data after preprocessing the multimodal data to obtain the pixel gray value corresponding to the first data; An image determination module, configured to determine a first image corresponding to the first data based on the pixel gray value; A probability calculation module, configured to calculate a probability value for the first image through a preset third model, and determine that the first data is normal when the probability value is within a preset range.

9. A computer-readable storage medium, characterized in that, A computer program is stored in the computer-readable storage medium, where the computer program is configured to execute the method described in any one of claims 1 to 4 when running.

10. An electronic device, comprising a memory and a processor, characterized in that, A computer program is stored in the memory, and the processor is configured to run the computer program to execute the method described in any one of claims 1 to 4.