A Breast Cancer Neoadjuvant Information Analysis System Based on Multimodal Fusion
The new adjuvant information analysis system for breast cancer, which integrates multimodal fusion and cross-attention mechanisms, solves the problems of inefficiency and information loss in data integration and analysis of traditional systems, and achieves efficient multimodal data fusion and accurate analysis results.
Patent Information
- Application Number
- CN202510951880.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-07-10
- Publication Date
- 2025-10-28
- Estimated Expiration
- 2045-07-10
AI Technical Summary
Traditional breast cancer auxiliary information analysis systems cannot fully capture tumor characteristics. The single data processing mode limits the accuracy of analysis and cannot fully explore the deep relationships between multidimensional data, resulting in the inability to fully utilize the synergistic effect between data when dealing with multidimensional data.
A multimodal fusion-based neoadjuvant information analysis system for breast cancer is adopted. It extracts features by combining 3D variable-scale convolution, 2D convolution and pyramid residual modules, and fuses multimodal data by combining triple cross-attention mechanism and spatiotemporal consistency-negative log loss function to generate high-quality feature maps and optimized analysis results.
It improved the accuracy and reliability of breast cancer analysis, enhanced the correlation between different data sources, reduced information redundancy, and significantly improved the accuracy and predictive ability of the analysis results.
Smart Images

Figure CN120452755B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of medical image analysis technology, and in particular to a neoadjuvant information analysis system for breast cancer based on multimodal fusion. Background Technology
[0002] With the continuous development of medical technology, the analysis of information for breast cancer has gradually shifted from traditional imaging and biomarker detection to comprehensive analysis of multimodal data. However, traditional breast cancer auxiliary information analysis systems still have the following shortcomings: First, due to the limitations of data sources, traditional systems cannot comprehensively capture all characteristics of the tumor; the single data processing mode of these systems limits their performance in breast cancer analysis and makes it difficult to provide more accurate analysis results; second, traditional systems often can only analyze from a single dimension and cannot fully explore the potential deep relationships between various data sources; this limitation leads to the inability of existing systems to fully utilize the synergistic effects between data when facing multidimensional data. Summary of the Invention
[0003] This invention provides a new auxiliary information analysis system for breast cancer based on multimodal fusion, aiming to improve the analysis accuracy of breast cancer-related data through the effective integration of multiple information sources. First, the system combines 3D variable-scale convolution, 2D convolution, and pyramid residual modules in its feature extraction module to perform multi-scale processing on the radiographic imaging data of breast cancer patients, extracting more representative image features and providing high-quality feature maps for subsequent analysis. Second, the system introduces a triple cross-attention mechanism in its multimodal fusion module, combined with a spatiotemporal consistency-negative logarithmic loss function, to complete the fusion processing of multimodal data. This module, through its refined fusion method, ensures the comprehensive utilization of various data types, not only improving data correlation but also effectively avoiding information redundancy and loss, thereby improving the accuracy and reliability of the analysis results. Through the above data processing and fusion methods, the breast cancer auxiliary information analysis system of this invention can efficiently integrate breast cancer-related multimodal data, ensuring maximum synergy between different data sources, thus providing more in-depth analytical tools and support for the field of breast cancer research.
[0004] This invention provides a breast cancer neoadjuvant information analysis system based on multimodal fusion. The system includes a data acquisition module, a data preprocessing module, a gene feature extraction module, an image feature extraction module, a multimodal fusion module, and a decision support module.
[0005] The data acquisition module acquires MRI, CT, and X-ray images as radiological imaging data, as well as genetic and clinical data.
[0006] The data preprocessing module preprocesses radiological imaging data, genetic data, and clinical data to generate preprocessed radiological imaging data, preprocessed genetic data, and preprocessed clinical data.
[0007] The gene feature extraction module extracts features from preprocessed gene data and generates gene feature data.
[0008] The image feature extraction module constructs a multi-scale variable convolutional pyramid residual model; the preprocessed radiographic image data is processed through the multi-scale variable convolutional pyramid residual model to generate multi-scale pyramid residual feature maps.
[0009] The multimodal fusion module constructs a cross-attention fusion model, initializes the parameters of the cross-attention fusion model, and inputs multi-scale pyramid residual feature maps, gene feature data, and preprocessed clinical data into the cross-attention fusion model to generate optimized breast cancer analysis results.
[0010] The decision support module performs disease risk assessment, prognosis prediction, personalized monitoring plan recommendations, disease progression trend analysis, patient grouping, and risk classification based on optimized breast cancer analysis results.
[0011] The multi-scale variable convolutional pyramid residual model includes 3D deep variable convolutional layers, 2D convolutional layers, and pyramid residual modules.
[0012] Furthermore, the image feature extraction module, in the process of generating multi-scale pyramid residual feature maps, specifically includes the following steps:
[0013] Step D1: Perform joint feature extraction on the preprocessed radiographic image data through a 3D deep variable convolutional layer to generate a multi-scale 3D convolutional feature map;
[0014] Step D2: Further extract detailed spatial features of the multi-scale 3D convolutional feature map through 2D convolutional layers to generate 2D convolutional feature maps;
[0015] Step D3: Increase the number of channels in the feature map layer by layer through the pyramid residual module, further extract high-level features from the 2D convolutional feature map, and generate multi-scale pyramid residual feature maps.
[0016] Furthermore, step D1 specifically includes the following steps:
[0017] Step D11: Select kernel size: Based on the complexity and characteristics of the preprocessed radiographic image data, adjust the kernel size of the 3D deep variable convolution to obtain a variable-scale convolution kernel;
[0018] Step D12: Convolution operation: Perform 3D convolution on the preprocessed radiographic image data using a variable-scale convolution kernel; generate a multi-scale 3D convolution feature map.
[0019] Furthermore, the multimodal fusion module, in generating optimized breast cancer analysis results, specifically includes the following steps:
[0020] Step E1: Use the multi-scale pyramid residual feature map as the query and the gene feature data as the key and value to generate image-gene association data;
[0021] Step E2: Use the multi-scale pyramid residual feature map as the query and clinical data as the key and value to generate image-clinical related data;
[0022] Step E3: Use gene feature data as the query and clinical data as the key and value to generate gene-clinical association data;
[0023] Step E4: Concatenate the image-gene association data, image-clinical association data, and gene-clinical association data to generate fused feature data;
[0024] Step E5: Process and fuse feature data through three fully connected layers to generate breast cancer analysis results;
[0025] Step E6: Optimize the parameters of the cross-attention fusion model using the spatiotemporal consistency-negative log loss function to generate optimized breast cancer analysis results.
[0026] Furthermore, step E6 specifically includes the following steps:
[0027] Step E61: Introduce the spatiotemporal consistency and negative log-likelihood loss functions, set adjustment coefficients, balance the weights between the loss terms of the spatiotemporal consistency and negative log-likelihood loss functions, and construct the spatiotemporal consistency-negative log loss function;
[0028] Step E62: Train the cross-attention fusion model using the spatiotemporal consistency-negative log loss function, optimize the parameters of the cross-attention fusion model through backpropagation, adjust the prediction output of the cross-attention fusion model using spatiotemporal consistency constraints, and generate optimized breast cancer analysis results.
[0029] By adopting the above solution, the beneficial effects achieved by the present invention are as follows:
[0030] This invention provides a new auxiliary information analysis system for breast cancer based on multimodal fusion, achieving efficient fusion of multimodal data and solving the problems of inefficiency and information loss in the data integration process of traditional breast cancer auxiliary information analysis systems. First, the feature extraction module combines 3D variable-scale convolution, 2D convolution, and pyramid residual modules to perform multi-scale processing on the radiographic imaging data of breast cancer patients. The system can capture the spatial distribution and detailed features of tumor regions, providing accurate information for subsequent analysis and improving the expression ability of tumor features. This processing can reveal the gene expression profile and gene mutation information of breast cancer patients, providing crucial basic data for subsequent multimodal fusion and analysis.
[0031] Furthermore, the multimodal fusion module based on a triple cross-attention mechanism, combined with the spatiotemporal consistency-negative logarithmic loss function, can deeply fuse multiple data sources, thereby enhancing the correlation between different data sources and maximizing the role of each data type in the analysis process. Through this refined fusion method, the system can eliminate redundant information between various types of data, reduce information interference, and improve the accuracy of the fused data. This technology significantly enhances the predictive ability of breast cancer analysis results, provides more efficient and reliable support for breast cancer research and risk assessment, and promotes the further development of the field of breast cancer auxiliary information analysis.
[0032] This invention significantly improves the performance and accuracy of breast cancer analysis systems by combining the aforementioned multimodal data fusion technology with advanced feature extraction methods and efficient data processing mechanisms. Through the effective fusion of radiological imaging data, genetic data, and clinical data, this invention not only improves the comprehensive utilization rate of various types of data but also ensures the synergistic effect of information across different data sources. This provides more comprehensive and accurate support for breast cancer-related research and data analysis, and promotes the development of breast cancer analysis systems towards greater efficiency and intelligence. Attached Figure Description
[0033] Figure 1 This is a schematic diagram of a breast cancer neoadjuvant information analysis system based on multimodal fusion proposed in this invention.
[0034] Figure 2 This is a diagram of the model structure of the multi-scale variable convolution pyramid residual model proposed in Example 2;
[0035] Figure 3 This is a model structure diagram of the cross-attention fusion model proposed in Example 5. Detailed Implementation
[0036] The technical solutions of the present invention will be clearly and completely described below with reference to the accompanying drawings of the embodiments of the present invention. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. All other embodiments obtained by those skilled in the art based on the embodiments of the present invention without creative effort are within the scope of protection of the present invention.
[0037] Example 1, according to Figure 1 This invention provides a new auxiliary information analysis system for breast cancer based on multimodal fusion. The system includes a data acquisition module, a data preprocessing module, a gene feature extraction module, an image feature extraction module, a multimodal fusion module, and a decision support module.
[0038] The data acquisition module collects MRI, CT, and X-ray images as radiological imaging data, as well as genetic and clinical data. Genetic data includes RNA sequencing data, gene mutation information, and gene expression profiles. Clinical data includes basic patient information, clinical characteristics of the tumor, and the patient's treatment background.
[0039] The data preprocessing module performs slice annotation, region selection, image cropping, image segmentation, and standardization on radiological image data to generate preprocessed radiological image data; it also performs standardization and data cleaning on genetic and clinical data to generate preprocessed genetic and clinical data.
[0040] The gene feature extraction module extracts features from preprocessed gene data and generates gene feature data.
[0041] The image feature extraction module combines 3D variable-scale convolution, 2D convolution, and pyramid residual modules to construct a multi-scale variable-scale convolution pyramid residual model; it processes preprocessed radiographic image data through the multi-scale variable-scale convolution pyramid residual model to generate multi-scale pyramid residual feature maps; it extracts features from preprocessed gene data to generate gene feature data; the multi-scale variable-scale convolution pyramid residual model includes 3D deep variable-scale convolution layers, 2D convolution layers, and pyramid residual modules;
[0042] The multimodal fusion module combines a triple cross-attention mechanism and a spatiotemporal consistency-negative log loss function to construct a cross-attention fusion model. It initializes the parameters of the cross-attention fusion model and inputs multi-scale pyramid residual feature maps, gene feature data, and preprocessed clinical data into the cross-attention fusion model to generate optimized breast cancer analysis results.
[0043] The decision support module performs disease risk assessment, prognosis prediction, personalized monitoring plan recommendations, disease progression trend analysis, patient grouping, and risk classification based on optimized breast cancer analysis results.
[0044] Example 2, according to Figure 2 This embodiment is based on Embodiment 1. In this embodiment, the image feature extraction module processes the preprocessed radiographic image data through a multi-scale variable convolutional pyramid residual model to generate a multi-scale pyramid residual feature map. The specific steps include:
[0045] Step D1: Extract joint features of spatial and spectral information from the preprocessed radiographic image data using a 3D deep variable convolutional layer to generate a multi-scale 3D convolutional feature map;
[0046] Step D2: Further extract detailed spatial features of multi-scale 3D convolutional feature maps through 2D convolutional layers, and reduce computational complexity through pooling layers to generate 2D convolutional feature maps;
[0047] Step D3: The number of channels in the feature map is increased layer by layer through the pyramid residual module to further extract high-level features from the 2D convolutional feature map. The residual blocks in each pyramid residual module optimize feature transfer through cross-layer connections and summation operations, reduce information loss, and enhance the expressive power of features to generate multi-scale pyramid residual feature maps.
[0048] Example 3, based on Example 1, describes the process of generating a multi-scale pyramid residual feature map using the feature extraction module, which includes the following steps:
[0049] Step R1: Extract joint features of spatial and spectral information from the preprocessed radiographic image data using 3D convolution to generate a 3D convolutional feature map;
[0050] Step R2: Further extract detailed spatial features of the 3D convolutional feature map through 2D convolutional layers, and reduce computational complexity through pooling layers to generate 2D convolutional feature maps;
[0051] Step R3: In the pyramid residual module, by increasing the number of channels of the feature map layer by layer, high-level features of the 2D convolutional feature map are further extracted. The residual blocks in each pyramid residual module optimize feature transfer through cross-layer connection and summation operations, reduce information loss, and enhance the expressive power of features, generating multi-scale pyramid residual feature maps.
[0052] Example 4, this example is based on Example 2. In this example, step D1 specifically includes the following steps:
[0053] Step D11: Select kernel size: Based on the complexity and characteristics of the preprocessed radiographic image data, adjust the kernel size of the 3D deep variable convolution to obtain a variable-scale convolution kernel;
[0054] Step D12: Convolution Operation: Perform 3D convolution on the preprocessed radiographic image data using a variable-scale convolution kernel; the variable-scale convolution kernel slides across the width, height, and depth dimensions of the preprocessed radiographic image data to calculate the features of each local region, generating a multi-scale 3D convolutional feature map. The formula used is as follows:
[0055] Variable-scale 3D convolution formula:
[0056] ;
[0057] in, This indicates the position of pixels in three-dimensional space within the preprocessed radiographic image data. This represents the index of the variable-scale convolutional kernel in the three dimensions of width, length, and depth. Indicates the convolutional layer index. Indicates that the variable-scale convolution kernel is in The size of the layer refers to the number of elements covered by the variable-scale convolution kernel in terms of width, length, and depth; This represents the pixel value after the convolution operation, i.e., the feature value at position (x, y, z); The weights represent the weights of the variable-scale convolution kernel. This represents the pixel value of the preprocessed radiographic image data at the current location.
[0058] Example 5, according to Figure 3 This embodiment is based on Embodiment 4. In this embodiment, the process of generating optimized breast cancer analysis results by the multimodal fusion module specifically includes the following steps:
[0059] Step E1: Using the multi-scale pyramid residual feature map as the query and gene feature data as the key and value, image-gene association data is generated through cross-attention calculation.
[0060] Step E2: Using the multi-scale pyramid residual feature map as the query and clinical data as the key and value, image-clinical association data is generated through cross-attention calculation;
[0061] Step E3: Using gene feature data as the query and clinical data as the key and value, gene-clinical association data is generated through cross-attention calculation;
[0062] Step E4: Concatenate the image-gene association data, image-clinical association data, and gene-clinical association data to generate fused feature data;
[0063] Step E5: Process and fuse feature data through three fully connected layers to generate breast cancer analysis results;
[0064] Step E6: Train the cross-attention fusion model using the spatiotemporal consistency-negative log loss function to improve the accuracy of the prediction results of the cross-attention fusion model and generate optimized breast cancer analysis results.
[0065] Example 6, based on Example 4, describes the process by which the multimodal fusion module generates optimized breast cancer analysis results, specifically including the following steps:
[0066] Step Q1: Using the multi-scale pyramid residual feature map as the query and gene feature data as the key and value, image-gene association data is generated through cross-attention calculation.
[0067] Step Q2: Using the multi-scale pyramid residual feature map as the query and clinical data as the key and value, image-clinical association data is generated through cross-attention calculation;
[0068] Step Q3: Concatenate the image-gene association data and the image-clinical association data to generate fused feature data;
[0069] Step Q4: Process and fuse feature data through three fully connected layers to generate breast cancer analysis results;
[0070] Step Q5: Train the cross-attention fusion model using the negative log-likelihood loss function to improve the accuracy of the prediction results of the cross-attention fusion model and generate optimized breast cancer analysis results.
[0071] Example 7, this example is based on Example 5. In this example, step E6 specifically includes the following steps:
[0072] Step E61: Introduce spatiotemporal consistency and negative log-likelihood loss function, set adjustment coefficients, balance the weights between the loss terms of spatiotemporal consistency and negative log-likelihood loss function, and construct the spatiotemporal consistency-negative log loss function. The formula used is as follows:
[0073] ;
[0074] in, This represents the spatiotemporal consistency loss value. and This represents the adjustment coefficient. Indicates a location index in time or space. Indicates the current sample index. Indicates an index along the time dimension. Indicates an index in a spatial dimension; Indicates the first Each sample at time point The risk value, Indicates the first Each sample at time point The risk value, Indicates the first The spatial location of each sample The risk value, Indicates the first The spatial location of each sample The risk value; Indicates time consistency items, Indicates spatial consistency terms;
[0075] ;
[0076] in, This represents the loss value of the spatiotemporal consistency-negative logarithmic loss function. Indicates the censored sample index. Indicates the first Risk value of a sample Indicates the first Risk value of a sample Indicates the first The right-censored set of samples; Indicates the first The natural logarithm of the risk values of a sample Indicates the first Risk value of a sample Perform exponentiation; Indicates the first Event indicator for each sample;
[0077] Step E62: Train the cross-attention fusion model using the spatiotemporal consistency-negative log loss function, optimize the parameters of the cross-attention fusion model through backpropagation, adjust the prediction output of the cross-attention fusion model using spatiotemporal consistency constraints, and generate optimized breast cancer analysis results.
[0078] The present invention and its embodiments have been described above. This description is not restrictive. The accompanying drawings are only one embodiment of the present invention, and the actual structure is not limited thereto. In short, if a person skilled in the art is inspired by this description and designs a similar structure and embodiment without departing from the spirit of the present invention, such design should fall within the protection scope of the present invention.
Claims
1. A neoadjuvant information analysis system for breast cancer based on multimodal fusion, comprising a data preprocessing module and a gene feature extraction module, wherein the data preprocessing module generates preprocessed radiological imaging data and preprocessed clinical data; and the gene feature extraction module generates gene feature data; characterized in that: The system also includes an image feature extraction module and a multimodal fusion module; The image feature extraction module constructs a multi-scale variable convolutional pyramid residual model; it processes the preprocessed radiographic image data through the multi-scale variable convolutional pyramid residual model to generate a multi-scale pyramid residual feature map. The multimodal fusion module constructs a cross-attention fusion model, initializes the parameters of the cross-attention fusion model, and inputs multi-scale pyramid residual feature maps, gene feature data, and preprocessed clinical data into the cross-attention fusion model to generate optimized breast cancer analysis results. The multimodal fusion module generates optimized breast cancer analysis results, specifically including the following steps: Step E1: Use the multi-scale pyramid residual feature map as the query and the gene feature data as the key and value to generate image-gene association data; Step E2: Use the multi-scale pyramid residual feature map as the query and clinical data as the key and value to generate image-clinical related data; Step E3: Use gene feature data as the query and clinical data as the key and value to generate gene-clinical association data; Step E4: Concatenate the image-gene association data, image-clinical association data, and gene-clinical association data to generate fused feature data; Step E5: Process the fused feature data to generate breast cancer analysis results; Step E6: Optimize the parameters of the cross-attention fusion model using the spatiotemporal consistency-negative log loss function to generate optimized breast cancer analysis results. The formula used is as follows: ; in, This represents the spatiotemporal consistency loss value. and This represents the adjustment coefficient. Indicates a location index in time or space. Indicates the current sample index. Indicates an index along the time dimension. Indicates an index in a spatial dimension; Indicates the first Each sample at time point The risk value, Indicates the first Each sample at time point The risk value, Indicates the first The spatial location of each sample The risk value, Indicates the first The spatial location of each sample The risk value; Indicates time consistency items, Indicates spatial consistency terms; ; in, This represents the loss value of the spatiotemporal consistency-negative logarithmic loss function. This indicates the censored sample index. Indicates the first Risk value of a sample Indicates the first Risk value of a sample Indicates the first The right-censored set of samples; Indicates the first The natural logarithm of the risk values of a sample Represents the risk value for the i-th sample. Perform exponentiation; Indicates the first Event indicator for each sample.
2. The breast cancer neoadjuvant information analysis system based on multimodal fusion according to claim 1, characterized in that: The multi-scale variable convolution pyramid residual model includes 3D deep variable convolution layers and pyramid residual modules.
3. The breast cancer neoadjuvant information analysis system based on multimodal fusion according to claim 2, characterized in that: The feature extraction module generates a multi-scale pyramid residual feature map, which specifically includes the following steps: Step D1: Perform joint feature extraction on the preprocessed radiographic image data through a 3D deep variable convolutional layer to generate a multi-scale 3D convolutional feature map; Step D2: Extract features from the multi-scale 3D convolutional feature map and generate a 2D convolutional feature map; Step D3: Extract high-level features from the 2D convolutional feature map using the pyramid residual module to generate a multi-scale pyramid residual feature map.
4. The breast cancer neoadjuvant information analysis system based on multimodal fusion according to claim 3, characterized in that: Step D1 specifically includes the following steps: Step D11: Based on the preprocessed radiographic image data, adjust the kernel size of the 3D deep variable convolutional layer to obtain a variable-scale convolutional kernel; Step D12: Perform 3D convolution on the preprocessed radiographic image data using variable-scale convolution kernels to generate multi-scale 3D convolution feature maps.
Citation Information
Patent Citations
Expression recognition method and system based on multi-feature fusion and three-cross attention mechanism
CN117152812A
Multi-modal breast tumor risk prediction method and system based on pathological image
CN118675618A