System and method for predicting risk of post-neoadjuvant therapy recurrence for triple-negative breast cancer based on multi-modal time-series medical imaging data

The prediction system based on multimodal temporal imaging data solves the problem of accurately assessing the risk of postoperative recurrence in TNBC patients, realizes dynamic recurrence risk monitoring for TNBC patients, improves the accuracy and specificity of prediction, and has end-to-end automation and robustness.

CN121617636BActive Publication Date: 2026-04-14THE FIRST AFFILIATED HOSPITAL OF WENZHOU MEDICAL UNIV
View PDF 2 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2026-02-02
Publication Date
2026-04-14

AI Technical Summary

Technical Problem

Existing methods for predicting the risk of breast cancer recurrence cannot accurately identify the risk of postoperative recurrence in patients with triple-negative breast cancer (TNBC). Traditional single imaging assessments and static assessments cannot fully reflect tumor heterogeneity and dynamic changes, and existing models lack the utilization of multimodal time-series data.

Method used

A prediction system based on multimodal temporal medical imaging data is constructed. Through the construction of multimodal datasets, adaptive lesion segmentation, deep learning feature extraction, radiomics feature extraction, tumor habitat feature extraction and feature fusion, combined with gene expression information, the Transformer architecture is used to process temporal data and construct a temporal prediction model to achieve dynamic recurrence risk assessment for TNBC patients.

Benefits of technology

It enables precise and dynamic monitoring of postoperative recurrence risk in TNBC patients, breaking through the diagnostic bottleneck of single imaging modality, improving the accuracy and specificity of recurrence risk prediction, and possessing end-to-end automation and robustness, thus having clinical application value.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121617636B_ABST
    Figure CN121617636B_ABST
Patent Text Reader

Abstract

The application discloses a triple negative breast cancer post-neoadjuvant therapy postoperative recurrence risk prediction system and method based on multi-modal time sequence medical image data, and belongs to the medical image analysis field.The system comprises: a data processing module for constructing a multi-modal data set; a multi-modal feature extraction and screening module for extracting deep learning, imageomics and tumor habitat features from the region, and performing feature screening; a model training module for constructing a time sequence model based on a Transformer architecture, and training through a multi-task learning strategy of integrating time consistency constraints and gene association auxiliary loss; and a recurrence risk prediction module for loading the trained model, and outputting a recurrence probability and a risk level.The application realizes dynamic and accurate quantification of the recurrence risk of triple negative breast cancer patients by fusing multi-modal time sequence images and gene information, and provides support for clinical individualized treatment decisions.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of postoperative prediction technology for malignant tumors, specifically to a system and method for predicting the risk of postoperative recurrence in triple-negative breast cancer after neoadjuvant therapy based on multimodal temporal medical imaging data. Background Technology

[0002] Breast cancer is the most common malignant tumor among women worldwide and the leading cause of cancer-related death. Although multidisciplinary treatment has improved the five-year survival rate to some extent, breast cancer recurrence and metastasis remain major challenges in clinical treatment, and patients with recurrence generally have a poor prognosis. Therefore, developing effective tools for predicting the risk of breast cancer recurrence is of great significance for guiding treatment and improving patient survival.

[0003] Triple-negative breast cancer (TNBC) is a subtype of breast cancer that lacks expression of estrogen receptor, progesterone receptor, and human epidermal growth factor receptor 2 (HER2). TNBC is difficult to target with therapy due to the lack of hormone receptors and HER2 expression; it is highly aggressive and heterogeneous, primarily relying on chemotherapy, and has a poorer prognosis compared to other subtypes. According to domestic and international treatment guidelines, neoadjuvant therapy (NAT) is recommended for TNBC patients with a certain tumor burden before surgery, followed by stratification based on whether NAT achieves pathological complete response (pCR) to select subsequent treatment options. However, current research evidence suggests that even without postoperative systemic therapy, approximately 50% of NAT patients with residual lesions do not experience recurrence. Furthermore, a meta-analysis also confirmed that pCR cannot serve as a surrogate endpoint for disease-free survival and overall survival in early breast cancer neoadjuvant trials. This suggests that relying solely on pCR or residual lesion assessment as the basis for postoperative adjuvant therapy decisions for TNBC cannot accurately stratify the risk of postoperative recurrence in TNBC patients, potentially leading to both undertreatment and overtreatment risks. In clinical practice, the risk of postoperative recurrence in breast cancer patients is typically assessed based on TNM staging combined with key clinicopathological markers, according to domestic and international guidelines. However, the accuracy of subjective evaluations based on pathological markers is currently not high, and their predictive efficacy is limited. Existing assessment methods based on clinicopathological markers are still insufficient to accurately identify the risk of postoperative recurrence in TNBC patients.

[0004] There are two main reasons for this inadequacy. First, TNBC is a highly heterogeneous tumor, and relying solely on a single clinicopathological marker or a single modality of imaging data is insufficient to fully quantify the tumor's biological characteristics. Second, pathological tissue obtained from a single or partial examination of breast cancer cannot represent the complete characteristics and changes of the tumor tissue, failing to comprehensively reflect tumor heterogeneity. Third, current imaging assessments are often limited to single examination methods, making it difficult to comprehensively utilize the complementary information carried by multimodal imaging such as breast ultrasound, mammography, digital mammography, and magnetic resonance imaging, neglecting the refined characterization of the tumor habitat and peritumoral microenvironment.

[0005] Secondly, static assessments are insufficient to accurately reflect a patient's long-term recurrence risk. The rate of metastasis in early-stage TNBC patients is as high as 25%-30% within 3-5 years of diagnosis. Current static assessments based on single-time-point data can only capture recurrence risk information at specific time points, failing to accurately assess long-term recurrence risk. The response of TNBC to neoadjuvant therapy is a dynamic evolutionary process. Existing prediction methods lack the ability to capture the dependence on time-series data throughout the entire treatment cycle (pre-treatment, during treatment, pre-operative, and follow-up), making it difficult to establish a dynamic model of recurrence risk evolution.

[0006] Although some research has been conducted on the prediction of breast cancer recurrence risk, existing models are limited by insufficient utilization of multimodal information and lack of temporal dynamic monitoring, and have not yet been widely used in clinical practice. Clinically, there is still a lack of effective methods to integrate multimodal temporal imaging and genetic information to comprehensively quantify TNBC tumor characteristics and accurately assess patients' long-term recurrence risk and survival prognosis. Summary of the Invention

[0007] To address the shortcomings of existing technologies, the present invention aims to provide a system and method for predicting the risk of postoperative recurrence in triple-negative breast cancer after neoadjuvant therapy based on multimodal temporal medical imaging data.

[0008] To achieve the above objectives, the present invention provides the following technical solution:

[0009] A predictive system for the risk of recurrence after neoadjuvant therapy in triple-negative breast cancer based on multimodal temporal imaging data, comprising:

[0010] The data processing module is used to construct multimodal medical datasets and perform structured preprocessing. This module specifically includes:

[0011] The multimodal dataset construction module is configured to collect multimodal imaging data, clinicopathological data, and gene expression information from N TNBC patients, and to simultaneously integrate and label recurrence. The multimodal imaging data covers four types of images: breast ultrasound (US), mammography (MG), digital mammography (DBT), and multiparameter magnetic resonance imaging (MRI).

[0012] The image preprocessing module is configured to perform standardization processing on the multimodal image data, the processing including at least one of image denoising, grayscale normalization, and histogram equalization; and is configured to perform data augmentation operations on the image data, the data augmentation including at least one of Gaussian blur, random rotation, vertical mirroring, and horizontal mirroring.

[0013] The adaptive lesion segmentation module is configured to automatically segment lesion regions in preprocessed image data using a hierarchical strategy based on the image modality dimension: for 2D modality US and MG images, a first segmentation model finely tuned based on the MedSAM 2D pre-trained model is loaded to perform segmentation; for 3D modality DBT and MRI images, a second segmentation model finely tuned based on the SAM-Med3D pre-trained model is loaded to perform segmentation.

[0014] The region of interest generation module is configured to receive the binary lesion mask output by the adaptive lesion segmentation module, perform morphological dilation operation on the lesion contour using a sphere as the structural element, thereby generating a peritumoral region with a preset width, and establish the lesion region and the peritumoral region together as a structured input for subsequent analysis.

[0015] A multimodal feature extraction and screening module is used to extract multidimensional features and screen for gene drivers. This module specifically includes:

[0016] The deep learning feature extraction module is used to input the region of interest from breast US, MG, DBT and MRI images of TNBC patients into a pre-trained deep learning large model, and automatically extract deep learning features that characterize the morphology, texture and edge of the lesion through the deep learning large model;

[0017] The radiomics feature extraction module is used to extract radiomics features from the regions of interest of the breast US, MG, DBT and MRI images using radiomics tools. The radiomics features include first-order histogram features, shape features, texture features and filter-transform-based features.

[0018] The tumor habitat feature extraction module obtains a tumor ecological diversity index feature set for constructing a recurrence risk quantification model based on the following steps:

[0019] C1) Perform supervoxel pre-segmentation on the tumor region within the region of interest to divide the tumor into multiple subregions with similar imaging features;

[0020] C2) Unsupervised clustering is performed based on the image features of the sub-regions, and the number of clusters K is optimized to quantify the ecological diversity within the tumor.

[0021] C3) Generate a tumor ecological diversity index feature set based on the optimal clustering results;

[0022] The feature fusion module is used to integrate the features extracted by the deep learning feature extraction module, the radiomics feature extraction module and the tumor habitat feature extraction module to construct a multi-dimensional, multimodal image feature matrix.

[0023] The feature selection module, based on gene expression data, performs feature selection on the fused multimodal image feature matrix and enhances feature representation capabilities through multi-source information fusion technology.

[0024] A model training module is used to build and train a time series prediction model. This module includes:

[0025] The temporal feature encoding unit is configured to receive a complete multimodal feature sequence and construct a temporal model based on the Transformer architecture. This unit uses a self-attention mechanism to capture the dependencies between different time points before neoadjuvant treatment, during treatment, the last time before treatment, and during follow-up. At the same time, it dynamically adjusts the attention weights in conjunction with the binary mask matrix transmitted by the interpolation unit to distinguish the influence of real observations and interpolated values ​​on temporal dependencies, and finally outputs a high-dimensional hidden layer feature representation that integrates spatiotemporal information.

[0026] The gene-driven and multi-task optimization unit is configured as a joint supervision module for the aforementioned units, employing a multi-task learning strategy to constrain the model as a whole. Backpropagation is used to collaboratively update the model parameters of the imputation and encoding units by calculating a weighted total loss function comprising the following three parts:

[0027] C1) Recurrence Prediction Main Loss: Based on the feature representation output by the coding unit, the binary cross-entropy loss function is used to measure the difference between the predicted risk and the true recurrence label;

[0028] C2) Time Consistency Constraint Loss: Constrains the coding unit to maintain the consistency of early prediction results with later prediction results in terms of trend when processing inputs at different time points;

[0029] C3) Gene association-assisted loss: A gene expression prediction branch is introduced into the feature extraction layer. The classification loss is calculated using gene expression data, which forces the coding unit to screen out image features that have a biologically significant association with relapse-related genes.

[0030] A relapse risk prediction module is used to perform the final prediction. This module mainly includes:

[0031] The risk probability calculation module is configured to load the optimized parameters of the model training module, input the complete feature sequence into the temporal feature encoding unit to extract high-dimensional hidden layer features, and input the features into the fully connected classification layer.

[0032] The quantitative scoring output module is configured to use the Sigmoid activation function to map the output of the fully connected layer to a recurrence probability value between 0 and 1; based on the preset clinical risk threshold, it generates a binary recurrence risk label and a visualized risk change trend curve as a basis for assisting clinical decision-making.

[0033] This invention also discloses a method for predicting the risk of recurrence after neoadjuvant therapy for triple-negative breast cancer based on multimodal temporal imaging data, comprising the following steps:

[0034] S1: Multimodal Data Construction and Structured Preprocessing: Four types of imaging data (breast US, MG, DBT, and MRI), clinicopathological data, and gene expression information were collected from N TNBC patients. The imaging data underwent denoising, normalization, and data augmentation. Subsequently, a hierarchical strategy was used for lesion segmentation based on the image modality: for 2D images (US, MG), a finely tuned MedSAM 2D model was used for segmentation; for 3D images (DBT, MRI), a finely tuned SAM-Med 3D model was used for segmentation. Based on this, morphological expansion of the lesion contour was performed using spheres as structural elements to generate a Region of Interest (ROI) containing both the lesion area and the peritumoral region.

[0035] S2: Multidimensional Feature Extraction and Gene-Driven Screening. For the region of interest, the following feature extraction operations are performed in parallel: deep learning features are extracted using a pre-trained large model; radiomics features, including shape, texture, and filtering transformations, are extracted using radiomics tools; tumor internal ecological diversity is quantified and tumor habitat features are generated through supervoxel pre-segmentation and unsupervised clustering techniques; after fusing the above features, the feature matrix is ​​screened and dimensionality reduced using the patient's gene expression data as prior knowledge, retaining image features that are highly correlated with gene expression, and constructing a multimodal temporal feature sequence.

[0036] S3: Construct and train a temporal risk prediction model. Construct a temporal model based on the Transformer architecture, inputting the feature sequence obtained in step S2 into the model; utilize the self-attention mechanism to capture the temporal dependencies of different stages of neoadjuvant therapy (pre-treatment, during treatment, pre-operative, and follow-up); train the model using a multi-task learning strategy, and collaboratively update the model parameters by calculating a weighted total loss function, which includes: the main loss for relapse prediction based on binary cross-entropy, the constraint loss for constraining temporal trend consistency, and the auxiliary loss based on gene expression classification; iteratively optimize through backpropagation until the model converges;

[0037] S4: Dynamic prediction and quantification of recurrence risk. The multimodal time-series image data of the patient to be predicted is processed according to steps S1 to S2 to generate corresponding feature sequences, which are then input into the trained time-series risk prediction model. The high-dimensional hidden layer features output by the model are processed by a fully connected layer and a sigmoid activation function to output a recurrence probability value between 0 and 1. Finally, based on a preset threshold, high-risk or low-risk binary labels are generated, and a risk trend curve over time is output to complete the prediction of recurrence risk.

[0038] The beneficial effects of this invention are:

[0039] (1) This invention breaks through the diagnostic bottleneck of a single imaging modality. By organically integrating data from four modalities—US, MG, DBT, and MRI—it constructs a multidimensional feature matrix encompassing deep learning, radiomics, and the tumor biosphere. In particular, by introducing features of the peritumoral microenvironment, it achieves comprehensive quantification of the high heterogeneity of TNBC, significantly improving the accuracy and specificity of recurrence risk prediction.

[0040] (2) This invention overcomes the shortcomings of traditional static models in capturing the trajectory of disease evolution. This invention uses the Transformer architecture to process time-series data of "the whole treatment cycle - postoperative follow-up". By mining the deep dependencies between time points and introducing time consistency constraints, dynamic monitoring of the long-term recurrence risk of patients is realized, effectively avoiding the one-sidedness of assessment at a single time point.

[0041] (3) This invention innovatively introduces an adaptive segmentation strategy based on MedSAM fine-tuning, which effectively solves the problem of accurate segmentation of lesions in cross-modal imaging. This strategy not only ensures the reliability of feature extraction, but also realizes end-to-end automation from original image input to recurrence risk assessment output, which has strong robustness and clinical application value. Attached Figure Description

[0042] Figure 1 This is a technical roadmap for the present invention.

[0043] Figure 2 This is a flowchart of the training process for the breast lesion segmentation model based on a deep learning large model according to the present invention.

[0044] Figure 3 This is a schematic diagram of the breast tumor subregion segmentation process and segmentation method of the present invention.

[0045] Figure 4 This is a schematic diagram of the image feature screening and multi-source information fusion based on gene information driven by the present invention.

[0046] Figure 5 This is a schematic diagram of the multimodal image feature screening and recurrence risk prediction model construction method of the present invention.

[0047] Figure 6 This is a schematic diagram illustrating the construction of the multimodal time-series prediction model of the present invention. Detailed Implementation

[0048] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only a part of the embodiments of the present invention, and not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative effort are within the scope of protection of the present invention.

[0049] This invention discloses a prediction system for the risk of postoperative recurrence in triple-negative breast cancer (TNBC) after neoadjuvant therapy based on multimodal temporal medical imaging data. The system includes a data processing module, a multimodal feature extraction and screening module, a model training module, and a recurrence risk prediction module. The specific process is as follows: Figure 1 As shown.

[0050] 1. The data processing module primarily employs a mathematically defined processing flow to transform heterogeneous raw medical data into standardized, structured input. The specific implementation steps are as follows:

[0051] 1) A multimodal dataset construction module is configured to collect multimodal imaging data, clinical pathological data, and gene expression information from N TNBC patients, and to simultaneously integrate and label the data with recurrence tags. The multimodal imaging data covers four types of images: breast ultrasound (US), mammography (MG), digital mammography (DBT), and multiparameter magnetic resonance imaging (MRI).

[0052] Specifically, it involves constructing a multi-dimensional TNBC time-series dataset. Dataset definition: Let the dataset be... ,in Let be the total number of patients. Sample vector definition: For the _th ... The patient's feature vector Defined as:

[0053]

[0054] in: This indicates discrete time points before neoadjuvant therapy, during treatment, before surgery, and during follow-up. This represents the set of image modalities. Represents the image tensor. Represents clinicopathological feature vectors. Represents a gene expression feature vector. Tag definition: It is a binary relapse label, determined based on disease-free survival (DFS).

[0055] 2) Image preprocessing module, whose main function is to eliminate device heterogeneity through mathematical transformations. This embodiment uses the following method as an example: grayscale normalization. To eliminate dimensional differences between different image acquisition devices, grayscale normalization is performed on all images... Perform zero-mean normalization. For any pixel in the image. Its standardized pixel values The calculation formula is:

[0056]

[0057] Simultaneously define the enhancement function set During the training phase, the input image With probability Random application of transformation ,in This expands the spatial distribution of samples, thereby increasing the robustness of sample training.

[0058] In other embodiments, the image processing further includes at least one of image denoising and histogram equalization; and is configured to perform data augmentation operations on the image data, the data augmentation including at least one of Gaussian blur, random rotation, vertical mirroring and horizontal mirroring.

[0059] 3) Adaptive lesion segmentation module, configured to automatically segment lesion regions in preprocessed image data using a hierarchical strategy based on image modality dimension.

[0060] Specifically, it constructs a mapping function based on the adaptive lesion segmentation module of a large model. The preprocessed image is mapped to a binary mask, and a layered segmentation strategy is employed, such as... Figure 2 As shown:

[0061] 2D Mapping (for US / MG): Defining the Segmentation Model The input is a 2D slice. The output is a mask. .

[0062] 3D Mapping (for DBT / MRI): Defining Segmentation Model The input is 3D data. The output is a mask. .

[0063] The final output is a binary image. ,in Indicates a potential tumor. Indicates the background.

[0064] 4) Region of Interest Generation Module: Specifically, this module receives the binary mask image of the lesion output from the adaptive lesion segmentation module. It automatically defines the analysis region containing the peritumoral microenvironment based on morphological dilatation calculations, and generates the dilatation region. For any coordinate point in the image space This operation can be formally defined as:

[0065]

[0066] This formula represents the result image. For each pixel position, take the pixel value in its neighborhood and the sphere structure element. The maximum value after superposition at corresponding positions. Here, the structural element is defined. It is a sphere (3D) or a circle (2D), whose physical radius corresponds to = 6mm, the pixel value within the structural element is set to 1, as shown in the formula:

[0067]

[0068] The final Region of Interest (ROI) is defined as the region generated after morphological dilation. This ROI will be used as an input mask for subsequent feature extraction modules, ensuring that the feature extraction scope covers not only the tumor body but also its surrounding peritumoral microenvironment.

[0069] In this embodiment, the multimodal feature extraction and screening module uses a parallel processing architecture to extract high-dimensional features from the ROIs of images from different modalities and optimizes these features using genetic information. It mainly includes a deep learning feature extraction module, a radiomics feature extraction module, a tumor habitat feature extraction module, a feature fusion module, and a feature screening module. The specific implementation methods of each module are as follows:

[0070] 1) Deep learning feature extraction module: This module is configured to receive preprocessed multimodal image data (breast US, MG, DBT, and MRI image data) and extract high-dimensional abstract features using a deep learning network (ResNet-50) pre-trained on large-scale medical images. First, it extracts ROI images containing lesions and peritumoral regions. Input deep neural network encoder Extracting the output of the deep hidden layers of the network as deep learning feature vectors The formula is as follows:

[0071]

[0072] 2) The radiomics feature extraction module is responsible for performing high-throughput feature computation on the ROI to generate a radiomics feature set. This module uses radiomics algorithms to calculate a comprehensive feature vector from the ROI. This vector contains the following feature subset:

[0073]

[0074] in: Features of the first-order histogram (mean, variance, skewness, kurtosis);

[0075] Shape characteristics (volume, surface area, compactness);

[0076] Texture features (calculated based on GLCM, GLRLM, and GLSZM matrices);

[0077] These are filtering features based on Gaussian and wavelet transforms.

[0078] 3) Tumor habitat feature extraction module: This module is configured to quantify the spatial heterogeneity within the tumor and generate an ecological diversity index feature set. Its processing flow includes three steps: supervoxel pre-segmentation, unsupervised clustering, and diversity index calculation.

[0079] C1) Hypervoxel pre-segmentation uses a linear iterative clustering algorithm to divide the tumor region into N hypervoxel subregions, such as Figure 3 As shown. Define the first The pixel and the The combined distance of each cluster center as follows:

[0080]

[0081]

[0082]

[0083] in Indicates the intensity value of multimodal image. Represents the spatial coordinates of a pixel. The distance is the image intensity value. Spatial distance; The preset grid step size, A compactness parameter is used to balance intensity similarity and spatial proximity.

[0084] C2) Habitat subregion clustering optimization: Based on the image features of subregions, the K-means algorithm is used for clustering, and the optimal number of clusters is determined by the Calinski-Harabasz (CH) index. (i.e., the optimal number of habitats). The CH index corresponding to different k values ​​is calculated, and the k value that maximizes the CH index is selected as the optimal number of clusters. The formula for calculating the CH index is:

[0085]

[0086] in: This represents the current number of clusters. This represents the total number of samples in the sub-region. Represents the trace of a matrix; The cluster scatter matrix is ​​the matrix of scatter within the cluster. Let be the inter-cluster scatter matrix, defined as follows:

[0087]

[0088]

[0089] Here, Indicates the first The sample set of each cluster For the first The center of each cluster, The global center of all samples. For the first The number of samples in each cluster.

[0090] C3) Feature set generation selection makes The largest Value as Based on the optimal clustering result, a tumor ecological diversity index feature set is generated. .

[0091] The feature fusion module is configured to concatenate and regularize the above feature vectors, such as... Figure 4 As shown, a multimodal image feature matrix is ​​constructed. As shown in the formula:

[0092]

[0093] in The modality count is (US, MG, DBT, MRI). This represents the total dimension of the merged features.

[0094] The feature filtering module is configured to receive a multimodal image feature matrix constructed by the feature fusion module. And by constructing a multi-task deep neural network that includes gene prediction branches and relapse prediction branches, such as Figure 5 As shown, the network weights are optimized using supervisory signals from gene expression data, thereby achieving better control during the feature encoding stage. Targeted screening.

[0095] First, construct a shared feature encoder. Its parameters are defined as follows: The multimodal image feature matrix The input encoder maps the selected high-dimensional latent feature vectors. :

[0096]

[0097] In this process The generated data will be updated based on the gradient of the gene prediction task during subsequent backpropagation, forcing the generated data to be updated accordingly. Only image information co-related with gene expression and relapse risk is retained to achieve feature screening. The specific model training process is shown in the model training module.

[0098] This embodiment details the internal structure of the model training module, which includes a temporal feature encoding unit and a gene-driven and multi-task optimization unit. The two work together to achieve in-depth analysis of multimodal temporal data and overall optimization of model parameters, as shown below:

[0099] 1) The temporal feature encoding unit is configured to receive the processed multimodal feature sequence and the corresponding mask matrix, and utilize a self-attention network with a masking mechanism to capture temporal dependencies, outputting high-dimensional hidden layer features that fuse spatiotemporal information. First, let the input multimodal feature sequence be... ,in For the first Feature vectors at each time point. First, positional encoding is introduced. To preserve temporal information, an input embedding vector is generated. :

[0100]

[0101] Then the sequence is processed using a Transformer encoder. Construct a query matrix Key matrix Sum matrix Introducing a binary mask matrix (in Indicates the interpolated value. (representing the true value), through the mask bias term Dynamically adjust the attention weights to eliminate the interference of interpolation noise on temporal dependency modeling, as shown in the formula:

[0102]

[0103] in, For the bias function, when hour ,when hour This mechanism ensures that the model focuses only on real observation data when aggregating information. After multiple layers of encoding, the output is a high-dimensional hidden layer feature sequence containing global temporal dependencies. .

[0104] 2) Gene-driven and multi-task optimization unit. This unit is configured as a joint supervision module. It adopts a multi-task learning strategy to construct a composite loss function that includes recurrence prediction, temporal consistency constraints and gene association prediction. The coding unit parameters are updated collaboratively through backpropagation.

[0105] C1) Recurrence predicts main loss ( ) Features based on the output of the coding unit A recurrence prediction branch is constructed. This branch calculates the recurrence prediction through a fully connected layer. Risk of recurrence at a specific time point And the binary cross-entropy loss function is used to measure the prediction bias:

[0106]

[0107] in, For sample batch, This is the set of actual observation time points where the sample exists. This is a label for genuine relapses.

[0108] C2) Time consistency constraint loss ( To ensure the model's risk assessment of the same patient at different time points exhibits trend stability, the consistency of prediction results at adjacent time points is constrained:

[0109]

[0110] This loss term forces the model to maintain a smooth prediction trend during time series extrapolation by minimizing the Euclidean distance between the predicted probabilities of adjacent time steps, thus preventing misjudgments caused by data fluctuations at a single time point.

[0111] C3) Gene-associated auxiliary loss ( This part is the concrete implementation of the aforementioned "feature selection module" during the training phase. A gene prediction branch is introduced at the end of the encoding unit, using gene expression data as an auxiliary supervision signal. This branch receives hidden layer features. Output gene expression prediction probability The loss function is defined as:

[0112]

[0113] Logical association: by minimizing The backpropagation algorithm passes the gradient to the temporal feature encoding unit at the front end, forcing the encoder parameters to... A targeted update occurs. This process mathematically "filters" the input features—that is, only those features that reflect both genetic biological characteristics and... It can also predict the risk of recurrence (reduce) The image features will be preserved and enhanced.

[0114] Finally, a weighted total loss function is constructed. The above three parts are jointly optimized:

[0115]

[0116] in These are hyperparameter weights. The system calculates... The model parameters are updated to enable multimodal temporal feature screening and accurate prediction driven by genes.

[0117] The recurrence risk prediction module described in this embodiment mainly includes a risk probability calculation module and a quantitative scoring output module. By loading and optimizing model parameters, it achieves accurate quantification and visualization of the patient's recurrence risk, such as... Figure 6 As shown, the specific implementation process is as follows.

[0118] 1) The risk probability calculation module is configured to load the globally optimal parameter set after the model training module has been optimized. Then, perform forward inference computation. First, the multimodal feature sequence of the patient to be predicted is... The input is given to a time-series feature encoding unit with fixed parameters. The optimized encoder parameters are then utilized. Extract high-dimensional hidden layer feature vectors containing spatiotemporal dependencies. :

[0119]

[0120] Then the high-dimensional hidden layer features The input is fed into a fully connected classification layer. The optimized weight matrix is ​​then used. and bias vector Calculate the unnormalized log-odds ratio of the risk .

[0121]

[0122] in, The dimensions of the output layer match the dimensions of the hidden layer features and the number of output categories.

[0123] 2) The quantitative scoring output module is configured to perform probability mapping and decision grading on the log-odds of the risks, generating clinically readable quantitative indicators. First, the Sigmoid activation function is used. Will Mapping to the (0, 1) interval yields the specific recurrence probability value. :

[0124]

[0125] Then preset clinical risk thresholds .according to and Relationship generation of binary relapse risk labels :

[0126]

[0127] To generate a risk change trend curve, the module performs analysis on the same patient at different follow-up time points. Calculate the recurrence probability separately

[0128]

[0129] The system is based on A dynamic trend curve with time on the horizontal axis and recurrence probability on the vertical axis is plotted to visually display the risk evolution trajectory of patients during the treatment and follow-up period, assisting doctors in making dynamic diagnosis and treatment decisions.

[0130] The aim of this technical solution is to construct a multimodal, multi-temporal, and multi-omics model that can dynamically predict the risk of TNBC recurrence at different time points, providing a scientific basis for personalized treatment of breast cancer patients and adjusting treatment or follow-up plans based on the prediction results.

[0131] This invention addresses the clinical challenge of accurately assessing the risk of recurrence after neoadjuvant therapy for tumor-associated neonatal nephropathy (TNBC). Traditional prediction methods struggle to fully reflect tumor heterogeneity and dynamic changes, and existing AI models based on single modalities and single time points suffer from insufficient accuracy. This invention intelligently quantifies multimodal imaging data and, driven by genetic information, screens key imaging features for recurrence risk, improving model accuracy. Utilizing multi-temporal imaging and clinical data, it constructs a longitudinal prediction model to dynamically and comprehensively assess patients' postoperative recurrence risk. This assists clinicians in making precise, individualized treatment decisions and allows for timely adjustments to treatment plans based on prediction results during follow-up.

[0132] This invention also provides a recurrence prediction method based on this prediction system. Specifically, it proposes a multimodal breast lesion segmentation scheme based on a pre-trained large model. A segmentation model is constructed for each modality, and fine-tuned using the powerful feature extraction capabilities of the pre-trained large model. Based on this, an image processing dilation algorithm is used to generate the peritumoral region, obtaining multimodal deep learning, radiomics, and tumor habitat feature sets, thereby constructing a multimodal image feature matrix. Furthermore, guided by gene data, supervised screening of multimodal image features is performed, and gene expression data, multimodal image features, and clinical features are organically integrated into a multi-threaded deep learning network. Finally, a Transformer-based multimodal temporal recurrence risk quantification model is constructed to dynamically predict the risk of TNBC recurrence at different time points. A comprehensive evaluation and interpretability analysis of the model are then conducted.

[0133] The embodiments should not be regarded as limitations on the present invention, but any improvements made based on the spirit of the present invention should be within the protection scope of the present invention.

Claims

1. A predictive system for the risk of postoperative recurrence in triple-negative breast cancer after neoadjuvant therapy based on multimodal temporal medical imaging data, characterized in that: It includes: The data processing module is used to construct a multimodal medical dataset and to perform data preprocessing and lesion segmentation. The data processing module first collects four types of imaging data from N triple-negative breast cancer patients—including breast ultrasound, mammography, digital mammography, and multiparameter magnetic resonance imaging—and simultaneously integrates their clinical pathology data and gene expression information to form a unified multimodal medical dataset. It also labels each patient with a recurrence prediction tag. Based on this, and using a deep learning large model architecture, it automatically and accurately segments the lesion area and peritumoral tissue in the multimodal imaging data, which serves as the structured input for feature extraction and recurrence risk prediction models. The multimodal feature extraction and screening module is used to extract deep learning features, radiomics features, and tumor habitat features from segmented multimodal image data, and combine them with clinicopathology to construct a multimodal image feature matrix, and use gene expression data to screen features from the multimodal image feature matrix. The model training module is based on the multimodal image feature data processed by the multimodal feature extraction and screening module. It uses a deep learning framework to train and optimize the model. During the training process, the model training module uses stochastic gradient descent as the optimization algorithm. It iteratively updates the model parameters by minimizing the difference between the prediction result and the recurrence label to obtain a well-trained multimodal temporal recurrence risk prediction model. The relapse risk prediction module predicts relapse risk based on a constructed multimodal time-series relapse risk prediction model. The model training module includes: The temporal feature encoding unit is configured to receive the complete multimodal feature sequence output and construct a temporal model based on the Transformer architecture. It uses a self-attention mechanism to capture the dependencies between different time points before neoadjuvant treatment, during treatment, the last time before treatment, and during follow-up. At the same time, it dynamically adjusts the attention weights in conjunction with the binary mask matrix transmitted by the interpolation unit to distinguish the influence of real observations and interpolated values ​​on temporal dependencies, and finally outputs a high-dimensional hidden layer feature representation that integrates spatiotemporal information. The gene-driven and multi-task optimization unit is configured as a joint supervision module serving as the temporal feature encoding unit, employing a multi-task learning strategy to constrain the model as a whole. Backpropagation is used to collaboratively update the model parameters of the imputation and encoding units by calculating a weighted total loss function comprising the following three parts: D1) Recurrence Prediction Main Loss: Based on the feature representation output by the coding unit, the binary cross-entropy loss function is used to measure the difference between the predicted risk and the true recurrence label; D2) Time Consistency Constraint Loss: Constrains the coding unit to maintain the consistency of early prediction results with later prediction results in terms of trend when processing inputs at different time points; D3) Gene association-assisted loss: Introduce a gene expression prediction branch in the feature extraction layer, use gene expression data to calculate classification loss, and force coding units to screen out image features that have biologically significant associations with relapse-related genes.

2. The prediction system for postoperative recurrence risk of triple-negative breast cancer after neoadjuvant therapy based on multimodal temporal medical imaging data according to claim 1, characterized in that, The data processing module includes: The multimodal dataset construction module is configured to collect multimodal imaging data, clinicopathological data, and gene expression information from N triple-negative breast cancer patients, and to simultaneously integrate and label recurrence. The multimodal imaging data covers four types of images: breast ultrasound, mammography, digital mammography, and multiparameter magnetic resonance imaging. The image preprocessing module is configured to perform standardization processing on the multimodal image data, the standardization processing including at least one of image denoising, grayscale normalization and histogram equalization; and is configured to perform data augmentation operations on the image data, the data augmentation including at least one of Gaussian blur, random rotation, vertical mirroring and horizontal mirroring. The adaptive lesion segmentation module is configured to automatically segment lesion regions from preprocessed multimodal image data using a hierarchical strategy based on the image modality dimension: for 2D modal breast ultrasound and mammography images, a first segmentation model finely tuned based on the MedSAM 2D pre-trained model is loaded for segmentation; for 3D modal digital mammography and multiparameter magnetic resonance imaging images, a second segmentation model finely tuned based on the SAM-Med 3D pre-trained model is loaded for segmentation. The region of interest generation module is configured to receive the binary lesion mask output by the adaptive lesion segmentation module, perform morphological dilation operation on the lesion contour using a sphere as the structural element, thereby generating a peritumoral region with a preset width, and establish the lesion region and the peritumoral region together as a structured input for subsequent analysis.

3. The prediction system for postoperative recurrence risk of triple-negative breast cancer after neoadjuvant therapy based on multimodal temporal medical imaging data according to claim 1, characterized in that: The multimodal feature extraction and screening module includes: The deep learning feature extraction module is used to input the region of interest from breast ultrasound, mammography, digital mammography and multi-parameter magnetic resonance imaging images of triple-negative breast cancer patients into a pre-trained deep learning model. The deep learning model automatically extracts deep learning features that characterize the morphology, texture and edge of the lesion. The radiomics feature extraction module is used to extract radiomics features from the regions of interest of the breast ultrasound, mammography, digital mammography and multi-parameter magnetic resonance imaging images using radiomics tools. The radiomics features include first-order histogram features, shape features, texture features and filter-transform-based features. The tumor habitat feature extraction module obtains a tumor ecological diversity index feature set for constructing a recurrence risk quantification model based on the following steps: C1) Perform supervoxel pre-segmentation on the tumor region within the region of interest to divide the tumor into multiple subregions with similar imaging features; C2) Unsupervised clustering is performed based on the image features of the sub-regions, and the number of clusters K is optimized to quantify the ecological diversity within the tumor. C3) Generate a tumor ecological diversity index feature set based on the optimal clustering results; The feature fusion module is used to integrate the features extracted by the deep learning feature extraction module, the radiomics feature extraction module and the tumor habitat feature extraction module to construct a multi-dimensional, multimodal image feature matrix. The feature selection module, based on gene expression data, performs feature selection on the fused multimodal image feature matrix and enhances feature representation capabilities through multi-source information fusion technology.

4. The prediction system for postoperative recurrence risk of triple-negative breast cancer after neoadjuvant therapy based on multimodal temporal medical imaging data according to claim 1, characterized in that: The multimodal time-series recurrence risk prediction model includes: The risk probability calculation module is configured to load the optimized parameters of the model training module, input the complete feature sequence into the temporal feature encoding unit to extract high-dimensional hidden layer features, and input the features into the fully connected classification layer. The quantitative scoring output module is configured to use the Sigmoid activation function to map the output of the fully connected layer to a recurrence probability value between 0 and 1; based on the preset clinical risk threshold, it generates a binary recurrence risk label and a visualized risk change trend curve as a basis for assisting clinical decision-making, wherein the binary recurrence risk label is high risk or low risk.

5. A method for predicting recurrence risk based on the prediction system according to any one of claims 1 to 4, characterized in that: It includes the following steps: S1: Multimodal data construction and structured preprocessing: Four types of imaging data, clinicopathological data, and gene expression information were collected from N patients with triple-negative breast cancer. The imaging data underwent denoising, normalization, and data augmentation. Subsequently, a hierarchical strategy was used for lesion segmentation based on the image modality: for 2D images, a finely tuned MedSAM 2D model was used for segmentation; for 3D images, a finely tuned SAM-Med 3D model was used for segmentation. Based on this, morphological expansion of the lesion contour was performed using spheres as structural elements to generate a region of interest containing both the lesion area and the peritumoral area. S2: Multidimensional Feature Extraction and Gene-Driven Screening. For the region of interest, the following feature extraction operations are performed in parallel: deep learning features are extracted using a pre-trained large model; radiomics features, including shape, texture, and filtering transformations, are extracted using radiomics tools; tumor internal ecological diversity is quantified and tumor habitat features are generated through supervoxel pre-segmentation and unsupervised clustering techniques; after fusing the above features, the feature matrix is ​​screened and dimensionality reduced using the patient's gene expression data as prior knowledge, retaining image features highly correlated with gene expression, and constructing a multimodal temporal feature sequence. S3: Construct and train a temporal risk prediction model, build a temporal model based on the Transformer architecture, and input the multimodal temporal feature sequence obtained in step S2 into the model; use the self-attention mechanism to capture the temporal dependence of different stages of neoadjuvant therapy; A multi-task learning strategy is used to train the model, and the model parameters are updated collaboratively by calculating a weighted total loss function. The total loss function includes: a recurrence prediction main loss based on binary cross-entropy, a constraint loss for constraining temporal trend consistency, and an auxiliary loss based on gene expression classification. The model is iteratively optimized through backpropagation until it converges. S4: Dynamic prediction and quantification of recurrence risk. The multimodal time-series image data of the patient to be predicted is processed according to steps S1 to S2 to generate corresponding feature sequences, which are then input into the trained time-series risk prediction model. The high-dimensional hidden layer features output by the model are processed by a fully connected layer and a sigmoid activation function to output a recurrence probability value between 0 and 1. Finally, based on a preset threshold, a high-risk or low-risk binary label is generated, and a risk trend curve over time is output to complete the prediction of recurrence risk.

6. The recurrence risk prediction method according to claim 5, characterized in that: Step S1 includes the following specific sub-steps: S11: Collect imaging data, clinical pathology reports, and gene testing data of N triple-negative breast cancer patients throughout the neoadjuvant therapy cycle; remove samples with severe image artifacts or missing key clinical information, and perform denoising, grayscale normalization, and histogram equalization on the image data. S12: Assign a binary recurrence label to each patient based on their follow-up records and pathological results: Set a preset follow-up period, mark patients who have local recurrence or distant metastasis during this period as "high-risk", and mark patients who have not recurred as "low-risk", and use this label as the supervised ground truth for model training; S13: Based on the image modality dimension, a hierarchical strategy is adopted for lesion segmentation: For 2D modal images, a finely tuned MedSAM 2D model is loaded, 2D slices are input, and lesion masks are output; for 3D modal images, a finely tuned SAM-Med 3D model is loaded, 3D volume data is input, and voxel-level lesion masks are output. S14: Receive the lesion mask mentioned above, perform morphological dilation operation on the lesion outline using a sphere as the structural element, and generate a ring-shaped peritumoral region with a preset pixel width; merge the original lesion region with the peritumoral region, and crop out the region of interest containing only the tumor and its microenvironment, as the standardized input for subsequent feature extraction.

7. The recurrence risk prediction method according to claim 5, characterized in that: Step S2 specifically includes the following sub-steps: S21: Perform three types of feature extraction in parallel for the region of interest: 1) Deep features: Input the region of interest into the pre-trained model feature extractor to extract high-dimensional abstract features; 2) Radiomics features: First-order statistical, shape, texture, and wavelet transform features were extracted using the PyRadiomics toolkit; 3) Tumor habitat characteristics: The region of interest is oversegmented using supervoxels. Based on the gray mean, the tumor is divided into multiple subregions. K-means clustering is used to determine the optimal number of clusters K. Shannon entropy and Simpson index are calculated to quantify the ecological heterogeneity within the tumor. S22: Vectorize and concatenate all the features of all modalities and types extracted above to construct an initial high-dimensional multimodal image feature matrix; S23: Introduce patient gene expression data as prior knowledge and calculate the Pearson correlation coefficient between each dimension of image feature vector and the expression level of key genes related to recurrence of triple-negative breast cancer; S24: Set a correlation threshold to remove redundant image features that are weakly correlated with gene expression, and retain a subset of image features with significant biological associations; then use principal component analysis to further reduce dimensionality and generate the final multimodal temporal feature sequence for model input.

8. The recurrence risk prediction method according to claim 6, characterized in that: Step S3 specifically includes the following sub-steps: S31: Input the complete feature sequence and mask matrix obtained after steps S1 and S2 into the Transformer model; in the self-attention calculation layer, use the mask matrix to adjust the attention score map, reset the weights of the missing modal positions to the minimum value, so that the model focuses on the temporal dependencies of the real observations and outputs high-dimensional hidden layer features that fuse spatiotemporal information. S32: Calculate the weighted total loss function To collaboratively update network parameters: in: : Recurrence prediction loss, calculated by the binary cross-entropy between the predicted probability and the actual recurrence label described in step S12; Time consistency loss: calculates the difference in predicted risk values ​​at adjacent time points to constrain the smoothness of risk change trends; Gene-assisted loss utilizes branches in the feature extraction layer to predict gene expression categories, enhancing the biological interpretability of features. For hyperparameter weights.

9. The recurrence risk prediction method according to claim 5, characterized in that: Step S4 specifically includes the following sub-steps: S41: Input the complete feature sequence into the time series prediction model. After the features are extracted by the multi-layer Transformer encoder, they are mapped through the fully connected layer and the Sigmoid activation function is used to output a value between 0 and 1, which represents the probability of postoperative recurrence in the patient. S42: Based on the preset clinical high-sensitivity threshold, the probability value is converted into a binary diagnostic suggestion of "high risk" or "low risk"; at the same time, the risk change trend curve of the patient over the course of treatment is output to help doctors judge the efficacy of neoadjuvant therapy and formulate postoperative adjuvant therapy plans.

Citation Information

Patent Citations

  • Tumor immunotherapy curative effect prediction method and system based on multi-modal data

    CN121354874A

  • Method and electronic device for predicting patch-level gene expression from histology image by using artificial intelligence model

    US20240257910A1