Breast cancer pCR early prediction system based on limited medical text
By combining limited medical text and multi-source image features in the early prediction system of breast cancer pCR, the problems of limited text and single image characteristics of traditional Chinese medicine are solved in the prior art, which improves the prediction accuracy and enhances the robustness of the model.
Patent Information
- Application Number
- CN202510099265.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-22
- Publication Date
- 2025-05-23
AI Technical Summary
The prior art faces the problems of limited medical text and unconsidered single image characteristics in the early prediction of breast cancer pCR, resulting in low prediction accuracy.
A breast cancer pCR early prediction system based on limited medical text is adopted, which includes a pre-training module, a training module and an inference module. Through the large language model optimization parameters, combining the morphological and hemodynamic characteristics of DWI and DCE-MRI images, and fusing medical text features to generate breast cancer pCR prediction results.
It improves the accuracy of early prediction of pCR in breast cancer, enhances the robustness and reliability of the model, avoids mispredictions caused by a single image, and provides a more effective auxiliary diagnostic tool.
Smart Images

Figure CN120032879A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of breast cancer prognostic response classification, and in particular to a breast cancer pCR early prediction system based on limited medical texts. Background Art
[0002] Breast cancer is one of the most common malignant tumors in the world. Its incidence has surpassed that of lung cancer and colorectal cancer, posing a serious threat to women's health. Neoadjuvant chemotherapy (NAC) is the first choice for patients with locally advanced breast cancer. It can reduce tumor size, thereby improving breast conservation rate and reducing the scope of axillary surgery. Medical studies have shown that the prognosis of breast cancer patients will be improved if they achieve pathologic complete response (pCR) after receiving neoadjuvant chemotherapy. Therefore, the efficacy of neoadjuvant chemotherapy can be evaluated by determining the pCR status of patients. Since breast cancer is a highly heterogeneous malignant tumor, the efficacy of NAC varies significantly among different patients. If pCR is classified after treatment, patients will suffer unnecessary pain. Therefore, a system that can classify and accurately predict pCR is urgently needed in clinical practice to assist clinicians in formulating personalized treatment plans.
[0003] In the research field of pCR prediction, the use of MRI image representation combined with deep learning technology has become a crucial research direction. As an advanced medical imaging technology, MRI is widely used in the diagnosis and efficacy monitoring of breast cancer due to its non-invasiveness, high resolution, and ability to display tissue structure in multiple dimensions, and is relatively easy to obtain. However, the problem is that if only a single type of image is used or the model design does not consider the characteristics of the image itself, it is easy to make wrong predictions. At the same time, although MRI image resources are abundant, MRI data with precise annotations is extremely scarce. Annotation work requires not only professional knowledge, but also a lot of time and effort. Especially in the complex scenario of NAC efficacy evaluation of breast cancer, the number and accuracy of annotations are directly related to the effect of deep learning model training. Therefore, how to efficiently and accurately extract tumor features from images with the assistance of limited medical text information, thereby improving the accuracy of early prediction of pCR, is still a huge challenge. Summary of the invention
[0004] The purpose of the present invention is to overcome the difficulty of limited medical text in early prediction of pCR, and to provide a breast cancer pCR early prediction system based on limited medical text, breaking through the limitations of traditional reliance on limited medical text information and a single type of image for pCR classification, improving the accuracy of early prediction, and providing a more effective auxiliary tool for doctors' diagnosis.
[0005] To achieve the above-mentioned purpose, the technical solution provided by the present invention is: a breast cancer pCR early prediction system based on limited medical text, comprising:
[0006] The pre-training module is used to obtain the parameters of the large language model adapted to the breast cancer pCR prediction task, input limited medical text into the large language model LLM, generate more text descriptions, extract feature descriptions from these text descriptions through the text encoder and generate text embeddings, and then pass the text embeddings through the classification head to generate pCR classification results, and align the pCR classification results with the actual category labels through the cross entropy loss function, so that the parameters of the large language model LLM are adapted to the pCR prediction task;
[0007] A training module, used to use the large language model parameters obtained by the pre-training module to train the pCR prediction model based on limited text to obtain an optimal model, wherein the pCR prediction model based on limited text first loads diffusion weighted imaging DWI and dynamic contrast enhanced magnetic resonance imaging DCE-MRI, and uses the morphological feature extraction module and the hemodynamic feature extraction module to extract morphological features and hemodynamic features therefrom respectively, and at the same time uses the text feature extraction module to expand the medical text and extract medical text features therefrom, and finally uses the feature fusion module to fuse the medical text features, morphological features and hemodynamic features, and inputs the fused features into the classification head to generate the breast cancer pCR prediction result;
[0008] The inference module is used to generate early prediction results of breast cancer pCR using the optimal model obtained in the training module.
[0009] Furthermore, the pre-training module inputs limited medical text, generates text description using a large language model, expands the limited medical text, inputs the expanded medical text into a text encoder and a classification head, and trains through a cross entropy loss, wherein:
[0010] Said limited medical text includes etiology, disease progression, clinical manifestations, and therapeutic interventions;
[0011] The text encoder uses BioBERT to generate text embeddings, which are then fed into the classification head;
[0012] The classification head includes two fully connected neural networks and a softmax function. Both fully connected neural networks have only one layer. The input is text embedding, and the output is the pCR classification result. The output dimension is 2, which indicates the probability that the patient's status after neoadjuvant chemotherapy is pCR or not reaching pCR. The training is performed by minimizing the cross entropy loss between the class probability and the class label, which is expressed as:
[0013]
[0014] Where P k Indicates limited medical text, L k represents the label associated with the category, Entropy represents the cross entropy loss, and loss is the calculated loss. For a limited medical text of a certain category k, a large language model is used to expand it into a richer text description G(P k ,L k ), then the text encoder E encodes it into a text embedding This embedding is then passed through the classification head S to generate class probabilities, and finally the class probabilities are compared with the input labels L using the cross entropy loss. k The difference between.
[0015] Furthermore, the pCR prediction model based on limited text includes a morphological feature extraction module, a hemodynamic feature extraction module, a text feature extraction module, a feature fusion module and a classification head. The model input is limited medical text and two images of DWI and DCE-MRI, and the output is the pCR classification result. The morphological feature extraction module is composed of ResNet50, and the output morphological feature M 1 The text feature extraction module consists of a large language model and BioBERT, and the classification head consists of two fully connected neural networks and a Softmax function, which outputs the medical text feature T 1 .
[0016] Furthermore, the hemodynamic feature extraction module includes 12 feature interaction modules, the output of the previous feature interaction module is the input of the next feature interaction module, wherein the feature interaction module includes a convolution block, a visual transformer (ViT) block and a layer-by-layer fusion module, the details of which are as follows:
[0017] The convolution block consists of a residual block, that is, first a 3×3×3 convolution layer and a batch normalization layer, followed by a ReLU activation function, and then another 3×3×3 convolution layer and a batch normalization layer. The features output by this layer are residually connected with the input of the convolution block, and then pass through a ReLU activation function to obtain the output of the i-th convolution block, that is, the tumor spatial feature representation B in DCE-MRI i ;
[0018] The ViT block first divides the input image sequence into multiple fixed-size image blocks, which are regarded as the basic units of model processing. After each image block is linearly transformed, it is flattened into a vector, and then a position code is added to the vector and input into the Transformer encoder. The position code is expressed as:
[0019]
[0020] Where, PE (pos,t) Indicates the position encoding of the specified position, pos represents each position, t=2j indicates that the even dimension uses the sine function to represent the position encoding, t=2j+1 indicates that the odd dimension uses the cosine function to represent the position encoding, j represents the dimension index of the position encoding vector, and dim represents the embedding dimension;
[0021] The vectors before and after the i-th ViT block are residually connected, and then a 1×1 convolution is performed to obtain the output of the i-th ViT block, that is, the sequence change feature representation V of DCE-MRI i ;
[0022] The layer-by-layer fusion module effectively fuses the tumor spatial feature representation and sequence change feature representation in DCE-MRI. The module first upsamples the output of the ViT block, which includes reshaping the vector shape, convolution calculation and batch normalization. Then, it is added to the output of the convolution block to obtain feature A. i+1 , the added feature A i+1 Input to the next feature interaction module, the layer-by-layer fusion module is expressed as:
[0023] A i+1 =BN(conv(reshape(V i )))+B i
[0024] Finally, the hemodynamic feature extraction module outputs a feature vector, which is the output of the last feature interaction module, representing the hemodynamic feature D of the tumor captured before treatment. 1 .
[0025] Furthermore, the feature fusion module uses the morphological feature M 1 , hemodynamic characteristics 1 and medical text features T 1 As input, the dynamic visual features fully integrated by the three are output, including the following steps:
[0026] Medical Text Features 1 and morphological characteristics 1 Perform cross attention calculation to obtain feature M 2 Then use the learnable parameter g 1 Weighted feature M 2 The different dimensions of feature M 3 , which means that feature M 3 Emphasize those features related to medical text, while reducing the influence of irrelevant features, expressed as:
[0027]
[0028] M 3 =g 1 *M 2
[0029] Where, d 1 Indicates M 1 Dimensions;
[0030] Medical Text Features 1 and hemodynamic characteristics 1 Perform cross attention calculation to obtain feature D 2 Then use the learnable parameter g 2 Weighted D 2 Different dimensions of D also extract features D that are highly relevant to medical text 3 , expressed as:
[0031]
[0032] D 3 =g 2 *D 2
[0033] Where, d 2 Indicates D 1 Dimensions;
[0034] With learnable parameters g 3 Weighted medical text features T 1 , then with M 3 and D 3 The fused comprehensive feature Y is obtained by splicing, which represents the medical text features and the visual features highly related to the medical text, and is used for result prediction, expressed as:
[0035] Y=concat(M 3 ,D 3 ,g 3 *T 1 )
[0036] In the formula, concat represents the vector concatenation operation.
[0037] Furthermore, the reasoning module uses the optimal model of the pCR prediction model based on limited text to make accurate predictions. The reasoning process first inputs limited medical text, DWI and DCE-MRI, and then the text feature extraction module, the morphological feature extraction module, and the hemodynamic feature extraction module extract features respectively. The feature fusion module dynamically fuses these features and generates the final features. The classification head predicts pCR based on the final features.
[0038] Compared with the prior art, the present invention has the following advantages and beneficial effects:
[0039] 1. Through pre-training, the parameters of the large language model are optimized. With the help of the parameters learned in the pre-training stage, the large language model can help the classification model to learn better in a limited data environment, improve the generalization ability of the model, and solve the problem of limited amount of high-quality labeled medical data.
[0040] 2. Through the morphological feature extraction module and the hemodynamic feature extraction module, the different characteristics of DWI and DCE-MRI are used to obtain the morphological and hemodynamic information in the two images respectively, and the multi-source data is comprehensively used to capture tumor characteristics more comprehensively, enhance the robustness and reliability of the prediction model, and significantly improve the accuracy of the prediction. Unlike other mass detection systems, the present invention avoids the problem of incorrect prediction caused by relying on a single image.
[0041] 3. The feature fusion module is used to adaptively fuse the three types of features, and the text features are used to guide the learning of image features. The fusion is given a certain degree of interpretability, which is conducive to providing a stronger basis for decision-making and also helps further verification of the results. BRIEF DESCRIPTION OF THE DRAWINGS
[0042] Figure 1 It is a flow chart of the system of the present invention.
[0043] Figure 2 It is a structural diagram of the pre-training module.
[0044] Figure 3 It is a schematic diagram of the structure of the pCR prediction model based on limited text.
[0045] Figure 4 It is a structural diagram of the feature fusion module. DETAILED DESCRIPTION
[0046] The present invention is further described in detail below with reference to specific embodiments.
[0047] This embodiment discloses a breast cancer pCR early prediction system based on limited medical text. It is a system developed using Python language and can run on multiple platforms. The implementation process is as follows: Figure 1 As shown, it includes:
[0048] A pre-training module is used to obtain parameters of a large language model adapted to the pCR prediction task of breast cancer, input limited medical text into the large language model (LLM), generate more text descriptions, extract feature descriptions from these text descriptions through a text encoder and generate text embeddings, then pass the text embeddings through a classification head to generate pCR classification results, and align the pCR classification results with the true category labels through a cross entropy loss function, so that the parameters of the large language model (LLM) are adapted to the pCR prediction task;
[0049] A training module is used to use the large language model parameters obtained by the pre-training module to train the pCR prediction model based on limited text to obtain an optimal model, wherein the pCR prediction model based on limited text first loads diffusion weighted imaging (DWI) and dynamic contrast enhancement magnetic resonance imaging (DCE-MRI), and uses the morphological feature extraction module and the hemodynamic feature extraction module to extract morphological features and hemodynamic features respectively, and uses the text feature extraction module to expand the medical text and extract medical text features from it, and finally uses the feature fusion module to fuse the medical text features, morphological features and hemodynamic features, and inputs the fused features into the classification head to generate the breast cancer pCR prediction result;
[0050] The inference module is used to generate early prediction results of breast cancer pCR using the optimal model obtained in the training module.
[0051] Specifically, Figure 2 As shown, the pre-training module inputs limited medical text, generates text description using a large language model, expands the limited medical text, inputs the expanded medical text into a text encoder and a classification head, and trains through a cross entropy loss, wherein:
[0052] Limited medical text includes etiology, disease progression, clinical presentation, and therapeutic interventions;
[0053] The text encoder uses BioBERT to generate text embeddings, which are then fed into the classification head;
[0054] The classification head contains two fully connected neural networks and a softmax function. Both fully connected neural networks have only one layer. The input is text embedding, and the output is the pCR classification result. The output dimension is 2, which indicates the probability of the patient's status being pCR or not reaching pCR after neoadjuvant chemotherapy. It is trained by minimizing the cross entropy loss between the class probability and the class label, which is expressed as:
[0055]
[0056] Where P k Indicates limited medical text, L k represents the label associated with the category, Entropy represents the cross entropy loss, and loss is the calculated loss. For a limited medical text of a certain category k, a large language model is used to expand it into a richer text description G(P k ,L k ), then the text encoder E encodes it into a text embedding This embedding is then passed through the classification head S to generate class probabilities, and finally the class probabilities are compared with the input labels L using the cross entropy loss. k The difference between.
[0057] Specifically, Figure 3 As shown in FIG. 1 , the pCR prediction model based on limited text includes a morphological feature extraction module, a hemodynamic feature extraction module, a text feature extraction module, a feature fusion module, and a classification head. The model input is limited medical text and two images, DWI and DCE-MRI, and the output is the pCR classification result. The morphological feature extraction module is composed of ResNet50, and the output morphological feature M 1 The text feature extraction module consists of a large language model and BioBERT, and the classification head consists of two fully connected neural networks and a Softmax function, which outputs the medical text feature T 1 .
[0058] The hemodynamic feature extraction module is as follows: Figure 3 As shown in the branch below, there are 12 feature interaction modules. The output of the previous feature interaction module is the input of the next feature interaction module. The feature interaction module includes convolution blocks, vision transformers (ViT blocks) and layer-by-layer fusion modules. The details are as follows:
[0059] The convolution block consists of a residual block, that is, first a 3×3×3 convolution layer and a batch normalization layer, followed by a ReLU activation function, and then another 3×3×3 convolution layer and a batch normalization layer. The features output by this layer are residually connected with the input of the convolution block, and then pass through a ReLU activation function to obtain the output of the i-th convolution block, that is, the tumor spatial feature representation B in DCE-MRI i ;
[0060] The ViT block first divides the input image sequence into multiple fixed-size image blocks, which are regarded as the basic units of model processing. After each image block is linearly transformed, it is flattened into a vector, and then a position code is added to the vector and input into the Transformer encoder. The position code is expressed as:
[0061]
[0062] Where, PE (pos,t) Indicates the position encoding of the specified position, pos represents each position, t=2j indicates that the even dimension uses the sine function to represent the position encoding, t=2j+1 indicates that the odd dimension uses the cosine function to represent the position encoding, j indicates the dimension index of the position encoding vector, dim indicates the embedding dimension, which is 512;
[0063] The vectors before and after the i-th ViT block are residually connected, and then a 1×1 convolution is performed to obtain the output of the i-th ViT block, that is, the sequence change feature representation V of DCE-MRI i ;
[0064] The layer-by-layer fusion module effectively fuses the tumor spatial feature representation and the sequence change feature representation in DCE-MRI. The module first upsamples the output of the ViT block, which includes changing the vector shape (reshape), convolution calculation and batch normalization. Then, it is added to the output of the convolution block to obtain feature A i+1 , the added feature A i+1 Input to the next feature interaction module, the layer-by-layer fusion module is expressed as:
[0065] A i+1 =BN(conv(reshape(V i )))+B i
[0066] Finally, the hemodynamic feature extraction module outputs a feature vector, which is the output of the last feature interaction module, representing the hemodynamic feature D of the tumor captured before treatment. 1 .
[0067] Specifically, the location of the feature fusion module is as follows: Figure 3 The specific results are shown in Figure 4 As shown, the morphological features M 1 , hemodynamic characteristics 1 and medical text features T 1 As input, the dynamic visual features fully integrated by the three are output, including the following steps:
[0068] Medical Text Features 1 and morphological characteristics 1 Perform cross attention calculation to obtain feature M 2 Then use the learnable parameter g 1 Weighted feature M 2 The different dimensions of feature M 3 , which means that feature M 3 Emphasize those features related to medical text, while reducing the influence of irrelevant features, expressed as:
[0069]
[0070] M 3 =g 1 *M 2
[0071] Where, d 1 Indicates M 1 Dimensions;
[0072] Medical Text Features 1 and hemodynamic characteristics 1 Perform cross attention calculation to obtain feature D 2 Then use the learnable parameter g 2 Weighted D 2 Different dimensions of D also extract features D that are highly relevant to medical text 3 , expressed as:
[0073]
[0074] D 3 =g 2 *D 2
[0075] Where, d 2 Indicates D 1 Dimensions;
[0076] With learnable parameters g 3 Weighted medical text features T 1 , then with M 3 and D 3The fused comprehensive feature Y is obtained by splicing, which represents the medical text features and the visual features highly related to the medical text, and is used for result prediction, expressed as:
[0077] Y=concat(M 3 ,D 3 ,g 3 *T 1 )
[0078] In the formula, concat represents the vector concatenation operation.
[0079] Specifically, the reasoning module uses the optimal model of the pCR prediction model based on limited text to make accurate predictions. The reasoning process first inputs limited medical text, DWI and DCE-MRI, and then the text feature extraction module, the morphological feature extraction module, and the hemodynamic feature extraction module extract features respectively. The feature fusion module dynamically fuses these features and generates the final features. The classification head predicts pCR based on the final features.
[0080] The above embodiments are preferred implementation modes of the present invention, but the implementation modes of the present invention are not limited to the above embodiments. Any other changes, modifications, substitutions, combinations, and simplifications that do not deviate from the spirit and principles of the present invention should be equivalent replacement methods and are included in the protection scope of the present invention.
Claims
1. An early prediction system for breast cancer pCR based on limited medical text, characterized by: include: The pre-training module is used to obtain the parameters of the large language model adapted to the breast cancer pCR prediction task, input limited medical text into the large language model LLM, generate more text descriptions, extract feature descriptions from these text descriptions through the text encoder and generate text embeddings, and then pass the text embeddings through the classification head to generate pCR classification results, and align the pCR classification results with the actual category labels through the cross entropy loss function, so that the parameters of the large language model LLM are adapted to the pCR prediction task; A training module, used to use the large language model parameters obtained by the pre-training module to train the pCR prediction model based on limited text to obtain an optimal model, wherein the pCR prediction model based on limited text first loads diffusion weighted imaging DWI and dynamic contrast enhanced magnetic resonance imaging DCE-MRI, and uses the morphological feature extraction module and the hemodynamic feature extraction module to extract morphological features and hemodynamic features therefrom respectively, and at the same time uses the text feature extraction module to expand the medical text and extract medical text features therefrom, and finally uses the feature fusion module to fuse the medical text features, morphological features and hemodynamic features, and inputs the fused features into the classification head to generate the breast cancer pCR prediction result; The inference module is used to generate early prediction results of breast cancer pCR using the optimal model obtained in the training module.
2. The breast cancer pCR early prediction system based on limited medical text according to claim 1, characterized in that: The pre-training module inputs limited medical text, generates text descriptions using a large language model, expands the limited medical text, inputs the expanded medical text into a text encoder and a classification head, and trains through a cross entropy loss, wherein: Said limited medical text includes etiology, disease progression, clinical manifestations, and therapeutic interventions; The text encoder uses BioBERT to generate text embeddings, which are then fed into the classification head; The classification head includes two fully connected neural networks and a softmax function. Both fully connected neural networks have only one layer. The input is text embedding, and the output is the pCR classification result. The output dimension is 2, which indicates the probability that the patient's status after neoadjuvant chemotherapy is pCR or not reaching pCR. The training is performed by minimizing the cross entropy loss between the class probability and the class label, which is expressed as: Where P k Indicates limited medical text, L k represents the label associated with the category, Entropy represents the cross entropy loss, and loss is the calculated loss. For a limited medical text of a certain category k, a large language model is used to expand it into a richer text description G(P k ,L k ), then the text encoder E encodes it into a text embedding This embedding is then passed through the classification head S to generate class probabilities, and finally the class probabilities are compared with the input labels L using the cross entropy loss. k The difference between.
3. The breast cancer pCR early prediction system based on limited medical text according to claim 1, characterized in that: The pCR prediction model based on limited text includes a morphological feature extraction module, a hemodynamic feature extraction module, a text feature extraction module, a feature fusion module and a classification head. The model input is limited medical text and two images, DWI and DCE-MRI, and the output is the pCR classification result. The morphological feature extraction module is composed of ResNet50 and outputs morphological features M1. The text feature extraction module is composed of a large language model and BioBERT. The classification head is composed of two fully connected neural networks and a Softmax function, and outputs medical text features T1.
4. The breast cancer pCR early prediction system based on limited medical text according to claim 3 is characterized in that: The hemodynamic feature extraction module includes 12 feature interaction modules, and the output of the previous feature interaction module is the input of the next feature interaction module. The feature interaction module includes convolutional blocks, visual transformers (ViT blocks) and layer-by-layer fusion modules. The details are as follows: The convolution block consists of a residual block, that is, first a 3×3×3 convolution layer and a batch normalization layer, followed by a ReLU activation function, and then another 3×3×3 convolution layer and a batch normalization layer. The features output by this layer are residually connected with the input of the convolution block, and then pass through a ReLU activation function to obtain the output of the i-th convolution block, that is, the tumor spatial feature representation B in DCE-MRI i ; The ViT block first divides the input image sequence into multiple fixed-size image blocks, which are regarded as the basic units of model processing. After each image block is linearly transformed, it is flattened into a vector, and then a position code is added to the vector and input into the Transformer encoder. The position code is expressed as: Where, PE (pos,t) Indicates the position encoding of the specified position, pos represents each position, t=2j indicates that the even dimension uses the sine function to represent the position encoding, t=2j+1 indicates that the odd dimension uses the cosine function to represent the position encoding, j represents the dimension index of the position encoding vector, and dim represents the embedding dimension; The vectors before and after the i-th ViT block are residually connected, and then a 1×1 convolution is performed to obtain the output of the i-th ViT block, that is, the sequence change feature representation V of DCE-MRI i ; The layer-by-layer fusion module effectively fuses the tumor spatial feature representation and sequence change feature representation in DCE-MRI. The module first upsamples the output of the ViT block, which includes reshaping the vector shape, convolution calculation and batch normalization. Then, it is added to the output of the convolution block to obtain feature A. i+1 , the added feature A i+1 Input to the next feature interaction module, the layer-by-layer fusion module is expressed as: A i+1 =BN(conv(reshape(V i )))+B i Finally, the hemodynamic feature extraction module outputs a feature vector, which is the output of the last feature interaction module, representing the hemodynamic feature D1 of the tumor captured before treatment.
5. The breast cancer pCR early prediction system based on limited medical text according to claim 3, characterized in that: The feature fusion module takes the morphological feature M1, the hemodynamic feature D1 and the medical text feature T1 as input, and takes the dynamic visual feature fully fused by the three as output, and includes the following steps: The medical text feature T1 and the morphological feature M1 are cross-attentioned to obtain feature M2, and then the different dimensions of feature M2 are weighted by the learnable parameter g1 to obtain feature M3, which means that feature M3 emphasizes those features related to the medical text while reducing the influence of irrelevant features, expressed as: M3=g1*M2 Where d1 represents the dimension of M1; The medical text feature T1 and the hemodynamic feature D1 are cross-attentionally calculated to obtain feature D2, and then the different dimensions of D2 are weighted by the learnable parameter g2, and feature D3 that is highly relevant to the medical text is also extracted, which is expressed as: D3=g2*D2 Where d2 represents the dimension of D1; The medical text feature T1 is weighted with the learnable parameter g3, and then concatenated with M3 and D3 to obtain the fused comprehensive feature Y, which represents the medical text feature and the visual feature highly related to the medical text, and is used for result prediction, expressed as: Y=concat(M3,D3,g3*T1) Where concat represents the vector concatenation operation.
6. The breast cancer pCR early prediction system based on limited medical text according to claim 1, characterized in that: The reasoning module uses the optimal model of the pCR prediction model based on limited text to make accurate predictions. The reasoning process first inputs limited medical text, DWI and DCE-MRI, then the text feature extraction module, the morphological feature extraction module, and the hemodynamic feature extraction module extract features respectively, the feature fusion module dynamically fuses these features and generates the final features, and the classification head predicts pCR based on the final features.