Neoadjuvant chemotherapy reaction prediction method and device, electronic equipment and storage medium

By combining multidimensional feature extraction and fusion of CT radiomics, WSI digital pathology, and clinical information, the problem of insufficient accuracy in predicting neoadjuvant chemotherapy response has been solved, enabling more accurate prediction of chemotherapy response and personalized treatment.

CN121905415APending Publication Date: 2026-04-21NEUSOFT GRP (WUHAN) CO LTD +1
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
NEUSOFT GRP (WUHAN) CO LTD
Filing Date
2025-12-31
Publication Date
2026-04-21

AI Technical Summary

Technical Problem

The accuracy of predicting neoadjuvant chemotherapy response in existing technologies is insufficient, mainly due to the limitations of single-dimensional data.

Method used

By combining CT radiomics, WSI digital pathology, and clinical information, multidimensional features are extracted and fused to construct a multidimensional prediction system, including CT feature extraction, WSI feature extraction, clinical feature extraction, and feature fusion. Finally, the prediction results of chemotherapy response are determined through a joint prediction module.

Benefits of technology

It improves the accuracy of predicting neoadjuvant chemotherapy response, enabling more precise identification of chemotherapy responses and optimization of individualized treatment plans.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121905415A_ABST
    Figure CN121905415A_ABST
Patent Text Reader

Abstract

The invention discloses a neoadjuvant chemotherapy reaction prediction method and device, electronic equipment and a storage medium. The method comprises the following steps: acquiring preprocessed CT data, WSI data and clinical data; performing feature extraction on the CT data to obtain CT features, and performing feature extraction on the WSI data to obtain WSI features; performing feature extraction on the clinical data through a clinical feature extraction module to obtain clinical features; fusing the CT features and the WSI features through a feature fusion module to obtain image fusion features; and determining a neoadjuvant chemotherapy reaction prediction result through a combined prediction module according to the image fusion features and the clinical features. According to the invention, the accuracy of the prediction result of the neoadjuvant chemotherapy reaction can be improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application belongs to the field of medical data processing technology, specifically relating to a method, device, electronic device, and storage medium for predicting the response to neoadjuvant chemotherapy. Background Technology

[0002] In the global disease spectrum, malignant and benign tumors together constitute a severe public health challenge. Data shows that nearly 20 million new cases of malignant tumors are diagnosed globally each year, and approximately 9.7 million people die from them, highlighting their high lethality. The threat of these diseases stems not only from their biological characteristics of unlimited proliferation and metastasis, but also from risk factors such as population aging and lifestyle changes. In contrast, benign tumors generally exhibit relatively mild biological behavior. However, this does not mean that their potential risks can be ignored. For example, benign intracranial lesions, although not malignant, continue to grow within the closed cranial cavity, compressing key brain functional areas and causing irreversible neurological damage, severely impacting patients' long-term quality of life and mental health. Therefore, the overall prevention and control of tumors, regardless of whether they are benign or malignant, must be taken seriously.

[0003] Against this backdrop, promoting standardized treatment pathways has become a core strategy for reducing the burden of cancer globally. Neoadjuvant chemotherapy is a crucial component of this strategy. Neoadjuvant chemotherapy refers to systemic chemotherapy administered before radical local treatments such as surgery, and it has now become a core strategy for the comprehensive treatment of various solid tumors. Its importance is reflected in three main aspects: First, it can effectively shrink the primary tumor lesion, thereby increasing the surgical resection rate and creating the possibility of preserving organ function (such as breast-conserving surgery for breast cancer); second, it can eliminate potential micrometastases in the body earlier, reducing the risk of recurrence and metastasis from the source; finally, it acts like an "in vivo drug sensitivity test," providing crucial evidence for assessing prognosis and developing subsequent treatment plans by observing the tumor's response to treatment.

[0004] Different cancer types exhibit vastly different responses and strategies to neoadjuvant chemotherapy. For example, in gastric cancer, neoadjuvant chemotherapy using the FLOT regimen has become the standard treatment for locally advanced patients, significantly improving surgical radicality and patient survival. Furthermore, different patients respond differently to neoadjuvant chemotherapy; therefore, predicting the response to neoadjuvant chemotherapy is necessary. Accurate prediction of pathological complete response (pCR) with neoadjuvant chemotherapy is crucial for optimizing individualized cancer treatment. Current technologies generally use single-dimensional data for neoadjuvant chemotherapy response prediction, which has limitations and leads to insufficient accuracy in the prediction results. Summary of the Invention

[0005] The purpose of this application is to provide a new method, device, electronic device, and storage medium for predicting adjuvant chemotherapy response, which can solve the problem of insufficient accuracy of prediction results.

[0006] To solve the above-mentioned technical problems, this application is implemented as follows: In a first aspect, embodiments of this application provide a method for predicting neoadjuvant chemotherapy response, the method comprising: Acquire preprocessed CT data, WSI data, and clinical data; The CT data is subjected to feature extraction to obtain CT features, and the WSI data is subjected to feature extraction to obtain WSI features; The clinical data is used to extract features through the clinical feature extraction module to obtain clinical features; The CT features and the WSI features are fused by the feature fusion module to obtain image fusion features; The neoadjuvant chemotherapy response prediction result is determined by the joint prediction module based on the image fusion features and the clinical features.

[0007] Secondly, embodiments of this application provide a neoadjuvant chemotherapy response prediction device, comprising: The data acquisition module is used to acquire preprocessed CT data, WSI data, and clinical data. The first feature extraction module is used to extract features from the CT data to obtain CT features and to extract features from the WSI data to obtain WSI features. The second feature extraction module is used to extract features from the clinical data through the clinical feature extraction module to obtain clinical features; The fusion module is used to fuse the CT features and the WSI features through the feature fusion module to obtain image fusion features; The prediction module is used to determine the neoadjuvant chemotherapy response prediction result based on the image fusion features and the clinical features through the joint prediction module.

[0008] Thirdly, embodiments of this application provide an electronic device including a processor, a memory, and a program or instructions stored in the memory and executable on the processor, wherein the program or instructions, when executed by the processor, implement the steps of the method described in the first aspect.

[0009] Fourthly, embodiments of this application provide a readable storage medium on which a program or instructions are stored, which, when executed by a processor, implement the steps of the method described in the first aspect.

[0010] Fifthly, embodiments of this application provide a chip, the chip including a processor and a communication interface, the communication interface being coupled to the processor, the processor being used to run programs or instructions to implement the method as described in the first aspect.

[0011] In this embodiment, CT features are obtained by extracting features from CT data, WSI features are obtained by extracting features from WSI data, and clinical features are obtained by extracting features from clinical data. The CT features and WSI features are then fused to obtain image fusion features. Based on the image fusion features and clinical features, the neoadjuvant chemotherapy response prediction result is determined. It can be seen that CT data, WSI data, and clinical data are fully utilized in the neoadjuvant chemotherapy response prediction. These data can complement each other and make up for the deficiencies of single-dimensional data, thereby making the input features for neoadjuvant chemotherapy response prediction more comprehensive and improving the accuracy of the neoadjuvant chemotherapy response prediction result. Attached Figure Description

[0012] Figure 1 This is a flowchart of a novel adjuvant chemotherapy response prediction method provided in an embodiment of this application; Figure 2 This is a schematic diagram of the neoadjuvant chemotherapy response prediction model in the embodiments of this application; Figure 3 This is a flowchart of the training process of the neoadjuvant chemotherapy response prediction model in the embodiments of this application; Figure 4 This is a schematic diagram of the structure of a novel adjuvant chemotherapy response prediction device provided in an embodiment of this application; Figure 5 A schematic diagram of the hardware structure of an electronic device to implement an embodiment of this application. Detailed Implementation

[0013] The technical solutions of the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of this application, not all embodiments. Based on the embodiments of this application, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of this application.

[0014] The terms "first," "second," etc., used in the specification and claims of this application are used to distinguish similar objects and not to describe a specific order or sequence. It should be understood that such use of data can be interchanged where appropriate so that embodiments of this application can be implemented in orders other than those illustrated or described herein. Furthermore, in the specification and claims, "and / or" indicates at least one of the connected objects, and the character " / " generally indicates that the preceding and following objects are in an "or" relationship.

[0015] The neoadjuvant chemotherapy response prediction method provided in this application will be described in detail below with reference to the accompanying drawings, through specific embodiments and application scenarios.

[0016] This application's embodiments focus on the fact that accurate prediction of pathological complete response (pCR) after neoadjuvant chemotherapy is key to optimizing personalized cancer treatment, and that integrating CT radiomics, WSI (Whole Slide Image), and clinical information can construct a multidimensional prediction system. Pre-treatment CT radiomics provides a macroscopic assessment by extracting quantitative features of the tumor (such as size, shape, texture heterogeneity, and enhancement pattern). High texture heterogeneity often indicates poor chemotherapy response, and these early changes are strong predictors of pCR. At the same time, the assessment of lymph node and distant metastatic responses can comprehensively reflect the systemic treatment effect. WSI digital pathology provides core biological information at the microscopic level: pre-treatment biopsies can assess traditional pathological indicators (such as the possibility of better response in the midgut type according to the Lauren classification of histology, differentiation degree, tumor budding, and lymphovascular invasion) as well as key molecular markers (such as HER2 positivity for trastuzumab-containing regimens and dMMR / MSI-H status predicting benefit from immunotherapy); while pathomics, through AI-quantified extraction of morphological features from WSI (nuclear morphology, arrangement patterns, stroma composition, and spatial distribution of immune cell infiltration), can uncover chemotherapy sensitivity biomarkers that go beyond manual interpretation. Clinical information, as the cornerstone of personalized treatment, encompasses patient factors (age, performance status, ECOG / PS, nutritional status such as albumin, and medical history), treatment parameters (specific regimen and dosage intensity), and laboratory indicators (CEA, CA19-9 baseline levels and their dynamic changes, reflecting tumor burden regression). These three types of information complement each other: CT provides macroscopic structural and dynamic response maps, WSI reveals microscopic tissue and molecular mechanisms, and clinical information anchors the individual treatment context. Single-dimensional prediction has limitations, while multimodal data fusion can significantly improve the accuracy of pCR prediction. It not only helps to screen high-response patients to avoid overtreatment, but also enables early identification of drug-resistant populations to adjust the treatment plan, ultimately achieving precise analysis of neoadjuvant chemotherapy.

[0017] Figure 1 This is a flowchart illustrating a method for predicting neoadjuvant chemotherapy response provided in an embodiment of this application. This method is applicable to scenarios involving predicting neoadjuvant chemotherapy response in tumors, such as... Figure 1 As shown, the method may include steps 110 to 150.

[0018] Step 110: Obtain preprocessed CT data, WSI data, and clinical data.

[0019] The preprocessed CT data, WSI data, and clinical data are from the same patient and the same tumor. The preprocessed CT data is of the first target size. The preprocessed WSI data includes multiple image patches of the second target size. The preprocessed clinical data is data of the target dimension.

[0020] Step 120: Extract features from CT data to obtain CT features, and extract features from WSI data to obtain WSI features.

[0021] CT and WSI features can be extracted using the image information extraction module. This module is part of the neoadjuvant chemotherapy response prediction model. It can use existing models for feature extraction, and its parameters do not need to be adjusted during the training of the neoadjuvant chemotherapy response prediction model.

[0022] Figure 2 This is a schematic diagram of the neoadjuvant chemotherapy response prediction model in an embodiment of this application. The image information extraction module may include a 3D image encoder. For example... Figure 2 As shown, for 3D CT data, a 3D image encoder such as M3D can be used to encode features to obtain CT features. The dimension of the CT features can be a preset CT feature dimension, such as 2048*768.

[0023] The image information extraction module can include a large-scale pathological model. For example... Figure 2 As shown, for the segmented WSI data, a large pathological model such as UNI is used for feature encoding. Each patch is encoded as a WSI feature of a preset WSI feature dimension. The preset WSI feature dimension can be, for example, 1*1024.

[0024] M3D and UNI can also be replaced by other similar large models, and the size of the encoded features can be changed accordingly. For example... Figure 2 As shown, the CT features after CT data encoding can be denoted as X[t*d1], and the WSI features after WSI data processing can be denoted as Y[p*d2], where p represents the number of image blocks in the WSI data. The size of p varies for different WSI data because the original size of different WSI data is different, resulting in a different number of image blocks after slicing.

[0025] Step 130: Extract clinical features from clinical data using the clinical feature extraction module to obtain clinical features.

[0026] The clinical feature extraction module is a component of the neoadjuvant chemotherapy response prediction model, used to extract features from clinical data. Preprocessed clinical data is input into the clinical feature extraction module, which then processes the data and extracts clinical features.

[0027] In some embodiments of this application, the clinical data includes multiple pieces of clinical information, and the clinical feature extraction module includes a graph transformation network; The clinical feature extraction module extracts features from the clinical data to obtain clinical features, including: treating each piece of clinical information as a node and constructing a fully connected graph structure with the nodes, initializing the weights of the edges in the graph structure as target values; and extracting features from the graph structure through a graph transformation network to obtain the clinical features.

[0028] like Figure 2 As shown, assuming there are m clinical information entries, the clinical data can be represented as a feature vector of size m*d. Here, each clinical information entry is treated as a node, and the 1*d feature vector is used as the feature of the current node. A fully connected graph structure is constructed with all nodes, and the weight of each edge in the graph structure is initialized to a target value (e.g., an integer 1). Therefore, there is a clinical information graph consisting of node feature data N[m*d] and edge data A[m*m]. To achieve efficient processing of the relationship information between clinical information entries in the graph structure and to capture long-range dependencies, a Graph Transformer 3 network can be used to extract clinical features. The extracted node features can be denoted as Z[m,d3].

[0029] Graph Transformation Networks (GNNs) are neural network models that apply the Transformer architecture to graph-structured data. By fusing the fundamental principles of Graph Neural Networks (GNNs) with the self-attention mechanism of the Transformer, this model effectively processes the relationship information between nodes in a graph and captures long-range dependencies. The node and edge features in the graph structure constitute the input data for GNNs. The self-attention mechanism in GNNs dynamically evaluates the strength of the association between nodes by calculating attention scores.

[0030] Graph transformation networks utilize Laplacian positional encoding to mathematically represent node positions using the eigenvectors of the graph's Laplacian matrix. This encoding method effectively captures the structural features of the graph, encoding connectivity and spatial relationships. Through Laplacian positional encoding, graph transformation networks can distinguish the structural characteristics of nodes, thus achieving efficient learning on unstructured or irregular graph data.

[0031] Message passing and aggregation mechanisms are technical components of graph transformation networks, with the following applications: message passing enables information exchange between nodes and their neighbors, while aggregation integrates the acquired information into effective feature representations. The synergistic effect of these two components allows graph transformation networks to learn deep representations of nodes, edges, and the overall graph structure, providing a technical foundation for solving complex graph tasks.

[0032] Feedforward networks, combined with nonlinear activation functions, play a key role in graph transformation networks, mainly used to optimize node embedding, introduce nonlinear characteristics, and enhance the model's pattern recognition capabilities.

[0033] Graph transformation networks support both local context (aggregated through neighborhood information) and global context (through full graph attention) modeling to capture long-range dependencies. Local context focuses on the immediate neighborhood information of nodes, including neighboring nodes and their connecting edges. Global context techniques aim to capture and process information from the entire graph structure or its main components.

[0034] When extracting features from a graph structure using a graph transformation network, the node data (clinical information) and edge data (edge ​​weights) of the graph structure are input into the graph transformation network. The graph transformation network performs Laplacian position encoding on the node data as a mathematical representation of the node position. The graph transformation network determines the deep representation of nodes, edges, and the overall graph structure through message passing and aggregation mechanisms. The graph transformation network optimizes node embedding and introduces nonlinear characteristics through a feedforward network. The graph transformation network obtains the final clinical features by processing the aforementioned graph structure data with local and global context. By constructing a fully connected graph structure based on each piece of clinical information and extracting features from the graph structure through a graph transformation network, the relationship information between different clinical information can be processed efficiently, long-range dependencies can be captured, the accuracy of clinical feature extraction can be improved, and the accuracy of neoadjuvant chemotherapy response prediction results can be further improved.

[0035] Step 140: The CT features and WSI features are fused using the feature fusion module to obtain image fusion features.

[0036] The feature fusion extraction module is a component of the neoadjuvant chemotherapy response prediction model, used to fuse CT and WSI features. CT and WSI features are input into the feature fusion module, which then fuses them into image fusion features.

[0037] In some embodiments of this application, CT features and WSI features are fused by a feature fusion module to obtain image fusion features, including: performing self-attention processing on CT features by the feature fusion module to obtain new CT features; performing self-attention processing on WSI features by the feature fusion module to obtain new WSI features; and fusing the new CT features and the new WSI features by the feature fusion module to obtain image fusion features.

[0038] like Figure 2 As shown, in feature fusion module 4, self-attention processing is performed on CT feature X[t*d1]. Feature fusion module 4 may include linear + tanh + linear... Figure 2 In neural networks, SA-Map stands for linear+tanh+linear. It maps the CT feature X[t*d1] to a first attention matrix Xa, and then applies the first attention matrix Xa to the CT feature X[t*d1], resulting in the attention-enhanced CT feature Xb[h*d1], which serves as the new CT feature. In neural networks, linear+tanh+linear refers to a specific multilayer perceptron (MLP) structure, consisting of three parts: a linear transformation layer, a hyperbolic tangent (tanh) activation function layer, and another linear transformation layer.

[0039] like Figure 2 As shown, in the feature fusion module 4, the WSI feature Y[p*d2] is subjected to self-attention processing. The WSI feature Y[p*d2] is mapped to the second attention matrix Ya by linear + tanh + linear. The second attention matrix Ya is applied to the WSI feature Y[p*d2] to obtain the WSI feature Yb[h*d2] with the attention mechanism, which is used as the new WSI feature.

[0040] The feature fusion module fuses new CT features and new WSI features to obtain image fusion features.

[0041] By applying self-attention processing to CT and WSI features, features that are beneficial for predicting neoadjuvant chemotherapy response can be amplified while features that are irrelevant to neoadjuvant chemotherapy response prediction can be suppressed. The resulting image fusion features fully integrate the features from CT and WSI that are beneficial for predicting neoadjuvant chemotherapy response, thereby further improving the accuracy of subsequent prediction results.

[0042] In some embodiments of this application, a feature fusion module fuses new CT features and new WSI features to obtain image fusion features, including: determining the correlation matrix between the new CT features and the new WSI features through the feature fusion module; determining the product of the new CT features and the correlation matrix through the feature fusion module to obtain CT features with fused WSI features; determining the product of the new WSI features and the correlation matrix through the feature fusion module to obtain WSI features with fused CT features; and stitching the CT features with fused WSI features and the WSI features with fused CT features together through the feature fusion module to obtain image fusion features.

[0043] like Figure 2 As shown, the feature fusion module 4 also includes a linear layer, which maps the new CT feature Xb[h*d1] to a one-dimensional CT feature Xc with the target dimension [1,h*k] through the linear layer. Similarly, the new WSI feature Yb[h*d2] is also mapped to a one-dimensional WSI feature Yc with the target dimension [1,h*k] through the linear layer. In order to fully interact with the information, the correlation matrix R[h*k,h*k]=Mix(Xc, Yc)=Trans(Xc)*Yc is calculated, where trans represents matrix transposition. Figure 2 In this context, Mix represents the calculation of the correlation matrix.

[0044] Calculate the product of the one-dimensional CT feature Xc and the correlation matrix to obtain the CT feature Xd=Xc*R that fuses the WSI features. Calculate the product of the one-dimensional WSI feature Yc and the correlation matrix to obtain the WSI feature Yd=Yc*R that fuses the CT features. This yields the image fusion feature I=cat(Xd,Yd) of size [h*2k].

[0045] By determining the correlation matrix and fusing new CT features and new WSI features based on the correlation matrix, CT features and WSI features can be fully interacted to complement each other, thereby making the image fusion features more accurate.

[0046] Step 150: The neoadjuvant chemotherapy response prediction result is determined by the joint prediction module based on image fusion features and clinical features.

[0047] The combined prediction module is one component of the neoadjuvant chemotherapy response prediction model. It is used to provide predictions of neoadjuvant chemotherapy response based on image fusion features and clinical features. The combined prediction module predicts the pathological complete response based on these features and clinical characteristics, obtaining a confidence level for complete remission. If this confidence level is greater than or equal to a confidence threshold, the predicted neoadjuvant chemotherapy response is determined to be a complete pathological remission; otherwise, if the confidence level is less than the confidence threshold, the predicted neoadjuvant chemotherapy response is determined to be a failure to achieve complete pathological remission.

[0048] In some embodiments of this application, the neoadjuvant chemotherapy response prediction result is determined by the joint prediction module based on image fusion features and clinical features, including: performing self-attention processing on the clinical features through the joint prediction module to obtain new clinical features; stitching the new clinical features and image fusion features through the joint prediction module to obtain multimodal fusion features; and performing neoadjuvant chemotherapy response prediction on the multimodal fusion features through the joint prediction module to obtain the neoadjuvant chemotherapy response prediction result.

[0049] The clinical feature Z[m,d3] is fused with the image fusion feature I[h*2k]. For example... Figure 2 As shown, in the joint prediction module 5, the clinical features are subjected to self-attention processing, that is, the third attention matrix Za of the clinical features is determined by linear + tanh + linear, and the third attention matrix Za is applied to the clinical features Z[m,d3] to obtain the clinical features Zb[h,d3] with attention mechanism, which is used as the new clinical features.

[0050] The new clinical features are mapped to a clinical feature Zc of the same dimension [h, 2k] as the image fusion feature to be fused using a linear layer. The clinical feature Zc to be fused and the image fusion feature I are then concatenated to obtain a multimodal fusion feature F=cat(I, Zc) of size [h, 4k]. The multimodal fusion feature is then processed through a flattening layer and a linear layer to obtain the neoadjuvant chemotherapy response prediction result Fp[1,1]. The flattening layer is a special layer that transforms multidimensional data into linear data for further processing. By applying self-attention processing to clinical features, features that are beneficial for predicting neoadjuvant chemotherapy responses can be amplified. Furthermore, by using multimodal fusion features that integrate new clinical and imaging features for neoadjuvant chemotherapy response prediction, clinical and imaging features can complement each other, fully utilizing their correlation and complementarity, thereby improving the accuracy of prediction results.

[0051] The neoadjuvant chemotherapy response prediction method provided in this application extracts features from CT data to obtain CT features, extracts features from WSI data to obtain WSI features, extracts features from clinical data to obtain clinical features, and fuses CT features and WSI features to obtain image fusion features. Based on the image fusion features and clinical features, the neoadjuvant chemotherapy response prediction result is determined. It can be seen that when predicting neoadjuvant chemotherapy response, CT data, WSI data, and clinical data are fully utilized. The data can complement each other and make up for the deficiencies of single-dimensional data, thereby making the input features for neoadjuvant chemotherapy response prediction more comprehensive and improving the accuracy of neoadjuvant chemotherapy response prediction results.

[0052] Based on the above technical solution, before acquiring the preprocessed CT data, WSI data and clinical data, the method further includes: adjusting the window width and window level of the initial CT data to the target range, and interpolating the adjusted CT data to the first target size to obtain the preprocessed CT data.

[0053] The target range can be set based on needs to distinguish between tumor-related information and background information. For example, the target range can be [0, 200]Hu.

[0054] For the initial CT data, the window width and window level can be adjusted to the target range (e.g., [0, 200] Hu). Positions smaller than the lower limit of the target range (e.g., 0) are assigned a value of 0, and positions larger than the upper limit of the target range (e.g., 200) are assigned a value of 200. This removes irrelevant background information and retains tumor-related information. Then, the data is adjusted to the first target size (e.g., 32*256*256) by nearest neighbor interpolation to obtain the preprocessed CT data.

[0055] Based on the above technical solution, before acquiring the preprocessed CT data, WSI data and clinical data, the method further includes: dividing the initial WSI data into image blocks of a second target size as preprocessed WSI data.

[0056] Because the initial WSI data is large and difficult to process, it is divided into image patches of a second target size, resulting in multiple image patches. These patches serve as the preprocessed WSI data. The second target size can be set based on requirements, for example, it can be 3*224*224.

[0057] By splitting the initial WSI data, the amount of data in subsequent processing can be reduced, making processing easier.

[0058] Based on the above technical solutions, clinical data includes numerical clinical data and text-descriptive clinical data; Before acquiring the preprocessed CT data, WSI data, and clinical data, the process includes: removing redundant information from the text-descriptive clinical data using a large language model, and encoding the redundant-free text-descriptive clinical data into a feature vector of the target dimension using a text encoder, which serves as the preprocessed text-descriptive clinical data; determining the labels corresponding to the numerical clinical data, encoding the labels, and interpolating the encoded data into the target dimension to obtain the preprocessed numerical clinical data.

[0059] Textual descriptive clinical data refers to tumor-related information described by patients, such as symptoms reported by the patient, medical history, and medication use. Numerical clinical data may include patient age, gender, blood test results, and gene immunohistochemical assays.

[0060] For text-descriptive clinical data, large language models can be used to process the data, removing redundant information, extracting key textual descriptions, and then using text encoders in large models such as Qwen-7B-Medical and PubMedBERT-large to encode the key textual descriptions into feature vectors of the target dimension (denoted as d3). For numerical clinical data, all possible values ​​are categorized into different intervals or types according to medical principles, with each interval or type corresponding to a category label. For example, positive and negative IDH (isocitric dehydrogenase) results correspond to 0 and 1 labels, respectively. More complex age groups can be divided into 5 categories: 0-12 year old children (0 label), 1-17 year old adolescents (1 label), 18-40 year old young adults (2 label), 41-65 year old middle-aged adults (3 label), and 65 year old or older adults (4 label). These labels are then mapped to 0 and 1 codes. For example, negative IDH is [0,1] and middle-aged adults are [0,0,0,1,0]. To better align with descriptive clinical data, the coded data can be interpolated to the target dimension based on nearest neighbor to obtain preprocessed numerical clinical data.

[0061] By preprocessing text-based descriptive clinical data and numerical clinical data in different ways, key features of the corresponding types of clinical data can be extracted, which facilitates subsequent processing.

[0062] Figure 3 This is a flowchart illustrating the training process of the neoadjuvant chemotherapy response prediction model in this application embodiment. For example... Figure 3 As shown, the training process for a neoadjuvant chemotherapy response prediction model, which includes a clinical feature extraction module, a feature fusion module, and a combined prediction module, includes: Step 310: Obtain the sample dataset. Each sample dataset includes CT samples, WSI samples, clinical data samples, and neoadjuvant chemotherapy response annotations.

[0063] CT samples, WSI samples, and clinical data samples are all preprocessed according to the above-mentioned preprocessing methods. The neoadjuvant chemotherapy response labeling represents the actual data on whether the pathology has completely resolved after neoadjuvant chemotherapy.

[0064] Step 320: Based on the sample dataset, keep the parameters of the initialized feature fusion module and the initialized joint prediction module fixed, and adjust the parameters of the initialized clinical feature extraction module to obtain the trained clinical feature extraction module.

[0065] The neoadjuvant chemotherapy response prediction model can be trained in two phases. In the first phase, after initializing the parameters of the neoadjuvant chemotherapy response prediction model, the parameters of the initialized feature fusion module and the initialized joint prediction module are kept fixed. CT samples, WSI samples, and clinical data samples are input into the neoadjuvant chemotherapy response prediction model. Based on the clinical prediction results and neoadjuvant chemotherapy response annotations output by the neoadjuvant chemotherapy response prediction model, the parameters of the initialized clinical feature extraction module are adjusted. This process is iteratively executed, and training ends when the end condition of the first phase is met, resulting in the trained clinical feature extraction module. The end condition of the first phase training may include convergence of the loss function value or convergence of the parameters of the clinical feature extraction module.

[0066] In some embodiments of this application, based on a sample dataset, the parameters of the initial feature fusion module and the initial joint prediction module are kept fixed, and the parameters of the initial clinical feature extraction module are adjusted to obtain a trained clinical feature extraction module, including: Feature extraction is performed on CT samples to obtain sample CT features, and feature extraction is performed on WSI samples to obtain sample WSI features; The clinical feature extraction module is initialized to extract features from the clinical data samples, thereby obtaining the clinical features of the samples. The sample CT features and sample WSI features are fused by the initialized feature fusion module to obtain the sample image fusion features; The sample clinical features are processed by self-attention through the initial joint prediction module to obtain new sample clinical features. Based on the sample image fusion features and the new sample clinical features, the multimodal fusion prediction results and clinical prediction results are determined. The first cross-entropy loss function value of the clinical prediction result and the neoadjuvant chemotherapy response label is determined, and the orthogonal projection loss function value of the new sample clinical features and sample image fusion features is determined. Based on the first cross-entropy loss function value and the orthogonal projection loss function value, the parameters of the initialized clinical feature extraction module are adjusted to obtain the trained clinical feature extraction module.

[0067] When training the neoadjuvant chemotherapy response prediction model, a multi-instance learning approach was used.

[0068] The feature extraction and fusion of CT samples, WSI samples, and clinical data samples are the same as those of CT data, WSI data, and clinical data in the above embodiments, and will not be repeated here.

[0069] The multimodal fusion prediction results are obtained by comprehensively predicting the fusion features of the sample influence and the new clinical features of the sample. The clinical prediction results are obtained by predicting the new clinical features of the sample.

[0070] like Figure 2 As shown, a new sample clinical feature Zb is mapped to a sample clinical feature Zc of the same dimension [h, 2k] as the image fusion feature through a linear layer. The sample clinical feature Zc and the sample image fusion feature I are then concatenated to obtain a sample multimodal fusion feature F=cat(I, Zc) of size [h, 4k]. The sample multimodal fusion feature is then processed through a flattening and linear layer to obtain the multimodal fusion prediction result Fp[1,1]. The sample clinical feature Zc is then mapped to the clinical prediction result Zp[1*1] through a flattening and linear layer. The flattening layer is a special layer that transforms multidimensional data into linear data for further processing.

[0071] The loss function for the first stage is expressed as follows:

[0072]

[0073] in, This represents the loss function for the first stage. Indicates information about clinical prediction results and neoadjuvant chemotherapy response labeling The first cross-entropy loss function, Indicates clinical characteristics of the target fusion Features fused with sample images The orthogonal projection loss function, where [h, 2k] represents the clinical features to be fused. Features of fusion with sample images Dimensions.

[0074] The loss function value for the first stage is calculated according to the above formula. Based on this loss function value, the parameters of the initialized clinical feature extraction model are adjusted. The process of performing feature extraction and adjusting the parameters of the clinical feature extraction module is iterated until the training termination condition of the first stage is met, and the training ends, resulting in the trained clinical feature extraction module.

[0075] By applying cross-entropy loss constraints based on clinical prediction results, the features obtained by the clinical feature extraction module are ensured to be relevant to the response to neoadjuvant chemotherapy, thus making more effective use of clinical features. By using orthogonal projection loss between sample clinical features and sample image fusion features, the complementarity between the two types of features can be increased, thereby further improving the accuracy of the prediction results of the trained model.

[0076] Step 330: Based on the sample dataset, keep the parameters of the trained clinical feature extraction module fixed, and adjust the parameters of the initialized feature fusion module and the initialized joint prediction module to obtain the trained neoadjuvant chemotherapy response prediction model.

[0077] In the second stage, the parameters of the trained clinical feature extraction module are kept fixed. CT samples, WSI samples, and clinical data samples are input into the neoadjuvant chemotherapy response prediction model. Based on the multimodal fusion prediction result Fp and neoadjuvant chemotherapy response labels output by the neoadjuvant chemotherapy response prediction model, the parameters of the initialized feature fusion module and the initialized joint prediction model are optimized using the binary cross-entropy loss function BCE(Fp,Label). This process is iteratively executed, and training ends when the training termination condition of the second stage is met. This results in the trained feature fusion module and the trained joint prediction module. The trained clinical feature extraction module, the trained feature fusion module, and the trained joint prediction module constitute the trained neoadjuvant chemotherapy response prediction model. The training termination condition of the second stage may include the convergence of the loss function value or the convergence of the parameters of the feature fusion module and the joint prediction module.

[0078] The loss function for the second stage is expressed as follows:

[0079] in, This indicates the results of multimodal fusion prediction. and neoadjuvant chemotherapy response labeling The second cross-entropy loss function.

[0080] During model training, an Adam optimizer with an initial learning rate of 2e-4 can be used, along with L2 regularization with weight decay of 1e-5. Due to the large amount of data per sample, training is performed in batches of 1, meaning one sample is trained at a time, to accommodate batches (sample data) of varying sizes over a maximum of 50 epochs. Weighted sampling and early stopping can enhance the training process and prevent overfitting. During the testing phase, a confidence threshold is determined based on the labeling of each sample and the confidence level of the multimodal fusion prediction result. This threshold facilitates the judgment of the model's output confidence during inference, leading to the final neoadjuvant chemotherapy response prediction result.

[0081] By adjusting the parameters of the clinical feature extraction module during the first training phase and fixing the parameters of the clinical feature extraction module during the second training phase while adjusting the parameters of the feature fusion module and the joint prediction module, the interference of irrelevant clinical information on the feature fusion mechanism can be reduced during training. This allows the model to focus more on useful information and improves the prediction accuracy of the trained model.

[0082] It should be noted that the neoadjuvant chemotherapy response prediction method provided in this application embodiment can be executed by a neoadjuvant chemotherapy response prediction device, or a control module in the neoadjuvant chemotherapy response prediction device for executing the neoadjuvant chemotherapy response prediction method. This application embodiment uses the execution of the neoadjuvant chemotherapy response prediction method by a neoadjuvant chemotherapy response prediction device as an example to illustrate the neoadjuvant chemotherapy response prediction method provided in this application embodiment.

[0083] Figure 4 This is a schematic diagram of the structure of a novel adjuvant chemotherapy response prediction device provided in an embodiment of this application, as shown below. Figure 4 As shown, the neoadjuvant chemotherapy response prediction device includes: The data acquisition module 410 is used to acquire preprocessed CT data, WSI data, and clinical data. The first feature extraction module 420 is used to extract features from the CT data to obtain CT features and to extract features from the WSI data to obtain WSI features. The second feature extraction module 430 is used to extract features from the clinical data through the clinical feature extraction module to obtain clinical features; The fusion module 440 is used to fuse the CT features and the WSI features through the feature fusion module to obtain image fusion features; The prediction module 450 is used to determine the neoadjuvant chemotherapy response prediction result based on the image fusion features and the clinical features through the joint prediction module.

[0084] Optionally, the device further includes: The CT preprocessing module is used to adjust the window width and window level of the initial CT data to the target range, and then interpolate the adjusted CT data to the first target size to obtain the preprocessed CT data.

[0085] Optionally, the device further includes: The WSI preprocessing module is used to divide the initial WSI data into image blocks of a second target size, which are then used as the preprocessed WSI data.

[0086] Optionally, the clinical data includes numerical clinical data and text-descriptive clinical data; The device further includes: The text-based clinical data preprocessing module is used to remove redundant information from the text-based descriptive clinical data through a large language model, and to encode the text-based descriptive clinical data after removing the redundant information into a feature vector of the target dimension through a text encoder, which serves as the preprocessed text-based descriptive clinical data. The numerical clinical data preprocessing module is used to determine the labels corresponding to the numerical clinical data, encode the labels, and interpolate the encoded data to the target dimension to obtain the preprocessed numerical clinical data.

[0087] Optionally, the clinical data includes multiple clinical information items, and the clinical feature extraction module includes a graph transformation network; The second feature extraction module is specifically used for: Each piece of clinical information is treated as a node, and a fully connected graph structure is constructed using these nodes. The weights of the edges in the graph structure are initialized to the target values. The clinical features are obtained by extracting features from the graph structure using a graph transformation network.

[0088] Optionally, the fusion module includes: The CT attention processing unit is used to perform self-attention processing on the CT features through the feature fusion module to obtain new CT features. The WSI attention processing unit is used to perform self-attention processing on the WSI features through the feature fusion module to obtain new WSI features. The fusion unit is used to fuse the new CT features and the new WSI features through the feature fusion module to obtain the image fusion features.

[0089] Optionally, the fusion unit is specifically used for: The feature fusion module determines the correlation matrix between the new CT features and the new WSI features; The feature fusion module determines the product of the new CT feature and the correlation matrix to obtain the CT feature fused with WSI features; The WSI feature with fused CT features is obtained by determining the product of the new WSI feature and the correlation matrix through the feature fusion module. The image fusion feature is obtained by stitching together the CT feature of the fused WSI feature and the WSI feature of the fused CT feature using the feature fusion module.

[0090] Optionally, the prediction module is specifically used for: New clinical features are obtained by performing self-attention processing on the clinical features through a joint prediction module. The new clinical features and the image fusion features are spliced ​​together by the joint prediction module to obtain multimodal fusion features; The neoadjuvant chemotherapy response prediction result is obtained by using the joint prediction module to predict the neoadjuvant chemotherapy response based on the multimodal fusion features.

[0091] Optionally, the training process for the neoadjuvant chemotherapy response prediction model, which includes the clinical feature extraction module, the feature fusion module, and the joint prediction module, includes: Obtain a sample dataset, wherein each sample data in the sample dataset includes CT samples, WSI samples, clinical data samples, and neoadjuvant chemotherapy response annotations; Based on the sample dataset, keep the parameters of the initial feature fusion module and the initial joint prediction module fixed, and adjust the parameters of the initial clinical feature extraction module to obtain the trained clinical feature extraction module. Based on the sample dataset, keeping the parameters of the trained clinical feature extraction module fixed, the parameters of the initialized feature fusion module and the initialized joint prediction module are adjusted to obtain the trained neoadjuvant chemotherapy response prediction model.

[0092] Optionally, based on the sample dataset, keeping the parameters of the initialized feature fusion module and the initialized joint prediction module fixed, and adjusting the parameters of the initialized clinical feature extraction module to obtain the trained clinical feature extraction module, includes: Feature extraction is performed on the CT samples to obtain sample CT features, and feature extraction is performed on the WSI samples to obtain sample WSI features; The clinical data sample is subjected to feature extraction by the initialized clinical feature extraction module to obtain the clinical features of the sample; The sample CT features and sample WSI features are fused by the initialized feature fusion module to obtain the sample image fusion features; The sample clinical features are processed by self-attention through the initial joint prediction module to obtain new sample clinical features. Based on the sample image fusion features and the new sample clinical features, the multimodal fusion prediction result and the clinical prediction result are determined. The first cross-entropy loss function value of the clinical prediction result and the neoadjuvant chemotherapy response label is determined, and the orthogonal projection loss function value of the new sample clinical features and the sample image fusion features is determined. Based on the first cross-entropy loss function value and the orthogonal projection loss function value, the parameters of the initialized clinical feature extraction module are adjusted to obtain the trained clinical feature extraction module.

[0093] The neoadjuvant chemotherapy response prediction device in this application embodiment can be a device, or a component, integrated circuit, or chip in a terminal. The device can be a mobile electronic device or a non-mobile electronic device. For example, mobile electronic devices can be mobile phones, tablets, laptops, PDAs, in-vehicle electronic devices, wearable devices, ultra-mobile personal computers (UMPCs), netbooks, or personal digital assistants (PDAs), etc., while non-mobile electronic devices can be servers, network-attached storage (NAS), personal computers (PCs), televisions (TVs), ATMs, or self-service machines, etc. This application embodiment does not impose specific limitations.

[0094] The neoadjuvant chemotherapy response prediction device in this application embodiment can be a device with an operating system. This operating system can be Android, iOS, or other possible operating systems; this application embodiment does not specifically limit it.

[0095] The neoadjuvant chemotherapy response prediction device provided in this application embodiment can achieve... Figures 1 to 3 The various processes implemented in the method implementation examples will not be described again here to avoid repetition.

[0096] The neoadjuvant chemotherapy response prediction device provided in this application extracts features from CT data to obtain CT features, extracts features from WSI data to obtain WSI features, extracts features from clinical data to obtain clinical features, and fuses CT features and WSI features to obtain image fusion features. Based on the image fusion features and clinical features, the neoadjuvant chemotherapy response prediction result is determined. It can be seen that when performing neoadjuvant chemotherapy response prediction, CT data, WSI data and clinical data are fully utilized. The data can complement each other and make up for the deficiencies of single-dimensional data, thereby making the input features for neoadjuvant chemotherapy response prediction more comprehensive and improving the accuracy of neoadjuvant chemotherapy response prediction results.

[0097] Optionally, embodiments of this application also provide an electronic device, including a processor, a memory, and a program or instructions stored in the memory and executable on the processor. When the program or instructions are executed by the processor 110, they implement the various processes of the above-described neoadjuvant chemotherapy response prediction method embodiments and achieve the same technical effects. To avoid repetition, they will not be described again here.

[0098] It should be noted that the electronic devices in the embodiments of this application include the mobile electronic devices and non-mobile electronic devices described above.

[0099] Figure 5 A schematic diagram of the hardware structure of an electronic device to implement an embodiment of this application. The electronic device 500 includes, but is not limited to, components such as: a radio frequency unit 501, a network module 502, an audio output unit 503, an input unit 504, a sensor 505, a display unit 506, a user input unit 507, an interface unit 508, a memory 509, and a processor 510. The input unit 504 may include an image processor 5041 and a microphone 5042; the display unit 506 may include a display panel 5061; and the user input unit 507 may include a touch panel 5071 and other input devices (such as a keyboard) 5072.

[0100] Those skilled in the art will understand that the electronic device 500 may also include a power supply (such as a battery) for supplying power to various components. The power supply may be logically connected to the processor 510 through a power management system, thereby enabling functions such as managing charging, discharging, and power consumption through the power management system. Figure 5 The electronic device structure shown does not constitute a limitation on the electronic device. The electronic device may include more or fewer components than shown, or combine certain components, or have different component arrangements, which will not be elaborated here. The memory 509 stores a program or instruction that can run on the processor. When the program or instruction is executed by the processor 510, it implements the various processes of the above-described neoadjuvant chemotherapy response prediction method embodiment and can achieve the same technical effect. To avoid repetition, it will not be described again here.

[0101] This application also provides a readable storage medium storing a program or instructions. When the program or instructions are executed by a processor, they implement the various processes of the above-described neoadjuvant chemotherapy response prediction method embodiments and achieve the same technical effect. To avoid repetition, they will not be described again here.

[0102] The processor is the processor in the electronic device described in the above embodiments. The readable storage medium includes computer-readable storage media, such as computer read-only memory (ROM), random access memory (RAM), magnetic disk, or optical disk.

[0103] This application embodiment also provides a chip, which includes a processor and a communication interface. The communication interface is coupled to the processor. The processor is used to run programs or instructions to implement the various processes of the above-described neoadjuvant chemotherapy response prediction method embodiment, and can achieve the same technical effect. To avoid repetition, it will not be described again here.

[0104] It should be understood that the chip mentioned in the embodiments of this application may also be referred to as a system-on-a-chip, system chip, chip system, or system-on-a-chip, etc.

[0105] It should be noted that, in this document, the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such a process, method, article, or apparatus. Without further limitations, an element defined by the phrase "comprising one..." does not exclude the presence of other identical elements in the process, method, article, or apparatus that includes that element. Furthermore, it should be noted that the scope of the methods and apparatuses in the embodiments of this application is not limited to performing functions in the order shown or discussed, but may also include performing functions substantially simultaneously or in the reverse order, depending on the functions involved. For example, the described methods may be performed in a different order than described, and various steps may be added, omitted, or combined. Additionally, features described with reference to certain examples may be combined in other examples.

[0106] Through the above description of the embodiments, those skilled in the art can clearly understand that the methods of the above embodiments can be implemented by means of software plus necessary general-purpose hardware platforms. Of course, they can also be implemented by hardware, but in many cases the former is a better implementation method. Based on this understanding, the technical solution of this application, in essence, or the part that contributes to the prior art, can be embodied in the form of a software product. This computer software product is stored in a storage medium (such as ROM / RAM, magnetic disk, optical disk) and includes several instructions to cause a terminal (which may be a mobile phone, computer, server, air conditioner, or network device, etc.) to execute the methods described in the various embodiments of this application.

[0107] The embodiments of this application have been described above with reference to the accompanying drawings. However, this application is not limited to the specific embodiments described above. The specific embodiments described above are merely illustrative and not restrictive. Those skilled in the art can make many other forms under the guidance of this application without departing from the spirit and scope of the claims, and all of these forms are within the protection scope of this application.

Claims

1. A method for predicting the response to neoadjuvant chemotherapy, characterized in that, include: Acquire preprocessed CT data, WSI data, and clinical data; Feature extraction is performed on the CT data to obtain CT features, and feature extraction is performed on the WSI data to obtain WSI features; The clinical data is used to extract features through the clinical feature extraction module to obtain clinical features; The CT features and the WSI features are fused by the feature fusion module to obtain image fusion features; The neoadjuvant chemotherapy response prediction result is determined by the joint prediction module based on the image fusion features and the clinical features.

2. The method according to claim 1, characterized in that, The clinical data includes numerical clinical data and textual descriptive clinical data; Before acquiring the preprocessed CT data, WSI data, and clinical data, the following steps are also included: Redundant information in text-based descriptive clinical data is removed by a large language model, and the text-based descriptive clinical data after removing the redundant information is encoded into a feature vector of the target dimension by a text encoder, which serves as the preprocessed text-based descriptive clinical data. The labels corresponding to the numerical clinical data are determined, the labels are encoded, and the encoded data is interpolated to the target dimension to obtain the preprocessed numerical clinical data.

3. The method according to claim 1, characterized in that, The clinical data includes multiple clinical information items, and the clinical feature extraction module includes a graph transformation network; The clinical feature extraction module extracts features from the clinical data to obtain clinical features, including: Each piece of clinical information is treated as a node, and a fully connected graph structure is constructed using these nodes. The weights of the edges in the graph structure are initialized to the target values. The clinical features are obtained by extracting features from the graph structure using a graph transformation network.

4. The method according to any one of claims 1-3, characterized in that, The process of fusing the CT features and the WSI features through a feature fusion module to obtain image fusion features includes: The feature fusion module performs self-attention processing on the CT features to obtain new CT features; The WSI features are processed by the feature fusion module to obtain new WSI features; The new CT features and the new WSI features are fused by the feature fusion module to obtain the image fusion features.

5. The method according to claim 4, characterized in that, The process of fusing the new CT features and the new WSI features through the feature fusion module to obtain the image fusion features includes: The feature fusion module determines the correlation matrix between the new CT features and the new WSI features; The feature fusion module determines the product of the new CT feature and the correlation matrix to obtain the CT feature fused with WSI features; The WSI feature with fused CT features is obtained by determining the product of the new WSI feature and the correlation matrix through the feature fusion module. The image fusion feature is obtained by stitching together the CT feature of the fused WSI feature and the WSI feature of the fused CT feature using the feature fusion module.

6. The method according to any one of claims 1-3, characterized in that, The step of determining the neoadjuvant chemotherapy response prediction result through the joint prediction module based on the image fusion features and the clinical features includes: New clinical features are obtained by performing self-attention processing on the clinical features through a joint prediction module. The new clinical features and the image fusion features are spliced ​​together by the joint prediction module to obtain multimodal fusion features; The neoadjuvant chemotherapy response prediction result is obtained by using the joint prediction module to predict the neoadjuvant chemotherapy response based on the multimodal fusion features.

7. The method according to any one of claims 1-3, characterized in that, The training process for the neoadjuvant chemotherapy response prediction model, which includes the clinical feature extraction module, the feature fusion module, and the combined prediction module, includes: Obtain a sample dataset, wherein each sample data in the sample dataset includes CT samples, WSI samples, clinical data samples, and neoadjuvant chemotherapy response annotations; Based on the sample dataset, keep the parameters of the initial feature fusion module and the initial joint prediction module fixed, and adjust the parameters of the initial clinical feature extraction module to obtain the trained clinical feature extraction module. Based on the sample dataset, keeping the parameters of the trained clinical feature extraction module fixed, the parameters of the initialized feature fusion module and the initialized joint prediction module are adjusted to obtain the trained neoadjuvant chemotherapy response prediction model.

8. The method according to claim 7, characterized in that, Based on the sample dataset, keeping the parameters of the initialized feature fusion module and the initialized joint prediction module fixed, the parameters of the initialized clinical feature extraction module are adjusted to obtain the trained clinical feature extraction module, including: Feature extraction is performed on the CT samples to obtain sample CT features, and feature extraction is performed on the WSI samples to obtain sample WSI features; The clinical data sample is subjected to feature extraction by the initialized clinical feature extraction module to obtain the clinical features of the sample; The sample CT features and sample WSI features are fused by the initialized feature fusion module to obtain the sample image fusion features; The sample clinical features are processed by self-attention through the initial joint prediction module to obtain new sample clinical features. Based on the sample image fusion features and the new sample clinical features, the multimodal fusion prediction result and the clinical prediction result are determined. The first cross-entropy loss function value of the clinical prediction result and the neoadjuvant chemotherapy response label is determined, and the orthogonal projection loss function value of the new sample clinical features and the sample image fusion features is determined. Based on the first cross-entropy loss function value and the orthogonal projection loss function value, the parameters of the initialized clinical feature extraction module are adjusted to obtain the trained clinical feature extraction module.

9. A novel adjuvant chemotherapy response prediction device, characterized in that, include: The data acquisition module is used to acquire preprocessed CT data, WSI data, and clinical data. The first feature extraction module is used to extract features from the CT data to obtain CT features and to extract features from the WSI data to obtain WSI features. The second feature extraction module is used to extract features from the clinical data through the clinical feature extraction module to obtain clinical features; The fusion module is used to fuse the CT features and the WSI features through the feature fusion module to obtain image fusion features; The prediction module is used to determine the neoadjuvant chemotherapy response prediction result based on the image fusion features and the clinical features through the joint prediction module.

10. An electronic device, characterized in that, It includes a processor, a memory, and a program or instructions stored in the memory and executable on the processor, wherein the program or instructions, when executed by the processor, implement the steps of the neoadjuvant chemotherapy response prediction method as described in any one of claims 1-8.

11. A readable storage medium, characterized in that, The readable storage medium stores a program or instructions that, when executed by a processor, implement the steps of the neoadjuvant chemotherapy response prediction method as described in any one of claims 1-8.