Cancer survival prediction system based on multi-modal feature decoupling, recombination and fusion
Through the feature decoupling method of combining regional cross attention and full connection, combined with the feature fusion of random feature recombination and hybrid expert model, the problem of feature extraction and fusion in multimodal cancer survival prediction is solved, and the prediction performance and model reliability are improved.
Patent Information
- Application Number
- CN202510410568.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-02
- Publication Date
- 2025-07-25
AI Technical Summary
The prior art is difficult to effectively extract and fuse the characteristics of MRI images and pathological slice data in multimodal cancer survival prediction, resulting in insufficient prediction performance, especially ignoring the interaction information between and within modals, affecting the overall performance of the model.
The feature decoupling method combined with regional cross attention and full connection is adopted, and the feature fusion strategy of random feature recombination and hybrid expert model is used to improve feature quality and dynamic fusion capabilities through feature decoupling modules and recombination modules. The image and pathological features are extracted using 3D ResNet50 and Transformer networks. The random feature recombination reduces the dependence of fixed feature combinations, and the dynamic fusion expert network enhances nonlinear expression.
The C-index indicators of the cancer survival prediction model have been significantly improved, the consistency between the prediction results and the actual situation has been improved, and the reliability and effectiveness of the model for different data sets and clinical scenarios.
Smart Images

Figure CN120376164A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of medical data processing, and in particular relates to a cancer survival prediction system based on multimodal feature decoupling, recombination and fusion. Background Art
[0002] Cancer seriously threatens human health and life. Accurate cancer survival prediction is crucial to optimizing treatment plans and improving the quality of life of patients. In the era of big data, medical data has shown explosive growth and diversified development. The innovation of imaging technology has enabled high-precision imaging data such as MRI and CT to fully capture the morphology, location and surrounding tissue information of cancer; the refined development of pathological diagnosis technology has enabled pathological images to reveal the characteristics of cancer at the cellular level. The large-scale collection and systematic organization of these imaging and pathological data have created favorable conditions for the in-depth application of machine learning and deep learning algorithms in the medical field. They can explore the potential laws in the data and assist doctors in making more accurate diagnoses and prognoses. However, the complexity and heterogeneity of multimodal data also bring huge challenges.
[0003] MRI images can clearly present the location, size, morphology of the tumor and its spatial relationship with important structures such as surrounding blood vessels. Through different imaging sequences, such as T1-weighted images, T2-weighted images and dynamic enhanced scanning, information such as the blood supply characteristics and tissue composition of the tumor can be obtained. In dynamic enhanced scanning, tumor lesions usually show obvious enhancement in the arterial phase, which helps to detect and diagnose tumors early, and can also provide a reference for evaluating the invasiveness of tumors; for pathological sections of cancer tissues, the morphology of cancer cells, the size and shape of the nucleus, the arrangement of cells, and whether there is capsule invasion, vascular invasion and other important pathological characteristics can be observed under a microscope after staining. This microscopic information plays a decisive role in determining the pathological type and degree of differentiation of cancer and judging the prognosis. However, there are significant differences and complementarities between MRI data and pathological imaging data of tumors. MRI data shows the tumor situation from a macroscopic anatomical level, while pathological images focus on the microscopic cell tissue level. The two are different in data format, resolution, feature expression, etc., and there are many challenges in fusing them for cancer survival prediction, such as how to accurately extract effective features of the two modal data, how to align data with different resolutions, and how to build a reasonable model to integrate these multi-source heterogeneous information. Therefore, how to better solve the above problems has become a key problem in the current cancer survival prediction research, and innovative methods and technologies are urgently needed to break through.
[0004] Among the existing methods for multimodal cancer survival prediction, the mainstream approach is to first extract features from multimodal data, then decouple the extracted features, and then fuse the decoupled data to generate the final prediction results.
[0005] Decoupling methods can be roughly divided into two categories: one focuses on only extracting and utilizing modality-shared or modality-specific features, but this method is prone to causing the model to ignore the potential interaction information between decoupled features, thereby limiting the overall performance of the model; the other attempts to extract and use all decoupled features, but there are still many difficulties in effectively extracting decoupled features. For example, when obtaining modality-shared features, a common approach is to concatenate modality features and then send them into a fully connected layer or calculate cross-attention. However, these methods tend to consider the interaction between modalities and lack sufficient learning of the interaction within modalities. Ignoring the interaction within modalities may result in insufficient representation of decoupled features.
[0006] In terms of decoupled feature fusion, current fusion methods mainly include concatenation, attention mechanism-based, or mixture of experts (MoE)-based methods. Some studies directly concatenate modality-shared and modality-specific features and then input them into a fully connected layer or an attention layer for prediction. However, this simple fusion method makes the model overly dependent on fixed feature combinations and lacks the ability to dynamically fuse decoupled features, making it difficult to flexibly fuse different decoupled features according to the importance of features. In the method using the MoE model, each expert network usually processes a single decoupled feature separately. Although a certain degree of dynamic fusion is achieved, there is a problem of information occlusion, that is, each expert network only considers the information of a specific decoupled feature and lacks comprehensive consideration of the useful information between decoupled features, which affects the learning of the relationship between decoupled features by the model and the improvement of the overall performance.
[0007] When evaluating the performance of a survival analysis model, the C-index metric is an important and commonly used measure. The C-index is mainly used to measure the consistency between the patient risk scores predicted by the model and the actual survival situation ranking. If the C-index value of the model is higher, it indicates that its prediction result is in better agreement with the actual situation, that is, it can more accurately stratify the survival risks of patients. For example, in the comparison of multiple groups of patients, the C-index can reflect the ability of the model to distinguish high-risk and low-risk patients, effectively evaluate the reliability and effectiveness of the model in different datasets and clinical scenarios, and provide a key basis for the improvement and optimization of the model. Summary of the Invention
[0008] In view of the above, the object of the present invention is to provide a cancer survival prediction system based on multi-modal feature decoupling, recombination, and fusion to effectively solve the technical problems existing in the multi-modal cancer survival prediction of existing methods, significantly improve the prediction performance, and provide more powerful decision-making support for cancer treatment.
[0009] To achieve the above invention objective, an embodiment provides a cancer survival prediction system based on multi-modal feature decoupling, recombination, and fusion, including:
[0010] A data acquisition and preprocessing unit, which is used to acquire multi-modal cancer data including image data and pathological section data and perform preprocessing;
[0011] A cancer survival prediction unit, which is used to perform survival prediction based on the preprocessed multi-modal data by using a cancer survival prediction model including a feature extraction module, a feature decoupling module, a feature recombination and fusion module, and a survival prediction module. Specifically, the preprocessed image data and pathological section data are respectively extracted by the feature extraction module to obtain an image feature representation and a pathological feature representation. These two types of feature representations are respectively subjected to feature interaction and decoupling by the feature decoupling module using regional cross-attention and a fully connected layer to obtain their respective specific-modal features, exploratory-modal features, and shared-modal features. These four types of modal features are recombined by the feature recombination and fusion module using a random feature recombination strategy and then fused by the feature fusion to obtain a fused feature. This fused feature is predicted by the survival prediction module to output a survival risk probability.
[0012] Preferably, the feature extraction module includes a 3D ResNet50 network, a global average pooling layer, a Transformer network, a ResNet network, and a fully connected network;
[0013] The preprocessed image data and pathological section data are respectively extracted by the feature extraction module to obtain an image feature representation and a pathological feature representation, including:
[0014] The preprocessed image data extracts the spatial features of cancer through a 3D ResNet50 network. After the spatial features are dimension-reduced by the global average pooling layer, the image feature representation is extracted through a Transformer network;
[0015] The preprocessed pathological section data extracts multiple instance features through a ResNet network. After the instance features are unified in feature dimension through a fully connected network, the pathological feature representation is extracted through a Transformer network.
[0016] Preferably, the feature decoupling module includes a specific-modal encoder 1 and a specific-modal encoder 2. These two types of feature representations are respectively encoded separately by the specific-modal encoder 1 and the specific-modal encoder 2 to obtain their respective specific-modal features. Each specific-modal encoder uses an MLP.
[0017] Preferably, the feature decoupling module further includes an exploratory-modal encoder and a shared-modal encoder with the same structure. In these two encoders, regional cross-attention and a fully connected layer are used for feature interaction and decoupling to obtain exploratory-modal features and shared-modal features, including:
[0018] These two types of feature representations are simultaneously input into the exploration modality encoder and the shared modality encoder. First, the two types of feature representations are respectively embedded through two fully connected layers to obtain new feature representations. Then, the two types of new feature representations are concatenated in different orders to obtain a concatenation result. Next, the two concatenation results are multiplied matrix by matrix to obtain an attention matrix. Then, sub-matrices of two different regions are obtained from the attention matrix and the attention results are calculated using the Softmax function. Finally, the attention results are averaged and output to obtain the shared modality features or the exploration modality features;
[0019] The mean squared error is used as a measure of the distance between the decoupled different modality features to train the feature decoupling module.
[0020] Preferably, the feature recombination and fusion module includes a random feature recombination sub-module. The four types of modality features are recombined using a random feature recombination strategy in the random feature recombination sub-module, which specifically includes:
[0021] First, a segmentation number array is set, and each decoupled modality feature is split into several equal sub-features. Then, the sub-features from the decoupled different modality features are combined. Each time during forward propagation, the segmentation number is randomly selected to recombine all the sub-features corresponding to the segmentation number. Finally, the recombined feature segments are connected along the channel dimension and used as the finally recombined feature.
[0022] Preferably, the feature recombination and fusion module further includes a feature fusion sub-module. The recombined feature undergoes feature fusion in the feature fusion sub-module to obtain a fusion feature, which specifically includes:
[0023] The feature fusion sub-module includes a mixture of experts model for dynamic fusion, which adopts a dense mixture of experts structure, activates all the expert networks and the gating networks. Each expert network extracts feature mappings based on the input recombined feature. The gating network generates dynamic weights for each expert network according to the input recombined feature, and the output features of the expert networks are weighted and concatenated based on the dynamic weights to obtain the fusion feature.
[0024] Preferably, each expert network consists of a fully connected network containing two Linear-ReLU layers with a Dropout mechanism added in the middle.
[0025] Preferably, the cancer survival prediction model is optimized in parameters before being applied. During optimization, the mean squared error and the negative log-likelihood are used as loss functions, and the concordance index C-Index is used to measure the prediction performance of the model.
[0026] To achieve the above invention objective, the embodiment also provides a computing device, including a memory and one or more processors. An executable code is stored in the memory. When the one or more processors execute the executable code, it is used to implement the survival prediction method using the above cancer survival prediction system based on multi-modal feature decoupling, recombination and fusion, including the following steps:
[0027] Use the data acquisition and preprocessing unit to acquire multi-modal cancer data including image data and pathological section data and perform preprocessing;
[0028] Use the cancer survival prediction unit to perform survival prediction based on the preprocessed multi-modal data.
[0029] To achieve the above invention objective, the embodiment also provides a computer-readable storage medium, on which a program is stored. When the program is executed by a processor, it implements the survival prediction method using the above cancer survival prediction system based on multi-modal feature decoupling, recombination and fusion, including the following steps:
[0030] Use the data acquisition and preprocessing unit to acquire multi-modal cancer data including image data and pathological section data and perform preprocessing;
[0031] Use the cancer survival prediction unit to perform survival prediction based on the preprocessed multi-modal data.
[0032] Compared with the prior art, the beneficial effects of the present invention at least include:
[0033] (1) In the feature decoupling module of the present invention, a novel and lightweight regional cross-attention is proposed, which can focus on the relationships between different modal features and within features, thereby improving the quality of specific modal decoupled features and exploratory modal decoupled features;
[0034] (2) In the feature recombination and fusion module of the present invention, feature fusion based on a mixture of experts model after feature recombination is proposed. Among them, feature recombination can reduce the over-reliance of the expert network on fixed feature connections, thereby enhancing the non-linear expression ability of the model. The expert network can analyze the recombined features from multiple aspects and achieve dynamic fusion. The final model effect is better than other existing models. Description of the Drawings
[0035] To more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the following will briefly introduce the drawings required for the description of the embodiments or the prior art. Obviously, the drawings in the following description are only some embodiments of the present invention. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.
[0036] Figure 1It is a schematic diagram of the structure and process of the cancer survival prediction system based on multi-modal feature decoupling, recombination and fusion provided by the embodiment;
[0037] Figure 2 It is a schematic diagram of the structure and process of the cancer survival prediction model provided by the embodiment;
[0038] Figure 3 It is a schematic diagram of feature recombination provided by the embodiment;
[0039] Figure 4 It is a flowchart of the cancer survival prediction method based on multi-modal feature decoupling, recombination and fusion provided by the embodiment. Detailed implementation manners
[0040] To make the objectives, technical solutions and advantages of the present invention clearer, the present invention will be further described in detail below with reference to the accompanying drawings and embodiments. It should be understood that the specific implementation manners described herein are only used to explain the present invention and do not limit the protection scope of the present invention.
[0041] As Figure 1 shown, the embodiment provides a cancer survival prediction system 100 based on multi-modal feature decoupling, recombination and fusion, including: a data acquisition and preprocessing unit 110 and a cancer survival prediction unit 120.
[0042] In the embodiment, the data acquisition and preprocessing unit 110 is used to acquire multi-modal cancer data including image data and pathological section data and perform preprocessing. Specifically, image data (such as MRI data) of the imaging equipment of a certain hospital and pathological image data of the pathology department are obtained to form multi-modal cancer data, and then these data are preprocessed. Specifically, for images such as MRI, there are five types of MRI for each patient, namely A, D, PMX, T1, and T2. First, a professional medical image processing software is used to crop and unify the sizes of the five types of MRI images according to the annotation information of the tumor region (ROI) to eliminate the influence regions irrelevant to the tumor and reduce the calculation amount of the subsequent system. Then, the images are normalized by Z-score to obtain the preprocessed image data. For pathological image data (WSI), the CLAM toolkit is used to segment it into multiple non-overlapping instances (patches with a magnification of ×20 and 256×256 pixels) to obtain the preprocessed pathological section data.
[0043] The cancer survival prediction unit 120 is used to perform survival prediction based on the preprocessed multi-modal data by using a cancer survival prediction model including a feature extraction module, a feature decoupling module, a feature recombination and fusion module, and a survival prediction module. The structure of the cancer survival prediction model is as Figure 2 shown, and specifically includes:
[0044] (1) The feature extraction module is used to extract image features and pathological features from the preprocessed image data and pathological section data. The feature extraction module includes a 3D ResNet50 network, a global average pooling layer, a Transformer network, a ResNet network, and a fully connected network. When extracting features based on this network structure, the preprocessed image data extracts the spatial features of cancer through the 3D ResNet50 network. After the spatial features pass through the global average pooling layer to reduce the features to an appropriate range, the dimensionality-reduced features of the image data are obtained. Then the dimensionality-reduced features are further passed through a Transformer network with 2 layers containing the Nystrom Attention mechanism to extract the image feature representation. I and C1 represent the number of patches and the embedding dimension respectively.
[0045] The preprocessed pathological section data extracts multiple instance features through the ResNet neural network pre-trained on the ImageNet dataset. Since the feature dimensions extracted from different modality data may vary, the instance features are passed through a fully connected network to unify the feature dimensions to obtain features. These features are further passed through a Transformer network with 2 layers containing the Nystrom Attention mechanism to extract the pathological feature representation.
[0046] The total survival period and survival status of the patient during the postoperative follow-up period are used as the final labels. The labels are the values in months and the survival status classified as 0 or 1. 0 represents the samples that are still alive during the follow-up period, and 1 represents the samples that died during the follow-up period.
[0047] In terms of effective feature extraction, the combination of ResNet and Transformer is usually adopted. For MRI data, ResNet can automatically learn features such as the texture, shape, and edge of tumors. Its hierarchical convolutional operations can gradually abstract higher-level semantic information from the original images, such as the general contour of the tumor area and the internal structure differences. When processing pathological image data, ResNet can identify key information such as the nuclear features of tumor cells and the cell arrangement pattern, and provide rich details for subsequent analysis through feature mappings at different levels. Transformer is good at capturing long-range dependencies and global information in the data. In cancer image data processing, it can associate different regions of the image and comprehensively consider the overall spatial layout information. For MRI data, Transformer can integrate image information at different levels to further improve the overall understanding of tumors. In pathological imaging, it can span different cell communities, analyze the interrelationships between cells and the overall architecture of tissues, making up for the deficiencies of convolutional neural networks in global information processing. The combination of the two, with ResNet responsible for feature extraction and Transformer focusing on exploring the relationships between features, complement each other and cooperate, enabling more comprehensive and accurate extraction of the features of tumor MRI and pathological image data, providing a solid foundation for subsequent survival prediction.
[0048] (2) Use the feature decoupling module to decouple the two types of feature representations with large heterogeneity (image feature representation and pathological feature representation). Specifically, regional cross-attention and fully connected layers are used for feature interaction and decoupling into four parts, namely their respective specific modality features V sp1 and V sp2 , explore the modality feature V share and the shared modality feature V explore .
[0049] As Figure 2 shown, the feature decoupling module includes a specific modality encoder 1 and a specific modality encoder 2. These two types of feature representations are separately encoded by the specific modality encoder 1 and the specific modality encoder 2 to obtain their respective specific modality features V sp1 and V sp2 . Each specific modality encoder uses a two-layer MLP network to further reduce the dimensionality of the feature vector and perform feature learning. Each layer of the MLP contains a Linear-ReLU module and a Dropout mechanism is added in the middle to enhance the non-linear expression ability of the network and reduce the risk of overfitting.
[0050] The feature decoupling module also includes an exploration modality encoder and a shared modality encoder with the same structure. In these two encoders, regional cross-attention and fully connected layers are used for feature interaction and decoupling to obtain the exploration modality feature and the shared modality feature, including:
[0051] These two types of features represent V m1 and V m2 which are simultaneously input into the exploration modality encoder and the shared modality encoder. First, embedding operations are performed on the two types of feature representations through two fully connected layers respectively to obtain new feature representations v m1 and v m2 ;
[0052] Then, the two types of new feature representations v m1 and v m2 are concatenated in different orders to obtain concatenation results [v m1 , v m2 , [v m2 , v m1 . Then, the two concatenation results are multiplied matrix-wise to obtain the attention matrix where M m1m2 , M m1m1 , M m2m2 , and M m2m1 represent sub-matrices of different regions. Then, sub-matrices of two different regions are obtained from the attention matrix and the Softmax function is used to calculate the attention results, that is:
[0053]
[0054] Finally, the attention results and are averaged and output to obtain the shared modality features or the exploration modality features
[0055] When training the feature decoupling module, the feature decoupling module is optimized by measuring the distance between the decoupled features. Among them, the distance between the modality-specific features should be as far as possible, the modality-specific features and the remaining decoupled features should have a certain distance, and there should be a certain similarity at the same time. Given the distance measurement function Dis(V i , V j ) and the decoupled modality features {V sp1 , V sp2 , V share , V explore}, the decoupling loss function is:
[0056]
[0057] where the mean squared error is selected as the distance measurement function.
[0058] In the feature decoupling module, for modality-specific features, only fully connected layers are used for extraction. However, for modality-shared features, a simple regional cross-attention method is used for extraction, which can fully consider the intra-modal and inter-modal interactions and improve the quality of the decoupled features.
[0059] (c) The feature recombination and fusion module recombines the four types of modality features using a random feature recombination strategy and then obtains the fused features through feature fusion.
[0060] As Figure 2 shown, the feature recombination and fusion module includes a random feature recombination sub-module. The four types of modality features are recombined using a random feature recombination strategy in the random feature recombination sub-module. Specifically, it includes: first, setting an array of segmentation numbers, splitting each decoupled modality feature into several equal sub-features, then combining the sub-features from different decoupled modality features, randomly selecting the segmentation number for recombination of all sub-features corresponding to the segmentation number during each forward propagation, and finally connecting the recombined feature segments along the channel dimension as the final recombined feature.
[0061] As Figure 3 shown, use V i , i = {1, 2, 3, 4} to represent the names of the four types of decoupled modality features {V sp1 , V sp2 , V share , V explore}. During the feature recombination operation, set an array of segmentation numbers where s t represents the length of the t-th segment in the array. Split each decoupled modality feature into L segments j is the index of the segment. Here, L = C2 / / s t , C2 represents the dimension (total length) of each modality feature, and the symbol / / represents integer division. Then each feature segment is recombined, that is Finally, connect the recombined feature segments along the channel dimension to obtain the final recombined feature
[0062] As Figure 2 shown, the feature recombination and fusion module also includes a feature fusion sub-module. The recombined features are fused in the feature fusion sub-module to obtain the fused features. Specifically, it includes:
[0063] The feature fusion sub-module includes a mixture of experts (MoE) model with dynamic fusion, which adopts a dense mixture of experts structure, activating N (e.g., 4) expert networks {E1, E2, E3, E4} and a gating network G. Each expert network consists of two layers of MLP networks, each layer contains a Linear-ReLU module and a Dropout mechanism is added in the middle, and the gating network consists of a single fully-connected network. Each expert network is based on the reorganized features of the input to extract feature mapping. The gating network G generates the dynamic weights G of each expert network according to the reorganized features of the input k (V fusion ) = Softmax(V fusion ·W gate ), where k ∈ {1,..., N}, representing the index of the expert network, W gate represents the gating weight, and the output features of the expert network are weighted and concatenated based on the dynamic weights to obtain the fused features
[0064] In the feature reorganization and fusion module, a feature reorganization algorithm is adopted to shuffle and reorganize the four decoupled features, reducing the dependence of the neural network on fixed feature combinations and enhancing the generalization ability of the network. At the same time, the reorganized features processed by each expert network contain all decoupled modal features, which to a certain extent enhances the comprehensive consideration of useful information between the decoupled modal features by the expert network.
[0065] (d) Using the survival prediction module to predict the output survival risk probability based on the fused features. In the embodiment, the survival prediction module consists of a simple fully-connected network Linear and an activation function Sigmoid, and the survival risk probability H output = Sigmoid(Linear(V expert )) is predicted based on this survival prediction module.
[0066] In the embodiment, the cancer survival prediction model is optimized before being applied. The optimization process is as follows:
[0067] (1) The pathological image data composed of the pathological examinations of cancer-related patients and through H&E staining technology, and the multi-period cancer site MRI image data composed of magnetic resonance examinations;
[0068] (2) The total survival period and survival status of the patient during the postoperative follow-up period are used as the final annotation. The labels are numerical values in months and survival status classified as 0 and 1. 0 represents the samples that are still alive during the follow-up period, and 1 represents the samples that died during the follow-up period;
[0069] (3) Based on the labeled data, divide the data into a training set and a validation set using five-fold cross-validation;
[0070] (4) Perform certain data transformation and normalization on the pathological images and MRI images in the training set and the validation set, without affecting the overall data size and numerical range distribution;
[0071] (5) Input the data processed in step (4) into the model for classification training, and iteratively update the model parameters for all training samples until convergence.
[0072] During training, use the mean squared error and the negative log-likelihood as the loss function where α is a balancing factor, and the negative log-likelihood
[0073] where, H output ={h1, h2,..., h n} represents the risk probabilities corresponding to n time periods predicted by the model, c represents the censoring status, and the discrete survival function Train the model using five-fold cross-validation. Each fold is trained for 30 epochs. Each time, input a batch of data, calculate the loss and backpropagate to update the model parameters until the training is completed. The model with the best validation effect will be saved during the iteration process. After training, use the concordance index C-Index to measure the prediction performance of the model.
[0074] In practical applications, the cancer survival prediction system of the present invention can predict the survival risk through the patient's MRI images and pathological section images. According to the prediction results, doctors can be guided to improve subsequent treatment plans.
[0075] Based on the same inventive concept, the embodiment also provides a computing device, including a memory and one or more processors. The memory stores executable code. When the one or more processors execute the executable code, it is used to implement the survival prediction method of the above-mentioned cancer survival prediction system based on multi-modal feature decoupling, recombination and fusion. This method uses the above-mentioned survival prediction system, as Figure 4 shown, specifically including the following steps:
[0076] S1, use the data acquisition and preprocessing unit to obtain multi-modal cancer data including image data and pathological section data and perform preprocessing;
[0077] S2, use the cancer survival prediction unit to perform survival prediction based on the preprocessed multi-modal data.
[0078] The computing device provided by the embodiment, at the hardware level, in addition to including a processor and a memory, also includes an internal bus, a network interface, a memory, and other hardware required for other services. The memory is a non-volatile memory. The processor reads the corresponding computer program from the non-volatile memory into the memory and then runs it to implement the survival prediction method of the cancer survival prediction system based on multi-modal feature decoupling, recombination, and fusion described in S1-S2 above. Of course, in addition to the software implementation method, the present invention does not exclude other implementation methods, such as logic devices or a combination of software and hardware, etc. That is to say, the execution subject of the following processing flow is not limited to each logic unit, and can also be hardware or a logic device.
[0079] Based on the same inventive concept, the embodiment also provides a computer-readable storage medium, on which a program is stored. When the program is executed by a processor, it implements the survival prediction method of the cancer survival prediction system based on multi-modal feature decoupling, recombination, and fusion described above. This method uses the above survival prediction system and specifically includes the following steps:
[0080] S1, using the data acquisition and preprocessing unit to acquire multi-modal cancer data including image data and pathological section data and perform preprocessing;
[0081] S2, using the cancer survival prediction unit to perform survival prediction based on the preprocessed multi-modal data.
[0082] In the embodiment, the computer-readable medium includes permanent and non-permanent, removable and non-removable media, and information storage can be implemented by any method or technology. The information can be computer-readable instructions, data structures, program modules, or other data.
[0083] The specific embodiments described above have detailed the technical solutions and beneficial effects of the present invention. It should be understood that the above is only the most preferred embodiment of the present invention and is not used to limit the present invention. Any modifications, supplements, equivalent replacements, etc. made within the scope of the principles of the present invention should be included in the protection scope of the present invention.
Claims
1. A cancer survival prediction system based on multi-modal feature decoupling, recombination and fusion, characterized in that Including: A data acquisition and preprocessing unit, which is used to acquire multi-modal cancer data including image data and pathological section data and perform preprocessing; A cancer survival prediction unit, which is used to perform survival prediction based on the preprocessed multi-modal data using a cancer survival prediction model including a feature extraction module, a feature decoupling module, a feature recombination and fusion module, and a survival prediction module. Specifically, the preprocessed image data and pathological section data are respectively extracted by the feature extraction module to obtain an image feature representation and a pathological feature representation. These two types of feature representations pass through the feature decoupling module to perform feature interaction and decoupling using regional cross-attention and a fully connected layer to obtain their respective specific-modal features, exploratory-modal features, and shared-modal features. These four types of modal features pass through the feature recombination and fusion module to be recombined using a random feature recombination strategy and then through feature fusion to obtain a fusion feature. This fusion feature passes through the survival prediction module to predict and output a survival risk probability.
2. The cancer survival prediction system based on multi-modal feature decoupling, recombination and fusion according to claim 1, characterized in that The feature extraction module includes a 3D ResNet50 network, a global average pooling layer, a Transformer network, a ResNet network, and a fully connected network; The preprocessed image data and pathological section data are respectively extracted by the feature extraction module to obtain an image feature representation and a pathological feature representation, including: The preprocessed image data extracts the spatial features of cancer through a 3D ResNet50 network. This spatial feature is dimension-reduced through the global average pooling layer and then the image feature representation is extracted through the Transformer network; The preprocessed pathological section data extracts multiple instance features through a ResNet network. This instance feature unifies the feature dimension through a fully connected network and then the pathological feature representation is extracted through the Transformer network.
3. The cancer survival prediction system based on multi-modal feature decoupling, recombination and fusion according to claim 1, wherein The feature decoupling module includes a specific-modal encoder 1 and a specific-modal encoder 2. These two types of feature representations are respectively encoded by the specific-modal encoder 1 and the specific-modal encoder 2 to obtain their respective specific-modal features. Each specific-modal encoder uses an MLP.
4. The cancer survival prediction system based on multi-modal feature decoupling, recombination and fusion according to claim 1, characterized in that, The feature decoupling module also includes an exploratory-modal encoder and a shared-modal encoder with the same structure. In these two encoders, regional cross-attention and a fully connected layer are used to perform feature interaction and decoupling to obtain exploratory-modal features and shared-modal features, including: These two types of feature representations are simultaneously input into the exploratory-modal encoder and the shared-modal encoder. First, embedding operations are respectively performed on the two types of feature representations through two fully connected layers to obtain new feature representations. Then, the two types of new feature representations are concatenated in different orders to obtain a concatenation result. Then, the two concatenation results are multiplied matrix-wise to obtain an attention matrix. Then, sub-matrices of two different regions are obtained from the attention matrix and the Softmax function is used to calculate the attention result. Finally, the attention result is averaged and output to obtain a shared-modal feature or an exploratory-modal feature; The mean squared error is used as a measure of the distance between the decoupled different-modal features to train the feature decoupling module.
5. The cancer survival prediction system based on multi-modal feature decoupling, recombination and fusion according to claim 1, wherein The feature recombination and fusion module includes a random feature recombination sub-module. The four types of modal features are recombined using a random feature recombination strategy in the random feature recombination sub-module, specifically including: First, set an array of segmentation numbers, split each decoupled modal feature into several equal sub-features, then combine the sub-features from different decoupled modal features. During each forward propagation, randomly select the segmentation number to recombine all the sub-features corresponding to the segmentation number, and finally connect the recombined feature segments along the channel dimension as the finally recombined feature.
6. The cancer survival prediction system based on multi-modal feature decoupling, recombination and fusion according to claim 1, characterized in that, The feature recombination and fusion module also includes a feature fusion sub-module. The recombined feature undergoes feature fusion in the feature fusion sub-module to obtain a fused feature, specifically including: The feature fusion sub-module includes a mixture-of-experts model for dynamic fusion, which adopts a dense mixture-of-experts structure, activates all expert networks and gating networks. Each expert network extracts feature maps based on the input recombined feature, and the gating network generates dynamic weights for each expert network according to the input recombined feature, and performs weighted concatenation on the output features of the expert networks based on the dynamic weights to obtain the fused feature.
7. The cancer survival prediction system based on multi-modal feature decoupling, recombination and fusion according to claim 6, wherein Each expert network consists of a fully connected network containing two Linear-ReLU layers with a Dropout mechanism added in the middle.
8. The cancer survival prediction system based on multi-modal feature decoupling, recombination and fusion according to claim 1, wherein Before being applied, the cancer survival prediction model undergoes parameter optimization. During optimization, the mean squared error and negative log-likelihood are used as loss functions, and the concordance index C-Index is used to measure the prediction performance of the model.
9. A computing device, comprising a memory and one or more processors, wherein executable code is stored in the memory, characterized in that, When the one or more processors execute the executable code, it is used to implement a survival prediction method using the cancer survival prediction system based on multi-modal feature decoupling, recombination and fusion according to any one of claims 1-8, including the following steps: Use the data acquisition and preprocessing unit to acquire multi-modal cancer data including image data and pathological section data and perform preprocessing; Use the cancer survival prediction unit to perform survival prediction based on the preprocessed multi-modal data.
10. A computer-readable storage medium, on which a program is stored. When the program is executed by a processor, it implements a survival prediction method using the cancer survival prediction system based on multi-modal feature decoupling, recombination and fusion according to any one of claims 1-8, including the following steps: Use the data acquisition and preprocessing unit to acquire multi-modal cancer data including image data and pathological section data and perform preprocessing; Use the cancer survival prediction unit to perform survival prediction based on the preprocessed multi-modal data.
Citation Information
Cited By
Acute myelogenous leukemia prognosis prediction method and device based on multi-omics fusion
CN120895214A
Lung cancer lifetime prediction system based on prognosis factor multi-data fusion
CN121812154A