Multi-modal survival prediction method and device
Through the multimodal survival prediction method, the preset survival prediction model is used to extract, interact and collaborate the pathological image and genomic data, which solves the problem of multimodal omic data fusion and improves the accuracy of survival prediction.
Patent Information
- Application Number
- CN202510052796.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-14
- Publication Date
- 2025-05-30
- Estimated Expiration
- 2045-01-14
AI Technical Summary
The prior art is difficult to effectively integrate and coordinate multimodal omics data, resulting in reduced accuracy of survival prediction.
Through a multimodal survival prediction method, pathological images and genomic data of cancer patients are obtained, and feature extraction, interaction and collaborative processing are used for feature extraction, interaction and collaborative processing are finally realized.
Effective fusion and coordination between pathologic images and genomic data is achieved, and the ability of survival prediction models to extract survival prediction-related features is improved, thereby improving the accuracy of survival prediction.
Smart Images

Figure CN120072042A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the technical field of data processing, and in particular, to a multi-modal survival prediction method and device. Background Art
[0002] Survival prediction aims to evaluate the relative death risk in cancer prognosis. In related technologies, the survival rate of cancer patients in a future time period is usually predicted by using the pathomics data and genomics data of cancer patients, so as to achieve survival prediction. However, related technologies often rely on the existing attention mechanism to integrate the features of pathomics data and genomics data, and it is difficult to effectively fuse and coordinate multi-modal omics data, resulting in a decrease in the accuracy of survival prediction. Summary of the Invention
[0003] Embodiments of this application provide a multi-modal survival prediction method and device for improving the accuracy of survival prediction.
[0004] On the one hand, embodiments of this application provide a multi-modal survival prediction method, including the following steps: Obtain the pathomic images and genomics data of cancer patients; Use a preset survival prediction model to perform survival prediction on the pathomic images and the genomics data to obtain the survival rate of the cancer patients in a future time period; Wherein, the survival prediction model includes: An isolation processing structure for extracting features from the pathomic images and the genomics data to obtain a pathomic isolation feature sequence and a gene isolation feature sequence; An interaction module for performing feature interaction on the pathomic isolation feature sequence and the gene isolation feature sequence to obtain a pathomic interaction feature sequence and a gene interaction feature sequence; A collaborative processing structure for extracting features from the pathomic interaction feature sequence and the gene interaction feature sequence to obtain a pathomic collaborative feature sequence and a gene collaborative feature sequence; A fusion prediction module for performing fusion-based survival prediction based on the pathomic collaborative feature sequence and the gene collaborative feature sequence to obtain the survival rate of the cancer patients in a future time period.
[0005] On the other hand, embodiments of this application provide a multi-modal survival prediction device, including: An acquisition module for acquiring the pathomic images and genomics data of cancer patients; A processing module for using a preset survival prediction model to perform survival prediction on the pathomic images and the genomics data to obtain the survival rate of the cancer patients in a future time period; Wherein, the survival prediction model includes: An isolation processing structure for extracting features from the pathological image and the genomics data to obtain a pathological isolation feature sequence and a gene isolation feature sequence; An interaction module for performing feature interaction on the pathological isolation feature sequence and the gene isolation feature sequence to obtain a pathological interaction feature sequence and a gene interaction feature sequence; A collaborative processing structure for extracting features from the pathological interaction feature sequence and the gene interaction feature sequence to obtain a pathological collaborative feature sequence and a gene collaborative feature sequence; A fusion prediction module for performing fusion-based survival prediction based on the pathological collaborative feature sequence and the gene collaborative feature sequence to obtain the survival rate of the cancer patient within a future time period.
[0006] The beneficial effects of this application are as follows: A multi-modal survival prediction method and device are provided. The pathological image and genomics data of a cancer patient are obtained, and a preset survival prediction model is used to perform survival prediction on the pathological image and genomics data to obtain the survival rate of the cancer patient within a future time period. In the prediction model, the isolation processing structure extracts features from the pathological image and genomics data to obtain a pathological isolation feature sequence and a gene isolation feature sequence. The interaction module performs feature interaction on the pathological isolation feature sequence and the gene isolation feature sequence to obtain a pathological interaction feature sequence and a gene interaction feature sequence. The collaborative processing structure extracts features from the pathological interaction feature sequence and the gene interaction feature sequence to obtain a pathological collaborative feature sequence and a gene collaborative feature sequence. The fusion prediction module performs fusion-based survival prediction based on the pathological collaborative feature sequence and the gene collaborative feature sequence to obtain the survival rate of the cancer patient within a future time period, thereby realizing the survival prediction of the cancer patient. In this way, the embodiments of this application can effectively fuse and collaborate between the pathological omics image and the genomics data, improve the feature extraction ability of the survival prediction model for survival prediction-related features, and thus improve the accuracy of the survival prediction model.
[0007] Other features and advantages of this application will be described in the subsequent specification. Moreover, some of them will become obvious from the specification, or can be understood by implementing this application. The objectives and other advantages of this application can be achieved and obtained through the structures specifically pointed out in the specification, claims, and drawings. Description of the Drawings
[0008] Figure 1 is a flowchart of the multi-modal survival prediction method provided by this application; Figure 2 is a schematic diagram of the principle of the multi-modal survival prediction method provided by this application; Figure 3 is a schematic diagram of the principle of progressive feature extraction provided by this application; Figure 4 It is the schematic diagram of optimized aggregation provided by this application; Figure 5 It is the statistical analysis verification diagram of the multi-modal survival prediction method provided by this application; Figure 6 It is the comparison diagram of the overhead and performance between the multi-modal survival prediction method provided by this application and other mainstream methods. Detailed implementation manners
[0009] In order to make the objectives, technical solutions and advantages of this application clearer, the following further elaborates on this application in combination with the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain this application and are not used to limit this application.
[0010] The following further illustrates this application in combination with the accompanying drawings of the specification and specific embodiments. The described embodiments should not be regarded as limitations on this application. All other embodiments obtained by those of ordinary skill in the art without creative efforts fall within the scope of protection of this application.
[0011] In the following descriptions, reference is made to "some embodiments", which describe a subset of all possible embodiments. However, it can be understood that "some embodiments" can be the same subset or different subsets of all possible embodiments, and can be combined with each other without conflict.
[0012] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by those skilled in the technical field to which this application belongs. The terms used herein are only for the purpose of describing the embodiments of this application and are not intended to limit this application.
[0013] Survival prediction aims to evaluate the relative risk of death in cancer prognosis, which provides crucial support for medical decision-making. Traditional survival prediction methods rely on single-modal data to achieve survival prediction. Taking pathomics data as an example, Multi-Instance Learning (MIL) has been widely used in recent years to process Whole Slide Images (WSIs), which are usually at the gigapixel level. In multi-instance learning, the entire pathomics image is regarded as an instance bag, and each instance in the instance bag represents a local area of the pathomics image. The survival prediction model can achieve survival prediction by identifying predictive local areas. Taking genomics data as another example, Self-Normalizing Neural Networks (SNNs) have been used in recent years to extract features from genomics data for survival prediction. This neural network can automatically standardize the data and extract features that are significant for survival prediction from gene expression data.
[0014] Although single-modal survival prediction methods have shown certain potential, analyzing only single-modal data cannot fully utilize the information of other modalities, making it difficult to fully leverage the advantages of multi-modal omics data, thus affecting the comprehensiveness and accuracy of survival prediction and limiting the effectiveness of survival prediction. In response, multi-modal technologies have gradually received more attention in survival prediction. Multi-modal survival prediction methods can effectively integrate the advantages of each modality by combining data from different modalities, thereby improving the accuracy of survival prediction and making up for the deficiencies of single-modal survival prediction methods. For example, in related technologies, pathomics data and genomics data of cancer patients are usually used to predict the survival rate of cancer patients in a future time period, thus achieving survival prediction. Although multi-modal survival prediction methods have made some progress, related technologies often rely on existing attention mechanisms such as the multi-head attention mechanism to integrate the features of pathomics data and genomics data, and it is difficult to achieve effective fusion and coordination of multi-modal omics data.
[0015] Specifically, the related technologies ignore the heterogeneity and sparsity of modalities, which limits the advantages of multi-modal omics data fusion. Regarding sparsity, multi-modal survival prediction methods usually need to process high-dimensional genomics data and gigapixel-level pathology images, but only a small amount of data in these data is closely related to survival prediction. Such data has a certain intra-modal sparsity, which not only increases the computational cost but also may cause the model to capture false associations, such as false associations between survival prediction results and non-pathological regions, and this kind of false association will significantly reduce the accuracy and reliability of survival prediction. It is difficult for the related technologies to reduce the impact of intra-modal sparsity on survival prediction only relying on the attention mechanism. In addition, affected by factors such as the distribution of tumors and insufficient blood samples, it is often difficult to obtain complete multi-modal data, which leads to the lack of some modal data, and the related technologies still cannot effectively handle this modal missing problem, thus affecting the effect and applicability of survival prediction. Regarding heterogeneity, on the one hand, clinical data often has significant intra-modal heterogeneity, which is mainly caused by factors such as genetic mutations and differences in the tumor microenvironment, making feature extraction complex and difficult. It is difficult for the related technologies to adapt to this intra-modal heterogeneity only relying on the attention mechanism, which in turn leads to problems such as low accuracy of survival prediction results and poor generalization ability of the model. On the other hand, inter-modal heterogeneity is also a major challenge. There are fundamental differences in the nature of pathology data and genomics data, which often requires professional technical personnel to process different modal omics data, thus increasing the cost of survival prediction and reducing the efficiency of survival prediction.
[0016] In summary, how to effectively address the intra-modal sparsity problem of multi-modal data in survival prediction, reduce the impact of false associations on survival prediction and improve the efficiency of survival prediction, and how to handle the heterogeneity problem of multi-modal data, achieve the effective fusion and coordination of different modal omics data, and then improve the accuracy of survival prediction, have become urgent problems to be solved.
[0017] In view of this, the present application provides a multi-modal survival prediction method and device, aiming to improve the accuracy of survival prediction.
[0018] First, the specific implementation steps of a multi-modal survival prediction method provided by the present application will be elaborated in detail below.
[0019] A multimodal survival prediction method provided by this application can be applied to a terminal, a server, or software running on a terminal or a server. The terminal can be a tablet computer, a laptop computer, a desktop computer, etc., but is not limited thereto. The server can be an independent physical server, a server cluster or a distributed system composed of multiple physical servers, or a cloud server providing basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communications, middleware services, domain name services, security services, content delivery network (CDN), and big data and artificial intelligence platforms. In addition, the server can also be a node server in a blockchain network, but is not limited thereto. Among them, the blockchain is a new application mode of computer technologies such as distributed data storage, peer-to-peer transmission, consensus mechanism, and encryption algorithms.
[0020] Referring to Figure 1 and Figure 2 , Figure 1 is a flowchart of the multimodal survival prediction method provided by this application, Figure 2 is a schematic diagram of the multimodal survival prediction method provided by this application, Figure 2 In [diagram reference], "genome" refers to genomics data, "pathological image" refers to a pathological image, "grouping and feature extraction", "isolation unit", and "ATSA" together indicate an isolation processing structure, "cross-attention mechanism" refers to an interaction module, "initialization" and "coordination unit" together indicate a coordination processing structure, and "low-rank fusion" and "classification head" refer to a fusion prediction module; the multimodal survival prediction method may include the following steps S101-S102.
[0021] S101, obtaining a pathological image and genomics data of a cancer patient.
[0022] It should be noted that the pathological image refers to a whole-slide digital section image generated by scanning a pathological section of a cancer patient, which is a two-dimensional optical image. In addition, the genomics data refers to gene expression data, which can be obtained through sequencing technologies such as deoxyribonucleic acid (DNA) microarray, RNA sequencing (RNA-seq) technology, or single-cell RNA sequencing (scRNA-seq) technology, but is not limited thereto.
[0023] In this step, the pathological images and genomics data of cancer patients are obtained through a preset database. It should be emphasized that the pathological images and genomics data of cancer patients are both data samples pre-stored in the database, rather than data samples directly sampled from a living body.
[0024] Optionally, the type of cancer can be selected according to the actual situation, and the embodiments of the present application do not make specific limitations in this regard. For example, the embodiments of the present application can be applicable to adenocarcinoma, glioma, lung squamous cell carcinoma, endometrial cancer, bladder cancer, breast cancer, etc., but are not limited thereto.
[0025] S102, using a preset survival prediction model to perform survival prediction on the pathological images and genomics data, and obtaining the survival rate of cancer patients within a future time period; Among them, the survival prediction model includes: An isolation processing structure for extracting features from the pathological images and genomics data to obtain a pathological isolation feature sequence and a gene isolation feature sequence; An interaction module for performing feature interaction on the pathological isolation feature sequence and the gene isolation feature sequence to obtain a pathological interaction feature sequence and a gene interaction feature sequence; A collaborative processing structure for extracting features from the pathological interaction feature sequence and the gene interaction feature sequence to obtain a pathological collaborative feature sequence and a gene collaborative feature sequence; A fusion prediction module for performing fusion-based survival prediction based on the pathological collaborative feature sequence and the gene collaborative feature sequence, and obtaining the survival rate of cancer patients within a future time period.
[0026] It should be noted that the survival prediction model is an improved neural network model pre-trained. Among them, when training the above survival prediction model, the data samples of the above survival prediction model include pathological image samples and genomics data samples of multiple patients, and the label of the above survival prediction model is the actual survival situation and event indicator variable of each patient at a certain moment. If a patient dies at a certain moment, the event indicator variable is the first value, otherwise it is the second value. Optionally, the first value and the second value can be set according to the actual situation, and the present embodiment does not make specific limitations in this regard. For example, the above first value can be one, and the above second value can be zero, but is not limited thereto.
[0027] It can be understood that the survival rate of cancer patients within a future time period can include the survival rate of cancer patients at at least one future moment, and the survival rate is the survival probability.
[0028] In this step, the pathological image and genomics data are input into the survival prediction model. In the survival prediction model, first, the pathological image and genomics data are subjected to feature extraction through an isolation processing structure, aiming to perform specific feature extraction for data of different modalities, capture the key features for survival prediction, and mask the redundant features irrelevant to survival prediction, so as to obtain a pathological isolation feature sequence and a gene isolation feature sequence; then, the pathological isolation feature sequence and the gene isolation feature sequence are subjected to feature interaction through an interaction module, aiming to promote cross-modal complementarity and fusion between pathological omics features and genomics features, so as to obtain a pathological interaction feature sequence and a gene interaction feature sequence; after that, the pathological interaction feature sequence and the gene interaction feature sequence are subjected to feature extraction through a collaborative processing structure, aiming to perform specific feature extraction for the fused pathological omics features and the fused genomics features, and further capture the key features for survival prediction from the fused omics features, so as to obtain a pathological collaborative feature sequence and a gene collaborative feature sequence; finally, a fusion-based survival prediction is performed on the pathological collaborative feature sequence and the gene collaborative feature sequence through a fusion prediction module. The fusion-based survival prediction means that feature fusion is first performed and then survival prediction is carried out, so as to obtain the survival rate of cancer patients in a future time period and achieve the survival prediction of cancer patients. In this way, the embodiments of the present application can achieve effective fusion and collaboration between pathological omics images and genomics data, improve the feature extraction ability of the survival prediction model for survival prediction-related features, and thus improve the accuracy of the survival prediction model.
[0029] In some embodiments, with reference to Figure 2 , the above isolation processing structure may include two parallel isolation processing branches. The input of one isolation processing branch is the pathological image, and the output of one isolation processing branch is the pathological isolation feature sequence. The input of the other isolation processing branch is the genomics data, and the output of the other isolation processing branch is the gene isolation feature sequence; the isolation processing branch includes a first input layer, an optimization module, and two sequentially connected isolation units; the first input layer is used to obtain a first input sequence according to the input of the isolation processing branch and a preset classification label; the input of the first isolation unit is the first input sequence, and the isolation unit is used to perform feature extraction on the input of the isolation unit to obtain the output of the isolation unit; the optimization module is used to optimize the output of the second isolation unit to obtain the output of the isolation processing branch.
[0030] In this embodiment, the isolation processing structure includes two parallel isolation processing branches, namely a pathological isolation processing branch and a gene isolation processing branch. The input of the pathological isolation processing branch is the pathological image, and the output is the pathological isolation feature sequence; the input of the gene isolation processing branch is the genomics data, and the output is the gene isolation feature sequence, as Figure 2As shown, the data flow branch where the "Pathological PREE" of the "Isolation Unit" is located is the pathological isolation processing branch, and the data flow branch where the "Genetic PREE" of the "Isolation Unit" is located is the genetic isolation processing branch. Both of these isolation processing branches include a first input layer, an optimization module, and two sequentially connected isolation units, where: Figure 2 "Group and extract features" in [reference] indicates the first input layer. The role of the first input layer is to obtain a first input sequence based on the input of the isolation processing branch and a preset classification token. Specifically, in the first input layer of the genetic isolation processing branch, first divide the genomics data into six functional groups: tumor suppression, protein kinase, protein kinase, cell differentiation, transcription, and cytokines and growth factors. Then, use a self-normalizing neural network to extract features from the data of each functional group to obtain genomics feature tokens for each functional group, such as Figure 2 As shown, "Genomic features" refers to genomics feature tokens. One functional group corresponds to at least one genomics feature token. After that, splice the genomics feature tokens of each functional group to obtain a genomics feature sequence. Finally, add a first classification token (Class token) corresponding to genomics at the starting position of the genomics feature sequence to obtain a first input sequence corresponding to genomics, such as Figure 2 As shown, "Global features" are the classification tokens. Among them, the first input sequence corresponding to genomics includes the first classification token and multiple genomics feature tokens. In the first input layer of the pathological isolation processing branch, first divide the pathological image into multiple non-overlapping image patches (Patches). Then, use a pre-trained model such as the Clustering-constrained Attention Multiple Instance Learning (CLAM) model to extract features from each image patch to obtain multiple pathomics feature tokens, such as Figure 2 As shown, "Pathological features" refers to pathomics feature tokens. After that, splice the pathomics feature tokens of each pathomics feature token to obtain a pathomics feature sequence. Finally, add a second classification token corresponding to pathomics at the starting position of the pathomics feature sequence to obtain a first input sequence corresponding to pathomics. Among them, the first input sequence corresponding to pathomics includes the second classification token and several pathomics feature tokens. It should be noted that the above first classification token and the above second classification token are both preset classification tokens, and different modalities have different classification tokens, that is, the above first classification token and the above second classification token are different.
[0031] The isolation units of both the genetic isolation processing branch and the pathological isolation processing branch are roughly the same in structure and function. Such as Figure 2As shown, two sequentially connected isolation units are arranged between the first input layer and the optimization module. That is, the input end of the first isolation unit is connected to the output end of the first input layer. The input of the first isolation unit is the first input sequence. The output end of the first isolation unit is connected to the input end of the second isolation unit. The output end of the second isolation unit is connected to the input end of the optimization module. The function of the isolation unit is to perform feature extraction on the input of the isolation unit to obtain the output of the isolation unit, aiming to perform specific feature extraction for different modal data and effectively address the problems of data diversity and modal missing.
[0032] The optimization modules of both the gene isolation processing branch and the pathology isolation processing branch are exactly the same in structure and function. Figure 2 "ATSA" in [reference] is the optimization module, also known as the Adaptive Tag Selection and Aggregation Module. The function of the optimization module is to optimize the output of the second isolation unit to obtain the output of the isolation processing branch, aiming to further extract the key features for survival prediction and mask the redundant features irrelevant to survival prediction, especially the spurious associations.
[0033] In some embodiments, referring to Figure 2 , the above isolation unit may include: A first encoder for encoding the input of the isolation unit to obtain a first encoded feature sequence; A first expert module for performing progressive feature extraction on the first encoded feature sequence to obtain the output of the isolation unit.
[0034] In this embodiment, the above isolation unit may include a first encoder and a first expert module. The first encoder is a Transformer encoder. As Figure 2 shown, the "Transformer block" in the "isolation unit" refers to the first encoder, and "PREE" refers to the first expert module, also known as the Progressive Residual Expert Expansion Module. The "gene PREE" in the "isolation unit" refers to the first expert module of the gene isolation processing branch, and the "pathology PREE" in the "isolation unit" refers to the first expert module of the pathology isolation processing branch.
[0035] In the isolation unit, first, the input of the isolation unit is encoded by the first encoder to obtain the first encoded feature sequence, which aims to calculate the correlation between omics feature tokens through the multi-head self-attention mechanism and retain the basic semantic information of the omics feature tokens in combination with the feed-forward network, so as to capture the global dependencies within the modality. Among them, the first encoded feature sequence includes a classification token and multiple omics feature tokens. Here, the classification token refers to the classification token after being processed by the first encoder, and the omics feature token here refers to the omics feature token after being processed by the first encoder. It should be noted that both the classification token and all omics feature tokens in the input of the isolation unit participate in the multi-head self-attention mechanism as queries, keys, and values at the same time. Then, the first expert module performs progressive feature extraction on the first encoded feature sequence output by the first encoder to obtain the output of the isolation unit. Progressive feature extraction refers to multi-scale feature extraction from shallow to deep, aiming to perform specific feature extraction on the data of the current modality and effectively address the problems of data diversity and modality missing.
[0036] Optionally, as Figure 2 shown, in the isolation unit of the pathological isolation processing branch, a preset positional encoding can be added to the first encoded feature sequence corresponding to the pathological omics, so as to capture the spatial information of the image and the relative positional relationship between pixels, thereby better retaining local features and enhancing the survival prediction model's understanding of the spatial context.
[0037] In some embodiments, referring to Figure 3 , for the problems of intra-modal sparsity and intra-modal heterogeneity, when the above-mentioned first expert module is used to perform progressive feature extraction on the first encoded feature sequence to obtain the output of the isolation unit, it is specifically used to perform the following operations: Process the first encoded feature sequence by using a preset first multi-layer perceptron and a preset first expert network to obtain a first residual feature sequence; Process the first residual feature sequence by using a preset second multi-layer perceptron and two preset second expert networks to obtain a second residual feature sequence; Process the second residual feature sequence by using a preset third multi-layer perceptron and four preset third expert networks to obtain the output of the isolation unit.
[0038] In this embodiment, the first expert module can be divided into a three-layer sequentially connected architecture, as Figure 3As shown, from bottom to top are the first - layer architecture, the second - layer architecture, and the third - layer architecture. "Expert" refers to the expert network. The first - layer architecture includes the first multi - layer perceptron and the first expert network. The second - layer architecture includes the second multi - layer perceptron and two second expert networks. The second - layer architecture selects the final expert network from the two second expert networks through a preset gating mechanism. The third - layer architecture includes the third multi - layer perceptron and four third expert networks. The third - layer architecture also selects the final expert network from the four third expert networks through a preset gating mechanism.
[0039] Specifically, in the first - layer architecture, the first encoded feature sequence is simultaneously input into the first multi - layer perceptron and the first expert network. The first multi - layer perceptron extracts the basic features of the current modality from the first encoded feature sequence, and the first expert network extracts more targeted features for the current modality from the first encoded feature sequence. Then, the output of the first multi - layer perceptron and the output of the first expert network are residually connected, aiming to fuse deep features and shallow features, so as to obtain the first residual feature sequence and input it into the second - layer architecture.
[0040] In the second - layer architecture, the second multi - layer perceptron extracts the basic features of the current modality from the first residual feature sequence. At the same time, the first residual feature sequence is combined with a preset gating mechanism to select the second expert network that best matches the output of the first - layer architecture from the two second expert networks as the first target network. After selecting and activating the first target network, the first target network extracts more targeted features for the current modality from the first residual feature sequence. Then, the output of the second multi - layer perceptron and the output of the first target network are residually connected, aiming to fuse deep features and shallow features, so as to obtain the second residual feature sequence and input it into the third - layer architecture.
[0041] In the third - layer architecture, the third multi - layer perceptron extracts the basic features of the current modality from the second residual feature sequence. At the same time, the second residual feature sequence is combined with a preset gating mechanism to select the third expert network that best matches the output of the second - layer architecture from the four third expert networks as the second target network. After selecting and activating the second target network, the second target network extracts more targeted features for the current modality from the second residual feature sequence. Then, the output of the third multi - layer perceptron and the output of the second target network are residually connected, aiming to fuse deep features and shallow features, so as to obtain the output of the isolation unit.
[0042] More specifically, the gating mechanism is a gating mechanism composed of a multi-layer perceptron and a Softmax function, which belongs to the existing gating mechanism. Briefly speaking, for the second-layer architecture and the third-layer architecture, in the gating mechanism, the features input to the current-layer architecture are mapped into score vectors of multiple expert networks through a multi-layer perceptron, and the Softmax function is used to score the classification task. For example, for the second-layer architecture, it has two expert networks and it is a binary classification task, while for the third-layer architecture, it has four expert networks and it is a four-classification task. The Softmax function can normalize the score vector of each expert network into a probability distribution, obtain the score information of each expert network, and select the expert network with the highest score information as the target network of the current-layer architecture. The target network will participate in feature extraction, while the other unselected expert networks will not participate in feature extraction.
[0043] It can be seen that for each layer of architecture in this embodiment, the number of selectable expert networks is one, two, and four respectively, that is, it gradually expands from one expert network to two expert networks and four expert networks, showing a progressive growth. In this way, it can match the gradually complex modal features and extract the modal multi-scale features from the shallow layer to the deep layer. Further, in this embodiment, the gating mechanism dynamically activates the expert networks targeted at specific modalities. For different modalities and different patient characteristics, different expert networks can be activated. This gating mechanism can select the most suitable expert network according to the input features, complete the optimal routing selection based on the current input, and achieve flexible and efficient feature processing. This expert selection helps to gradually extract more complex deep modal features, and then uses the selected expert network for feature extraction and introduces residual connections to fuse the shallow features and deep features of specific modalities, realizing progressive feature extraction. In this way, this embodiment can gradually extract the modal multi-scale features from the shallow layer to the deep layer for a specific modality, capture more and more complex modal features, and effectively address the data diversity problem (i.e., intra-modal heterogeneity) caused by factors such as genetic mutations and differences in the tumor microenvironment, and the modal missing problem (i.e., intra-modal sparsity) caused by factors such as the distribution of tumors and insufficient blood samples, thereby improving the feature extraction ability of the survival prediction model and the accuracy of survival prediction.
[0044] Optionally, for the isolation unit of the pathological isolation processing branch, the expert network of its first expert module can be a Convolutional Neural Network (CNN); while for the isolation unit of the gene isolation processing branch, the expert network of its first expert module can be a self-normalizing neural network, but it is not limited to this. In addition, during the training of the survival prediction model, for the first expert module of the isolation unit, the hyperparameters of all multi-layer perceptrons are frozen, that is, the hyperparameters of all multi-layer perceptrons are not updated, and they are initialized with the model parameters in the preprocessing process; while the hyperparameters of all expert networks are trainable, that is, the hyperparameters of all expert networks are updated.
[0045] In some embodiments, referring to Figure 4 , for the problem of intra-modal sparsity, when the above optimization module is used to optimize the output of the second isolation unit to obtain the output of the isolation processing branch, it is specifically used to perform the following operations: Concatenate the classification markers and multiple omics feature markers in the output of the second isolation unit to obtain multiple concatenated markers and the score information of each concatenated marker; Classify the multiple concatenated markers based on the score information of each concatenated marker to obtain a first category group and a second category group. The first category group includes multiple concatenated markers with score information greater than or equal to a preset first threshold, and the second category group includes multiple concatenated markers with score information less than the first threshold; Selectively improve the first category group to obtain multiple first selection markers and multiple second selection markers; Perform adaptive pooling based on the multiple second selection markers and the second category group to obtain multiple pooled markers; Obtain the output of the isolation processing branch according to the multiple first selection markers and the multiple pooled markers.
[0046] In this embodiment, the above optimization module aims to calculate the importance of the input high-dimensional feature markers and select features with different scores for screening and aggregation, so as to reduce redundant information and retain key information. Specifically, the output of the second isolation unit is a feature sequence of a specific modality, and this feature sequence can be expressed as , which can include classification markers and N omics feature markers . Here, the classification marker refers to the classification marker introduced from the first input layer and processed by two isolation units, and the omics feature marker refers to the omics feature marker processed by two isolation units. The output of the second isolation unit is input to the optimization module, and in the optimization module: First, a multi-layer perceptron is used to reduce the dimension of the output of the second isolation unit to obtain the output after dimension reduction. Each omics feature label in the output after dimension reduction is concatenated with the classification label in the output after dimension reduction to obtain N concatenated labels. Since the classification label represents global information, concatenating the classification label with the omics feature label enables the omics feature label to incorporate global information. For ease of understanding, the th concatenated label is expressed as the following formula (1): , (1); In formula (1), represents the th concatenated label; represents the th omics feature label in the output of the second isolation unit; represents the classification label in the output of the second isolation unit; represents the multi-layer perceptron.
[0047] After completing the label concatenation, the importance score of each concatenated label is calculated through a gating mechanism composed of a multi-layer perceptron and a Softmax function, and then the score information of multiple concatenated labels is obtained. It can be understood that the principle of the gating mechanism here is the same as that of the gating mechanism in the aforementioned expert module, and will not be elaborated here. For ease of understanding, the score information of the th concatenated label is expressed as the following formula (2): (2); In formula (2), represents the score information of the th concatenated label; represents the Softmax function.
[0048] Then, based on the score information of each concatenated label, the N concatenated labels are sorted from largest to smallest. It can be understood that the higher the score information of the concatenated label, the smaller the sorting number of the concatenated label, that is, the higher the ranking, and the lower the score information of the concatenated label, the larger the sorting number of the concatenated label, that is, the lower the ranking. The first category group is constructed by K concatenated labels whose score information is greater than or equal to the first threshold. These concatenated labels whose score information is greater than or equal to the first threshold can be understood as the top K concatenated labels, denoted as TopK labels, which are the most predictive omics feature labels. At the same time, the second category group is constructed by N - K concatenated labels whose score information is less than the first threshold. These concatenated labels whose score information is less than the first threshold can be understood as the last K + 1 concatenated labels, denoted as non - TopK labels, which are redundant omics feature labels. For ease of understanding, the first category group satisfies the following formula (3): , (3); In formula (3), represents the first category group; represents the set composed of the score information of all splicing markers; represents the set obtained by operating on the set using the TopK function. This operation means sorting all splicing markers in descending order of score information and selecting the top K splicing markers; represents the th splicing marker in the first category group, which belongs to the set .
[0049] Optionally, the number of splicing markers included in the above first category group and the above second category group, as well as the above first threshold, can all be set according to the actual situation, and this embodiment does not make specific limitations on this.
[0050] After that, selective refinement is performed. Specifically, first, the K TopK markers are divided to obtain first selection markers and second selection markers to achieve selective refinement. It can be understood that the proportion of the first selection markers among all TopK markers is , and the proportion of the second selection markers among all TopK markers is .
[0051] Optionally, the division method of selective refinement can be set according to the actual situation, and this embodiment does not make specific limitations on this. For example, is a preset value, and the first category group is divided into first selection markers and second selection markers according to the preset ratio. Another example is to classify the first category group based on the score information of each splicing marker in the first category group. The specific implementation is the same as the specific implementation of classifying multiple splicing markers. Among them, the top K splicing markers in the first category group are the first selection markers, and the remaining splicing markers are the second selection markers, but it is not limited to this.
[0052] After completing the selective refinement, on the marker dimension, second selection markers and N - K non - TopK markers are spliced to obtain a splicing sequence, and the splicing sequence is pooled to generate supplementary information, that is, A pooling marker is used to achieve adaptive pooling. In this way, by aggregating redundant omics feature markers using a part of the most predictive omics feature markers, redundant features irrelevant to survival prediction can be effectively masked, feature information falsely associated with survival prediction can be reduced, and the computational complexity of feature extraction can be lowered. Thus, the impact of within-modal sparsity on survival prediction can be reduced. At the same time, features associated with survival prediction hidden in redundant features can be effectively captured, which is beneficial to improving feature comprehensiveness. For ease of understanding, multiple pooling markers are represented by the following formula (4): (4); In formula (4), represents the number of pooling markers; represents pooling processing; represents the number of second selection markers; represents N-K splicing markers in the second category group.
[0053] Optionally, the above pooling processing can be flexibly set according to the actual situation. For example, the above pooling processing can be average pooling processing, but it is not limited to this.
[0054] Finally, add the number of first selection markers and the number of pooling markers to obtain K refined TopK marker representations, which are the outputs of the isolation processing branch. In this way, by integrating the optimized features with another part of the most predictive omics feature markers, key features with predictive value can be effectively retained, and the comprehensiveness and accuracy of the key features for survival prediction can be improved. For ease of understanding, the output of the isolation processing branch is represented by the following formula (5): (5); In formula (5), represents the output of the isolation processing branch; represents the number of first selection markers.
[0055] It can be seen that the optimization module provided in this embodiment first interacts the classification markers and multiple omics feature markers in the output of the second isolation unit, aiming to make the omics feature markers fuse with global information, so as to obtain multiple splicing markers. Then, a gating mechanism is used to calculate the score information of each splicing marker, and all splicing markers are divided into a first category group and a second category group based on the score information of each splicing marker. The first category group has high predictive value, while the second category group is regarded as redundant features. Then, a part of the features with high predictive value is used to aggregate the redundant features, which can effectively shield the redundant features irrelevant to survival prediction, reduce the feature information falsely associated with survival prediction, and reduce the computational complexity of feature extraction, thereby reducing the impact of in-modal sparsity on survival prediction. At the same time, it can effectively capture the features associated with survival prediction hidden in the redundant features, which is beneficial to improving the comprehensiveness of features. Finally, the optimized features are integrated with another part of the most predictive omics feature markers to obtain the output of the isolation processing branch, which can effectively retain the key information with predictive value and improve the comprehensiveness and accuracy of the key features for survival prediction. In this way, this embodiment can effectively address the problem of in-modal sparsity and improve the feature extraction ability of the survival prediction model.
[0056] In some embodiments, referring to Figure 2 , for the problem of inter-modal heterogeneity, when the above interaction module is used to perform feature interaction on the pathological isolation feature sequence and the gene isolation feature sequence to obtain the pathological interaction feature sequence and the gene interaction feature sequence, it is specifically used to perform the following operations: Use a cross-modal attention mechanism to process the pathological isolation feature sequence and the gene isolation feature sequence to obtain the pathological interaction feature sequence and the gene interaction feature sequence.
[0057] In this embodiment, in the isolation processing structure, the output of the pathological isolation processing branch is the pathological isolation feature sequence, and the pathological isolation feature sequence includes multiple pathological isolation feature markers. The output of the gene isolation processing branch is the gene isolation feature sequence, and the gene isolation feature sequence includes multiple gene isolation feature markers. Here, the isolation feature marker refers to the omics feature marker after being processed by the isolation processing structure, such as Figure 2As shown, the "improved feature" is the isolation feature marker. The cross-modal attention mechanism performs feature interaction on the pathological isolation feature sequence and the gene isolation feature sequence, aiming to capture the correlation between pathological images and genomics data, promote the interaction and fusion between different modalities, and address the problem of inter-modal heterogeneity, thereby obtaining the pathological interaction feature sequence and the gene interaction feature sequence, and realizing the effective fusion and collaboration between pathomics images and genomics data. Among them, the pathological interaction feature sequence includes multiple pathological interaction feature markers, and the pathological interaction feature marker refers to the pathological isolation feature marker after being processed by the interaction module, while the gene interaction feature sequence includes multiple gene interaction feature markers, and the gene interaction feature marker refers to the gene isolation feature marker after being processed by the interaction module, such as Figure 2 As shown, the "collaborated feature" output by the "cross-attention mechanism" is the gene interaction feature marker.
[0058] Specifically, the cross-modal attention mechanism calculates the fusion result of multi-modal features, and uses multiple trainable weights to perform interaction between input features. Trainable means updating the weights during the training process of the survival prediction model. The output of the cross-modal attention mechanism contains two types of important information, namely the pathological interaction feature sequence and the gene interaction feature sequence. The process is as follows: First, the pathological isolation feature sequence is dimensionally transformed through a trainable first weight to obtain the key vector (key) of the cross-modal attention mechanism. At the same time, the gene isolation feature sequence is dimensionally transformed through a trainable second weight to obtain the query vector (query) of the cross-modal attention mechanism, as shown in the following formula (6): , (6); In formula (6), represents the key vector of the cross-modal attention mechanism; represents the trainable first weight; represents the pathological isolation feature sequence; represents the query vector of the cross-modal attention mechanism; represents the trainable second weight; represents the gene isolation feature sequence. It can be understood that the functions of the first weight and the second weight are to perform dimensional transformation on the feature sequence for the calculation of the cross-modal attention mechanism.
[0059] Then, perform the Softmax operation on the key vector and query vector of the cross-modal attention mechanism in the column direction to prompt the pathomics features to focus on the information associated with survival prediction in the genomics features, promote the fusion of the pathomics modality and the genomics modality, obtain the attention weight matrix of the key vector, then transpose the attention weight matrix of the key vector, and finally obtain the pathological interaction feature sequence based on the transposed attention weight matrix of the key vector, the pathological isolation feature sequence, and the trainable third weight, thereby reducing the impact of inter-modal heterogeneity on survival prediction, as shown in the following formula (7): (7); In formula (7), represents the pathological interaction feature sequence; represents the trainable third weight; represents the ColumnSoftmax function, that is, perform the Softmax operation in the column direction. It can be understood that the role of the third weight is to map the fused pathomics features.
[0060] Meanwhile, perform the Softmax operation on the key vector and query vector of the cross-modal attention mechanism in the row direction to prompt the genomics features to focus on the information associated with survival prediction in the pathomics features, promote the fusion of the pathomics modality and the genomics modality, obtain the attention weight matrix of the query vector, then transpose the attention weight matrix of the query vector, and finally obtain the gene interaction feature sequence based on the transposed attention weight matrix of the query vector, the gene isolation feature sequence, and the trainable fourth weight, thereby reducing the impact of inter-modal heterogeneity on survival prediction, as shown in the following formula (8): (8); In formula (8), represents the gene interaction feature sequence; represents the trainable fourth weight; represents the RowSoftmax function, that is, perform the Softmax operation in the row direction. It can be understood that the role of the fourth weight is to map the fused genomics features.
[0061] In some embodiments, refer to Figure 2, the above collaborative processing structure may include two parallel collaborative processing branches. The input of one collaborative processing branch is the pathological interaction feature sequence, and the output of one collaborative processing branch is the pathological collaborative feature sequence. The input of the other collaborative processing branch is the gene interaction feature sequence, and the output of the other collaborative processing branch is the gene collaborative feature sequence. The collaborative processing branch includes a second input layer and two sequentially connected collaborative units. The second input layer is used to obtain a second input sequence according to the input of the collaborative processing branch and a preset classification label. The input of the first collaborative unit is the second input sequence, and the output of the second collaborative unit is the output of the collaborative processing branch. The collaborative unit is used to extract features from the input of the collaborative unit to obtain the output of the collaborative unit.
[0062] In this embodiment, the collaborative processing structure aims to further extract the deep interaction features between modalities, especially the key features associated with survival prediction between modalities, and promote the deep extraction and interaction of features of different modalities, improving the integrity and deep representation of survival prediction features without introducing additional interference or redundant features. Similar to the isolation processing structure in the foregoing embodiment, the collaborative processing structure includes two parallel pathological collaborative processing branches and gene collaborative processing branches. The input of the pathological collaborative processing branch is the pathological interaction feature sequence, and the output is the pathological collaborative feature sequence. The input of the gene collaborative processing branch is the gene interaction feature sequence, and the output is the gene collaborative feature sequence. As Figure 2 shown, the data flow branch where "Pathological PREE" of the "collaborative unit" is located is the pathological collaborative processing branch, and the data flow branch where "Gene PREE" of the "collaborative unit" is located is the gene collaborative processing branch. Both of these collaborative processing branches include a second input layer and two sequentially connected collaborative units, where: Figure 2 "Initialization" in indicates the second input layer. The role of the second input layer is to obtain a second input sequence according to the input of the collaborative processing branch and a preset classification label. Specifically, in the second input layer of the gene collaborative processing branch, a third classification label corresponding to genomics is added at the starting position of the gene interaction feature sequence to obtain a second input sequence corresponding to genomics. Among them, the second input sequence corresponding to genomics includes the third classification label and multiple gene interaction feature labels. In the second input layer of the pathological collaborative processing branch, a fourth classification label corresponding to pathomics is added at the starting position of the pathological interaction feature sequence to obtain a second input sequence corresponding to pathomics. Among them, the second input sequence corresponding to pathomics includes the fourth classification label and multiple pathological interaction feature labels. It should be noted that the above third classification label and the above fourth classification label are both preset classification labels, and different modalities have different classification labels, that is, the above third classification label and the above fourth classification label are different.
[0063] The input end of the first collaborative unit is connected to the output end of the second input layer. The input of the first collaborative unit is the second input sequence. The output end of the first collaborative unit is connected to the input end of the second collaborative unit. The output of the second collaborative unit is the output of the collaborative processing branch. The role of the collaborative unit is to perform feature extraction on the input of the collaborative unit to obtain the output of the collaborative unit, aiming to perform specific feature extraction on the fused pathomics features and the fused genomics features, and further capture the key features for survival prediction from the fused omics features.
[0064] In some embodiments, referring to Figure 2 , the above-mentioned collaborative unit may include: A second encoder for encoding the input of the collaborative unit to obtain a second encoded feature sequence; A second expert module for performing progressive feature extraction on the second encoded feature sequence to obtain the output of the collaborative unit.
[0065] In this embodiment, the above-mentioned collaborative unit may include a second encoder and a second expert module. The second encoder is a Transformer encoder. As Figure 2 shown, the "Transformer block" in the "collaborative unit" refers to the second encoder, and "PREE" refers to the second expert module. Among them, "Gene PREE" refers to the second expert module of the gene collaborative processing branch, and "Pathology PREE" refers to the second expert module of the pathology collaborative processing branch. In the collaborative unit, first, the input of the collaborative unit is encoded by the second encoder to obtain a second encoded feature sequence, which aims to calculate the correlation between interactive feature tokens through the multi-head self-attention mechanism and retain the basic semantic information of the interactive feature tokens by combining with the feed-forward network, so as to capture the global dependencies between modalities. Among them, the second encoded feature sequence includes classification tokens and multiple interactive feature tokens. Here, the classification tokens refer to the classification tokens after being processed by the second encoder, and the interactive feature tokens refer to the interactive feature tokens after being processed by the second encoder. It should be noted that both the classification tokens and all interactive feature tokens in the input of the collaborative unit participate in the multi-head self-attention mechanism as queries, keys, and values at the same time. Then, the second expert module performs progressive feature extraction on the second encoded feature sequence output by the second encoder to obtain the output of the collaborative unit. Progressive feature extraction refers to multi-scale feature extraction from shallow to deep, aiming to perform specific feature extraction on the fused pathomics features and the fused genomics features, further capture the key features for survival prediction from the fused omics features, and effectively address the problems of data diversity and modality loss.
[0066] Optionally, as Figure 2As shown, in the collaboration unit of the pathological collaborative processing branch, a preset position encoding can be added to the second encoded feature sequence corresponding to pathomics, so that the spatial information of the image and the relative positional relationship between pixels can be captured, thereby better retaining local features and enhancing the survival prediction model's understanding of the spatial context.
[0067] In some embodiments, with reference to Figure 3 , for the problems of intra-modal sparsity and intra-modal heterogeneity, when the second expert module is used to perform progressive feature extraction on the second encoded feature sequence to obtain the output of the collaboration unit, it is specifically used to perform the following operations: Process the second encoded feature sequence using a preset fourth multi-layer perceptron and a preset fourth expert network to obtain a third residual feature sequence; Process the third residual feature sequence using a preset fifth multi-layer perceptron and two preset fifth expert networks to obtain a fourth residual feature sequence; Process the fourth residual feature sequence using a preset sixth multi-layer perceptron and four preset sixth expert networks to obtain the output of the collaboration unit.
[0068] In this embodiment, the structure of the second expert module is the same as that of the first expert module. The second expert module can be divided into a three-layer architecture connected in sequence. Among them, the first-layer architecture includes a fourth multi-layer perceptron and a fourth expert network; the second-layer architecture includes a fifth multi-layer perceptron and two fifth expert networks. The second-layer architecture selects the final expert network from the two fifth expert networks through a preset gating mechanism; the third-layer architecture includes a sixth multi-layer perceptron and four sixth expert networks. The third-layer architecture also selects the final expert network from the four sixth expert networks through a preset gating mechanism.
[0069] Specifically, in the first-layer architecture, the second encoded feature sequence is simultaneously input into the fourth multi-layer perceptron and the fourth expert network. The fourth multi-layer perceptron extracts the basic features of the current modality that interact with other modality information from the second encoded feature sequence, and the fourth expert network extracts more targeted features for the current modality that interacts with other modality information from the second encoded feature sequence. Then, the output of the fourth multi-layer perceptron and the output of the fourth expert network are subjected to residual connection, aiming to fuse deep features and shallow features, thereby obtaining a third residual feature sequence and inputting it into the second-layer architecture.
[0070] In the second - layer architecture, the fifth multi - layer perceptron is used to extract the basic features of the current modality that interact with other modality information from the third residual feature sequence. At the same time, a preset gating mechanism is combined with the third residual feature sequence to select the fifth expert network that best matches the output of the first - layer architecture from two fifth expert networks as the third target network. After selecting and activating the third target network, the third target network is used to extract more targeted features for the current modality that interacts with other modality information from the third residual feature sequence. Then, the output of the fifth multi - layer perceptron and the output of the third target network are subjected to residual connection, aiming to fuse deep features and shallow features, so as to obtain the fourth residual feature sequence and input it into the third - layer architecture.
[0071] In the third - layer architecture, the sixth multi - layer perceptron is used to extract the basic features of the current modality that interact with other modality information from the fourth residual feature sequence. At the same time, a preset gating mechanism is combined with the fourth residual feature sequence to select the sixth expert network that best matches the output of the second - layer architecture from four sixth expert networks as the fourth target network. After selecting and activating the fourth target network, the fourth target network is used to extract more targeted features for the current modality that interacts with other modality information from the fourth residual feature sequence. Then, the output of the sixth multi - layer perceptron and the output of the fourth target network are subjected to residual connection, aiming to fuse deep features and shallow features, so as to obtain the output of the collaborative unit.
[0072] More specifically, the gating mechanism is a gating mechanism composed of a multi - layer perceptron and a Softmax function, which belongs to the existing gating mechanism. Briefly speaking, for the second - layer architecture and the third - layer architecture, in the gating mechanism, the features input to the current - layer architecture are mapped into score vectors of multiple expert networks through the multi - layer perceptron, and the Softmax function is used to score the classification task. For example, for the second - layer architecture, it has two expert networks, which is a binary - classification task, while for the third - layer architecture, it has four expert networks, which is a four - classification task. The Softmax function can normalize the score vector of each expert network into a probability distribution, obtain the score information of each expert network, and select the expert network with the highest score information as the target network of the current - layer architecture. The target network will participate in feature extraction, while the other unselected expert networks will not participate in feature extraction.
[0073] It can be seen that for each layer of architecture in this embodiment, the number of selectable expert networks is one, two, and four respectively, that is, gradually expanding from one expert network to two expert networks and four expert networks for progressive growth. In this way, it can match the gradually complex modal features and extract modal multi-scale features from shallow to deep. Further, in this embodiment, the gating mechanism dynamically activates the expert networks targeted at specific modalities, and different expert networks can be activated for different modalities and different patient characteristics. This gating mechanism can select the most suitable expert network according to the input features to complete the optimal routing selection based on the current input, realizing flexible and efficient feature processing. This expert selection helps to gradually extract more complex deep modal features, and then uses the selected expert network for feature extraction and introduces residual connections to fuse shallow features and deep features to achieve progressive feature extraction. In this way, this embodiment can gradually extract modal multi-scale features from shallow to deep for a specific modality interacting with other modal information, capture more and more complex modal features, and effectively address the data diversity problem (i.e., intra-modal heterogeneity) caused by factors such as genetic mutations and differences in the tumor microenvironment, and the modal missing problem (i.e., intra-modal sparsity) caused by factors such as the distribution of tumors and insufficient blood samples, thereby improving the feature extraction ability of the survival prediction model and the accuracy of survival prediction.
[0074] Optionally, for the isolation unit of the pathological collaborative processing branch, the expert network of its second expert module is a convolutional neural network; while for the isolation unit of the gene collaborative processing branch, the expert network of its second expert module is a self-normalizing neural network. In addition, during the training of the survival prediction model, for the second expert module of the collaborative unit, the hyperparameters of all multi-layer perceptrons are frozen, that is, the hyperparameters of all multi-layer perceptrons are not updated, and it is initialized with the model parameters in the preprocessing process; while the hyperparameters of all expert networks are trainable, that is, the hyperparameters of all expert networks are updated.
[0075] In some embodiments, referring to Figure 2 , when the above fusion prediction module is used for fusion-based survival prediction based on the pathological collaborative feature sequence and the gene collaborative feature sequence to obtain the survival rate of cancer patients in a future time period, it is specifically used to perform the following operations: Fuse multiple pathological collaborative feature markers in the pathological collaborative feature sequence and multiple gene collaborative feature markers in the gene collaborative feature sequence to obtain target local features; Fuse the classification markers in the pathological collaborative feature sequence and the classification markers in the gene collaborative feature sequence to obtain target global features; Fuse the target local features and the target global features to obtain target fusion features; Survival prediction is performed based on the target fusion features to obtain the survival rate of cancer patients within a future time period.
[0076] In this embodiment, the above-mentioned fusion prediction module aims to extract local features through a low-rank fusion method, and then fuse the global features and local features to generate the final survival prediction feature representation. Then, based on this feature representation, a prediction is made to obtain the survival rate of cancer patients within a future time period. Specifically, the pathological collaborative feature sequence is the output of the pathological collaborative processing branch in the collaborative processing structure, which includes a classification label and multiple pathological collaborative feature labels; the gene collaborative feature sequence is the output of the gene collaborative processing branch in the collaborative processing structure, which includes a classification label and multiple gene collaborative feature labels. Here, the classification label refers to the classification label introduced and processed by the collaborative processing structure, and the collaborative feature label refers to the interaction feature label processed by the collaborative processing structure. As Figure 2 shown, the "collaborative feature after collaboration" output by the "collaborative unit" is the collaborative feature label.
[0077] In the fusion prediction module, as Figure 2 shown in the "low-rank fusion" in First, a low-rank fusion method is used to fuse multiple pathological collaborative feature labels in the pathological collaborative feature sequence and multiple gene collaborative feature labels in the gene collaborative feature sequence. Specifically, it is to decompose multiple pathological feature labels in the pathological collaborative feature sequence into a low-rank subspace to obtain pathological parameter matrices, and at the same time decompose multiple gene feature labels in the gene collaborative feature sequence into a low-rank subspace to obtain gene parameter matrices. Then, based on pathological parameter matrices, (9); In formula (9), ; represents the target local feature; represents multiple pathological feature labels in the pathological collaborative feature sequence; represents multiple gene feature labels in the gene collaborative feature sequence; represents the th pathological parameter matrix; represents the th gene parameter matrix.
[0078] AsFigure 2 As shown in the "classification header" in Figure 2 , after low-rank fusion is completed, the classification markers in the pathological collaborative feature sequence and the classification markers in the gene collaborative feature sequence are processed by a multi-layer perceptron to fuse the global information of the two modalities, thereby obtaining the target global feature. For ease of understanding, the target global feature is represented by the following formula (10): (10); In formula (10), represents the target global feature; represents the classification marker in the gene collaborative feature sequence; represents the classification marker in the pathological collaborative feature sequence.
[0079] Next, the target local feature and the target global feature are fused by means of linear combination to obtain the target fusion feature, effectively improving the comprehensive representation ability of the key features for survival prediction, which is represented by the following formula (11): (11); In formula (11), is the target fusion feature; is a parameter used to control the fusion ratio of the local feature and the global feature.
[0080] After obtaining the target fusion feature, the target fusion feature is input into a multi-layer perceptron, and the probability distribution of the survival risk of cancer patients is generated through the multi-layer perceptron and the Sigmoid function, and the survival probability at each future time is calculated, so as to obtain the survival rate of cancer patients in the future time period, as shown in the following formula (12): (12); In formula (12), represents the survival rate of cancer patients in the future time period, and this future time period includes future time , future time , …, future time ; represents the sigmoid function.
[0081] In some embodiments, the above method may further include: During the training of the survival prediction model, the negative log-likelihood loss function is used to update the parameters of the survival prediction model.
[0082] In this embodiment, in order to improve the performance of the survival prediction model, during model training, the loss of the survival prediction model is calculated by the negative log-likelihood loss function shown in the following formula (13), and gradient backpropagation is performed based on this loss to update the model parameters of the survival prediction model, such as the hyperparameters of the above isolation processing structure, the above interaction module, the above collaborative processing structure, and the above fusion prediction module: (13); In formula (13), represents the loss of the survival prediction model; is an event indicator variable. If the patient has experienced the target event (i.e., death), the event indicator variable is one, otherwise it is zero; represents at the time when the sample is used as input the survival rate; represents at the time when the sample is used as input the survival rate; represents the probability of having cancer at the time when the sample is used as input. Among them, the survival rate at the time when the sample is used as input satisfies the following formula (14): In formula (14), represents the probability of having cancer at the time when the sample .
[0083] For the convenience of understanding the above multi-modal survival prediction method in the embodiments of the present application, the principle of the above multi-modal survival prediction method in the embodiments of the present application will be described below with an application scenario.
[0084] In this application scenario, the training data samples include pathological image samples and genomics data samples of multiple patients. The label of the training data samples is the actual survival situation and event indicator variable of each patient at a certain moment. If the patient dies at a certain moment, the event indicator variable is the first value, otherwise it is the second value. The survival prediction model is trained through the training data samples and their labels. Among them, the loss function shown in the above formulas (13)-(14) is used to update the model parameters during model training. In real-time prediction, the patient to be tested is a lung adenocarcinoma patient. The whole-slide digital section image and genomics data of the lung adenocarcinoma patient are obtained through a preset database, and the whole-slide digital section image and genomics data of the lung adenocarcinoma patient are input into the survival prediction model to predict the survival probability of the lung adenocarcinoma patient in the future time period.
[0085] Refer to Figure 2 , in the survival prediction model, the following operations are specifically performed: S01, the whole-field digital slice image and genomics data are input into the isolation processing structure. The isolation processing structure includes parallel pathological isolation processing branches and gene isolation processing branches. In the gene isolation processing branch, the genomics data is first divided into six functional groups, namely tumor suppression, protein kinase, protein kinase, cell differentiation, transcription, and cytokines and growth factors, through the first input layer. Then, the data of each functional group is subjected to feature extraction through a self-normalizing neural network to obtain genomics feature markers for each functional group and splice them to obtain a genomics feature sequence. Finally, a first classification marker corresponding to genomics is added at the starting position of the genomics feature sequence to obtain a first input sequence corresponding to genomics. The first input sequence corresponding to genomics includes the first classification marker and multiple genomics feature markers; the first input sequence corresponding to genomics is subjected to feature extraction through two isolation units; the output of the second isolation unit is aggregated and optimized through an optimization module to obtain a gene isolation feature sequence, and the gene isolation feature sequence includes multiple gene isolation feature markers. In the pathological isolation processing branch, the pathological image is first divided into multiple non-overlapping image patches through the first input layer, and then a clustering-constrained attention multi-instance learning model is used to perform feature extraction on each image patch to obtain multiple pathomics feature markers and splice them to obtain a pathomics feature sequence. Finally, a second classification marker corresponding to pathomics is added at the starting position of the pathomics feature sequence to obtain a first input sequence corresponding to pathomics. The first input sequence corresponding to pathomics includes the second classification marker and several pathomics feature markers; the first input sequence corresponding to pathomics is subjected to feature extraction through two isolation units; the output of the second isolation unit is aggregated and optimized through an optimization module to obtain a pathological isolation feature sequence, and the pathological isolation feature sequence includes multiple pathological isolation feature markers.
[0086] Furthermore, in the isolation unit, first, the input of the isolation unit is encoded through a Transformer encoder to obtain a first encoded feature sequence. In particular, in the isolation unit of the pathological isolation processing branch, a position encoding is added to the first encoded feature sequence after the encoding process; then, the first encoded feature sequence is subjected to progressive feature extraction through the first expert module to obtain the output of the isolation unit, as Figure 3As shown, that is: the first encoded feature sequence is processed by the first multi-layer perceptron, the first encoded feature sequence is processed by the first expert network, the output of the first multi-layer perceptron and the output of the first expert network are subjected to residual connection to obtain the first residual feature sequence; the first residual feature sequence is processed by the second multi-layer perceptron, the first target network is determined from two second expert networks through a gating mechanism composed of a multi-layer perceptron and a Softmax function, the first residual feature sequence is processed by the first target network, and the output of the second multi-layer perceptron and the output of the first target network are subjected to residual connection to obtain the second residual feature sequence; the second residual feature sequence is processed by the third multi-layer perceptron, the second target network is determined from four third expert networks through a gating mechanism composed of a multi-layer perceptron and a Softmax function, the second residual feature sequence is processed by the second target network, and the output of the third multi-layer perceptron and the output of the second target network are subjected to residual connection to obtain the output of the isolation unit. It can be understood that the output of the isolation unit includes the classification label after being processed by the isolation unit and multiple omics feature labels after being processed by the isolation unit.
[0087] Further, referring to Figure 4 , in the optimization module, first, the output of the second isolation unit is dimensionally reduced by a multi-layer perceptron to obtain the dimensionally reduced output, and each omics feature label in the dimensionally reduced output is respectively concatenated with the classification label in the dimensionally reduced output to obtain N concatenated labels; then, the importance scores of each concatenated label are calculated through a gating mechanism composed of a multi-layer perceptron and a Softmax function to obtain the score information of the N concatenated labels, and the N concatenated labels are sorted from large to small based on the score information of each concatenated label. The first category group is constructed by K concatenated labels whose score information is greater than or equal to the first threshold, and the second category group is constructed by N-K concatenated labels whose score information is less than the first threshold; then, the K concatenated labels of the first category group are divided to obtain first selected labels and second selected labels; furthermore, on the label dimension, second selected labels and the N-K concatenated labels in the second category group are concatenated to obtain a concatenated sequence and subjected to pooling processing to generate pooled labels; finally, first selected labels and pooled labels are added to obtain K refined features, that is, the output of the isolation processing branch.
[0088] S02. The pathological isolation feature sequence and the gene isolation feature sequence are subjected to feature interaction through the cross-modal attention mechanism shown in the above formulas (6)-(8) to obtain a pathological interaction feature sequence and a gene interaction feature sequence. Among them, the pathological interaction feature sequence includes multiple pathological interaction feature markers, and the gene interaction feature sequence includes multiple gene interaction feature markers.
[0089] S03. The pathological interaction feature sequence and the gene interaction feature sequence are input into a collaborative processing structure, which includes a parallel pathological collaborative processing branch and a gene collaborative processing branch. In the gene collaborative processing branch, a third classification marker corresponding to genomics is added to the starting position of the gene interaction feature sequence through the second input layer to obtain a second input sequence corresponding to genomics, and the second input sequence corresponding to genomics includes the third classification marker and multiple gene interaction feature markers; two collaborative units are used to extract features from the second input sequence corresponding to genomics, and the second collaborative unit outputs a gene collaborative feature sequence, which includes the third classification marker processed by the collaborative unit and multiple gene collaborative feature markers. In the pathological collaborative processing branch, a fourth classification marker corresponding to pathomics is added to the starting position of the pathological interaction feature sequence through the second input layer to obtain a second input sequence corresponding to pathomics, and the second input sequence corresponding to pathomics includes the fourth classification marker and multiple pathological interaction feature markers; two collaborative units are used to extract features from the second input sequence corresponding to pathomics, and the second collaborative unit outputs a pathological collaborative feature sequence, which includes the fourth classification marker processed by the collaborative unit and multiple pathological collaborative feature markers. The processing principle of the collaborative unit is the same as that of the isolation unit, which will not be elaborated here.
[0090] S04. Multiple pathological feature markers in the pathological collaborative feature sequence and multiple gene feature markers in the gene collaborative feature sequence are fused through a low-rank fusion method, as shown in the above formula (9), to obtain a target local feature; at the same time, the classification markers in the pathological collaborative feature sequence and the classification markers in the gene collaborative feature sequence are processed through a multi-layer perceptron, as shown in the above formula (10), to obtain a target global feature; the target local feature and the target global feature are fused through a linear combination method, as shown in the above formula (11), to obtain a target fusion feature; the target fusion feature is input into a multi-layer perceptron, and the probability distribution of the survival risk of cancer patients is generated through the multi-layer perceptron and the Sigmoid function, and the survival probability at each future moment is calculated, so as to obtain the survival rate of cancer patients in the future time period, as shown in the above formula (12).
[0091] Next, the performance of the survival prediction method of the embodiment of the present application will be verified through Examples 1 to 3.
[0092] Example 1: On the publicly available bladder cancer (BLCA) dataset, breast cancer (BRCA) dataset, uterine corpus endometrial carcinoma (UCEC) dataset, glioblastoma multiforme (GBMLGG) dataset, and lung adenocarcinoma (LUAD) dataset, the survival prediction method of this application was compared with mainstream survival prediction methods in terms of performance. The results are shown in Table 1 below. In Table 1, P refers to pathological images, G refers to genomics data, the tick mark indicates the selection of data of this modality, the metric is accuracy, "AdaMHF" is the survival prediction method of this application, and the rest are mainstream methods. As can be seen from Table 1, this application achieved the best values in the survival prediction tasks of bladder cancer patients, breast cancer patients, uterine corpus endometrial carcinoma patients, and glioblastoma multiforme lung adenocarcinoma patients.
[0093] Table 1: Performance table of the survival prediction method of this application and mainstream survival prediction methods in different survival prediction tasks
[0094] Example 2: On the publicly available lung adenocarcinoma dataset, the survival prediction method of this application was verified through statistical analysis methods. The results are as Figure 5 shown. "Low Risk" represents the low-risk group, "High Risk" represents the high-risk group. The vertical coordinate "Overall Survival" represents the proportion of individuals still alive in the entire data at the current time point, and the horizontal coordinate "Time (Months)" represents the passage of time, with its unit being months. Through Figure 5 it can be seen that the p-value of this application reached , and it can accurately predict the survival probability of cancer patients in the future time period.
[0095] Example 3: On the publicly available lung adenocarcinoma dataset, the survival prediction method of this application was comprehensively compared with mainstream survival prediction methods in terms of overhead and performance. The results are as Figure 6 shown. The horizontal coordinate "FLOPs (G)" represents the computational complexity, that is, the number of floating-point operations per second, with its unit being gigabytes. The vertical coordinate "C-index" represents the concordance index, whose full English name is concordance index, and it is a performance evaluation metric. "AdaMHF" is the survival prediction method of this application, and the rest are mainstream methods. Through Figure 6 it can be seen that this application can significantly reduce the computational complexity of the survival prediction task and at the same time achieve the best survival prediction effect.
[0096] In addition, the present application also provides a survival prediction device, which may include: an acquisition module and a processing module. The acquisition module is used to acquire the pathological images and genomics data of cancer patients; the processing module is used to perform survival prediction on the pathological images and genomics data by using a preset survival prediction model to obtain the survival rate of cancer patients within a future time period. Among them, the survival prediction model includes an isolation processing structure, an interaction module, a collaborative processing structure, and a fusion prediction module. The isolation processing structure is used to extract features from the pathological images and genomics data to obtain a pathological isolation feature sequence and a gene isolation feature sequence; the interaction module is used to perform feature interaction on the pathological isolation feature sequence and the gene isolation feature sequence to obtain a pathological interaction feature sequence and a gene interaction feature sequence; the collaborative processing structure is used to extract features from the pathological interaction feature sequence and the gene interaction feature sequence to obtain a pathological collaborative feature sequence and a gene collaborative feature sequence; the fusion prediction module is used to perform fusion-based survival prediction based on the pathological collaborative feature sequence and the gene collaborative feature sequence to obtain the survival rate of cancer patients within a future time period.
[0097] The content in the above method embodiments is applicable to the device embodiments of the present application. The functions specifically implemented by the device embodiments of the present application are the same as those in the above method embodiments, and the beneficial effects achieved are also the same as those in the above method embodiments.
[0098] In summary, the present application can effectively solve the problems of sparsity and heterogeneity in multimodal survival prediction, and improve the accuracy and robustness of survival prediction. By introducing an expert module and an optimization module, the present application realizes feature extraction and optimization for intra-modal and inter-modal heterogeneity. Among them, the expert module effectively addresses the problems of data diversity and missingness by dynamically activating the expert network to perform specific feature extraction for different modal data; while the optimization module retains key information and significantly reduces the computational complexity by selecting the most informative markers and aggregating redundant markers. In addition, the fusion prediction module combines local information and global information, and effectively improves the comprehensive representation ability of data through low-rank fusion and global marker extraction. A large number of experimental results show that the present application exhibits superior performance under both complete-modal and missing-modal conditions, outperforming existing mainstream methods. At the same time, the present application realizes efficient feature extraction and fusion in resource-constrained scenarios, significantly reducing the computational overhead, and is applicable to large-scale medical data analysis and survival prediction tasks, with strong practicality and promotion value.
[0099] In some alternative embodiments, the functions / operations recited in the block diagrams may not occur in the order presented in the operational illustrations. For example, depending on the functions / operations involved, two blocks shown in succession may actually be executed substantially simultaneously or the blocks can sometimes be executed in the reverse order. Additionally, the embodiments presented and described in the flowcharts of the present application are provided by way of example in order to provide a more comprehensive understanding of the technology. The disclosed methods are not limited to the operations and logical flows presented herein. Alternative embodiments are contemplated in which the order of various operations is altered and in which sub-operations described as part of a larger operation are executed independently. Further, although the present application has been described in the context of functional modules, it should be understood that, unless otherwise stated to the contrary, one or more of the functions and / or features may be integrated in a single physical device and / or software module, or one or more functions and / or features may be implemented in separate physical devices or software modules. It should also be understood that a detailed discussion of the actual implementation of each module is not necessary for an understanding of the present application. Rather, given the attributes, functions, and internal relationships of the various functional modules in the devices disclosed herein, the actual implementation of the modules will be understood within the ordinary skill of an engineer. Accordingly, those of ordinary skill in the art can implement the present application as set forth in the claims without undue experimentation. It should also be understood that the particular concepts disclosed are illustrative only and are not intended to limit the scope of the present application, which is determined by the full scope of the appended claims and their equivalents.
[0100] If the described functions are implemented in the form of software functional units and sold or used as independent products, they may be stored in a computer-readable storage medium. Based on such understanding, the technical solution of the present application, in essence or the part that contributes to the prior art, or a part of the technical solution, may be embodied in the form of a software product. The computer software product is stored in a storage medium and includes several programs for causing a computer device (which may be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the methods described in the various embodiments of the present application. The aforementioned storage medium includes various media that can store program codes, such as USB flash drives, mobile hard disks, read-only memories (ROMs), random access memories (RAMs), magnetic disks, or optical discs.
[0101] The logic and / or steps represented in the flowchart or otherwise described herein, for example, can be considered as a definitional list of executable instructions for implementing logical functions, which can be embodied in any computer-readable medium for use by or in connection with a program execution system, apparatus, or device, such as a computer-based system, a system including a processor, or other systems that can fetch and execute the program from the program execution system, apparatus, or device. As used in this specification, a "computer-readable medium" can be any device that can contain, store, communicate, propagate, or transport the program for use by or in connection with the program execution system, apparatus, or device.
[0102] More specific examples of computer-readable media (a non-exhaustive list) include the following: an electrical connection having one or more wires (electronic device), a portable computer diskette (magnetic device), a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or Flash memory), an optical fiber device, and a portable compact disc read-only memory (CDROM). Additionally, the computer-readable medium can even be paper or other suitable media on which the program can be printed, as the program can be obtained electronically, for example, by optically scanning the paper or other media, followed by editing, interpretation, or other appropriate processing as necessary, and then stored in a computer memory. It should be understood that various parts of the present application can be implemented in hardware, software, firmware, or a combination thereof. In the above-described embodiments, multiple steps or methods can be implemented using software or firmware stored in a memory and executed by a suitable program execution system. For example, if implemented in hardware, as in another embodiment, any one of the following technologies known in the art or a combination thereof can be used: discrete logic circuits having logic gates for implementing logical functions on data signals, application specific integrated circuits having appropriate combinational logic gates, programmable gate arrays (PGA), field programmable gate arrays (FPGA), etc.
[0103] In the foregoing description of this specification, the description with reference to terms such as "one embodiment / example", "another embodiment / example" or "certain embodiments / examples" means that the specific features, structures, materials or characteristics described in connection with the embodiment or example are included in at least one embodiment or example of the present application. In this specification, the schematic representations of the above terms do not necessarily refer to the same embodiment or example. Moreover, the specific features, structures, materials or characteristics described may be combined in any one or more embodiments or examples in a suitable manner. Although the embodiments of the present application have been shown and described, those of ordinary skill in the art can understand that various changes, modifications, substitutions and variations can be made to these embodiments without departing from the principles and purposes of the present application, and the scope of the present application is defined by the claims and their equivalents. The above is a specific description of the preferred embodiments of the present application, but the present application is not limited to the described embodiments. Those skilled in the art can also make various equivalent deformations or substitutions without violating the spirit of the present application, and these equivalent deformations or substitutions are all included in the scope defined by the claims of the present application.
Claims
1. A multimodal survival prediction method, characterized in that: The following steps are involved: Access pathology images and genomic data of cancer patients; Using a preset survival prediction model to perform survival prediction on the pathological image and the genomics data to obtain the survival rate of the cancer patient in a future time period; Wherein, the survival prediction model includes: An isolation processing structure, used for extracting features from the pathological image and the genomics data to obtain a pathological isolation feature sequence and a gene isolation feature sequence; An interaction module, used for performing feature interaction between the pathological isolation feature sequence and the gene isolation feature sequence to obtain a pathological interaction feature sequence and a gene interaction feature sequence; A collaborative processing structure, used for performing feature extraction on the pathology interaction feature sequence and the gene interaction feature sequence to obtain a pathology collaborative feature sequence and a gene collaborative feature sequence; The fusion prediction module is used to perform fusion survival prediction based on the pathology collaborative feature sequence and the gene collaborative feature sequence to obtain the survival rate of the cancer patient in a future time period.
2. The multimodal survival prediction method according to claim 1, characterized in that: The isolation processing structure comprises two parallel isolation processing branches, wherein the input of one of the isolation processing branches is the pathological image, the output of one of the isolation processing branches is the pathological isolation feature sequence, the input of another isolation processing branch is the genomics data, and the output of another isolation processing branch is the gene isolation feature sequence; The isolation processing branch includes a first input layer, an optimization module and two isolation units connected in sequence; The first input layer is used to obtain a first input sequence according to the input of the isolation processing branch and a preset classification label; The input of the first isolation unit is the first input sequence, and the isolation unit is used to perform feature extraction on the input of the isolation unit to obtain the output of the isolation unit; The optimization module is used to optimize the output of the second isolation unit to obtain the output of the isolation processing branch.
3. The multimodal survival prediction method according to claim 2, characterized in that: The isolation unit comprises: A first encoder, used for encoding the input of the isolation unit to obtain a first encoding feature sequence; The first expert module is used to perform progressive feature extraction on the first coding feature sequence to obtain the output of the isolation unit.
4. The multimodal survival prediction method according to claim 3, characterized in that: The step of performing progressive feature extraction on the first coding feature sequence to obtain the output of the isolation unit includes: Processing the first encoding feature sequence using a preset first multilayer perceptron and a preset first expert network to obtain a first residual feature sequence; Processing the first residual feature sequence using a preset second multilayer perceptron and two preset second expert networks to obtain a second residual feature sequence; The second residual feature sequence is processed by using a preset third multilayer perceptron and four preset third expert networks to obtain the output of the isolation unit.
5. The multimodal survival prediction method according to claim 2, characterized in that: The step of optimizing the output of the second isolation unit to obtain the output of the isolation processing branch includes: splicing the classification mark and the plurality of omics feature marks in the output of the second isolation unit to obtain a plurality of splicing marks and score information of each of the splicing marks; Classifying the plurality of splicing marks based on the score information of each splicing mark to obtain a first category group and a second category group, wherein the first category group includes a plurality of splicing marks whose score information is greater than or equal to a preset first threshold, and the second category group includes a plurality of splicing marks whose score information is less than the first threshold; performing selective refinement on the first category group to obtain a plurality of first selection markers and a plurality of second selection markers; Performing adaptive pooling based on the plurality of second selection markers and the second category group to obtain a plurality of pooled markers; The output of the isolation processing branch is obtained according to the plurality of the first selection marks and the plurality of the pooling marks.
6. The multimodal survival prediction method according to claim 1, characterized in that: The step of performing feature interaction on the pathological isolation feature sequence and the gene isolation feature sequence to obtain a pathological interaction feature sequence and a gene interaction feature sequence comprises: The pathology isolation feature sequence and the gene isolation feature sequence are processed by using a cross-modal attention mechanism to obtain the pathology interaction feature sequence and the gene interaction feature sequence.
7. The multimodal survival prediction method according to claim 1, characterized in that: The collaborative processing structure includes two parallel collaborative processing branches, wherein the input of one of the collaborative processing branches is the pathology interaction feature sequence, the output of one of the collaborative processing branches is the pathology collaborative feature sequence, the input of another collaborative processing branch is the gene interaction feature sequence, and the output of another collaborative processing branch is the gene collaborative feature sequence; The collaborative processing branch includes a second input layer and two sequentially connected collaborative units; the second input layer is used to obtain a second input sequence based on the input of the collaborative processing branch and a preset classification label; the input of the first collaborative unit is the second input sequence, and the output of the second collaborative unit is the output of the collaborative processing branch; the collaborative unit is used to perform feature extraction on the input of the collaborative unit to obtain the output of the collaborative unit.
8. The multimodal survival prediction method according to claim 1, characterized in that: The fusion survival prediction based on the pathological collaborative feature sequence and the gene collaborative feature sequence to obtain the survival rate of the cancer patient in a future time period includes: Fusing a plurality of pathology collaborative feature markers in the pathology collaborative feature sequence and a plurality of gene collaborative feature markers in the gene collaborative feature sequence to obtain a target local feature; Fusion of the classification markers in the pathology collaborative feature sequence and the classification markers in the gene collaborative feature sequence to obtain a target global feature; Fusing the target local feature with the target global feature to obtain a target fused feature; Survival prediction is performed based on the target fusion feature to obtain the survival rate of the cancer patient in a future time period.
9. The multimodal survival prediction method according to claim 1, characterized in that: The method further comprises the following steps: During the training of the survival prediction model, the parameters of the survival prediction model are updated using a negative log-likelihood loss function.
10. A multimodal survival prediction device, characterized in that: include: An acquisition module, used to acquire pathological images and genomic data of cancer patients; A processing module, used to perform survival prediction on the pathological image and the genomics data using a preset survival prediction model to obtain the survival rate of the cancer patient in a future time period; Wherein, the survival prediction model includes: An isolation processing structure, used for extracting features from the pathological image and the genomics data to obtain a pathological isolation feature sequence and a gene isolation feature sequence; An interaction module, used for performing feature interaction between the pathological isolation feature sequence and the gene isolation feature sequence to obtain a pathological interaction feature sequence and a gene interaction feature sequence; A collaborative processing structure, used for performing feature extraction on the pathology interaction feature sequence and the gene interaction feature sequence to obtain a pathology collaborative feature sequence and a gene collaborative feature sequence; The fusion prediction module is used to perform fusion survival prediction based on the pathology collaborative feature sequence and the gene collaborative feature sequence to obtain the survival rate of the cancer patient in a future time period.
Citation Information
Patent Citations
Multi-modal fusion survival prognosis method and device based on pathology and genes
CN117594225A
Multi-modal fusion survival prediction method based on Sinkhorn algorithm
CN117952966A
Patient survival prognosis prediction method suitable for various cancers
CN118039162A
HLJ1 gene expression
US20060194235A1
Systems and methods for deep orthogonal fusion for multimodal prognostic biomarker discovery
US20220292674A1