Fusible incomplete multimodal and cross-medical site nasopharyngeal carcinoma prognosis prediction method

By combining a fully connected reconstruction network and a multi-layer survival analysis network with a cross-site biased contrastive learning model, the incomplete multimodal and data bias problems in nasopharyngeal carcinoma prognosis prediction were solved, achieving high-performance cross-site prognosis prediction.

CN115938532BActive Publication Date: 2025-11-11SUN YAT SEN UNIV
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202211538034.7
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-12-02
Publication Date
2025-11-11
Estimated Expiration
2042-12-02

AI Technical Summary

Technical Problem

Existing prognostic prediction technologies for nasopharyngeal carcinoma cannot cope with incomplete multimodal data and data offset between medical sites, resulting in poor prediction performance of the models across sites.

Method used

A fully connected reconstruction network is used for multimodal data fusion and cross-site feature transformation. Combined with a multi-layer fully connected survival analysis network and a cross-site partial contrastive learning model, gradient optimization is performed using PyTorch to train the reconstruction network and survival analysis network, generate fused features and survival risks, and achieve cross-site nasopharyngeal carcinoma prognosis prediction.

Benefits of technology

It effectively solves the problems of incomplete multimodal fusion and data offset, and improves the accuracy of nasopharyngeal carcinoma prognosis prediction and cross-site prediction performance.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115938532B_ABST
    Figure CN115938532B_ABST
Patent Text Reader

Abstract

The present application relates to the technical field of prognosis prediction, and discloses a nasopharyngeal carcinoma prognosis prediction method which can fuse incomplete multi-modal data and can be used across medical sites, comprising the following steps: S1. receiving original data and extracting multi-modal data; S2. introducing a full connection reconstruction network to obtain fusion features of the multi-modal data through the reconstruction network; S3. introducing a multi-layer full connection survival analysis network to obtain survival risk; S4. constructing a cross-site partial contrast learning model; the partial contrast learning model performs contrast learning according to different sources of positive paired samples; S5. obtaining a total loss function based on the reconstruction network, the survival analysis network and the partial contrast learning model, and training the reconstruction network and the survival analysis network; and S6. performing nasopharyngeal carcinoma prognosis prediction through the trained reconstruction network and survival analysis network in step S5. The present application solves the problem that the prior art cannot deal with incomplete multi-modal data and data offset between medical sites.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of prognostic prediction technology, and more specifically, to a prognostic prediction method for nasopharyngeal carcinoma that can integrate incomplete multimodal data and is applicable across medical sites. Background Technology

[0002] Nasopharyngeal carcinoma has a complex pathogenesis and its anatomical location near the skull base makes surgery challenging. With the widespread adoption of intensity-modulated radiotherapy (IMRT), the long-term survival rate of nasopharyngeal carcinoma patients has improved to some extent, and the 5-year local recurrence rate has decreased to below 10%. However, distant metastasis and local recurrence remain the main causes of treatment failure in nasopharyngeal carcinoma. Therefore, establishing accurate and robust prognostic prediction models can better guide the increase (or decrease) of treatment intensity, thereby reducing the risk of recurrence and death in high-risk patients (or reducing treatment toxicity in low-risk patients).

[0003] To address the prognostic prediction problem of nasopharyngeal carcinoma (NPC) patients, existing literature proposes a deep learning method utilizing multimodal data including MRI images and clinical characteristics. This method extracts deep features from T1-weighted, T2-weighted, and T1-weighted contrast-enhanced MRI images, then combines these features with clinical factors such as EBV quantitative detection results, TNM staging indicators, gender, age, BMI, lactate dehydrogenase, and albumin levels. The results are then input into the XGBoost machine learning algorithm to obtain the final risk index. This deep learning method for NPC prognostic prediction has achieved certain predictive results in survival analysis tasks at multiple medical sites (i.e., multiple hospitals).

[0004] This method has two significant shortcomings: First, the model is simply a concatenation of multimodal data from the treatment and examination process. This approach requires that all modal data in the input model be complete. However, incomplete multimodal data exists in clinical practice, which limits the application of this method. The problem of incomplete multimodal fusion refers to the fact that the diagnosis and treatment of nasopharyngeal carcinoma requires multiple examinations, such as multi-parameter MRI and blood routine examinations, which yield multimodal data such as MRI images and clinical characteristics. However, due to personalized treatment and the availability of examination equipment, it is extremely common in clinical practice to lack certain modal data.

[0005] On the other hand, the model suffers from data offset between medical sites in prediction experiments at multiple medical sites. The external validation effect is far inferior to the internal validation effect. The data offset between medical sites refers to the statistical distribution offset between the collected image data due to differences in imaging equipment, imaging protocols, etc., which affects the model's performance in cross-site prediction scenarios.

[0006] Existing prognostic prediction technologies for nasopharyngeal carcinoma patients cannot cope with incomplete multimodal data and data offsets between medical sites. Therefore, how to invent a prognostic prediction method for nasopharyngeal carcinoma that can cope with incomplete multimodal data and data offsets between medical sites is a technical problem that urgently needs to be solved in this field. Summary of the Invention

[0007] To address the limitations of existing technologies in handling incomplete multimodal data and data offsets between medical sites, this invention provides a method for predicting the prognosis of nasopharyngeal carcinoma that can fuse incomplete multimodal data and span across medical sites.

[0008] To achieve the above-mentioned objectives of this invention, the technical solution adopted is as follows:

[0009] A prognostic prediction method for nasopharyngeal carcinoma that can integrate incomplete multimodal approaches and span multiple medical sites includes the following steps:

[0010] S1. Receive raw data including MRI images and clinical features, and extract their radiomics features to obtain multimodal data;

[0011] S2. Introduce a fully connected reconstruction network to perform multimodal data fusion and cross-site feature transformation to obtain the fused features of multimodal data;

[0012] S3. Introduce a multi-layer fully connected survival analysis network, input the fused features into the survival analysis network, and obtain the survival risk;

[0013] S4. Construct a cross-site partial contrastive learning model; the sites with survival events and survival time data labels in the multimodal data are called the source domain, and the sites without data labels and to be predicted by the model in the multimodal data are called the target domain; the partial contrastive learning model performs contrastive learning on the source domain, target domain, and cross-domain features of the fusion features of the multimodal data according to the different sources of the positively paired samples;

[0014] S5. Based on the reconstruction network, survival analysis network, and partial contrastive learning model, the total loss function is obtained, and gradient optimization is performed using PyTorch to train the reconstruction network and survival analysis network.

[0015] S6. Obtain fusion features and survival risk through the reconstructed network and survival analysis network trained in step S5, and further calculate and output similar sample pairing and individualized survival curves. Predict the prognosis of nasopharyngeal carcinoma through similar sample pairing and individualized survival curves.

[0016] Preferably, in step S1, the magnetic resonance imaging includes a head scan before induction chemotherapy, a head scan after induction chemotherapy, a neck scan before induction chemotherapy, and a neck scan after induction chemotherapy; the head scan before induction chemotherapy includes T1-weighted MRI and T1-weighted contrast-enhanced MRI; the head scan after induction chemotherapy includes T1-weighted MRI or T1-weighted contrast-enhanced MRI; the neck scan before induction chemotherapy includes T2-weighted MRI; the neck scan after induction chemotherapy includes T2-weighted MRI; the clinical data includes prognostic biomarkers, complete blood count, and history of smoking and alcohol consumption; the prognostic biomarkers include pEBV-DNA levels before and after induction chemotherapy, T stage, N stage, total stage, and age; the complete blood count includes hemoglobin, albumin, C-reactive protein, and lactate dehydrogenase.

[0017] Furthermore, in step S1, raw data including MRI images and clinical features are received, and their radiomics features are extracted to obtain multimodal data. The specific steps are as follows:

[0018] S101. The head scan before induction chemotherapy, the head scan after induction chemotherapy, the neck scan before induction chemotherapy, and the neck scan after induction chemotherapy are segmented to obtain MRI images with segmentation masks. The MRI images with segmentation masks include the primary lesion before induction chemotherapy, the primary lesion after induction chemotherapy, the lymph nodes before induction chemotherapy, and the lymph nodes after induction chemotherapy.

[0019] S102. Input the MRI images with segmentation masks into the Python package Pyradiomics to extract radiomics features; the radiomics features include three sets of features: shape features, first-order statistical features, and second-order statistical features. The shape features and first-order statistical features use default settings, while the second-order statistical features include features extracted using gray-level co-occurrence matrix, gray-level run-length matrix, gray-level size region matrix, and gray-level correlation matrix.

[0020] S103. Combine the patient's radiomics characteristics with their clinical data to form multimodal data, denoted as... Where M is the number of modalities, i is the patient number; let x be the sum of the modalities. i The corresponding modal indicator variable is a i =[a i1 ;…;a iM ], where a im =1 indicates that patient i has data from the m-th modality, a im =0 indicates that patient i is missing data for the m-th modality.

[0021] Furthermore, in step S2, a fully connected reconstruction network is introduced to perform multimodal data fusion and cross-site feature transformation through the reconstruction network to obtain the fused features of the multimodal data. The specific steps are as follows:

[0022] S201. Let n be the source domain of the multimodal data. s The dataset consists of samples. n on the target domain of multimodal data t The dataset consists of samples. δ i =1 indicates that patient i at the time of data collection has experienced local recurrence or died, and the corresponding y i Survival time; δ i =0 indicates that patient i at the time of data collection has neither carcinoma in situ nor invasive disease, and the corresponding y i For the time of deletion;

[0023] S202. Construct M parallel fully connected reconstruction networks for multimodal samples. and Reconstructing the network {R} using M parallel fully connected networks (1) ,…,R (M)}, to find the corresponding fusion feature f in the fusion space of the source domain dataset and the target domain dataset of multimodal data. i The objective function for reconstructing the network is:

[0024]

[0025] Where F = [f1,…,f n ] T For fusion feature f i The matrix formed, θ R For M parallel fully connected reconstructed networks, reconstruct all parameters to be optimized;

[0026] S203. Repeat S202 for all samples in the source and target domains of the multimodal data to obtain the fusion feature matrix F of the source domain of the multimodal data. s The fusion feature matrix F of the target domain and multimodal data t ,Right now and

[0027] S204. Based on the fusion feature matrices of the source and target domains of the multimodal data, calculate the cross-site attention matrix:

[0028]

[0029] Where d is f iThe dimension is , where softmax indicates that the matrix is ​​normalized by row;

[0030] S205. Based on the cross-site attention matrix A, the fusion feature matrix F of the target domain of the multimodal data. t The transformation is performed to obtain the fusion feature matrix of the source domain after transformation:

[0031]

[0032] The fusion feature matrix of the transformed source domain and the fusion feature matrix of the target domain of the multimodal data are used as the fusion features of the multimodal data.

[0033] Furthermore, in step S2, the objective function for reconstructing the network is also obtained; specifically:

[0034] The fused features of the transformed source domain are input into M parallel fully connected reconstruction networks {R}. (1) ,…,R (M) In}, the objective function for reconstructing the network is obtained:

[0035]

[0036] Furthermore, in step S3, a multi-layer fully connected survival analysis network is introduced, and the fused features are input into the survival analysis network to obtain the survival risk. The specific steps are as follows:

[0037] S301. Introduce a multi-layer fully connected survival analysis network to fuse the transformed source domain features. Fusion features of target domains and multimodal data The data are input into a multi-layer fully connected survival analysis network G to obtain the survival risk, i.e., the survival risk of the source domain. Survival risk of the target domain

[0038] S302. Construct a set U from the samples in the source domain of the multimodal data that have undergone local relapse or death, and input the sample features in this set into the negative partial log-likelihood function to obtain the objective function for survival analysis:

[0039]

[0040] Where Ω i Let U be the set of samples whose survival time is greater than that of sample i of type U.

[0041] Furthermore, in step S4, a cross-site partial contrastive learning model is constructed. The partial contrastive learning model performs comparative learning on the source domain, target domain, and cross-domain features of the fusion features of multimodal data according to the different sources of the positively paired samples, thereby enhancing the discriminativeness of the fusion features. The specific steps are as follows:

[0042] S401. Incorporate the fusion features of multimodal data In source domain samples i that have experienced local recurrence or death, the survival time exceeds The source domain samples are sorted in ascending order of survival time. In each contrastive learning process, pairs with small differences in survival time are identified as positive sample pairs, while pairs determined by the survival time inequality are identified as negative sample pairs.

[0043]

[0044] in and These correspond to the survival times of positive and negative samples, respectively.

[0045] S402. Obtain the fusion features of multimodal data through a fully connected survival analysis network G. The survival risk in source domain sample i and source domain positive sample j, which have experienced local recurrence or death, is lower than that of the sample i. target domain samples k is the negative sample in the target domain;

[0046] S403. Construct a cross-site partial contrastive learning model; combine positive sample pairs, negative sample pairs, and negative sample k to perform partial contrastive learning through the partial contrastive learning model. The objective function of the contrastive learning is:

[0047]

[0048] Where D s To find the number of terms in the summation, The similarity between two feature vectors is given by τ, where τ is the temperature parameter.

[0049] S404. For a positive sample pair of the target domain in a positive sample pair, the inequality 0≤r is used. i -r j <r i -r k The positive and negative sample pairs in the target domain are obtained by direct comparison; the positive and negative sample pairs in the target domain are then input into the partial contrastive learning model for partial contrastive learning.

[0050]

[0051] Where D t To find the number of terms in the summation;

[0052] S405. For cross-domain positive sample pairs in positive sample pairs, the source domain pairing of the cross-domain positive sample pair is identified by the survival time inequality, and the target domain pairing and cross-domain pairing of the cross-domain positive sample pairing are also identified; the identified source domain pairing, target domain pairing, and cross-domain pairing of the cross-domain positive sample pairing are input into the partial contrastive learning model for partial contrastive learning:

[0053]

[0054]

[0055] Where D st and D ts This is the number of terms to be summed.

[0056] Furthermore, in step S5, the total loss function obtained based on the reconstructed network, survival analysis network, and partial contrastive learning model is specifically as follows:

[0057] The total loss function is obtained based on the reconstructed network, survival analysis network, and partial contrastive learning model:

[0058]

[0059] Where λ1 and λ2 are non-negative hyperparameters.

[0060] Furthermore, in step S5, gradient optimization is performed using PyTorch to train the reconstruction network and the survival analysis network, specifically as follows:

[0061] We select a three-layer fully connected survival analysis network G with d=128 and dimensions of 64, 64, and 32, and the dimension is min{128, 32d}. m}、min{128,16d m}、min{128,8d m}、min{128,4d m The four-layer fully connected reconstructed network R (m) A learning rate of 0.0001, β1 = 0.5, β2 = 0.999 is used to optimize the AMS-Grad network with d1 = ... = d4 = 100, d5 = 6, d6 = 4, d7 = 2, τ = 0.5, λ1 = λ2 = 1000, and β1 = 0.5, β2 = 0.999. The backpropagation process is completed using machine learning libraries such as PyTorch to update the parameters of the survival analysis network and the fully connected reconstruction network.

[0062] Furthermore, in step S6, the fusion features and survival risk are obtained through the trained reconstruction network and survival analysis network, and the similar sample pairing and individualized survival curve are further calculated and output. The specific steps are as follows:

[0063] S601. Given source domain multimodal data Target domain multimodal data The input is fed into the fully connected reconstruction network trained in step S5 to obtain the fused features of the multimodal data. and calculate and Find the cosine similarity and select the pair with the largest similarity. As a pairing of similar samples across sites, using Survival annotation information guidance Treatment; calculation and Find the cosine similarity and select the pair with the largest similarity. As a pairing of similar samples across sites, using Survival annotation information guidance Treatment;

[0064] S602. Denote the survival times of samples that have experienced local relapse or death in the source domain of the multimodal data fusion features as K time points: y( 1) <… <y (K) Record time point y (i) The number of relapses or deaths was d i The risk function is calculated using the trained reconstruction network and survival analysis network:

[0065]

[0066] S603. Calculate the survival baseline function at any time point x. The individualized survival curve of the target sample j, which has the fusion features of multimodal data, is as follows: Where x>0 is the time variable.

[0067] The beneficial effects of this invention are as follows:

[0068] This invention designs a prognostic prediction method for nasopharyngeal carcinoma (NPC) that can fuse incomplete multimodal data across medical sites, addressing the issues of incomplete multimodal fusion and data offset between medical sites in NPC prognostic prediction tasks. This invention obtains multimodal data by receiving raw data including MRI images and clinical features. It introduces a fully connected reconstruction network to find fusion features from the fusion space, enabling these fusion features to be reconstructed into original multimodal features through multiple parallel fully connected reconstruction networks. This ensures that the fusion features contain multimodal information, achieving multimodal feature fusion based on cross-site attention. This method does not rely on the assumption of multimodal completeness, addressing the challenge of missing partial modality data in clinical practice. This invention also introduces a multi-layer fully connected survival analysis network to obtain survival risk. Furthermore, this invention proposes using a cross-site attention-based partial contrastive learning model to learn the similarity between cross-domain samples through an attention matrix and transform the target domain fusion feature matrix, ensuring that the transformed multi-site fusion features are in the same space, thus solving the problem of data offset between medical sites. Therefore, this invention effectively utilizes survival events, survival time, and survival risk information in the samples to improve the similarity of positive sample pairs and reduce the similarity of negative sample pairs, so that the cross-site fusion features are distributed in an orderly manner according to the risk level, ultimately improving the discriminativeness of the fusion feature space and achieving high-performance nasopharyngeal carcinoma prognosis prediction. Attached Figure Description

[0069] Figure 1 This is a flowchart of the nasopharyngeal carcinoma prognosis prediction method of the present invention, which can integrate incomplete multimodal approaches and is applicable across medical sites.

[0070] Figure 2 This is a schematic diagram of the module flow of the nasopharyngeal carcinoma prognosis prediction system of the present invention, which can integrate incomplete multimodal approaches and can be implemented across medical sites.

[0071] Figure 3 This is a simplified network structure diagram of the system in Example 3 of the present invention, which can integrate incomplete multimodal and cross-medical site nasopharyngeal carcinoma prognosis prediction methods. Detailed Implementation

[0072] The present invention will now be described in detail with reference to the accompanying drawings and specific embodiments.

[0073] Example 1

[0074] like Figure 1 As shown, a prognostic prediction method for nasopharyngeal carcinoma that can integrate incomplete multimodal approaches and span medical sites includes the following steps:

[0075] S1. Receive raw data including MRI images and clinical features, and extract their radiomics features to obtain multimodal data;

[0076] S2. Introduce a fully connected reconstruction network to perform multimodal data fusion and cross-site feature transformation to obtain the fused features of multimodal data;

[0077] S3. Introduce a multi-layer fully connected survival analysis network, input the fused features into the survival analysis network, and obtain the survival risk;

[0078] S4. Construct a cross-site partial contrastive learning model; the sites with survival events and survival time data labels in the multimodal data are called the source domain, and the sites without data labels and to be predicted by the model in the multimodal data are called the target domain; the partial contrastive learning model performs contrastive learning on the source domain, target domain, and cross-domain features of the fusion features of the multimodal data according to the different sources of the positively paired samples;

[0079] S5. Based on the reconstruction network, survival analysis network, and partial contrastive learning model, the total loss function is obtained, and gradient optimization is performed using PyTorch to train the reconstruction network and survival analysis network.

[0080] S6. Obtain fusion features and survival risk through the reconstructed network and survival analysis network trained in step S5, and further calculate and output similar sample pairing and individualized survival curves. Predict the prognosis of nasopharyngeal carcinoma through similar sample pairing and individualized survival curves.

[0081] Example 2

[0082] More specifically, in this embodiment, the prognosis of nasopharyngeal carcinoma patients A, B, and C from two sites, Hospital A and Hospital B, was predicted using the method of the present invention. The data from Hospital A included survival annotation information, while the data from Hospital B did not. The data from Hospital B is referred to as test data.

[0083] In one specific embodiment, in step S1, the magnetic resonance imaging includes a head scan before induction chemotherapy, a head scan after induction chemotherapy, a neck scan before induction chemotherapy, and a neck scan after induction chemotherapy; the head scan before induction chemotherapy includes T1-weighted MRI images and T1-weighted contrast-enhanced images; the head scan after induction chemotherapy includes T1-weighted MRI images or T1-weighted contrast-enhanced images; the neck scan before induction chemotherapy includes T2-weighted images; the neck scan after induction chemotherapy includes T2-weighted images; the clinical data includes prognostic biomarkers, complete blood count, and history of smoking and alcohol consumption; the prognostic biomarkers include pEBV-DNA levels before and after induction chemotherapy, T stage, N stage, total stage, and age; the complete blood count includes hemoglobin, albumin, C-reactive protein, and lactate dehydrogenase.

[0084] In one specific embodiment, step S1 involves receiving raw data including MRI images and clinical features, and extracting their radiomics features to obtain multimodal data. The specific steps are as follows:

[0085] S101. The head scan before induction chemotherapy, the head scan after induction chemotherapy, the neck scan before induction chemotherapy, and the neck scan after induction chemotherapy are segmented to obtain MRI images with segmentation masks. The MRI images with segmentation masks include the primary lesion before induction chemotherapy, the primary lesion after induction chemotherapy, the lymph nodes before induction chemotherapy, and the lymph nodes after induction chemotherapy.

[0086] S102. Input the MRI images with segmentation masks into the Python package Pyradiomics to extract radiomics features; the radiomics features include three sets of features: shape features, first-order statistical features, and second-order statistical features. The shape features and first-order statistical features use default settings, while the second-order statistical features include features extracted using gray-level co-occurrence matrix, gray-level run-length matrix, gray-level size region matrix, and gray-level correlation matrix.

[0087] S103. Combine the patient's radiomics characteristics with their clinical data to form multimodal data, denoted as... Where M is the number of modalities, i is the patient number; let x be the sum of the modalities. i The corresponding modal indicator variable is a i =[a i1 ;…;a iM ], where a im =1 indicates that patient i has data from the m-th modality, a im =0 indicates that patient i is missing data for the m-th modality.

[0088] In one specific embodiment, in step S2, a fully connected reconstruction network is introduced, and multimodal data fusion and cross-site feature transformation are performed through the reconstruction network to obtain the fused features of multimodal data. The specific steps are as follows:

[0089] S201. Let n be the source domain of the multimodal data. s The dataset consists of samples. n on the target domain of multimodal data t The dataset consists of samples. δ i =1 indicates that patient i at the time of data collection has experienced local recurrence or died, and the corresponding y i Survival time; δ i =0 indicates that patient i at the time of data collection has neither carcinoma in situ nor invasive disease, and the corresponding y i For the time of deletion;

[0090] S202. Construct M parallel fully connected reconstruction networks for multimodal samples. and Reconstructing the network {R} using M parallel fully connected networks (1) ,…,R (M)}, to find the corresponding fusion feature f in the fusion space of the source domain dataset and the target domain dataset of multimodal data. i The objective function for reconstructing the network is:

[0091]

[0092] Where F = [f1,…,f n ] T For fusion feature f i The matrix formed, θ R For M parallel fully connected reconstructed networks, reconstruct all parameters to be optimized;

[0093] S203. Repeat S202 for all samples in the source and target domains of the multimodal data to obtain the fusion feature matrix F of the source domain of the multimodal data. s The fusion feature matrix F of the target domain and multimodal data t ,Right now and

[0094] S204. Based on the fusion feature matrices of the source and target domains of the multimodal data, calculate the cross-site attention matrix:

[0095]

[0096] Where d is f i The dimension is , where softmax indicates that the matrix is ​​normalized by row;

[0097] S205. Based on the cross-site attention matrix A, the fusion feature matrix F of the target domain of the multimodal data. t The transformation is performed to obtain the fusion feature matrix of the source domain after transformation:

[0098]

[0099] The fusion feature matrix of the transformed source domain and the fusion feature matrix of the target domain of the multimodal data are used as the fusion features of the multimodal data.

[0100] In one specific embodiment, step S2 further yields the objective function for reconstructing the network; specifically:

[0101] The fused features of the transformed source domain are input into M parallel fully connected reconstruction networks {R}.(1) ,…,R (M) In}, the objective function for reconstructing the network is obtained:

[0102]

[0103] In one specific embodiment, in step S3, a multi-layer fully connected survival analysis network is introduced, and the fused features are input into the survival analysis network to obtain the survival risk. The specific steps are as follows:

[0104] S301. Introduce a multi-layer fully connected survival analysis network to fuse the transformed source domain features. Fusion features of target domains and multimodal data The data are input into a multi-layer fully connected survival analysis network G to obtain the survival risk, i.e., the survival risk of the source domain. Survival risk of the target domain

[0105] S302. Construct a set U from the samples in the source domain of the multimodal data that have undergone local relapse or death, and input the sample features in this set into the negative partial log-likelihood function to obtain the objective function for survival analysis:

[0106]

[0107] Where Ω i Let U be the set of samples whose survival time is greater than that of sample i of type U.

[0108] In one specific embodiment, in step S4, a cross-site partial contrastive learning model is constructed; the partial contrastive learning model performs comparative learning on the source domain, target domain, and cross-domain features of the fusion features of multimodal data according to the different sources of the positively paired samples, thereby enhancing the discriminativeness of the fusion features. The specific steps are as follows:

[0109] S401. Incorporate the fusion features of multimodal data In source domain samples i that have experienced local recurrence or death, the survival time exceeds The source domain samples are sorted in ascending order of survival time. In each contrastive learning process, pairs with small differences in survival time are identified as positive sample pairs, while pairs determined by the survival time inequality are identified as negative sample pairs.

[0110]

[0111] in and These correspond to the survival times of positive and negative samples, respectively.

[0112] S402. Obtain the fusion features of multimodal data through a fully connected survival analysis network G. The survival risk in source domain sample i and source domain positive sample j, which have experienced local recurrence or death, is lower than that of the sample i. target domain samples k is the negative sample in the target domain;

[0113] S403. Construct a cross-site partial contrastive learning model; combine positive sample pairs, negative sample pairs, and negative sample k to perform partial contrastive learning through the partial contrastive learning model. The objective function of the contrastive learning is:

[0114]

[0115] Where D s To find the number of terms in the summation, The similarity between two feature vectors is given by τ, where τ is the temperature parameter.

[0116] S404. For a positive sample pair of the target domain in a positive sample pair, the inequality 0≤r is used. i -r j <r i -r k The positive and negative sample pairs in the target domain are obtained by direct comparison; the positive and negative sample pairs in the target domain are then input into the partial contrastive learning model for partial contrastive learning.

[0117]

[0118] Where D t To find the number of terms in the summation;

[0119] S405. For cross-domain positive sample pairs in positive sample pairs, the source domain pairing of the cross-domain positive sample pair is identified by the survival time inequality, and the target domain pairing and cross-domain pairing of the cross-domain positive sample pairing are also identified; the identified source domain pairing, target domain pairing, and cross-domain pairing of the cross-domain positive sample pairing are input into the partial contrastive learning model for partial contrastive learning:

[0120]

[0121]

[0122] Where D st and D ts This is the number of terms to be summed.

[0123] In one specific embodiment, in step S5, the total loss function is obtained based on the reconstruction network, the survival analysis network, and the partial contrastive learning model, and gradient optimization is performed using PyTorch to train the reconstruction network and the survival analysis network.

[0124] S501. The total loss function is obtained based on the reconstruction network, survival analysis network, and partial contrastive learning model:

[0125]

[0126] Where λ1 and λ2 are non-negative hyperparameters.

[0127] S502. Select a three-layer fully connected survival analysis network G with d = 128 and dimensions of 64, 64, and 32, where the dimension is min{128, 32d}. m}、min{128,16d m}、min{128,8d m}、min{128,4d m The four-layer fully connected reconstructed network R (m) A learning rate of 0.0001, β1 = 0.5, β2 = 0.999 is used to optimize the AMS-Grad network with d1 = ... = d4 = 100, d5 = 6, d6 = 4, d7 = 2, τ = 0.5, λ1 = λ2 = 1000, and β1 = 0.5, β2 = 0.999. The backpropagation process is completed using machine learning libraries such as PyTorch to update the parameters of the survival analysis network and the fully connected reconstruction network.

[0128] In one specific embodiment, step S6 involves obtaining fusion features and survival risk through the trained reconstruction network and survival analysis network, and further calculating and outputting similar sample pairing and individualized survival curves. The specific steps are as follows:

[0129] S601. Given source domain multimodal data Target domain multimodal data The input is fed into the fully connected reconstruction network trained in step S5 to obtain the fused features of the multimodal data. and calculate and Find the cosine similarity and select the pair with the largest similarity. As a pairing of similar samples across sites, using Survival annotation information guidance Treatment; calculation and Find the cosine similarity and select the pair with the largest similarity. As a pairing of similar samples across sites, using Survival annotation information guidance Treatment;

[0130] S602. Record the survival times of samples that have experienced local relapse or death in the source domain of the multimodal data fusion features as K time points: y(1) <… <y (K) Record time point y (i) The number of relapses or deaths was d i The risk function is calculated using the trained reconstruction network and survival analysis network:

[0131]

[0132] S603. Calculate the survival baseline function at any time point x. The individualized survival curve of the target sample j, which has the fusion features of multimodal data, is as follows: Where x>0 is the time variable.

[0133] This invention designs a prognostic prediction method for nasopharyngeal carcinoma (NPC) that can fuse incomplete multimodal data across medical sites, addressing the issues of incomplete multimodal fusion and data offset between medical sites in NPC prognostic prediction tasks. This invention obtains multimodal data by receiving raw data including MRI images and clinical features. It introduces a fully connected reconstruction network to find fusion features from the fusion space, enabling these fusion features to be reconstructed into original multimodal features through multiple parallel fully connected reconstruction networks. This ensures that the fusion features contain multimodal information, achieving multimodal feature fusion based on cross-site attention. This method does not rely on the assumption of multimodal completeness, addressing the challenge of missing partial modality data in clinical practice. This invention also introduces a multi-layer fully connected survival analysis network to obtain survival risk. Furthermore, this invention proposes using a cross-site attention-based partial contrastive learning model to learn the similarity between cross-domain samples through an attention matrix and transform the target domain fusion feature matrix, ensuring that the transformed multi-site fusion features are in the same space, thus solving the problem of data offset between medical sites. Therefore, this invention effectively utilizes survival events, survival time, and survival risk information in the samples to improve the similarity of positive sample pairs and reduce the similarity of negative sample pairs, so that the cross-site fusion features are distributed in an orderly manner according to the risk level, ultimately improving the discriminativeness of the fusion feature space and achieving high-performance nasopharyngeal carcinoma prognosis prediction.

[0134] Example 3

[0135] A prognostic prediction method for nasopharyngeal carcinoma that can integrate incomplete multimodal approaches and span multiple medical sites includes the following steps:

[0136] S1. Receive raw data including MRI images and clinical features, and extract their radiomics features to obtain multimodal data;

[0137] S2. Introduce a fully connected reconstruction network to perform multimodal data fusion and cross-site feature transformation to obtain the fused features of multimodal data;

[0138] S3. Introduce a multi-layer fully connected survival analysis network, input the fused features into the survival analysis network, and obtain the survival risk;

[0139] S4. Construct a cross-site partial contrastive learning model; the sites with survival events and survival time data labels in the multimodal data are called the source domain, and the sites without data labels and to be predicted by the model in the multimodal data are called the target domain; the partial contrastive learning model performs contrastive learning on the source domain, target domain, and cross-domain features of the fusion features of the multimodal data according to the different sources of the positively paired samples;

[0140] S5. Based on the reconstruction network, survival analysis network, and partial contrastive learning model, the total loss function is obtained, and gradient optimization is performed using PyTorch to train the reconstruction network and survival analysis network.

[0141] S6. Obtain fusion features and survival risk through the reconstructed network and survival analysis network trained in step S5, and further calculate and output similar sample pairing and individualized survival curves. Predict the prognosis of nasopharyngeal carcinoma through similar sample pairing and individualized survival curves.

[0142] In this embodiment, as Figure 2 As shown, based on the nasopharyngeal carcinoma prognostic prediction method that can fuse incomplete multimodal features and cross medical sites, a nasopharyngeal carcinoma prognostic prediction system that can fuse incomplete multimodal features and cross medical sites is also constructed, including a data input module, a multimodal feature fusion module, a deep survival analysis module, a cross-site partial contrastive learning module, a loss module, and a data output module.

[0143] In this embodiment, the prognostic prediction system of the present invention, which can integrate incomplete multimodal data and can be used across medical sites, was applied to nasopharyngeal carcinoma patients at two sites, Hospital A and Hospital B, for prognostic prediction. The data from Hospital A includes survival labeling information, while the data from Hospital B does not. The data from Hospital B is referred to as test data.

[0144] like Figure 3 As shown, the data input module is used to receive raw data including MRI images and clinical features, and extract their radiomics features to obtain multimodal data;

[0145] The multimodal feature fusion module is used to introduce a fully connected reconstruction network, and through the reconstruction network, multimodal data fusion and cross-site feature transformation are performed to obtain the fused features of the multimodal data;

[0146] The deep survival analysis module is used to introduce a multi-layer fully connected survival analysis network, input the fused features into the survival analysis network, and obtain the survival risk.

[0147] The cross-site partial contrastive learning module is used to construct a cross-site partial contrastive learning model. Sites with survival events and survival time data labels in multimodal data are called source domains, and sites without data labels and to be predicted by the model are called target domains. The partial contrastive learning model performs comparative learning on the source domain, target domain, and cross-domain features of the fusion features of multimodal data according to the different sources of positively paired samples, thereby enhancing the discriminativeness of the fusion features.

[0148] The loss module is used to obtain the total loss function based on the reconstruction network, survival analysis network, and partial contrastive learning model, and to use PyTorch for gradient optimization to train the reconstruction network and survival analysis network.

[0149] The data output module is used to obtain fusion features and survival risk through the trained reconstruction network and survival analysis network, and further calculate and output similar sample pairing and individualized survival curves. The prognosis of nasopharyngeal carcinoma is predicted through similar sample pairing and individualized survival curves.

[0150] To address the common challenges of incomplete multimodal fusion and data offset between medical sites in nasopharyngeal carcinoma prognostic prediction tasks, a multimodal feature fusion module based on cross-site attention and a cross-site biased contrastive learning module were designed and embedded into a deep survival analysis network.

[0151] In this embodiment, the nasopharyngeal carcinoma prognostic prediction system of the present invention, which can integrate incomplete multimodal data and can be used across medical sites, was used to predict the prognosis of nasopharyngeal carcinoma patients at three sites: Hospital A, Hospital B, and Hospital C.

[0152] The specific implementation method is as follows:

[0153] P1. Extract radiomics features from MRI images from multiple sites through the data input module, and then cascade them with clinical features to form multimodal features;

[0154] P2. Multimodal data fusion and cross-site feature transformation are performed through the multimodal feature fusion module to obtain the fused features of the multimodal data;

[0155] P3. Survival risk is obtained by performing survival analysis on the fused features of multimodal data through the deep survival analysis module;

[0156] P4. Multimodal fusion features are learned through a cross-site partial contrastive learning module;

[0157] P5. The loss module is used to train the reconstruction network and the survival analysis network based on the total loss function obtained from the reconstruction network, the survival analysis network, and the partial contrastive learning model.

[0158] P6. The data output module obtains the final similar sample pairing and individualized survival curves through the trained reconstruction network and survival analysis network.

[0159] The results of implementing this invention and traditional deep survival analysis methods in a cross-medical site nasopharyngeal carcinoma prognostic prediction task are shown in the table below:

[0160]

[0161] Where A->B represents the cross-site prognostic prediction task from source domain A to target domain B; the numbers in the table are C exponents, and the larger the value, the better the corresponding method implementation results.

[0162] The table above leads to the following conclusions: The deep survival analysis method based on reconstruction fusion outperforms the deep survival analysis method based on cascade fusion, verifying that the present invention can better integrate multimodal information in the diagnosis and treatment process; compared with other cross-domain prognostic prediction methods (i.e., the combination of deep survival analysis methods with MMD, DA, CKB, and ATM regularization terms), the present invention achieves a higher C-index, indicating that the present invention better solves the problem of data offset between medical sites across domains; the present invention greatly improves the prognostic prediction effect of traditional deep survival analysis, which verifies the effectiveness of the multimodal feature fusion module based on cross-site attention and the cross-site biased contrastive learning module.

[0163] The system constructed based on this invention does not rely on the assumption of completeness of multimodal data during diagnosis and treatment, and can naturally solve the situation of missing one or more modal data in clinical practice. Experiments show that compared with the fusion method of direct cascading after zero-padding, the multimodal fusion module improves the C-index by 1.5%-7.3%. It uses attention matrix to realize cross-domain fusion feature transformation, and through partial contrastive learning, it makes cross-site fusion features distributed in an orderly manner according to risk level. Therefore, it can learn cross-site, highly discriminative fusion features, which improves the C-index by 6.2%-19.3% compared with traditional deep survival analysis methods.

[0164] This invention achieved excellent prognostic prediction results for nasopharyngeal carcinoma in cross-site data validation experiments, and has good clinical application value. Compared with other cross-domain prognostic prediction methods, this invention improved the C-index by 2.4%-14.8%.

[0165] Obviously, the above embodiments of the present invention are merely examples for clearly illustrating the present invention, and are not intended to limit the implementation of the present invention. Any modifications, equivalent substitutions, and improvements made within the spirit and principles of the present invention should be included within the protection scope of the claims of the present invention.

Claims

1. A prognostic prediction method for nasopharyngeal carcinoma that can integrate incomplete multimodal data and is applicable across medical sites, characterized by: Includes the following steps: S1. Receive raw data including MRI images and clinical features, and extract their radiomics features to obtain multimodal data; S2. Introduce a fully connected reconstruction network to perform multimodal data fusion and cross-site feature transformation to obtain the fused features of multimodal data; The specific steps are as follows: S201. Denote the source domain of multimodal data. The dataset consists of samples. On the target domain of multimodal data The dataset consists of samples. ; This indicates the patients whose collection time was recorded. The corresponding state is one of local recurrence or death. For survival time; This indicates the patients whose collection time was recorded. The state of having neither carcinoma in situ nor invasive disease, corresponding to For the time of deletion; S202. Construction A parallel fully connected reconstruction network for multimodal samples and ,use A parallel fully connected reconfiguration network To find the corresponding fusion features in the fusion space of the source domain dataset and the target domain dataset of multimodal data. The objective function for reconstructing the network is: in For fusion features The matrix formed for All parameters to be optimized in a parallel, fully connected reconstructed network; S203. Repeat S202 for all samples in the source and target domains of the multimodal data to obtain the fusion feature matrix F of the source domain of the multimodal data. s The fusion feature matrix F of the target domain and multimodal data t ,Right now and ; S204. Based on the fusion feature matrices of the source and target domains of the multimodal data, calculate the cross-site attention matrix: in for The dimension is , where softmax indicates that the matrix is ​​normalized by row; S205. Based on the cross-site attention matrix The fusion feature matrix of the target domain for multimodal data The transformation is performed to obtain the fusion feature matrix of the source domain after transformation: The fusion feature matrix of the transformed source domain and the fusion feature matrix of the target domain of the multimodal data are used as the fusion features of the multimodal data. S3. Introduce a multi-layer fully connected survival analysis network, input the fused features into the survival analysis network, and obtain the survival risk; S4. Construct a cross-site partial contrastive learning model; the sites in the multimodal data that have survival events and survival time data labels are called the source domain, and the sites in the multimodal data that do not have data labels and are to be predicted by the model are called the target domain; the partial contrastive learning model performs contrastive learning on the source domain, target domain, and cross-domain features of the fusion features of the multimodal data according to the different sources of the positively paired samples. The specific steps are as follows: S401. Incorporate the fusion features of multimodal data Source domain samples that have experienced local recurrence or death Survival time in the middle exceeds The source domain samples are sorted in ascending order of survival time. In each contrastive learning process, pairs with small differences in survival time are identified as positive sample pairs, while pairs determined by the survival time inequality are identified as negative sample pairs. in and These correspond to the survival times of positive and negative samples, respectively. ; S402. Through a fully connected survival analysis network In obtaining the fusion features of multimodal data Source domain samples that have experienced local recurrence or death and source domain positive samples The risk of survival is lower than target domain samples As negative samples of the target domain ; S403. Construct a cross-site partial contrastive learning model; Combining positive sample pairs, negative sample pairs, and negative samples Partial contrastive learning is performed using a partial contrastive learning model. The objective function for contrastive learning is: in To find the number of terms in the summation, The similarity between two feature vectors. For temperature parameters; S404. For positive sample pairs of the target domain in a positive sample pair, the inequality is used to... The positive and negative sample pairs in the target domain are obtained by direct comparison; the positive and negative sample pairs in the target domain are then input into the partial contrastive learning model for partial contrastive learning. in To find the number of terms in the summation; S405. For cross-domain positive sample pairs in positive sample pairs, the source domain pairing of the cross-domain positive sample pair is identified by the survival time inequality, and the target domain pairing and cross-domain pairing of the cross-domain positive sample pairing are also identified; the identified source domain pairing, target domain pairing, and cross-domain pairing of the cross-domain positive sample pairing are input into the partial contrastive learning model for partial contrastive learning: in and To find the number of terms in the summation; S5. Based on the reconstruction network, survival analysis network, and partial contrastive learning model, the total loss function is obtained, and gradient optimization is performed using PyTorch to train the reconstruction network and survival analysis network. S6. Obtain fusion features and survival risk through the reconstructed network and survival analysis network trained in step S5, and further calculate and output similar sample pairing and individualized survival curves. Predict the prognosis of nasopharyngeal carcinoma through similar sample pairing and individualized survival curves.

2. The method for predicting the prognosis of nasopharyngeal carcinoma according to claim 1, characterized in that: In step S1, the MRI images include a head scan before induction chemotherapy, a head scan after induction chemotherapy, a neck scan before induction chemotherapy, and a neck scan after induction chemotherapy; the head scan before induction chemotherapy includes T1-weighted MRI images and T1-weighted contrast-enhanced images; the head scan after induction chemotherapy includes either T1-weighted MRI images or T1-weighted contrast-enhanced images; the neck scan before induction chemotherapy includes T2-weighted images; the neck scan after induction chemotherapy includes T2-weighted images; the clinical data include prognostic biomarkers, complete blood count (CBC) tests, and smoking and alcohol history; the prognostic biomarkers include pEBV-DNA levels before and after induction chemotherapy, T stage, N stage, total stage, and age; the CBC tests include hemoglobin, albumin, C-reactive protein, and lactate dehydrogenase.

3. The method for predicting the prognosis of nasopharyngeal carcinoma according to claim 2, characterized in that: In step S1, raw data including MRI images and clinical features are received, and radiomics features are extracted to obtain multimodal data. The specific steps are as follows: S101. The head scan before induction chemotherapy, the head scan after induction chemotherapy, the neck scan before induction chemotherapy, and the neck scan after induction chemotherapy are segmented to obtain MRI images with segmentation masks. The MRI images with segmentation masks include the primary lesion before induction chemotherapy, the primary lesion after induction chemotherapy, the lymph nodes before induction chemotherapy, and the lymph nodes after induction chemotherapy. S102. Input the NMR images with segmentation masks into the Python package Pyradiomics to extract radiomics features; the radiomics features include three sets of features: shape features, first-order statistical features, and second-order statistical features. The shape features and first-order statistical features use default settings, while the second-order statistical features include features extracted using gray-level co-occurrence matrix, gray-level run-length matrix, gray-level size region matrix, and gray-level correlation matrix. S103. Combine the patient's radiomics characteristics with their clinical data to form multimodal data, denoted as... ,in For the number of modes, Number the patients; record and The corresponding modal indicator variable is ,in Indicates the patient Having the first Data for each modality, Indicates the patient Missing the first Data for each modality.

4. The method for predicting the prognosis of nasopharyngeal carcinoma according to claim 3, characterized in that: In step S2, the objective function for reconstructing the network is also obtained; specifically: The fused features of the transformed source domain are input into... A parallel fully connected reconfiguration network In this process, the objective function for reconstructing the network is obtained: 。 5. The method for predicting the prognosis of nasopharyngeal carcinoma according to claim 4, characterized in that: In step S3, a multi-layer fully connected survival analysis network is introduced, and the fused features are input into the survival analysis network to obtain the survival risk. The specific steps are as follows: S301. Introduce a multi-layer fully connected survival analysis network to fuse the transformed source domain features. Fusion features of target domains and multimodal data The inputs are fed into a multi-layer fully connected survival analysis network. In this process, we obtain the survival risk, which is the survival risk of the source domain. Survival risk of the target domain ; S302. Construct a set of samples from the source domain of the multimodal data that have experienced local recurrence or death. The sample features from this set are then input into the negative partial log-likelihood function to obtain the objective function for survival analysis: in For a survival time greater than the set Type of sample A sample set of survival times.

6. The method for predicting the prognosis of nasopharyngeal carcinoma according to claim 5, characterized in that: In step S5, the total loss function is obtained based on the reconstruction network, survival analysis network, and partial contrastive learning model, specifically as follows: The total loss function is obtained based on the reconstructed network, survival analysis network, and partial contrastive learning model: in and For non-negative hyperparameters, , .

7. The method for predicting the prognosis of nasopharyngeal carcinoma according to claim 6, characterized in that: In step S5, gradient optimization is performed using PyTorch to train the reconstruction network and the survival analysis network, specifically as follows: Select , dimension is , , Three-layer fully connected survival analysis network , dimension , , , Four-layer fully connected reconstructed network , , , , , , The learning rate is , , The AMS-Grad optimizer utilizes machine learning libraries such as PyTorch to complete the backpropagation process and update the parameters of the survival analysis network and the fully connected reconstruction network.

8. The method for predicting the prognosis of nasopharyngeal carcinoma according to claim 7, characterized in that: Step S6 involves obtaining fusion features and survival risk through the trained reconstruction network and survival analysis network, and further calculating and outputting similar sample pairing and individualized survival curves. The specific steps are as follows: S601. Given source domain multimodal data Target domain multimodal data The data is input into the fully connected reconstruction network trained in step S5 to obtain the fused features of the multimodal data. and ;calculate and Find the cosine similarity and select the pair with the largest similarity. As a pairing of similar samples across sites, using Survival annotation information guidance Treatment; S602. The survival time of samples in the source domain of multimodal data fusion features that have experienced local relapse or death is denoted as... Points in time: Record the time points The number of relapses or deaths was The risk function is calculated using the trained reconstruction network and survival analysis network: , ; S603. Calculate any time point Survival baseline function Target samples of multimodal data fusion features are obtained. The individualized survival curve is ,in It is a time variable.

Citation Information

Patent Citations

  • Multi-modal target detection method and system suitable for modal deficiency

    CN114359586A

  • Detecting anomalies during operation of a computer system based on multimodal data

    US20200336500A1