Multi-sequence fusion network model training method and device based on cell habitat and medium
By constructing a multi-sequence fusion network model, using cell habitat input data to train and screen the best cell habitat, combined with different MRI sequence characteristics, the problem of uncaptured internal tumor heterogeneity in the prior art is solved, and high-accurate image recognition of glioma grading and IDH mutation status prediction is achieved.
Patent Information
- Application Number
- CN202510321194.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-18
- Publication Date
- 2025-07-22
AI Technical Summary
The prior art ignores the internal biological heterogeneity of the tumor in glioma grading and IDH mutation status prediction based on MRI imaging omists, resulting in insufficient image recognition specificity and high misjudgment rate. Parameter analysis relying on specific imaging protocols cannot fully characterize the tumor microenvironment, affecting the accuracy rate.
By constructing a multi-sequence fusion network model, using cell habitat input data to train and screen the best cell habitat, combining different MRI sequence characteristics, and using a weighted objective function optimization model, the decoupling and complementary enhancement of cross-sequence common features and sequence-specific information is achieved, and image recognition accuracy is improved.
The image recognition accuracy of glioma grading and IDH mutation status prediction was improved, spatial localization and biological explanatory limitations were overcome, and a high robustness and strong generalization model was constructed, and the most biologically discriminant cell habitat-sequence combination was adaptively selected.
Smart Images

Figure CN120356027A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of medical image processing, and particularly to a method, device, and medium for training a multi-sequence fusion network model based on cell niche. Background Art
[0002] With the progress of medical imaging technology and the maturity of radiomics analysis methods, it has become possible to non-invasively grade gliomas and evaluate molecular subtypes based on MRI image features. Achieving accurate grading of adult diffuse gliomas (ADG) and predicting IDH mutation status through non-invasive means has important clinical value in avoiding the risks of invasive biopsies, guiding the formulation of individualized treatment plans, and improving patient prognosis assessment.
[0003] However, the image recognition methods based on traditional MRI radiomics mainly face two key problems: 1. Only using the information of the whole tumor region and easily ignoring the internal biological heterogeneity of the tumor, resulting in insufficient specificity and a high misjudgment rate in image recognition; 2. Existing niche imaging technologies overly rely on the parameter analysis of special imaging protocols (such as DKI, DSC-MRI), and the simple statistical features extracted by them are difficult to comprehensively characterize the tumor microenvironment, resulting in limited feature expression ability and ultimately affecting the accuracy of image recognition. Summary of the Invention
[0004] The present invention provides a method, device, and medium for training a multi-sequence fusion network model based on cell niche, which can improve the accuracy of image recognition in ADG grading and IDH mutation status prediction.
[0005] In a first aspect, an embodiment of the present invention provides a method for training a model of a multi-sequence fusion network, including:
[0006] Constructing cell niche input data corresponding to each cell niche according to the first MRI sequence input data; wherein, the first MRI sequence input data includes a plurality of MRI sequences;
[0007] Using each of the cell niche input data as an input, respectively iteratively training the corresponding first multi-sequence fusion network model, outputting the first multi-sequence fusion network models corresponding to each of the cell niche input data, and forming a first multi-sequence fusion network model set with all the first multi-sequence fusion network models;
[0008] Screening the first multi-sequence fusion network model set to determine the optimal cell niche;
[0009] Performing free combination on each MRI sequence in the first cell niche input data according to the first cell niche input data corresponding to the optimal cell niche to obtain a first MRI sequence combination data set;
[0010] Taking each piece of the first MRI sequence combination data in the first MRI sequence combination data set as an input, iteratively training the corresponding second multi-sequence fusion network model respectively, and outputting the second multi-sequence fusion network model corresponding to each piece of the first MRI sequence combination data;
[0011] Screening each of the second multi-sequence fusion network models, and outputting an optimal multi-sequence fusion network model.
[0012] In the embodiments of the present application, by constructing cell habitat input data corresponding to each habitat, the internal heterogeneity of tumors can be better captured, providing a basis for model training. Then, based on multiple cell habitat input data, the present application trains and screens to determine the optimal cell habitat, so as to accurately lock in the core region of tumor heterogeneity expression. Next, the present application optimally selects the optimal MRI sequence combination in the optimal cell habitat, which can break through the limitation of the information dimension of a single sequence and at the same time avoid the noise interference of low-correlation sequences. Compared with the prior art, the embodiments of the present application adaptively select the most biologically discriminative cell habitat-sequence combination on the basis of ensuring the accurate analysis of the internal heterogeneity of tumors, thus combining the spatial distribution law of cell density and the characteristics of different MRI sequences, constructing a model with high robustness and strong generalization ability, and further improving the accuracy of image recognition.
[0013] As a preferred example of the first aspect, the constructing cell habitat input data corresponding to each cell habitat according to the first MRI sequence input data includes:
[0014] Performing region extraction on the first T2-FLAIR data according to a preset region of interest drawing rule to obtain second T2-FLAIR data; wherein, the first MRI sequence input data includes the first T2-FLAIR data;
[0015] Aligning the first MRI sequence input data and the preset first DWI sequence data to the T2-FLAIR space of the second T2-FLAIR data through a non-linear registration method to obtain second MRI sequence input data and second DWI sequence data;
[0016] Determining cell habitat input data corresponding to each cell habitat according to the second DWI sequence data, the second MRI sequence input data, and the K-means clustering algorithm.
[0017] In this preferred example, through T2-FLAIR space registration, the spatial consistency of MRI data is ensured. Secondly, constructing data under different cell habitats can better capture the internal heterogeneity of tumors, overcoming the limitations of the prior art in terms of spatial positioning and biological interpretation.
[0018] As a preferred example of the first aspect, when any one of the cell microscopy input data in each of the cell microscopy input data is used as input, the step of using each of the cell microscopy input data as input and iteratively training the corresponding first multi-sequence fusion network model respectively to output the first multi-sequence fusion network model corresponding to each of the cell microscopy input data includes:
[0019] Construct independent encoders for each MRI sequence in the cell microscopy input data, and generate a set of shared representations between sequences and a set of unique representations within sequences through each of the encoders;
[0020] Calculate the average value of all elements in the set of shared representations between sequences to obtain a fused shared representation;
[0021] Input the set of unique representations within sequences and the fused shared representation into a decoder network and a two-layer fully connected prediction network, and calculate the first total loss through a weighted objective function. When the current iteration meets the first preset condition, output the first multi-sequence fusion network model corresponding to the cell microscopy input data.
[0022] In this preferred example, by splitting the set of latent representation distributions into a set of shared representation distributions between sequences and a set of unique representation distributions within sequences, the decoupling and complementary enhancement of MRI sequence information can be achieved: capturing cross-sequence common features through the shared representation distribution, eliminating redundant noise while strengthening the expression of biological essential features; while the unique representation distribution retains the specific pathological information of each sequence, preventing key detailed information from being diluted during the multi-modal fusion process, thereby improving the accuracy of image recognition.
[0023] As a preferred example of the first aspect, when any one of the first MRI sequence combination data sets in the first MRI sequence combination data set is used as input, the step of using each of the first MRI sequence combination data in the first MRI sequence combination data set as input and iteratively training the corresponding second multi-sequence fusion network model respectively includes:
[0024] Construct independent encoders for each MRI sequence in the first MRI sequence combination data, and generate a set of shared representations between sequences and a set of unique representations within sequences through each of the encoders;
[0025] Calculate the average value of all elements in the set of shared representations between sequences to obtain a fused shared representation;
[0026] Input the set of unique representations within sequences and the fused shared representation into a decoder network and a two-layer fully connected prediction network, and calculate the second total loss through a weighted objective function. When the current iteration meets the second preset condition, output the second multi-sequence fusion network model corresponding to the first MRI sequence combination data.
[0027] In this preferred example, by splitting the set of latent representation distributions into a set of inter-sequence shared representation distributions and a set of intra-sequence specific representation distributions, the decoupling and complementary enhancement of MRI sequence information can be achieved: the cross-sequence common features are captured by the shared representation distributions, while the redundant noise is eliminated and the expression of the biological essential features is strengthened; the specific representation distributions, on the other hand, retain the specific pathological information of each sequence, preventing the dilution of key detail information during the multi-modal fusion process, thereby improving the accuracy of image recognition.
[0028] As a preferred example of the first aspect, the weighted objective function includes:
[0029] The calculation formula of the weighted objective function is specifically as follows:
[0030]
[0031] where is the total loss, is the K-L divergence loss, λ1 is the weight coefficient of the K-L divergence loss, is the disentanglement loss, λ2 is the weight coefficient of the disentanglement loss, is the reconstruction loss, λ3 is the weight coefficient of the reconstruction loss, is the classification loss, λ4 is the weight coefficient of the classification loss.
[0032] In this preferred example, through the design of a weighted loss function for multi-objective collaborative optimization, the K-L divergence loss is used to constrain the canonical distribution of the latent space, the disentanglement loss strengthens the decoupling ability of the inter-sequence shared features and the intra-sequence specific features, the reconstruction loss ensures the integrity retention of key radiomics information, and the classification loss focuses on optimizing the decision boundary of the clinical diagnosis task. The four form a dynamically balanced optimization mechanism, effectively improving the parsing accuracy of the model for the heterogeneous features of the tumor microenvironment, enhancing the feature representation specificity of the key habitat regions, and finally constructing a model with strong generalization ability and high interpretability.
[0033] As a preferred example of the first aspect, the screening of the first multi-sequence fusion network model set to determine the best cell habitat includes:
[0034] Obtain the AUC value corresponding to each model in the first multi-sequence fusion network model set;
[0035] Take the cell habitat corresponding to the first multi-sequence fusion network model with the largest AUC value as the best cell habitat.
[0036] In this preferred example, through the AUC value screening mechanism, the objective quantitative selection of the best cell habitat is realized, effectively overcoming the subjective deviation of the traditional manual experience in selecting cell habitats, and ensuring that the model focuses on the cell habitats that are most discriminative for ADG grading and IDH mutation prediction.
[0037] As a preferred example of the first aspect, screening each of the second multi-sequence fusion network models and outputting an optimal multi-sequence fusion network model includes:
[0038] Obtaining the AUC value corresponding to each of the second multi-sequence fusion network models;
[0039] Outputting the second multi-sequence fusion network model with the largest AUC value as the optimal multi-sequence fusion network model.
[0040] In this preferred example, through the screening mechanism of the AUC value, after determining the best cell habitat, the optimal multi-sequence fusion network model is further determined, ensuring that the model adaptively integrates the most biologically relevant radiomics features, thereby improving the accuracy of the image recognition result.
[0041] In a second aspect, an embodiment of the present invention further provides an image recognition method based on a multi-sequence fusion network model, including:
[0042] Obtaining test data to be processed;
[0043] According to the test data and the optimal multi-sequence fusion network model, outputting an image recognition result corresponding to the test data, where the optimal multi-sequence fusion network model is obtained according to the multi-sequence fusion network model training method described in any one of the embodiments of the first aspect.
[0044] In this preferred example, image recognition is performed through the optimal multi-sequence fusion network model, which can retain the key features of tumor heterogeneity and deeply couple multi-sequence radiomics features, improving the accuracy of the image recognition result.
[0045] In a third aspect, an embodiment of the present invention further provides a multi-sequence fusion network model training device based on cell habitats, including: a first construction module, a first iterative training module, a cell habitat screening module, a second construction module, a second iterative training module, and an optimal model screening module;
[0046] Among them, the first construction module is used to respectively construct cell habitat input data corresponding to each cell habitat according to the first MRI sequence input data; where the first MRI sequence input data includes a plurality of MRI sequences;
[0047] The first iterative training module is configured to use each of the cell niche input data as input, iteratively train the corresponding first multi-sequence fusion network model respectively, output the first multi-sequence fusion network model corresponding to each of the cell niche input data, and form a first multi-sequence fusion network model set with all the first multi-sequence fusion network models;
[0048] The cell niche screening module is configured to screen the first multi-sequence fusion network model set to determine the optimal cell niche;
[0049] The second construction module is configured to freely combine each MRI sequence in the first cell niche input data according to the first cell niche input data corresponding to the optimal cell niche to obtain a first MRI sequence combination data set;
[0050] The second iterative training module is configured to use each first MRI sequence combination data in the first MRI sequence combination data set as input, iteratively train the corresponding second multi-sequence fusion network model respectively, and output the second multi-sequence fusion network model corresponding to each first MRI sequence combination data;
[0051] The optimal model screening module is configured to screen each of the second multi-sequence fusion network models and output an optimal multi-sequence fusion network model.
[0052] As a preferred example of the third aspect, the first construction module includes: a data extraction unit, a spatial registration unit, and a data determination unit;
[0053] The data extraction unit is configured to perform region extraction on the first T2-FLAIR data according to a preset region of interest drawing rule to obtain second T2-FLAIR data; wherein, the first MRI sequence input data includes the first T2-FLAIR data;
[0054] The spatial registration unit is configured to align the first MRI sequence input data and the preset first DWI sequence data to the T2-FLAIR space of the second T2-FLAIR data through a non-linear registration method to obtain second MRI sequence input data and second DWI sequence data;
[0055] The data determination unit is configured to determine the cell niche input data corresponding to each cell niche according to the second DWI sequence data, the second MRI sequence input data, and the K-means clustering algorithm.
[0056] As a preferred example of the third aspect, the first iterative training module includes: a first feature splitting unit, a first acquisition unit, and a first output unit;
[0057] The first feature splitting unit is used to construct independent encoders for each MRI sequence in the input data of the cyto-microscope, and generate a set of inter-sequence shared representations and a set of intra-sequence specific representations through each of the encoders;
[0058] The first obtaining unit is used to obtain a fused shared representation by calculating the average value of all elements in the set of inter-sequence shared representations;
[0059] The first output unit is used to input the set of intra-sequence specific representations and the fused shared representation into a decoder network and a two-layer fully connected prediction network, and calculate a first total loss through a weighted objective function. When the current iteration meets the first preset condition, the first multi-sequence fusion network model corresponding to the input data of the cyto-microscope is output.
[0060] As a preferred example of the third aspect, the second iterative training module includes: a second feature splitting unit, a second obtaining unit, and a second output unit;
[0061] The second feature splitting unit is used to construct independent encoders for each MRI sequence in the first MRI sequence combination data, and generate a set of inter-sequence shared representations and a set of intra-sequence specific representations through each of the encoders;
[0062] The second feature splitting unit is used to obtain a fused shared representation by calculating the average value of all elements in the set of inter-sequence shared representations;
[0063] The second output unit is used to input the set of intra-sequence specific representations and the fused shared representation into a decoder network and a two-layer fully connected prediction network, and calculate a second total loss through a weighted objective function. When the current iteration meets the second preset condition, the second multi-sequence fusion network model corresponding to the first MRI sequence combination data is output.
[0064] As a preferred example of the third aspect, the multi-sequence fusion network model training device further includes a weighted objective function calculation module;
[0065] The weighted objective function calculation module is used to calculate the total loss through a weighted objective function. The total loss is a weighted superposition of K-L divergence loss, disentanglement loss, reconstruction loss, and classification loss. The specific formula for calculating the total loss is as follows:
[0066]
[0067] Where, is the total loss, is the K-L divergence loss, and λ1 is the weight coefficient of the K-L divergence loss, is the disentanglement loss, λ2 is the weight coefficient of the disentanglement loss, is the reconstruction loss, λ3 is the weight coefficient of the reconstruction loss, is the classification loss, λ4 is the weight coefficient of the classification loss.
[0068] As a preferred example of the third aspect, the cell niche screening module includes: a first AUC value acquisition unit and a cell niche screening unit;
[0069] The first AUC value acquisition unit is used to acquire the AUC value corresponding to each model in the first multi-sequence fusion network model set;
[0070] The cell niche screening unit is used to use the cell niche corresponding to the first multi-sequence fusion network model with the largest AUC value as the optimal cell niche.
[0071] As a preferred example of the third aspect, the optimal model screening module includes: a second AUC value acquisition unit and a model screening unit;
[0072] The second AUC value acquisition unit is used to acquire the AUC value corresponding to each of the second multi-sequence fusion network models;
[0073] The model screening unit is used to output the second multi-sequence fusion network model with the largest AUC value as the optimal multi-sequence fusion network model.
[0074] Fourth aspect, an embodiment of the present invention further provides an image recognition device based on a multi-sequence fusion network model, including: a test data acquisition module and a result generation module;
[0075] The test data acquisition module is used to acquire test data to be processed;
[0076] The result generation module is used to output the image recognition result corresponding to the test data according to the test data and the optimal multi-sequence fusion network model, where the optimal multi-sequence fusion network model is obtained according to the method for training a multi-sequence fusion network model based on cell niche as described in any one of the embodiments of the first aspect.
[0077] Fifth aspect, the present application further provides a computer storage medium, characterized in that a computer program is stored on the computer-readable storage medium, and when the computer program is executed by a processor, it implements the method for training a multi-sequence fusion network model based on cell niche described in the present application. Description of the Drawings
[0078] Figure 1Schematic flowchart of an embodiment of the method for training a multi-sequence fusion network model based on cellular niche provided by the present invention;
[0079] Figure 2 Schematic flowchart of an embodiment of the method for image recognition based on a multi-sequence fusion network model provided by the present invention;
[0080] Figure 3 Schematic structural diagram of an embodiment of the device for training a multi-sequence fusion network model based on cellular niche provided by the present invention;
[0081] Figure 4 Schematic structural diagram of an embodiment of the device for image recognition based on a multi-sequence fusion network model provided by the present invention. Detailed implementation manners
[0082] In the prior art, glioma is the most common malignant tumor of the central nervous system (CNS). Among them, adult diffuse glioma (ADG) is classified into WHO grades 2 to 4 according to its different histological and molecular characteristics. WHO grade 2 represents low-grade glioma (LGG) with a relatively low degree of malignancy and a good response to radiotherapy and chemotherapy; WHO grades 3-4 are classified as high-grade glioma (HGG) with stronger invasiveness and a poor prognosis. In addition, the isocitrate dehydrogenase (IDH) gene status is an important molecular marker for ADG, and IDH mutant (IDH-MT) gliomas show stronger treatment sensitivity and a better prognosis. Currently, the grading of ADG and the assessment of IDH mutations rely on invasive procedures such as biopsy or surgical resection, which is challenging for cases where the tumor is located deep or the patient's physical condition is poor. This emphasizes the urgent need for non-invasive methods for accurate preoperative assessment of ADG grading and IDH status for better diagnosis and treatment decisions. MRI radiomics, which utilizes quantitative imaging features, has been successfully applied to glioma grading and IDH mutation status prediction. However, traditional radiomics studies mainly rely on the voxel relationship metrics of the whole tumor, ignoring the inherent biological heterogeneity within the tumor, which limits its biological interpretability and clinical application. To address this issue, habitat imaging, as an emerging technology, has received extensive attention in the clinic. By segmenting the tumor into sub-regions with different pathological and physiological characteristics, habitat imaging can more accurately characterize the tumor microenvironment.
[0083] Currently, habitat imaging techniques based on advanced MRI sequences, such as diffusion kurtosis imaging (DKI) and dynamic susceptibility contrast-enhanced (DSC) MRI, have shown certain potential in predicting IDH mutation status. However, these methods are only based on simple sequence parameter analysis, such as only calculating the minimum value, maximum value, average value of the parameters within the habitat, and the percentage of habitat volume, etc., which may not comprehensively capture the heterogeneity of tumors; in addition, due to relying on specific imaging protocols, it limits their universality in routine clinical practice. In contrast, diffusion-weighted imaging (DWI) is a widely used physiological MRI sequence that quantifies the diffusion of water molecules through the apparent diffusion coefficient (ADC) and provides valuable information about cell density and tissue structure. The present invention uses the ADC-based cell habitat to better capture the internal heterogeneity of tumors and overcomes the limitations of traditional radiomics in terms of spatial localization and biological interpretability. By combining these cell habitats with the radiomic features extracted from conventional MRI sequences, the accuracy of image recognition in ADG grading and IDH mutation status prediction can be improved.
[0084] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of the present invention.
[0085] Embodiment 1
[0086] Please refer to Figure 1 , for the method for training a multi-sequence fusion network model based on cell habitats provided by the embodiments of the present invention, including S11 to S16, specifically as follows:
[0087] S11: According to the first MRI sequence input data, construct the cell habitat input data corresponding to each cell habitat; wherein, the first MRI sequence input data includes a plurality of MRI sequences;
[0088] Further, in some embodiments of the present application, the constructing the cell habitat input data corresponding to each cell habitat according to the first MRI sequence input data includes steps S111 to S113, and the specific steps are as follows:
[0089] S111: According to the preset rule for drawing the region of interest, perform region extraction on the first T2-FLAIR data to obtain the second T2-FLAIR data; wherein, the first MRI sequence input data includes the first T2-FLAIR data;
[0090] S112: Align the first MRI sequence input data and the preset first DWI sequence data to the T2-FLAIR space of the second T2-FLAIR data through a non-linear registration method to obtain the second MRI sequence input data and the second DWI sequence data;
[0091] S123: Determine the cell niche input data corresponding to each cell niche according to the second DWI sequence data, the second MRI sequence input data, and the K-means clustering algorithm.
[0092] Specifically, the first MRI sequence input data in step S111 and the first DWI sequence data in step S112 can be obtained and processed in the following ways:
[0093] The image data was collected from the UCSF-PDGM public database from 2015 to 2021 and the BJTTH internal dataset from 2009 to 2021 (ethical approval number: KY2022-078-04). 625 ADG patients diagnosed by WHO CNS5 (2021) were included, all of whom had complete preoperative multi-parametric MRI sequences: T1WI, T2WI, T2-FLAIR, CE-T1WI, and DWI. The exclusion criteria included missing IDH mutation data and substandard image quality. Finally, 476 cases (UCSF) were divided into the training set and 149 cases (BJTTH) into the test set. Image acquisition was completed using a clinical scanner. The T1WI, T2WI, T2-FLAIR, and CE-T1WI sequence images constituted the first MRI sequence input data, and the DWI sequence images were the first DWI sequence data.
[0094] For the BJTTH dataset, use the non-linear registration method provided by the Advanced Normalization Tools (ANT) to register other sequences to the T2-FLAIR sequence, and use the linear interpolation method to resample the contrast of all sequences to 1mm 3 , and use the z-score normalization method to normalize the voxels of all sequences to the range of [0, 1]. For the UCSF-PDGM dataset, use a set of segmentation algorithms to automatically segment the images to obtain the segmentation masks. Image preprocessing involves using the non-linear registration method of the Advanced Normalization Tools to register and resample the contrast of each image to the space where the T2-FLAIR image is located (isotropic resolution of 1mm), and at the same time perform z-score normalization on the voxels of all sequences.
[0095] Specifically, the division of cell niches in step S123 can be as follows:
[0096] The K-means clustering algorithm was used to divide the data into different cell habitats. The optimal number of clusters ("k") for K-means was determined by applying the elbow method to the aggregated voxels of all patients in the training set. The specific calculation process of the elbow method is as follows:
[0097] (1) First, a reasonable range of cluster numbers is selected. In the present invention, the range is set from 2 to 10.
[0098] (2) The voxel data of the ADC sequences of all patients in the training set are input into the K-means algorithm together, and the voxel data are divided into k clusters according to the preset number of clusters.
[0099] (3) Calculate the sum of squared errors (SSE) of the k clusters obtained under each cluster number. The calculation formula is as follows:
[0100]
[0101] Among them, C i represents the i-th cluster, μ i represents the center point of the i-th cluster, and x belongs to cluster C i The data points in ‖x-μ i ‖ is the distance from the data point x to the cluster center point μ i The Euclidean distance of .
[0102] (4) With k as the horizontal coordinate and SSE as the vertical coordinate, each k value and its corresponding SSE value are plotted on a two-dimensional plane to obtain a curve, and then the "elbow" point on the curve is found. Usually, before the elbow point, SSE decreases rapidly as k increases; after the elbow point, the downward trend of SSE slows down significantly. In the present invention, the elbow point observed is k=3, so k=3 is the optimal number of clusters sought. Therefore, the data is divided into three cell habitats: cell habitat 1 (H1, low ADC value), cell habitat 2 (H2, medium ADC value) and cell habitat 3 (H3, high ADC value). In order to explore the potential interactions between different ADC value regions, a total of seven cell habitats were obtained, including three separate cell habitats (H1, H2, H3), the combination of each pair (H1; 2, H1; 3, H2; 3) and the combination of three cell habitats (H1; 2; 3), representing the entire tumor.
[0103] Compared with the prior art, the above embodiment has the following beneficial effects: through T2-FLAI R spatial registration, the spatial consistency of MRI data is ensured; secondly, data under different cell habitats are constructed to better capture the internal heterogeneity of the tumor, overcoming the limitations of the prior art in spatial positioning and biological subject interpretability.
[0104] S12: Using each of the cell niche input data as input, iteratively train the corresponding first multi-sequence fusion network model respectively, output the first multi-sequence fusion network models corresponding to each of the cell niche input data, and form a first multi-sequence fusion network model set with all the first multi-sequence fusion network models;
[0105] Further, in some embodiments of the present application, when using any one of the cell niche input data in each of the cell niche input data as input, the step of using each of the cell niche input data as input, iteratively training the corresponding first multi-sequence fusion network model respectively, and outputting the first multi-sequence fusion network models corresponding to each of the cell niche input data includes steps S121 to S123, and the specific steps are as follows:
[0106] S121: Construct independent encoders for each MRI sequence in the cell niche input data, and generate a set of inter-sequence shared representations and a set of intra-sequence specific representations through each of the encoders;
[0107] S122: Calculate the average value of all elements in the set of inter-sequence shared representations to obtain a fused shared representation;
[0108] S123: Input the set of intra-sequence specific representations and the fused shared representation into a decoder network and a two-layer fully connected prediction network, and calculate a first total loss through a weighted objective function. When the current iteration meets a first preset condition, output the first multi-sequence fusion network model corresponding to the cell niche input data.
[0109] Specifically, the method for generating the set of inter-sequence shared representations and the set of intra-sequence specific representations in step S121 can be as follows:
[0110] First, extract radiomics features X i (i = 1,..., L) from each of the L MRI sequences, and input them into a four-layer encoder network to learn the distribution of the latent representation Z i . The structure of the encoder consists of the following parts: an input layer with a dimension of m; followed by 3 fully connected layers in sequence, with dimensions of p, q, and d (the activation function is ReLU for all). At the same time, batch normalization is performed between any two adjacent fully connected layers. These representations are split into 2 parts: the distribution of the inter-sequence shared representation and the distribution of the intra-sequence specific representation . Then, using the reparameterization trick, sample from the above distributions of each sequence respectively to obtain the inter-sequence shared representation (the set of inter-sequence shared representations) and the intra-sequence specific representation . (Set of specific representations within a sequence). Among them, the reparameterization technique is a method for dealing with random latent variables Z in training a generative model. Its purpose is to convert the non-differentiable operation of randomly sampling from a specific distribution into a differentiable operation, so as to ensure that the gradient descent algorithm can work properly. The latent variable Z usually follows a normal distribution N(μ,σ 2 ), and the implementation method of the reparameterization technique is that the operation of sampling a Z from this normal distribution is equivalent to randomly drawing a noise ε from the standard normal distribution N(0,1), and then calculating Z by ε×σ + μ. The specific formula is as follows:
[0111] Z = ε×σ + μ, ε~N(0,1)
[0112] Among them, μ and σ 2 are the parameter mean and variance of the normal distribution automatically learned by the network respectively.
[0113] Specifically, the structure of the decoder in step S123 can be:
[0114] The structure of the decoder is as follows: Connect 3 fully connected layers in sequence, with dimensions d, q, and p respectively (the activation function is ReLU for all). At the same time, batch normalization is performed between any two adjacent fully connected layers. Finally, connect a full input layer with a dimension of m (the activation function is Tanh).
[0115] Compared with the prior art, the above embodiments have the following beneficial effects: By splitting the set of latent representation distributions into the set of inter-sequence shared representation distributions and the set of intra-sequence specific representation distributions, the decoupling and complementary enhancement of MRI sequence information can be achieved: capturing cross-sequence common features through the shared representation distribution, eliminating redundant noise while strengthening the expression of biological essential features; while the specific representation distribution retains the specific pathological information of each sequence, preventing key detail information from being diluted during the multi-modal fusion process, thereby improving the accuracy of image recognition.
[0116] Furthermore, in some embodiments of the present application, the weighted objective function includes:
[0117] The specific calculation formula of the weighted objective function is as follows:
[0118]
[0119] Among them, is the total loss, is the K-L divergence loss, λ1 is the weight coefficient of the K-L divergence loss, is the disentanglement loss, λ2 is the weight coefficient of the disentanglement loss, is the reconstruction loss, λ3 is the weight coefficient of the reconstruction loss, is the classification loss, and λ4 is the weight coefficient of the classification loss.
[0120] Specifically, the weighted objective function can be composed of the losses of the following four parts:
[0121] (1) Kullback-Leibler (K-L) divergence loss. K-L divergence is used to regularize the distribution of the latent representation generated by the encoder to approximate the prior distribution p θ (Z i ), and the calculation formula is as follows:
[0122]
[0123] The prior distribution p θ (Z i ) is usually set to the standard normal distribution N(0, 1) to optimize the latent space. Similarly, the distribution of the latent representation should also be assumed to be a normal distribution Since the latent representation of each sequence has been split into the components shared among sequences and the components specific to each sequence Therefore, the total K-L divergence loss is defined as follows:
[0124]
[0125] (2) Disentanglement loss. The disentanglement loss ensures that the specific representations in each sequence are different from each other, while promoting the similarity of the shared representations among sequences. The disentanglement loss has the following calculation formula:
[0126]
[0127] (3) Reconstruction loss. The reconstruction loss encourages the output of the decoder to be as close as possible to the original input, including the reconstruction of the specific representations and the shared representations in each sequence, as well as the fusion of the shared representations and the reconstruction of the specific representations across sequences. The reconstruction loss and the shared representations , and the reconstruction of the shared representation fusion across sequences and the specific representations . The calculation formula of the reconstruction loss is as follows:
[0128]
[0129] where D i represents the decoder.
[0130] (4) Classification loss. The classification loss is used to evaluate the classification performance. The classification loss has the following calculation formula:
[0131]
[0132] Among them, n represents the total number of samples, and y i is the true label of the i-th sample (positive class is 1, negative class is 0), and p i represents the predicted probability that the i-th sample belongs to the positive class.
[0133] Total loss is the weighted sum of the above four losses, and the calculation formula is as follows:
[0134]
[0135] Among them, λ1, λ2, λ3, and λ4 respectively represent the weight coefficients of each loss.
[0136] Specifically, the parameter settings of the multi-sequence fusion network model can be:
[0137] The parameter L represents the number of MRI sequences and is set to 4. The parameter m represents the dimension of the original radiomics features, and the value is 109. The parameters p, q, and d correspond to the dimensions after dimensionality reduction and are set to 80, 40, and 10 respectively. The parameter d is further divided into a 4-dimensional within-sequence specific representation and a 6-dimensional between-sequence shared representation. The parameter K represents the joint dimension of the latent representation across multiple sequences. When all four sequences are fused, the value of K is 22. Finally, the parameter l represents the dimension of the last output fully connected layer and is set to 1. This network is implemented using TensorFlow 2.1.0 and trained on an NVIDIA GeForce GTX 1080 with 8GB of memory. The training is performed for at most 100 iterations, using the Adam optimizer with a learning rate of 0.01. To prevent overfitting, an early stopping strategy is applied. When the total loss on the validation set has not improved for 10 consecutive iterations, the training process is stopped. The regularization parameters are set empirically, specifically: λ1 = 0.01, λ2 = 0.0001, λ3 = 0.001, λ4 = 0.01.
[0138] Compared with the prior art, the above embodiments have the following beneficial effects: Through the design of a weighted loss function for multi-objective collaborative optimization, using the K-L divergence loss to constrain the canonical distribution of the latent space, the disentanglement loss strengthens the decoupling ability of the shared features between sequences and the specific features within sequences, the reconstruction loss ensures the integrity retention of key radiomics information, and the classification loss focuses on optimizing the decision boundary of the clinical diagnosis task. The four form a dynamically balanced optimization mechanism, effectively improving the analysis accuracy of the model for the heterogeneous characteristics of the tumor microenvironment, enhancing the feature representation specificity of key habitat regions, and finally constructing a model with strong generalization ability and high interpretability.
[0139] S13: Screen the first multi-sequence fusion network model set to determine the optimal cell habitat;
[0140] Further, in some embodiments of the present application, the screening of the first multi-sequence fusion network model set to determine the optimal cell habitat includes:
[0141] Obtain the AUC value corresponding to each model in the first multi-sequence fusion network model set;
[0142] Take the cell habitat corresponding to the first multi-sequence fusion network model with the largest AUC value as the optimal cell habitat.
[0143] Preferably, to obtain the AUC value corresponding to each model in the first multi-sequence fusion network model set, the data example shown in Table 1 below can be referred to:
[0144] Table 1
[0145]
[0146] Based on Table 1 above, for the ADG grading tasks (Grade 2 vs. Grade 3+4 and Grade 2+3 vs. Grade 4), habitat 12 (H1; 2) produced the highest average AUC values, which were 0.8678 and 0.8668 respectively. In the prediction of IDH mutation status, the average AUC of habitat 23 (H2; 3) was the highest, which was 0.8365. Therefore, H1; 2 will be used for subsequent experiments on grading, while H2; 3 will be used for IDH mutation prediction.
[0147] Compared with the prior art, the above embodiments have the following beneficial effects: Through the AUC value screening mechanism, an objective and quantitative selection of the optimal cell habitat is achieved, effectively overcoming the subjective deviation of the traditional manual experience in selecting the cell habitat, and ensuring that the model focuses on the cell habitat that is most discriminative for ADG grading and IDH mutation prediction.
[0148] S14: According to the first cell habitat input data corresponding to the optimal cell habitat, freely combine each MRI sequence in the first cell habitat input data to obtain a set of first MRI sequence combination data;
[0149] S15: Take each first MRI sequence combination data in the set of first MRI sequence combination data as input, and iteratively train the corresponding second multi-sequence fusion network model respectively, and output the second multi-sequence fusion network model corresponding to each first MRI sequence combination data;
[0150] Further, in some embodiments of the present application, when any first MRI sequence combination data in the first MRI sequence combination data set is used as input, training the corresponding second multi-sequence fusion network model by using each first MRI sequence combination data in the first MRI sequence combination data set as input respectively through iteration includes steps S151 to S153, and the specific steps are as follows;
[0151] S151: Construct independent encoders for each MRI sequence in the first MRI sequence combination data, and generate a set of inter-sequence shared representations and a set of intra-sequence specific representations through each of the encoders;
[0152] S152: Calculate the average value of all elements in the set of inter-sequence shared representations to obtain a fused shared representation;
[0153] S153: Input the set of intra-sequence specific representations and the fused shared representation into a decoder network and a two-layer fully connected prediction network, and calculate the second total loss through a weighted objective function. When the current iteration meets the second preset condition, output the second multi-sequence fusion network model corresponding to the first MRI sequence combination data.
[0154] Further, in some embodiments of the present application, the weighted objective function includes:
[0155] The specific calculation formula of the weighted objective function is as follows:
[0156]
[0157] where, is the total loss, is the K-L divergence loss, λ1 is the weight coefficient of the K-L divergence loss, is the disentanglement loss, λ2 is the weight coefficient of the disentanglement loss, is the reconstruction loss, λ3 is the weight coefficient of the reconstruction loss, is the classification loss, λ4 is the weight coefficient of the classification loss.
[0158] S16: Screen each of the second multi-sequence fusion network models, and output an optimal multi-sequence fusion network model.
[0159] Further, in some embodiments of the present application, the screening of each of the second multi-sequence fusion network models to output an optimal multi-sequence fusion network model includes:
[0160] Obtain the AUC value corresponding to each of the second multi-sequence fusion network models;
[0161] Output the second multi-sequence fusion network model with the largest AUC value as the optimal multi-sequence fusion network model.
[0162] Preferably, to obtain the AUC value corresponding to each of the second multi-sequence fusion network models, the data example shown in Table 2 below can be referred to:
[0163] Table 2
[0164]
[0165] Based on Table 2 above, the optimal MRI sequence combinations for all three tasks are T1WI, CE-T1WI, and T2-FLAIR, achieving AUCs of 0.9360, 0.9605, and 0.8721 respectively in the training set.
[0166] Compared with the prior art, the above embodiments have the following beneficial effects: Through the screening mechanism of the AUC value, the optimal multi-sequence fusion network model is further determined after the best cell habitat is determined, ensuring that the model adaptively integrates the most biologically relevant radiomics features, thereby improving the accuracy of the image recognition result.
[0167] In summary, it can be seen that by constructing the cell habitat input data corresponding to each habitat in the embodiments of the present application, the internal heterogeneity of the tumor can be better captured, providing a basis for model training. Then, based on multiple cell habitat input data, the present application trains and screens to determine the best cell habitat, which can accurately lock the core region of tumor heterogeneity expression. Next, the present application optimally selects the optimal MRI sequence combination under the best cell habitat, which can break through the limitation of the information dimension of a single sequence and avoid the noise interference of low-correlation sequences at the same time. Compared with the prior art, the embodiments of the present application adaptively select the most biologically discriminative cell habitat-sequence combination on the basis of ensuring the accurate analysis of the internal heterogeneity of the tumor, combining the spatial distribution law of cell density and the characteristics of different MRI sequences, constructing a model with high robustness and strong generalization ability, and thus improving the accuracy of image recognition.
[0168] Embodiment 2
[0169] Please refer to Figure 2 , an image recognition method based on a multi-sequence fusion network model provided by an embodiment of the present application, including S21 to S22, specifically:
[0170] S21: Obtain the test data to be processed;
[0171] S22: According to the test data and the optimal multi-sequence fusion network model, output the image recognition result corresponding to the test data, where the optimal multi-sequence fusion network model is obtained according to any of the multi-sequence fusion network model training methods based on cell habitat in the embodiments of the present invention.
[0172] Preferably, according to the test data and the optimal multi-sequence fusion network model, the image recognition result corresponding to the test data is output. The data example shown in Table 3 below can be referred to:
[0173] Table 3
[0174]
[0175] Based on Table 3 above, the recognition ability of the optimal multi-sequence fusion network model on the test data is summarized. For Grade 2 vs. Grade 3+4, the AUC of the optimal multi-sequence fusion network model is 0.7985; for Grade 2+3 vs. Grade 4, its AUC is 0.8240; for IDH mutation status prediction, the AUC of the optimal multi-sequence fusion network model is 0.8190. The performance of the optimal multi-sequence fusion network model is always better than that of the single-sequence model, and there is a significant statistical improvement in terms of AUC (p<0.05 for almost all tasks except F2(CE-T1WI) in IDH prediction).
[0176] Embodiment III
[0177] Please refer to Figure 3 , which is a schematic structural diagram of an embodiment of the multi-sequence fusion network model training device provided by the present invention. As Figure 3 shown, the device includes a first construction module 31, a first iterative training module 32, a cell niche screening module 33, a second construction module 34, a second iterative training module 35, and an optimal model screening module 36.
[0178] Further, in some embodiments of the present application, a first construction module 31 is configured to construct cell habitat input data corresponding to each cell habitat according to the first MRI sequence input data; wherein the first MRI sequence input data includes a plurality of MRI sequences; a first iterative training module 32 is configured to use each of the cell habitat input data as an input, iteratively train the corresponding first multi-sequence fusion network model respectively, output the first multi-sequence fusion network model corresponding to each of the cell habitat input data, and form a first multi-sequence fusion network model set with all the first multi-sequence fusion network models; a cell habitat screening module 33 is configured to screen the first multi-sequence fusion network model set to determine the optimal cell habitat; a second construction module 34 is configured to freely combine each MRI sequence in the first cell habitat input data according to the first cell habitat input data corresponding to the optimal cell habitat to obtain a first MRI sequence combination data set; a second iterative training module 35 is configured to use each first MRI sequence combination data in the first MRI sequence combination data set as an input, iteratively train the corresponding second multi-sequence fusion network model respectively, and output the second multi-sequence fusion network model corresponding to each first MRI sequence combination data; an optimal model screening module 36 is configured to screen each of the second multi-sequence fusion network models and output an optimal multi-sequence fusion network model.
[0179] Further, in some embodiments of the present application, the first construction module 31 includes a data extraction unit, a spatial registration unit, and a data determination unit; the data extraction unit is configured to perform region extraction on the first T2-FLAIR data according to a preset rule for drawing an interested region to obtain second T2-FLAIR data; wherein the first MRI sequence input data includes the first T2-FLAIR data; the spatial registration unit is configured to align the first MRI sequence input data and the preset first DWI sequence data to the T2-FLAIR space of the second T2-FLAIR data through a non-linear registration method to obtain second MRI sequence input data and second DWI sequence data; the data determination unit is configured to determine the cell habitat input data corresponding to each cell habitat according to the second DWI sequence data, the second MRI sequence input data, and the K-means clustering algorithm.
[0180] Further, in some embodiments of the present application, the first iterative training module 32 includes a first feature splitting unit, a first obtaining unit, and a first output unit; the first feature splitting unit is configured to construct independent encoders for each MRI sequence in the input data of the cell microscopy, and generate a set of inter-sequence shared representations and a set of intra-sequence specific representations through each of the encoders; the first obtaining unit is configured to obtain a fused shared representation by calculating the average value of all elements in the set of inter-sequence shared representations; the first output unit is configured to input the set of intra-sequence specific representations and the fused shared representation into a decoder network and a two-layer fully connected prediction network, and calculate a first total loss through a weighted objective function, and output a first multi-sequence fusion network model corresponding to the input data of the cell microscopy until the current iteration meets a first preset condition.
[0181] Further, in some embodiments of the present application, the second iterative training module 35 includes: a second feature splitting unit, a second obtaining unit, and a second output unit; the second feature splitting unit is configured to construct independent encoders for each MRI sequence in the first MRI sequence combined data, and generate a set of inter-sequence shared representations and a set of intra-sequence specific representations through each of the encoders; the second feature splitting unit is configured to obtain a fused shared representation by calculating the average value of all elements in the set of inter-sequence shared representations; the second output unit is configured to input the set of intra-sequence specific representations and the fused shared representation into a decoder network and a two-layer fully connected prediction network, and calculate a second total loss through a weighted objective function, and output a second multi-sequence fusion network model corresponding to the first MRI sequence combined data until the current iteration meets a second preset condition.
[0182] Further, in some embodiments of the present application, the multi-sequence fusion network model training device further includes a weighted objective function calculation module; the weighted objective function calculation module is configured to calculate a total loss through a weighted objective function, and the specific formula of the weighted objective function is as follows:
[0183]
[0184] Wherein, is the total loss, is the K-L divergence loss, λ1 is the weight coefficient of the K-L divergence loss, is the disentanglement loss, λ2 is the weight coefficient of the disentanglement loss, is the reconstruction loss, λ3 is the weight coefficient of the reconstruction loss, is the classification loss, λ4 is the weight coefficient of the classification loss.
[0185] Further, in some embodiments of the present application, the cell niche screening module 33 includes: a first AUC value acquisition unit and a cell niche screening unit; the first AUC value acquisition unit is configured to acquire the AUC value corresponding to each model in the first multi-sequence fusion network model set; the cell niche screening unit is configured to use the cell niche corresponding to the first multi-sequence fusion network model with the largest AUC value as the optimal cell niche.
[0186] Further, in some embodiments of the present application, the optimal model screening module 36 includes: a second AUC value acquisition unit and a model screening unit; the second AUC value acquisition unit is configured to acquire the AUC value corresponding to each of the second multi-sequence fusion network models; the model screening unit is configured to output the second multi-sequence fusion network model with the largest AUC value as the optimal multi-sequence fusion network model.
[0187] For a more detailed step flow and working principle of this embodiment, reference may be made to, but not limited to, the relevant records in Embodiment 1.
[0188] Embodiment 4
[0189] Reference Figure 4 , which is a schematic structural diagram of an embodiment of an image recognition device based on a multi-sequence fusion network model provided by the present invention. As Figure 4 shown, the image recognition device based on the multi-sequence fusion network model includes a test data acquisition module 41 and a result generation module 42.
[0190] Further, in some embodiments of the present application, the test data acquisition module 41 is configured to acquire test data to be processed; the result generation module 42 is configured to output an image recognition result corresponding to the test data according to the test data and the optimal multi-sequence fusion network model, where the optimal multi-sequence fusion network model is obtained according to the cell niche-based multi-sequence fusion network model training device described in any one of the above embodiments.
[0191] Embodiment 5
[0192] Based on the embodiments of the above multi-sequence fusion network model training method, another embodiment of the present application provides a storage medium, which includes a stored computer program. When the computer program runs, it controls the device where the storage medium is located to execute the cell niche-based multi-sequence fusion network model training method of any embodiment of the present application.
[0193] In this embodiment, the above storage medium is a computer-readable storage medium, and the computer program includes computer program code, which may be in the form of source code, object code, executable file or some intermediate form, etc. The computer-readable medium may include: any entity or device capable of carrying the computer program code, recording medium, USB flash drive, mobile hard disk, magnetic disk, optical disk, computer memory, read-only memory (ROM), random access memory (RAM), electrical carrier signal, telecommunication signal, and software distribution medium, etc. It should be noted that the content included in the computer-readable medium may be appropriately increased or decreased according to the requirements of legislation and patent practice within the jurisdiction. For example, in some jurisdictions, according to legislation and patent practice, the computer-readable medium does not include electrical carrier signals and telecommunication signals.
[0194] The specific embodiments described above further elaborate on the purpose, technical solution, and beneficial effects of the present application. It should be understood that the above are only specific embodiments of the present application and are not used to limit the protection scope of the present application. In particular, for those skilled in the art, any modifications, equivalent replacements, improvements, etc. made within the spirit and principle of the present application shall be included in the protection scope of the present application.
Claims
1. A method for training a multi-sequence fusion network model based on cell habitats, characterized in that Including: According to the first MRI sequence input data, cell habitat input data corresponding to each cell habitat is respectively constructed; wherein, the first MRI sequence input data includes a plurality of MRI sequences; Taking each of the cell habitat input data as input, respectively iteratively training corresponding first multi-sequence fusion network models, outputting the first multi-sequence fusion network models corresponding to each of the cell habitat input data, and forming a first multi-sequence fusion network model set with all the first multi-sequence fusion network models; Screening the first multi-sequence fusion network model set to determine the optimal cell habitat; According to the first cell habitat input data corresponding to the optimal cell habitat, freely combining each MRI sequence in the first cell habitat input data to obtain a first MRI sequence combination data set; Taking each first MRI sequence combination data in the first MRI sequence combination data set as input, respectively iteratively training corresponding second multi-sequence fusion network models, and outputting the second multi-sequence fusion network models corresponding to each first MRI sequence combination data; Screening each of the second multi-sequence fusion network models to output an optimal multi-sequence fusion network model.
2. The method for training a multi-sequence fusion network model based on a cellular niche according to claim 1, wherein The constructing of the cell habitat input data corresponding to each cell habitat according to the first MRI sequence input data includes: Performing region extraction on the first T2-FLAIR data according to a preset region of interest drawing rule to obtain second T2-FLAIR data; wherein, the first MRI sequence input data includes the first T2-FLAIR data; Aligning the first MRI sequence input data and the preset first DWI sequence data to the T2-FLAIR space of the second T2-FLAIR data through a non-linear registration method to obtain second MRI sequence input data and second DWI sequence data; Determining the cell habitat input data corresponding to each cell habitat according to the second DWI sequence data, the second MRI sequence input data and the K-means clustering algorithm.
3. The method for training a multi-sequence fusion network model based on a cell niche according to claim 1, wherein When taking any one of the cell habitat input data in each of the cell habitat input data as input, the taking each of the cell habitat input data as input, respectively iteratively training corresponding first multi-sequence fusion network models, and outputting the first multi-sequence fusion network models corresponding to each of the cell habitat input data includes: Respectively constructing independent encoders for each MRI sequence in the cell habitat input data, and generating a set of shared representations between sequences and a set of unique representations within sequences through each of the encoders; Calculating the average value of all elements in the set of shared representations between sequences to obtain a fused shared representation; Inputting the set of unique representations within sequences and the fused shared representation into a decoder network and a two-layer fully connected prediction network, and calculating a first total loss through a weighted objective function until the current iteration meets a first preset condition, and outputting the first multi-sequence fusion network model corresponding to the cell habitat input data.
4. The method for training a multi-sequence fusion network model based on a cell niche according to claim 1, wherein When any first MRI sequence combination data in the first MRI sequence combination data set is used as input, iteratively training the corresponding second multi-sequence fusion network model by using each first MRI sequence combination data in the first MRI sequence combination data set as input respectively, including: Constructing independent encoders for each MRI sequence in the first MRI sequence combination data, and generating a set of inter-sequence shared representations and a set of intra-sequence specific representations through each encoder; Obtaining a fused shared representation by calculating the average value of all elements in the set of inter-sequence shared representations; Inputting the set of intra-sequence specific representations and the fused shared representation into a decoder network and a two-layer fully connected prediction network, and calculating a second total loss through a weighted objective function, and outputting the second multi-sequence fusion network model corresponding to the first MRI sequence combination data until the current iteration meets a second preset condition; 5. The method for training a multi-sequence fusion network model based on a cell niche according to claim 3 or 4, characterized in that, The weighted objective function includes: The specific calculation formula of the weighted objective function is as follows: wherein, is the total loss, is the K-L divergence loss, and λ1 is the weight coefficient of the K-L divergence loss, is the disentanglement loss, and λ2 is the weight coefficient of the disentanglement loss, is the reconstruction loss, and λ3 is the weight coefficient of the reconstruction loss, is the classification loss, and λ4 is the weight coefficient of the classification loss.
6. The method for training a multi-sequence fusion network model based on a cellular niche according to claim 1, wherein Screening the first multi-sequence fusion network model set to determine the optimal cell niche, including: Obtaining the AUC value corresponding to each model in the first multi-sequence fusion network model set; Taking the cell niche corresponding to the first multi-sequence fusion network model with the largest AUC value as the optimal cell niche; 7. The method for training a multi-sequence fusion network model based on a cell niche according to claim 1, wherein, Screening each of the second multi-sequence fusion network models and outputting an optimal multi-sequence fusion network model, including: Obtaining the AUC value corresponding to each of the second multi-sequence fusion network models; Outputting the second multi-sequence fusion network model with the largest AUC value as the optimal multi-sequence fusion network model; 8. An image recognition method based on a multi-sequence fusion network model, characterized in that, including: Obtaining test data to be processed; Outputting an image recognition result corresponding to the test data according to the test data and the optimal multi-sequence fusion network model, where the optimal multi-sequence fusion network model is obtained according to the method for training a multi-sequence fusion network model based on cell niche as claimed in claims 1 to 7; 9. A training device for a multi-sequence fusion network model based on cell habitats, characterized in that including: A first construction module, a first iterative training module, a cell niche screening module, a second construction module, a second iterative training module, and an optimal model screening module; The first construction module is configured to respectively construct cell niche input data corresponding to each cell niche according to first MRI sequence input data; wherein the first MRI sequence input data includes a plurality of MRI sequences; The first iterative training module is configured to use each of the cell niche input data as input, iteratively train the corresponding first multi-sequence fusion network model respectively, output the first multi-sequence fusion network model corresponding to each of the cell niche input data, and form a first multi-sequence fusion network model set with all the first multi-sequence fusion network models; The cell niche screening module is configured to screen the first multi-sequence fusion network model set to determine the optimal cell niche; The second construction module is configured to freely combine each MRI sequence in the first cell niche input data according to the first cell niche input data corresponding to the optimal cell niche to obtain a first MRI sequence combination data set; The second iterative training module is used to take each first MRI sequence combination data in the first MRI sequence combination data set as input, iteratively train the corresponding second multi-sequence fusion network model respectively, and output the second multi-sequence fusion network model corresponding to each first MRI sequence combination data. The optimal model screening module is used to screen each of the second multi-sequence fusion network models and output the optimal multi-sequence fusion network model.
10. The apparatus for training a multi-sequence fusion network model based on a cellular niche according to claim 9, wherein The first construction module includes: a data extraction unit, a spatial registration unit, and a data determination unit. The data extraction unit is used to perform region extraction on the first T2-FLAIR data according to the preset region of interest drawing rule to obtain the second T2-FLAIR data; wherein, the first MRI sequence input data includes the first T2-FLAIR data. The spatial registration unit is used to align the first MRI sequence input data and the preset first DWI sequence data to the T2-FLAIR space of the second T2-FLAIR data through a non-linear registration method to obtain the second MRI sequence input data and the second DWI sequence data. The data determination unit is used to determine the cell habitat input data corresponding to each cell habitat according to the second DWI sequence data, the second MRI sequence input data, and the K-means clustering algorithm.
11. The training device for the multi-sequence fusion network model based on cell niche according to claim 9, characterized in that, The first iterative training module includes: a first feature splitting unit, a first acquisition unit, and a first output unit. The first feature splitting unit is used to construct independent encoders for each MRI sequence in the cell habitat input data, and generate a set of inter-sequence shared representations and a set of intra-sequence specific representations through each encoder. The first acquisition unit is used to calculate the average value of all elements in the set of inter-sequence shared representations to obtain the fused shared representation. The first output unit is used to input the set of intra-sequence specific representations and the fused shared representation into a decoder network and a two-layer fully connected prediction network, and calculate the first total loss through a weighted objective function. Until the current iteration meets the first preset condition, output the first multi-sequence fusion network model corresponding to the cell habitat input data.
12. The training device for the multi-sequence fusion network model based on cell niche according to claim 9, wherein The second iterative training module includes: a second feature splitting unit, a second acquisition unit, and a second output unit. The second feature splitting unit is used to construct independent encoders for each MRI sequence in the first MRI sequence combination data, and generate a set of inter-sequence shared representations and a set of intra-sequence specific representations through each encoder. The second feature splitting unit is used to calculate the average value of all elements in the set of inter-sequence shared representations to obtain the fused shared representation. The second output unit is used to input the set of intra-sequence specific representations and the fused shared representation into a decoder network and a two-layer fully connected prediction network, and calculate the second total loss through a weighted objective function. Until the current iteration meets the second preset condition, output the second multi-sequence fusion network model corresponding to the first MRI sequence combination data.
13. The training device for the multi-sequence fusion network model based on cell habitats according to claim 9, characterized in that, The multi-sequence fusion network model training device further includes a weighted objective function calculation module; The weighted objective function calculation module is used to calculate the total loss through a weighted objective function, and the specific calculation formula of the weighted objective function is as follows: wherein, is the total loss, is the K-L divergence loss, and λ1 is the weight coefficient of the K-L divergence loss, is the disentanglement loss, and λ2 is the weight coefficient of the disentanglement loss, is the reconstruction loss, and λ3 is the weight coefficient of the reconstruction loss, is the classification loss, and λ4 is the weight coefficient of the classification loss.
14. The training device for the multi-sequence fusion network model based on cell habitat according to claim 9, wherein The cell niche screening module includes: a first AUC value acquisition unit and a cell niche screening unit; The first AUC value acquisition unit is used to acquire the AUC value corresponding to each model in the first multi-sequence fusion network model set; The cell niche screening unit is used to use the cell niche corresponding to the first multi-sequence fusion network model with the largest AUC value as the optimal cell niche.
15. The apparatus for training a multi-sequence fusion network model based on cell habitats according to claim 9, wherein The optimal model screening module includes: a second AUC value acquisition unit and a model screening unit; The second AUC value acquisition unit is used to acquire the AUC value corresponding to each of the second multi-sequence fusion network models; The model screening unit is used to output the second multi-sequence fusion network model with the largest AUC value as the optimal multi-sequence fusion network model.
16. An image recognition device based on a multi-sequence fusion network model, characterized in that, It includes: a test data acquisition module and a result generation module; The test data acquisition module is used to acquire the test data to be processed; The result generation module is used to output the image recognition result corresponding to the test data according to the test data and the optimal multi-sequence fusion network model, where the optimal multi-sequence fusion network model is obtained according to the method for training a multi-sequence fusion network model based on cell niche as claimed in claims 1 to 7.
17. A computer storage medium, characterized in that, A computer program is stored on the computer storage medium, and when the computer program is executed by a processor, it implements the method for training a multi-sequence fusion network model based on cell niche as claimed in any one of claims 1 to 8.