SegTCLIP glioma segmentation method and system based on data amplification and semantic map and application
By adopting the SegTCLIP method based on data amplification and semantic map in glioma image segmentation, the problem of glioma image deletion and insufficient fusion of image and text information is solved, and high-accuracy segmentation of complex glioma areas is achieved, and the diagnosis and segmentation effect of glioma is improved.
Patent Information
- Application Number
- CN202510276149.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-10
- Publication Date
- 2025-06-06
AI Technical Summary
Glioma image segmentation faces many challenges, including diversity in tumor morphology and size, blurred boundaries, low contrast and noise interference, as well as difficulties in data annotation and insufficient data volume, especially when image modality is missing or image quality is low, the performance of the model may significantly decrease.
The glioma images were preprocessed by spatial alignment normalization using SegTCLIP glioma segmentation method based on data amplification and semantic maps, and missing modal filling was performed using improved diffusion models. Then, the fill-processed image and text data are extracted by the SegTCLIP large-model encoder, and the graph-text mixed feature map is constructed based on the knowledge graph mixing and improved cross attention mechanism. Finally, the pixel-text score map is calculated based on the multimodal pixel feature map, and input it to the multimodal decoding network for feature decoding to obtain the segmentation result.
This method effectively solves the problems of glioma image deletion and insufficient fusion of image and text information, improves the accuracy of segmentation of complex glioma regions, ensures the integrity of multimodal data, and thus improves the accuracy of glioma diagnosis and segmentation.
Smart Images

Figure CN120107596A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of medical image processing and image segmentation, and in particular to a SegTCLIP glioma segmentation method, system and application based on data amplification and semantic graph. Background Art
[0002] Glioma is a highly invasive tumor originating from the central nervous system. Its segmentation is a technology that accurately annotates tumor areas by analyzing medical imaging data. It is widely used to assist in diagnosis and treatment planning. The core of glioma segmentation is to accurately extract tumor morphology and location information from magnetic resonance imaging (MRI) images, providing key support for tumor grading assessment, surgical planning, and efficacy monitoring. However, glioma image segmentation faces many challenges, including the diversity of tumor morphology and size, blurred boundaries, low contrast and noise interference, as well as difficulties in data annotation and insufficient data volume. In addition, some scanned images may lack T1C, T2 or enhanced modality image data, resulting in incomplete information. The heterogeneity of tumors and the anatomical differences of individual patients also make the performance of segmentation models unstable in different patient images, further increasing the difficulty of glioma segmentation.
[0003] Many researchers have tried to use different methods to improve the accuracy of glioma segmentation. For example, convolutional neural networks (CNNs) are widely used in the automatic segmentation of gliomas. Through deep learning methods, CNNs can automatically extract features of tumor regions from multimodal MRI images. However, due to the complex morphology and blurred boundaries of tumors, CNNs still face certain challenges in dealing with these problems. In order to improve the segmentation performance, some researchers have fused multimodal MRI images and used the complementary information between different modalities to enhance the recognition ability of the model. This method can improve the robustness of segmentation, but due to the feature differences and noise problems between different modalities, it may still lead to unstable segmentation accuracy, especially when a certain modality is missing or the image quality is low, the performance of the model may drop significantly. To solve this problem, in recent years, some researchers have proposed introducing generative adversarial networks (GANs) into glioma segmentation. By using GANs for data augmentation, diverse image samples can be effectively generated to alleviate the problems of insufficient data and missing modalities. However, the accuracy and clinical applicability of GAN-generated images in medical imaging remain a challenge. In addition, some researchers proposed to apply Transformers and Self-Attention Mechanism to image segmentation tasks, using the global feature learning ability of Transformer to help the model better capture the complex information of the tumor area, while the self-attention mechanism helps to improve the model's sensitivity to local details, thereby improving segmentation accuracy. However, it still cannot completely solve the problems of complex tumor morphology, blurred boundaries, and inconsistency between modalities. Some researchers proposed a framework based on the combination of graph convolutional networks (GCN) and transformer networks, which improves the model's sensitivity to local details of the tumor area by embedding spatial information in the graph structure. By using GCN, the model can effectively capture the complex structural information in the image, while the transformer network helps to strengthen the learning of global features, thereby improving segmentation accuracy. Although this method improves the segmentation accuracy of the tumor area to a certain extent, it still has problems such as high computational complexity, long training time, and possible information loss when processing high-dimensional data.
[0004] Although glioma image segmentation has broad application prospects in many fields, it still faces several challenges. For example, due to equipment limitations or different scanning conditions, some image modalities may be missing, affecting the integrity of the image and the comprehensiveness of the information, thereby affecting the accuracy of tumor segmentation. Medical image segmentation tasks often require a large amount of labeled data. However, due to the difficulty of medical image labeling and resource limitations, the labeled data is often very limited, which makes it possible for existing deep learning models to fail to learn sufficient tumor feature information during training, especially in new patients or rare cases. The model performs poorly. In addition, glioma segmentation requires not only pixel-level image information, but also the semantic information of the tumor.
[0005] Currently, no effective solution has been proposed for the problems in the related technologies. Summary of the invention
[0006] In view of this, the present invention provides a SegTCLIP glioma segmentation method, system and application based on data amplification and semantic graph to solve the above-mentioned problems.
[0007] In order to solve the above problems, the specific technical solutions adopted by the present invention are as follows:
[0008] According to a first aspect of the present invention, a SegTCLIP glioma segmentation method based on data amplification and semantic graph is provided, comprising:
[0009] S1. Preprocess the pre-collected glioma images by spatial alignment normalization, and perform missing modality filling on the preprocessed glioma images by using an improved diffusion model;
[0010] S2. The SegTCLIP large model encoder is used to extract features from the padded glioma images and the corresponding text data and map them into embedding vectors. The cross-attention mechanism with improved hybrid knowledge graph is combined to construct a mixed feature graph of images and texts.
[0011] S3. Calculate the pixel-text score map based on the multimodal pixel feature map, and input the image-text mixed feature map and the pixel-text score map into the multimodal decoding network for feature decoding to obtain the segmentation result.
[0012] Preferably, the preprocessing of the pre-collected glioma images by spatial alignment normalization and the missing modality filling processing of the preprocessed glioma images by an improved diffusion model include:
[0013] S11, collecting glioma images, performing spatial deformation elimination processing on the glioma images by using a spatial alignment method, and cropping non-glioma areas in the glioma images;
[0014] S12, performing scale-unified processing on the grayscale values of the cropped glioma image by using a normalization technique to obtain a normalized glioma image;
[0015] S13. Determine whether the normalized glioma image is missing a modality through multimodal image registration, and perform modality filling processing on the glioma image with missing modality through the improved diffusion model.
[0016] Preferably, judging whether the normalized glioma image is missing a modality by multimodal image registration, and performing modality filling processing on the glioma image missing a modality by the improved diffusion model comprises:
[0017] S131, collecting a benchmark reference template image, and registering the normalized glioma image using the benchmark reference template image;
[0018] S132, dynamically determining the size of the cropping frame according to the distribution range of the tumor and cropping the registered glioma image to obtain an image containing a key area of the tumor;
[0019] S133. Determine whether the patient contains all expected modalities based on the image containing the key area of the tumor. If not, predict the missing modal map through the improved diffusion model, and fill the missing modal map into the key area image containing the tumor that does not contain all expected modalities.
[0020] Preferably, the predicting of the missing modal diagram by using the improved diffusion model comprises the following steps:
[0021] S1331. For an image containing a key tumor area, determine the existing modality of the image and use a convolutional neural network of a diffusion model to extract features, and obtain a fused feature map by fusing the extracted features;
[0022] S1332, gradually adding standard normally distributed noise to the fused feature map to obtain a pure noise map;
[0023] S1333, for an image containing a key tumor area, calculating the image gradient of the image, and calculating the standard deviation of the image gradient using a standard deviation calculation formula, comparing the standard deviation of the image gradient with a preset threshold, and determining the number of denoising steps according to the comparison result;
[0024] The standard deviation calculation formula is:
[0025]
[0026] In the formula, σ represents the standard deviation of the image gradient, M represents the number of rows in the image, N represents the number of columns in the image, G(i,j) represents the comprehensive gradient amplitude of the current pixel (i,j), and G avgRepresents the average gradient magnitude of the image;
[0027] S1334, generating a spatial attention map according to the standard deviation of the image gradient and the fused feature map, and generating a dynamic spatial attention score map in combination with the dynamic adjustment factor;
[0028] S1335. Based on the dynamic spatial attention score map and in combination with the modality consistency constraint loss function, perform preliminary denoising on the pure noise image to obtain an image feature map of the missing modality predicted after the noise is removed, and decode the predicted missing modality image feature map through a decoder to obtain a predicted missing modality image;
[0029] The expression of the modal consistency constraint loss function is:
[0030]
[0031] In the formula, represents the modal consistency constraint loss function, X n represents the dynamic spatial attention score map, Φ fusion Represents the fused feature map.
[0032] Preferably, the method of extracting features from the padded glioma image and the corresponding text data through the SegTCLIP large model encoder and mapping them into embedded vectors, and constructing a mixed image-text feature map by combining the knowledge graph hybrid improved cross attention mechanism includes:
[0033] S21. Configure the network structure of the SegTCLIP large model, optimize the network structure of the SegTCLIP large model through the weight of the pre-trained model CLIP, and obtain the SegTCLIP large model;
[0034] S22, obtaining a text description of the padded glioma image, and inputting the text description and the padded glioma image into the SegTCLIP large model;
[0035] S23, extracting the image spatial information features and the text description semantic features respectively through the image encoding module and the text encoding module in the SegTCLIP large model, and mapping them into embedding vectors to obtain the image embedding vector and the text embedding vector;
[0036] S24. Use knowledge graph technology to fuse the image embedding vector and the text embedding vector, and combine the fused text embedding vector and image embedding vector through an improved cross-attention mechanism to obtain a mixed image-text feature map.
[0037] Preferably, configuring the network structure of the SegTCLIP large model, optimizing the network structure of the SegTCLIP large model by the weight of the pre-trained model CLIP, and obtaining the SegTCLIP large model comprises:
[0038] S211. Establish the initial SegTCLIP large model, determine the number of image encoding modules, text encoding modules, and fusion modules and their connection order, and define the overall network structure;
[0039] S212. Optimizing core parameters of the image encoding module, the text encoding module, and the fusion module according to the overall network structure;
[0040] S213. Use the weights of the pre-trained model CLIP to initialize the image and text encoding modules to obtain the final SegTCLIP large model.
[0041] Preferably, the method of fusing the image embedding vector and the text embedding vector by using the knowledge graph technology, and combining the fused text embedding vector and the image embedding vector by using an improved cross attention mechanism to obtain a mixed image-text feature map includes:
[0042] S241, mapping the image embedding vector and the text embedding vector into query, key and value;
[0043] S242, calculating the attention weight by calculating the dot product between the image embedding query and the text embedding key, and constructing an attention weight matrix;
[0044] S243, using the attention weight matrix to perform weighted summation on the value vector of the text embedding vector to obtain an initial image-text mixed feature map;
[0045] The expression for weighted summation of the value vector of the text embedding vector using the attention weight matrix is:
[0046] e fusion =α i ·A·V text ;
[0047] In the formula, e fusion represents the initial image-text mixed feature map, α i represents the weight of each layer, A represents the attention weight, V text Represents a value;
[0048] S244, using a level dynamic weight control mechanism, adjusting the initial image-text mixed feature map to obtain a final image-text mixed feature map;
[0049] The expression for adjusting the initial image-text mixed feature map using the level-level dynamic weight control mechanism is:
[0050]
[0051] In the formula, e final represents the final image-text mixed feature map, β i represents the dynamically learned features, e fusion Represents the initial image-text mixed feature map, and i represents the index value.
[0052] Preferably, the pixel-text score map is calculated based on the multimodal pixel feature map, and the image-text mixed feature map and the pixel-text score map are input into the multimodal decoding network for feature decoding to obtain the segmentation result, which includes:
[0053] S31, dividing the glioma image after the missing modality filling process into a number of pixel blocks, extracting pixel features from each pixel block through a convolutional neural network, and fusing the pixel features with those at the same position of other modalities to obtain a multi-modal pixel feature map;
[0054] S32, mapping the pixel features of each pixel block in the multimodal pixel feature map into a pixel embedding vector;
[0055] S33, calculating a pixel-text similarity score according to the pixel embedding vector and the text embedding vector to obtain a pixel-text score graph;
[0056] The calculation formula of the pixel-text similarity score is:
[0057]
[0058] In the formula, S pixel-text represents the pixel-text similarity score, I p represents the pixel embedding vector of the pixel block at the pth position, e text Represents the text embedding vector;
[0059] S34, inputting the pixel-text score map and the image-text mixed feature map into the adaptive multimodal fusion network, using an adaptive strategy, and dynamically adjusting the weights of the two features according to the learning, and generating a fused multimodal feature map through decoding;
[0060] S35. The final glioma segmentation map is obtained through multiple decoding.
[0061] Preferably, the step of inputting the pixel-text score map and the image-text mixed feature map into an adaptive multimodal fusion network, using an adaptive strategy, and dynamically adjusting the weights of the features of the two according to learning, and generating a fused multimodal feature map by decoding includes:
[0062] S341, using the image-text mixed feature map as the input of the adaptive multimodal fusion network, and decoding it through the decoder of the adaptive multimodal fusion network to obtain a decoded feature map;
[0063] S342, using an adaptive strategy and dynamically adjusting a hyperparameter of the influence of the pixel-text score map on decoding according to the learning;
[0064] S343, according to the hyperparameter of the influence degree of the pixel-text score map on decoding, weightedly fuse the decoded decoding feature map and the pixel-text score map to obtain a fused multimodal feature map;
[0065] The expression for weighted fusion of the decoded feature map and the pixel-text score map is:
[0066]
[0067] In the formula, represents the fused multimodal feature map, represents the decoding feature map obtained by the current decoder, E represents the hyperparameter of the influence of the pixel-text score map on decoding, Represents the gradient of the current decoded feature map, S pixel-text (i,j) represents the value of the pixel-text score map.
[0068] According to a second aspect of the present invention, there is provided an application of a SegTCLIP glioma segmentation method based on data augmentation and semantic graph in updating semantic associations and inference rules of a knowledge graph, including:
[0069] Perform semantic segmentation on the glioma segmentation map through the segmentation model to obtain a semantic segmentation map, and extract semantic information of key areas in the semantic segmentation map;
[0070] Using predefined semantic mapping rules, the extracted semantic information is mapped with the entities and relationship nodes in the knowledge graph;
[0071] Update semantic associations in the knowledge graph and optimize inference rules.
[0072] According to a third aspect of the present invention, a SegTCLIP glioma segmentation system based on data amplification and semantic graph is provided, comprising:
[0073] An image processing module is used to pre-process the pre-collected glioma images by spatial alignment normalization, and to perform missing modality filling processing on the pre-processed glioma images by using an improved diffusion model;
[0074] The feature map construction module is used to extract features from the padded glioma images and the corresponding text data through the SegTCLIP large model encoder and map them into embedded vectors, and to construct a mixed image-text feature map by combining the cross-attention mechanism with the knowledge graph hybrid improvement;
[0075] The feature decoding module is used to calculate the pixel-text score map based on the multimodal pixel feature map, and input the image-text mixed feature map and the pixel-text score map into the multimodal decoding network for feature decoding to obtain the segmentation result.
[0076] The beneficial effects of the present invention are:
[0077] 1. The present invention realizes a large glioma segmentation model based on SegTCLIP, which can be used for glioma image segmentation, solves the problems of missing glioma image modality and insufficient fusion of glioma image and text information, and improves the accuracy of segmentation of complex glioma regions.
[0078] 2. The present invention uses a diffusion model to complete the missing modalities of gliomas, solves the problem of missing image data modalities, ensures the integrity of multimodal data, and thus improves the accuracy of glioma diagnosis and segmentation.
[0079] 3. The present invention realizes the fusion of image and text embeddings using spatial-semantic graph and semantic alignment graph respectively, which solves the problem that traditional visual methods are difficult to capture fuzzy or complex features, and at the same time improves the model's ability to understand features.
[0080] 4. The present invention realizes the combination of the fused image and text embedding through an improved cross-attention mechanism to obtain a mixed image-text feature map, which solves the problems of insufficient feature fusion, insufficient information complementarity and weak fine-grained feature capture capability of traditional models.
[0081] 5. The present invention realizes the fusion of image-text mixed feature map and pixel-text score map through an adaptive multimodal fusion network, which solves the problem that key information cannot be effectively captured by the model due to information redundancy or information loss encountered by traditional models in multimodal task learning.
[0082] 6. The present invention realizes the use of model segmentation results to update the semantic associations and reasoning rules of the spatial-semantic map and the semantic alignment map, solving the problem of insufficient utilization of semantic information in traditional model image segmentation results. BRIEF DESCRIPTION OF THE DRAWINGS
[0083] In order to more clearly illustrate the embodiments of the present invention or the technical solutions in the prior art, the following briefly introduces the drawings required for use in the embodiments. Obviously, the drawings described below are only some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without creative work. In the drawings:
[0084] Figure 1is a flow chart of a SegTCLIP glioma segmentation method based on data amplification and semantic graph according to an embodiment of the present invention;
[0085] Figure 2 is an application diagram of a SegTCLIP glioma segmentation method based on data amplification and semantic graph according to an embodiment of the present invention;
[0086] Figure 3 It is a flowchart of using an improved diffusion model to fill in missing modalities in a SegTCLIP glioma segmentation method based on data amplification and semantic graph according to an embodiment of the present invention;
[0087] Figure 4 It is a flowchart of generating mixed image and text features by improving the cross-attention mechanism in a SegTCLIP glioma segmentation method based on data amplification and semantic graph according to an embodiment of the present invention;
[0088] Figure 5 It is a flowchart of performing feature fusion through an adaptive multimodal fusion network in a SegTCLIP glioma segmentation method based on data amplification and semantic graph according to an embodiment of the present invention;
[0089] Figure 6 It is a flowchart of semantic association and inference rules of updating the knowledge graph through model segmentation results in a SegTCLIP glioma segmentation method based on data amplification and semantic graph according to an embodiment of the present invention;
[0090] Figure 7 It is a principle block diagram of a SegTCLIP glioma segmentation system based on data amplification and semantic graph according to an embodiment of the present invention.
[0091] In the figure:
[0092] 1. Image processing module; 2. Feature map construction module; 3. Feature decoding module. DETAILED DESCRIPTION
[0093] In order to enable those skilled in the art to better understand the technical solutions in this application, the technical solutions in the embodiments of this application will be clearly and completely described below in conjunction with the drawings in the embodiments of this application. Obviously, the described embodiments are only part of the embodiments of this application, not all of the embodiments. Based on the embodiments in this application, all other embodiments obtained by ordinary technicians in this field without creative work should fall within the scope of protection of this application.
[0094] According to an embodiment of the present invention, a SegTCLIP glioma segmentation method, system and application based on data amplification and semantic graph are provided.
[0095] Since the traditional glioma segmentation model mainly relies on image features and ignores the potential information of clinical text data, the accuracy of glioma segmentation is not high. To address the above problems, first, the image is preprocessed and the missing modalities are supplemented. Secondly, the glioma image and the corresponding text data are input into the SegTCLIP large model for feature extraction and mapping to embedding vectors. Then, the knowledge graph is fused with the image and text embeddings, and the fused image and text embeddings are combined through the improved cross-attention mechanism to obtain a mixed image and text feature map. Next, the pixel-text score map is calculated. For the feature vector of each image pixel, the similarity between it and the text feature vector is calculated to obtain the pixel-text score map, and the mixed image and text feature map and the pixel-text score map are input into the adaptive multimodal fusion network for feature fusion, and the fused feature map is sent to the decoder for decoding to obtain the final segmentation result. Finally, the model segmentation results are used to update the semantic associations and reasoning rules of the knowledge graph. The semantic segmentation map obtained by the segmentation model is used to extract the semantic information of key areas and convert it into structured data that can be used to update the knowledge graph. The predefined semantic mapping rules are used to correspond the semantic information extracted from the segmentation results to the entities in the knowledge graph, enhance the instance data of the knowledge graph, supplement the newly extracted information to the knowledge graph, create or update entities and their attribute values, and update the relationship strength between existing nodes, thereby improving the segmentation performance of the model.
[0096] The traditional glioma segmentation model has the following problems: 1) The traditional model lacks cross-modal information fusion; 2) The traditional model cannot fill the missing modality; 3) The traditional model lacks semantic understanding and reasoning capabilities; 4) The traditional model lacks alignment and complementarity between images and texts; 5) The traditional model is difficult to dynamically update and optimize.
[0097] To address the above issues, a large glioma segmentation model based on SegTCLIP is used to segment glioma images. Traditional glioma segmentation models may process image data or text data separately, and usually cannot perform effective image-text alignment. To address this issue, the SegTCLIP large model uses an improved cross-attention mechanism to align images and texts, and generates mixed feature maps of images and texts, giving full play to the complementary advantages of images and texts in spatial information and semantic information.
[0098] During the training process of traditional glioma segmentation models, due to the problem of missing glioma image modality, the quality of the trained model is often uncertain. To address the above problems, a glioma segmentation model based on SegTCLIP is used to segment glioma images, and the modality is completed through an improved diffusion model. The image and text data are feature extracted and mapped into embedding vectors through the improved CLIP model, and the knowledge graph is fused with the image and text embedding to obtain the fused image and text embedding. The image and text embeddings are aligned through the improved cross-attention mechanism, and the image-text mixed feature map is output. By calculating the cosine similarity between each pixel of the image and the text description, a pixel-text score map is generated and the score map and the image-text mixed feature map are sent to the adaptive multimodal fusion network for feature fusion. Finally, the fused feature map is sent to the decoder for decoding to obtain the final segmentation result.
[0099] The present invention will now be further described with reference to the accompanying drawings and specific embodiments. Figure 1 As shown, according to a first embodiment of the present invention, a SegTCLIP glioma segmentation method based on data amplification and semantic graph is provided, comprising:
[0100] S1. Preprocess the pre-collected glioma images by spatial alignment normalization, and perform missing modality filling on the preprocessed glioma images by using an improved diffusion model;
[0101] It should be noted that if Figure 3 As shown in the figure, the original glioma images are collected. Since the original glioma images may cause spatial deformation due to patient posture, equipment differences or image acquisition angle problems, the glioma images need to be spatially aligned. Through spatial alignment, the spatial deformation caused by patient posture, equipment differences or image acquisition angles is eliminated; through normalization technology, the image grayscale values of different patients are adjusted to a uniform scale. Determine whether the glioma image is missing modality; for the image data of patients with missing modalities, the improved diffusion model is used to fill the modality.
[0102] As a preferred implementation, the preprocessing of the pre-collected glioma images by spatial alignment normalization and the missing modality filling processing of the preprocessed glioma images by an improved diffusion model include:
[0103] S11, collecting glioma images, performing spatial deformation elimination processing on the glioma images by using a spatial alignment method, and cropping non-glioma areas in the glioma images;
[0104] S12, performing scale-unified processing on the grayscale values of the cropped glioma image by using a normalization technique to obtain a normalized glioma image;
[0105] S13. Determine whether the normalized glioma image is missing a modality through multimodal image registration, and perform modality filling processing on the glioma image with missing modality through the improved diffusion model.
[0106] As a preferred implementation, the method of determining whether the normalized glioma image is missing a modality by multimodal image registration, and performing modality filling processing on the glioma image with missing modality by the improved diffusion model includes:
[0107] S131, collecting a benchmark reference template image, and registering the normalized glioma image using the benchmark reference template image;
[0108] S132, dynamically determining the size of the cropping frame according to the distribution range of the tumor and cropping the registered glioma image to obtain an image containing a key area of the tumor;
[0109] S133. Determine whether the patient contains all expected modalities based on the image containing the key area of the tumor. If not, predict the missing modal map through the improved diffusion model, and fill the missing modal map into the key area image containing the tumor that does not contain all expected modalities.
[0110] As a preferred embodiment, the method of predicting the missing modal diagram by using the improved diffusion model comprises the following steps:
[0111] S1331. For an image containing a key tumor area, determine the existing modality of the image and use a convolutional neural network of a diffusion model to extract features, and obtain a fused feature map by fusing the extracted features;
[0112] S1332, gradually adding standard normally distributed noise to the fused feature map to obtain a pure noise map;
[0113] S1333, for an image containing a key tumor area, calculating the image gradient of the image, and calculating the standard deviation of the image gradient using a standard deviation calculation formula, comparing the standard deviation of the image gradient with a preset threshold, and determining the number of denoising steps according to the comparison result;
[0114] The standard deviation calculation formula is:
[0115]
[0116] In the formula, σ represents the standard deviation of the image gradient, M represents the number of rows in the image, N represents the number of columns in the image, G(i,j) represents the comprehensive gradient amplitude of the current pixel (i,j), and G avg Represents the average gradient magnitude of the image;
[0117] S1334, generating a spatial attention map according to the standard deviation of the image gradient and the fused feature map, and generating a dynamic spatial attention score map in combination with the dynamic adjustment factor;
[0118] S1335. Based on the dynamic spatial attention score map and in combination with the modality consistency constraint loss function, perform preliminary denoising on the pure noise image to obtain an image feature map of the missing modality predicted after the noise is removed, and decode the predicted missing modality image feature map through a decoder to obtain a predicted missing modality image;
[0119] The expression of the modal consistency constraint loss function is:
[0120]
[0121] In the formula, represents the modal consistency constraint loss function, X n represents the dynamic spatial attention score map, Φ fusion represents the fused feature map, ||·|| represents the weighted L2 consistency norm, which is used to measure X n With Φ fusion The consistency between them.
[0122] Specifically, by judging whether each patient contains all expected modalities; using existing modality information as the diffusion model input, and introducing patient clinical labels to enhance model understanding; extracting spatial and semantic features of all available modalities through a shared encoder to provide guidance for modality completion; introducing global and local features of existing modalities in the diffusion process, generating realistic modalities through modality consistency constraints, and focusing on the lesion area through the spatial attention module to improve the generation accuracy; for complex lesion areas, dynamically adjusting the number of diffusion steps, and using a high-frequency texture loss function to retain details and avoid overly smooth generation.
[0123] In addition, the principle of model consistency constraint is to introduce the fusion feature map of the existing modality into each denoising step of the diffusion model, and design a consistency loss function to ensure that the generated missing modality is semantically consistent with the existing modality. The specific method is as follows:
[0124] First, assuming that there are modalities T1 and T2, the purpose is to generate FLAIR modal images, and input the spatially aligned T1 modal image I T1 , T2 modality image I T2 , the generated FLAIR image To calculate the texture complexity of the input image, first calculate the gradient of the input image:
[0125] G x (i,j)=I(i+1,j)-I(i-1,j);
[0126] G y (i,j)=I(i,j+1)-I(i,j-1);
[0127]
[0128] Among them, G x (i, j) represents the horizontal gradient of the image, which is the difference between the current pixel (i, j) and its adjacent pixels in the horizontal direction. y (i, j) represents the vertical gradient of the image, which means the difference between the current pixel and the adjacent pixels in the vertical direction. G(i, j) represents a comprehensive gradient amplitude obtained by combining the horizontal and vertical gradients, which represents the total degree of change of the image. avg is the average gradient amplitude of the image, M and N are the number of rows and columns of the image, and then the standard deviation σ of the image gradient is calculated. The larger the standard deviation, the higher the texture complexity of the image.
[0129] Then the standard deviation σ of the image gradient is compared with the preset standard deviation threshold T to determine the number of steps in the denoising process. The convolutional neural network is used to extract features from the image to obtain the feature maps of the two modal images: Φ T1 , Φ T2 .
[0130] The fused feature map of the two feature maps is obtained by element-wise addition:
[0131] Φ fusion =Φ T1 +Φ T2 ;
[0132] P fusion Generate spatial attention map:
[0133] A spatial =κ(W s ·Conv(Φ fusion ));
[0134] Among them, A spatial Represents the generated spatial attention map, Conv(Φ fusion ) means that the fused feature map passes through a 1×1 convolution operation layer, W s represents the weight parameter matrix, and κ represents the Sigmoid activation function.
[0135] P fusion Gradually add standard normal distribution noise ε t , until Φ fusion It becomes a pure noise image, and the noise at each time step t is added as follows:
[0136]
[0137] Among them, ε t represents the noise added in step t, represents an attenuation coefficient, expressed as:
[0138]
[0139] Among them, α i represents the noise control parameter associated with each time step, set to α i =1-β t , β t Represents the noise intensity added at each time step. In this process, as t increases, the noise gradually becomes larger and the feature map begins to be polluted until t = T, when the feature map becomes a pure noise map. The pure noise feature map is gradually denoised. First, the standard deviation of the image gradient is compared with the threshold T to determine the number of denoising steps:
[0140]
[0141] Among them, K 10% Indicates that 10% noise is removed in each step during the denoising process, K 5% It means that 5% noise is removed in each step during the denoising process. Assuming that 10% noise is removed in each step, the specific denoising formula is as follows:
[0142]
[0143] Where n∈{1, 2, ..., 10} represents the number of steps, x n represents the current noise map, ε n represents the current noise level, D represents the convolutional neural network model used for denoising, and α k Indicates the denoising ratio of the kth step, and H indicates the total number of denoising steps. In the denoising process, the dynamic spatial attention module is introduced to make the model pay more attention to the tumor area in the image during the denoising process:
[0144]
[0145] in, Represents the dynamically adjusted spatial attention score map, A spatial= represents the generated spatial attention map, and χ represents the dynamic adjustment factor. Since the feature map usually contains a lot of noise in the early stage of denoising, directly multiplying the spatial attention map with these noisy feature maps will not immediately produce significant results. Therefore, the χ dynamic adjustment factor is introduced to gradually strengthen the weight of the spatial attention map as the denoising progresses. In this way, in the early stage of denoising, the spatial attention map will not exert too much influence on the noise feature map, and as the denoising process deepens, the attention map gradually strengthens its guiding role on the tumor area.
[0146] In each step of denoising, the feature map Φ of T1 and T2 modal fusion is fusion It is introduced and the modality consistency constraint loss function is added. The final output is the image feature map of the missing modality predicted after noise removal. The decoder is used to decode it to obtain the predicted missing modality image.
[0147] S2. The SegTCLIP large model encoder is used to extract features from the padded glioma images and the corresponding text data and map them into embedding vectors. The cross-attention mechanism with improved hybrid knowledge graph is combined to construct a mixed feature graph of images and texts.
[0148] As a preferred implementation, the SegTCLIP large model encoder is used to extract features from the padded glioma image and the corresponding text data and map them into embedded vectors, and the cross-attention mechanism combined with the knowledge graph hybrid improvement is used to construct a mixed image-text feature graph, including:
[0149] S21. Configure the network structure of the SegTCLIP large model, optimize the network structure of the SegTCLIP large model through the weight of the pre-trained model CLIP, and obtain the SegTCLIP large model;
[0150] As a preferred embodiment, the configuration of the network structure of the SegTCLIP large model, optimizing the network structure of the SegTCLIP large model by the weight of the pre-trained model CLIP, and obtaining the SegTCLIP large model includes:
[0151] S211. Establish the initial SegTCLIP large model, determine the number and connection order of the image coding (Patch Embedding) module, text coding (Text Embedding) module, convolution layer, pooling layer, batch normalization (Batch Normalization, BN) layer and fusion module, and define the overall network structure;
[0152] S212. According to the overall network structure, optimize the core parameters of the image encoding module, text encoding module and fusion module; such as the number of convolution layers, convolution kernel size and activation function type of the image encoding module, the number of Transformer layers and attention heads of the text encoding module, the parameters of the improved cross attention mechanism, the stride of the pooling layer, the scaling factor of the BN layer, etc.;
[0153] S213. Use the weights of the pre-trained model CLIP to initialize the image and text encoding modules to obtain the final SegTCLIP large model; based on the optimized parameters, the existing cross-modal feature alignment capability of the pre-trained model CLIP can be inherited.
[0154] S22, obtaining a text description of the padded glioma image, and inputting the text description and the padded glioma image into the SegTCLIP large model;
[0155] S23, extract the image spatial information features and text description semantic features respectively through the image encoding module and text encoding module in the SegTCLIP large model, and map them into embedding vectors to obtain the image embedding vector and the text embedding vector.
[0156] Specifically, after obtaining the embedding vector, the embedding vector is passed through the convolution layer to extract local features, the feature space dimension is reduced through the pooling layer, and the data is normalized through the BN layer to avoid the feature distribution being too dispersed or skewed;
[0157] S24. Use knowledge graph technology to fuse the image embedding vector and the text embedding vector, and combine the fused text embedding vector and image embedding vector through an improved cross-attention mechanism to obtain a mixed image-text feature map.
[0158] As a preferred implementation, the image embedding vector and the text embedding vector are fused by using the knowledge graph technology, and the fused text embedding vector and the image embedding vector are combined by an improved cross attention mechanism to obtain a mixed image-text feature map including:
[0159] S241, mapping the image embedding vector and the text embedding vector into query, key and value;
[0160] S242, calculating the attention weight by calculating the dot product between the image embedding query and the text embedding key, and constructing an attention weight matrix;
[0161] S243, using the attention weight matrix to perform weighted summation on the value vector of the text embedding vector to obtain an initial image-text mixed feature map;
[0162] S244, using a level dynamic weight control mechanism, adjusting the initial image-text mixed feature map to obtain a final image-text mixed feature map;
[0163] Specifically, if Figure 4 As shown, first, the input text is extracted through the named entity recognition (NER) model to extract key medical terms or entities (such as "glioma", "T1 enhanced image", "unclear border", etc.). Suppose the input text is: the patient's glioma is located in the left parietal lobe, and the T1 enhanced image shows that the tumor boundary is unclear. The entities extracted by the NER model are: disease: glioma, image modality: T1 enhanced image, description feature: unclear boundary. Then, search for supplementary information from the knowledge graph, use the extracted entities as query conditions, obtain information related to these entities from the knowledge graph, and combine the retrieved information with the original text description to form a supplemented text description. The specific search and combination formulas are as follows:
[0164] Assume that the input text is T = {W 1 ,W 2 ,...,W n}, the NER model identifies the entity category as y = {y 1 ,y 2 ,...,y n}, y represents the set of all possible entity labels, P(y i |w 1 ,w 2 ,...,w n ) is the probability distribution calculated by the NER model, which is used to represent the possibility of entity labels.
[0165] Among them, W n Indicates the nth word of the input text, y n represents the entity tag predicted by the nth word, and Y represents all possible entity tag categories.
[0166] Constructed glioma knowledge graph:
[0167] G = {V, U};
[0168] Among them, G represents the glioma knowledge graph, V represents the node set in the graph, representing the concepts in the knowledge graph (such as tumor type, image modality, etc.), and U represents the edge set between nodes, representing the semantic relationship between nodes (such as contains, located in, etc.), where:
[0169] W related = Qurey(y);
[0170] Where W related represents the relevant supplementary information queried by entity category, Qurey(y) represents the query of node y, and W relatedThe image and the supplemented text are supplemented to form a supplemented text description. Then, the image and the supplemented text are extracted and fused through the improved cross attention mechanism to obtain a mixed image-text feature map. The specific method is as follows:
[0171] First, the four modalities of the input glioma image are weighted averaged to perform feature extraction and feature fusion to obtain a multimodal image feature map:
[0172]
[0173] Among them, f image represents the multimodal image feature map, w T1 represents the T1 modal weight, w T2 represents the T2 modal weight, w FLAIR represents the Flair modal weight, w T1_enhanced Represents T1_enhanced modality weight.
[0174] Then perform feature extraction on the input text to obtain the text feature map:
[0175] f text =TextEncoder(T);
[0176] In the formula, f text Represents a text feature map, and TextEncoder(T) represents text encoding of T;
[0177] Then two fully connected layers are used to map the image features and text features to the same dimension:
[0178] e image =W image ·f image +b image , e text =W text ·E text +b text ;
[0179] Among them, W image , W test Represents the weight matrices of image feature and text feature mapping, b image , b test Represent the image feature and text feature mapping bias items respectively. By introducing the cross attention of the multi-level dynamic weight control mechanism, the image and text embeddings are combined to obtain the image-text mixed feature map. First, the image embedding and text embedding are mapped to the query Q image , key K test Sum value V test :
[0180] Q image =WQ ·e image +b Q ; K text =W K ·e text +b K ; V text =W V ·e text +b V ;
[0181] Among them, W Q , W K , W V Represents the weight matrices of query, key and value respectively, b Q , b K , b V Denote the bias terms of query, key and value respectively. The attention weight A is then calculated by calculating the dot product between the image embedding query and the text embedding key:
[0182]
[0183] Where d represents the dimension of the embedding vector, Represents a text embed key.
[0184] Then, the attention weight matrix A is used to perform weighted summation on the value vector of the text to obtain a mixed feature map of text and image, and a multi-level dynamic weight control mechanism is referenced in it:
[0185]
[0186] Among them, e fusion represents the initial image-text mixed feature map, α i represents the weight of each layer, A represents the attention weight, V text Indicates the value, β i Represents the dynamically learned features, which are used to adjust the contribution of the i-layer features. final Represents the final mixed image-text feature map. Dynamic weights are introduced in different cross-attention layers, and these weights can be dynamically adjusted according to the complexity of the image and text and the characteristics of the current input.
[0187] S3, calculating a pixel-text score map based on the multimodal pixel feature map, and inputting the image-text mixed feature map and the pixel-text score map into the multimodal decoding network for feature decoding to obtain a segmentation result;
[0188] As a preferred implementation, the pixel-text score map is calculated based on the multimodal pixel feature map, and the image-text mixed feature map and the pixel-text score map are input into the multimodal decoding network for feature decoding, and the segmentation result obtained includes:
[0189] S31, dividing the glioma image after the missing modality filling process into a number of pixel blocks, extracting pixel features from each pixel block through a convolutional neural network, and fusing the pixel features with those at the same position of other modalities to obtain a multi-modal pixel feature map;
[0190] S32, mapping the pixel features of each pixel block in the multimodal pixel feature map into a pixel embedding vector;
[0191] S33, calculating a pixel-text similarity score according to the pixel embedding vector and the text embedding vector to obtain a pixel-text score graph;
[0192] The calculation formula of the pixel-text similarity score is:
[0193]
[0194] In the formula, S pixel-text represents the pixel-text similarity score, I p represents the pixel embedding vector of the pixel block at the pth position, e text Represents the text embedding vector;
[0195] S34, inputting the pixel-text score map and the image-text mixed feature map into the adaptive multimodal fusion network, using an adaptive strategy, and dynamically adjusting the weights of the two features according to the learning, and generating a fused multimodal feature map through decoding;
[0196] As a preferred implementation, the pixel-text score map and the image-text mixed feature map are input into an adaptive multimodal fusion network, an adaptive strategy is used, and the weights of the features of the two are dynamically adjusted according to learning, and the fused multimodal feature map is generated by decoding, including:
[0197] S341, using the image-text mixed feature map as the input of the adaptive multimodal fusion network, and decoding it through the decoder of the adaptive multimodal fusion network to obtain a decoded feature map;
[0198] S342, using an adaptive strategy and dynamically adjusting a hyperparameter of the influence of the pixel-text score map on decoding according to the learning;
[0199] S343, according to the hyperparameter of the influence degree of the pixel-text score map on decoding, weightedly fuse the decoded decoding feature map and the pixel-text score map to obtain a fused multimodal feature map;
[0200] The expression for weighted fusion of the decoded feature map and the pixel-text score map is:
[0201] S35. The final glioma segmentation map is obtained through multiple decoding.
[0202] Specifically, if Figure 5 As shown in the figure, the four input modality images T1, T2, Flair and T1_enhanced are first divided into L pixel blocks, and each pixel block is extracted through a convolutional neural network. The features are fused with the pixel features at the same position of other modalities to obtain a feature map that integrates the pixel information of the four modalities. At the same time, the features of each pixel block in the feature map are mapped to an embedded vector: I n ,n∈(1,2,...,L).
[0203] Then the cosine similarity between each pixel block and the text embedding is calculated to obtain a pixel-text similarity score map, which reflects the degree of match between each pixel block and the relevant information in the text description.
[0204] The pixel-text score map and the image-text mixed feature map are fed into the multimodal decoder. The pixel-text score map guides the decoding process of the image-text mixed feature map, ensuring that the decoding result is spatially and semantically consistent with the input text description.
[0205] Specifically, in each decoding step, the pixel-text score map extracts local features through the convolution layer, restores the feature space dimension through the upsampling layer, normalizes the data through the BN layer, and is fused with the decoder output through a weighted mechanism. According to the similarity information of each pixel in the score map, the decoding direction of the image feature map is adjusted. Finally, after several steps of decoding, the network outputs the final glioma region segmentation map. The specific steps are as follows: First, the image-text mixed feature map e final As input, the decoder gets the decoded feature map F d Then, at each decoding step, the pixel-text score map is weightedly fused with the current decoder output:
[0206]
[0207] In the formula, represents the fused multimodal feature map, represents the decoding feature map obtained by the current decoder, E represents the hyperparameter of the influence of the pixel-text score map on decoding, Represents the gradient of the current decoded feature map, indicating the change of the decoder, S pixel-text (i, j) represents the value of the pixel-text score map, which indicates the similarity between each pixel and the text description.
[0208] After multiple decoding, the final glioma segmentation map is obtained
[0209] like Figure 2As shown, according to a second embodiment of the present invention, a SegTCLIP glioma segmentation method based on data augmentation and semantic graph is provided for use in updating semantic associations and inference rules of a knowledge graph, including:
[0210] Perform semantic segmentation on the glioma segmentation map through the segmentation model to obtain a semantic segmentation map, and extract semantic information of key areas in the semantic segmentation map;
[0211] Using predefined semantic mapping rules, the extracted semantic information is mapped with the entities and relationship nodes in the knowledge graph;
[0212] Update the semantic associations in the knowledge graph and optimize the inference rules so that it can process and analyze the newly added semantic information. Use the updated knowledge graph to more accurately guide the cross-modal fusion of image features and text features, thereby improving the segmentation performance of the model.
[0213] Specifically, if Figure 6 As shown, firstly, the glioma segmentation map Extract spatial information. It contains the following areas: S core (tumor core area), S enhanced (enhancing tumor area), S edema (edema area), S normal (Normal area).
[0214] These regions contain spatial location information, and the geometric features of these regions are extracted from the segmentation map, including:
[0215] Regional boundaries
[0216] Area shape
[0217] Location
[0218] size
[0219] in, represents the i-th area in the image, express The centroid represents the centroid of the tumor area, and Area represents the area of the tumor.
[0220] Then the segmentation map Extract semantic information. Contains the following semantic information:
[0221] Tumor type tumor =classify(S core ,Senhanced ,S edema );
[0222] Here, classify represents the classification operation. Based on the segmentation results and existing medical knowledge, the type of tumor (such as low-grade, malignant, high-grade, etc.) can be inferred.
[0223] Tumor vascular permeabilityvascular_invasion=bool(mean(S enhanced )>T vascular );
[0224] Among them, mean(S enhanced ) represents the average strength of the enhanced signal, bool represents whether it exceeds the threshold, T vascular It represents the preset signal threshold, which indicates the lower limit of the enhanced signal strength. If it exceeds the threshold, it is judged as enhanced vascular permeability.
[0225] Relationship between edema area and tumor core area
[0226] Among them, S core ⌒S edema represents the overlapping area of the tumor core and edema area, S core ∪S edema is the combined area of the two regions.
[0227] Next, the extracted spatial and semantic information is mapped to the corresponding entities in the knowledge graph:
[0228] E tumor ={r1:location(r 1 ),size(r 1 ),type(r 1 )};
[0229] Among them, E tumor represents the tumor entity in the knowledge graph, r 1 Represents the extracted tumor region entity, including information such as location, size, and type.
[0230] In the atlas, tumor entities have multiple attributes, such as morphology, malignancy, vascular permeability, etc. The semantic information of the segmentation results is used to update the attributes of the tumor: Based on the spatial relationships extracted from the segmentation results, the relationships between different entities are added or updated in the knowledge graph:
[0231] R relation ={hasCore(r 1 ,r 2 ),enhances(r2 ,r 3 )};
[0232] Among them, R relation represents the relationship between tumor regions, hasCore represents the inclusion relationship between the tumor core region and the enhancement region, enhances represents the extension relationship between the enhancement region and the edema region, and r 2 ,r 3 Represents different tumor region entities. By extracting spatial information and semantic information from the segmentation results and using them to update the semantic associations and reasoning rules of the knowledge graph, after each segmentation is completed, the knowledge graph is dynamically updated based on the reasoning results, thereby improving the accuracy of glioma segmentation.
[0233] In addition, in the present invention, by completing the missing modalities and combining the semantic information of images and texts, the segmentation accuracy of the model is improved under conditions of multimodal data missing, blurred tumor boundaries, and data scarcity.
[0234] At the same time, by designing a large glioma segmentation model based on SegTCLIP, image and text semantic information are intelligently integrated for intelligent glioma segmentation, which solves the problems of missing multimodal data modality, blurred tumor boundaries and low segmentation accuracy under data scarcity conditions, while improving the model's generalization ability and adaptability to complex tumor morphologies.
[0235] First, the diffusion model is used to complete the missing modalities, and then the text-driven CLIP model that fuses the image knowledge graph and the corresponding text knowledge graph is used to extract features and understand the semantics of the fused multimodal data. At the same time, a multimodal fusion network is used to fuse the image-text mixed feature map and the pixel-text score map, so that the network can focus on global semantic consistency while retaining local detail features.
[0236] It should be noted that convolutional neural networks (CNNs) have been widely used in various fields, but they still face uncertainties and challenges in the learning process. In image processing tasks, since different images present diverse feature expressions, simply relying on the extraction of image features may ignore the semantic information behind the image. The fusion of images and texts can effectively make up for this defect. By combining the semantic information in the text, the model can more comprehensively understand the context and details of the image, thereby improving the overall cognition of the image content. This fusion not only helps the model better understand the objects, scenes and relationships in the image, but also provides a richer and more accurate semantic background for tasks such as image segmentation and target recognition. With the rise of large models, these models can capture more complex and rich features through large-scale parameter spaces and deep network structures, thereby more comprehensively extracting details and semantic information when processing images. Large models can play an important role in the learning and extraction of image feature information through higher-dimensional semantic representations, and enhance cross-modal understanding capabilities. Large models can accurately capture the semantic connection between images and texts by deeply learning the relationship between text and images in large corpora, especially in multimodal learning scenarios, and can effectively solve complex semantic matching problems. Therefore, by utilizing large models for image feature learning, the limitations of traditional convolutional neural networks in complex tasks are overcome, and the accuracy of image understanding is significantly improved.
[0237] In image segmentation models, reasoning usually relies only on information about the image itself, while ignoring the potential impact of the segmentation results on the semantic structure in the knowledge graph, which may lead to inaccurate semantic matching between the image and text knowledge graphs. To solve this problem, a technology for updating image and text knowledge graphs based on image segmentation results is proposed, aiming to dynamically adjust entities and relationships in the knowledge graph through detailed information of image segmentation. In this way, the segmentation results can provide new contextual information for the knowledge graph, thereby achieving a close integration of image information and text knowledge, so that the knowledge graph can more accurately reflect the actual content and semantic information in the image. The application of this method is expected to improve the effect of cross-modal learning, enhance the semantic understanding of images and texts, and provide more accurate and rich knowledge support for subsequent reasoning tasks.
[0238] like Figure 7 As shown, according to a third embodiment of the present invention, a SegTCLIP glioma segmentation system based on data amplification and semantic graph is provided, comprising:
[0239] An image processing module 1 is used to pre-process the pre-collected glioma image by spatial alignment normalization, and to perform missing modality filling processing on the pre-processed glioma image by an improved diffusion model;
[0240] Feature map construction module 2 is used to extract features from the padded glioma image and the corresponding text data through the SegTCLIP large model encoder and map them into embedded vectors, and construct a mixed image-text feature map by combining the cross-attention mechanism with the knowledge graph hybrid improvement;
[0241] The feature decoding module 3 is used to calculate the pixel-text score map based on the multimodal pixel feature map, and input the image-text mixed feature map and the pixel-text score map into the multimodal decoding network for feature decoding to obtain the segmentation result.
[0242] In summary, with the help of the above technical solutions of the present invention, the present invention realizes a large glioma segmentation model based on SegTCLIP, which can be used for glioma image segmentation, solves the problem of missing glioma image modality and insufficient fusion of glioma image and text information, and improves the accuracy of complex glioma region segmentation. The present invention realizes the use of diffusion model to complete the missing modality of glioma, solves the problem of missing image data modality, ensures the integrity of multimodal data, and thus improves the diagnosis and segmentation accuracy of glioma. The present invention realizes the use of spatial-semantic graph and semantic alignment graph to fuse image and text embedding respectively, solves the problem that traditional visual methods are difficult to capture fuzzy or complex features, and improves the ability of the model to understand features. The present invention realizes the combination of the fused image and text embedding through an improved cross-attention mechanism to obtain a mixed feature map of images and texts, which solves the problems of insufficient feature fusion, insufficient information complementarity and weak fine-grained feature capture in traditional models. The present invention achieves the fusion of image-text mixed feature maps and pixel-text score maps through an adaptive multimodal fusion network, solving the problem that key information cannot be effectively captured by the model due to information redundancy or information loss encountered by traditional models in multimodal task learning. The present invention achieves the use of model segmentation results to update the semantic associations and inference rules of the spatial-semantic map and the semantic alignment map, solving the problem of insufficient utilization of semantic information in the image segmentation results of traditional models.
[0243] It will be appreciated by those skilled in the art that embodiments of the present invention may be provided as methods, systems or computer program products. Therefore, the present invention may take the form of a complete hardware embodiment, a complete software embodiment or an embodiment combining software and hardware. Moreover, the present invention may take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk storage, optical storage, etc.) containing computer-usable program code.
[0244] The specific embodiments described above further illustrate the objectives, technical solutions and beneficial effects of the present invention in detail. It should be understood that the above description is only a specific embodiment of the present invention and is not intended to limit the scope of protection of the present invention. Any modifications, equivalent substitutions, improvements, etc. made within the spirit and principles of the present invention should be included in the scope of protection of the present invention.
Claims
1. A SegTCLIP glioma segmentation method based on data amplification and semantic graph, characterized in that: include: S1. Preprocess the pre-collected glioma images by spatial alignment normalization, and perform missing modality filling on the preprocessed glioma images by using an improved diffusion model; S2. The SegTCLIP large model encoder is used to extract features from the padded glioma images and the corresponding text data and map them into embedding vectors. The cross-attention mechanism with improved hybrid knowledge graph is combined to construct a mixed feature graph of images and texts. S3. Calculate the pixel-text score map based on the multimodal pixel feature map, and input the image-text mixed feature map and the pixel-text score map into the multimodal decoding network for feature decoding to obtain the segmentation result.
2. The SegTCLIP glioma segmentation method based on data amplification and semantic graph according to claim 1, characterized in that: The preprocessing of the pre-collected glioma image by spatial alignment normalization and the missing modality filling processing of the preprocessed glioma image by the improved diffusion model include: S11, collecting glioma images, performing spatial deformation elimination processing on the glioma images by using a spatial alignment method, and cropping non-glioma areas in the glioma images; S12, performing scale-unified processing on the grayscale values of the cropped glioma image by using a normalization technique to obtain a normalized glioma image; S13. Determine whether the normalized glioma image is missing a modality through multimodal image registration, and perform modality filling processing on the glioma image with missing modality through the improved diffusion model.
3. The SegTCLIP glioma segmentation method based on data amplification and semantic graph according to claim 2, characterized in that: The method of judging whether the normalized glioma image is missing a modality by multimodal image registration, and performing modality filling processing on the glioma image missing a modality by the improved diffusion model includes: S131, collecting a benchmark reference template image, and registering the normalized glioma image using the benchmark reference template image; S132, dynamically determining the size of the cropping frame according to the distribution range of the tumor and cropping the registered glioma image to obtain an image containing a key area of the tumor; S133. Determine whether the patient contains all expected modalities based on the image containing the key area of the tumor. If not, predict the missing modal map through the improved diffusion model, and fill the missing modal map into the key area image containing the tumor that does not contain all expected modalities.
4. The SegTCLIP glioma segmentation method based on data amplification and semantic graph according to claim 3, characterized in that: The method of predicting the missing modal diagram by using the improved diffusion model comprises the following steps: S1331. For an image containing a key tumor area, determine the existing modality of the image and use a convolutional neural network of a diffusion model to extract features, and obtain a fused feature map by fusing the extracted features; S1332, gradually adding standard normally distributed noise to the fused feature map to obtain a pure noise map; S1333, for an image containing a tumor key area, calculating the image gradient of the image, and calculating the standard deviation of the image gradient using a standard deviation calculation formula, comparing the standard deviation of the image gradient with a preset threshold, and determining the number of denoising steps according to the comparison result; The standard deviation calculation formula is: In the formula, σ represents the standard deviation of the image gradient, M represents the number of rows in the image, N represents the number of columns in the image, G(i,j) represents the comprehensive gradient amplitude of the current pixel (i,j), and G avg Represents the average gradient magnitude of the image; S1334, generating a spatial attention map according to the standard deviation of the image gradient and the fused feature map, and generating a dynamic spatial attention score map in combination with the dynamic adjustment factor; S1335. Based on the dynamic spatial attention score map and in combination with the modality consistency constraint loss function, perform preliminary denoising on the pure noise image to obtain an image feature map of the missing modality predicted after the noise is removed, and decode the predicted missing modality image feature map through a decoder to obtain a predicted missing modality image; The expression of the modal consistency constraint loss function is: In the formula, represents the modal consistency constraint loss function, X n represents the dynamic spatial attention score map, Φ fusion Represents the fused feature map.
5. The SegTCLIP glioma segmentation method based on data amplification and semantic graph according to claim 1, characterized in that: The SegTCLIP large model encoder is used to extract features from the padded glioma image and the corresponding text data and map them into embedded vectors, and the cross-attention mechanism combined with the knowledge graph hybrid improvement is used to construct a mixed image-text feature graph, including: S21. Configure the network structure of the SegTCLIP large model, optimize the network structure of the SegTCLIP large model through the weight of the pre-trained model CLIP, and obtain the SegTCLIP large model; S22, obtaining a text description of the padded glioma image, and inputting the text description and the padded glioma image into the SegTCLIP large model; S23, extracting the image spatial information features and the text description semantic features respectively through the image encoding module and the text encoding module in the SegTCLIP large model, and mapping them into embedding vectors to obtain the image embedding vector and the text embedding vector; S24. Use knowledge graph technology to fuse the image embedding vector and the text embedding vector, and combine the fused text embedding vector and image embedding vector through an improved cross-attention mechanism to obtain a mixed image-text feature map.
6. The SegTCLIP glioma segmentation method based on data amplification and semantic graph according to claim 5, characterized in that: The configuration of the network structure of the SegTCLIP large model, optimizing the network structure of the SegTCLIP large model by the weight of the pre-trained model CLIP, and obtaining the SegTCLIP large model includes: S211. Establish the initial SegTCLIP large model, determine the number of image encoding modules, text encoding modules, and fusion modules and their connection order, and define the overall network structure; S212. Optimizing core parameters of the image encoding module, the text encoding module, and the fusion module according to the overall network structure; S213. Use the weights of the pre-trained model CLIP to initialize the image and text encoding modules to obtain the final SegTCLIP large model.
7. The SegTCLIP glioma segmentation method based on data amplification and semantic graph according to claim 5, characterized in that: The image embedding vector and the text embedding vector are fused by using the knowledge graph technology, and the fused text embedding vector and the image embedding vector are combined by using the improved cross attention mechanism to obtain the image-text mixed feature map including: S241, mapping the image embedding vector and the text embedding vector into query, key and value; S242, calculating the attention weight by calculating the dot product between the image embedding query and the text embedding key, and constructing an attention weight matrix; S243, using the attention weight matrix to perform weighted summation on the value vector of the text embedding vector to obtain an initial image-text mixed feature map; The expression for weighted summation of the value vector of the text embedding vector using the attention weight matrix is: e fusion =a i ·A·V text ; In the formula, e fusion represents the initial image-text mixed feature map, α i represents the weight of each layer, A represents the attention weight, V text Represents a value; S244, using a level dynamic weight control mechanism, adjusting the initial image-text mixed feature map to obtain a final image-text mixed feature map; The expression for adjusting the initial image-text mixed feature map using the level-by-level dynamic weight control mechanism is: In the formula, e final represents the final image-text mixed feature map, β i represents the dynamically learned features, e fusion Represents the initial image-text mixed feature map, and i represents the index value.
8. The SegTCLIP glioma segmentation method based on data amplification and semantic graph according to claim 1, characterized in that: The pixel-text score map is calculated based on the multimodal pixel feature map, and the image-text mixed feature map and the pixel-text score map are input into the multimodal decoding network for feature decoding, and the segmentation result obtained includes: S31, dividing the glioma image after the missing modality filling process into a number of pixel blocks, extracting pixel features from each pixel block through a convolutional neural network, and fusing the pixel features with those at the same position of other modalities to obtain a multi-modal pixel feature map; S32, mapping the pixel features of each pixel block in the multimodal pixel feature map into a pixel embedding vector; S33, calculating a pixel-text similarity score according to the pixel embedding vector and the text embedding vector to obtain a pixel-text score graph; The calculation formula of the pixel-text similarity score is: In the formula, S pixel-text represents the pixel-text similarity score, I p represents the pixel embedding vector of the pixel block at the pth position, e text Represents the text embedding vector; S34, inputting the pixel-text score map and the image-text mixed feature map into the adaptive multimodal fusion network, using an adaptive strategy, and dynamically adjusting the weights of the two features according to the learning, and generating a fused multimodal feature map through decoding; S35. The final glioma segmentation map is obtained through multiple decoding.
9. The SegTCLIP glioma segmentation method based on data amplification and semantic graph according to claim 8, characterized in that: The pixel-text score map and the image-text mixed feature map are input into the adaptive multimodal fusion network, and the weights of the features of the two are adjusted dynamically according to the learning using the adaptive strategy, and the fused multimodal feature map is generated by decoding, including: S341, using the image-text mixed feature map as the input of the adaptive multimodal fusion network, and decoding it through the decoder of the adaptive multimodal fusion network to obtain a decoded feature map; S342, using an adaptive strategy and dynamically adjusting a hyperparameter of the influence of the pixel-text score map on decoding according to the learning; S343, according to the hyperparameter of the influence degree of the pixel-text score map on decoding, weightedly fuse the decoded decoding feature map and the pixel-text score map to obtain a fused multimodal feature map; The expression for weighted fusion of the decoded feature map and the pixel-text score map is: In the formula, represents the fused multimodal feature map, represents the decoding feature map obtained by the current decoder, E represents the hyperparameter of the influence of the pixel-text score map on decoding, Represents the gradient of the current decoded feature map, S pixel-text (i,j) represents the value of the pixel-text score map.
10. Application of the SegTCLIP glioma segmentation method based on data amplification and semantic graph in updating the semantic association and inference rules of the knowledge graph according to any one of claims 1 to 9, characterized in that: include: Perform semantic segmentation on the glioma segmentation map through the segmentation model to obtain a semantic segmentation map, and extract semantic information of key areas in the semantic segmentation map; Using predefined semantic mapping rules, the extracted semantic information is mapped with the entities and relationship nodes in the knowledge graph; Update semantic associations in the knowledge graph and optimize inference rules.
Citation Information
Cited By
Medical image segmentation method and system based on image-text interaction
CN120707586A
Unbiased missing modal learning method based on multi-stage double diffusion network
CN121117499A
An unbiased missing modal learning method based on a multi-stage double diffusion network
CN121117499B