PET-CT image fusion segmentation method based on prompt information

Through the PET-CT image fusion segmentation method based on prompt information, the PMC-CLIP model and nnUNet network are used to solve the missegment problem caused by the similarity between normal organ tissues and lymphoma in PET images, and achieve higher segmentation accuracy and accuracy.

CN120163831APending Publication Date: 2025-06-17ZHEJIANG UNIV OF TECH
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510086796.5
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-01-20
Publication Date
2025-06-17

AI Technical Summary

Technical Problem

The missegment problem caused by the similarity between high metabolic activity and lymphoma of normal organ tissues in PET images makes it difficult to accurately distinguish the source of metabolic signals.

Method used

Using the PET-CT image fusion segmentation method based on the prompt information, the initial target prompt information is generated and refined through the PMC-CLIP model and the nnUNet convolutional neural network. Combined with the CT image and PET image features, the missegment interference of highly metabolic active organs on the target is reduced.

Benefits of technology

It effectively suppresses interference in related areas in PET images, improves the accuracy and accuracy of image segmentation, and reduces the situation of missegment.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120163831A_ABST
    Figure CN120163831A_ABST
Patent Text Reader

Abstract

The invention discloses a PET-CT image fusion segmentation method based on prompt information, and the method comprises the steps: extracting a target text feature and a CT image feature through a PMC-CLIP model, carrying out the foreground and background division through the cosine similarity and a threshold value, generating initial prompt information, carrying out the feature learning of PET and CT images through employing an nnUNet convolutional neural network, and carrying out the segmentation of the features of the PET and CT images. And in combination with a CT image second feature extracted by the nnUNet convolutional neural network and prompt information, obtaining a target foreground feature. And then, extracting a prototype vector of the region through K-means clustering, and further refining the prompt information. And finally, splicing the optimized prompt information with the PET image features and the CT image second features to realize accurate image segmentation. According to the method, the interference of high-metabolic-activity organs on target wrong segmentation is reduced, so that the precision of automatic segmentation is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application belongs to the field of computer vision technology, and particularly relates to a PET-CT image fusion and segmentation method based on prompt information. Background Art

[0002] Medical image processing has extremely important applications and significance in the medical field. Through digital image processing technology, doctors can observe the lesion area more clearly.

[0003] Positron emission tomography / computed tomography (PET / CT) is widely regarded as an effective imaging tool for segmenting lymphoma. The standardized uptake value (SUV) in PET images can reflect the metabolic activity of tissues and has high sensitivity for the detection of lymphoma. Since PET has low resolution and high signal-to-noise ratio, it usually needs to be combined with CT images to make up for the deficiencies. CT images mainly provide anatomical structure information, such as the shape, position, and density of organs, but cannot directly reflect the metabolic activity of tissues. While PET images show the tissue activity level through metabolic signals (such as SUV values), these metabolic signals often cover multiple anatomical structures, making it difficult to accurately distinguish whether the source of the metabolic signal is lymphoma or normal organ tissue even with the assistance of CT anatomical information. Summary of the Invention

[0004] The purpose of this application is to provide a PET-CT image fusion and segmentation method based on prompt information to solve the problem of missegmentation caused by the similarity between the high metabolic activity of normal organ tissues and lymphoma in PET images and improve the image segmentation effect.

[0005] To achieve the above purpose, the technical solution of this application is as follows: A PET-CT image fusion and segmentation method based on prompt information, comprising: Input the CT image and the text description corresponding to the target into a pre-trained PMC-CLIP model respectively, and combine the first feature of the CT image and the target text feature extracted to generate initial target prompt information; Input the CT image into the nnUNet convolutional neural network, extract the second feature of the CT image, and combine the second feature of the CT image with the initial target prompt information to generate target foreground features; Cluster the target foreground features into regional features, and perform average pooling on each regional feature to obtain the target prototype vectors corresponding to the respective regional features; Calculate the cosine similarity between the target prototype vectors corresponding to the respective regional features and the second feature of the CT image, and average the results to obtain refined target prompt information; Input the PET image into the nnUNet convolutional neural network to extract the PET image features. Refine the target prompt information, concatenate the PET image features and the second CT image features, generate the final output features, and obtain the segmentation result through the segmentation head.

[0006] Further, combining the extracted first CT image feature and the target text feature to generate the initial target prompt information, including: Calculate the cosine similarity between the first CT image feature and the target text feature; Compare the calculated cosine similarity with the first threshold. If it is greater than the first threshold, mark it as the foreground to generate the initial target prompt information.

[0007] Further, combining the second CT image feature and the initial target prompt information to generate the target foreground feature, including: Multiply the initial target prompt information by the second CT image feature to obtain the target foreground feature.

[0008] Further, averaging the results to obtain the refined target prompt information. The formula is as follows: Where, represents the cosine similarity between the target prototype vector and the second CT image feature, M is the total number of regional features, and the finally obtained represents the refined target prompt information.

[0009] Further, the segmentation head adopts a convolutional layer.

[0010] Further, the loss of the nnUNet convolutional neural network includes Dice coefficient loss and cross-entropy loss.

[0011] This application proposes a PET-CT image fusion segmentation method based on prompt information. Using the PMC-CLIP model and its pre-trained weights, combining the CT image and the text information corresponding to the target to generate the initial target prompt for assisting the segmentation of the PET image, thereby effectively suppressing the interference in the relevant regions of the PET image. Subsequently, refine the generated initial target prompt information, extract the prototypes of multiple regions, and use these prototypes to reactivate the original image to generate the optimized target prompt information, thereby further reducing interference and improving the segmentation effect. This application reduces the interference of highly metabolically active organs on the mis-segmentation of the target, thereby improving the accuracy of automatic segmentation. Brief Description of the Drawings

[0012] Figure 1 It is the flow chart of the PET-CT image fusion segmentation method based on prompt information of this application; Figure 2 This is a schematic diagram of the image segmentation network structure of the present application. Specific implementation manners

[0013] In order to make the objectives, technical solutions and advantages of the present application clearer, the present application will be further described in detail below with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present application and are not used to limit the present application.

[0014] One embodiment of the present application, as Figure 1 shown, provides a PET-CT image fusion segmentation method based on prompt information, including: Step S1: Input the CT image and the text description corresponding to the target into a pre-trained PMC-CLIP model respectively, and generate initial target prompt information by combining the first feature of the CT image and the target text feature extracted.

[0015] The PMC-CLIP model is a pre-trained model based on contrastive learning, aiming to use biomedical literature for language-image learning. This model has achieved leading results in multiple downstream tasks through pre-training on the open-access PubMed Central (PMC) database.

[0016] In the present application, the CT image and the text description corresponding to the target are input into the pre-trained PMC-CLIP model respectively, and the weight parameters of the PMC-CLIP model are frozen for feature extraction. The PMC-CLIP model does not participate in the training and update of the subsequent entire network parameters.

[0017] The input CT image is input into the PMC-CLIP model, and the first feature of the CT image is extracted; the text description corresponding to the target (such as "The liver is located in the upper right abdomen, is wedge-shaped, and the vertical diameter is about 9-11 cm") is input into the PMC-CLIP model, and the target text feature is extracted.

[0018] In a specific embodiment, generating the initial target prompt information by combining the first feature of the CT image and the target text feature includes: Calculating the cosine similarity between the first feature of the CT image and the target text feature, and the specific formula is as follows: Wherein, represents the first feature of the CT image extracted from the CT image, represents the extracted target text feature, It is the result of the cosine similarity between them. The cosine similarity generates a correlation matrix, indicating the potential distribution area of the target in the CT image. Through threshold segmentation, the matrix is divided into foreground and background regions, obtaining initial target hint information, which provides rough target position information for the segmentation task;

[0019] The calculated cosine similarity is compared with the first threshold. If it is greater than the first threshold, it is marked as the foreground, generating initial target hint information.

[0020] In this embodiment, the first threshold is set to the value of the 80th percentile in the current hint information. The part greater than the first threshold is marked as the foreground (value is 1), and the rest is marked as the background (value is 0).

[0021] The specific formula is as follows: ; ; Among them, represents the initial target hint information, Per represents the value at the percentile position of the corresponding value extracted from the data, represents the value of the current 80th percentile.

[0022] Step S2: Input the CT image into the nnUNet convolutional neural network, extract the second feature of the CT image, and combine the second feature of the CT image with the initial target hint information to generate the target foreground feature.

[0023] ‌nnUNet‌ is a deep learning-based medical image segmentation framework. It can automatically adjust all hyperparameters according to the attributes of the given dataset, and the whole process requires no manual intervention. It relies on the simple U-Net structure and robust training scheme, and has achieved state-of-the-art performance in multiple medical image segmentation challenges.

[0024] In this embodiment, the CT image is input into the nnUnet convolutional neural network for feature learning to obtain the second feature of the CT image. Then, the initial target hint information is multiplied by the second feature of the CT image to obtain the target foreground feature.

[0025] The specific formula is as follows: Among them, represents the second feature of the CT image, represents the initial target hint information, is the obtained target foreground feature.

[0026] Step S3: Cluster the target foreground features into regional features, and perform average pooling on each regional feature to obtain the target prototype vector corresponding to each regional feature.

[0027] In this embodiment, K-means clustering is applied to divide the target foreground features into multiple regional features. Then, average pooling is performed on each regional feature to obtain the target prototype vector corresponding to the regional feature.

[0028] The specific formula is as follows: Among them, is the th regional feature obtained by Kmeans clustering, and is the prototype vector of this regional feature after average pooling.

[0029] Step S4: Calculate the cosine similarity between the target prototype vector corresponding to each regional feature and the second feature of the CT image, and average the results to obtain the refined target prompt information.

[0030] In this embodiment, the cosine similarity between multiple regional prototype vectors and the second feature of the CT image is calculated, and the results are averaged.

[0031] The specific formula is as follows: Among them, represents the cosine similarity between the target prototype vector and the second feature of the CT image, M is the total number of regional features, and the finally obtained represents the refined target prompt information.

[0032] Step S5: Input the PET image into the nnUNet convolutional neural network to extract the PET image features, splice the refined target prompt information, the PET image features and the second feature of the CT image to generate the final output features, and obtain the segmentation result through the segmentation head.

[0033] In this embodiment, the refined target prompt information is spliced with the PET image features and the second feature of the CT image to generate the final output features, and then the final output features are input into the segmentation head to obtain the final segmentation prediction result.

[0034] Specifically, a joint loss function composed of the Dice coefficient loss and the cross-entropy loss is used for training to adjust the weight parameters in the network. Specifically, the weight parameters of the nnUNet convolutional neural network are trained. The training of the network is a relatively mature technology in this field and will not be elaborated here. After training the weight parameters of the network, the final output features are input into the segmentation head to obtain the final segmentation prediction result.

[0035] The segmentation head Cls adopts a convolutional layer to obtain the final segmentation prediction result.

[0036] The splicing formula is as follows: .

[0037] Among them, Concat represents the splicing of feature channels, Cls represents the use of a convolutional segmentation head, represents the PET image feature.

[0038] The formula for the combined loss function is as follows: ; where N is the number of pixels; K represents the number of categories; and respectively represent the true label of the one-hot encoding and the probability output of category k at pixel n.

[0039] In the technical solution of this application, the PET and CT images are respectively input into the nnUNet convolutional neural network for feature learning to obtain the multi-scale context information of the images. nnUNet extracts deep semantic features through the encoder, and the decoder gradually restores the detailed features. Apply K-means clustering to the target foreground features and divide them into multiple sub-regions, and each region corresponds to different parts of the target. The clustering process simplifies the complex feature representation into several regions by optimizing the feature similarity. Perform average pooling operation on the features of each region to compress the high-dimensional features into prototype vectors of a fixed dimension. These prototype vectors represent the global features of each region of the target. Calculate the cosine similarity between the prototype vectors and the second feature of the CT image to generate the correlation map of each prototype. Subsequently, average all the correlation maps to obtain the refined target hint information. Compared with the initial target hint information, the refined target hint information combines the high-quality information of the prototype vectors and the shallow features, enhances the distinguishability of the target region, and effectively suppresses the interference of background noise and irrelevant tissues at the same time. The final output features combine the multi-modal information of the anatomical structure, metabolic activity and hint information, and then use the segmentation head to perform convolutional processing on the final output features to output the segmentation result. Through the joint optimization of the Dice coefficient loss and the cross-entropy loss, the segmentation accuracy of the model for the target region and the overall pixel classification performance are improved, and accurate segmentation is achieved.

[0040] This application also verifies the above technical solution through experiments. The data set contains PET and CT data (both in DICOM format) from 174 users, of which the training set contains the data of 139 users and the validation set contains the data of 35 users. The sizes of the CT and PET data are both , where, Indicates the resolution of the cross-section (X and Y axes), and 256 indicates the number of slices in the depth (Z axis). All experiments were conducted in the Ubuntu 20.04.04LTS system environment, with the hardware configuration being a single NVIDIA A6000 graphics card. The software environment includes CUDA 11.3 and PyTorch 2.0. The segmentation task is based on the nnUNet open-source semantic segmentation framework. The optimizer selected is SGD, and the initial learning rate is set to 0.01. The experimental results are shown in Table 1 as follows:

[0041] Table 1

[0042] Table 1 compares the technical indicators of the technical solution of this application with various existing traditional technical solutions. The Dice coefficient is an index for measuring the similarity between two sets and is widely used in segmentation tasks to evaluate the overlap between the predicted result and the ground truth label. Its value ranges from 0 to 100, and the larger the value, the more similar the prediction is to the ground truth label. Recall measures the proportion of positive examples (i.e., lesions, regions of interest, etc.) correctly detected by the model among all true positive examples. The Jaccard index is used to measure the similarity between two sets and is usually used to evaluate the performance of segmentation tasks. Similar to the Dice coefficient, Jaccard also measures the overlap between the prediction and the ground truth label, but it has a higher penalty for false detections. Precision measures the proportion of the regions predicted as positive examples by the model that are truly positive examples, and it measures the accuracy of the prediction result.

[0043] The results in Table 1 show that the technical solution of this application performs better in all indicators (Dice, Recall, Jaccard, Precision), especially with obvious advantages in Precision (65.48) and Dice (63.23), indicating that the technical solution of this application can not only segment images well but also has a more accurate prediction result. The MedNext model performs better in terms of recall and Jaccard coefficient, but its accuracy is slightly lower than that of the model of this application. UX-Net and MAML are at a medium level in all indicators, showing a relatively balanced performance but not reaching the level of this application. The performance of the Unet model is relatively weak, especially in terms of the Dice coefficient (56.46) and Jaccard coefficient (42.69), indicating its deficiencies in segmentation accuracy and prediction accuracy.

[0044] The above-described embodiments merely represent several implementation manners of the present application. The description thereof is relatively specific and detailed, but it should not be construed as a limitation on the scope of the invention patent. It should be noted that for those of ordinary skill in the art, without departing from the concept of the present application, several modifications and improvements can still be made, and these all fall within the protection scope of the present application. Therefore, the protection scope of the patent of the present application shall be subject to the appended claims.

Claims

1. A PET-CT image fusion segmentation method based on prompt information, characterized in that: The PET-CT image fusion segmentation method based on prompt information includes: The CT image and the text description corresponding to the target are respectively input into the pre-trained PMC-CLIP model, and the initial target prompt information is generated by combining the extracted first feature of the CT image and the target text feature; Input the CT image into the nnUNet convolutional neural network, extract the second feature of the CT image, combine the second feature of the CT image with the initial target prompt information, and generate the target foreground feature; The target foreground features are clustered into regional features, and each regional feature is averagely pooled to obtain the target prototype vector corresponding to each regional feature; Calculate the cosine similarity between the target prototype vector corresponding to each regional feature and the second feature of the CT image, and average the results to obtain refined target prompt information; The PET image is input into the nnUNet convolutional neural network to extract the PET image features, refine the target prompt information, splice the PET image features and the second features of the CT image to generate the final output features, and then pass through the segmentation head to obtain the segmentation result.

2. The PET-CT image fusion segmentation method based on prompt information according to claim 1, characterized in that: The step of combining the extracted first feature of the CT image and the target text feature to generate initial target prompt information includes: Calculate the cosine similarity between the first feature of the CT image and the target text feature; The calculated cosine similarity is compared with the first threshold, and if it is greater than the first threshold, it is marked as a foreground to generate initial target prompt information.

3. The PET-CT image fusion segmentation method based on prompt information according to claim 1, characterized in that: The step of combining the second feature of the CT image with the initial target prompt information to generate a target foreground feature includes: The initial target prompt information is multiplied with the second feature of the CT image to obtain the target foreground feature.

4. The PET-CT image fusion segmentation method based on prompt information according to claim 1, characterized in that: The results are averaged to obtain refined target prompt information, and the formula is as follows: in, represents the cosine similarity between the target prototype vector and the second feature of the CT image, M is the total number of regional features, and the final Indicates detailed target prompt information.

5. The PET-CT image fusion segmentation method based on prompt information according to claim 1, characterized in that: The segmentation head adopts Convolutional layer.

6. The PET-CT image fusion segmentation method based on prompt information according to claim 1, characterized in that: The loss of the nnUNet convolutional neural network includes Dice coefficient loss and cross entropy loss.