Pseudo-label Selection Method, System, Device, and Medium for Computer Tomography Image Segmentation Based on Multimodal Model

The multimodal model BiomedCLIP filters pseudo-label quality problem in CT image segmentation is solved, and efficient pseudo-label quality judgment without manual intervention and high-precision medical image segmentation are achieved.

CN119888233BActive Publication Date: 2025-07-25JIANGSU OPEN UNIVERSITY (THE CITY VOCATIONAL COLLEGE OF JIANGSU)
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202510210509.7
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-02-25
Publication Date
2025-07-25
Estimated Expiration
2045-02-25

AI Technical Summary

Technical Problem

In the prior art, computed tomography image segmentation faces the problems of noise, artifacts and low contrast, resulting in a lack of labeling data and uneven pseudo-label quality, which affects the training effect of the model. Especially in CT images, noise and inconsistent pseudo-labels will reduce the model's ability to recognize medical lesion characteristics.

Method used

The multimodal model BiomedCLIP is used to construct a joint loss function through the combination of image encoder and text encoder, and use visual features and text feature similarity to judge the quality of pseudo-labels, and filter out high-quality pseudo-labels, including noise pseudo-label generation, degradation operations and prompt word training to realize explicit quality judgment of pseudo-labels.

Benefits of technology

Without adding manual annotation, high-quality training data is automatically generated to improve the accuracy of the model's medical image segmentation, avoiding the inefficiency caused by artificially setting thresholds and subjectivity, and directly filtering out high-quality pseudo-labels for training, improving the segmentation effect of the model.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119888233B_ABST
    Figure CN119888233B_ABST
Patent Text Reader

Abstract

The present invention discloses a method, system, device, and medium for selecting pseudo-labels for computer tomography image segmentation based on a multimodal model. The method includes: using Totalsegmentator to label unlabeled anatomical organ types in the original CT image dataset to obtain noisy pseudo-labels; degrading the manually labeled labels; forming a training set with the original CT images, manually labeled labels, and degraded manually labeled labels, and inputting it into the image encoder of BiomedCLIP to extract visual features; initializing the prompt words and inputting them into the text encoder of BiomedCLIP to extract text features; constructing a joint loss function to supervise and train the prompt words, superimposing the category information of the noisy pseudo-labels on the trained prompt words, and then inputting them into the text encoder to obtain text features, and inputting the original CT images and noisy pseudo-labels after transformation into the image encoder to extract visual features; judging the quality of the noisy pseudo-labels based on the similarity between the text features and visual features. The present invention improves the efficiency of screening high-quality pseudo-labels.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the fields of artificial intelligence and computer vision, and particularly relates to a method, system, device, and medium for selecting pseudo-labels for computer tomography image segmentation based on a multi-modal model. Background Art

[0002] Applying deep learning to medical imaging, especially using image segmentation technology to assist doctors in diagnosis, has become a current research hotspot. However, there are noise, artifacts, and low contrast in Computed Tomography (CT) images. Coupled with complex anatomical structures and a lack of labeled data, CT image segmentation faces many challenges.

[0003] Although a large amount of CT data is generated daily in major hospitals, due to the fact that medical images involve patient privacy and interpreting medical images requires professional knowledge and must be labeled by professional doctors, it is costly to obtain a large amount of accurately labeled medical images. To solve the problem of insufficient labeled data, pseudo-labels are widely used in semi-supervised learning and unsupervised learning. By using the prediction results generated by the model as the labels of the training data, the data scale is thus expanded. However, the uneven quality of pseudo-labels has a significant impact on the training effect of the model. Especially in CT images, noise and inconsistent pseudo-labels will reduce the model's ability to recognize medical lesion features. Therefore, efficiently screening out high-quality pseudo-labels has become a key problem to be solved urgently.

[0004] The development of artificial intelligence technology has triggered a profound transformation in the field of medical image processing. Automated methods can reduce the subjective differences among observers and achieve the quantification of anatomical features. The literature ("Multi-modal contrastive mutual learning and pseudo-label re-learning for semi-supervised medical image segmentation", Medical Image Analysis, 2023.) proposed a method of multi-modal contrastive mutual learning and pseudo-label re-learning for semi-supervised medical image segmentation. In another study, the literature ("Pseudo-label guided contrastive learning for semi-supervised medical image segmentation", Proceedings of the IEEE / CVF Conference on Computer Vision and Pattern Recognition, 2023.) proposed a pseudo-label guided contrastive learning framework for semi-supervised medical image segmentation. In addition, the literature ("Semi-supervised medical image segmentation via cross teaching between CNN and Transformer", Proceedings of the IEEE / CVF International Conference on Computer Vision, 2021.) proposed a cross-learning method between CNN and Transformer for semi-supervised medical image segmentation. However, these studies mainly focus on the fusion of multi-modal data and the generation of pseudo-labels, and have not fully solved the problem of uneven quality of pseudo-labels. It can be seen from the above work that there are still few studies on the selection of pseudo-labels using multi-modal models at present. Summary of the Invention

[0005] Aiming at the deficiencies in the prior art, the present invention provides a method, system, device and medium for pseudo-label selection in computer tomography image segmentation based on a multi-modal model to explicitly measure the quality of labels, so as to expand as many reliable labels as possible without introducing new training data.

[0006] To achieve the above object, the present invention adopts the following technical solutions:

[0007] A pseudo label selection method for computed tomography image segmentation based on a multimodal model comprises the following steps:

[0008] S1, use the original CT image dataset with manually annotated labels to train the segmentation model Totalsegmentator;

[0009] S2. Use the trained segmentation model Totalsegmentator to label the unlabeled anatomical organ types in the original CT image dataset. The labels labeled by the segmentation model Totalsegmentator are noise pseudo labels.

[0010] S3, degenerating the manually annotated labels to obtain degraded manually annotated labels;

[0011] S4, preparing a pre-trained multimodal model BiomedCLIP, wherein the multimodal model BiomedCLIP comprises an image encoder and a text encoder, and the original CT image, the manually annotated labels and the degraded manually annotated labels form a training data set, and the training data set is transformed and input into the image encoder of the multimodal model BiomedCLIP to extract visual features;

[0012] S5, construct a learnable prompt word, initialize the prompt word and input it into the text encoder of the multimodal model BiomedCLIP to extract text features;

[0013] S6. Construct a joint loss function using visual features and text features to supervise and train learnable prompt words, with the training goal of minimizing the joint loss. After the training is completed, freeze the prompt word parameters and the parameters of the multimodal model BiomedCLIP.

[0014] S7, superimposing the category information of the noise pseudo-label generated by the segmentation model Totalsegmentator onto the trained prompt word, and then inputting it into the text encoder in the multimodal model BiomedCLIP to obtain text features, transforming the original CT image and its corresponding noise pseudo-label and inputting it into the image encoder in the multimodal model BiomedCLIP to extract visual features;

[0015] S8. Calculate the similarity between the text features and the visual features obtained in S7, and judge the quality of the noise pseudo-label based on the similarity between the text features and the visual features.

[0016] To optimize the above technical solutions, the specific measures taken also include:

[0017] Furthermore, in S3, the step of degrading the manually annotated labels is as follows:

[0018] Randomly select multiple degradation operations from the following degradation operations to degrade the manually annotated labels: adding random noise, Gaussian blur, randomly moving pixel positions, random affine transformation, non-linear deformation, simulating missing regions, adding random shapes, and random morphological transformation.

[0019] Further, in S4, the training dataset is expressed as:

[0020] ;

[0021] In the formula, D T is the training dataset, is the i-th CT image, is the manually annotated label corresponding to the i-th CT image, is the degraded manually annotated label corresponding to the i-th CT image, and N is the number of samples in the training dataset;

[0022] The transformation of the training dataset is specifically:

[0023] Take corresponding or First perform maximum projection to obtain Subsequently, corresponding and Convert from 3D to 2D images, and the obtained 2D images are uniformly denoted as Then, perform edge extraction operation on to obtain edge information Take , and Stack them together to form three-channel data;

[0024] The image encoder of the multimodal model BiomedCLIP is expressed as the following formula:

[0025] ;

[0026] In the formula, is the mapping function of the image encoder, is the input image set, represents the D-dimensional feature space.

[0027] Further, in S5, the text encoder of the multimodal model BiomedCLIP is expressed as the following formula:

[0028] ;

[0029] In the formula, is the mapping function of the text encoder, is the input text set, representing the D-dimensional feature space.

[0030] Furthermore, in S6, the training objective is to minimize the joint loss, which is expressed by the formula:

[0031]

[0032] where, is the joint loss function, is the cross-entropy loss, is the weight of the contrastive loss, is the contrastive loss;

[0033] The calculation formula of the cross-entropy loss is as follows:

[0034] ;

[0035] where, N is the number of samples in the training dataset, is the indicator function, is the i-th CT image, is the manually annotated label corresponding to the i-th CT image, is the degraded manually annotated label corresponding to the i-th CT image, represents the probability that the label of the i-th CT image is the manually annotated label, represents the probability that the label of the i-th CT image is the degraded manually annotated label;

[0036] The calculation formula of the contrastive loss is as follows:

[0037] ;

[0038] where, is the visual feature of the i-th manually annotated label, is the text feature of the i-th manually annotated label, and m represents the lower limit of the distance between the visual feature of the manually annotated label and the text feature of the degraded manually annotated label; represents the text feature of the i-th degraded manually annotated label.

[0039] Furthermore, in S8, the similarity between the text feature and the visual feature is specifically measured by the cosine similarity.

[0040] The present invention also proposes a pseudo-label selection system for computer tomography image segmentation based on a multi-modal model, including:

[0041] A noise pseudo-label generation module, which is used to label the unlabeled anatomical organ types in the original CT image dataset using the trained segmentation large model Totalsegmentator to generate noise pseudo-labels;

[0042] A degradation module, which is used to degrade the manually labeled labels to obtain degraded manually labeled labels;

[0043] A multimodal model BiomedCLIP, where the multimodal model BiomedCLIP includes an image encoder and a text encoder;

[0044] A multi-view construction module, which is used to form a training dataset with the original CT image, the manually labeled labels, and the degraded manually labeled labels, and after transforming the training dataset, input it into the image encoder of the multimodal model BiomedCLIP;

[0045] The image encoder is used to extract visual features; the text encoder is used to extract text features;

[0046] A prompt training module, which is used to initialize the prompt and then input it into the text encoder of the multimodal model BiomedCLIP to extract text features; construct a joint loss function using the visual features and the text features to supervise and train the learnable prompt, and the training objective is to minimize the joint loss; after the training is completed, freeze the prompt parameters and the parameters of the multimodal model BiomedCLIP,

[0047] A noise pseudo-label quality judgment module, which is used to superimpose the category information of the noise pseudo-labels generated by the segmentation large model Totalsegmentator on the trained prompt, and then input it into the text encoder in the multimodal model BiomedCLIP to obtain text features, and transform the original CT image and its corresponding noise pseudo-labels and input them into the image encoder of the multimodal model BiomedCLIP to extract visual features; judge the quality of the noise pseudo-labels based on the similarity between the text features and the visual features.

[0048] The present invention also proposes an electronic device, including: a memory, a processor, and a computer program stored on the memory and executable on the processor. When the processor executes the computer program, it implements the above-mentioned pseudo-label selection method for computer tomography image segmentation based on a multimodal model.

[0049] The present invention also proposes a computer-readable storage medium storing a computer program, and the computer program causes a computer to execute the above-mentioned pseudo-label selection method for computer tomography image segmentation based on a multimodal model.

[0050] The beneficial effects of the present invention are:

[0051] The present invention can explicitly judge a large number of pseudo - labels, directly screen out high - quality pseudo - labels without any manual intervention, and avoid problems such as low efficiency and high subjectivity caused by artificially setting thresholds or judging label quality by human eyes. The high - quality labels after screening are used to train the segmentation model of computed tomography (CT) images, and the trained segmentation model of CT images is used to segment the organs in CT images. This method can automatically generate high - quality training data to guide the segmentation model to generate high - precision organ segmentation results without increasing manual annotation. BRIEF DESCRIPTION OF THE DRAWINGS

[0052] Figure 1 It is the overall framework diagram of the pseudo - label selection method for CT image segmentation based on a multi - modal model proposed by the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0053] Next, the technical solutions in the embodiments of the present application will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present application. Obviously, the described embodiments are only a part of the embodiments of the present application, rather than all the embodiments. Based on the embodiments in the present application, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of the present application.

[0054] Embodiment 1

[0055] The present invention proposes a pseudo - label selection method for CT image segmentation based on a multi - modal model. The overall framework of this method is as Figure 1 shown and includes the following steps:

[0056] S1. Use the original CT image dataset containing manually annotated labels to train the segmentation large model Totalsegmentator; Totalsegmentator can segment 117 organ categories. The original CT image dataset contains N types of manually annotated labels, denoted by . The manually annotated labels are also called high - quality labels in this embodiment.

[0057] S2. Expand the dataset used for the downstream task. Use the trained segmentation large model Totalsegmentator to annotate the unannotated anatomical organ types in the original CT image dataset. The labels annotated by the segmentation large model Totalsegmentator are noise pseudo - labels; expand the N types of annotated data contained in the original data to M types of annotated data; at this time, in addition to the N types of high - quality annotated data manually annotated in the original CT image dataset, there are also M - N types of noise pseudo - label data generated by Totalsegmentator, denoted by Represent;

[0058] S3. Degrade the manually annotated labels to obtain degraded manually annotated labels. This part of the data is represented by ; The degraded manually annotated labels are also called low-quality labels in this embodiment. Specifically, S3 is as follows:

[0059] Design a degradation module. The degradation module includes three categories of transformations, namely: 1) noise and interference, 2) geometric transformation, and 3) morphological change. The specific transformations included in the three categories are as follows:

[0060] 1) Noise and interference

[0061] Add random noise: Simulate the degradation of the signal by adding random noise.

[0062] Gaussian blur: Reduce the image details through blurring to simulate the situation of reduced image resolution or out-of-focus.

[0063] 2) Geometric transformation

[0064] Random displacement: Randomly move the pixel positions in the label image to simulate image registration errors.

[0065] Random affine transformation: Perform random affine transformations such as scaling, rotation, and translation on the label image to simulate geometric distortions during label annotation.

[0066] Elastic deformation: Perform non-linear deformation on the label image to simulate the natural changes of tissue structures.

[0067] 3) Morphological processing

[0068] Simulate missing areas: Randomly delete some areas in the label image to simulate the situation where some parts of the label are occluded or lost.

[0069] Add random shapes: Randomly add some irrelevant shapes in the label image to simulate mislabeling or interference in the label.

[0070] Random morphological transformation: Perform random morphological operations such as dilation and erosion on the label image to simulate irregular changes in the image edges.

[0071] For each manually annotated label, at least 2 - 5 of these transformations will be selected, and the selection is completely random. After a total of three rounds of manual degradation for all the data, is obtained. .

[0072] S4. Prepare the pre-trained multimodal model BiomedCLIP. The multimodal model BiomedCLIP includes an image encoder and a text encoder. Compose a training dataset from the original CT images, manually annotated labels, and degraded manually annotated labels. The training dataset is expressed as:

[0073] ;

[0074] where D T is the training dataset, is the i-th CT image, is the manually annotated label corresponding to the i-th CT image, is the degraded manually annotated label corresponding to the i-th CT image, and N is the number of samples in the training dataset.

[0075] After transforming the training dataset, input it into the image encoder of the multimodal model BiomedCLIP to extract visual features. The specific process of transforming the training dataset is as follows: For the corresponding or , first perform maximum projection to obtain . Subsequently, for the corresponding and , convert from 3D to 2D images. The obtained 2D images are uniformly denoted as . Then, perform edge extraction operation on to obtain edge information . Stack , and together to form a three-channel data.

[0076] The image encoder of the multimodal model BiomedCLIP is expressed as the following formula:

[0077] ;

[0078] where is the mapping function of the image encoder, is the input image set, represents a D-dimensional feature space.

[0079] The text encoder of the multimodal model BiomedCLIP is expressed as the following formula:

[0080] ;

[0081] where is the mapping function of the text encoder, is the input text set, Represents a D-dimensional feature space.

[0082] S5. Construct a learnable prompt, where the prompt is set to 128 tokens. A token is the smallest unit after splitting text in the field of natural language processing, which can be a word, a sub-word, or a special symbol. After initializing the prompt, input it into the text encoder of the multimodal model BiomedCLIP to extract text features;

[0083] S6. Use the visual features and text features to construct a joint loss function to supervise and train the learnable prompt. After several rounds of training on the learnable prompt, the prompt at this time can effectively represent high-quality labels and low-quality labels. The training objective is to minimize the joint loss; expressed by the formula:

[0084]

[0085] In the formula, is the joint loss function, is the cross-entropy loss, is the weight of the contrastive loss, is the contrastive loss;

[0086] The cross-entropy loss is used to ensure that the features output by the image encoder are aligned with the features (labels) output by the correct text encoder. The calculation formula of the cross-entropy loss is as follows:

[0087] ;

[0088] In the formula, N is the number of samples in the training dataset, is the indicator function, is the i-th CT image, is the manually annotated label corresponding to the i-th CT image, is the degraded manually annotated label corresponding to the i-th CT image, represents the probability that the label of the i-th CT image is the manually annotated label, represents the probability that the label of the i-th CT image is the degraded manually annotated label;

[0089] To enhance the ability of the prompt to judge high-quality labels and low-quality labels, ensure that the image features of high-quality labels are close to their corresponding text features, and at the same time keep a distance from low-quality text features, a contrastive loss is designed to further supervise the training of the learnable prompt. The calculation formula of the contrastive loss is as follows:

[0090] ;

[0091] In the formula, is the visual feature of the i-th manually annotated label, is the text feature of the i-th manually annotated label, and m represents the lower limit of the distance between the visual feature of the manually annotated label and the text feature of the degraded manually annotated label; represents the text feature of the i-th degraded manually annotated label.

[0092] After the training is completed, freeze the prompt word parameters and the parameters of the multimodal model BiomedCLIP.

[0093] S7. The noise pseudo-labels generated by the segmentation large model Totalsegmentator actually have category information. To further improve the robustness of the model to judge labels, the category information of the noise pseudo-labels generated by the segmentation large model Totalsegmentator is superimposed on the trained prompt words, and then input into the text encoder in the multimodal model BiomedCLIP to obtain text features. The original CT image and its corresponding noise pseudo-labels are transformed and then input into the image encoder of the multimodal model BiomedCLIP to extract visual features;

[0094] S8. Calculate the similarity between the text features and visual features obtained in S7, and judge the quality of the noise pseudo-labels based on the similarity between the text features and visual features. In this embodiment, the similarity between the text features and visual features is measured by cosine similarity, and the expression of cosine similarity is as follows:

[0095] ;

[0096] In the formula, represents the cosine similarity between the text feature and the visual feature, and are the input image and text, represents the visual feature, represents the text feature.

[0097] Embodiment 2

[0098] The present invention proposes a pseudo-label selection system for computer tomography image segmentation based on a multimodal model corresponding to the method of Embodiment 1, including:

[0099] A noise pseudo-label generation module, which is used to label the unannotated anatomical organ types in the original CT image dataset using the trained segmentation large model Totalsegmentator to generate noise pseudo-labels;

[0100] A degradation module, which is used to degrade the manually annotated labels to obtain degraded manually annotated labels;

[0101] The multimodal model BiomedCLIP, where the multimodal model BiomedCLIP includes an image encoder and a text encoder;

[0102] A multi-view construction module for forming a training data set from the original CT image, the manually annotated label, and the degraded manually annotated label, transforming the training data set and inputting it into the image encoder of the multimodal model BiomedCLIP;

[0103] The image encoder is used to extract visual features; the text encoder is used to extract text features;

[0104] A prompt training module for initializing the prompt and inputting it into the text encoder of the multimodal model BiomedCLIP to extract text features; constructing a joint loss function using the visual features and text features to supervise and train the learnable prompt, with the training objective of minimizing the joint loss; after training is completed, freezing the prompt parameters and the parameters of the multimodal model BiomedCLIP,

[0105] A noise pseudo-label quality judgment module for superimposing the class information of the noise pseudo-labels generated by the segmentation large model Totalsegmentator on the trained prompt, then inputting it into the text encoder in the multimodal model BiomedCLIP to obtain text features, transforming the original CT image and its corresponding noise pseudo-labels and inputting them into the image encoder of the multimodal model BiomedCLIP to extract visual features; judging the quality of the noise pseudo-labels based on the similarity between the text features and visual features.

[0106] The implementation methods of each module and module functions in the system are exactly the same as the steps of the method in Embodiment 1, so they will not be elaborated here.

[0107] Embodiment 3

[0108] The present invention proposes an electronic device, including: a memory, a processor, and a computer program stored on the memory and executable on the processor. When the processor executes the computer program, it implements the pseudo-label selection method for computer tomography image segmentation based on a multimodal model as described in Embodiment 1.

[0109] Embodiment 4

[0110] The present invention proposes a computer-readable storage medium storing a computer program, and the computer program causes a computer to execute the pseudo-label selection method for computer tomography image segmentation based on a multimodal model as described in Embodiment 1.

[0111] The present invention can explicitly judge a large number of pseudo-labels, directly screen out high-quality pseudo-labels without any manual intervention, and avoid problems such as low efficiency and high subjectivity caused by artificially setting thresholds or judging label quality by human eyes. The high-quality labels after screening are used to train the segmentation model of computed tomography images, and the trained segmentation model of computed tomography images is used to segment the organs in computed tomography images. The present invention can automatically generate high-quality training data to guide the segmentation model to generate high-precision organ segmentation results without increasing manual annotation.

[0112] In the embodiments disclosed in the present application, the computer storage medium can be a tangible medium that can contain or store a program for use by or in conjunction with an instruction execution system, apparatus, or device. The computer storage medium can include, but is not limited to, electronic, magnetic, optical, electromagnetic, infrared, or semiconductor systems, apparatus, or devices, or any suitable combination of the foregoing. More specific examples of the computer storage medium would include electrical connections based on one or more wires, portable computer disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fibers, portable compact disc read-only memory (CD-ROM), optical storage devices, magnetic storage devices, or any suitable combination of the foregoing.

[0113] Those of ordinary skill in the art can realize that the units and algorithm steps of each example described in combination with the embodiments disclosed in the present application can be implemented by electronic hardware or a combination of computer software and electronic hardware. Whether these functions are executed in a hardware or software manner depends on the specific application and design constraints of the technical solution. Professional technicians can use different methods to implement the described functions for each specific application, but such implementation should not be considered to exceed the scope of the present application.

[0114] The above are only the preferred embodiments of the present invention, and the protection scope of the present invention is not limited to the above embodiments. All technical solutions falling within the concept of the present invention belong to the protection scope of the present invention. It should be noted that for those of ordinary skill in the art in the technical field, several improvements and refinements made without departing from the principle of the present invention should be regarded as within the protection scope of the present invention.

Claims

1. A method for selecting pseudo-labels for computer tomography image segmentation based on a multi-modal model, characterized in that, The following steps are involved: S1, use the original CT image dataset with manually annotated labels to train the segmentation model Totalsegmentator; S2. Use the trained segmentation model Totalsegmentator to label the unlabeled anatomical organ types in the original CT image dataset. The labels labeled by the segmentation model Totalsegmentator are noise pseudo labels. S3, degenerating the manually annotated labels to obtain degraded manually annotated labels; S4, preparing a pre-trained multimodal model BiomedCLIP, wherein the multimodal model BiomedCLIP comprises an image encoder and a text encoder, and the original CT image, the manually annotated labels and the degraded manually annotated labels form a training data set, and the training data set is transformed and input into the image encoder of the multimodal model BiomedCLIP to extract visual features; S5, construct a learnable prompt word, initialize the prompt word and input it into the text encoder of the multimodal model BiomedCLIP to extract text features; S6. Construct a joint loss function using visual features and text features to supervise and train learnable prompt words, with the training goal of minimizing the joint loss. After the training is completed, freeze the prompt word parameters and the parameters of the multimodal model BiomedCLIP. S7, superimposing the category information of the noise pseudo-label generated by the segmentation model Totalsegmentator onto the trained prompt word, and then inputting it into the text encoder in the multimodal model BiomedCLIP to obtain text features, transforming the original CT image and its corresponding noise pseudo-label and inputting it into the image encoder in the multimodal model BiomedCLIP to extract visual features; S8. Calculate the similarity between the text features and the visual features obtained in S7, and judge the quality of the noise pseudo-label based on the similarity between the text features and the visual features.

2. The pseudo-label selection method for computer tomography image segmentation based on a multi-modal model according to claim 1, wherein, In S3, the degradation of the manually annotated labels is specifically as follows: Multiple degradation operations are randomly selected from the following degradation operations to degrade the manually annotated labels: adding random noise, Gaussian blur, randomly moving pixel positions, random affine transformation, nonlinear deformation, simulating missing areas, adding random shapes and random morphological transformation.

3. The pseudo-label selection method for computer tomography image segmentation based on a multi-modal model according to claim 1, wherein, In S4, the training data set is represented as: ; where D T is the training data set, is the i-th CT image, is the manually annotated label corresponding to the i-th CT image, is the degraded manually annotated label corresponding to the i-th CT image, and N is the number of samples in the training data set; The transformation of the training data set is specifically as follows: Convert corresponding or First, perform maximum projection to obtain , then corresponding and Convert from 3D to 2D image, and the resulting 2D images are uniformly denoted as , then perform edge extraction operation on to obtain edge information Convert , and stack them together to form three-channel data; The image encoder of the multimodal model BiomedCLIP is expressed as follows: ; In the formula, is the mapping function of the image encoder, is the input image set, represents the D-dimensional feature space.

4. The pseudo-label selection method for computer tomography image segmentation based on a multi-modal model according to claim 1, wherein In S5, the text encoder of the multimodal model BiomedCLIP is expressed as follows: ; wherein, is the mapping function of the text encoder, is the input text set, represents a D-dimensional feature space.

5. The pseudo-label selection method for computer tomography image segmentation based on a multi-modal model according to claim 1, wherein, In S6, the training objective is to minimize the joint loss and is expressed as: ; In the formula, is the combined loss function, is the cross-entropy loss, is the weight of the contrast loss, is the contrast loss; The calculation formula for cross entropy loss is as follows: ; where N is the number of samples in the training dataset, is the indicator function, is the i-th CT image, is the manually annotated label corresponding to the i-th CT image, is the degraded manually annotated label corresponding to the i-th CT image, represents the probability that the label of the i-th CT image is the manually annotated label, represents the probability that the label of the i-th CT image is the degraded manually annotated label; The calculation formula of contrast loss is as follows: ; Wherein, is the visual feature of the i-th manually labeled tag, is the text feature of the i-th manually labeled tag, and m represents the lower limit of the distance between the visual feature of the manually labeled tag and the text feature of the degraded manually labeled tag; represents the text feature of the i-th degraded manually labeled tag.

6. The pseudo-label selection method for computer tomography image segmentation based on a multi-modal model according to claim 1, wherein, In S8, the similarity between the text feature and the visual feature is specifically measured using cosine similarity.

7. A pseudo-label selection system for computer tomography image segmentation based on a multi-modal model, characterized in that, include: The noise pseudo-label generation module is used to use the trained segmentation model Totalsegmentator to annotate the unlabeled anatomical organ types in the original CT image dataset and generate noise pseudo-labels; A degradation module for degrading the manually annotated labels to obtain degraded manually annotated labels; A multimodal model BiomedCLIP, which includes an image encoder and a text encoder; A multi-view construction module for forming a training dataset with the original CT image, the manually annotated labels, and the degraded manually annotated labels, and inputting the transformed training dataset into the image encoder of the multimodal model BiomedCLIP; The image encoder is used to extract visual features; The text encoder is used to extract text features; A prompt training module for initializing the prompt and inputting it into the text encoder of the multimodal model BiomedCLIP to extract text features; constructing a joint loss function using the visual features and the text features, supervising and training the learnable prompt, and the training objective is to minimize the joint loss; after the training is completed, freeze the prompt parameters and the parameters of the multimodal model BiomedCLIP, A noise pseudo-label quality judgment module for superimposing the class information of the noise pseudo-labels generated by the segmentation large model Totalsegmentator on the trained prompt, and then inputting it into the text encoder in the multimodal model BiomedCLIP to obtain text features, and inputting the original CT image and its corresponding noise pseudo-labels after transformation into the image encoder of the multimodal model BiomedCLIP to extract visual features; Judge the quality of the noise pseudo-labels based on the similarity between the text features and the visual features.

8. An electronic device, characterized in that, Including: A memory, a processor, and a computer program stored on the memory and executable on the processor. When the processor executes the computer program, it implements the pseudo-label selection method for computer tomography image segmentation based on a multimodal model according to any one of claims 1-6.

9. A computer-readable storage medium storing a computer program, characterized in that, The computer program causes the computer to execute the pseudo-label selection method for computer tomography image segmentation based on a multimodal model according to any one of claims 1-6.

Citation Information

Patent Citations

  • Medical image segmentation method

    WO2024098318A1