Image segmentation model generation method and device, image segmentation method and device and electronic equipment
By constructing a model training method for image representation, text representation and feature fine alignment module, the problem of inconsistent alignment of medical images and text information is solved, and the accuracy and efficiency of medical image segmentation are improved.
Patent Information
- Application Number
- CN202510102850.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-22
- Publication Date
- 2025-06-03
- Estimated Expiration
- 2045-01-22
AI Technical Summary
In the prior art, text-guided medical image segmentation methods are difficult to effectively learn joint embedding representations and accurately establish consistency between images and text, resulting in poor segmentation performance.
Build an initial model, including image representation module, text representation module, feature fine alignment module and mask segmentation module, train the model to align medical images with text information, use the loss function to optimize model parameters, and improve segmentation accuracy.
By enhancing the alignment of medical images with text information, the accuracy and efficiency of image segmentation are improved, and the problem of poor segmentation performance is effectively solved.
Smart Images

Figure CN120088795A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical fields of medical image processing and computer vision, and particularly to the generation of an image segmentation model, and an image segmentation method, device and electronic device. Background Art
[0002] Medical image segmentation is an important task in medical image analysis. Text-guided medical image segmentation aims to simultaneously utilize medical images and corresponding medical texts to segment medical images, so as to assist doctors in diagnosis and treatment.
[0003] Compared with single-modal image segmentation, text-guided medical image segmentation can make full use of medical texts to make up for the quality defects of medical images and does not require additional costs to obtain. Medical texts are usually generated together with medical images and have natural complementarity. Medical texts can provide the quantity and spatial information of feature regions, so that feature regions can be segmented more accurately. This is particularly important for clinical diagnosis and medical research, such as polyp segmentation and ultrasound image segmentation.
[0004] Due to the huge differences between different modalities, the main problem of text-guided medical image segmentation is how to learn a joint embedding representation and accurately establish the consistency between images and texts. Many works have proposed some deep learning segmentation methods, but most of them fail to effectively capture the discriminant regions of images, and the rough alignment of images and texts will affect the utilization of text information and ultimately affect the segmentation performance. Therefore, how to effectively align medical images with corresponding medical texts has become a very important task in text-guided image segmentation. Summary of the Invention
[0005] In view of this, it is necessary to provide a method for generating an image segmentation model, an image segmentation method, device and electronic device to solve the technical problem of poor image segmentation performance in the prior art.
[0006] To solve the above technical problem, on the one hand, the present invention provides a method for generating an image segmentation model, including: Construct an initial model, the structure of the initial model sequentially includes: an image representation module for converting a medical image into a first feature vector, a text representation module for converting a medical text into a second feature vector, a feature fine alignment module for aligning based on the first feature vector and the second feature vector to obtain a text-enhanced feature, and a mask segmentation module for obtaining a prediction result based on the text-enhanced feature and calculating the loss function value of the prediction result; Train the initial model based on the obtained medical image data and medical text data to obtain a trained and complete image segmentation model.
[0007] In a possible implementation, the image representation module includes: a plurality of downsampling convolutional modules; Wherein, the downsampling convolutional module includes: a max pooling layer, a convolutional layer, a normalization module, and an activation layer; The text representation module includes: a linear layer, an activation layer, and a BERT model pre-trained based on the medical text data and the medical vocabulary, wherein the medical vocabulary includes: medical terms and corresponding index values; The feature fine alignment module includes: a scale attention module and a spatial attention module both composed of an average pooling layer and a Sigmoid function; The mask segmentation module includes: an upsampling convolutional module and a preset loss function.
[0008] In a possible implementation, training the initial model based on the obtained medical image data and medical text data includes: Performing downsampling convolution on the medical image data based on the image representation module to obtain a first feature vector including a plurality of scales; Converting the medical text data into a second feature vector based on the text representation module; Aligning the first feature vector and the second feature vector on the scale level and the spatial level in sequence based on the scale attention module and the spatial attention module to obtain the text enhanced feature; Performing prediction based on the mask segmentation module and the text enhanced feature to obtain a prediction result, and calculating the loss function value of the prediction result; Adjusting the model parameters based on the loss function value until the loss function value is less than a preset threshold or reaches a preset maximum number of iterations.
[0009] In a possible implementation, before converting the medical text data into a second feature vector based on the text representation module, it further includes: Performing word segmentation on the medical text data to obtain a sub-word sequence; Adding preset markers before and after the sub-word sequence; Converting each sub-word of the sub-word sequence into an ID value based on the medical vocabulary to obtain a Token text sequence, wherein if the sub-word appears in the medical vocabulary, it is converted into the corresponding index value, otherwise, it is converted into the marker; Generating position encoding and attention mask for the Token text sequence.
[0010] In a possible implementation, converting the medical text data into a second feature vector based on the text representation module includes: Input the ID value, positional encoding, and attention mask into the BERT model for forward propagation to obtain the hidden states of each Token text sequence, and organize to obtain the output features; Input the output features into the linear layer and activation layer in sequence to obtain the second feature vector.
[0011] In a possible implementation manner, the feature fine alignment module aligns the first feature vector and the second feature vector at the scale level and the spatial level to obtain the text enhancement feature, including: Calculate the first similarity between the first feature vector and the second feature vector based on the average pooling layer; Generate the first attention weight based on the Sigmoid function and the first similarity, and generate the third feature vector enhanced at the scale level based on the first attention weight and the first feature vector; Calculate the second similarity between the third feature vector and the second feature vector based on the average pooling layer; Generate the second attention weight based on the Sigmoid function and the second similarity, and generate the text enhancement feature enhanced at the scale level and the spatial level based on the second attention weight and the first feature vector.
[0012] In a possible implementation manner, adjusting the model parameters based on the loss function value includes: Perform backpropagation based on the loss function value, and optimize the connection weights inside the initial model based on a preset optimizer and the gradient.
[0013] The present invention also provides an image segmentation method, including: Obtain an image segmentation model based on the generation method of the image segmentation model according to any one of the above method items; Input the medical image to be segmented and the medical text description corresponding to the medical image into the image segmentation model; Obtain the segmented masked medical image.
[0014] The present invention also provides a generation device for an image segmentation model, including: A model construction module for constructing an initial model, the structure of the initial model sequentially includes: an image representation module for converting a medical image into a first feature vector, a text representation module for converting a medical text into a second feature vector, a feature fine alignment module for aligning based on the first feature vector and the second feature vector to obtain a text enhancement feature, and a mask segmentation module for obtaining a prediction result based on the text enhancement feature and calculating the loss function value of the prediction result; A training optimization module for training an initial model based on the acquired medical image data and medical text data to obtain a fully trained image segmentation model.
[0015] The present invention also provides an electronic device, including: A memory for storing programs; A processor coupled to the memory for executing the program stored in the memory to implement the steps in the generation method or image segmentation method of the image segmentation model described in any one of the above method items.
[0016] The beneficial effects of the present invention are as follows: The present invention provides a method for generating an image segmentation model. First, by designing a new text information fusion module to learn important feature information in medical images, and at the same time using a feature fine alignment module to capture the detailed correspondence between text descriptions and medical images, it enables better alignment of semantic representations of different modalities. By introducing text information to enhance the features captured by medical images, missing detailed information is supplemented in terms of spatial position and feature scale size, effectively enhancing the accuracy of image segmentation. On this basis, by introducing a loss function, the similarity between the segmentation result and the true label is maintained, the detailed features of the medical image and the semantic content of the text information are retained, and the differences between different modalities are eliminated. Finally, the efficiency and accuracy of image segmentation are improved, effectively solving the technical problem of poor performance of image segmentation in the prior art. BRIEF DESCRIPTION OF THE DRAWINGS In order to more clearly illustrate the technical solutions in the embodiments of the present invention, the following will briefly introduce the drawings required for the description of the embodiments. Obviously, the following drawings are only some embodiments of the present invention. For those skilled in the art, without creative efforts, other drawings can be obtained based on these drawings.
[0017] Figure 1 It is a schematic flowchart of an embodiment of the method for generating an image segmentation model provided by the present invention; Figure 2 For the present invention Figure 1 It is a schematic flowchart of an embodiment of S102 in the present invention; Figure 3 For the present invention Figure 2 It is a schematic flowchart of an embodiment of other steps before S202 in the present invention; Figure 4 For the present invention Figure 2 It is a schematic flowchart of an embodiment of S202 in the present invention; Figure 5 For the present invention Figure 2 It is a schematic flowchart of an embodiment of S203 in the present invention; Figure 6Schematic flowchart of an embodiment of the image segmentation method provided by the present invention; Figure 7 Schematic structural diagram of an embodiment of the apparatus for generating an image segmentation model provided by the present invention; Figure 8 Schematic structural diagram of an embodiment of an electronic device provided by the present invention. Detailed implementation manners
[0018] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative efforts shall fall within the protection scope of the present invention.
[0019] In the description of the embodiments of the present invention, unless otherwise specified, the meaning of "a plurality of" is two or more. "And / or" describes the association relationship of associated objects, indicating that three relationships may exist. For example, A and / or B may represent: A exists alone, A and B exist simultaneously, and B exists alone.
[0020] The descriptions such as "first" and "second" involved in the embodiments of the present invention are only for descriptive purposes, and cannot be understood as indicating or implying their relative importance or implicitly indicating the quantity of the indicated technical features. Therefore, the technical features defined with "first" and "second" may explicitly or implicitly include at least one such feature.
[0021] Referring to "embodiment" in this article means that the specific features, structures or characteristics described in connection with the embodiment may be included in at least one embodiment of the present invention. The phrase appears in various positions in the specification does not necessarily refer to the same embodiment, nor is it an independent or alternative embodiment mutually exclusive with other embodiments. Those skilled in the art explicitly and implicitly understand that the embodiments described herein may be combined with other embodiments.
[0022] The present invention provides a method, apparatus and electronic device for generating an image segmentation model and image segmentation, which will be described separately below.
[0023] Figure 1 Schematic flowchart of an embodiment of the method for generating an image segmentation model provided by the present invention, as Figure 1 shown, the method for generating an image segmentation model includes: Step S101, constructing an initial model; Specifically, the structure of the initial model sequentially includes: an image representation module for converting a medical image into a first feature vector, a text representation module for converting a medical text into a second feature vector, a feature fine alignment module for aligning based on the first feature vector and the second feature vector to obtain text-enhanced features, and a mask segmentation module for obtaining a prediction result based on the text-enhanced features and calculating the loss function value of the prediction result; Step S102: Train the initial model based on the obtained medical image data and medical text data to obtain a trained and complete image segmentation model.
[0024] In a possible implementation, the image representation module includes: a plurality of downsampling convolutional modules; It should be noted that the downsampling convolutional module is used to sample the input medical image, and the image features with the same center and different scales obtained by sampling are used as the multi-scale feature representation of the image: F v = F v 0 , F v 1 , F v 2 , F v 3 ; Preferably, the downsampling convolutional module is implemented by a max pooling layer, a convolutional layer, a normalization module, and an activation layer.
[0025] The text representation module includes: a linear layer, an activation layer, and a BERT model pre-trained based on medical text data and a medical vocabulary, where the medical vocabulary includes: medical terms and corresponding index values; The feature fine alignment module includes: a scale attention module and a spatial attention module both composed of an average pooling layer and a Sigmoid function; The mask segmentation module includes: an upsampling convolutional module and a preset loss function.
[0026] Such as Figure 2 , in a possible implementation, step S102 includes: Step S201: Perform downsampling convolution on the medical image data based on the image representation module to obtain a first feature vector including several scales; Step S202: Convert the medical text data into a second feature vector based on the text representation module; Step S203: Align the first feature vector and the second feature vector successively at the scale level and the spatial level based on the scale attention module and the spatial attention module to obtain text enhanced features; Step S204: Make a prediction based on the mask segmentation module and the text enhanced features to obtain a prediction result, and calculate the loss function value of the prediction result; Specifically, calculate the total loss function value of the model L including the Focal loss function L f and the Dice loss function L d and the similarity loss function L c in three parts: (1) Focal loss function L l , aiming to solve the class imbalance problem, by giving higher weights to difficult-to-classify samples, thereby improving the performance of the model. Its calculation formula is as follows: (1) In formula (1), p i represents the predicted probability that the i -th sample belongs to the correct class, is the balance factor used to control the weights of positive and negative samples, is the adjustment factor used to adjust the weights of easy-to-classify and difficult-to-classify samples, N represents the total number of samples.
[0027] (2) Dice loss function L d , and its formula is as follows: (2) In formula (2), p i is the predicted value of the i-th pixel in the segmentation result predicted by the model. In a binary classification problem, it is usually 0 or 1, but in practical applications, it can be a probability value between 0 and 1, indicating the probability that the pixel belongs to the target class. g i represents the true value of the i-th pixel in the ground truth label. It is usually 0 or 1, indicating whether the pixel belongs to the target class.
[0028] Similarity loss function L c has the following formula: (3) Furthermore, the calculation formula of the total loss function L of the model is: L = L f + L d + L c (4) In formula (4), represents a weight parameter, represents a hyperparameter value that controls the proportion of the loss function for reducing modal differences L c in the loss function.
[0029] Step S205: Adjust the model parameters based on the loss function value until the loss function value is less than a preset threshold or the preset maximum number of iterations is reached.
[0030] For example Figure 3 , in a possible implementation, before step S202, it further includes: Step S301: Perform word segmentation on the medical text data to obtain a sub-word sequence; Step S302: Add preset markers before and after the sub-word sequence; Exemplarily, when splitting the input medical text into words or sub-words, the WordPiece word segmentation method can be used. Special markers [CLS] and [SEP] are added before and after the segmented text. The [CLS] marker is used to represent the features of the entire sentence, and the [SEP] marker is used to separate sentences or mark the end of a sentence. Suppose the input medical text is T, and its segmented result is { t 1 , t 2 , t 3 , …, t n}. The result after adding special markers is { |CLS| , t 1 , t 2 , t 3 , …, t n , |SEP|}.
[0031] Step S303: Convert each sub-word in the sub-word sequence to an ID value based on the medical vocabulary to obtain a Token text sequence; Among them, if the sub-word appears in the medical vocabulary, it is converted to the corresponding index value; otherwise, it is converted to a marker. Specifically, the pre-trained vocabulary of the BERT model is used to convert the tokenized text into indices (token IDs) in the vocabulary.
[0032] Furthermore, using the pre-trained vocabulary of the BERT model, each token is converted into its corresponding ID. The result of converting them into Token IDs is { ID CLS , ID 1 , ID 2 , ID 3 , …, ID n , ID SEP}.
[0033] Step S304, generate position encoding and attention mask for the Token text sequence.
[0034] It should be noted that the position information for each token is mainly used to represent the position of each token in the sentence. The position encoding is P = { P CLS , P 1 , P 2 , P 3 , …, P n , P SEP}. Create a mask to indicate which tokens are actual inputs (1) and which are padding parts (0). The attention mask is M = { M CLS , M 1 , M 2 , M 3 , …, M n , M SEP}.
[0035] For example Figure 4 , in a possible implementation, step S202 includes: Step S401, input the ID value, position encoding, and attention mask into the BERT model for forward propagation to obtain the hidden state of each Token text sequence and organize the output features; Specifically, input Token IDs, position encodings, and attention masks into the BERT model to obtain the hidden states H of each token = { H CLS , H 1 , H 2 , H 3 , … , H n , H SEP}.
[0036] Step S402: Input the output features into the linear layer and activation layer in sequence to obtain the second feature vector.
[0037] Specifically, extract the output features of the BERT model, obtain the hidden state of the last layer, and extract the feature vector corresponding to the [CLS] token. This hidden state contains the feature vectors of each token. Finally, obtain the second feature vector after processing F t。
[0038] For example Figure 5 , in one possible implementation, step S203 includes: Step S501: Calculate the first similarity between the first feature vector processed by the average pooling layer and the second feature vector; Specifically, the calculation formula is as follows: (5) In formula (5), represents the first feature vector after processing F va and the second feature vector F ta 's first similarity.
[0039] Step S502: Generate the first attention weight based on the Sigmoid function and the first similarity, and generate the third feature vector enhanced at the scale level based on the first attention weight and the first feature vector; Specifically, the calculation formula is as follows: F v' = F va × θ ( S ( F va, F ta )) (6) In formula (6), θRepresents the Sigmoid function operation. F v’ Represents the third feature vector.
[0040] Step S503: Calculate the second similarity between the third feature vector and the second feature vector based on the average pooling layer. Specifically, the calculation formula is as follows: (7) In formula (7), Represents the second feature vector F t And the third feature vector F v’ Of the second similarity.
[0041] Step S504: Generate the second attention weight based on the Sigmoid function and the second similarity, and generate the text enhancement feature enhanced at the scale level and the spatial level based on the second attention weight and the first feature vector.
[0042] Specifically, the calculation formula is as follows: F vt = F v × θ ( S ( F v', F t ))(8) In formula (8), F vt Represents the text enhancement feature, F v Represents the first feature vector.
[0043] In a possible implementation manner, step S205 includes: Perform backpropagation based on the loss function value, and optimize the connection weights inside the initial model based on the preset optimizer and the gradient.
[0044] As Figure 6 , the present invention also provides an image segmentation method, including: Step S601: Obtain an image segmentation model based on the generation method of the image segmentation model described in any one of the above method items; Step S602: Input the medical image to be segmented and the medical text description corresponding to the medical image into the image segmentation model; Step S603: Obtain the segmented masked medical image.
[0045] As Figure 7, the present invention also provides a generating device 70 for an image segmentation model, including: A model construction module 710, configured to construct an initial model, the structure of the initial model sequentially includes: an image representation module for converting a medical image into a first feature vector, a text representation module for converting a medical text into a second feature vector, a feature fine alignment module for aligning based on the first feature vector and the second feature vector to obtain a text-enhanced feature, and a mask segmentation module for obtaining a prediction result based on the text-enhanced feature and calculating a loss function value of the prediction result; A training and optimization module 720, configured to train the initial model based on the obtained medical image data and medical text data to obtain a trained and complete image segmentation model.
[0046] Such as Figure 8 , the present invention also provides an electronic device 80, including: A memory 810, configured to store a program; A processor 820, coupled to the memory 810, configured to execute the program stored in the memory 810 to implement the steps in the generating method of the image segmentation model or the image segmentation method described in any one of the above method items.
[0047] Those skilled in the art can understand that all or part of the processes of implementing the methods in the above embodiments can be completed by instructing relevant hardware (such as a processor, a controller, etc.) through a computer program, and the computer program can be stored in a computer-readable storage medium. Among them, the computer-readable storage medium is a disk, an optical disc, a read-only memory, or a random access memory, etc.
[0048] The above has introduced in detail a generating method and an image segmentation method, a device, and an electronic device for an image segmentation model provided by the present invention. Specific examples are used in this article to elaborate on the principle and implementation manner of the present invention. The description of the above embodiments is only used to help understand the method and its core idea of the present invention; at the same time, for those skilled in the art, according to the idea of the present invention, there will be changes in the specific implementation manner and application scope. In summary, the content of this specification should not be construed as a limitation to the present invention.
Claims
1. A method for generating an image segmentation model, characterized in that: include: Constructing an initial model, the structure of the initial model sequentially includes: an image representation module for converting a medical image into a first feature vector, a text representation module for converting a medical text into a second feature vector, a feature fine alignment module for aligning the first feature vector and the second feature vector to obtain a text enhancement feature, and a mask segmentation module for obtaining a prediction result based on the text enhancement feature and calculating a loss function value of the prediction result; The initial model is trained based on the acquired medical image data and medical text data to obtain a fully trained image segmentation model.
2. The method for generating an image segmentation model according to claim 1, characterized in that: The image representation module includes: a plurality of downsampling convolution modules; Wherein, the downsampling convolution module includes: a maximum pooling layer, a convolution layer, a normalization module and an activation layer; The text representation module includes: a linear layer, an activation layer, and a BERT model pre-trained based on the medical text data and a medical vocabulary, wherein the medical vocabulary includes: medical vocabulary and corresponding index values; The feature fine alignment module includes: a scale attention module and a spatial attention module, both of which are composed of an average pooling layer and a sigmoid function; The mask segmentation module includes: an upsampling convolution module and a preset loss function.
3. The method for generating an image segmentation model according to claim 2, characterized in that: The training of the initial model based on the acquired medical image data and medical text data includes: Performing down-sampling convolution on the medical image data based on the image representation module to obtain a first feature vector including a plurality of scales; Converting the medical text data into a second feature vector based on the text representation module; Based on the scale attention module and the space attention module, the first feature vector and the second feature vector are aligned in turn at the scale level and the space level to obtain the text enhancement feature; Predicting based on the mask segmentation module and the text enhancement feature to obtain a prediction result, and calculating a loss function value of the prediction result; The model parameters are adjusted based on the loss function value until the loss function value is less than a preset threshold or reaches a preset maximum number of iterations.
4. The method for generating an image segmentation model according to claim 3, characterized in that: Before converting the medical text data into a second feature vector based on the text representation module, the method further includes: Performing word segmentation processing on the medical text data to obtain a subword sequence; Adding preset marks before and after the subword sequence; Based on the medical vocabulary, each subword of the subword sequence is converted into an ID value to obtain a Token text sequence, wherein if the subword appears in the medical vocabulary, it is converted into a corresponding index value, otherwise, it is converted into the token; Generate position encoding and attention mask for the Token text sequence.
5. The method for generating an image segmentation model according to claim 4, characterized in that: The converting the medical text data into a second feature vector based on the text representation module includes: The ID value, position encoding, and attention mask are input into the BERT model for forward propagation to obtain the hidden state of each Token text sequence, and the output features are obtained by sorting. The output features are sequentially input into the linear layer and the activation layer to obtain the second feature vector.
6. The method for generating an image segmentation model according to claim 3, characterized in that: The step of aligning the first feature vector and the second feature vector at the scale level and the space level based on the feature fine alignment module to obtain the text enhancement feature includes: Calculate a first similarity between the first feature vector and the second feature vector based on the average pooling layer; Generate a first attention weight based on the Sigmoid function and the first similarity, and generate a third eigenvector enhanced at the scale level based on the first attention weight and the first eigenvector; Calculating a second similarity between the third feature vector and the second feature vector based on the average pooling layer; A second attention weight is generated based on the Sigmoid function and the second similarity, and a text enhancement feature enhanced at the scale level and the spatial level is generated based on the second attention weight and the first feature vector.
7. The method for generating an image segmentation model according to claim 3, characterized in that: The adjusting the model parameters based on the loss function value comprises: Back propagation is performed based on the loss function value, and the connection weights inside the initial model are optimized based on a preset optimizer and the gradient.
8. An image segmentation method, characterized in that: include: Obtaining an image segmentation model based on the method for generating an image segmentation model according to any one of claims 1 to 7; Inputting a medical image to be segmented and a medical text description corresponding to the medical image into the image segmentation model; Get the segmented mask medical image.
9. A device for generating an image segmentation model, characterized in that: include: A model building module, used to build an initial model, the structure of the initial model sequentially includes: an image representation module for converting a medical image into a first feature vector, a text representation module for converting a medical text into a second feature vector, a feature fine alignment module for aligning the first feature vector and the second feature vector to obtain a text enhancement feature, and a mask segmentation module for obtaining a prediction result based on the text enhancement feature and calculating a loss function value of the prediction result; The training optimization module is used to train the initial model based on the acquired medical image data and medical text data to obtain a fully trained image segmentation model.
10. An electronic device, characterized in that: include: Memory, used to store programs; A processor, coupled to the memory, is used to execute the program stored in the memory to implement the method for generating an image segmentation model as described in any one of claims 1 to 7 or the steps in the image segmentation method as described in claim 8.
Citation Information
Patent Citations
Medical image segmentation method and device, electronic equipment and storage medium
CN115375698A
Anaphora segmentation method based on multi-scale feature selective fusion
CN116152265A
Training method, device and equipment for medical image text alignment model
CN118411504A
Training method of medical image segmentation model, image segmentation method and related device
CN119131522A
Semi-supervised medical image segmentation method for eye movement guided hybrid data enhancement
CN119205802A