Image segmentation method and device, electronic equipment and storage medium
This image segmentation method, trained using semi-supervised learning and generative adversarial models, addresses the problem of low segmentation accuracy in multi-organ segmentation of medical images, achieving high-precision segmentation results in multi-class image segmentation.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- SHENZHEN INST OF ADVANCED TECH CHINESE ACAD OF SCI
- Filing Date
- 2022-11-16
- Publication Date
- 2026-04-28
AI Technical Summary
Existing neural network models suffer from poor segmentation accuracy and low target segmentation image precision in multi-organ segmentation of medical images, especially when organ segmentation labels are insufficient.
A semi-supervised learning approach is used to train the image segmentation model. By acquiring the first indication information of the image to be segmented and the second indication information of the sample segmentation image, the initial segmentation model is trained using a generative adversarial model. The feature extraction capability of the model is improved by combining one-hot encoding, which is suitable for multi-class image segmentation.
It improves the accuracy of image segmentation, enabling the training of accurate image segmentation models with limited labeled data, and is suitable for image segmentation in various application scenarios.
Smart Images

Figure CN115760864B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of computer application technology, and in particular to an image segmentation method, apparatus, electronic device, and storage medium. Background Technology
[0002] Currently, image segmentation methods using trained neural network models are widely used. However, in the medical field, these methods have limitations. For example, medical images may suffer from insufficient segmentation labels for certain organs, or all organs may not be represented in a single image, resulting in inconsistent training sample quality.
[0003] Among related technologies, the segmentation model obtained by using neural networks has poor accuracy, and the accuracy of the obtained target segmentation image is low. Summary of the Invention
[0004] This invention provides an image segmentation method, apparatus, electronic device, and storage medium to solve the technical problem of low accuracy in the acquired target segmented image.
[0005] According to one aspect of the present invention, an image segmentation method is provided, wherein the method includes:
[0006] Obtain first indication information corresponding to the image to be segmented, wherein the first indication information is used to indicate the object to be segmented in the image to be segmented;
[0007] Based on the image to be segmented, the first indication information, and the pre-trained image segmentation model, a target segmented image corresponding to the object to be segmented is obtained;
[0008] The image segmentation model is trained on an initial segmentation model using a semi-supervised learning method based on sample segmentation images and second indication information corresponding to the sample segmentation images. The second indication information is used to indicate the sample segmentation object of the sample segmentation image.
[0009] According to another aspect of the present invention, an image segmentation apparatus is provided, wherein the apparatus comprises:
[0010] An image acquisition module is used to acquire a first indication information corresponding to the image to be segmented, wherein the first indication information is used to indicate the object to be segmented in the image to be segmented;
[0011] The image segmentation module is used to obtain a target segmented image corresponding to the object to be segmented based on the image to be segmented, the first indication information, and a pre-trained image segmentation model;
[0012] The image segmentation model is trained on an initial segmentation model using a semi-supervised learning method based on sample segmentation images and second indication information corresponding to the sample segmentation images. The second indication information is used to indicate the sample segmentation object of the sample segmentation image.
[0013] According to another aspect of the present invention, an electronic device is provided, the electronic device comprising:
[0014] At least one processor; and
[0015] A memory communicatively connected to the at least one processor; wherein,
[0016] The memory stores a computer program that can be executed by the at least one processor, the computer program being executed by the at least one processor to enable the at least one processor to perform the image segmentation method according to any embodiment of the present invention.
[0017] According to another aspect of the present invention, a computer-readable storage medium is provided, the computer-readable storage medium storing computer instructions for causing a processor to execute and implement the image segmentation method according to any embodiment of the present invention.
[0018] The technical solution of this invention involves obtaining a first indication information corresponding to an image to be segmented. This first indication information indicates the object to be segmented in the image to be segmented. In an image to be segmented that includes multiple objects to be segmented, the current object can be identified, providing a basis for accurate segmentation of the object by the model. Based on the image to be segmented, the first indication information, and a pre-trained image segmentation model, a target segmented image corresponding to the object to be segmented is obtained. The image segmentation model is trained using a semi-supervised learning method based on sample segmented images and second indication information corresponding to those sample segmented images. The second indication information indicates the sample segmented object in the sample segmented image. The semi-supervised learning method allows for training a more accurate image segmentation model with less labeled data. This technical solution is applicable to image segmentation in various application scenarios and can improve the accuracy of image segmentation.
[0019] It should be understood that the description in this section is not intended to identify key or essential features of the embodiments of the present invention, nor is it intended to limit the scope of the invention. Other features of the invention will become readily apparent from the following description. Attached Figure Description
[0020] To more clearly illustrate the technical solutions in the embodiments of the present invention, the accompanying drawings used in the description of the embodiments will be briefly introduced below. Obviously, the accompanying drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0021] Figure 1 This is a flowchart of an image segmentation method provided according to Embodiment 1 of the present invention;
[0022] Figure 2 This is a flowchart of an image segmentation method provided according to Embodiment 2 of the present invention;
[0023] Figure 3 This is a schematic diagram illustrating the training process of a discrimination model provided according to an embodiment of the present invention;
[0024] Figure 4 This is a schematic diagram of the training process of a generative model according to an embodiment of the present invention;
[0025] Figure 5 This is a schematic diagram of the structure of an image segmentation device according to Embodiment 3 of the present invention;
[0026] Figure 6 This is a schematic diagram of the structure of an electronic device that implements the image segmentation method of this invention. Detailed Implementation
[0027] To enable those skilled in the art to better understand the present invention, the technical solutions of the present invention will be clearly and completely described below with reference to the accompanying drawings of the embodiments. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort should fall within the scope of protection of the present invention.
[0028] It should be noted that the terms "first," "second," etc., in the specification, claims, and accompanying drawings of this invention are used to distinguish similar objects and are not necessarily used to describe a specific order or sequence. It should be understood that such data can be interchanged where appropriate so that the embodiments of the invention described herein can be implemented in orders other than those illustrated or described herein. Furthermore, the terms "comprising" and "having," and any variations thereof, are intended to cover a non-exclusive inclusion; for example, a process, method, system, product, or apparatus that comprises a series of steps or units is not necessarily limited to those steps or units explicitly listed, but may include other steps or units not explicitly listed or inherent to such processes, methods, products, or apparatus.
[0029] Example 1
[0030] Figure 1 This is a flowchart illustrating an image segmentation method according to Embodiment 1 of the present invention. This embodiment is applicable to image processing scenarios. The method can be executed by an image segmentation device, which can be implemented in hardware and / or software and can be configured in a computer. Figure 1 As shown, the method includes:
[0031] S110. Obtain the first indication information corresponding to the image to be segmented and the image to be segmented.
[0032] The first indication information is used to indicate the category corresponding to the object to be segmented in the image to be segmented, and the image to be segmented includes multiple categories of objects to be segmented. The image to be segmented can be understood as an image to be segmented. In this embodiment of the invention, the image to be segmented can be set according to the needs of the scenario, and is not specifically limited here.
[0033] For example, in a medical scenario, the image to be segmented can be a multi-organ image or a tumor image, etc.; in a traffic scenario, the image to be segmented can be a multi-vehicle image, etc.; in a home scenario, the image to be segmented can be a multi-furniture image, etc. For example, in a medical scenario, the image to be segmented can be a head and neck organ image, an abdominal multi-organ image, a liver, pancreas, and kidney multi-organ image, or a tumor image, etc. In a traffic scenario, the image to be segmented can be a highway multi-vehicle image or an intersection multi-vehicle image, etc. In a home scenario, the image to be segmented can be a living room multi-furniture image, a kitchen multi-furniture image, and a bedroom multi-furniture image, etc.
[0034] The first indication information can be understood as information used to distinguish multiple objects to be segmented in the image to be segmented. In this embodiment of the invention, the first indication information can be used to indicate the category corresponding to the object to be segmented in the image to be segmented, and the image to be segmented includes multiple categories of objects to be segmented. Optionally, in a medical scenario, the first indication information can be information distinguishing the category of the eye, brain, or neck in a head and neck organ image, or information distinguishing the category of the liver, kidney, or pancreas in a liver, pancreas, and kidney multi-organ image. In a traffic scenario, the first indication information can be information distinguishing the category of motorcycles, trucks, or cars in a high-speed multi-vehicle image. In a home scenario, the first indication information can be information distinguishing the category of bowls, chopsticks, or spoons in a kitchen multi-furniture image. It is understood that each type of first indication information can indicate one category of objects to be segmented, and one category of objects to be segmented can include one or more objects to be segmented. For example, when the first indication information indicates information about the kidney in a liver, pancreas, and kidney multi-organ image, the first indication information can indicate two objects to be segmented.
[0035] Optionally, the first indication information is preset encoding information corresponding to the object to be segmented, or the first indication information is information obtained by concatenating the preset encoding information corresponding to the object to be segmented with the image to be segmented. In this embodiment of the invention, the image to be segmented can be segmented by the image segmentation model based on the indication relationship between different first indication information and different objects to be segmented in the image to be segmented, which can improve the accuracy and applicability of image segmentation.
[0036] The preset encoding information can be understood as information pre-set and encoded in a certain way to distinguish between different categories of objects to be segmented. In this embodiment of the invention, the preset encoding information can be preset according to scenario requirements, and is not specifically limited here. Specifically, optionally, the preset encoding information includes encoding information generated based on one-hot encoding.
[0037] The one-hot encoding method uses an N-bit state register to encode N states. In this embodiment, the encoding information generated based on one-hot encoding can solve the problem of classifiers' poor handling of attribute data, thus expanding the features. Optionally, the preset encoding information can be first encoding information, second encoding information, and third encoding information, etc. For example, in a multi-organ image of the liver, pancreas, and kidneys in a medical scene, the first encoding information can be used to indicate the liver to be segmented, the second encoding information can be used to indicate the pancreas to be segmented, and the third encoding information can be used to indicate the kidneys to be segmented, etc. In a high-speed multi-vehicle image in a traffic scene, the first encoding information can be used to indicate motorcycles to be segmented, the second encoding information can be used to indicate trucks to be segmented, and the third encoding information can be used to indicate cars to be segmented, etc.
[0038] The object to be segmented can be understood as the object to be segmented in the image to be segmented, or as the region of interest to be segmented from the image to be segmented. In this embodiment of the invention, the object to be segmented can be set according to the needs of the scenario, and is not specifically limited here. Optionally, the object to be segmented can be an organ or tumor to be segmented in the image to be segmented. For example, in a multi-organ image of the liver, pancreas, and kidneys in a medical scenario, the object to be segmented can be the liver, pancreas, or kidney in the multi-organ image of the liver, pancreas, and kidneys. In a multi-furniture image of a bedroom in a home scenario, the object to be segmented can be the bed, wardrobe, or lamp in the multi-furniture image of the bedroom.
[0039] S120. Based on the image to be segmented, the first indication information, and the pre-trained image segmentation model, a target segmentation image corresponding to the object to be segmented is obtained.
[0040] The image segmentation model can be understood as an artificial intelligence model used to segment the image to be segmented. The target segmentation image can be understood as the image determined by the image segmentation model based on the image to be segmented and the first indication information. Optionally, the target segmentation image can be an image from which the object to be segmented is segmented from the image to be segmented. It is understood that the target segmentation image can be the image to be segmented with the object to be segmented marked, or it can be an image obtained by segmenting the image region containing the object to be segmented from the image to be segmented, that is, a portion of the image to be segmented.
[0041] Optionally, the image segmentation model is trained by semi-supervised learning based on the sample segmented image and the second indication information corresponding to the sample segmented image.
[0042] The sample segmentation image can be understood as a sample image used to train the image segmentation model. In this embodiment of the invention, the sample segmentation image can be set according to the scenario requirements, and is not specifically limited here. Optionally, the sample segmentation image may include training images of one or more sample segmentation objects.
[0043] The second indication information can be understood as information used to indicate the sample segmentation objects in the sample segmentation image. In this embodiment of the invention, the second indication information can be information of the same type as the first indication information. In this embodiment of the invention, the second indication information corresponding to the sample segmentation image can accurately indicate each of the sample segmentation objects. Moreover, for sample segmentation images including multiple sample segmentation objects, different second indication information can be combined as training samples, which is particularly suitable for training scenarios with scarce sample size. A highly accurate image segmentation model can be obtained by training with a small number of sample segmentation images.
[0044] The sample segmentation object can be understood as the object in the sample segmentation image that corresponds to the second indication information. In this embodiment of the invention, the indication relationship between the second indication information and the object to be segmented can be preset according to scenario requirements, and is not specifically limited here.
[0045] Optionally, the initial segmentation model is a generative adversarial model, which includes a generative model and a discriminative model. The image segmentation model is trained in the following manner:
[0046] Obtain a semi-supervised sample set, wherein the semi-supervised sample set includes a first number of labeled sample segmentation images and a second number of unlabeled sample segmentation images;
[0047] Determine second indication information corresponding to the sample segmentation image, wherein the second indication information is used to indicate the sample segmentation object of the sample segmentation image, and the label is the expected segmentation image corresponding to the second indication information;
[0048] Based on the sample segmentation images in the semi-supervised sample set and the second indication information corresponding to the sample segmentation images, the generative adversarial model is trained, and the trained generative model is used as the image segmentation model.
[0049] The semi-supervised sample set can be understood as the sample set required to train the initial segmentation model using a semi-supervised learning approach. Optionally, the semi-supervised sample set may include a first number of labeled sample segmentation images and a second number of unlabeled sample segmentation images.
[0050] Wherein, the first quantity can be understood as the number of labeled sample segmentation images in the semi-supervised sample set. Optionally, the first quantity and the second quantity can be in the thousands or tens of thousands, etc. For example, the first quantity can be 1,000, 5,000, or 10,000, etc. The second quantity can be understood as the number of unlabeled sample segmentation images in the semi-supervised sample set. In this embodiment of the invention, the first quantity and the second quantity can be preset according to the scenario requirements, and are not specifically limited here. The first quantity and the second quantity can be the same or different.
[0051] The label can be understood as the expected segmentation image corresponding to the second indication information. The expected segmentation image can be understood as the image expected to be output by the initial segmentation model based on the second indication information corresponding to the sample segmentation image. Considering that in practical application scenarios, such as in medical scenarios, it may be difficult to obtain sample labels, the first number may optionally be less than the second number.
[0052] The generative adversarial model can be understood as a network model composed of a generative model and a discriminative model. The generative model can be understood as a model used to generate a segmented image corresponding to the sample segmented image. The discriminative model can be understood as determining whether an image generated based on the generative model that is similar to the sample segmented image is a true sample segmented image. In this embodiment of the invention, the discriminative model can be trained based on the output image of the first model, second indication information corresponding to the output image of the first model, and a desired segmentation image corresponding to the output image of the second model.
[0053] Specifically, based on the alternating iteration of the generative model and the discriminative model in the generative adversarial model, the model parameters of the generative model are adjusted to obtain the image segmentation model.
[0054] Optionally, the step of training the generative adversarial model based on the sample segmentation images in the semi-supervised sample set and the second indication information corresponding to the sample segmentation images, and using the trained generative model as the image segmentation model, includes:
[0055] The labeled sample segmentation image in the semi-supervised sample set and the second indication information corresponding to the sample segmentation image are input into the generative model in the generative adversarial model to obtain the first model output image.
[0056] Based on the first model output image, the sample segmentation image, and the expected segmentation image corresponding to the sample segmentation image, the model parameters of the generated model are adjusted to obtain an initial segmentation model;
[0057] The unlabeled sample segmentation images in the semi-supervised sample set and the second indication information corresponding to the sample segmentation images are input into the initial segmentation model to obtain the second model output image;
[0058] The model parameters of the generating model are adjusted based on the discrimination results of the first model output image and the second model output image after training. The discrimination model is trained based on the first model output image, the second indication information corresponding to the first model output image, and the expected segmentation image corresponding to the second model output image.
[0059] If the training termination condition is met, the trained generative model is used as the image segmentation model.
[0060] The first model output image can be understood as an image generated by a generation model based on a labeled sample segmentation image and the second indication information corresponding to the sample segmentation image. It can be understood that the first model output image can be an image output by the generation model after inputting the labeled sample segmentation image, resulting in the segmentation of the sample object corresponding to the second indication information. Similarly, the second model output image can be understood as an image generated by the initial segmentation model based on an unlabeled sample segmentation image and the second indication information corresponding to the sample segmentation image.
[0061] The discrimination result can be understood as the probability of judging the first model output image and the second model output image as true or false using the discrimination model. In other words, the discrimination result can be the result of judging whether the image input to the discrimination model is the model output image or the desired segmented image.
[0062] The model parameters can be understood as the configuration parameters of the generated model. These model parameters generally include the weights and offsets between layers.
[0063] The training termination condition can be understood as the condition under which the training of the generative adversarial model can be terminated. In this embodiment of the invention, the training termination condition can be preset according to the scenario requirements, and is not specifically limited here. Optionally, the training termination condition may be that the model parameters of the generative model no longer change in the same direction, that is, the model parameters converge; or, a preset number of iterations is reached; or, the discrimination error rate of the discrimination model on the output image of the generative model reaches a preset value, etc.
[0064] Specifically, the training methods for generative adversarial models can be:
[0065] 1. Training the discriminative model of the generative adversarial model. First, the discriminative model of the generative adversarial model is trained based on the second indication information, the expected segmentation image corresponding to the second indication information, and the Gaussian noise image. Specifically, the second indication information is input into the first sub-model of the discriminative model in the generative adversarial model to obtain initial network parameters, and the network parameters of the second sub-model of the discriminative model are updated based on the initial network parameters.
[0066] 2. Training the generative model of the generative adversarial model. Specifically, the second instruction information is first input into the first sub-model of the generative model in the generative adversarial model to obtain initial network parameters. The initial network parameters are then used as the network parameters for the attention mechanism in the second sub-model of the generative model. Then, the sample segmentation image is input into the second sub-model of the generative model to obtain the model segmentation image output by the generative model.
[0067] Furthermore, in the process of inputting the sample segmentation image into the second sub-model of the generative model, the labeled sample segmentation image is first input into the second sub-model of the generative model to obtain the output image of the first model. Then, the unlabeled sample segmentation image is input into the generative model to obtain the output image of the second model.
[0068] 3. Adjust the model parameters of the generative model based on the discrimination results. Specifically, the second-type output image and the corresponding expected segmentation image can be input into the discrimination model to obtain the discrimination result corresponding to the second-type output image; the model parameters of the generative model are then adjusted based on the discrimination result. Furthermore, the discrimination model can be further trained based on the discrimination result to obtain a trained discrimination model. Additionally, the model parameters of the generative model can be further adjusted based on the discrimination result of the newly trained discrimination model.
[0069] Finally, if the model parameters converge, the trained generative model can be used as the image segmentation model.
[0070] In this embodiment of the invention, a generative adversarial model is trained based on the sample segmentation images in the semi-supervised sample set and the second indication information corresponding to the sample segmentation images. The trained generative model is used as the image segmentation model, which solves the problem of insufficient sample set labels. While training the generative adversarial model based on semi-supervised learning, the accuracy of the obtained image segmentation model is improved.
[0071] The technical solution of this invention involves obtaining a first indication information corresponding to an image to be segmented. This first indication information indicates the object to be segmented in the image to be segmented. In an image to be segmented that includes multiple objects to be segmented, the current object can be identified, providing a basis for accurate segmentation of the object by the model. Based on the image to be segmented, the first indication information, and a pre-trained image segmentation model, a target segmented image corresponding to the object to be segmented is obtained. The image segmentation model is trained using a semi-supervised learning method based on sample segmented images and second indication information corresponding to those sample segmented images. The second indication information indicates the sample segmented object in the sample segmented image. The semi-supervised learning method allows for training a more accurate image segmentation model with less labeled data. This technical solution is applicable to image segmentation in various application scenarios and can improve the accuracy of image segmentation.
[0072] Example 2
[0073] Figure 2 This is a flowchart of an image segmentation method provided in Embodiment 2 of the present invention. This embodiment is based on the above embodiment, which refines the target segmentation image corresponding to the object to be segmented, based on the image to be segmented, the first indication information, and the pre-trained image segmentation model.
[0074] like Figure 2 As shown, the method includes:
[0075] S210. Obtain the first indication information corresponding to the image to be segmented and the image to be segmented, wherein the first indication information is used to indicate the object to be segmented in the image to be segmented.
[0076] S220. Input the first indication information into the first sub-model of the pre-trained image segmentation model to obtain the initial network parameters of the second sub-model of the image segmentation model, wherein the initial network parameters include at least initial weights and offset values.
[0077] In this embodiment of the invention, the image segmentation model may include a first sub-model and a second sub-model.
[0078] The first sub-model can be understood as a model that obtains the initial network parameters of the second sub-model of the image segmentation model based on the first indication information.
[0079] The initial network parameters can be understood as the initial values of the initial weights and offsets of each node before the second sub-model is trained. The initial weights can be understood as the initial values of the weights in the initial network parameters. The offsets can be understood as the difference between the logical address of the program and the beginning of the segment.
[0080] Optionally, the first sub-model includes multiple convolutional layers, wherein at least two convolutional layers are connected based on non-linear activation function layers, and the second sub-model includes an attention mechanism;
[0081] Updating the network parameters of the second sub-model based on the initial network parameters includes:
[0082] The network parameters of the attention mechanism are updated based on the initial network parameters.
[0083] The convolutional layer can be understood as a network layer that extracts different features of the first indication information input. The nonlinear activation function layer can be understood as a nonlinear network layer that adds nonlinearity to the image segmentation model. The attention mechanism can be understood as a mechanism that enables the second sub-model to focus on the object to be segmented corresponding to the first indication information input, i.e., selects a specific input. In this embodiment of the invention, the attention mechanism can allocate resources under limited computing power, allocating computing resources to more important tasks and solving the problem of information overload. This improves computing speed and the efficiency of acquiring the target segmented image.
[0084] S230. Update the network parameters of the second sub-model based on the initial network parameters, and input the image to be segmented into the second sub-model after the network parameters have been updated to obtain the target segmentation image corresponding to the object to be segmented.
[0085] The second sub-model can be understood as a model that obtains the target segmentation image corresponding to the object to be segmented based on the image to be segmented and the initial network parameters.
[0086] Specifically, the first indication information is input into the first sub-model of the pre-trained image segmentation model to obtain the initial weights and offset values of the second sub-model of the image segmentation model; the network parameters of the second sub-model are updated based on the initial weights and offset values, and the image to be segmented is input into the second sub-model after the network parameters are updated to obtain the target segmented image corresponding to the object to be segmented.
[0087] The technical solution of this invention involves inputting the first indication information into a first sub-model of a pre-trained image segmentation model to obtain the initial network parameters of a second sub-model of the image segmentation model. The initial network parameters include at least initial weights and offset values. Based on the initial network parameters, the network parameters of the second sub-model are updated, and the image to be segmented is input into the updated second sub-model to obtain a target segmented image corresponding to the object to be segmented. By processing the first indication information corresponding to the image to be segmented, the target segmented image is accurately obtained.
[0088] Optional, Figure 3 This is a schematic diagram illustrating the training process of a discrimination model according to an embodiment of the present invention; as shown below. Figure 3 As shown, the training process of the discriminative model can be as follows:
[0089] 1. Input the second indication information into the first sub-model of the discriminant model to obtain the initial network parameters of the second sub-model. Based on one-hot encoding, the input organ number is encoded to generate an m-dimensional one-hot code, which is the second indication information; then, the second indication information is input into the first sub-model of the discriminant model to obtain the weights and biases of the n convolutional layers, which are the initial network parameters.
[0090] 2. Use the obtained initial network parameters as the initial network parameters of the second sub-model of the discriminant model.
[0091] 3. The second sub-model is trained based on image labels (expected segmented image) and Gaussian noise map (pseudo-labeled image). The image labels and Gaussian noise are input into the second sub-model of the discriminative model to obtain a discrimination result of 0 or 1, where 0 can represent the discrimination result as a fake image and 1 can represent the discrimination result as a real image.
[0092] Optional, Figure 4 This is a schematic diagram illustrating the training process of a generative model according to an embodiment of the present invention; as shown below. Figure 4 As shown, the training process for a generative model can be as follows:
[0093] 1. Input the second instruction information into the first sub-model of the generative model to obtain the initial network parameters of the second sub-model. Then, input the second instruction information into the first sub-model of the generative model to obtain the weights and biases of the n convolutional layers of the generative model, i.e., the initial network parameters.
[0094] 2. Use the obtained initial network parameters as the initial network parameters for the attention mechanism in the second sub-model of the generative model.
[0095] 3. The second sub-model is trained based on labeled sample segmentation images, unlabeled sample images, and a discriminative model to obtain the model's output image. The labeled sample segmentation images and unlabeled sample images are then input into the second sub-model of the generative model to obtain the model's output image. During this process, the model parameters of the second sub-model are adjusted based on the discrimination results of the discriminative model to obtain the trained generative model, i.e., the image segmentation model.
[0096] 4. Input the image to be segmented and the first indication information into the trained image segmentation model to obtain the target segmented image.
[0097] This technical solution proposes a multi-task semi-supervised segmentation method based on object recognition, applicable to multi-class object segmentation using single or multiple slices. Compared to related multi-class object segmentation techniques, this invention does not require all labels for all objects to be segmented, nor does it require all labels for all objects to be segmented to be on a single image. Besides more flexibly addressing the label shortage problem in multi-class object segmentation, it is also more suitable for solving multi-class object segmentation problems on large datasets.
[0098] Example 3
[0099] Figure 5 This is a schematic diagram of the structure of an image segmentation device provided in Embodiment 3 of the present invention. Figure 5 As shown, the device includes an image acquisition module 310 and an image segmentation module 320.
[0100] The image acquisition module 310 is used to acquire the image to be segmented and the first indication information corresponding to the image to be segmented, wherein the first indication information is used to indicate the object to be segmented in the image to be segmented; the image segmentation module 320 is used to obtain a target segmented image corresponding to the object to be segmented based on the image to be segmented, the first indication information and a pre-trained image segmentation model; wherein the image segmentation model is trained on an initial segmentation model using a semi-supervised learning method based on a sample segmented image and the second indication information corresponding to the sample segmented image, wherein the second indication information is used to indicate the sample segmented object in the sample segmented image.
[0101] The technical solution of this invention involves obtaining a first indication information corresponding to an image to be segmented. This first indication information indicates the object to be segmented in the image to be segmented. In an image to be segmented that includes multiple objects to be segmented, the current object can be identified, providing a basis for accurate segmentation of the object by the model. Based on the image to be segmented, the first indication information, and a pre-trained image segmentation model, a target segmented image corresponding to the object to be segmented is obtained. The image segmentation model is trained using a semi-supervised learning method based on sample segmented images and second indication information corresponding to those sample segmented images. The second indication information indicates the sample segmented object in the sample segmented image. The semi-supervised learning method allows for training a more accurate image segmentation model with less labeled data. This technical solution is applicable to image segmentation in various application scenarios and can improve the accuracy of image segmentation.
[0102] Optionally, the image segmentation model includes a first sub-model and a second sub-model;
[0103] Image segmentation module 320, used for:
[0104] The first indication information is input into the first sub-model of the pre-trained image segmentation model to obtain the initial network parameters of the second sub-model of the image segmentation model, wherein the initial network parameters include at least initial weights and offset values;
[0105] The network parameters of the second sub-model are updated based on the initial network parameters, and the image to be segmented is input into the second sub-model after the network parameters are updated to obtain the target segmentation image corresponding to the object to be segmented.
[0106] Optionally, the first indication information is the preset encoding information corresponding to the object to be segmented, or the first indication information is the information obtained by splicing the preset encoding information corresponding to the object to be segmented with the image to be segmented.
[0107] Optionally, the preset encoding information includes encoding information generated based on one-hot encoding.
[0108] Optionally, the first sub-model includes multiple convolutional layers, wherein at least two convolutional layers are connected based on a non-linear activation function layer.
[0109] Optionally, the initial segmentation model is a generative adversarial model, which includes a generative model and a discriminative model; the image segmentation model can be trained based on a model training module, wherein the model training module includes: a sample set acquisition submodule, an indication information determination submodule, and a model training submodule.
[0110] The sample set acquisition submodule is used to acquire a semi-supervised sample set, wherein the semi-supervised sample set includes a first number of labeled sample segmentation images and a second number of unlabeled sample segmentation images.
[0111] The indication information determination submodule is used to determine the second indication information corresponding to the sample segmentation image, wherein the second indication information is used to indicate the sample segmentation object of the sample segmentation image, and the label is the expected segmentation image corresponding to the second indication information;
[0112] The model training submodule is used to train the generative adversarial model based on the sample segmentation images in the semi-supervised sample set and the second indication information corresponding to the sample segmentation images, and to use the trained generative model as the image segmentation model.
[0113] Optionally, the model training submodule is used for:
[0114] The labeled sample segmentation image in the semi-supervised sample set and the second indication information corresponding to the sample segmentation image are input into the generative model in the generative adversarial model to obtain the first model output image.
[0115] Based on the first model output image, the sample segmentation image, and the expected segmentation image corresponding to the sample segmentation image, the model parameters of the generated model are adjusted to obtain an initial segmentation model;
[0116] The unlabeled sample segmentation images in the semi-supervised sample set and the second indication information corresponding to the sample segmentation images are input into the initial segmentation model to obtain the second model output image;
[0117] The model parameters of the generating model are adjusted based on the discrimination results of the first model output image and the second model output image after training. The discrimination model is trained based on the first model output image, the second indication information corresponding to the first model output image, and the expected segmentation image corresponding to the second model output image.
[0118] If the training termination condition is met, the trained generative model is used as the image segmentation model.
[0119] The image segmentation apparatus provided in the embodiments of the present invention can execute the image segmentation method provided in any embodiment of the present invention, and has the corresponding functional modules and beneficial effects of executing the method.
[0120] Example 4
[0121] Figure 6A schematic diagram of an electronic device 10 that can be used to implement embodiments of the present invention is shown. The electronic device is intended to represent various forms of digital computers, such as laptop computers, desktop computers, workstations, personal digital assistants, servers, blade servers, mainframe computers, and other suitable computers. The electronic device can also represent various forms of mobile devices, such as personal digital processors, cellular phones, smartphones, wearable devices (e.g., helmets, glasses, watches, etc.), and other similar computing devices. The components shown herein, their connections and relationships, and their functions are merely illustrative and are not intended to limit the implementation of the invention described and / or claimed herein.
[0122] like Figure 6 As shown, the electronic device 10 includes at least one processor 11 and a memory, such as a read-only memory (ROM) 12 or a random access memory (RAM) 13, communicatively connected to the at least one processor 11. The memory stores computer programs executable by the at least one processor. The processor 11 can perform various appropriate actions and processes based on the computer program stored in the ROM 12 or loaded from storage unit 18 into the RAM 13. The RAM 13 may also store various programs and data required for the operation of the electronic device 10. The processor 11, ROM 12, and RAM 13 are interconnected via a bus 14. An input / output (I / O) interface 15 is also connected to the bus 14.
[0123] Multiple components in electronic device 10 are connected to I / O interface 15, including: input unit 16, such as keyboard, mouse, etc.; output unit 17, such as various types of displays, speakers, etc.; storage unit 18, such as disk, optical disk, etc.; and communication unit 19, such as network card, modem, wireless transceiver, etc. Communication unit 19 allows electronic device 10 to exchange information / data with other devices through computer networks such as the Internet and / or various telecommunications networks.
[0124] Processor 11 can be a variety of general-purpose and / or special-purpose processing components with processing and computing capabilities. Some examples of processor 11 include, but are not limited to, a central processing unit (CPU), a graphics processing unit (GPU), various special-purpose artificial intelligence (AI) computing chips, various processors running machine learning model algorithms, a digital signal processor (DSP), and any suitable processor, controller, microcontroller, etc. Processor 11 performs the various methods and processes described above, such as image segmentation methods.
[0125] In some embodiments, the image segmentation method may be implemented as a computer program tangibly contained in a computer-readable storage medium, such as storage unit 18. In some embodiments, part or all of the computer program may be loaded and / or mounted on electronic device 10 via ROM 12 and / or communication unit 19. When the computer program is loaded into RAM 13 and executed by processor 11, one or more steps of the image segmentation method described above may be performed. Alternatively, in other embodiments, processor 11 may be configured to perform the image segmentation method by any other suitable means (e.g., by means of firmware).
[0126] Various embodiments of the systems and techniques described above herein can be implemented in digital electronic circuit systems, integrated circuit systems, field-programmable gate arrays (FPGAs), application-specific integrated circuits (ASICs), application-specific standard products (ASSPs), systems-on-a-chip (SoCs), payload-programmable logic devices (CPLDs), computer hardware, firmware, software, and / or combinations thereof. These various embodiments may include implementations in one or more computer programs that can be executed and / or interpreted on a programmable system including at least one programmable processor, which may be a dedicated or general-purpose programmable processor, capable of receiving data and instructions from a storage system, at least one input device, and at least one output device, and transmitting data and instructions to the storage system, the at least one input device, and the at least one output device.
[0127] Computer programs used to implement the methods of the present invention may be written in any combination of one or more programming languages. These computer programs may be provided to a processor of a general-purpose computer, a special-purpose computer, or other programmable data processing device, such that when executed by the processor, the computer programs cause the functions / operations specified in the flowcharts and / or block diagrams to be performed. The computer programs may be executed entirely on a machine, partially on a machine, or as a standalone software package, partially on a machine and partially on a remote machine, or entirely on a remote machine or server.
[0128] In the context of this invention, a computer-readable storage medium can be a tangible medium that may contain or store a computer program for use by or in conjunction with an instruction execution system, apparatus, or device. A computer-readable storage medium may include, but is not limited to, electronic, magnetic, optical, electromagnetic, infrared, or semiconductor systems, apparatus, or devices, or any suitable combination thereof. Alternatively, a computer-readable storage medium may be a machine-readable signal medium. More specific examples of machine-readable storage media include electrical connections based on one or more wires, portable computer disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fibers, portable compact disk read-only memory (CD-ROM), optical storage devices, magnetic storage devices, or any suitable combination thereof.
[0129] To provide interaction with a user, the systems and techniques described herein can be implemented on an electronic device having: a display device for displaying information to the user (e.g., a CRT (cathode ray tube) or LCD (liquid crystal display) monitor); and a keyboard and pointing device (e.g., a mouse or trackball) through which the user provides input to the electronic device. Other types of devices can also be used to provide interaction with the user; for example, feedback provided to the user can be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and input from the user can be received in any form (including sound input, voice input, or tactile input).
[0130] The systems and technologies described herein can be implemented in computing systems that include backend components (e.g., as data servers), or computing systems that include middleware components (e.g., application servers), or computing systems that include frontend components (e.g., user computers with graphical user interfaces or web browsers through which users can interact with implementations of the systems and technologies described herein), or any combination of such backend, middleware, or frontend components. The components of the system can be interconnected via digital data communication of any form or medium (e.g., communication networks). Examples of communication networks include local area networks (LANs), wide area networks (WANs), blockchain networks, and the Internet.
[0131] A computing system can include clients and servers. Clients and servers are generally located far apart and typically interact through communication networks. The client-server relationship is created by computer programs running on the respective computers and having a client-server relationship with each other. The server can be a cloud server, also known as a cloud computing server or cloud host, which is a hosting product within the cloud computing service system to address the shortcomings of traditional physical hosts and VPS services, such as high management difficulty and weak business scalability.
[0132] It should be understood that the various forms of processes shown above can be used, with steps reordered, added, or deleted. For example, the steps described in this invention can be executed in parallel, sequentially, or in different orders, as long as the desired result of the technical solution of this invention can be achieved, and this is not limited herein.
[0133] The specific embodiments described above do not constitute a limitation on the scope of protection of this invention. Those skilled in the art should understand that various modifications, combinations, sub-combinations, and substitutions can be made according to design requirements and other factors. Any modifications, equivalent substitutions, and improvements made within the spirit and principles of this invention should be included within the scope of protection of this invention.
Claims
1. An image segmentation method, characterized in that, include: Obtain first indication information corresponding to the image to be segmented, wherein the first indication information is used to indicate the category corresponding to the object to be segmented in the image to be segmented, and the image to be segmented includes multiple categories of objects to be segmented; Based on the image to be segmented, the first indication information, and the pre-trained image segmentation model, a target segmented image corresponding to the object to be segmented is obtained; The image segmentation model is trained by semi-supervised learning based on the sample segmentation image and the second indication information corresponding to the sample segmentation image. The second indication information is used to indicate the sample segmentation object of the sample segmentation image. The image segmentation model includes a first sub-model and a second sub-model. The step of obtaining a target segmented image corresponding to the object to be segmented based on the image to be segmented, the first indication information, and a pre-trained image segmentation model includes: The first indication information is input into the first sub-model of the pre-trained image segmentation model to obtain the initial network parameters of the second sub-model of the image segmentation model, wherein the initial network parameters include at least initial weights and offset values; The network parameters of the second sub-model are updated based on the initial network parameters, and the image to be segmented is input into the second sub-model after the network parameters are updated to obtain the target segmentation image corresponding to the object to be segmented; The initial segmentation model is a generative adversarial model, which includes a generative model and a discriminative model. The image segmentation model is trained in the following manner: Obtain a semi-supervised sample set, wherein the semi-supervised sample set includes a first number of labeled sample segmentation images and a second number of unlabeled sample segmentation images; Determine second indication information corresponding to the sample segmentation image, wherein the second indication information is used to indicate the sample segmentation object of the sample segmentation image, and the label is the expected segmentation image corresponding to the second indication information; Based on the sample segmentation images in the semi-supervised sample set and the second indication information corresponding to the sample segmentation images, the generative adversarial model is trained, and the trained generative model is used as the image segmentation model.
2. The method according to claim 1, characterized in that, The first sub-model includes multiple convolutional layers, wherein at least two convolutional layers are connected based on a non-linear activation function layer, and the second sub-model includes an attention mechanism. Updating the network parameters of the second sub-model based on the initial network parameters includes: The network parameters of the attention mechanism are updated based on the initial network parameters.
3. The method according to claim 1, characterized in that, The first indication information is the preset encoding information corresponding to the object to be segmented, or the first indication information is the information obtained by splicing the preset encoding information corresponding to the object to be segmented with the image to be segmented.
4. The method according to claim 3, characterized in that, The preset encoding information includes encoding information generated based on one-hot encoding.
5. The method according to claim 1, characterized in that, The step of training a generative adversarial model based on sample segmentation images from the semi-supervised sample set and corresponding second indication information, and using the trained generative model as the image segmentation model, includes: The labeled sample segmentation image in the semi-supervised sample set and the second indication information corresponding to the sample segmentation image are input into the generative model in the generative adversarial model to obtain the first model output image. Based on the first model output image, the sample segmentation image, and the expected segmentation image corresponding to the sample segmentation image, the model parameters of the generated model are adjusted to obtain an initial segmentation model; The unlabeled sample segmentation images in the semi-supervised sample set and the second indication information corresponding to the sample segmentation images are input into the initial segmentation model to obtain the second model output image; The model parameters of the generating model are adjusted based on the discrimination results of the first model output image and the second model output image after training. The discrimination model is trained based on the first model output image, the second indication information corresponding to the first model output image, and the expected segmentation image corresponding to the first model output image. If the training termination condition is met, the trained generative model is used as the image segmentation model.
6. An image segmentation apparatus, characterized in that, include: An image acquisition module is used to acquire a first indication information corresponding to the image to be segmented, wherein the first indication information is used to indicate the object to be segmented in the image to be segmented; The image segmentation module is used to obtain a target segmented image corresponding to the object to be segmented based on the image to be segmented, the first indication information, and a pre-trained image segmentation model; The image segmentation model is trained by semi-supervised learning based on the sample segmentation image and the second indication information corresponding to the sample segmentation image. The second indication information is used to indicate the sample segmentation object of the sample segmentation image. The image segmentation model includes a first sub-model and a second sub-model. The image segmentation module is used for: The first indication information is input into the first sub-model of the pre-trained image segmentation model to obtain the initial network parameters of the second sub-model of the image segmentation model, wherein the initial network parameters include at least initial weights and offset values; The network parameters of the second sub-model are updated based on the initial network parameters, and the image to be segmented is input into the second sub-model after the network parameters are updated to obtain the target segmentation image corresponding to the object to be segmented; The initial segmentation model is a generative adversarial model, which includes a generative model and a discriminative model. The image segmentation model is trained based on a model training module, which includes a sample set acquisition submodule, an indication information determination submodule, and a model training submodule. The sample set acquisition submodule is used to acquire a semi-supervised sample set, wherein the semi-supervised sample set includes a first number of labeled sample segmentation images and a second number of unlabeled sample segmentation images. The indication information determination submodule is used to determine the second indication information corresponding to the sample segmentation image, wherein the second indication information is used to indicate the sample segmentation object of the sample segmentation image, and the label is the expected segmentation image corresponding to the second indication information; The model training submodule is used to train the generative adversarial model based on the sample segmentation images in the semi-supervised sample set and the second indication information corresponding to the sample segmentation images, and to use the trained generative model as the image segmentation model.
7. An electronic device, characterized in that, The electronic device includes: At least one processor; and A memory communicatively connected to the at least one processor; wherein, The memory stores a computer program that can be executed by the at least one processor, the computer program being executed by the at least one processor to enable the at least one processor to perform the image segmentation method according to any one of claims 1-5.
8. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores computer instructions that cause a processor to execute the image segmentation method according to any one of claims 1-5.
Citation Information
Patent Citations
Image segmentation model generation method and device and image segmentation method and device
CN114862878A