Method, apparatus, device, and storage medium for splitting an image

Through the portrait segmentation model combined with the first and second training sets, structured human feature information is extracted, which solves the problem of missegment of deep learning models and achieves efficient and accurate image segmentation.

CN114937047BActive Publication Date: 2025-07-08UBTECH ROBOTICS CORP LTD
View PDF 4 Cites 0 Cited by

Patent Information

Application Number
CN202210545425.5
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-05-19
Publication Date
2025-07-08
Estimated Expiration
2042-05-19

AI Technical Summary

Technical Problem

In the prior art, the deep learning-based segmentation model does not consider structured human body feature information when segmenting images, resulting in frequent missegmentation, inaccurate segmentation results, and high cost and time-consuming training of image annotation.

Method used

The portrait segmentation model is adopted, and the portrait detection model is trained using the first training set and the second training set. The first training set is manually annotated, and the second training set is generated by the third training set through the graph cutting algorithm to extract structured human feature information, reduce the labeling cost and improve training efficiency and accuracy.

Benefits of technology

Effectively avoid missegment, improve the accuracy of segmentation results, reduce training costs and time, improve training efficiency, and enhance model accuracy.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114937047B_ABST
    Figure CN114937047B_ABST
Patent Text Reader

Abstract

This application is applicable to the field of image processing technology, and provides a method, device, equipment and storage medium for segmenting images, including: obtaining an image to be segmented; inputting the image to be segmented into a trained human portrait segmentation model for processing to obtain a segmentation result of the image to be segmented. The human portrait segmentation model is obtained by training a human portrait detection model using a first training set and a second training set, and the second training set is obtained by processing a third training set used for training the human portrait detection model using a graph cut algorithm. In the above solution, the human portrait segmentation model is obtained by training the human portrait detection model using the first training set and the second training set. Since the human body detection model searches for the human body in the image in the form of a detection frame, the human portrait segmentation model can learn the structured human body feature information in the human portrait detection model. When the image to be segmented is processed by this human portrait segmentation model, missegmentation can be effectively avoided, thereby improving the accuracy of the segmentation result.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application belongs to the technical field of image processing, and particularly relates to a method, device, equipment and storage medium for segmenting images. Background Art

[0002] As a common image processing method, human image segmentation technology can be widely applied to various fields such as virtual reality, video surveillance, and behavior analysis.

[0003] In the prior art, when a segmentation model based on deep learning processes an image, the processing mainly classifies each pixel point in the image individually, without considering the structured human feature information. Therefore, it is prone to mis-segmentation, resulting in inaccurate image segmentation results. Summary of the Invention

[0004] In view of this, embodiments of this application provide a method, device, equipment and storage medium for segmenting images, so as to solve the problem that in the prior art, when segmenting an image, mis-segmentation is prone to occur, resulting in inaccurate image segmentation results.

[0005] The first aspect of the embodiments of this application provides a method for segmenting an image, and the method includes:

[0006] Obtain an image to be segmented;

[0007] Input the image to be segmented into a trained human portrait segmentation model for processing, and obtain the segmentation result of the image to be segmented. The human portrait segmentation model is obtained by training a human portrait detection model using a first training set and a second training set. The first training set includes a plurality of first sample images and the corresponding first sample segmentation results of each of the first sample images. The second training set is obtained by processing a third training set used for training the human portrait detection model using a graph cut algorithm.

[0008] In the above solution, the image to be segmented is processed by a trained human portrait segmentation model. The human portrait segmentation model is obtained by training a human portrait detection model using a first training set and a second training set. Since the human body detection model searches for the human body in the image in the form of a detection frame, structured human feature information can be extracted. This enables the human portrait segmentation model to learn the structured human feature information in the human portrait detection model during the training process. Furthermore, when the trained human portrait segmentation model processes the image to be segmented, the occurrence of mis-segmentation is effectively avoided, thereby improving the accuracy of the segmentation result.

[0009] Meanwhile, the second training set is obtained by processing the third training set used to train the human portrait detection model with the graph cut algorithm. The third training set used to train the human portrait detection model has a low annotation cost and is easy to obtain. Processing the third training set with the graph cut algorithm to obtain the second training set required for training the human portrait segmentation model can reduce the cost of training the human portrait segmentation model, reduce the time for training the human portrait segmentation model, and improve the efficiency of training the human portrait segmentation model. Moreover, using a large number of second training sets for training can improve the accuracy of the human portrait segmentation model.

[0010] Optionally, the human portrait segmentation model includes a human body image detection network and a human body image segmentation network. The segmentation result includes a human body detection frame and a human body image. Inputting the image to be segmented into the trained human portrait segmentation model for processing to obtain the segmentation result of the image to be segmented includes:

[0011] Performing detection processing on the image to be segmented through the human body image detection network to obtain the human body detection frame, where the human body detection frame is used to predict the area where the human body is located in the image to be segmented;

[0012] Performing segmentation processing on the image to be segmented through the human body image segmentation network to obtain the human body image.

[0013] Optionally, after performing segmentation processing on the image to be segmented through the human body image segmentation network to obtain the human body image, the method further includes:

[0014] Using the human body detection frame to inspect the human body image;

[0015] Determining the final segmentation result of the image to be segmented according to the inspection result.

[0016] Optionally, determining the final segmentation result of the image to be segmented according to the inspection result includes:

[0017] When the inspection result is that the area corresponding to the human body image exceeds the area corresponding to the human body detection frame, deleting the image of the area exceeding the area corresponding to the human body detection frame in the human body image to obtain the final segmentation result.

[0018] Optionally, the second training set includes a plurality of second sample images, second sample segmentation results corresponding to each of the second sample images, and sample detection results corresponding to each of the second sample images. Before obtaining the image to be segmented, the method further includes:

[0019] Training an initial model with the third training set and a preset loss function to obtain the human portrait detection model. The third training set includes a plurality of third sample images and target detection frames corresponding to each of the third sample images;

[0020] Using the first training set, the second training set, and the loss function, train the human face detection model to obtain the human face segmentation model.

[0021] Optionally, before using the first training set, the second training set, and the loss function to train the human face detection model to obtain the human face segmentation model, the method further includes:

[0022] According to the target detection box corresponding to each third sample image, determine a target image in each third sample image;

[0023] Use the graph cut algorithm to classify each target image to obtain a classification result corresponding to each target image;

[0024] Generate the second training set according to the respective classification results and the respective third sample images.

[0025] Optionally, during the process of training the human face segmentation model, when it is detected that the training set is the first training set, adjust the first parameter in the loss function; and / or, when it is detected that the training set is the second training set, adjust the second parameter in the loss function.

[0026] A second aspect of the embodiments of the present application provides an apparatus for segmenting an image, including:

[0027] An acquisition unit, configured to acquire an image to be segmented;

[0028] A processing unit, configured to input the image to be segmented into a trained human face segmentation model for processing to obtain a segmentation result of the image to be segmented. The human face segmentation model is obtained by training a human face detection model using a first training set and a second training set. The first training set includes a plurality of first sample images and first sample segmentation results corresponding to the respective first sample images. The second training set is obtained by processing a third training set used for training the human face detection model using the graph cut algorithm.

[0029] A third aspect of the embodiments of the present application provides a device for segmenting an image, including a memory, a processor, and a computer program stored in the memory and executable on the processor. The processor, when executing the computer program, implements the steps of the method for segmenting an image as described in the first aspect above.

[0030] The fourth aspect of the embodiments of the present application provides a computer-readable storage medium storing a computer program, and when the computer program is executed by a processor, the steps of the method for segmenting an image as described in the first aspect above are implemented.

[0031] The fifth aspect of the embodiments of the present application provides a computer program product, and when the computer program product runs on a device for segmenting an image, the device for segmenting an image is enabled to execute the steps of the method for segmenting an image as described in the first aspect above. Description of the Drawings

[0032] In order to more clearly illustrate the technical solutions in the embodiments of the present application, the following will briefly introduce the drawings required for use in the embodiments or the description of the prior art. Obviously, the drawings in the following description are only some embodiments of the present application. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.

[0033] Figure 1 It is a schematic flowchart of a method for segmenting an image provided by an exemplary embodiment of the present application;

[0034] Figure 2 It is a schematic diagram of an image processing result provided by the present application;

[0035] Figure 3 It is a schematic diagram of another image processing result provided by the present application;

[0036] Figure 4 It is a specific flowchart of step S102 of a method for segmenting an image shown in another exemplary embodiment of the present application;

[0037] Figure 5 It is a schematic diagram of yet another image processing result provided by the present application;

[0038] Figure 6 It is a schematic diagram of still another image processing result provided by the present application;

[0039] Figure 7 It is a schematic flowchart of a method for training a portrait segmentation model provided by the present application;

[0040] Figure 8 It is a schematic diagram of a device for segmenting an image provided by an embodiment of the present application;

[0041] Figure 9 It is a schematic diagram of a device for segmenting an image provided by another embodiment of the present application. Detailed Embodiments

[0042] In order to make the objectives, technical solutions, and advantages of this application clearer, the following further elaborates on this application in conjunction with the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain this application and are not used to limit this application.

[0043] As a common image processing method, human body image segmentation technology can be widely applied in various fields such as virtual reality, video surveillance, and behavior analysis. In recent years, due to its strong learning ability, wide coverage, and good adaptability, the deep learning method has achieved excellent performance in the image segmentation task.

[0044] In the prior art, when the segmentation model constructed based on deep learning processes an image, its processing process mainly classifies each pixel point in the image individually, without considering the structured human body feature information. In this case, missegmentation is very likely to occur, that is, the background in the image is missegmented as the human body, or the human body in the image is missegmented as the background, resulting in inaccurate image segmentation results.

[0045] In addition, when constructing the segmentation model, the labeling cost of the training images is high and time-consuming. For example, the time required to label a human body segmentation image is often about 10 times that of labeling a human body detection image. Therefore, the data for training the segmentation model is often scarce, resulting in high cost and long time consumption for constructing the segmentation model.

[0046] In view of this, the embodiment of this application provides a method for segmenting an image. First, obtain the image to be segmented; input the image to be segmented into a trained human portrait segmentation model for processing to obtain the segmentation result of the image to be segmented. The human portrait segmentation model is obtained by training a human body detection model using a first training set and a second training set. The first training set includes multiple first sample images and the corresponding first sample segmentation results of each first sample image. The second training set is obtained by processing the third training set used for training the human body detection model using the graph cut algorithm.

[0047] In this solution, the image to be segmented is processed by a trained human portrait segmentation model. The human portrait segmentation model is obtained by training a human body detection model using a first training set and a second training set. Since the human body detection model searches for the human body in the image in the form of a detection frame, structured human body feature information can be extracted. This enables the human portrait segmentation model to learn the structured human body feature information in the human body detection model during the training process. Furthermore, when the image to be segmented is processed by the trained human portrait segmentation model, the occurrence of missegmentation is effectively avoided, thereby improving the accuracy of the segmentation result.

[0048] Meanwhile, the second training set is obtained by processing the third training set used to train the human portrait detection model using the graph cut algorithm. The third training set used to train the human portrait detection model has low annotation cost, large quantity, and is easy to obtain. Processing the third training set using the graph cut algorithm to obtain the second training set required for training the human portrait segmentation model can reduce the cost of training the human portrait segmentation model, reduce the time for training the human portrait segmentation model, and improve the efficiency of training the human portrait segmentation model. Moreover, using a large amount of the second training set for training can improve the accuracy of the human portrait segmentation model.

[0049] Please refer to Figure 1 , Figure 1 which is a schematic flowchart of a method for segmenting an image provided by an exemplary embodiment of the present application. The execution subject of the method for segmenting an image provided by the present application is a device for segmenting an image. Among them, the device includes, but is not limited to, in-vehicle computers, tablet computers, computers, personal digital assistants (PDAs), and other devices, and may also include various types of servers. For example, the server can be an independent server or a cloud service that provides basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communications, middleware services, domain name services, security services, content delivery networks (CDNs), and big data and artificial intelligence platforms.

[0050] Such as Figure 1 The method for segmenting an image shown may include: S101 to S102, specifically as follows:

[0051] S101: Obtain the image to be segmented.

[0052] Exemplarily, the image to be segmented may contain one or more human bodies. It should be noted that the human body contained in the image to be segmented can be the whole body of the human body or a part of the human body. For example, it may include one or more parts of the human body such as the head, neck, shoulders, hands, chest and abdomen, legs, and feet. This is only an exemplary description and is not limited thereto.

[0053] When the device for segmenting an image detects an image segmentation instruction, it obtains the image to be segmented. The image segmentation instruction is used to instruct the device to process the image to be segmented. The image segmentation instruction can be triggered by the user, such as when the user clicks the image segmentation option in the device.

[0054] Obtaining the image to be segmented can be one or more images to be segmented uploaded by the user to the device, or the device can obtain the file corresponding to the file identifier according to the file identifier included in the image segmentation instruction and extract one or more images to be segmented from the file.

[0055] S102: Input the image to be segmented into the trained human portrait segmentation model for processing to obtain the segmentation result of the image to be segmented.

[0056] The trained human portrait segmentation model is obtained by training the human portrait detection model using the first training set and the second training set. Among them, the human portrait detection model is similar to the general human portrait detection model in the prior art. When processing an image, this human portrait detection model searches for the human body in the image in the form of a detection frame. Exemplarily, by processing the image to be segmented with the human portrait detection model, the human body in the image to be segmented can be framed with a rectangular box.

[0057] Please refer to Figure 2 , Figure 2 which is a schematic diagram of an image processing result provided by this application. Figure 2 The human body framed by the rectangular box shown in

[0058] is the result obtained by processing the image to be segmented with the human portrait detection model.

[0059] The first training set is the training set used when training the traditional human portrait segmentation model. This first training set can include multiple first sample images and the corresponding first sample segmentation results of each first sample image.

[0060] Among them, the first sample image is similar to the image to be segmented, that is, the first sample image can contain one or more human bodies. The human body contained in the first sample image can be the whole body of a human or a part of a human body. For example, it can include one or more parts of the head, neck, shoulders, hands, chest, abdomen, legs, feet, etc. of a human body. This is only for exemplary illustration and is not limited thereto.

[0061] The corresponding first sample segmentation result of each first sample image is obtained by manually annotating each first sample image. For example, for any first sample image, each pixel point of the image corresponding to the human body in the first sample image is marked as 0, so that the image corresponding to the human body in the first sample image is displayed in black. Each pixel point of the image corresponding to the remaining parts in the first sample image except this human body is marked as 255, so that the image corresponding to the remaining parts in the first sample image except the human body is displayed in white. After marking, the corresponding first sample segmentation result of this first sample image is obtained.

[0062] Please refer to Figure 3 , Figure 3It is a schematic diagram of another image processing result provided by this application. Figure 3 That is, it is the first sample segmentation result corresponding to a certain first sample image. Figure 3 The human body shown in black in it is the human body segmented from the first sample image, and the rest except the human body is white.

[0063] Since it takes a long time to annotate a human body segmentation image, in the implementation provided by this application, in order to reduce the cost and time of training the human portrait segmentation model, the number of training samples in the first training set is much smaller than the number of training samples in the second training set.

[0064] The second training set is obtained by processing the third training set used to train the human portrait detection model using the graph cut algorithm. The third training set includes multiple third sample images and the target detection boxes corresponding to each third sample image. The third training set is used to train the initial model to obtain a trained human portrait detection model.

[0065] The training samples in the third training set can be collected from the network or obtained by manual annotation. For example, for any third sample image, the area where the human body is located in the third sample image is marked with a rectangular box to obtain the target detection box corresponding to the third sample image. Since it takes a very short time to mark the detection box in the image and the training samples in the third training set can also be obtained from the network, a large number of training samples in the third training set can be obtained.

[0066] By processing the third training set using the graph cut algorithm, a large number of training samples in the second training set can be obtained. The graph cut algorithm can include algorithms such as Graph Cut algorithm, Grab Cut algorithm, onecut algorithm, etc. This graph cut algorithm only needs to obtain the target detection box, and it can classify the foreground (such as the human body) and background (other parts except the human body) in the target detection box, so as to obtain a rough annotation of the human body segmentation image. Therefore, as long as the number of training samples in the third training set is large enough, a large number of training samples in the second training set can be obtained.

[0067] Using the first training set and the second training set to train the human portrait detection model, although the number of training samples in the first training set is less than the number of training samples in the second training set, the training samples in the first training set are obtained by manual annotation and have high accuracy, which is conducive to monitoring the human portrait detection model during the training process. Although the accuracy of the training samples in the second training set is not as high as that of the training samples in the first training set, the number of training samples in the second training set is much larger than the number of training samples in the first training set, which ensures the quantity and diversity of the training samples for training the human portrait detection model, thereby improving the accuracy of the human portrait detection model.

[0068] In this embodiment, the human segmentation model can be pre-trained by this device, or can be pre-trained by other devices and then the file corresponding to the human segmentation model is transplanted into this device. That is to say, the execution entity for training the human segmentation model and the execution entity for performing image segmentation using the human segmentation model can be the same or different. For example, when using other devices to train the human segmentation model, after other devices finish training the human segmentation model, fix the model parameters of the human segmentation model to obtain the file corresponding to the trained human segmentation model, and then transplant this file into this device.

[0069] Input the image to be segmented into the trained human segmentation model for processing to obtain the segmentation result of the image to be segmented.

[0070] In this embodiment, the image to be segmented is processed by the trained human segmentation model. The human segmentation model is obtained by training the human detection model using the first training set and the second training set. Since the human detection model searches for humans in the image in the form of detection frames, structured human feature information can be extracted, which enables the human segmentation model to learn the structured human feature information in the human detection model during the training process. Furthermore, when the trained human segmentation model processes the image to be segmented, the occurrence of mis-segmentation is effectively avoided, thereby improving the accuracy of the segmentation result.

[0071] At the same time, the second training set is obtained by processing the third training set used to train the human detection model using the graph cut algorithm. The third training set used to train the human detection model has a low annotation cost and is easy to obtain. Processing the third training set using the graph cut algorithm to obtain the second training set required for training the human segmentation model can reduce the cost of training the human segmentation model, reduce the time for training the human segmentation model, and improve the efficiency of training the human segmentation model. Moreover, using a large number of second training sets for training can improve the accuracy of the human segmentation model.

[0072] Please refer to Figure 4 , Figure 4 which is a specific flowchart of step S102 of a method for segmenting an image shown in another exemplary embodiment of the present application; optionally, in some possible implementation manners of the present application, the above S102 may include S1021 to S1022, specifically as follows:

[0073] S1021: Detect and process the image to be segmented through a human body image detection network to obtain a human detection frame.

[0074] The human portrait segmentation model may include a human body image detection network and a human body image segmentation network. The segmentation result may include a human body detection box and a human body image. Among them, the human body image detection network is used to perform detection processing on the image to be segmented, and obtain the human body detection box corresponding to the image to be segmented. The human body detection box is used to predict the area where the human body is located in the image to be segmented.

[0075] The function and structure of the human body image detection network are similar to those of the human portrait detection model. For example, the human body image detection network may include an encoder and a detection decoder. Among them, the encoder is composed of multiple convolutional layers and pooling layers, and the detection decoder also contains multiple convolutional layers.

[0076] Exemplarily, the encoder in the human body image detection network encodes the image to be segmented. For example, the image to be segmented is continuously convolved through multiple convolutional layers, and the convolution result is input into the pooling layer. The pooling layer performs partition sampling on the convolution result to obtain the high-dimensional feature corresponding to the image to be segmented. The high-dimensional feature may include high-resolution human body feature information and low-resolution human body feature information. Among them, the high-resolution human body feature information may include high-precision edge information, and the low-resolution human body feature information may include distinct semantic information.

[0077] The encoding result of the encoder is input into the detection decoder for decoding processing to obtain the human body detection box corresponding to the image to be segmented. For example, the high-dimensional feature is input into the detection decoder, and the convolutional layer in the detection decoder performs deconvolution operation and convolution operation on the high-dimensional feature to restore these high-dimensional features, so as to obtain the human body detection box corresponding to the image to be segmented.

[0078] It should be noted that the human body detection box output by the human body image detection network is represented in the form of pixel points and the coordinates corresponding to the pixel points. According to the pixel points and the coordinates corresponding to the pixel points, the human body detection box can be marked in the image to be segmented, and the Figure 2 image shown in can be obtained. This is only an exemplary illustration here, and the comparison is not limited.

[0079] S1022: Perform segmentation processing on the image to be segmented through the human body image segmentation network to obtain a human body image.

[0080] The human portrait segmentation model may include a human body image detection network and a human body image segmentation network. The segmentation result may include a human body detection box and a human body image. Among them, the human body image segmentation network is used to perform segmentation processing on the image to be segmented to obtain the human body image corresponding to the image to be segmented.

[0081] The human body detection box mainly uses a rectangular box to enclose the area where the human body is located in the image to be segmented, as Figure 2As shown, in addition to the human body, the rectangular box also includes a part of the background. The human body image is to accurately segment the human body in the image to be segmented, such as Figure 3 the image shown in

[0082] The human body image segmentation network may include an encoder and a segmentation decoder. The encoder of the human body image segmentation network and the encoder in the human body image detection network may be the same one, or two encoders with the same structure.

[0083] For example, when the encoder of the human body image segmentation network and the encoder in the human body image detection network are the same one, when this encoder works with the detection decoder, it constitutes the human body image detection network; when this encoder works with the segmentation decoder, it constitutes the human body image segmentation network.

[0084] This encoder is composed of multiple convolutional layers and pooling layers, and the segmentation decoder also contains multiple convolutional layers.

[0085] Exemplarily, when the encoder of the human body image segmentation network and the encoder in the human body image detection network are the same one, the encoding result of the encoder in the human body image detection network for the image to be segmented can be directly used. For example, obtain the high-dimensional features output by the encoder in the human body image detection network, and input these high-dimensional features into the segmentation decoder for decoding processing to obtain the human body image corresponding to the image to be segmented.

[0086] Optionally, when the encoder of the human body image segmentation network and the encoder in the human body image detection network are not the same one, since the encoder in the human body image segmentation network is the same as the encoder in the human body image detection network, the process of encoding the image to be segmented by the encoder in the human body image segmentation network is the same as the process of encoding the image to be segmented by the encoder in the human body image detection network. For example, continuously convolve the image to be segmented through multiple convolutional layers, input the convolution result into the pooling layer, and the pooling layer performs partition sampling on the convolution result to obtain the high-dimensional features corresponding to the image to be segmented.

[0087] Input the encoding result of the encoder into the segmentation decoder for decoding processing to obtain the human body image corresponding to the image to be segmented. For example, input the high-dimensional features into the segmentation decoder, and the convolutional layers in this segmentation decoder perform deconvolution operations and convolution operations on the high-dimensional features to restore these high-dimensional features, thereby obtaining the human body image corresponding to the image to be segmented.

[0088] Although the encoder of the human body image segmentation network is the same as that of the human body image detection network, the segmentation decoder in the human body image segmentation network is different from the detection decoder in the human body image detection network. During the training process of the portrait segmentation model, the segmentation decoder and the detection decoder learn different decoding capabilities. Therefore, a human body image can be obtained by decoding through the segmentation decoder, and a human body detection box can be obtained by decoding through the detection decoder.

[0089] It should be noted that in this application, in order to prominently display the segmented human body image, the human body image is shown in black, and the rest of the background is shown in white. In the actual implementation process, the human body image can be multiplied by the image to be segmented, so as to restore the black human body image to a color human body image. This is only an exemplary illustration and is not limited in comparison.

[0090] In this embodiment, since the portrait segmentation model is trained by using the first training set and the second training set for the portrait detection model, on the basis of retaining the function of the portrait detection model, it also learns the function of portrait segmentation. Therefore, the trained portrait segmentation model includes a human body image detection network and a human body image segmentation network. By processing the image to be segmented through this portrait segmentation model, two results, namely a human body detection box and a human body image, can be obtained simultaneously.

[0091] The two results output by the portrait segmentation model do not interfere with each other, and the segmentation result needed can be selected according to the actual situation. For example, when it is necessary to know the approximate position of the human body in the image to be segmented, the human body detection box output by the portrait segmentation model can be selected; when it is necessary to know the specific position of the human body in the image to be segmented, the human body image output by the portrait segmentation model can be selected. This is only an exemplary illustration and is not limited in comparison.

[0092] Optionally, in some possible implementation manners of this application, after step S1022, S1023 to S1024 may further be included, which are specifically as follows:

[0093] S1023: Use the human body detection box to inspect the human body image.

[0094] Exemplarily, determine the area corresponding to the human body detection box according to the human body detection box, and determine the area corresponding to the human body image according to the human body image. Compare the two areas to obtain an inspection result.

[0095] Please refer to Figure 5 , Figure 5 which is a schematic diagram of another image processing result provided by this application. As Figure 5 shown, Figure 5 the area enclosed by the rectangular box in Figure 5 is the area corresponding to the human body detection box, and Figure 5The human image in consists of two parts, one is the human body in the human detection frame, and the other is the black ellipse outside the human detection frame.

[0096] S1024: Determine the final segmentation result of the image to be segmented according to the inspection result.

[0097] Exemplarily, the test result may include two results. One is that the area corresponding to the human image exceeds the area corresponding to the human detection frame, and the other is that the area corresponding to the human image does not exceed the area corresponding to the human detection frame. The human image is processed differently according to different test results, thereby obtaining the final segmentation result of the image to be segmented.

[0098] For example, when the inspection result is that the area corresponding to the human body image does not exceed the area corresponding to the human body detection frame, the current human body image is determined as the final segmentation result of the image to be segmented.

[0099] See also Figure 6 , Figure 6 FIG. 1 is a schematic diagram of another image processing result provided by the present application. Figure 6 As shown, Figure 6 The area selected by the rectangular box in is the area corresponding to the human body detection box. Figure 6 The black image in is the human body image. It can be clearly seen that Figure 6 The area corresponding to the human image in does not exceed the area corresponding to the human detection frame. This proves that the segmented human image is very accurate, so there is no need to do additional processing on the current human image, and the current human image is directly used as the final segmentation result of the image to be segmented.

[0100] Optionally, in some possible implementations of the present application, when the inspection result is that the area corresponding to the human image exceeds the area corresponding to the human detection frame, the image exceeding the area corresponding to the human detection frame is deleted from the human image to obtain the final segmentation result.

[0101] like Figure 5 As shown, it can be clearly seen that Figure 5 The area corresponding to the human image in exceeds the area corresponding to the human detection frame, that is Figure 5 The black ellipse area in the figure exceeds the area corresponding to the human body detection frame. This proves that the human body image obtained by the current segmentation is not accurate enough and there is a mis-segmentation, that is, the background is mis-segmented as a human body. At this time, the image that exceeds the area corresponding to the human body detection frame is deleted from the human body image. For example, delete Figure 5 The image corresponding to the black ellipse in the figure is obtained to obtain the final segmentation result of the image to be segmented.

[0102] In this embodiment, the human body detection frame is used to inspect the human body image, and the current human body image is adjusted according to the inspection result, further improving the accuracy of the segmentation result and making the final segmentation result of the image to be segmented cleaner.

[0103] Please refer to Figure 7 , Figure 7 which is a schematic flowchart of a method for training a portrait segmentation model provided by this application. This method for training a portrait segmentation model can be executed before step S101. As Figure 7 shown, the method for training a portrait segmentation model may include: S201 - S202, specifically as follows:

[0104] S201: Use a third training set and a preset loss function to train an initial model to obtain a human body detection model.

[0105] Exemplarily, the initial model may include an encoder, a detection decoder, and a segmentation decoder. At this time, the encoder, detection decoder, and segmentation decoder in this initial model are all untrained.

[0106] Since the human body detection model needs to be trained first, only the encoder and the detection decoder are required in the human body detection model, and the segmentation decoder is not used. Therefore, when training the human body detection model, the training of the segmentation decoder can be turned off. In this way, by using the third training set and the preset loss function to train the initial model, a human body detection model can be obtained. Specifically, the training of the segmentation decoder can be turned off or the training of the segmentation decoder can be turned on by adjusting the parameters in the loss function.

[0107] The third training set may include multiple third sample images and the target detection frames corresponding to each third sample image. The preset loss function may be:

[0108]

[0109] where is used to represent the loss when training the detection decoder, is used to represent the loss when training the segmentation decoder.

[0110] Adjusting value can turn on the training of the segmentation decoder or turn off the training of the segmentation decoder. For example, by adjusting the value of λ such that is 0, the training of the detection decoder will not be affected, which is equivalent to turning off the training of the segmentation decoder at this time.

[0111] After the training of the closed segmentation decoder, the third sample image in the third training set is input into the current initial model for processing, and the initial model outputs the actual detection box corresponding to the third sample image. According to the current loss function (i.e., the loss function after adjusting the value), calculate the first loss value between the actual detection box corresponding to the third sample image and the target detection box corresponding to the third sample image.

[0112] Judge the magnitude relationship between the first loss value and the first preset threshold. When the device executing the training process (e.g., the device for segmenting images, or other devices) confirms that the current first loss value is greater than the first preset threshold, it is determined that the initial model being trained does not meet the requirements. At this time, adjust the model parameters (such as weight values) of the initial model being trained, and continue to train the initial model using the third training set.

[0113] When the device executing the training process (e.g., the device for segmenting images, or other devices) confirms that the current first loss value is less than or equal to the first preset threshold, it is determined that the initial model being trained meets the requirements. At this time, fix the model parameters in the initial model, and determine the initial model with fixed model parameters as the trained human portrait detection model.

[0114] S202: Use the first training set, the second training set, and the loss function to train the human portrait detection model to obtain a human portrait segmentation model.

[0115] The second training set includes multiple second sample images, the second sample segmentation results corresponding to each second sample image, and the sample detection results corresponding to each second sample image.

[0116] Optionally, before executing S202, use the graph cut algorithm to process the third training set to obtain the second training set. Specifically, according to the target detection box corresponding to each third sample image, determine the target image in each third sample image; use the graph cut algorithm to classify each target image to obtain the classification result corresponding to each target image; generate the second training set according to each classification result and each third sample image.

[0117] Exemplarily, for each third sample image, the image framed by the target detection box in its corresponding third sample image is used as the target image, and the graph cut algorithm is used to classify the target image into foreground (such as the human body) and background (other parts except the human body) to obtain the classification result corresponding to the target image. In the classification result, the foreground in the target image can be displayed in color, and the background in the target image can be displayed in black.

[0118] Among them, the graph cut algorithm may include the Graph Cut algorithm, the Grab Cut algorithm, the onecut algorithm, etc. The process of using the graph cut algorithm to classify the target image can refer to the prior art and will not be elaborated here. Since the graph cut algorithm is used to roughly classify the target image, the classification result is not very accurate and has a large amount of noise compared with the sample segmentation result manually labeled. For example, in the classification result, a small part of the background may be mis-segmented as the foreground, or a small part of the foreground may be mis-segmented as the background. However, the classification speed is very fast, so a large number of classification results can be obtained in a short time.

[0119] Take a third sample image, the target detection box corresponding to the third sample image, and the classification result corresponding to the third sample image as a set of training samples in the second training set. For example, the third sample image is equivalent to the second sample image, the target detection box corresponding to the third sample image is equivalent to the sample detection result corresponding to the second sample image, and the classification result corresponding to the third sample image is equivalent to the second sample segmentation result corresponding to the second sample image. Multiple sets of training samples in the second training set can be obtained through this method.

[0120] Optionally, in a possible implementation manner, to facilitate the subsequent training of the human portrait segmentation model, the classification result corresponding to the third sample image and the target detection box corresponding to the third sample image can be labeled in the third sample image, and the unlabeled third sample image and the labeled third sample image are used as a set of training samples in the second training set. Multiple sets of training samples in the second training set can be obtained through this method.

[0121] In the above implementation manner, the second training set is obtained by processing the third training set used to train the human portrait detection model with the graph cut algorithm. The third training set used to train the human portrait detection model has a low labeling cost, a large quantity, and is easy to obtain. Processing the third training set with the graph cut algorithm to obtain the second training set required for training the human portrait segmentation model can reduce the cost of training the human portrait segmentation model, reduce the time for training the human portrait segmentation model, and improve the efficiency of training the human portrait segmentation model. Moreover, using a large number of second training sets for training can improve the accuracy of the human portrait segmentation model.

[0122] Execute S202 after generating the second training set. Exemplarily, based on the human portrait detection model trained in S201, start the training of the segmentation decoder, and use the first training set, the second training set, and the loss function to train the human portrait detection model with the segmentation decoder enabled to obtain the human portrait segmentation model.

[0123] In the process of training the portrait segmentation model, since there are only multiple first sample images and the corresponding first sample segmentation results in the first training set, while the second training set not only contains multiple second sample images, the corresponding second sample segmentation results of each second sample image, but also contains the sample detection results corresponding to each second sample image.

[0124] That is to say, the first training set lacks the sample detection results corresponding to the sample images compared to the second training set. Therefore, when it is detected that the training set is the first training set, the first parameter in the loss function needs to be adjusted to turn off the training of the detection decoder.

[0125] Exemplarily, the first parameter can be By adjusting the value, the training of the detection decoder is turned on or off. For example, by adjusting the value of α, when α is 0, it makes the value 0, which has no impact on the training of the segmentation decoder. At this time, it is equivalent to turning off the training of the detection decoder.

[0126] After turning off the training of the detection decoder, the first sample images in the first training set are input into the current portrait detection model for processing, and the current portrait detection model outputs the actual segmentation results corresponding to the first sample images. According to the current loss function (i.e., the loss function after adjusting the value), the second loss value between the actual segmentation results corresponding to the first sample images and the first sample segmentation results corresponding to the first sample images is calculated.

[0127] The size of the second loss value and the second preset threshold are judged. When the device executing the training process (for example, the device for segmenting images, or other devices) confirms that the current second loss value is greater than the second preset threshold, it is determined that the currently training portrait detection model does not meet the requirements. At this time, the model parameters (such as weight values) of the currently training portrait detection model are adjusted, and the portrait detection model is continued to be trained using the first training set.

[0128] When the device executing the training process (for example, the device for segmenting images, or other devices) confirms that the current second loss value is less than or equal to the second preset threshold, it is determined that the currently training portrait detection model meets the requirements. At this time, the model parameters in the portrait detection model are fixed.

[0129] Since the second training set contains both the second sample segmentation results corresponding to the second sample images and the sample detection results corresponding to the second sample images, after fixing the model parameters in the portrait detection model, the training of the detection decoder is turned on, and the portrait detection model with fixed model parameters is continued to be trained using the second training set.

[0130] Input the second sample images in the second training set into the human portrait detection model with fixed model parameters for processing, and output the actual detection results corresponding to the second sample images and the segmentation results corresponding to the second sample images. Calculate the loss value between the actual detection results corresponding to the second sample images and the sample detection results corresponding to the second sample images according to the current loss function, and at the same time calculate the loss value between the segmentation results corresponding to the second sample images and the second sample segmentation results corresponding to the second sample images. Calculate the sum of the two loss values to obtain the third loss value.

[0131] Determine the magnitude relationship between the third loss value and the third preset threshold. When the device performing the training process (for example, the device for segmenting images, or other devices) confirms that the current third loss value is greater than the third preset threshold, it is determined that the human portrait detection model currently under training does not meet the requirements. At this time, adjust the model parameters (such as weight values) of the human portrait detection model under training, and continue to train the human portrait detection model using the second training set.

[0132] When the device performing the training process (for example, the device for segmenting images, or other devices) confirms that the current third loss value is less than or equal to the third preset threshold, it is determined that the human portrait detection model currently under training meets the requirements. At this time, fix the model parameters in the human portrait detection model, and determine the current human portrait detection model as the trained human portrait segmentation model.

[0133] In this embodiment, on the one hand, a large number of low-cost third training sets are used to generate the second training set to train the human portrait segmentation model, thereby reducing the cost of training the human portrait segmentation model, shortening the training time of the human portrait segmentation model, and improving the training efficiency of the human portrait segmentation model. On the other hand, the segmentation encoder is supervised by the detection encoder, enabling the trained human portrait segmentation model to pay more attention to structured human feature information, effectively avoiding the occurrence of missegmentation, and thus improving the accuracy of the segmentation results.

[0134] Optionally, in some possible implementation manners of the present application, since the second sample segmentation results in the second training set have more noise compared with the sample segmentation results manually labeled. To reduce the impact of this noise on the human portrait segmentation model, when it is detected that the training set is the second training set, adjust the second parameter in the loss function. Exemplarily, the second parameter can be λ. For example, when it is detected that the training set is the second training set, adjust the value of λ to 0.01.

[0135] Optionally, to further improve the accuracy of the human portrait segmentation model, when it is detected that the training set is the first training set, the second parameter in the loss function can also be adjusted. For example, when it is detected that the training set is the second training set, adjust the value of λ to 1. This is only for illustrative purposes and is not limited by comparison.

[0136] In this embodiment, by adjusting the value of the second parameter in the loss function, the influence of noise on the human portrait segmentation model is effectively reduced, thereby improving the accuracy of the human portrait segmentation model. When using the trained human portrait segmentation model to process the image to be segmented, the occurrence of mis-segmentation is effectively avoided, thereby improving the accuracy of the segmentation result.

[0137] Please refer to Figure 8 , Figure 8 which is a schematic diagram of a device for segmenting an image provided by an embodiment of the present application. Each unit included in the device for segmenting an image is used to execute Figure 1 , Figure 4 , Figure 7 the respective steps in the corresponding embodiments. For details, please refer to Figure 1 , Figure 4 , Figure 7 the relevant descriptions in their respective corresponding embodiments. For the sake of convenience of description, only the parts related to this embodiment are shown. Refer to Figure 8 , including:

[0138] An acquisition unit 310, configured to acquire an image to be segmented;

[0139] A processing unit 320, configured to input the image to be segmented into a trained human portrait segmentation model for processing, to obtain a segmentation result of the image to be segmented. The human portrait segmentation model is obtained by training a human portrait detection model using a first training set and a second training set. The first training set includes a plurality of first sample images and first sample segmentation results corresponding to each of the first sample images. The second training set is obtained by processing a third training set used for training the human portrait detection model using a graph cut algorithm.

[0140] Optionally, the human portrait segmentation model includes a human body image detection network and a human body image segmentation network, and the segmentation result includes a human body detection frame and a human body image.

[0141] The processing unit 320 is specifically configured to:

[0142] Perform detection processing on the image to be segmented through the human body image detection network to obtain the human body detection frame, and the human body detection frame is used to predict the area where the human body is located in the image to be segmented;

[0143] Perform segmentation processing on the image to be segmented through the human body image segmentation network to obtain the human body image.

[0144] Optionally, the device further includes a verification unit, configured to:

[0145] Verify the human body image using the human body detection frame;

[0146] Determine the final segmentation result of the image to be segmented according to the inspection result.

[0147] Optionally, the inspection unit is further configured to: when the inspection result indicates that the area corresponding to the human body image exceeds the area corresponding to the human body detection frame, delete the image of the area exceeding the area corresponding to the human body detection frame in the human body image to obtain the final segmentation result.

[0148] Optionally, the device further includes a training unit, configured to:

[0149] Train the initial model using the third training set and a preset loss function to obtain the human portrait detection model, where the third training set includes a plurality of third sample images and the target detection frame corresponding to each of the third sample images;

[0150] Train the human portrait detection model using the first training set, the second training set, and the loss function to obtain the human portrait segmentation model.

[0151] Optionally, the device further includes a generation unit, configured to:

[0152] Determine a target image in each of the third sample images according to the target detection frame corresponding to each of the third sample images;

[0153] Classify each of the target images using the graph cut algorithm to obtain the classification result corresponding to each of the target images;

[0154] Generate the second training set according to each of the classification results and each of the third sample images.

[0155] Optionally, the training unit is further configured to: during the process of training the human portrait segmentation model, when it is detected that the training set is the first training set, adjust the first parameter in the loss function; and / or, when it is detected that the training set is the second training set, adjust the second parameter in the loss function.

[0156] Please refer to Figure 9 , Figure 9 which is a schematic diagram of a device for segmenting an image provided in another embodiment of the present application. As Figure 9 shown, the device 4 for segmenting an image in this embodiment includes: a processor 40, a memory 41, and a computer program 42 stored in the memory 41 and executable on the processor 40. When the processor 40 executes the computer program 42, the steps in the method embodiments for segmenting an image described above are implemented, such as Figure 1 S101 to S102 shown. Or, when the processor 40 executes the computer program 42, the functions of each unit in the above embodiments are implemented, such asFigure 8 Functions of the units 310 to 320 shown

[0157] Exemplarily, the computer program 42 may be divided into one or more units. The one or more units are stored in the memory 41 and executed by the processor 40 to complete this application. The one or more units may be a series of computer instruction segments capable of performing specific functions, and these instruction segments are used to describe the execution process of the computer program 42 in the device 4 for splitting images. For example, the computer program 42 may be divided into an acquisition unit and a processing unit, and the specific functions of each unit are as described above.

[0158] The device may include, but is not limited to, a processor 40 and a memory 41. Those skilled in the art can understand that Figure 9 This is merely an example of the device 4 for splitting images and does not constitute a limitation on the device. It may include more or fewer components than shown in the figure, or combine certain components, or different components. For example, the device may also include input / output devices, network access devices, buses, etc.

[0159] The so-called processor 40 may be a central processing unit (CPU), or may also be other general-purpose processors, digital signal processors (DSPs), application specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. The general-purpose processor may be a microprocessor or the processor may also be any conventional processor, etc.

[0160] The memory 41 may be an internal storage unit of the device, such as the hard disk or memory of the device. The memory 41 may also be an external storage terminal of the device, such as a plug-in hard disk, a smart media card (SMC), a secure digital (SD) card, a flash card, etc. equipped on the device. Further, the memory 41 may also include both the internal storage unit and the external storage terminal of the device. The memory 41 is used to store the computer instructions and other programs and data required by the terminal. The memory 41 may also be used to temporarily store the data that has been output or will be output.

[0161] The embodiments of the present application also provide a computer storage medium, which can be non-volatile or volatile. The computer storage medium stores a computer program, and when the computer program is executed by a processor, the steps in the method embodiments of the above-mentioned various image segmentation are implemented.

[0162] The present application also provides a computer program product. When the computer program product runs on a device, the device is enabled to execute the steps in the method embodiments of the above-mentioned various image segmentation.

[0163] The embodiments of the present application also provide a chip or integrated circuit, which includes: a processor for calling and running a computer program from a memory, so that a device installed with the chip or integrated circuit executes the steps in the method embodiments of the above-mentioned various image segmentation.

[0164] Those skilled in the art can clearly understand that, for the convenience and conciseness of description, only the above-mentioned division of each functional unit and module is used as an example. In practical applications, the above-mentioned functions can be allocated to different functional units and modules according to needs, that is, the internal structure of the device is divided into different functional units or modules to complete all or part of the functions described above. Each functional unit and module in the embodiments can be integrated into one processing unit, or each unit can exist physically alone, or two or more units can be integrated into one unit. The above-mentioned integrated unit can be implemented in the form of hardware or in the form of a software functional unit. In addition, the specific names of the functional units and modules are only for the convenience of mutual distinction and do not limit the protection scope of the present application. The specific working processes of the units and modules in the above system can refer to the corresponding processes in the foregoing method embodiments and will not be elaborated here.

[0165] In the above embodiments, the descriptions of the various embodiments have their own emphases. For the parts not detailed or recorded in a certain embodiment, reference can be made to the relevant descriptions of other embodiments.

[0166] Those of ordinary skill in the art can realize that the units and algorithm steps of the examples described in combination with the embodiments disclosed herein can be implemented by electronic hardware, or by a combination of computer software and electronic hardware. Whether these functions are executed in a hardware or software manner depends on the specific application and design constraints of the technical solution. Those skilled in the art can use different methods for each specific application to implement the described functions, but such implementation should not be considered to exceed the scope of the present application.

[0167] The embodiments described above are only used to illustrate the technical solutions of the present application, rather than to limit it; although the present application has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that: they can still modify the technical solutions recorded in the foregoing embodiments, or perform equivalent replacements on some of the technical features; and these modifications or replacements do not cause the essence of the corresponding technical solutions to deviate from the spirit scope of the technical solutions of the various embodiments of the present application, and should all be included within the protection scope of the present application.

Claims

1. A method for segmenting an image, characterized in that, Including: Obtain the image to be segmented; Perform detection processing on the image to be segmented through the human body image detection network in the trained human portrait segmentation model to obtain a human body detection frame; perform segmentation processing on the image to be segmented through the human body image segmentation network in the human portrait segmentation model to obtain a human body image; Use the human body detection frame to inspect the human body image; When the inspection result is that the area corresponding to the human body image does not exceed the area corresponding to the human body detection frame, determine the human body image as the final segmentation result of the image to be segmented; When the inspection result is that the area corresponding to the human body image exceeds the area corresponding to the human body detection frame, delete the image of the area exceeding the area corresponding to the human body detection frame in the human body image to obtain the final segmentation result; The human portrait segmentation model is obtained by training a human portrait detection model using a first training set and a second training set. The first training set includes a plurality of first sample images and first sample segmentation results corresponding to each of the first sample images. The second training set is obtained by processing a third training set used for training the human portrait detection model using a graph cut algorithm.

2. The method according to claim 1, wherein The second training set includes a plurality of second sample images, second sample segmentation results corresponding to each of the second sample images, and sample detection results corresponding to each of the second sample images. Before obtaining the image to be segmented, the method further includes: Train an initial model using the third training set and a preset loss function to obtain the human portrait detection model. The third training set includes a plurality of third sample images and target detection frames corresponding to each of the third sample images; Train the human portrait detection model using the first training set, the second training set, and the loss function to obtain the human portrait segmentation model.

3. The method according to claim 2, wherein Before training the human portrait detection model using the first training set, the second training set, and the loss function to obtain the human portrait segmentation model, the method further includes: Determine a target image in each of the third sample images according to the target detection frame corresponding to each of the third sample images; Classify each of the target images using the graph cut algorithm to obtain a classification result corresponding to each of the target images; Generate the second training set according to the respective classification results and the respective third sample images.

4. The method according to claim 2, wherein During the process of training the human portrait segmentation model, when it is detected that the training set is the first training set, adjust the first parameter in the loss function; and / or when it is detected that the training set is the second training set, adjust the second parameter in the loss function.

5. An apparatus for segmenting an image, characterized in that, Including: An acquisition unit for acquiring the image to be segmented; A processing unit for performing detection processing on the image to be segmented through the human body image detection network in the trained human portrait segmentation model to obtain a human body detection frame; performing segmentation processing on the image to be segmented through the human body image segmentation network in the human portrait segmentation model to obtain a human body image; Use the human body detection frame to inspect the human body image; when the inspection result is that the area corresponding to the human body image does not exceed the area corresponding to the human body detection frame, determine the human body image as the final segmentation result of the image to be segmented; When the inspection result is that the area corresponding to the human body image exceeds the area corresponding to the human body detection frame, delete the image of the area exceeding the area corresponding to the human body detection frame in the human body image to obtain the final segmentation result; The portrait segmentation model is obtained by training a portrait detection model using a first training set and a second training set. The first training set includes a plurality of first sample images and first sample segmentation results corresponding to each of the first sample images. The second training set is obtained by processing a third training set used for training the portrait detection model using a graph cut algorithm.

6. An apparatus for segmenting an image, comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein, When the processor executes the computer program, it implements the method according to any one of claims 1 to 4.

7. A computer-readable storage medium storing a computer program, characterized in that, When the computer program is executed by a processor, it implements the method according to any one of claims 1 to 4.

Citation Information

Patent Citations

  • Multi-modal medical image segmentation method and terminal equipment

    CN111145147A

  • Image segmentation method, and training method and device of image segmentation model

    CN112489063A

  • Image recognition method, device, electronic equipment and storage medium

    CN113011409A

  • Portrait segmentation model training method, storage medium and terminal equipment

    CN113256643A