Image Segmentation Method, Apparatus, Device, and Storage Medium

By performing image segmentation model processing and smoothing processing on each frame of the online video data, the problem that the prior art cannot accurately segment online video data is solved, real-time and accurate image segmentation is realized, and it is suitable for complex images.

CN115349139BActive Publication Date: 2025-06-20GUANGZHOU SHIYUAN ELECTRONICS CO LTD
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202080099096.5
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2020-12-21
Publication Date
2025-06-20
Estimated Expiration
2040-12-21

AI Technical Summary

Technical Problem

The prior art cannot accurately segment online video data in real time, simple and accurate image.

Method used

By acquiring each frame image in the video data and inputting it to the trained image segmentation model, the first segmented image is obtained and then smoothing is performed to obtain the second segmented image.

Benefits of technology

Real-time and accurate image segmentation of online video data is realized, suitable for complex images, without the need for artificial prior information, simplifying the complexity of image segmentation.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115349139B_ABST
    Figure CN115349139B_ABST
Patent Text Reader

Abstract

An embodiment of the present application provides an image segmentation method, apparatus, device, and storage medium, which relate to the technical field of image processing. The method includes: obtaining a current frame image in video data, where a target object is displayed in the video data; inputting the current frame image into a trained image segmentation model to obtain a first segmentation image based on the target object; performing smoothing processing on the first segmentation image to obtain a second segmentation image based on the target object; taking the next frame image in the video data as the current frame image, and returning to execute the operation of inputting the current frame image into the trained image segmentation model until corresponding second segmentation images are obtained for each frame image in the video data. The above method can solve the technical problem that the prior art cannot accurately perform image segmentation on online video data.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The embodiments of the present application relate to the technical field of image processing, and in particular, to an image segmentation method, apparatus, device, and storage medium. Background Art

[0002] Image segmentation is one of the common techniques in image processing. It is used to accurately extract the region of interest in the image to be processed and use the region of interest as the target region image for subsequent processing of the target region image (such as background replacement, cropping the target region image, etc.). Image segmentation based on human figures is an important application in the field of image segmentation. Image segmentation based on human figures means accurately separating the human figure region and the background region in the image to be processed. Currently, with the development of computer and network technologies, it is of great significance to perform image segmentation based on human figures on online video data. For example, in scenarios such as online meetings or online live broadcasts, image segmentation is performed on the online video data to accurately separate the human figure region and the background region in the video data. Then, the background image of the background region is replaced to achieve the purpose of protecting user privacy.

[0003] In the process of implementing the present application, the inventors found that some image segmentation techniques have the following defects: Image segmentation mainly includes methods based on thresholds, regions, edges, and graph theory and energy functionals. Among them, the method based on thresholds needs to perform segmentation according to the gray-scale features in the image, and its defect is that it is only applicable to images where the gray-scale value of the human figure region is uniformly distributed outside the gray-scale value of the background region. The region-based method segments the image into different regions according to the similarity criterion of the spatial neighborhood, and its defect is that it cannot process complex images. The edge-based method mainly uses the discontinuity of the local features of the image (such as the pixel mutation at the edge of the face) to obtain the boundary of the human figure region, and its defect is that the computational complexity is high. The method based on graph theory and energy functional mainly uses the energy functional of the image for human figure segmentation, and its defect is that the amount of calculation is huge and artificial prior information is required. Due to the defects of the above technologies, they are not applicable to the scenario of real-time, simple, and accurate image segmentation of online video data.

[0004] In summary, how to perform image segmentation on any online video data in real time, simply, and accurately has become a technical problem that needs to be solved urgently. Summary of the Invention

[0005] The embodiments of the present application provide an image segmentation method, apparatus, device, and storage medium to solve the technical problem that the above technologies cannot accurately perform image segmentation on online video data.

[0006] In a first aspect, the embodiments of the present application provide an image segmentation method, including:

[0007] Obtain the current frame image in the video data, where the target object is displayed in the video data;

[0008] Input the current frame image into the trained image segmentation model to obtain a first segmentation image based on the target object;

[0009] Perform smoothing processing on the first segmentation image to obtain a second segmentation image based on the target object;

[0010] Take the next frame image in the video data as the current frame image, and return to execute the operation of inputting the current frame image into the trained image segmentation model until a corresponding second segmentation image is obtained for each frame image in the video data.

[0011] In a second aspect, an embodiment of the present application further provides an image segmentation device, including:

[0012] A data acquisition module, configured to obtain the current frame image in the video data, where the target object is displayed in the video data;

[0013] A first segmentation module, configured to input the current frame image into the trained image segmentation model to obtain a first segmentation image based on the target object;

[0014] A second segmentation module, configured to perform smoothing processing on the first segmentation image to obtain a second segmentation image based on the target object;

[0015] A repeated segmentation module, configured to take the next frame image in the video data as the current frame image, and return to execute the operation of inputting the current frame image into the trained image segmentation model until a corresponding second segmentation image is obtained for each frame image in the video data.

[0016] In a third aspect, an embodiment of the present application further provides an image segmentation device, including:

[0017] One or more processors;

[0018] A memory, configured to store one or more programs;

[0019] When the one or more programs are executed by the one or more processors, the one or more processors implement the image segmentation method as described in the first aspect.

[0020] In a fourth aspect, an embodiment of the present application further provides a computer-readable storage medium, on which a computer program is stored, and when the program is executed by a processor, the image segmentation method as described in the first aspect is implemented.

[0021] The above image segmentation method, apparatus, device, and storage medium solve the technical problem that some image segmentation technologies cannot accurately segment online video data by means of obtaining video data containing a target object, inputting each frame image of the video data into an image segmentation model to obtain a corresponding first segmentation image, and then performing smoothing processing on the first segmentation image to obtain a second segmentation image. By using an image segmentation model based on an autoencoder and smoothing processing, online video data can be segmented in real time and accurately. Moreover, due to the self-learning nature of the image segmentation model, it can be applied to online video data with complex images. During the application process, only the image segmentation model needs to be deployed and can be directly applied without prior human information, simplifying the complexity of image segmentation and expanding the application scenarios of the image segmentation method. Description of the Drawings

[0022] Figure 1 It is a flowchart of an image segmentation method provided by an embodiment of the present application;

[0023] Figure 2 It is a flowchart of another image segmentation method provided by an embodiment of the present application;

[0024] Figure 3 It is a schematic structural diagram of an image segmentation model provided by an embodiment of the present application;

[0025] Figure 4 It is a schematic diagram of an original image provided by an embodiment of the present application;

[0026] Figure 5 It is a schematic diagram of a segmentation result image provided by an embodiment of the present application;

[0027] Figure 6 It is a schematic diagram of an edge result image provided by an embodiment of the present application;

[0028] Figure 7 It is a schematic structural diagram of another image segmentation model provided by an embodiment of the present application;

[0029] Figure 8 It is a schematic structural diagram of an image segmentation apparatus provided by an embodiment of the present application;

[0030] Figure 9 It is a schematic structural diagram of an image segmentation device provided by an embodiment of the present application. Detailed Embodiments

[0031] The present application will be further described in detail below with reference to the drawings and embodiments. It can be understood that the specific embodiments described herein are used to explain the present application rather than to limit the present application. Additionally, it should be noted that for the sake of description, only parts related to the present application are shown in the drawings rather than all structures.

[0032] It should be noted that in this text, relational terms such as first and second are only used to distinguish one entity or operation or object from another entity or operation or object, and do not necessarily require or imply any such actual relationship or order between these entities or operations or objects. For example, the "first" and "second" in the first segmented image and the second segmented image are used to distinguish two different segmented images.

[0033] The image segmentation method provided by the embodiments of this application can be executed by an image segmentation device. The image segmentation device can be implemented in software and / or hardware. The image segmentation device can be composed of two or more physical entities or one physical entity. For example, the image segmentation device can be an intelligent device with data operation and analysis capabilities such as a computer, a mobile phone, a tablet, or an interactive intelligent tablet.

[0034] Figure 1 It is a flowchart of an image segmentation method provided by the embodiments of this application. Refer to Figure 1 , the image segmentation method specifically includes:

[0035] Step 110, obtain the current frame image in the video data, and a target object is displayed in the video data.

[0036] The video data is the video data that needs to be segmented currently, and it can be online video data or offline video data. The video data contains multiple frames of images, and each frame of image displays a target object. The target object can be considered as an object that needs to be separated from the background image. Optionally, the background images of each frame of image in the video data can be the same or different, and the embodiments do not limit this. And the target object can change as the video data is played, but during the change process, the type of the target object remains unchanged. For example, when the target object is a human, the human images in the video data can change (such as changing people or adding new people, etc.), but the target object in the video data is always a human. In the following embodiments, the target object is taken as an example of a human. Optionally, the source of the video data is not limited in the embodiments. For example, the video data is a video captured by an image acquisition device (such as a camera, a camera, etc.) connected to the image segmentation device. Another example is that the video data is a conference screen obtained from the network in a video conference scenario. Another example is that the video data is a live broadcast screen obtained from the network in a live broadcast scenario.

[0037] Exemplarily, image segmentation of video data refers to separating the region where the target object is located in each frame of the video data. In the embodiment, the target object is exemplified as a human. Exemplarily, the processing of video data is performed on a frame-by-frame basis, that is, the images in the video data are obtained frame by frame, and the images are processed to obtain the final image segmentation result. In the embodiment, the currently processed image is denoted as the current frame image, and the description is made by taking the processing of the current frame image as an example.

[0038] Step 120: Input the current frame image into the trained image segmentation model to obtain a first segmentation image based on the target object.

[0039] The image segmentation model is a pre-trained neural network model, which is used to segment the target object in the current frame image and output the segmentation result corresponding to the current frame image. In the embodiment, the segmentation result is denoted as the first segmentation image. Through the first segmentation image, the portrait area and the background area of the current frame image can be determined. Among them, the portrait area can be considered as the area where the target object (human) is located. In one embodiment, the first segmentation image is a binary image, and its pixel values include two types: 0 and 1. Among them, the area with a pixel value of 0 belongs to the background area of the current frame image, and the area with a pixel value of 1 belongs to the portrait area of the current frame image. It can be understood that for the convenience of visual display of the first segmentation image, the pixel values are converted into two types: 0 and 255 before displaying the first segmentation image. Among them, the area with a pixel value of 0 belongs to the background area, and the area with a pixel value of 255 belongs to the portrait area. The resolution of the first segmentation image is the same as that of the current frame image. It can be understood that when the image segmentation model has a resolution requirement for the input image, that is, an image with a fixed resolution needs to be input, it is necessary to first determine whether the resolution of the current frame image meets the resolution requirement. If it does not meet the resolution requirement, the resolution of the current frame image is converted to obtain a current frame image that meets the resolution requirement. At this time, after obtaining the first segmentation image, the resolution of the first segmentation image is also converted so that the resolution of the first segmentation image is the same as that of the original current frame image (that is, the current frame image before resolution conversion). When the image segmentation model has no resolution requirement for the input image, the current frame image can be directly input into the image segmentation model to obtain a first segmentation image with the same resolution.

[0040] Exemplarily, the structure and parameters of the image segmentation model can be set according to the actual situation. In the embodiment, the image segmentation model adopts an autoencoder structure. Among them, an autoencoder is a type of artificial neural network used in semi-supervised learning and unsupervised learning, and its function is to perform representation learning on the input information by taking the input information as the learning target. The autoencoder includes two parts: an encoder and a decoder. Among them, the encoder is used to extract features in the image, and the decoder is used to decode the extracted features to obtain the learning result (such as the first segmented image in the embodiment). Optionally, the encoder adopts a lightweight network to reduce the amount of data processing and calculation when extracting features and speed up the processing speed. The decoder can be implemented by processes such as combining residual blocks with channel shuffling and upsampling to achieve fully automatic real-time segmentation of images. In the embodiment, features at different resolutions of the current frame image can be extracted through the encoder, and then, through the decoder, operations such as upsampling, fusion, and decoding are performed on each feature to reuse each feature, thereby obtaining an accurate first segmented image.

[0041] Optionally, the image segmentation model is deployed under a forward inference framework. The specific type of the forward inference framework can be set according to the actual situation. For example, the forward inference framework is the openvino framework. Among them, when deployed under the forward inference framework, the image segmentation model has a low dependence on the GPU and is relatively lightweight, without occupying a large amount of storage space.

[0042] Step 130: Smooth the first segmented image to obtain a second segmented image based on the target object.

[0043] In the embodiment, there are edge jaggednesses of different degrees in the first segmented image. Among them, the edge jaggedness can be understood as the edge between the portrait area and the background area being jagged, which makes the separation between the portrait area and the background area appear too rigid. In the embodiment, in order to reduce the influence of the edge jaggedness, the first segmented image is smoothed, that is, the edge jaggedness in the first segmented image is smoothed to obtain a segmented image with smoother edges. In the embodiment, the segmented image after smoothing is denoted as the second segmented image. The second segmented image can also be regarded as the final segmentation result of the current frame image. It can be understood that the second segmented image is also a binary image, and its pixel values include two types: 0 and 1. Among them, the area with a pixel value of 0 belongs to the background area of the current frame image, and the area with a pixel value of 1 belongs to the portrait area of the current frame image.

[0044] Among them, the technical means adopted during the smoothing process can be set according to the actual situation. In the embodiment, the smoothing process is implemented by Gaussian smoothing filtering. Exemplarily, in Gaussian smoothing filtering, the Gaussian kernel function is used to process the first segmented image to obtain the second segmented image. Among them, the Gaussian kernel function is a commonly used kernel function. At this time, the smoothing process can be expressed as: S2 = S1 * G, where S2 represents the second segmented image, S1 represents the first segmented image, and G represents the Gaussian kernel function.

[0045] Step 140: Take the next frame image in the video data as the current frame image, and return to execute the operation of inputting the current frame image into the trained image segmentation model until the second segmented image corresponding to each frame image in the video data is obtained.

[0046] Exemplarily, after obtaining the second segmented image, it can be considered that the image segmentation of the current frame image has been completed. Therefore, the next frame image in the video data can be processed. Among them, the processing process is to take the next frame image as the current frame image, and repeat steps 110 - 130 to obtain the second segmented image of the current frame image again. After that, obtain the next frame image again and repeat the above process until the second segmented image corresponding to each frame image in the video data is obtained, so as to achieve image segmentation.

[0047] It can be understood that after obtaining the second segmented image, the current frame image can be processed according to the actual requirements. In the embodiment, taking background replacement as an example, at this time, after performing smoothing processing on the first segmented image to obtain the second segmented image based on the target object, it further includes: obtaining a target background image, where the target background image contains the target background; performing background replacement on the current frame image according to the target background image and the second segmented image to obtain a new current frame image.

[0048] Among them, the target background refers to the new background used after background replacement. The target background image refers to the image containing the target background. Optionally, the target background image and the second segmentation image have the same resolution. The target background image can be an image selected by the user of the image segmentation device or the default image of the image segmentation device. Exemplarily, after obtaining the target background image, the background of the current frame image is replaced to obtain the replaced image. In the embodiment, the image after background replacement is denoted as the new current frame image. Exemplarily, the background replacement method is as follows: determine the pixel points in the portrait area and the background area of the current frame image through the second segmentation image, and then retain the portrait area and replace the corresponding background area with the relevant target background in the target background image to obtain the new current frame image. The background replacement can be expressed as: I’ = I × S2 + (1 - S2) × B, where S2 represents the second segmentation image, I’ represents the new current frame image, I represents the current frame image, and B represents the target background image. In the above formula, the portrait area can be retained through I × S2 (that is, after multiplying the current frame image by the second segmentation image, the pixel points in the current frame image corresponding to the pixel points with a pixel value of 1 in the second segmentation image are retained), and the background area can be replaced through (1 - S2) × B (that is, after multiplying the target background image by the second segmentation image, the pixel points in the target background image corresponding to the pixel points with a pixel value of 0 in the second segmentation image are retained). It can be understood that when performing image segmentation on video data, after obtaining the second segmentation image of each frame, the new current frame image corresponding to the current frame image can be obtained through the second segmentation image. Then, each new current frame image can form new video data after background replacement.

[0049] It can be understood that in the embodiment, for the convenience of understanding the technical solution, it is defined that the video data contains the target object. In practical applications, the video data may not contain the target object. At this time, the first segmentation image obtained according to the above method is a segmentation image with all pixel values of 0.

[0050] As described above, by obtaining the video data containing the target object, inputting each frame image of the video data into the image segmentation model to obtain the corresponding first segmentation image, and then performing smoothing processing on the first segmentation image to obtain the second segmentation image, the technical problem that some image segmentation technologies cannot accurately segment online video data is solved. By using the image segmentation model of the autoencoder and the smoothing processing, the video data can be accurately segmented, especially for online video data, ensuring the processing speed of the online video data. And due to the self-learning nature of the image segmentation model, it can be applied to video data with complex images. And in the application process, only the image segmentation model needs to be deployed and can be directly applied without prior human information, simplifying the complexity of image segmentation and expanding the application scenario of the image segmentation method.

[0051] It can be understood that the above image segmentation method can be regarded as the application process of an image segmentation model. In practical applications, the performance of the image segmentation model can directly affect the result of image segmentation. Therefore, in addition to applying the image segmentation model, the training process of the image segmentation model is also an important link. Exemplarily, Figure 2 FIG. is a flowchart of another image segmentation method provided by an embodiment of the present application. This image segmentation method is an exemplary illustration of the training process of the image segmentation model on the basis of the above image segmentation method. Refer to Figure 2 and the image segmentation method specifically includes:

[0052] Step 210, obtain a training data set, where the training data set includes multiple original images.

[0053] Training data refers to the data that enables the image segmentation model to learn when training the image segmentation model. In the embodiment, the training data is in the form of images. Therefore, the training data is called the original image, and the original image and the video data contain the same type of target object. Exemplarily, the training data set refers to a data set containing a large number of original images. That is, during the training process, a large number of original images are selected from the training data set for the image segmentation model to learn, so as to improve the accuracy of the image segmentation model.

[0054] Exemplarily, the video data contains a large number of images. If the original images are collected according to the video data, it is necessary to collect images frame by frame based on the video data, which will consume a large amount of workload and production cost, and the collected original images will contain a large amount of duplicate content, which is not conducive to the training of the image segmentation model. Therefore, in the embodiment, the training data set is constructed by using independent original images instead of video data. At this time, the constructed training data set can include original images with different human poses in different scenarios. Among them, the scenario is preferably a natural scenario. That is, multiple natural scenarios are pre-selected, and multiple images containing humans are taken as original images by using an image acquisition device in each natural scenario, where the poses of humans in the multiple images are different. Optionally, in order to reduce the influence of the parameters of the image acquisition device during shooting (such as the position, aperture size, focusing degree, etc. of the image acquisition device) and the illumination in the natural environment on the performance of the image segmentation model, when constructing the training data set, multiple original images under different illuminations and different shooting parameters are collected in the same natural scenario and the same human pose, so as to ensure the performance of the image segmentation model when processing video data under different scenarios, different human poses, different illuminations, and different shooting parameters.

[0055] It can be understood that existing public image data sets can also be used as the training data set, such as using the public data set Supervisely as the training data set, or using the public data set EG1800 as the training data set.

[0056] Step 220: Construct a label dataset according to the training dataset. The label dataset includes multiple segmentation label images and multiple edge label images. One original image corresponds to one segmentation label image and one edge label image.

[0057] Exemplarily, the label data can be understood as reference data for determining whether the image segmentation model is accurate, which plays a supervisory role. If the result output by the image segmentation model is more similar to the corresponding label data, it indicates that the accuracy of the image segmentation model is higher, that is, the performance is better. Otherwise, it indicates that the accuracy of the image segmentation model is lower. It can be understood that the process of training the image segmentation model is a process of making the result output by the image segmentation model more similar to the corresponding label data.

[0058] In one embodiment, after the original image is input into the image segmentation model, the image segmentation model outputs a segmentation image and an edge image corresponding to the original image. Among them, the segmentation image is a binary image obtained by performing image segmentation on the target object in the original image. In the embodiment, the segmentation image output by the image segmentation model during the training process is denoted as the segmentation result image. The edge image is a binary image representing the edge between the portrait area and the background area in the original image. In the embodiment, the edge image output by the image segmentation model during the training process is denoted as the edge result image. In order to accurately train the image segmentation model, in the embodiment, according to the output result of the image segmentation module, the label data is set to include a segmentation label image and an edge label image. Among them, the segmentation label image corresponds to the segmentation result image and is used as a reference for the segmentation result image. The edge label image corresponds to the edge result image and is used as a reference for the edge result image. Each original image has a corresponding segmentation label image and edge label image, and the segmentation label images and edge label images together form the label dataset.

[0059] Exemplarily, both the edge label image and the segmentation label image can be obtained from the above-mentioned original images. For example, by using the method of manual annotation, the portrait area, background area, and edge area are marked in each original image, and then the edge label image and the segmentation label image are obtained according to the portrait area, background area, and edge area. Another example is to use the method of manual annotation to mark the portrait area and the background area in each original image, and then obtain the segmentation label image according to the portrait area and the background area, and obtain the edge label image according to the segmentation label image.

[0060] In the embodiment, the method of obtaining the segmentation label image by manual annotation and obtaining the edge label image from the segmentation label image is used for exemplary description. In this embodiment, step 220 includes steps 221 - 225:

[0061] Step 221: Obtain the annotation result for the original image.

[0062] The annotation result refers to the result obtained after marking the portrait area and the background area in the original image. In the embodiment, the annotation result is obtained by manual annotation, that is, the portrait area and the background area are manually marked in the above original image, and then the image segmentation device obtains the annotation result according to the marked portrait area and background area.

[0063] Step 222: Obtain a corresponding segmentation label image according to the annotation result.

[0064] Exemplarily, according to the annotation result, the pixel values of the pixel points included in the portrait area in the original image are changed to 255, and the pixel values of the pixel points included in the background area in the original image are changed to 0, so as to obtain the segmentation label image. It can be understood that the segmentation label image is a binary image.

[0065] Step 223: Perform an erosion operation on the segmentation label image to obtain an eroded image.

[0066] The erosion operation can be understood as reducing and refining the white area (i.e., the portrait area) with a pixel value of 255 in the segmentation label image. In the embodiment, the image obtained after performing the erosion operation on the segmentation label image is denoted as the eroded image. It can be understood that the number of pixel points occupied by the white area in the eroded image is less than the number of pixel points occupied by the white area in the segmentation label image, and the white area in the segmentation label image can completely cover the white area in the eroded image.

[0067] Step 224: Perform a Boolean operation on the segmentation label image and the eroded image to obtain an edge label image corresponding to the original image.

[0068] Boolean operations include union, intersection, and subtraction. The multiple objects for performing Boolean operations are the operands. In an embodiment, the operands include a segmentation label image and an erosion image, and more specifically, the white regions in the segmentation label image and the erosion image. The result obtained through the Boolean operation can be denoted as a Boolean object. In an embodiment, the Boolean object is an edge label image. Exemplarily, union means that the obtained Boolean object contains the volumes of the two operands. Since the white region in the segmentation label image can completely cover the white region in the erosion image, the Boolean object obtained by performing the union on the segmentation label image and the erosion image is the white region in the segmentation label image. Intersection means that the obtained Boolean object only contains the common volume of the two operands (i.e., only contains the overlapping positions). Since the white region in the segmentation label image can completely cover the white region in the erosion image, the Boolean object obtained by performing the intersection on the segmentation label image and the erosion image is the white region in the erosion image. Subtraction means that the Boolean object contains the volume of the operand from which the intersecting volume is subtracted. For example, the Boolean object obtained by performing the subtraction on the segmentation label image and the erosion image is the white region obtained by removing the white region corresponding to the erosion image from the white region of the segmentation label image. It can be understood that since the erosion image is an image obtained by shrinking the white region in the segmentation label image, the edges of the white regions in the erosion image and the segmentation label image are highly similar. Therefore, after performing the subtraction of the Boolean operation on the segmentation label image and the erosion image, a white region representing only the edge can be obtained, that is, an edge label image is obtained. At this time, the edge label image can be expressed as: GT edge = GT - GT erode , where GT edge represents the edge label image, GT represents the segmentation label image, and GT erode represents the erosion image. It can be understood that the edge label image is a binary image, and its resolution is equal to that of the segmentation label image.

[0069] Step 225: Obtain a label data set according to the segmentation label image and the edge label image.

[0070] After obtaining the segmentation label images and edge label images of the respective original images according to the above steps, a label data set is composed of the respective segmentation label icons and the respective edge label images. It can be understood that the segmentation label image and the edge label image can be regarded as Ground Truth, that is, the correct label.

[0071] Step 230: Train an image segmentation model according to the training data set and the label data set.

[0072] Exemplarily, an original image is input into an image segmentation model, and a loss function is constructed based on the result output by the image segmentation model and the corresponding label data in the label dataset. Then, the model parameters of the image segmentation model are updated according to the loss function. After that, another original image is input into the updated image segmentation model to construct the loss function again, and the model parameters of the image segmentation model are updated again according to the loss function. The above training process is repeated until the loss function converges. Among them, when the numerical values of the loss function calculated continuously for a certain number of times are within a set range, it can be considered that the loss function converges, and further, it is determined that the accuracy of the output result of the image segmentation model is stable. Therefore, it can be considered that the training of the image segmentation model is completed.

[0073] Exemplarily, the specific structure of the image segmentation model can be set according to the actual situation. In the embodiment, the image segmentation model includes: a normalization module, an encoding module, a channel shuffling module, a residual module, a multiple upsampling module, an output module, and an edge module as an example for description. For the convenience of understanding, the Figure 3 structure shown is used to exemplarily describe the image segmentation model. Among them, Figure 3 is a schematic structural diagram of an image segmentation model provided by an embodiment of the present application. Refer to Figure 3 , the image segmentation model includes a normalization module 21, an encoding module 22, four channel shuffling modules 23, three residual modules 24, four multiple upsampling modules 25, an output module 26, and an edge module 27. In this embodiment, step 230 includes steps 231 - step 2310:

[0074] Step 231: Input the original image into the normalization module to obtain a normalized image.

[0075] In the embodiment, the resolution of the original image is taken as 224×224 as an example for description. For example, Figure 4 is a schematic diagram of the original image provided by an embodiment of the present application. Refer to Figure 4 , the original image contains a portrait area. It should be noted that Figure 4 the original image used herein is from the public dataset Supervisely.

[0076] Exemplarily, normalization refers to a process of performing a series of standard processing transformations on an image to transform the image into a fixed standard form. At this time, the obtained standard image is called a normalized image. Normalization is divided into linear normalization and non-linear normalization. In the embodiment, the original image is processed in a linear normalization manner. Among them, linear normalization normalizes the pixel values in each image from [0, 255] to [-1, 1], and the resolution of the obtained normalized image is equal to the resolution of the image before linear normalization. It can be understood that the normalization module is a module for implementing linear normalization operations. After inputting the original image into the normalization module, the normalization module outputs a normalized image with pixel values of [-1, 1].

[0077] Step 232: Use the encoding module to obtain multi-layer image features of the normalized image, and the resolution of each layer of image features is different.

[0078] The encoding module is used to extract features in the normalized image. In the embodiment, the extracted features are denoted as image features. It can be understood that the image features can reflect information such as color features, texture features, shape features, and spatial relationship features in the normalized image, including global information and / or local information. Exemplarily, the encoding module is a lightweight network. Among them, the lightweight network refers to a neural network with a small number of parameters, a small amount of computation, and a short inference time. The type of lightweight network adopted by the encoding module can be selected according to the actual situation. In the embodiment, refer to Figure 3 , and take the encoding module 12 as the MobileNetV2 network as an example for description.

[0079] In one embodiment, after passing through MobileNetV2, the normalized image can output multi-layer image features. Among them, the resolutions of each layer of image features are different and there is a multiple relationship. Optionally, the resolutions of each layer of image features are all smaller than the resolution of the original image. In one embodiment, each layer of image features is arranged from top to bottom in the order of decreasing resolution, that is, the image feature with the highest resolution is located at the highest layer, and the image feature with the lowest resolution is located at the lowest layer. It can be understood that the number of layers of image features output by the encoding module can be set according to the actual situation. For example, when the resolution of the original image is 224×224, the encoding module outputs four layers of image features. At this time, refer to Figure 3 , among the four layers of image features output by the encoding module 22, the resolution of the highest layer (the first layer) image feature is 112×112( Figure 3 the image feature of this layer is denoted as Feature112×112 in), the resolution of the second highest layer (the second layer) image feature is 56×56( Figure 3 the image feature of this layer is denoted as Feature56×56 in), the resolution of the second lowest layer (the third layer) image feature is 28×28( Figure 3In this layer, the image feature is denoted as Feature28×28), and the resolution of the image feature of the lowest layer (the fourth layer) is 14×14( Figure 3 In this layer, the image feature is denoted as Feature14×14). It can be understood that the amount of information contained in each layer of image features increases from bottom to top. There is the same multiple relationship between the image features of adjacent layers, and the resolution of each layer of image features is less than that of the original image. It can be understood that the correspondence between the resolution and the layer in the embodiment is only for explaining the image segmentation model, rather than limiting the image segmentation model.

[0080] It should be noted that the number of channels contained in each layer of image features is not limited in the embodiment.

[0081] It can be understood that the encoding module can be regarded as the encoder in the image segmentation model.

[0082] Step 233: Input each layer of image features into the corresponding channel mixing module respectively to obtain multiple layers of mixed features, and each layer of image features corresponds to a channel mixing module.

[0083] The channel mixing module is used to fuse the features between channels in the layer to enrich the information contained in each layer of image features and ensure the accuracy of the image segmentation model without increasing the subsequent calculation amount. It can be understood that each layer of image features corresponds to a channel mixing module. For example, Figure 3 in the four layers of image features, there are four channel mixing modules 23, and each channel mixing module 23 is used to fuse the image features between multiple channels in the corresponding layer.

[0084] In one embodiment, the channel mixing module is composed of a 1×1 convolutional layer, a batch normalization (Batch Normalization, BN) layer, and an activation function layer. Among them, the Relu activation function is used in the activation function layer. Among them, the confusion of image features between channels is realized through the 1×1 convolutional layer, and the confused image features can be made more stable through the BN layer + activation function layer. It can be understood that the above structure of the channel mixing module is only an exemplary description. In actual applications, other structures can also be set for the channel mixing module.

[0085] Exemplarily, the feature output by the channel mixing module is denoted as the mixed feature. It can be understood that each layer of image features has a corresponding mixed feature, and the resolution of the mixed feature and the image feature in the same layer is the same. In one embodiment, except for the mixed feature with the lowest resolution, other mixed features are the central layer features, that is, other layers can be regarded as the network central layer. For Figure 3For example, after passing through their respective channel confusion modules 23, the confusion features of the lowest layer are represented as Decode 14×14, and the confusion features of other layers are represented as Center28×28, Center56×56, and Center112×112 respectively. Among them, the numerical part represents the resolution.

[0086] It can be understood that the confusion features output by the channel confusion module can also be considered as the features obtained after decoding the image features, that is, in addition to the confusion features, the channel confusion module can also play the role of decoding.

[0087] Step 234: Except for the confusion features of the highest-resolution layer, upsample the confusion features of each other layer and fuse them with the confusion features of the higher-resolution level to obtain the fused features corresponding to the higher-resolution level.

[0088] Upsampling can be understood as enlarging the features to increase the resolution of the features. In the embodiment, upsampling is implemented by the linear interpolation method, that is, appropriate interpolation algorithms are used to insert new elements between the confusion features to increase the resolution of the confusion features.

[0089] In this step, the resolution of the confusion features can be increased through upsampling so that the increased resolution is equal to the higher-resolution level. Among them, the higher-resolution level refers to the resolution that is higher than and only higher than the resolution currently being upsampled. At this time, the resolution being upsampled can be considered as the lower-resolution level of its higher-resolution level. For example, Figure 3Among them, except for the lowest layer, the resolution of each layer above is the next higher resolution than that of the layer below it. It can be understood that since the resolution of the confusion features of any layer and its next higher resolution are in a multiple relationship, therefore, the upsampling multiple can be determined according to this multiple. For example, if the resolution of the confusion features of a certain layer is 0.5 times that of the next higher resolution, then the resolution of the confusion features of this layer can be enlarged by means of double upsampling. Then, the confusion features of the next higher resolution are fused with the upsampled confusion features of the corresponding lower resolution through a skip connection to reuse the confusion features and ensure that more informative features are used in the subsequent processing. It can be understood that image segmentation is a type of dense prediction, so more informative features are required in the original image segmentation model. In the embodiment, the fused features are denoted as the fused features. At this time, except for the confusion features with the lowest resolution, each layer of confusion features has corresponding fused features. Among them, the operation of feature fusion can be understood as a concatenate (vector splicing) operation. It can be understood that the size of the fused features of each layer is the sum of the size of the confusion features of this layer and the size of the upsampled confusion features of the lower resolution. For example, if C in the [NCHW] before fusion of the confusion features of this layer is 3, and C in the [NCHW] before fusion of the upsampled confusion features of the lower resolution is 3, then C in the [NCHW] of the fused features after fusion is 6, and the values of N, H, and W remain unchanged. Among them, N is the quantity, C is the number of channels, H is the height, and W is the width. H×W can be understood as the resolution. It should be noted that since there is no next higher resolution for the highest resolution, there is no need to upsample the confusion features with the highest resolution.

[0090] For example, referring to Figure 3 the image segmentation model shown, the resolution of the confusion features Decode 14×14 of the lowest layer is doubled after double upsampling, that is, a feature with a resolution of 28×28 is obtained. Then, the confusion features Center28×28 of the next higher resolution (i.e., the second lowest layer) of the lowest layer are fused with the 28×28 feature after double upsampling of the lowest layer through a skip connection to obtain the fused features of the second lowest layer. Similarly, the resolution of the confusion features Center28×28 of the second lowest layer is doubled after double upsampling, that is, a feature with a resolution of 56×56 is obtained. Then, the confusion features Center56×56 of the next higher resolution (i.e., the second highest layer) of the second lowest layer are fused with the 56×56 feature after double upsampling of the second lowest layer through a skip connection to obtain the fused features of the second highest layer. Similarly, the fused features of the highest layer are obtained.

[0091] Step 235: Input each layer of fused features into the corresponding residual module to obtain multiple layers of first decoded features. Each layer of fused features corresponds to a residual module, and the lowest-resolution confused feature serves as the lowest-resolution first decoded feature.

[0092] The residual module is used to further extract and decode the fused features. The residual module can include one or more residual blocks (Residual Block, RS Block). In the embodiment, the case where the residual module includes one residual block is described as an example, and the structure of the residual block can be set according to the actual situation. It can be understood that each layer of fused features corresponds to a residual module, and the features output after the residual module processes have the same resolution as the fused features of that layer. Since the residual module can further extract and decode the fused features, that is, the features output by the residual module are decoded features, therefore, in the embodiment, the features output by the residual module are denoted as the first decoded features.

[0093] It can be understood that since there is no corresponding fused feature for the lowest-resolution confused feature, there is no need to set a residual module in the lowest-resolution layer. At this time, the lowest-resolution confused feature can be directly regarded as the first decoded feature of that layer. Correspondingly, after the fused features of other layers pass through the corresponding residual modules, the corresponding first decoded features can be obtained.

[0094] Take Figure 3 as an example. It includes 3 residual modules 24. The first decoded feature output after the fused features of the second-lowest layer are input into the residual module is denoted as RS Block28×28, that is, the resolution of the first decoded feature is 28×28. The first decoded feature output after the fused features of the second-highest layer are input into the residual module is denoted as RS Block56×56, that is, the resolution of the first decoded feature is 56×56. The first decoded feature output after the fused features of the highest layer are input into the residual module is denoted as RS Block112×112, that is, the resolution of the first decoded feature is 112×112. And the first decoded feature of the lowest layer is Decode14×14.

[0095] Step 236: Input each layer of first decoded features into the corresponding multi-fold upsampling module to obtain multiple layers of second decoded features. Each layer of first decoded features corresponds to a multi-fold upsampling module, and the resolutions of the second decoded features are the same as that of the original image.

[0096] Exemplarily, the multi-fold upsampling module is used to perform multi-fold upsampling on the first decoded feature so that the resolution after multi-fold upsampling is equal to the resolution of the original image. Among them, the specific multiple during multi-fold upsampling can be determined according to the resolution of the first decoded feature and the resolution of the original image. For example, if the resolution of the first decoded feature is 14×14 and the resolution of the original image is 224×224, then the first decoded feature needs to be upsampled 16 times to obtain a decoded feature with a resolution of 224×224.

[0097] It can be understood that for a binary image segmentation model, the final output binary image (segmentation result image) is used to distinguish the foreground (such as the portrait area) and the background. Therefore, the segmentation task of the image segmentation model belongs to a binary segmentation task. At this time, before obtaining the segmentation result image, a decoded feature with 2 channels needs to be obtained first. In the embodiment, in addition to performing multi-fold upsampling on the first decoded feature, the multi-fold upsampling module also needs to change the number of channels of the first decoded feature after multi-fold upsampling to 2. For each layer of the first decoded feature, only the resolution changes after multi-fold upsampling, and the number of channels does not change. Therefore, in the embodiment, a 1×1 convolutional layer is set in the multi-fold upsampling module, that is, a 1×1 convolutional layer is connected after performing multi-fold upsampling on the first decoded feature to change the number of channels of the first decoded feature after multi-fold upsampling to 2. In practical applications, the image segmentation model can also perform multi-class segmentation tasks. At this time, before obtaining the final output image, a decoded feature with the number of channels equal to the number of classifications also needs to be obtained. For example, if the image segmentation model performs a five-class segmentation task, then before obtaining the 5-class segmentation result image, a 5-channel decoded feature needs to be obtained first. It should be noted that when using the corresponding segmentation label image for supervision, in order to facilitate the calculation of the loss function, the pixel values of the pixel points in the segmentation label image need to be converted from 0 and 255 to 0 and 1, that is, the pixel points with a pixel value of 0 are converted to 0, and the pixel points with a pixel value of 255 are converted to 1. At this time, when training the segmentation network model, in order to make the image segmentation model finally output a decoded feature with 2 channels, the segmentation label image needs to change the Ground Truth into a one-hot encoding form, that is, each category has one channel, and the pixel points of each channel have a value of 1 when belonging to the current category, and the values of other channels are 0.

[0098] In the embodiment, for the convenience of description, the feature output by the multi-fold upsampling module is denoted as the second decoded feature. It can be understood that each layer of the first decoded feature corresponds to a multi-fold upsampling module, and through the multi-fold upsampling module, a second decoded feature with 2 channels and the same resolution as the original image resolution can be obtained. The second decoded feature can be considered as the network prediction output obtained after decoding the image features of the current layer.

[0099] For example, refer toFigure 3 After the four - layer first decoded features pass through the corresponding multiple up - sampling modules 25 respectively, four second decoded features with a resolution of 224×224 and 2 channels can be obtained. Figure 3 In [reference], the 4 second decoded features are all denoted as 224×224.

[0100] It can be understood that the second decoded feature of each layer can be regarded as the temporary output result obtained by decoding the image feature of that layer. The final segmentation result image can be obtained through the temporary output result.

[0101] Step 237: Combine the multi - layer second decoded features and input them into the output module to obtain the segmentation result image.

[0102] Since the image segmentation model finally needs to output a segmentation result image, after obtaining the second decoded features, the output module integrates the second decoded features of each layer to obtain a segmentation result image (i.e., a binary image). Exemplarily, first, the second decoded features of each layer are fused (i.e., Concatenate) to facilitate the output module to obtain richer features, so as to recover a more accurate image. Then, the output module uses the fused second decoded features to obtain the segmentation result image. At this time, the specific process of the output module is as follows: Connect the fused second decoded features to a 1×1 convolutional layer to obtain a decoded feature with 2 channels. It can be understood that the fused second decoded features are only combined together, and through the 1×1 convolutional layer in the decoding module, the fused second decoded features can be further decoded to output the final decoded feature with reference to the second decoded features of each layer. This decoded feature has 2 channels and is used to describe the result of binary classification, that is, to describe whether each pixel point in the original image is a portrait area or a background area. Then, after passing the decoded feature through the softmax function and the argmax function, the segmentation result image is obtained. That is, the output module consists of a 1×1 convolutional layer and an activation function layer. Among them, the activation function layer consists of a softmax function and an argmax function. Among them, the data processed by the softmax function can be understood as the output data of the logic layer, that is, to interpret the meaning represented by the decoded feature output by the 1×1 convolutional layer and obtain the description of the logic layer. When the label is in one - hot form, the argmax function is a common function to obtain the output result, that is, the corresponding segmentation result image is output by the argmax function.

[0103] For example, referring to Figure 3 , after fusing the four second decoded features and inputting them into the output module 26, at this time, first pass through a 1×1 convolutional layer to obtain a decoded feature with 2 channels ( Figure 3 denoted as Refine224×224 in [reference]), and then, pass through the activation function layer to obtain the segmentation result image ( Figure 3Denoted as output224×224 in the record.

[0104] It can be understood that the pixel values of each pixel point in the segmentation result image output by the image segmentation model are 0 or 1. Among them, the pixel points with a pixel value of 0 are the pixel points in the background area, and the pixel points with a pixel value of 1 are the pixel points in the portrait area. For the convenience of visualizing the segmentation result image, when displaying the segmentation result image, the pixel value of each pixel point is multiplied by 255. For example, Figure 5 This is a schematic diagram of the segmentation result image provided by the embodiment of the present application. Figure 4 The training data shown is input into Figure 3 the image segmentation model shown, and after obtaining the segmentation result image, by multiplying the pixel values of the segmentation result image by 255, the Figure 5 segmentation result image shown can be obtained.

[0105] Step 238: Input the first decoded feature with the highest resolution into the edge module to obtain an edge result image.

[0106] To improve the learning ability of the image segmentation model for the edge between the portrait area and the background area, in the embodiment, an edge module is set in the image segmentation model to perform additional supervision on the first decoded feature with the highest resolution through the edge module, that is, to play a role of regularization constraint to improve the ability of the image segmentation model to learn edges. The specific structural embodiment of the edge module is not limited. In the embodiment, an example of the edge module being a 1×1 convolutional layer is described. Exemplarily, after inputting the first decoded feature with the highest resolution into the edge module, an edge feature with 2 channels and the same resolution as the original image resolution can be obtained. Through this edge feature, a binary image representing only the edge can be obtained. In the embodiment, the binary image representing the edge is denoted as the edge result image. It can be understood that the pixel values of each pixel point in the edge result image are 0 or 1. Among them, the pixel points with a pixel value of 1 represent the pixel points where the edge is located, and the pixel points with a pixel value of 0 represent the pixel points where the edge is not located. It should be noted that the first decoded feature with the highest resolution has richer detailed information. Therefore, a more accurate edge feature can be obtained through the first decoded feature with the highest resolution.

[0107] For example, as Figure 3 shown, the first decoded feature RS Block112×112 of the highest layer passes through the edge module 27, and an edge feature with a resolution of 224×224 can be obtained, Figure 3 denoted as edge224×224 in the record.

[0108] For the convenience of visualizing the edge result image, when displaying the edge result image, the pixel value of each pixel point is multiplied by 255. For example, Figure 6Schematic diagram of the edge result image provided by the embodiment of the present application. After inputting the training data shown in Figure 4 into the image segmentation model shown in Figure 3 , an edge result image is obtained. After that, by multiplying each pixel value of the edge result image by 255, the edge result image shown in Figure 6 can be obtained.

[0109] It can be understood that except for the normalization module and the encoding module, other modules can be considered as the modules that make up the decoder.

[0110] Step 239: Construct a loss function according to each second decoding feature, the edge result image, the corresponding segmentation label image, and the edge label image, and update the model parameters of the image segmentation model according to the loss function.

[0111] The loss function of the segmentation network model is composed of a segmentation loss function and an edge loss function. Among them, the segmentation loss function can reflect the segmentation ability of the segmentation network model, and the segmentation loss function is obtained according to the second decoding features of each layer and the segmentation label image. At this time, a sub-loss function can be obtained based on the second decoding feature and the segmentation label image of each layer. After combining the sub-loss functions of each layer, the segmentation loss function can be obtained. It can be understood that the calculation methods of the sub-loss functions are the same. In one embodiment, the sub-loss function is calculated by the Iou function, and the Iou function can be defined as: the ratio of the area of the intersection of the predicted pixel region (i.e., the second decoding feature) and the label pixel region (i.e., the segmentation label image) to the area of the union, that is, the Iou function can reflect the overlapping similarity between the binary image corresponding to the second decoding feature and the segmentation label image. At this time, the sub-loss function calculated by the Iou function can reflect the loss of the overlapping similarity. Exemplarily, the edge loss function can reflect the ability of the segmentation network model to learn edges, and the edge loss function is obtained from the edge result image and the edge label image. In one embodiment, since the proportion of the pixel points of the edge in the pixel points of the entire original image is very low, the edge loss function adopts the Focal loss, and the Focal loss is a common loss function, which can reduce the weight of a large number of simple negative samples in the training, and can also be understood as a kind of hard sample mining.

[0112] Exemplarily, the loss function of the segmentation network model is expressed as:

[0113] Among them, Loss represents the loss function of the segmentation network model, n represents the total number of layers corresponding to the second decoding feature, represents the sub-loss function calculated according to the second decoding feature with the highest resolution and the corresponding segmentation label image, represents the sub-loss function calculated according to the second decoding feature with the lowest resolution and the segmentation label image, A n represents the second decoding feature with the lowest resolution, B represents the corresponding segmented label image, and Iou n represents the overlap similarity between A n and B, and loss edge is the Focal loss function.

[0114] Exemplarily, the image segmentation model has n layers (n≥2), that is, the second decoding feature has n layers. At this time, n sub-loss functions can be obtained according to the n-layer second decoding feature and the segmented label image. The first layer has the highest resolution, and its corresponding sub-loss function is denoted as The second layer has the second highest resolution, and its corresponding sub-loss function is denoted as And so on, the nth layer has the lowest resolution, and its corresponding sub-loss function is denoted as Since the calculation methods of the sub-loss functions are the same, in the embodiment, the sub-loss function of the nth layer is taken as an example for description. Exemplarily, that is represents the loss of the Iou function of the nth layer. A n represents the second decoding feature of the nth layer, B represents the corresponding segmented label image, and A n ∩B represents the intersection of A n and B, and A n ∪B represents the union of A n and B, and Iou n represents the overlap similarity between A n and B. At this time, represents the loss of the overlap similarity. It can be understood that the more similar the binary image corresponding to the second decoding feature is to the segmented label image, the smaller the corresponding sub-loss function is, the better the segmentation ability of the image segmentation model is, and the higher the segmentation accuracy is. Exemplarily, loss edge represents the edge loss function. In the embodiment, loss edge is the Focal loss function. loss edge (p t ) = -α t (1 - p t ) γ log(p t ). Among them, p t represents the predicted probability value that the pixel point in the edge result image is an edge, α t represents the balance weight coefficient, which is used to balance positive and negative samples, and γ represents the modulation coefficient, which is used to control the weights of difficult and easy classification samples. The values of α t and γ can be set according to the actual situation. According to the loss edge (pt ) The loss can be obtained edge , specifically, the loss of each pixel edge (p t ) is added and the mean value is calculated, and the calculated mean value is used as the loss edge .

[0115] After obtaining the loss function, the model parameters of the image segmentation model can be updated according to the loss function to make the performance of the updated image segmentation model higher.

[0116] Step 2310: Select the next original image, and return to perform the operation of inputting the original image into the normalization module until the loss function converges.

[0117] It can be understood that after modifying the model parameters of the image segmentation model through the loss function, it can be considered that one training ends. At this time, select another original image and the corresponding segmentation label image and edge label image to train the image segmentation model to calculate the loss function again and modify the model parameters according to the loss function. After multiple trainings, if the values of the loss function calculated in the current consecutive times are within the preset value range, it means that the loss function converges, that is, the image segmentation model is stable. It can be understood that the specific value of the preset value range can be set according to the actual situation.

[0118] After the image segmentation model is stable, it is determined that the training ends. After that, the image segmentation model can be applied to segment the portraits in the video data.

[0119] Based on the above embodiments, after training the image segmentation model according to the training data set and the label data set, it further includes: when the image segmentation model is not a network model recognizable by the forward inference framework, converting the image segmentation model into a network model recognizable by the forward inference framework.

[0120] The image segmentation model is trained in the corresponding framework, which is usually frameworks such as tensorflow and pytorch. In the embodiment, the pytorch framework is taken as an example for description. The pytorch framework is mainly used for the design, training and testing of the model. Since the image segmentation model runs in real time in the image segmentation device, and the pytorch framework occupies a large amount of memory. If the image segmentation model under the pytorch framework is run in an application program of the image segmentation device, it will greatly increase the storage space occupied by the application program. At the same time, when running the image segmentation model under the pytorch framework, it has a relatively high dependence on the Graphics Processing Unit (GPU). If there is no GPU installed in the image segmentation device, the image segmentation model will have a slow processing speed. The forward inference framework is generally for a specific platform (such as an embedded platform), and the hardware configurations of different platforms are different. When the forward inference framework is deployed on the platform, it can combine the hardware configuration of the platform, make reasonable use of resources, and perform optimization and acceleration. That is, the forward inference framework can perform optimization and acceleration when running the model deployed inside it. The forward inference model is mainly used for the prediction process of the model. Among them, the prediction process includes the testing process and the prediction process (application process) of the model, but does not include the training process of the model. And the forward inference framework has a low dependence on the GPU and is relatively lightweight, and will not cause the application program to occupy a large amount of storage space. Therefore, when applying the image segmentation model, the image segmentation model is run in the forward inference framework. In one embodiment, before applying the image segmentation model, it is first determined whether the image segmentation model runs in the forward inference framework. If the image segmentation model runs in the forward inference framework, the image segmentation model is directly applied. If the image segmentation model does not run in the forward inference framework, the image segmentation model is converted into a network model recognizable in the forward inference framework. Exemplarily, the specific type of the forward inference framework can be set according to the actual situation. For example, the forward inference framework is the openvino framework. At this time, when converting the image segmentation model under the pytorch framework into the image segmentation model under the openvino framework, the specific means can be: using the existing pytorch conversion tool to convert the image segmentation model into an Open Neural Network Exchange (ONNX) model, and then using the openvino conversion tool to convert the ONNX model into the image segmentation model under the openvino framework. Among them, ONNX is a standard for representing deep learning models, which can enable the model to be transferred between different frameworks.

[0121] Based on the above embodiment, after the loss function of the image segmentation model converges, it further includes: deleting the edge module.

[0122] It is understandable that the advantage of setting the edge module during the training process is to improve the learning ability of the image segmentation model for edges, thereby ensuring the accuracy of the segmented result image. During the application process of the image segmentation model, since only the first segmented image needs to be output, without the need to output the edge result image, and the image segmentation model already has the learning ability for edges, therefore, when applying the image segmentation model, the edge module therein can be deleted, that is, the data processing process of the edge module is cancelled during the application of the image segmentation model, so as to reduce the data processing volume of the image segmentation model and improve the processing speed.

[0123] As described above, by collecting the original images in different scenarios, the problems of consuming a large amount of workload and production cost when collecting the original images frame by frame based on video data can be avoided, and there is less duplicate content among the original images in different scenarios, which is beneficial to improving the learning ability of the image segmentation model. The encoding module of the image segmentation model adopts a lightweight network, which can reduce the data processing volume during encoding. At the same time, through the channel shuffling module, the image features between channels can be shuffled without significantly increasing the computational amount, so as to enrich the feature information in the channels, thereby ensuring the accuracy of the image segmentation model. And by upsampling the shuffled features and fusing the shuffled features of a higher resolution, the detailed features at different resolutions can be enriched, further ensuring the accuracy of the image segmentation model. In addition, by using the fused features and the second decoded features, the features of each layer can be reused and deeply supervised, improving the utilization rate of the information contained in the features, strengthening the information transmission efficiency, and enhancing the role of label data supervision. By setting the edge module, the learning ability of the image segmentation model for edges is improved, further ensuring the accuracy of the image segmentation model, and during the application process, the edge module is deleted to reduce the computational amount of the image segmentation model. Converting the image segmentation model into an image segmentation model under the forward inference framework can reduce the dependence of the image segmentation model on the GPU and reduce the storage space occupied by the application program running the image segmentation model. During the application process of the trained image segmentation model, no artificial prior or interaction is required, and it can accurately segment the portrait area in the video data. After testing, in the environment of a common PC integrated graphics card, the processing time for each frame of image in the video data is only about 20 ms, and real-time automatic portrait segmentation can be achieved.

[0124] Based on the above embodiments, the image segmentation model further includes: a decoding module. Correspondingly, after step 235, it further includes: inputting the first decoded feature with the highest resolution into the decoding module to obtain a corresponding new first decoded feature.

[0125] Figure 7 This is a schematic structural diagram of another image segmentation model provided by the embodiments of the present application. Compared with Figure 3 the image segmentation model shown, Figure 7The image segmentation model shown in further includes a decoding module 28.

[0126] Exemplarily, after obtaining the first decoded feature with the highest resolution through the residual module, it passes through a decoding module to further decode the first decoded feature with the highest resolution, that is, a new first decoded feature is obtained. At this time, the new first decoded feature can be considered as the first decoded feature finally obtained at the level with the highest resolution. After that, the new first decoded feature is input into the multiple upsampling module and the edge module set in the level with the highest resolution. It can be understood that the number of channels and the resolution of the new first decoded feature are the same as those of the original first decoded feature. For example, Figure 7 The first decoded feature after passing through the decoding module 28 in is denoted as Refine112×112, and its resolution is the same as that of RS Block112×112. In one embodiment, the decoding module is a convolutional network, and the number of convolutional layers and the structural embodiment are not limited. By the decoding module, the accuracy of the first decoded feature at the highest layer can be improved, and thus the accuracy of the image segmentation model can be improved. It should be noted that for the first decoded feature, the lower the resolution, the higher the semantic features it has, and the higher the resolution, the richer the detailed features it has. For the first decoded feature with the highest resolution, directly upsampling it will result in a sawtooth phenomenon, that is, the detailed features appear as sawtooth. Therefore, a decoding module is added to make the finally obtained new first decoded feature transition more evenly and avoid the appearance of the sawtooth phenomenon. However, after upsampling the first decoded features included in other layers, there is basically no sawtooth phenomenon. Even if a decoding module is set for them, it has little impact on the accuracy of the image segmentation model. Therefore, there is no need to set a decoding module for other layers. It can be understood that in practical applications, if a sawtooth phenomenon appears after upsampling the first decoded features of other layers, a decoding module can also be set for them to improve the accuracy of the image segmentation model.

[0127] It can be understood that in the above image segmentation methods, the target object is described as a human. In practical applications, the target object can also be any other object.

[0128] Figure 8 The following is a schematic structural diagram of an image segmentation device provided by an embodiment of the present application. Refer to Figure 8 , this image segmentation device includes: a data acquisition module 301, a first segmentation module 302, a second segmentation module 303, and a repeated segmentation module 304.

[0129] Among them, a data acquisition module 301 is configured to acquire a current frame image in video data, where a target object is displayed in the video data; a first segmentation module 302 is configured to input the current frame image into a trained image segmentation model to obtain a first segmentation image based on the target object; a second segmentation module 303 is configured to perform smoothing processing on the first segmentation image to obtain a second segmentation image based on the target object; a repeated segmentation module 304 is configured to use the next frame image in the video data as the current frame image, and return to perform the operation of inputting the current frame image into the trained image segmentation model until corresponding second segmentation images are obtained for each frame image in the video data.

[0130] Based on the above embodiments, it further includes: a training acquisition module, configured to acquire a training data set, where the training data set includes multiple original images; a label construction module, configured to construct a label data set according to the training data set, where the label data set includes multiple segmentation label images and multiple edge label images, and one original image corresponds to one segmentation label image and one edge label image; a model training module, configured to train an image segmentation model according to the training data set and the label data set.

[0131] Based on the above embodiments, taking Figure 3 the shown image segmentation model as an example, the image segmentation model includes: a normalization module 21, an encoding module 22, a channel confusion module 23, a residual module 24, a multi-fold upsampling module 25, an output module 26, and an edge module 27. At this time, the above model training module includes: a normalization unit, configured to input the original image into the normalization module 21 to obtain a normalized image; an encoding unit, configured to use the encoding module 22 to obtain multi-layer image features of the normalized image, and the resolution of each layer of image features is different; a channel confusion unit, configured to input each layer of image features into the corresponding channel confusion module 23 to obtain multi-layer confused features, and each layer of the image features corresponds to a channel confusion module 23; a fusion unit, configured to, except for the confused feature with the highest resolution, perform upsampling on other layers of confused features and fuse them with the confused feature of a higher resolution level to obtain a fusion feature corresponding to the higher resolution level, as Figure 3Except for the confusion features of the highest layer, the confusion features of each other layer are upsampled and then fused with the confusion features of the previous layer to obtain the fused features of the previous layer; a residual unit, which is used to input each layer of fused features into the corresponding residual module 24 respectively to obtain multiple layers of first decoded features, each layer of fused features corresponds to a residual module 24, and the confusion features with the lowest resolution are used as the first decoded features with the lowest resolution; a multi-fold upsampling unit, which is used to input each layer of first decoded features into the corresponding multi-fold upsampling module 25 respectively to obtain multiple layers of second decoded features, each layer of first decoded features corresponds to a multi-fold upsampling module, and the resolutions of the second decoded features are the same as those of the original image; a segmentation output unit, which is used to jointly input the multiple layers of second decoded features into the output module 26 to obtain a segmented result image; an edge output unit, which is used to input the first decoded feature with the highest resolution into the edge module 27 to obtain an edge result image; a parameter update unit, which is used to construct a loss function according to each of the second decoded features, the edge result image, the corresponding segmentation label image, and the edge label image, and update the model parameters of the image segmentation model according to the loss function; an image selection unit, which is used to select the next original image and return to execute the operation of inputting the original image into the normalization module until the loss function converges.

[0132] Based on the above embodiments, referring to Figure 7 , the image segmentation model further includes: a decoding module 28. Correspondingly, the above model training module further includes: a decoding unit, which is used to input each layer of the fused features into the corresponding residual module 24 respectively to obtain multiple layers of first decoded features, and then input the first decoded feature with the highest resolution into the decoding module 28 to obtain the corresponding new first decoded feature.

[0133] Based on the above embodiments, the encoding module includes a MobileNetV2 network.

[0134] Based on the above embodiments, the loss function is expressed as: where Loss represents the loss function, n represents the total number of layers corresponding to the second decoded features, represents the sub-loss function calculated according to the second decoded feature with the highest resolution and the corresponding segmentation label image, represents the sub-loss function calculated according to the second decoded feature with the lowest resolution and the segmentation label image, A n represents the second decoded feature with the lowest resolution, B represents the corresponding segmentation label image, and Iou n represents the overlapping similarity between A n and B, and loss edge is the Focal loss function.

[0135] Based on the above embodiments, it further includes: an edge deletion module, which is used to, after the loss function of the image segmentation model converges, further include: deleting this edge module.

[0136] Based on the above embodiments, it further includes: a framework conversion module, which is used to, after training an image segmentation model according to the training data set and the label data set, when the image segmentation model is not a network model recognizable by the forward inference framework, convert the image segmentation model into a network model recognizable by the forward inference framework.

[0137] Based on the above embodiments, the label construction module includes: an annotation acquisition unit, which is used to acquire an annotation result for the original image; a segmentation label acquisition unit, which is used to obtain a corresponding segmentation label image according to the annotation result; an erosion unit, which is used to perform an erosion operation on the segmentation label image to obtain an eroded image; a Boolean unit, which is used to perform a Boolean operation on the segmentation label image and the eroded image to obtain an edge label image corresponding to the original image; a data set construction unit, which is used to obtain a label data set according to the segmentation label image and the edge label image.

[0138] Based on the above embodiments, it further includes: a target background acquisition module, which is used to, after performing smoothing processing on the first segmentation image to obtain a second segmentation image based on the target object, acquire a target background image, where the target background image contains a target background; a background replacement module, which is used to perform background replacement on the current frame image according to the target background image and the second segmentation image to obtain a new current frame image.

[0139] The above-provided image segmentation device can be used to execute the image segmentation method provided in any of the above embodiments, and has corresponding functions and beneficial effects.

[0140] It should be noted that in the embodiments of the above image segmentation device, the various units and modules included are only divided according to functional logic, but are not limited to the above division, as long as the corresponding functions can be realized; in addition, the specific names of the functional units are only for the convenience of mutual distinction and do not limit the protection scope of this application.

[0141] Figure 9 It is a schematic structural diagram of an image segmentation device provided in an embodiment of this application. As Figure 9 shown, the image segmentation device includes a processor 40, a memory 41, an input device 42, and an output device 44; the number of processors 40 in the image segmentation device can be one or more, Figure 9 Taking one processor 40 as an example. The processor 40, the memory 41, the input device 42, and the output device 43 in the image segmentation device can be connected through a bus or other means,Figure 9 Take the bus connection as an example.

[0142] The memory 41, as a computer-readable storage medium, can be used to store software programs, computer-executable programs, and modules, such as the program instructions / modules corresponding to the image segmentation method in the embodiments of the present application (for example, the data acquisition module 301, the first segmentation module 302, the second segmentation module 303, and the repeated segmentation module 304 in the image segmentation device). The processor 40 executes various functional applications and data processing of the image segmentation device by running the software programs, instructions, and modules stored in the memory 41, that is, implements the above-mentioned image segmentation method.

[0143] The memory 41 may mainly include a program storage area and a data storage area. Among them, the program storage area can store an operating system and application programs required for at least one function; the data storage area can store data created according to the use of the image segmentation device, etc. In addition, the memory 41 may include high-speed random access memory and may also include non-volatile memory, such as at least one magnetic disk storage device, a flash memory device, or other non-volatile solid-state storage devices. In some instances, the memory 41 may further include a memory remotely set relative to the processor 40, and these remote memories can be connected to the image segmentation device through a network. Examples of the above network include but are not limited to the Internet, an enterprise intranet, a local area network, a mobile communication network, and combinations thereof.

[0144] The input device 42 can be used to receive input digital or character information, and generate key signal inputs related to the user settings and function controls of the image segmentation device. The output device 43 may include a display device such as a display screen.

[0145] The above-mentioned image segmentation device includes an image segmentation device, which can be used to execute any image segmentation method and has corresponding functions and beneficial effects.

[0146] In addition, the embodiments of the present application also provide a storage medium containing computer-executable instructions, and the computer-executable instructions are used to execute relevant operations in the image segmentation method provided by any embodiment of the present application when executed by a computer processor, and have corresponding functions and beneficial effects.

[0147] Those skilled in the art should understand that the embodiments of the present application can be provided as a method, a system, or a computer program product.

[0148] Thus, the present application may take the form of a complete hardware embodiment, a complete software embodiment, or an embodiment combining software and hardware aspects. Moreover, the present application may take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk memory, CD-ROM, optical memory, etc.) that contain computer-usable program code. The present application is described with reference to the flowcharts and / or block diagrams of methods, apparatus (systems), and computer program products according to the embodiments of the present application. It should be understood that each flow and / or block in the flowchart and / or block diagram, and the combination of flows and / or blocks in the flowchart and / or block diagram, can be implemented by computer program instructions. These computer program instructions can be provided to the processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing devices to generate a machine, such that the instructions executed by the processor of the computer or other programmable data processing devices generate means for implementing the functions specified in the process Figure 1 one process or multiple processes and / or blocks Figure 1 These computer program instructions can also be stored in a computer-readable memory that can direct a computer or other programmable data processing device to work in a specific manner, such that the instructions stored in the computer-readable memory generate a manufactured article including instruction means that implement the functions specified in the process Figure 1 one process or multiple processes and / or blocks Figure 1 These computer program instructions can also be loaded onto a computer or other programmable data processing device, such that a series of operation steps are executed on the computer or other programmable device to generate a computer-implemented process, so that the instructions executed on the computer or other programmable device provide steps for implementing the functions specified in the process Figure 1 one process or multiple processes and / or blocks Figure 1 one block or multiple blocks

[0149] In a typical configuration, a computing device includes one or more processors (CPUs), an input / output interface, a network interface, and memory. The memory may include non-permanent memory in the form of computer-readable media, random access memory (RAM), and / or non-volatile memory, such as read-only memory (ROM) or flash memory (flash RAM). The memory is an example of computer-readable media.

[0150] Computer readable media include permanent and non-permanent, removable and non-removable media that can be implemented by any method or technology to store information. Information can be computer readable instructions, data structures, program modules or other data. Examples of computer storage media include, but are not limited to, phase change memory (PRAM), static random access memory (SRAM), dynamic random access memory (DRAM), other types of random access memory (RAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory or other memory technology, compact disk read-only memory (CD-ROM), digital versatile disk (DVD) or other optical storage, magnetic cassettes, magnetic tape magnetic disk storage or other magnetic storage devices or any other non-transmission media that can be used to store information that can be accessed by a computing device. As defined herein, computer readable media does not include temporary computer readable media (transitory media), such as modulated data signals and carrier waves.

[0151] It should also be noted that the terms "include", "comprises" or any other variations thereof are intended to cover non-exclusive inclusion, so that a process, method, commodity or device including a series of elements includes not only those elements, but also other elements not explicitly listed, or also includes elements inherent to such process, method, commodity or device. In the absence of more restrictions, the elements defined by the sentence "comprises a ..." do not exclude the existence of other identical elements in the process, method, commodity or device including the elements.

[0152] Note that the above are only preferred embodiments of the present application and the technical principles used. Those skilled in the art will understand that the present application is not limited to the specific embodiments described herein, and that various obvious changes, readjustments and substitutions can be made by those skilled in the art without departing from the scope of protection of the present application. Therefore, although the present application is described in more detail through the above embodiments, the present application is not limited to the above embodiments, and may include more other equivalent embodiments without departing from the concept of the present application, and the scope of the present application is determined by the scope of the appended claims.

Claims

1. An image segmentation method, characterized in that, Including: Obtain the current frame image in the video data, where the target object is shown in the video data; Input the current frame image into the trained image segmentation model to obtain a first segmentation image based on the target object; Perform smoothing processing on the first segmentation image to obtain a second segmentation image based on the target object; Take the next frame image in the video data as the current frame image, and return to perform the operation of inputting the current frame image into the trained image segmentation model until corresponding second segmentation images are obtained for each frame image in the video data; The image segmentation model includes: a normalization module, an encoding module, multiple channel confusion modules, multiple residual modules, a multi-fold upsampling module, an output module, and an edge module; the normalization module is used to transform the input image into a normalized image, the encoding module is used to obtain multi-layer image features of the normalized image, and the resolution of each layer of the image features is different; each channel confusion module corresponds to input one layer of image features respectively to obtain multi-layer confusion features; each residual module is used to input the corresponding fusion feature to obtain multi-layer first decoding features, and the fusion feature is obtained by upsampling the confusion features of each layer except the confusion feature with the highest resolution and fusing them with the confusion feature with a higher resolution, and the confusion feature with the lowest resolution is used as the first decoding feature with the lowest resolution; the multi-fold upsampling module is used to perform multi-fold upsampling on the first decoding features to obtain multi-layer second decoding features with a resolution equal to that of the image input into the normalization module; the output module is used to input the combined multi-layer second decoding features to obtain a segmentation result image; the edge module is used to input the first decoding feature with the highest resolution to obtain an edge result image; the loss function of the image segmentation model is constructed based on the segmentation label image, edge label image, second decoding features, and edge result image of the image input into the normalization module.

2. The image segmentation method according to claim 1, characterized in that, Also including: Obtain a training data set, where the training data set contains multiple original images; Construct a label data set according to the training data set, where the label data set contains multiple segmentation label images and multiple edge label images, and one original image corresponds to one segmentation label image and one edge label image; Train the image segmentation model according to the training data set and the label data set.

3. The image segmentation method according to claim 2, characterized in that, The training of the image segmentation model according to the training data set and the label data set includes: Input the original image into the normalization module to obtain a normalized image; Use the encoding module to obtain multi-layer image features of the normalized image, and the resolution of each layer of the image features is different; Input each layer of the image features into the corresponding channel confusion module respectively to obtain multi-layer confusion features, and each layer of the image features corresponds to one channel confusion module; Except for the confusion feature with the highest resolution, perform upsampling on the confusion features of each other layer and fuse them with the confusion feature with a higher resolution to obtain the fusion feature corresponding to the higher resolution; Input each layer of the fused features into the corresponding residual module to obtain multiple layers of first decoded features. Each layer of the fused features corresponds to a residual module, and the confused feature with the lowest resolution is used as the first decoded feature with the lowest resolution. Input each layer of the first decoded features into the corresponding multi-fold upsampling module to obtain multiple layers of second decoded features. Each layer of the first decoded features corresponds to a multi-fold upsampling module, and the resolutions of the second decoded features are the same as that of the original image. Jointly input the multiple layers of the second decoded features into the output module to obtain a segmented result image. Input the first decoded feature with the highest resolution into the edge module to obtain an edge result image. Construct a loss function based on each of the second decoded features, the edge result image, the corresponding segmentation label image, and the edge label image, and update the model parameters of the image segmentation model according to the loss function. Select the next original image and return to perform the operation of inputting the original image into the normalization module until the loss function converges.

4. The image segmentation method according to claim 3, characterized in that, The image segmentation model further includes: a decoding module. After inputting each layer of the fused features into the corresponding residual module to obtain multiple layers of first decoded features, it further includes: Input the first decoded feature with the highest resolution into the decoding module to obtain a corresponding new first decoded feature.

5. The image segmentation method according to claim 3, wherein, The encoding module includes a MobileNetV2 network.

6. The image segmentation method according to claim 3, characterized in that, The loss function is expressed as: Among them, Loss represents the loss function, and n represents the total number of layers corresponding to the second decoded feature. represents the sub-loss function calculated according to the second decoded feature with the highest resolution and the corresponding segmentation label image. represents the sub-loss function calculated according to the second decoded feature with the lowest resolution and the segmentation label image. A n represents the second decoded feature with the lowest resolution, B represents the corresponding segmentation label image, and Iou n represents A n and the overlapping similarity of B, and loss edge is the Focal loss function.

7. The image segmentation method according to claim 3, characterized in that, After the loss function of the image segmentation model converges, it further includes: Delete the edge module.

8. The image segmentation method according to claim 2, characterized in that, After training the image segmentation model according to the training data set and the label data set, it further includes: When the image segmentation model is not a network model recognizable by the forward inference framework, convert the image segmentation model into a network model recognizable by the forward inference framework.

9. The image segmentation method according to claim 2, wherein The constructing the label data set according to the training data set includes: Obtain the annotation result for the original image; Obtain the corresponding segmentation label image according to the annotation result; Perform an erosion operation on the segmentation label image to obtain an eroded image; Perform a Boolean operation on the segmentation label image and the eroded image to obtain the edge label image corresponding to the original image; Obtain the label data set according to the segmentation label image and the edge label image.

10. The image segmentation method according to claim 1, wherein After performing smoothing processing on the first segmented image to obtain a second segmented image based on the target object, it further includes: Obtain a target background image, where the target background image contains the target background; Perform background replacement on the current frame image according to the target background image and the second segmented image to obtain a new current frame image.

11. An image segmentation device, wherein Includes: A data acquisition module, configured to acquire a current frame image in video data, where the video data shows a target object; A first segmentation module, configured to input the current frame image into a trained image segmentation model to obtain a first segmented image based on the target object; A second segmentation module, configured to perform smoothing processing on the first segmented image to obtain a second segmented image based on the target object; A repeated segmentation module, configured to use the next frame image in the video data as the current frame image, and return an operation of inputting the current frame image into a trained image segmentation model until a corresponding second segmentation image is obtained for each frame image in the video data; The image segmentation model includes: a normalization module, an encoding module, a plurality of channel shuffle modules, a plurality of residual modules, a multi-fold upsampling module, an output module, and an edge module; the normalization module is configured to transform an input image into a normalized image, the encoding module is configured to obtain multi-layer image features of the normalized image, and the resolution of each layer of the image features is different; each of the channel shuffle modules corresponds to inputting one layer of image features to obtain multi-layer shuffled features; each of the residual modules is configured to input corresponding fused features to obtain multi-layer first decoding features, the fused features are obtained by upsampling each layer of the shuffled features except the shuffled feature with the highest resolution, and fusing with the shuffled feature with a higher resolution, and the shuffled feature with the lowest resolution is used as the first decoding feature with the lowest resolution; the multi-fold upsampling module is configured to perform multi-fold upsampling on the first decoding features to obtain multi-layer second decoding features with a resolution equal to that of the image input to the normalization module; the output module is configured to input the combined multi-layer second decoding features to obtain a segmentation result image; the edge module is configured to input the first decoding feature with the highest resolution to obtain an edge result image; the loss function of the image segmentation model is constructed based on the segmentation label image, the edge label image, the second decoding features, and the edge result image of the image input to the normalization module.

12. An image segmentation device, wherein Comprising: One or more processors A memory for storing one or more programs; When the one or more programs are executed by the one or more processors, the one or more processors implement the image segmentation method according to any one of claims 1-10.

13. A computer-readable storage medium, on which a computer program is stored, wherein When the program is executed by a processor, it implements the image segmentation method according to any one of claims 1-10.

Citation Information

Patent Citations

  • Video instance segmentation method and device based on convolutional neural network

    CN110827292A