Image processing method, mask image output model training method and device

By using masked images to output the model in the cutout technology, processing the feature data of the target image and gradually adjusting the deletion mask, the problem of insufficient fineness of cutouts in the existing technology is solved, and high-precision cutout and rapid processing of fine texture edges is achieved.

CN114066907BActive Publication Date: 2025-06-06BEIJING KINGSOFT CLOUD NETWORK TECH CO LTD
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202010907289.0
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Priority Date
2020-08-10
Filing Date
2020-09-01
Publication Date
2025-06-06
Estimated Expiration
2040-09-01

AI Technical Summary

Technical Problem

The prior art is difficult to achieve high fineness in the process of cutting pictures, especially when dealing with fine texture edges such as hair, the fineness is insufficient and it is difficult to meet the high fineness needs.

Method used

By obtaining the feature data of the target image and inputting it into the pre-trained mask image output model, the mask image corresponding to the target image is output. Based on this mask image, the target area is extracted from the target image, and the initial mask is gradually adjusted through multiple splicing and convolution operations to obtain a more refined mask image.

Benefits of technology

It realizes high-precision cutout of fine texture edges such as hair, improves the fineness of cutouts, improves the cutout effect of the picture, and significantly improves the processing speed, and supports real-time cutout results.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114066907B_ABST
    Figure CN114066907B_ABST
Patent Text Reader

Abstract

The present invention provides a method for processing an image, a method for training a mask image output model and a device, wherein the characteristic data of a target image is input into a pre-trained mask image output model, and a mask image corresponding to the target image is output; based on the mask image, a target area is extracted from the target image; wherein the mask image output model obtains an initial mask of the target image through the characteristic data; obtains an edge area of ​​the target area to be extracted through the target image and the initial mask; and adjusts the initial mask based on the edge area and the target image to obtain the mask image. In the process of processing the characteristic data of the image, the output model makes a secondary adjustment to the obtained initial mask, further refines the edge area of ​​the image, can achieve the cutout accuracy for the edge of fine textures such as hair, improves the fineness of the cutout, and thus improves the cutout effect of the image.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of image processing, and in particular to an image processing method, a training method and a device for a mask image output model. Background Art

[0002] Cutting out is to separate a part of the image area from the original image. The separated image area can include the main objects such as people, animals, still life or background in the original image; these image areas can be used for subsequent creation of images. The cutout algorithm in the related technology is usually implemented based on the calibrated trimap, but the degree of precision is poor; the degree of precision of the cutout can be improved to a certain extent through the semantic segmentation algorithm or the instance segmentation algorithm, but the precision of the processing of fine texture edges such as hair is still poor, and it is difficult to meet the higher precision requirements. Summary of the invention

[0003] In view of this, the purpose of the present invention is to provide an image processing method, a training method and device for a mask image output model to improve the precision of the cutout and thereby enhance the cutout effect of the image.

[0004] In a first aspect, an embodiment of the present disclosure provides an image processing method, the method comprising: obtaining feature data of a target image; inputting the feature data into a pre-trained mask image output model, and outputting a mask image corresponding to the target image; based on the mask image, extracting a target area from the target image; wherein the mask image output model is used to: obtain an initial mask of the target image through the feature data; obtain an edge area of ​​the target area to be extracted through the target image and the initial mask; and adjust the initial mask based on the edge area and the target image to obtain the mask image.

[0005] Furthermore, the step of obtaining the edge area of ​​the target area to be extracted through the target image and the initial mask includes: performing a first splicing process on the target image and the initial mask to obtain first spliced ​​data; performing a first convolution calculation on the first spliced ​​data to obtain a first data feature of the first spliced ​​data; and extracting the edge area of ​​the target area from the target image based on the first data feature.

[0006] Furthermore, the step of extracting the edge area of ​​the target area from the target image based on the first data feature includes: identifying the foreground area and the background area in the target image based on the first data feature; corroding and dilating the edge of the foreground area to obtain the edge area; generating a region calibration image of the target image based on the foreground area, the background area and the edge area; wherein the region calibration image is used to indicate the position of the foreground area, the background area and the edge area in the target image; and extracting the edge area of ​​the target area from the target image according to the position of the edge area indicated by the region calibration image.

[0007] Furthermore, the step of adjusting the initial mask based on the edge area and the target image to obtain the mask image includes: performing a second splicing process on the target image and the initial mask to obtain second spliced ​​data; performing a second convolution calculation on the second spliced ​​data to obtain second data features of the second spliced ​​data; and obtaining the mask image through the second data features and the edge area.

[0008] Furthermore, the step of obtaining a mask image through the second data feature and the edge area includes: for each feature point in the second data feature, superimposing the feature value of the feature point in the second data feature and the value at the position corresponding to the position of the feature point in the regional calibration image used to indicate the position of the edge area to obtain the superimposed feature value of the feature point; the superimposed feature values ​​of each feature point are combined into superimposed features, and a third convolution operation is performed on the superimposed features through a preset residual network to obtain the mask image.

[0009] Furthermore, the step of obtaining an initial mask of the target image through feature data includes: extracting multi-level features from the feature data; dividing the multi-level features into multiple level groups, performing a first fusion process on the level features in each level group to obtain an initial fused feature; wherein the levels in each level group are adjacent; the number of levels of the initial fused feature matches the number of levels of the decoding layer; performing a second fusion process on the initial fused feature in a preset order to output an initial mask.

[0010] Furthermore, the step of performing a second fusion processing on the initial fusion features in a preset order and outputting an initial mask includes: for the highest level decoding layer, inputting the initial fusion features of the highest level in the initial fusion features to the highest level decoding layer; for the decoding layers of levels other than the highest level, performing a third fusion processing on the initial fusion features of the current level and all the initial fusion features of levels higher than the current level to obtain a third fusion feature; inputting the third fusion feature to the decoding layer of the current level; and performing a fourth fusion processing on the features input to each decoding layer according to the arrangement order of each decoding layer to obtain an initial mask.

[0011] In a second aspect, an embodiment of the present disclosure provides a training method for a mask image output model, the method comprising: determining a sample image based on a preset training sample set; the sample image carries a mask label; obtaining image features of the sample image; training a preset mask image output model based on the image features and the mask label to obtain a trained mask image output model; wherein the mask image output model is used to: obtain an initial mask of the sample image through image features; obtain an edge area of ​​a target area to be extracted through the sample image and the initial mask; adjust the initial mask based on the edge area and the sample image to obtain a mask image of the sample image.

[0012] Furthermore, the steps of training a preset mask image output model based on image features and mask labels to obtain a trained mask image output model include: extracting multi-level features from image features; dividing the multi-level features into multiple level groups, performing a first fusion process on the level features in each level group to obtain an initial fusion feature; wherein the levels in each level group are adjacent; the number of levels of the initial fusion feature matches the number of levels of the decoding layer; performing a second fusion process on the initial fusion feature in a preset order to obtain a processing result; determining a loss value based on the processing result and the mask label; for the highest level decoding layer, updating the parameters of the highest level decoding layer based on the loss value until the mask image output model converges to obtain the parameters of the highest level decoding layer; for decoding layers other than the highest level, updating the parameters of the current decoding layer based on the parameters of all decoding layers with a higher level than the current decoding layer and the loss value until the mask image output model converges to obtain the parameters of the current decoding layer.

[0013] Furthermore, for decoding layers other than the highest level, the parameters of the current decoding layer are updated based on the parameters of all decoding layers at a level higher than the current decoding layer and the loss value until the mask image output model converges, and the step of obtaining the parameters of the current decoding layer includes: while fixing the parameters of all decoding layers at a level higher than the current decoding layer, updating the parameters of the current decoding layer based on the loss value until the mask image output model converges; updating the parameters of decoding layers at a higher level than the current decoding layer and the parameters of the current decoding layer based on the loss value until the mask image output model converges.

[0014] In a third aspect, an embodiment of the present disclosure provides an image processing device, which includes: an acquisition module for acquiring feature data of a target image; an output module for inputting the feature data into a pre-trained mask image output model, and outputting a mask image corresponding to the target image; an extraction module for extracting a target area from the target image based on the mask image; wherein the mask image output model is used to: obtain an initial mask of the target image through the feature data; obtain an edge area of ​​the target area to be extracted through the target image and the initial mask; and adjust the initial mask based on the edge area and the target image to obtain the mask image.

[0015] In a fourth aspect, an embodiment of the present disclosure provides a training device for a mask image output model, the device comprising: a determination module for determining a sample image based on a preset training sample set; the sample image carries a mask label; a training module for obtaining image features of the sample image; training the preset mask image output model based on the image features and the mask label to obtain a trained mask image output model; wherein the mask image output model is used to: obtain an initial mask of the sample image through image features; obtain an edge area of ​​a target area to be extracted through the sample image and the initial mask; adjust the initial mask based on the edge area and the sample image to obtain a mask image of the sample image.

[0016] In a fifth aspect, an embodiment of the present disclosure provides an electronic device, comprising a processor and a memory, wherein the memory stores machine executable instructions that can be executed by the processor, and the processor executes the machine executable instructions to implement the image processing method of any one of the first aspect, or the training method of the mask image output model of any one of the second aspect.

[0017] In a sixth aspect, an embodiment of the present disclosure provides a machine-readable storage medium, which stores machine-executable instructions. When the machine-executable instructions are called and executed by a processor, the machine-executable instructions prompt the processor to implement the image processing method of any one of the first aspects, or the training method of the mask image output model of any one of the second aspects.

[0018] The embodiments of the present invention bring the following beneficial effects:

[0019] The embodiment of the present invention provides an image processing method, a training method and device for a mask image output model, wherein the feature data of a target image is input into a pre-trained mask image output model, and a mask image corresponding to the target image is output; based on the mask image, a target area is extracted from the target image; wherein the mask image output model obtains an initial mask of the target image through the feature data; obtains an edge area of ​​the target area to be extracted through the target image and the initial mask; and adjusts the initial mask based on the edge area and the target image to obtain the mask image. In the process of processing the feature data of the image, the output model makes a secondary adjustment to the obtained initial mask, further refines the edge area of ​​the image, can achieve the accuracy of cutting out the edges of fine textures such as hair, improves the fineness of cutting out, and thus improves the cutting effect of the image.

[0020] Other features and advantages of the present invention will be described in the following description, and partly become apparent from the description, or understood by practicing the present invention. The purpose and other advantages of the present invention are realized and obtained by the structures particularly pointed out in the description, claims and drawings.

[0021] In order to make the above-mentioned objects, features and advantages of the present invention more obvious and easy to understand, preferred embodiments are given below and described in detail with reference to the accompanying drawings. BRIEF DESCRIPTION OF THE DRAWINGS

[0022] In order to more clearly illustrate the specific implementation methods of the present invention or the technical solutions in the prior art, the drawings required for use in the specific implementation methods or the description of the prior art will be briefly introduced below. Obviously, the drawings described below are some implementation methods of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without paying any creative work.

[0023] Figure 1 A flowchart of a method for processing an image provided by an embodiment of the present invention;

[0024] Figure 2 A structural schematic diagram of a mask image output model provided by an embodiment of the present invention;

[0025] Figure 3 A schematic diagram of the structure of another mask image output model provided by an embodiment of the present invention;

[0026] Figure 4 A flowchart of a method for training a mask image output model provided by an embodiment of the present invention;

[0027] Figure 5 A schematic diagram of the structure of a picture processing device provided by an embodiment of the present invention;

[0028] Figure 6 A schematic diagram of the structure of a training device for a mask image output model provided by an embodiment of the present invention;

[0029] Figure 7 A schematic diagram of an electronic device provided by an embodiment of the present invention. DETAILED DESCRIPTION

[0030] In order to make the purpose, technical solution and advantages of the embodiments of the present invention clearer, the technical solution of the present invention will be clearly and completely described below in conjunction with the accompanying drawings. Obviously, the described embodiments are part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative work are within the scope of protection of the present invention.

[0031] In the related art, there are mainly three automatic cutout algorithms:

[0032] The first is the traditional method, which implements the cutout based on sampling algorithm, pixel affinity algorithm or a combination of the two. However, these methods need to solve a large linear system and require the user to manually calibrate a trimap map, which is used to calibrate the target area of ​​the cutout, and then use the similarity between adjacent pixels to perform extended matching combination to obtain the final cutout result; the processing speed of this method is proportional to the number of unknown points.

[0033] The second method is based on the CNN (Convolutional Neural Networks) model. The RGB original image and the manually annotated Trimap image are combined as the input of the CNN model. The Trimap image is used to calibrate the key areas of the model. The model finally outputs the Alpha image. Finally, the cutout result is obtained based on the Alpha image.

[0034] The third method is also based on the CNN model. This method usually uses a semantic segmentation model to extract the target body, and then indirectly obtains the Trimap image by corrosion and expansion. Then, the RGB original image and the Trimap original image are combined, and the Alpha image is finally obtained using the second method mentioned above.

[0035] However, the above three methods of cutting out images have certain disadvantages. Among them, the first traditional cutting out method usually has a processing speed proportional to the size of the image. For an image of about 1 meter in size, the processing time of this method varies from a few seconds to a few minutes depending on the computer hardware conditions. Moreover, this method usually does not produce ideal cutting effects in slightly complex scenes. In terms of processing effect and processing speed, this method is far from meeting the requirements of lowering the threshold for creation.

[0036] For the second method, the trimap image is obtained directly or indirectly, and the trimap image is combined with the RGB original image and input into the CNN model. This CNN model is usually deployed in an enterprise-level GPU (Graphics Processing Unit) server. Due to the consumption of IO (Input / Output) ports, it cannot meet the real-time requirements. In addition, the user's manual calibration of the trimap will also increase the complexity of the system and reduce the user experience.

[0037] For the third method, the semantic segmentation algorithm or the instance segmentation algorithm can improve the precision of the cutout to a certain extent. However, for the processing of fine texture edges such as hair, the precision is still poor and it is difficult to meet the higher precision requirements.

[0038] In summary, the embodiments of the present invention provide a method for processing an image, a method for training a mask image output model, and a device. The technology can be applied to the application scenario of extracting a target area from an original image, such as extracting a main object, a foreground image, etc. in the original image. To facilitate understanding of the present embodiment, a method for processing an image disclosed in the embodiment of the present disclosure is first introduced in detail.

[0039] like Figure 1 As shown, the method comprises the steps of:

[0040] Step S102, obtaining feature data of the target image;

[0041] The above-mentioned target image usually refers to the image to be processed, which can be an image containing people, animals, buildings, commodities and other objects or subjects, such as portraits, cats, buildings, apples, etc.; the target image usually includes the detailed features of the object or subject, such as the hair features on the edge of the person; the above-mentioned feature data usually includes the detailed features of the target image; specifically, the target image can be input into the feature extraction network, and the feature data of the target image can be output; the feature extraction network can specifically be a backbone network based on a convolutional neural network, also known as a CNN backbone. The number of channels of the feature data is usually high, for example, it can be 16 channels or 32 channels; the scale of the feature data is usually related to the scale of the target image, for example, the height of the feature data is half of the height of the target image, and the width of the feature data is half of the width of the target image, which can be set according to actual needs.

[0042] Step S104, inputting the feature data into a pre-trained mask image output model, and outputting a mask image corresponding to the target image;

[0043] Among them, the mask image output model is used to: obtain the initial mask of the target image through feature data; obtain the edge area of ​​the target area to be extracted through the target image and the initial mask; adjust the initial mask based on the edge area and the target image to obtain the mask image.

[0044] The above-mentioned mask image output model may include multiple modules, each module has different functions. For example, the mask image output model may include an initial mask output module, through which multi-level features are extracted from feature data, and based on the multi-level features, the initial mask of the target image is output through fusion processing; for another example, the mask image output model may include an edge area extraction module, through which the edge area of ​​the target area is obtained based on the initial mask and the target image; for another example, the mask image output model may include an adjustment module, through which the mask image is obtained based on the edge area of ​​the target area, the initial mask and the target image.

[0045] The mask image output model may include an initial mask output module, which may include multiple layers of sequentially connected computing layers, which may be convolutional layers, fully connected layers, pooling layers, etc.; each computing layer may output a hierarchical feature; the width, height, number of channels, etc. of the hierarchical features output by different computing layers may be the same or different, which may be determined according to the operators in the computing layer.

[0046] The mask image output model may include an edge region extraction module, which may include a stitching function and a convolution layer connected in sequence. Through the initial mask of the target image and the target image, the stitching function may output a stitching data with the same size and multiple channels as the target image, and the convolution layer may perform convolution calculation on the data to focus on the edge region of the object in the image to obtain the edge region of the target area. The edge region of the target area may be an image with an edge position indicator.

[0047] The mask image output model may include an adjustment module, which may include a stitching function, a convolution layer, and an overlay calculation layer connected in sequence. The convolution layer first passes through the initial mask of the target image and the target image. The stitching function can output a stitching data with the same size and multiple channels as the target image. The convolution layer performs convolution calculation on the stitching data to obtain feature data; the feature data and the edge area of ​​the target image are input into the overlay calculation layer for overlay calculation to obtain overlay features, and then the mask image is obtained through convolution calculation; the main purpose of the adjustment module is to enhance the lost detail information, so as to obtain the most refined output.

[0048] Step S106: extracting the target area from the target image based on the mask image.

[0049] The pixel values ​​of the pixels in the mask image usually include two types, 1 and 0; the mask image and the target image are multiplied by corresponding pixels, and the area of ​​the target image corresponding to the area with a pixel value of 1 in the mask image can be extracted, that is, the above-mentioned target area.

[0050] The embodiment of the present invention provides a method for processing an image, inputting feature data of a target image into a pre-trained mask image output model, outputting a mask image corresponding to the target image; extracting a target area from the target image based on the mask image; wherein the mask image output model obtains an initial mask of the target image through the feature data; obtains an edge area of ​​the target area to be extracted through the target image and the initial mask; and adjusts the initial mask based on the edge area and the target image to obtain the mask image. In the process of processing the feature data of the image, the output model makes a secondary adjustment to the obtained initial mask, further refines the edge area of ​​the image, can achieve the accuracy of cutting out the edges of fine textures such as hair, improves the fineness of cutting out, and thus improves the cutting effect of the image.

[0051] The embodiment of the present invention also provides another image processing method, which is implemented on the basis of the above embodiment. This embodiment mainly describes the structure of the mask image output model, the specific implementation process of the step of obtaining the edge area of ​​the target area to be extracted through the target image and the initial mask, and the specific implementation process of the step of adjusting the initial mask based on the edge area and the target image to obtain the mask image. Figure 2 As shown, the mask image output model includes a boundary area extraction module; wherein the boundary area extraction module can also be called a Boundary Attentation Module module. The step of obtaining the boundary area of ​​the target area to be extracted through the target image and the initial mask includes:

[0052] The target image and the initial mask are subjected to a first splicing process to obtain first splicing data; a first convolution calculation is performed on the first splicing data to obtain a first data feature of the first splicing data; and an edge area of ​​the target area is extracted from the target image based on the first data feature.

[0053] The target image and the initial mask have the same size; the first splicing process can be implemented by the CONCATENATE function, specifically, the features of the target image and the initial mask are combined in one dimension, and the finer-grained features are merged together to obtain the first spliced ​​data, which can be a feature map with the same size as the target image and a channel number of 4; the first convolution calculation can be calculated by the convolution layer in the edge area extraction module, focusing the calculation attention on the edge area of ​​the object in the image, and obtaining the first data feature of the first spliced ​​data, which usually includes the features of multiple areas of the target image, mainly including the edge area features of the target image; then based on the edge area features in the first data features, the edge area of ​​the target area is extracted from the target image.

[0054] The step of extracting the edge area of ​​the target area from the target image based on the first data feature includes:

[0055] (1) identifying a foreground area and a background area in a target image based on a first data feature;

[0056] The first data feature generally includes a foreground region feature and a background region feature, wherein the foreground region feature may also include an unknown region feature; specifically, the foreground region of the target image may be identified according to the foreground region feature in the first data feature, and the background region of the target image may be identified according to the background region feature in the first data feature. The foreground region includes an unknown edge region of the foreground region, and the unknown edge region may be referred to as an unknown region to be found.

[0057] (2) performing corrosion and expansion processing on the edge of the foreground area to obtain the edge area;

[0058] Usually, the unknown edge area of ​​the foreground area includes the edge area of ​​the target image; the above-mentioned corrosion processing is usually a process of eliminating the boundary points of the foreground area and shrinking the boundary inward. By corroding the foreground area, small and meaningless edge objects can be eliminated. The above-mentioned expansion processing is usually a process of merging all the points of the background area that are in contact with the foreground area into the foreground area, so that the boundary of the foreground area expands outward. Through the expansion processing, small holes in the foreground area image and small concave parts at the edge of the image can be filled. Specifically, the edge of the foreground area, that is, the above-mentioned edge area, can be obtained according to the corrosion processing and the expansion processing.

[0059] (3) generating a region calibration image of the target image based on the foreground region, the background region, and the edge region; wherein the region calibration image is used to indicate the position of the foreground region, the background region, and the edge region in the target image;

[0060] (4) According to the position of the edge region indicated by the region calibration image, the edge region of the target region is extracted from the target image.

[0061] The region calibration image of the target image includes a foreground region, a background region and an edge region, as well as marks indicating the positions of the respective regions. The edge region can be marked with a designated mark, for example, the position mark of the edge region is displayed with a line filled with a color, and then the edge region of the target region is extracted from the target image. The region calibration image can also be called a trimap image, which can also include 4 channels.

[0062] Further, such as Figure 2 The mask image output model includes an adjustment module, wherein the adjustment module may also be referred to as a Refine Module. The step of adjusting the initial mask based on the edge region and the target image to obtain the mask image includes:

[0063] The target image and the initial mask are subjected to a second splicing process to obtain second splicing data; a second convolution calculation is performed on the second splicing data to obtain second data features of the second splicing data; and a mask image is obtained through the second data features and the edge area.

[0064] The above-mentioned second stitching processing can be implemented by the CONCATENATE function, specifically, the features of the target image and the initial mask are combined in one dimension, and the finer-grained features are merged together to obtain the second stitching data, which can be a feature map with the same size as the target image and a channel number of 4; the above-mentioned second convolution calculation can be calculated by the convolution layer in the adjustment module to strengthen the detail information of the lost edge area to obtain the second data feature of the second stitching data, which usually includes more detailed features of multiple areas of the target image, mainly including finer-grained edge area features in the target image; then the finer-grained edge area features included in the second data features and the edge area output by the edge area extraction module are processed to obtain a mask image, and the pixel values ​​of the pixels in the mask image usually include two types, 1 and 0, wherein the pixel with a pixel value of 0 is the foreground image, and the pixel with a pixel value of 1 is the background image.

[0065] The step of obtaining a mask image by using the second data feature and the edge area includes:

[0066] (1) for each feature point in the second data feature, superimpose a feature value of the feature point in the second data feature with a value at a position corresponding to the position of the feature point in the region calibration image for indicating the position of the edge region to obtain a superimposed feature value of the feature point;

[0067] The second data feature may include multiple channels, and the number of channels of the region calibration image may be set according to actual requirements, and is usually the same as the number of channels of the second data feature, so as to achieve superposition processing of each feature point. The superposition processing of the feature points may be achieved through the concat function, which is mainly used to connect two or more data features. Specifically, the feature value of each feature point in the second data feature is superimposed with the value at the position corresponding to the position of the feature point in the edge region of the region calibration image to obtain a superimposed feature value.

[0068] (2) The superimposed feature values ​​of each feature point are combined into superimposed features, and the superimposed features are subjected to a third convolution operation through a preset residual network to obtain a mask image.

[0069] After obtaining the superimposed feature value of each feature point, the superimposed feature value of each feature point can be combined into a superimposed feature to make the details of the edge area marked in the regional calibration image more refined, so as to achieve the processing of fine texture edges such as hair. The above residual network can be implemented by several ResBlocks, and the number of channels of the ResBlock can be 4 channels or 8 channels. The third convolution operation is performed on the superimposed feature through the residual network, and the difference between the predicted value and the observed value of the superimposed feature is calculated to obtain the final refined mask image, including the fine hair edge.

[0070] Further, such as Figure 3 As shown, the mask image output model may include an encoding module, a fusion module and a decoding module; wherein the step of obtaining the initial mask of the target image through the feature data includes:

[0071] (1) Extract multi-level features from feature data;

[0072] (2) Dividing the multi-level features into multiple level groups, performing a first fusion process on the level features in each level group to obtain an initial fusion feature; wherein the levels in each level group are adjacent; and the number of levels of the initial fusion feature matches the number of levels of the decoding layer;

[0073] The decoding layer usually refers to a computing layer in the decoding module. Specifically, the data features of the target image are first input into the encoding module, and the encoding module extracts multi-level features from the data features; Figure 3The example of eight hierarchical features included in the example is used for explanation. The fusion module divides the eight hierarchical features into five hierarchical groups, among which hierarchical features 1 and hierarchical features 2 in the multi-level features are divided into one hierarchical group, and after the first fusion processing of hierarchical features 1 and hierarchical features 2, fusion feature 1 in the initial fusion feature is obtained; hierarchical feature 3 in the multi-level features is divided into one hierarchical group, and the hierarchical group does not need to be subjected to feature fusion processing, and fusion feature 2 is obtained directly based on hierarchical feature 3; similarly, hierarchical feature 4 in the multi-level features is divided into one hierarchical group, and the hierarchical group does not need to be subjected to feature fusion processing, and fusion feature 3 is obtained directly based on hierarchical feature 4; hierarchical features 5 and hierarchical features 6 in the multi-level features are divided into one hierarchical group, and after the first fusion processing of hierarchical features 5 and hierarchical features 6, fusion feature 4 in the initial fusion feature is obtained; hierarchical features 7 and hierarchical features 8 in the multi-level features are divided into one hierarchical group, and after the first fusion processing of hierarchical features 7 and hierarchical features 8, fusion feature 5 in the initial fusion feature is obtained.

[0074] Of course, there are other ways to divide the hierarchical groups in the fusion module. For example, hierarchical features 3 and 4 can be divided into one hierarchical group for the first fusion process; hierarchical features 2 and 3 can also be divided into one hierarchical group for the first fusion process. The hierarchical group division method can be set according to demand. For example, multiple division methods can be pre-set, and experiments can be conducted for each division method. The most accurate division method for the initial mask of the final output target image is determined as the final division method.

[0075] In order to facilitate feature fusion or make the fused features more effective, no matter which division method is used, the levels in the same level are adjacent; the initial fused features also include multi-level features, such as Figure 3 In the example, there are five levels of fusion features; usually, the number of levels of the initial fusion features matches the number of calculation layers in the decoding module, so that the fusion features of each level in the initial fusion features are input into the corresponding calculation layer in the decoding module for re-fusion calculation.

[0076] In addition, in order to facilitate the fusion of hierarchical features in the same hierarchical group, the above-mentioned mask image output model also includes a compression module; after the encoding module outputs the multi-level features, the multi-level features are first input into the compression module; through the compression module, the number of channels of the hierarchical features of the specified level in the multi-level features is reduced so that the number of channels of the hierarchical features in the same hierarchical group are the same. In addition, the two hierarchical features for the first fusion calculation need to have the same number of channels, as well as the same width and height. If the width and height are different, they also need to be adjusted. The fused features after the first fusion processing of the two hierarchical features have the same number of channels, width and height as the two hierarchical features involved in the fusion calculation.

[0077] (3) Perform a second fusion process on the initial fusion features in a preset order and output an initial mask.

[0078] For example, Figure 3 The order of the initial fusion features shown is that the fusion feature 5 is first input into the calculation layer matching the features of this level in the decoding module for processing, and then the fusion feature 5 is fused with the fusion feature 4; the fusion result is input into the calculation layer matching the fusion feature 4, and then the fusion result of the fusion feature 1 at the lowest level and the fusion feature 2 at the previous level is input into the calculation layer matching the fusion feature 1, and the initial template can be output after certain calculations.

[0079] The step of performing a second fusion process on the initial fusion features in a preset order and outputting an initial mask includes:

[0080] (a) For the highest level decoding layer, the highest level initial fusion features in the initial fusion features are input to the highest level decoding layer;

[0081] Continue to refer Figure 3 As shown in the figure, the encoding module includes 8 layers of calculation layers connected in sequence, and a total of 8 layers of hierarchical features are output; each hierarchical feature has two parameters, for example, in the hierarchical feature 1, in "X / 2", X represents the width or height of the feature data; "X / 2" represents that the width of the hierarchical feature 1 is 1 / 2 of the width of the feature data, and the height is 1 / 2 of the height of the feature data; the other parameter 16 represents that the number of channels of the hierarchical feature 1 is 16; the parameters of other hierarchical features and fusion features have the same meaning.

[0082] As an example, the calculation layer 5 in the decoding module is the calculation layer of the highest level mentioned above, and the fusion feature 5 is the initial fusion feature of the highest level; the fusion feature 5 of the highest level in the initial fusion feature is input into the calculation layer 5 of the highest level.

[0083] (b) for the decoding layers of the layers other than the highest layer, performing a third fusion process on the initial fusion features of the current layer and all the initial fusion features of the layers higher than the current layer to obtain a third fusion feature; and inputting the third fusion feature into the decoding layer of the current layer;

[0084] As an example, for the calculation layer 4 in the decoding module, the fusion feature 4 corresponding to the calculation layer 4 and the fusion feature 5 of the calculation layer 5 higher than the calculation layer 4 are subjected to the third fusion processing to obtain the third fusion feature of the calculation layer 4, and the third fusion feature is input to the calculation layer 4. Figure 3The calculation layer 3 in the calculation layer 3, the fusion feature 3 of the calculation layer 3, the fusion feature 4 of the calculation layer 4 higher than the calculation layer 3, the fusion feature 5 of the calculation layer 5, and the third fusion processing are performed to obtain the third fusion feature of the calculation layer 3, and the third fusion feature is input to the calculation layer 3; for the lowest level calculation layer 1, the fusion feature 1 corresponding to the calculation layer 1 and the fusion feature 2 of the calculation layer 2 higher than the calculation layer 1, the fusion feature 3 of the calculation layer 3, the fusion feature 4 of the calculation layer 4, and the fusion feature 5 of the calculation layer 5 are subjected to the third fusion processing to obtain the third fusion feature of the calculation layer 1, and the third fusion feature is input to the calculation layer 1.

[0085] The fusion features of each layer represent different semantic features. According to specific actual needs, the initial fusion features of the current layer and some initial fusion features of a layer higher than the current layer can be subjected to a third fusion process to obtain the third fusion features; the third fusion features are input into the calculation layer of the current layer; this method can enhance different features and make the detail features of the output initial mask more refined.

[0086] In addition, for all levels except the highest level, the scale of the features needs to be adjusted before the third fusion process. Specifically, the output features of the calculation layer of the middle level of the decoding module and the previous calculation layer adjacent to the current calculation layer are interpolated to match the scale of the fused features input to the current calculation layer. Taking calculation layer 4 as an example, the output features of calculation layer 5 are interpolated, and the width and height of the output features are doubled. At this time, the width and height of the output features are the same as the width and height of the fused feature 4. The output features after interpolation are then convolved with a convolution kernel of 3*3 and fused with the fused feature 4.

[0087] (c) According to the arrangement order of each decoding layer, the features input to each decoding layer are subjected to a fourth fusion process to obtain an initial mask.

[0088] As an example, follow Figure 3 The arrangement order between the various calculation layers in the decoding module shown is that the fused feature 5 input to the calculation layer 5 is first convolved; then the third fused feature input to the calculation layer 4 and the feature output by the calculation layer 5 are convolved; then the third fused feature input to the calculation layer 3 and the feature output by the calculation layer 4 are convolved, and the same processing is performed on the features input to the calculation layer 2 and the calculation layer 1 according to this calculation method and calculation order; finally, the feature output by the calculation layer 1 of the lowest level is input to the specified calculation layer for convolution calculation, and then the output feature is scaled, such as doubling the length and width of the output feature to obtain the initial mask of the target image.

[0089] In the above-mentioned transmission method of each feature, it can be understood that the low-level features in the encoding module are fused and input into a higher-level layer and transmitted in the encoding module through a jump transmission method.

[0090] In addition, in order to facilitate the calculation of the model, the feature data can be preprocessed before being input into the mask image output model. First, determine whether the scale of the feature data meets the preset conditions; wherein the preset conditions are related to the scale of the features output by each calculation layer in the mask image output model; if the scale of the feature data does not meet the preset conditions, adjust the scale of the feature data until it meets the preset conditions. For example, when the encoding module is calculating, the scale of the hierarchical features output by each calculation layer is unchanged and reduced, such as Figure 3 In the example, the scale of the hierarchical feature 8 output by the computing layer 8 is 1 / 32 of the scale of the feature data. Based on this, the scale of the feature data can be adjusted to a multiple of 32 to facilitate the operation of the encoding module.

[0091] It should also be noted that in order to enable the mask image output model to run on mobile terminals such as mobile phones, the operators of each computing layer in the mask image output model all use basic operators, such as convolution kernels, upsampling, downsampling and other operators, thereby obtaining a lightweight mask image output model and facilitating the migration of the model.

[0092] In the above method, the mask image output model makes a secondary adjustment to the initial mask obtained during the process of processing the feature data of the image, further refining the edge area of ​​the image, and can achieve the accuracy of the cutout of fine texture edges such as hair, improve the fineness of the cutout, and thus improve the cutout effect of the image. In addition, the mask image output model can significantly improve the speed of image processing while ensuring the processing effect. During the cutout process, there is no need for the user to manually mark the trimap, and the cutout speed is in milliseconds, which improves the user experience.

[0093] Based on the above-mentioned embodiment of the image processing method, this embodiment also provides a training method for a mask image output model, such as Figure 4 As shown, the method comprises the following steps:

[0094] Step S402, determining a sample image based on a preset training sample set; the sample image carries a mask label;

[0095] Step S404, obtaining the image features of the sample image; training the preset mask image output model based on the image features and the mask label to obtain a trained mask image output model;

[0096] Among them, the mask image output model is used to: obtain the initial mask of the sample image through image features; obtain the edge area of ​​the target area to be extracted through the sample image and the initial mask; adjust the initial mask based on the edge area and the sample image to obtain the mask image of the sample image.

[0097] In the process of processing image features, the mask image output model performs secondary fusion of the features, thereby reducing the amount of feature data processed by the model while preserving the features, thereby improving the cutout processing speed while ensuring the cutout effect; since the amount of data processed by the mask image output model is relatively low, the mask image output model can run on terminal devices with limited hardware conditions, and can provide users with real-time cutout results, thereby improving user experience; in addition, in the process of processing the feature data of the image, the mask image output model performs secondary adjustments to the obtained initial mask, further refining the edge area of ​​the image, and can achieve cutout accuracy for the edges of subtle textures such as hair, thereby improving the fineness of cutouts and thus improving the cutout effect of the image.

[0098] The above-mentioned step of training a preset mask image output model based on image features and mask labels to obtain a trained mask image output model includes: extracting multi-level features from image features; dividing the multi-level features into multiple level groups, and performing a first fusion process on the level features in each level group to obtain an initial fusion feature; wherein the levels in each level group are adjacent; the number of levels of the initial fusion feature matches the number of levels of the decoding layer; performing a second fusion process on the initial fusion feature in a preset order to obtain a processing result; determining a loss value based on the processing result and the mask label; for the highest level decoding layer, updating the parameters of the highest level decoding layer based on the loss value until the mask image output model converges to obtain the parameters of the highest level decoding layer; for decoding layers other than the highest level, updating the parameters of the current decoding layer based on the parameters of all decoding layers with a higher level than the current decoding layer and the loss value until the mask image output model converges to obtain the parameters of the current decoding layer.

[0099] The above-mentioned step of updating the parameters of the current decoding layer based on the parameters of all decoding layers at a higher level than the current decoding layer and the loss value until the mask image output model converges for the decoding layers other than the highest level, and obtaining the parameters of the current decoding layer includes: updating the parameters of the current decoding layer based on the loss value while fixing the parameters of all decoding layers at a higher level than the current decoding layer until the mask image output model converges; updating the parameters of the decoding layers at a higher level of the current decoding layer and the parameters of the current decoding layer based on the loss value until the mask image output model converges.

[0100] The above-mentioned mask image output model usually includes an edge area extraction module to be trained, an adjustment module to be trained, a pre-trained encoding module, a pre-trained fusion module, and a decoding module to be trained; the edge area extraction module includes multiple convolutional layers, the adjustment module includes multiple convolutional layers, and the decoding module includes multiple levels of computing layers; during the training process of the mask image output model, the encoding model and the fusion module can be trained first, and after the encoding model and the fusion module are trained, the decoding module, the edge area extraction module, and the adjustment module are trained.

[0101] When training the encoding model and fusion module, you can first fix the parameters of the decoding module, input the sample image into the encoding module, and obtain the output result after processing by the encoding module, fusion module and decoding module; calculate the loss value based on the output result and the label of the sample image, and adjust the parameters of the encoding module and fusion module according to the loss value; then continue to input the sample image into the encoding module, and repeat this cycle multiple times until the loss value converges and the training of the encoding module and fusion module is completed.

[0102] The following focuses on describing the training process of the decoding module, edge area extraction module, and adjustment module. First, the multi-level features are extracted from the feature data of the sample image through the encoding module; the multi-level features are divided into multiple level groups through the fusion module, and the level features in each level group are subjected to the first fusion process to obtain the initial fusion feature; wherein the levels in each level group are adjacent; the number of levels of the initial fusion feature matches the number of levels of the calculation layer in the decoding module; for each level of the level features in the initial fusion feature, the initial fusion feature is subjected to the second fusion process in a preset order through the decoding module to obtain the initial mask; based on the initial mask and the mask label, the first loss value is determined.

[0103] Through the mask image output model, the target image and the initial mask are subjected to a first splicing process to obtain first splicing data; a first convolution calculation is performed on the first splicing data to obtain a first data feature of the first splicing data; an edge region of the target region is extracted from the target image based on the first data feature; and a second loss value of the edge is calculated based on the edge region;

[0104] The target image and the initial mask are subjected to a second splicing process through the adjustment module to obtain second splicing data; a second convolution calculation is performed on the second splicing data to obtain a second data feature of the second splicing data; and a mask image is obtained through the second data feature and the edge area. A third loss value is determined based on the mask image and the mask label. Based on the first loss value, the second loss value, and the third loss value, the parameters of the edge area extraction module, the adjustment module, and the decoding module included in the trained mask image output model are adjusted to obtain a trained mask image output model.

[0105] For the training of the decoding module, in order to compress the feature quantity of each layer as much as possible to reduce the parameters of the mask image output model and ensure the effectiveness of each layer of features, the decoding module adopts a step-by-step output method for supervised training. When the low-scale feature output is stable, the high-scale feature output training is carried out, and the parameters of the calculation layer that outputs the low-scale features serve as the basis for the parameter training of the calculation layer that outputs the high-scale features.

[0106] Specifically, starting from the highest-level computing layer, training is performed one computing layer at a time; for the highest-level computing layer of the decoding module, the parameters of the highest-level computing layer are updated based on the above-mentioned loss value until the mask image output model converges, and the parameters of the highest-level computing layer are obtained; when training the parameters of the highest-level computing layer, the parameters of other levels remain unchanged, and the parameters of the highest-level computing layer are adjusted during the training process until the output results of the mask image output model converge, and the parameters of the highest-level computing layer are obtained.

[0107] At this time, continue to execute the aforementioned process of determining the sample image based on the preset training sample set to obtain new first loss values, second loss values, and third loss values, and then train the calculation layer of the next level of the highest level based on the first loss value; for the calculation layers of the decoding module other than the highest level, update the parameters of the current calculation layer based on the parameters of all calculation layers higher than the current calculation layer and the first loss value until the mask image output model converges and the parameters of the current calculation layer are obtained. The decoding module starts from the highest level and trains step by step. After obtaining the parameters of the calculation layer of the highest level in the above manner, the parameters of the calculation layer of the highest level can be fixed, and the parameters of the next calculation layer of the highest level can be trained until the mask image output model converges and the parameters of the calculation layer are obtained; and so on, until the training of the calculation layer of the lowest level is completed.

[0108] For the calculation layers other than the highest level, each calculation layer can be trained in the following manner: while fixing the parameters of all the calculation layers higher than the current calculation layer, update the parameters of the current calculation layer based on the first loss value until the mask image output model converges; update the parameters of the calculation layer of the high level of the current calculation layer and the parameters of the current calculation layer based on the loss value until the mask image output model converges. That is, the training of each calculation layer is divided into two steps. The first step is to fix the parameters of the calculation layer of the high level of the current calculation layer, and only train the parameters of the current calculation layer. After the mask image output model converges, the parameters of the calculation layer of the high level of the current calculation layer and the parameters of the current calculation layer are trained simultaneously until the mask image output model converges and the current calculation layer is trained. Then, the next calculation layer of the current calculation layer is trained in the same way until the training of the calculation layer of the lowest level is completed.

[0109] As an example, continue to refer to the above Figure 3 In the masked image output model, taking the computing layer 3 in the decoding module as an example, when the output of the trained masked image output model of computing layer 4 converges, fix all the parameters of computing layer 4 and all the computing layers higher than computing layer 4, and train computing layer 3 with a small learning rate. When the output of the masked image output model converges again, release the fixed state of the parameters of layer_4 and all the computing layers higher than computing layer 4, and train computing layer 3 as a whole until the output of the masked image output model converges.

[0110] The above loss function can be a SAD (Sum of Absolute Difference) function, Where Pa and Pr represent the sample's mask label Label and the model's output mask image, respectively, and the value range is [0,1]. The lower the SAD, the closer the model's output is to the expected Label, and the better the model's output effect is. Otherwise, the worse the model's output Alpha is. The above loss function can also be pixel IOU (Intersection-over-Union), that is, the intersection-over-union ratio of the sample's mask label Label and the model's output mask image, which is used to measure the semantic integrity of the mask image output model. The higher the IOU value, the more complete the semantics.

[0111] The above mask image output model can significantly improve the speed of image processing while ensuring the processing effect. During the cutout process, there is no need for users to manually mark the trimap, and the cutout speed is at the millisecond level, which improves the user experience.

[0112] Corresponding to the above-mentioned picture processing method embodiment, this embodiment provides a structural schematic diagram of a picture processing device, such as Figure 5 As shown, the device comprises:

[0113] An acquisition module 51 is used to acquire feature data of a target image;

[0114] An output module 52 is used to input the feature data into a pre-trained mask image output model and output a mask image corresponding to the target image;

[0115] An extraction module 53, used for extracting a target area from a target image based on a mask image;

[0116] Among them, the mask image output model is used to: obtain the initial mask of the target image through feature data; obtain the edge area of ​​the target area to be extracted through the target image and the initial mask; adjust the initial mask based on the edge area and the target image to obtain the mask image.

[0117] The embodiment of the present invention provides an image processing device, which inputs the feature data of a target image into a pre-trained mask image output model, outputs a mask image corresponding to the target image; based on the mask image, extracts a target area from the target image; wherein the mask image output model obtains an initial mask of the target image through the feature data; obtains an edge area of ​​the target area to be extracted through the initial mask; and adjusts the initial mask based on the edge area to obtain the mask image. In the process of processing the feature data of the image, the output model makes a secondary adjustment to the obtained initial mask, further refines the edge area of ​​the image, can achieve the accuracy of cutting out the edges of fine textures such as hair, improves the fineness of cutting out, and thus improves the cutting effect of the image.

[0118] Furthermore, the above-mentioned mask image output model includes an edge area extraction module; the edge area extraction module is used to: perform a first splicing process on the target image and the initial mask to obtain first splicing data; perform a first convolution calculation on the first splicing data to obtain a first data feature of the first splicing data; and extract the edge area of ​​the target area from the target image based on the first data feature.

[0119] Furthermore, the edge area extraction module is also used to: identify the foreground area and the background area in the target image based on the first data feature; perform corrosion and expansion processing on the edge of the foreground area to obtain the edge area; generate a region calibration image of the target image based on the foreground area, the background area and the edge area; wherein the region calibration image is used to indicate the position of the foreground area, the background area and the edge area in the target image; and extract the edge area of ​​the target area from the target image according to the position of the edge area indicated by the region calibration image.

[0120] Furthermore, the mask image output model includes an adjustment module; the adjustment module is used to: perform a second splicing process on the target image and the initial mask to obtain second splicing data; perform a second convolution calculation on the second splicing data to obtain second data features of the second splicing data; and obtain a mask image through the second data features and the edge area.

[0121] Furthermore, the adjustment module is also used to: perform feature superposition processing on the second data feature and the regional calibration image used to indicate the position of the edge area feature point by feature point to obtain superimposed features; perform a third convolution operation on the superimposed features through a preset residual network to obtain a mask image.

[0122] Furthermore, the mask image output model includes an encoding module, a fusion module and a decoding module; the encoding module is used to extract multi-level features from feature data; the fusion module is used to divide the multi-level features into multiple level groups, and perform a first fusion process on the level features in each level group to obtain initial fused features; wherein the levels in each level group are adjacent; the number of levels of the initial fused features matches the number of levels of the calculation layers in the decoding module; the decoding module is used to perform a second fusion process on the initial fused features in a preset order to output an initial mask.

[0123] Furthermore, the decoding module is also used for: for the highest-level computing layer in the decoding module, inputting the highest-level initial fusion features in the initial fusion features to the computing layer of the highest level; for the computing layers of levels other than the highest level in the decoding module, performing a third fusion process on the initial fusion features of the current level and all initial fusion features of levels higher than the current level to obtain a third fusion feature; inputting the third fusion feature to the computing layer of the current level; and performing a fourth fusion process on the features input to each computing layer according to the arrangement order of the computing layers in the decoding module to obtain an initial mask.

[0124] The image processing device provided in the embodiment of the present invention has the same technical features as the image processing method provided in the above embodiment, so it can also solve the same technical problems and achieve the same technical effects.

[0125] Corresponding to the above-mentioned training method embodiment of the mask image output model, this embodiment provides a structural schematic diagram of a training device for a mask image output model, such as Figure 6 As shown, the device comprises:

[0126] A determination module 61 is used to determine a sample image based on a preset training sample set; the sample image carries a mask label; a training module 62 is used to obtain image features of the sample image; a preset mask image output model is trained based on the image features and the mask label to obtain a trained mask image output model; wherein the mask image output model is used to: obtain an initial mask of the sample image through image features; obtain an edge area of ​​a target area to be extracted through the sample image and the initial mask; adjust the initial mask based on the edge area and the sample image to obtain a mask image of the sample image.

[0127] Furthermore, the above-mentioned mask image output model is also used to: extract multi-level features from image features; divide the multi-level features into multiple level groups, and perform a first fusion process on the level features in each level group to obtain an initial fusion feature; wherein the levels in each level group are adjacent; the number of levels of the initial fusion feature matches the number of levels of the decoding layer; perform a second fusion process on the initial fusion feature in a preset order to obtain a processing result; determine a loss value based on the processing result and the mask label; for the highest level decoding layer, update the parameters of the highest level decoding layer based on the loss value until the mask image output model converges to obtain the parameters of the highest level decoding layer; for decoding layers other than the highest level, update the parameters of the current decoding layer based on the parameters of all decoding layers with higher levels than the current decoding layer and the loss value until the mask image output model converges to obtain the parameters of the current decoding layer.

[0128] Furthermore, the above-mentioned mask image output model is also used to: update the parameters of the current decoding layer based on the loss value while fixing the parameters of all decoding layers higher than the current decoding layer until the mask image output model converges; update the parameters of the high-level decoding layers of the current decoding layer and the parameters of the current decoding layer based on the loss value until the mask image output model converges.

[0129] The training device for the mask image output model provided in the embodiment of the present invention has the same technical features as the training method for the mask image output model provided in the above embodiment, so it can also solve the same technical problems and achieve the same technical effects.

[0130] This embodiment also provides an electronic device, including a processor and a memory, wherein the memory stores machine executable instructions that can be executed by the processor, and the processor executes the machine executable instructions to implement the above-mentioned image processing method, or the training method of the mask image output model.

[0131] See also Figure 7 As shown, the electronic device includes a processor 100 and a memory 101, the memory 101 stores machine executable instructions that can be executed by the processor 100, and the processor 100 executes the machine executable instructions to implement the above-mentioned image processing method, or the training method of the mask image output model. The electronic device includes but is not limited to a user terminal and a server. When the electronic device is a server, after receiving the target image to be processed sent by the user terminal, the server can extract the target area from the target image by executing the above-mentioned image processing method, and feed back the target image after extracting the target area to the user terminal.

[0132] Further, Figure 7The electronic device shown further includes a bus 102 and a communication interface 103 , and the processor 100 , the communication interface 103 and the memory 101 are connected via the bus 102 .

[0133] Among them, the memory 101 may include a high-speed random access memory (RAM), and may also include a non-volatile memory (non-volatile memory), such as at least one disk storage. The communication connection between the system network element and at least one other network element is realized through at least one communication interface 103 (which can be wired or wireless), and the Internet, wide area network, local area network, metropolitan area network, etc. can be used. The bus 102 can be an ISA bus, a PCI bus or an EISA bus, etc. The bus can be divided into an address bus, a data bus, a control bus, etc. For ease of representation, Figure 7 Only one bidirectional arrow is used in the diagram, but this does not mean that there is only one bus or only one type of bus.

[0134] The processor 100 may be an integrated circuit chip with signal processing capabilities. In the implementation process, each step of the above method can be completed by the hardware integrated logic circuit or software instructions in the processor 100. The above processor 100 can be a general-purpose processor, including a central processing unit (CPU), a network processor (NP), etc.; it can also be a digital signal processor (DSP), an application-specific integrated circuit (ASIC), a field-programmable gate array (FPGA) or other programmable logic devices, discrete gates or transistor logic devices, discrete hardware components. The methods, steps and logic block diagrams disclosed in the embodiments of the present invention can be implemented or executed. The general-purpose processor can be a microprocessor or the processor can also be any conventional processor, etc. The steps of the method disclosed in conjunction with the embodiments of the present invention can be directly embodied as a hardware decoding processor for execution, or a combination of hardware and software modules in the decoding processor for execution. The software module may be located in a storage medium mature in the art, such as a random access memory, a flash memory, a read-only memory, a programmable read-only memory, or an electrically erasable programmable memory, a register, etc. The storage medium is located in the memory 101, and the processor 100 reads the information in the memory 101 and completes the steps of the method of the above embodiment in combination with its hardware.

[0135] This embodiment also provides a machine-readable storage medium, which stores machine-executable instructions. When the machine-executable instructions are called and executed by a processor, the machine-executable instructions prompt the processor to implement the above-mentioned image processing method, or the training method of the mask image output model.

[0136] The computer program product of the image processing method, mask image output model training method and device provided in the embodiments of the present invention includes a computer-readable storage medium storing program code. The instructions included in the program code can be used to execute the methods described in the previous method embodiments. The specific implementation can be found in the method embodiments, which will not be repeated here.

[0137] Those skilled in the art can clearly understand that, for the convenience and brevity of description, the specific working process of the system and device described above can refer to the corresponding process in the aforementioned method embodiment, and will not be repeated here.

[0138] In addition, in the description of the embodiments of the present invention, unless otherwise clearly specified and limited, the terms "installed", "connected", and "connected" should be understood in a broad sense, for example, it can be a fixed connection, a detachable connection, or an integral connection; it can be a mechanical connection or an electrical connection; it can be a direct connection or an indirect connection through an intermediate medium, and it can be the internal communication of two components. For those skilled in the art, the specific meanings of the above terms in the present invention can be understood according to specific circumstances.

[0139] If the functions are implemented in the form of software functional units and sold or used as independent products, they can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present invention, or the part that contributes to the prior art or the part of the technical solution, can be embodied in the form of a software product. The computer software product is stored in a storage medium, including several instructions for a computer device (which can be a personal computer, a server, or a network device, etc.) to perform all or part of the steps of the methods described in each embodiment of the present invention. The aforementioned storage medium includes: U disk, mobile hard disk, read-only memory (ROM, Read-Only Memory), random access memory (RAM, Random Access Memory), disk or optical disk, etc., which can store program codes.

[0140] In the description of the present invention, it should be noted that the terms "center", "upper", "lower", "left", "right", "vertical", "horizontal", "inner", "outer", etc., indicating the orientation or positional relationship, are based on the orientation or positional relationship shown in the drawings, and are only for the convenience of describing the present invention and simplifying the description, rather than indicating or implying that the device or element referred to must have a specific orientation, be constructed and operated in a specific orientation, and therefore cannot be understood as limiting the present invention. In addition, the terms "first", "second", and "third" are used for descriptive purposes only, and cannot be understood as indicating or implying relative importance.

[0141] Finally, it should be noted that the above embodiments are only specific implementations of the present invention, which are used to illustrate the technical solutions of the present invention, rather than to limit them. The protection scope of the present invention is not limited thereto. Although the present invention is described in detail with reference to the above embodiments, those skilled in the art should understand that any person skilled in the art can still modify the technical solutions recorded in the above embodiments within the technical scope disclosed by the present invention, or can easily think of changes, or make equivalent replacements for some of the technical features therein; and these modifications, changes or replacements do not make the essence of the corresponding technical solutions deviate from the spirit and scope of the technical solutions of the embodiments of the present invention, and should be included in the protection scope of the present invention. Therefore, the protection scope of the present invention shall be based on the protection scope of the claims.

Claims

1. A method for processing an image, It is characterized in that The method comprises: Get the feature data of the target image; Inputting the feature data into a pre-trained mask image output model, and outputting a mask image corresponding to the target image; Based on the mask image, extracting a target area from the target image; The mask image output model is used to: obtain an initial mask of the target image through the feature data; obtain an edge area of ​​the target area to be extracted through the target image and the initial mask; adjust the initial mask based on the edge area and the target image to obtain the mask image; The step of obtaining the edge area of ​​the target area to be extracted through the target image and the initial mask includes: performing a first splicing process on the target image and the initial mask to obtain first splicing data; performing a first convolution calculation on the first splicing data to obtain a first data feature of the first splicing data; identifying the foreground area and the background area in the target image based on the first data feature; performing corrosion and expansion processing on the edge of the foreground area to obtain the edge area; generating a region calibration image of the target image based on the foreground area, the background area and the edge area; wherein the region calibration image is used to indicate the position of the foreground area, the background area and the edge area in the target image; and extracting the edge area of ​​the target area from the target image according to the position of the edge area indicated by the region calibration image.

2. The method according to claim 1, It is characterized in that The step of adjusting the initial mask based on the edge area and the target image to obtain the mask image includes: The target image and the initial mask are subjected to a second splicing process to obtain second spliced ​​data; a second convolution operation is performed on the second spliced ​​data to obtain a second data feature of the second spliced ​​data; the mask image is obtained through the second data feature and the edge area, including: for each feature point in the second data feature, a feature value of the feature point in the second data feature and a value at a position corresponding to the position of the feature point in a regional calibration image used to indicate the position of the edge area are superimposed to obtain a superimposed feature value of the feature point; the superimposed feature values ​​of each of the feature points are combined into superimposed features, and a third convolution operation is performed on the superimposed features through a preset residual network to obtain the mask image.

3. The method according to claim 1, It is characterized in that The step of obtaining an initial mask of the target image through the feature data comprises: Extracting multi-level features from the feature data; Dividing the multi-level features into a plurality of level groups, performing a first fusion process on the level features in each level group to obtain an initial fusion feature; wherein the levels in each level group are adjacent; and the number of levels of the initial fusion feature matches the number of levels of the decoding layer; Perform a second fusion process on the initial fusion features in a preset order and output the initial mask.

4. The method according to claim 3, It is characterized in that The step of performing a second fusion process on the initial fusion features in a preset order and outputting the initial mask comprises: For the highest level decoding layer, the highest level initial fusion features among the initial fusion features are input to the highest level decoding layer; For decoding layers of levels other than the highest level, performing a third fusion process on the initial fusion features of the current level and all initial fusion features of levels higher than the current level to obtain a third fusion feature; and inputting the third fusion feature into the decoding layer of the current level; According to the arrangement order of each decoding layer, the fourth fusion process is performed on the features input to each decoding layer to obtain the initial mask.

5. A training method for a mask image output model, It is characterized in that The method comprises: Determining a sample image based on a preset training sample set; the sample image carries a mask label; Acquire the image features of the sample image; train a preset mask image output model based on the image features and the mask label to obtain a trained mask image output model; The mask image output model is used to: obtain an initial mask of the sample image through the image features; obtain an edge area of ​​the target area to be extracted through the sample image and the initial mask; adjust the initial mask based on the edge area and the sample image to obtain a mask image of the sample image; The step of obtaining the edge area of ​​the target area to be extracted through the sample image and the initial mask includes: performing a first splicing process on the sample image and the initial mask to obtain first splicing data; performing a first convolution calculation on the first splicing data to obtain a first data feature of the first splicing data; identifying the foreground area and the background area in the sample image based on the first data feature; performing corrosion and expansion processing on the edge of the foreground area to obtain the edge area; generating a region calibration image of the sample image based on the foreground area, the background area and the edge area; wherein the region calibration image is used to indicate the position of the foreground area, the background area and the edge area in the sample image; and extracting the edge area of ​​the target area from the sample image according to the position of the edge area indicated by the region calibration image.

6. The method according to claim 5, It is characterized in that The step of training a preset mask image output model based on the image features and the mask label to obtain a trained mask image output model includes: Extracting multi-level features from the image features; Dividing the multi-level features into a plurality of level groups, performing a first fusion process on the level features in each level group to obtain an initial fusion feature; wherein the levels in each level group are adjacent; and the number of levels of the initial fusion feature matches the number of levels of the decoding layer; Performing a second fusion process on the initial fusion features in a preset order to obtain a processing result; Determining a loss value based on the processing result and the mask label; For a decoding layer at a highest level, updating parameters of the decoding layer at the highest level based on the loss value until the mask image output model converges, thereby obtaining parameters of the decoding layer at the highest level; For decoding layers other than the highest level, the parameters of the current decoding layer are updated based on the parameters of all decoding layers at a level higher than the current decoding layer and the loss value until the mask image output model converges to obtain the parameters of the current decoding layer.

7. The method according to claim 6, It is characterized in that For decoding layers other than the highest level, based on parameters of all decoding layers at a level higher than the current decoding layer and the loss value, updating the parameters of the current decoding layer until the mask image output model converges, and obtaining the parameters of the current decoding layer comprises: Under the condition that parameters of all decoding layers higher than the current decoding layer are fixed, updating the parameters of the current decoding layer based on the loss value until the mask image output model converges; Based on the loss value, the parameters of the high-level decoding layer of the current decoding layer and the parameters of the current decoding layer are updated until the mask image output model converges.

8. A picture processing device, It is characterized in that The device comprises: An acquisition module is used to obtain feature data of a target image; An output module, used to input the feature data into a pre-trained mask image output model, and output a mask image corresponding to the target image; An extraction module, used for extracting a target area from the target image based on the mask image; The mask image output model is used to: obtain an initial mask of the target image through the feature data; obtain an edge area of ​​the target area to be extracted through the target image and the initial mask; adjust the initial mask based on the edge area and the target image to obtain the mask image; The step of obtaining the edge area of ​​the target area to be extracted through the target image and the initial mask includes: performing a first splicing process on the target image and the initial mask to obtain first splicing data; performing a first convolution calculation on the first splicing data to obtain a first data feature of the first splicing data; identifying the foreground area and the background area in the target image based on the first data feature; performing corrosion and expansion processing on the edge of the foreground area to obtain the edge area; generating a region calibration image of the target image based on the foreground area, the background area and the edge area; wherein the region calibration image is used to indicate the position of the foreground area, the background area and the edge area in the target image; and extracting the edge area of ​​the target area from the target image according to the position of the edge area indicated by the region calibration image.

9. A training device for a mask image output model, It is characterized in that The device comprises: A determination module, used to determine a sample image based on a preset training sample set; the sample image carries a mask label; A training module, used to obtain the image features of the sample image; train a preset mask image output model based on the image features and the mask label to obtain a trained mask image output model; The mask image output model is used to: obtain an initial mask of the sample image through the image features; obtain an edge area of ​​the target area to be extracted through the sample image and the initial mask; adjust the initial mask based on the edge area and the sample image to obtain a mask image of the sample image; The step of obtaining the edge area of ​​the target area to be extracted through the sample image and the initial mask includes: performing a first splicing process on the sample image and the initial mask to obtain first splicing data; performing a first convolution calculation on the first splicing data to obtain a first data feature of the first splicing data; identifying the foreground area and the background area in the sample image based on the first data feature; performing corrosion and expansion processing on the edge of the foreground area to obtain the edge area; generating a region calibration image of the sample image based on the foreground area, the background area and the edge area; wherein the region calibration image is used to indicate the position of the foreground area, the background area and the edge area in the sample image; and extracting the edge area of ​​the target area from the sample image according to the position of the edge area indicated by the region calibration image.

10. An electronic device, It is characterized in that It includes a processor and a memory, the memory stores machine executable instructions that can be executed by the processor, and the processor executes the machine executable instructions to implement the image processing method described in any one of claims 1 to 4, or the training method of the mask image output model described in any one of claims 5 to 7.

11. A machine-readable storage medium, It is characterized in that The machine-readable storage medium stores machine-executable instructions. When the machine-executable instructions are called and executed by the processor, the machine-executable instructions prompt the processor to implement the image processing method described in any one of claims 1 to 4, or the training method of the mask image output model described in any one of claims 5 to 7.

Citation Information

Patent Citations

  • Automatic fine segmentation method for image

    CN110706234A