Image processing method, device, medium and computing equipment

By fusing the target image and the preamble image mask input image processing model for feature fusion, the jitter and flicker problems caused by low accuracy of the image mask in the prior art are solved, and image mask prediction with higher accuracy is achieved.

CN114444599BActive Publication Date: 2025-05-13HANGZHOU NETEASE ZHIQI TECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202210100685.1
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-01-27
Publication Date
2025-05-13
Estimated Expiration
2042-01-27

AI Technical Summary

Technical Problem

In the prior art, when extracting images of portrait areas from images, the predicted image mask has low accuracy, resulting in problems of jitter and flickering when video playback.

Method used

By inputting the target image to be processed and the preamble image mask predicted from the preamble image to be trained image processing model, the feature fusion combined with spatial attention is performed, so that the target image mask corresponding to the portrait area in the target image is predicted.

Benefits of technology

Improve the accuracy of the target image mask and reduce the jitter and flickering of images in portrait areas during video playback.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114444599B_ABST
    Figure CN114444599B_ABST
Patent Text Reader

Abstract

The embodiments of the present disclosure provide an image processing method, apparatus, medium and computing device. The method includes: obtaining a target image to be processed; determining a preceding image corresponding to the target image; wherein the preceding image is an image that is located before the target image in time sequence; based on a trained first image processing model, predicting a preceding image mask corresponding to a portrait area in the preceding image from the preceding image; inputting the preceding image mask and the target image into the first image processing model, so that the first image processing model performs feature fusion of the preceding image mask and the target image in combination with spatial attention, and predicting a target image mask corresponding to a portrait area in the target image from the target image based on the fusion result. The present disclosure can improve the accuracy of an image mask corresponding to a portrait area in an image predicted from an image.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The embodiments of the present disclosure relate to the field of computer application technology. More specifically, the embodiments of the present disclosure relate to an image processing method, apparatus, medium and computing device. Background Art

[0002] This section is intended to provide a background or context to the embodiments of the disclosure that are recited in the claims. No description herein is admitted to be prior art by inclusion in this section.

[0003] With the development of image processing technology, image processing technology has been widely used in both civil and commercial business scenarios, especially in scenarios such as video calls, video surveillance, and virtual reality, playing an increasingly important role. Summary of the invention

[0004] In this context, embodiments of the present disclosure are intended to provide an image processing method, apparatus, medium, and computing device.

[0005] In a first aspect of the embodiments of the present disclosure, there is provided an image processing method, the method comprising:

[0006] Obtaining a target image to be processed;

[0007] Determine a preceding image corresponding to the target image; wherein the preceding image is an image that precedes the target image in time sequence;

[0008] Based on the trained first image processing model, predicting a preceding image mask corresponding to the portrait area in the preceding image from the preceding image;

[0009] The preceding image mask and the target image are input into the first image processing model, so that the first image processing model performs feature fusion on the preceding image mask and the target image in combination with spatial attention, and predicts a target image mask corresponding to a portrait area in the target image from the target image based on the fusion result.

[0010] In a second aspect of the embodiments of the present disclosure, there is provided an image processing device, the device comprising:

[0011] An acquisition module, used for acquiring a target image to be processed;

[0012] A determination module, used to determine a preceding image corresponding to the target image; wherein the preceding image is an image that precedes the target image in time sequence;

[0013] A first prediction module, configured to predict, from the preceding image, a preceding image mask corresponding to a portrait region in the preceding image based on a trained first image processing model;

[0014] The second prediction module is used to input the previous image mask and the target image into the first image processing model, so that the first image processing model performs feature fusion on the previous image mask and the target image in combination with spatial attention, and predicts a target image mask corresponding to the portrait area in the target image from the target image based on the fusion result.

[0015] In a third aspect of the embodiments of the present disclosure, a medium is provided, on which a computer program is stored, and when the computer program is executed by a processor, any of the above-mentioned image processing methods is implemented.

[0016] In a fourth aspect of the embodiments of the present disclosure, there is provided a computing device, comprising:

[0017] processor;

[0018] a memory for storing a processor executable program;

[0019] The processor implements any of the above-mentioned image processing methods by running the executable program.

[0020] According to an embodiment of the present disclosure, a target image to be processed and a preceding image mask corresponding to a portrait area in a preceding image that is temporally preceding the target image based on a trained first image processing model can be input into the first image processing model, so that the first image processing model performs feature fusion of the preceding image mask and the target image in combination with spatial attention, and predicts a target image mask corresponding to the portrait area in the target image from the target image based on the fusion result.

[0021] By adopting the above method, for the above target image, spatial attention can help focus on the important information in the target image. Since the above previous image mask and the target image are combined with the feature fusion of spatial attention, and the previous image mask corresponds to the portrait area in the above previous image, therefore, in this case, spatial attention can help focus on the portrait area in the target image, that is, with the help of the position of the previous image mask in the previous image, the image shape of the previous image mask and other information, focus on the portrait area in the target image. By using the previous image mask as auxiliary information to predict the target image mask corresponding to the portrait area in the target image from the target image, the accuracy of the predicted target image mask can be improved.

[0022] In addition, since the previous image mask is used as an aid when predicting the target image mask from the target image, the problem of flickering caused by a large difference between the target image mask and the previous image mask can be avoided. BRIEF DESCRIPTION OF THE DRAWINGS

[0023] The above and other objects, features and advantages of the exemplary embodiments of the present disclosure will become readily understood by reading the detailed description below with reference to the accompanying drawings. In the accompanying drawings, several embodiments of the present disclosure are shown in an exemplary and non-limiting manner, in which:

[0024] Figure 1 A schematic diagram schematically shows an application scenario of image processing according to an embodiment of the present disclosure;

[0025] Figure 2 A flowchart of an image processing method according to an embodiment of the present disclosure is schematically shown;

[0026] Figure 3 A schematic diagram schematically shows a first image processing model according to an embodiment of the present disclosure;

[0027] Figure 4 A flowchart of an image prediction method according to an embodiment of the present disclosure is schematically shown;

[0028] Figure 5 A schematic diagram of a module for performing feature fusion combined with attention according to an embodiment of the present disclosure is schematically shown;

[0029] Figure 6 Another flow chart of an image prediction method according to an embodiment of the present disclosure is schematically shown;

[0030] Figure 7 Schematically shows another schematic diagram of a module for performing feature fusion combined with attention according to an embodiment of the present disclosure;

[0031] Figure 8 A flowchart of a training method for a first image processing model according to an embodiment of the present disclosure is schematically shown;

[0032] Fig. 9 Another flow chart of an image prediction method according to an embodiment of the present disclosure is schematically shown;

[0033] Fig.10 A flowchart of a target image mask optimization method according to an embodiment of the present disclosure is schematically shown;

[0034] Fig.11A schematic diagram of a medium according to an embodiment of the present disclosure is schematically shown;

[0035] Fig.12 A block diagram schematically shows an image processing device according to an embodiment of the present disclosure;

[0036] Fig.13 A schematic diagram of a computing device according to an embodiment of the present disclosure is schematically shown.

[0037] In the drawings, the same or corresponding reference numerals represent the same or corresponding parts. DETAILED DESCRIPTION

[0038] The principles and spirit of the present disclosure will be described below with reference to several exemplary embodiments. It should be understood that these embodiments are provided only to enable those skilled in the art to better understand and implement the present disclosure, and are not intended to limit the scope of the present disclosure in any way. On the contrary, these embodiments are provided to make the present disclosure more thorough and complete, and to fully convey the scope of the present disclosure to those skilled in the art.

[0039] Those skilled in the art will appreciate that the embodiments of the present disclosure may be implemented as a system, device, apparatus, method or computer program product. Therefore, the present disclosure may be specifically implemented in the following forms, namely: complete hardware, complete software (including firmware, resident software, microcode, etc.), or a combination of hardware and software.

[0040] According to an embodiment of the present disclosure, an image processing method, apparatus, medium and computing device are proposed.

[0041] It should be understood herein that any number of elements in the drawings is for illustration rather than limitation, and any naming is only for distinction rather than having any limiting meaning.

[0042] The principle and spirit of the present disclosure are explained in detail below with reference to several representative embodiments of the present disclosure. SUMMARY OF THE INVENTION

[0044] In practical applications, for an image containing a portrait, it is very likely that there will be a need to extract the image of the portrait area in the image for subsequent operations such as portrait recognition and background replacement.

[0045] In the related art, usually only one image is processed to predict an image mask corresponding to the portrait area in the image, and based on the predicted image mask, the image of the portrait area is extracted from the image. However, the image mask predicted from the image in this way has low accuracy.

[0046] In addition, in the related art, there is often a need to extract images of portrait areas from images in an image sequence such as a video.

[0047] Taking a video as an example, assuming that the person in the video is stationary or has very small movements, the actual difference between the images of the portrait areas in the two adjacent frames in the video will be relatively subtle and difficult to detect with the naked eye. However, when the images of the portrait areas are extracted from the two frames, the accuracy of the image mask predicted from the image is low because only one of the images is processed at a time. Therefore, it is very likely that the difference between the images of the portrait areas extracted from the two frames is greater than the actual difference, resulting in visual jitter and flickering of the images of the portrait areas when the images of the portrait areas extracted from the two frames are played in the form of a video.

[0048] In order to solve the above problems, the present disclosure provides a technical solution for image processing. In the technical solution, a target image to be processed and a preceding image mask corresponding to a portrait area in the preceding image predicted from a preceding image temporally preceding the target image based on a trained first image processing model can be input into the first image processing model, so that the first image processing model performs feature fusion of the preceding image mask and the target image in combination with spatial attention, and predicts a target image mask corresponding to the portrait area in the target image from the target image based on the fusion result.

[0049] Usually, the introduction of the attention mechanism is to filter out a small amount of important information from a large amount of information, focus on these important information, and ignore most of the unimportant information. Among them, the attention weight represents the importance of the corresponding information.

[0050] For the above target image, spatial attention can help focus on important information in the target image. Since the above previous image mask and the target image are combined with spatial attention feature fusion, and the previous image mask corresponds to the portrait area in the above previous image, in this case, spatial attention can help focus on the portrait area in the target image, that is, with the help of the position of the previous image mask in the previous image, the image shape of the previous image mask and other information, focus on the portrait area in the target image. By using the previous image mask as auxiliary information to predict the target image mask corresponding to the portrait area in the target image from the target image, the accuracy of the predicted target image mask can be improved.

[0051] In addition, since the previous image mask is used as an aid when predicting the target image mask from the target image, the problem of flickering caused by a large difference between the target image mask and the previous image mask can be avoided.

[0052] After introducing the basic principles of the present disclosure, various non-limiting embodiments of the present disclosure are described in detail below.

[0053] Application Scenario Overview

[0054] First reference Figure 1 , Figure 1 A schematic diagram schematically shows an application scenario of image processing according to an embodiment of the present disclosure.

[0055] like Figure 1 As shown, in the application scenario of image processing, a server and at least one client (eg, client 1-N) connected to the server may be included.

[0056] Usually, users can install the client corresponding to a certain application in the device they use; the device can be a terminal device such as a smart phone, tablet computer, PDA, laptop, PC (Personal Computer), smart wearable device, smart car device or game console. The above-mentioned server can be deployed on a single server or server cluster; or, the server can also be a server built based on cloud computing services.

[0057] The device equipped with the above-mentioned server or the above-mentioned client can also be installed with embedded cameras, external cameras and other camera hardware for collecting images or videos, so that these camera hardware can be called to realize image or video collection.

[0058] Exemplary Methods

[0059] Combine the following Figure 1 For application scenarios, refer to Figure 2-Figure 10 To describe the method for image processing according to an exemplary embodiment of the present disclosure. It should be noted that the above application scenarios are only shown to facilitate understanding of the spirit and principles of the present disclosure, and the embodiments of the present disclosure are not limited in this regard. On the contrary, the embodiments of the present disclosure can be applied to any applicable scenario.

[0060] refer to Figure 2 , Figure 2 The flowchart of an image processing method according to an embodiment of the present disclosure is schematically shown.

[0061] The above-mentioned image processing method can be applied to the above-mentioned server or the above-mentioned client equipped with a machine learning model for image prediction (which can be called the first image processing model).

[0062] For the above-mentioned server and the above-mentioned client, usually, the hardware resources of the single server or server cluster and other devices where the server is located are richer than the hardware resources of the terminal device where the above-mentioned client is located. Therefore, the computing power of the server is stronger than that of the client, and the complexity of the machine learning model that can run in the server is also higher than the complexity of the machine learning model that can run in the client.

[0063] In addition, for machine learning models that need to perform real-time computing tasks (for example, real-time image processing tasks during video calls, real-time audio processing tasks during voice calls, etc.), the computing rate requirements for such machine learning models are usually high, that is, such machine learning models are required to be able to output calculation results in a shorter time.

[0064] In order to enable the first image processing model to run in both the server and the client, and also to improve the calculation rate of the first image processing model, the complexity of the first image processing model can be appropriately reduced.

[0065] Specifically, in one embodiment shown, in addition to the above-mentioned first image processing model, another machine learning model for image prediction (which may be referred to as a second image processing model) may be pre-set; wherein the model structure of the first image processing model may be simpler than the model structure of the second image processing model, for example: the first image processing model may be obtained by deleting the convolutional layers in the second image processing model, that is, the number of convolutional layers in the first image processing model may be less than the number of convolutional layers in the second image processing model.

[0066] In practical applications, both the first image processing model and the second image processing model can be deeplabV3+ models. The backbone network of the second image processing model can adopt mobilenetV3, and the first image processing model can be obtained by cutting out the repeated structure in the second image processing model.

[0067] In the above case, the second image processing model can be trained first. Subsequently, the first image processing model can learn the knowledge of the trained second image processing model through knowledge distillation, so that the model effect of the first image processing model is similar to the model effect of the second image processing model. That is, the model parameters migrated from the second image processing model through knowledge distillation can be used as the model parameters of the first image processing model.

[0068] In one example, the first image processing model and the second image processing model can be pre-set in the server. After completing knowledge distillation in the server, the first image processing model obtained by knowledge distillation can be redeployed in the server or the client to perform subsequent image processing tasks.

[0069] In another example, the first image processing model can be pre-set in the client, and the second image processing model can be pre-set in the server. After knowledge distillation is completed through data exchange between the client and the server, subsequent image processing tasks can be performed based on the first image processing model obtained through knowledge distillation in the client.

[0070] The above image processing method may include the following steps:

[0071] Step 201: Acquire a target image to be processed.

[0072] Step 202: Determine a preceding image corresponding to the target image; wherein the preceding image is an image that precedes the target image in time sequence.

[0073] In this embodiment, an image to be processed (which may be referred to as a target image) may be acquired, and a preceding image corresponding to the target image may be determined; wherein the preceding image may be an image that is located before the target image in terms of time sequence.

[0074] In one embodiment shown, the target image may be any frame image except the first frame in any image sequence, and the preceding image corresponding to the target image may be an image in the image sequence that is located before the target image and has a frame number that is a preset threshold (referred to as the first threshold) between the image and the target image.

[0075] In practical applications, the first threshold may be preset by a technician, and the present disclosure does not impose any limitation on this.

[0076] For example, the above-mentioned image sequence can be a video. Generally, multiple images can be collected at a certain time interval within a period of time, and these images can be sorted according to the order of the collection time, so as to combine these arranged images into a video. In this case, the above-mentioned target image can be any frame image in the video except the first frame, and the previous image corresponding to the target image can be an image in the video that is located before the target image and the number of frames between the target image and the image is the above-mentioned first threshold. For example, assuming that the video is collected at a time interval of 50 milliseconds within 2 seconds, 40 images can be collected, and these 40 images can be combined into a video in the order of the collection time; assuming that the above-mentioned target image is the 20th frame image in the video, and the above-mentioned first threshold is 1, then the previous image corresponding to the target image is the 19th frame image in the video.

[0077] It should be noted that, for the first frame image in the above image sequence, a blank image can be used as a preceding image corresponding to the image.

[0078] In another embodiment shown, the target image may be any image that does not belong to any image sequence. Since the target image does not belong to any image sequence and is an independent single image, it is impossible to select a frame of image from the image sequence as a preceding image corresponding to the target image. In this case, an affine transformation may be performed on the target image based on a preset image transformation relationship, and the image obtained by the affine transformation may be used as a preceding image corresponding to the target image.

[0079] For a person in motion, if the person is photographed, the position and posture of the person in the images taken at different times may be different. Therefore, for the above target image and the previous image corresponding to the target image obtained by affine transformation, two frames of images recording the person in motion can be simulated.

[0080] In practical applications, the above-mentioned image transformation relationship may include parameters such as linear transformation parameters and translation parameters. The specific values ​​of these parameters may be preset by technical personnel, and the present disclosure does not impose any limitation on this.

[0081] Step 203: Based on the trained first image processing model, a preceding image mask corresponding to the portrait area in the preceding image is predicted from the preceding image.

[0082] Step 204: Input the previous image mask and the target image into the first image processing model, so that the first image processing model performs feature fusion on the previous image mask and the target image in combination with spatial attention, and predicts a target image mask corresponding to the portrait area in the target image from the target image based on the fusion result.

[0083] In this embodiment, for the preceding image corresponding to the above-mentioned target image, before performing image processing on the target image, the preceding image may also be processed first, that is, based on the trained first image processing model, an image mask (which may be called a preceding image mask) corresponding to the portrait area in the preceding image may be predicted from the preceding image; wherein the portrait area is the image area containing the portrait.

[0084] It should be noted that the image mask corresponding to the portrait region in a certain image can represent information such as the shape, size, and position of the portrait region in the image. In this case, the image mask can be used to extract an image of a partial region from the image, and the extracted image of the partial region can be considered as the image of the portrait region in the image.

[0085] When the preceding image is determined, the preceding image mask may be further obtained, and the preceding image mask and the obtained target image may be input into the first image processing model. In this case, the first image processing model may perform feature fusion combining spatial attention on the preceding image mask and the target image, and predict an image mask (referred to as a target image mask) corresponding to the portrait region in the target image from the target image based on the fusion result.

[0086] Similar to the above process, if the target image is an image in the image sequence, then for a certain image (which may be referred to as a subsequent image) that is temporally located after the target image in the image sequence, the target image may continue to be used as a preceding image corresponding to the subsequent image. In this case, when performing image processing on the subsequent image, the target image mask and the subsequent image may be input into the first image processing model, so that the first image processing model performs feature fusion of the target image mask and the subsequent image in combination with spatial attention, and predicts an image mask corresponding to the portrait area in the subsequent image from the subsequent image based on the fusion result. By analogy, the image processing of each image in the image sequence may be completed.

[0087] In order to use the image mask to extract the image of the portrait area from the image, in one embodiment shown, after obtaining the target image mask, the target image mask can be further binarized to obtain a corresponding binary mask. Subsequently, the image of the portrait area can be extracted from the target image based on the binary mask.

[0088] In practical applications, the extracted image of the portrait area can be fused to the background image to obtain a fused image; wherein the background image can be different from the image of the background area in the target image to achieve background replacement for the target image. In this case, if the target image is an image in the image sequence, then the background replacement for each image in the image sequence can be completed; further, since the complexity of the first image processing model is relatively low and the calculation time is relatively short, real-time background replacement for the image sequence can be achieved.

[0089] According to the above embodiment, a preceding image mask corresponding to the portrait area in a preceding image that is temporally preceding the target image can be feature fused with the target image in combination with spatial attention, and a target image mask corresponding to the portrait area in the target image can be predicted from the target image based on the fusion result.

[0090] Usually, the introduction of the attention mechanism is to filter out a small amount of important information from a large amount of information, focus on these important information, and ignore most of the unimportant information. Among them, the attention weight represents the importance of the corresponding information.

[0091] For the above target image, spatial attention can help focus on important information in the target image. Since the above previous image mask and the target image are combined with spatial attention feature fusion, and the previous image mask corresponds to the portrait area in the above previous image, in this case, spatial attention can help focus on the portrait area in the target image, that is, with the help of the position of the previous image mask in the previous image, the image shape of the previous image mask and other information, focus on the portrait area in the target image. By using the previous image mask as auxiliary information to predict the target image mask corresponding to the portrait area in the target image from the target image, the accuracy of the predicted target image mask can be improved.

[0092] In addition, since the previous image mask is used as an aid when predicting the target image mask from the target image, the problem of flickering caused by a large difference between the target image mask and the previous image mask can be avoided.

[0093] The following is an explanation of the process in step 204 where the first image processing model predicts the target image mask from the target image.

[0094] In one embodiment shown, reference Figure 3 , Figure 3 A schematic diagram of a first image processing model according to an embodiment of the present disclosure is schematically shown.

[0095] The first image processing model may include at least two convolutional layers, which may be referred to as a first convolutional layer and a second convolutional layer. Figure 3 As shown, the first image processing model may include five convolutional layers, an ASPP (Atrous Spatial Pyramid Pooling) module and a module for feature fusion combined with attention.

[0096] In the process of machine learning, the features learned by a convolutional layer are usually local; and the higher the number of convolutional layers, the more global the learned features are. In the above-mentioned first image processing model, the above-mentioned first convolutional layer can be used to extract feature data related to the image details of the image input to the first image processing model (which can be called first feature data), and the above-mentioned second convolutional layer can be used to extract feature data related to the image semantics of the image input to the first image processing model (which can be called second feature data).

[0097] Combined with Figure 3 The first image processing model shown above, refer to Figure 4 , Figure 4 The figure schematically shows a flowchart of an image prediction method according to an embodiment of the present disclosure.

[0098] The above image prediction method may include the following steps:

[0099] Step 401: Input the target image and the previous image mask into the first image processing model so that the first image processing model executes the following steps 402-405.

[0100] Step 402: Obtain the first feature data extracted by the first convolutional layer for the target image, and obtain the second feature data extracted by the second convolutional layer for the target image.

[0101] Step 403: performing feature fusion on the first feature data and the second feature data to obtain first fused feature data.

[0102] Step 404: Based on the spatial attention mechanism and the previous image mask, calculate the spatial attention weight corresponding to the target image, and multiply the spatial attention weight by the first fused feature data to obtain weighted first fused feature data.

[0103] Step 405: performing feature fusion on the preceding image mask and the weighted first fused feature data to obtain second fused feature data, and further performing feature extraction on the second fused feature data to obtain a target image mask corresponding to the portrait area in the target image.

[0104] For the first image processing model, when the preceding image mask and the target image are input into the first image processing model, first, on the one hand, the first convolution layer in the first image processing model can extract the first feature data related to the image details of the target image for the target image; on the other hand, the second convolution layer in the first image processing model can extract the second feature data related to the image semantics of the target image for the target image. Subsequently, the first feature data and the second feature data can be input into the module for performing feature fusion combined with attention, so that the module performs feature fusion combined with spatial attention on the first feature data and the second feature data.

[0105] Specifically, refer to Figure 5 , Figure 5 A schematic diagram of a module for performing feature fusion combined with attention according to an embodiment of the present disclosure is schematically shown.

[0106] like Figure 5 As shown, the module for performing feature fusion combined with attention can first perform feature fusion on the first feature data and the second feature data (i.e., the process shown in 501 in the figure) to obtain fused feature data (which can be referred to as first fused feature data). Then, based on the spatial attention mechanism and the previous image mask, the spatial attention weight corresponding to the target image can be calculated (i.e., the process shown in 502 in the figure), and the calculated spatial attention weight can be multiplied by the first fused feature data (i.e., the process shown in 503 in the figure) to obtain the weighted first fused feature data. Finally, the previous image mask and the weighted first fused feature data can be feature fused (i.e., the process shown in 504 in the figure) to obtain fused feature data (which can be referred to as second fused feature data).

[0107] When the module for performing feature fusion combined with attention obtains the second fused feature data, the second fused feature data can be input into the last convolution layer (i.e., Figure 3The fifth convolution layer shown in FIG. 1 is used to further extract features from the second fused feature data. In this case, the data output by the last convolution layer can be used as the target image mask.

[0108] Usually, an image can be viewed as a matrix of data; the pixel values ​​of each pixel in the image are usually in the range of [0, 255]. In this case, continue to refer to Figure 5 When calculating the above-mentioned spatial attention weight based on the spatial attention mechanism and the above-mentioned previous image mask, the module for performing feature fusion combined with attention can specifically divide the pixel value of each pixel in the previous image mask by 255 to obtain the corresponding normalized matrix, and determine the normalized matrix as the spatial attention weight.

[0109] Continue to refer Figure 5 When the module for feature fusion combined with attention performs feature fusion on the first feature data and the second feature data, it can specifically calculate the sum of the first feature data and the second feature data, and determine the sum of the calculated feature data as the first fused feature data; similarly, when performing feature fusion on the previous image mask and the weighted first fused feature data, it can specifically calculate the sum of the previous image mask and the weighted first fused feature data, and determine the sum of the calculated feature data as the second fused feature data.

[0110] That is, if X represents the first feature data, Y represents the second feature data, and Z represents the previous image mask, then Figure 5 The above-mentioned second fusion feature data shown = Z + (Z / 255) * (X + Y).

[0111] For a certain convolution layer, after a certain image is input into the convolution layer for feature extraction, the number of images output by the convolution layer is equal to the number of convolution channels of the convolution layer. That is, each convolution channel of the convolution layer will output feature data obtained by extracting features from the image through the convolution channel. In this case, channel attention can help focus on the convolution channel that outputs important information. For the above-mentioned target image, in addition to combining the feature fusion of the previous image mask and the target image with spatial attention, the target image can also be combined with channel attention feature fusion according to the number of convolution channels of the above-mentioned first convolution layer and the number of convolution channels of the above-mentioned second convolution layer in the above-mentioned first image processing model, so as to further improve the accuracy of the predicted target image mask.

[0112] In another embodiment shown, in combination with Figure 3 The first image processing model shown above, refer to Figure 6 , Figure 6 The flowchart of another image prediction method according to an embodiment of the present disclosure is schematically shown.

[0113] The above image prediction method may include the following steps:

[0114] Step 601: Input the target image and the previous image mask into the first image processing model so that the first image processing model executes the following steps 602-606.

[0115] Step 602: Obtain the first feature data extracted by the first convolutional layer for the target image, and obtain the second feature data extracted by the second convolutional layer for the target image.

[0116] Step 603: performing feature fusion on the first feature data and the second feature data to obtain first fused feature data.

[0117] Step 604: Based on the spatial attention mechanism and the previous image mask, calculate the spatial attention weight corresponding to the target image, and multiply the spatial attention weight by the first fused feature data to obtain weighted first fused feature data.

[0118] Step 605: Based on the channel attention mechanism and the first fused feature data, calculate the channel attention weights corresponding to the first feature data and the second feature data, and multiply the first feature data and the second feature data by the channel attention respectively to obtain the weighted first feature data and the weighted second feature data.

[0119] Step 606: Perform feature fusion on the previous image mask, the weighted first fused feature data, the weighted first feature data and the weighted second feature data to obtain second fused feature data, and further perform feature extraction on the second fused feature data to obtain a target image mask corresponding to the portrait area in the target image.

[0120] The specific implementation of the above steps 601-604 and 606 can refer to the above steps 401-405. Figure 4 Similar to the image prediction method shown, the first feature data and the second feature data can be input into the module for performing feature fusion combining attention, so that the module performs feature fusion combining spatial attention and channel attention on the first feature data and the second feature data.

[0121] Specifically, refer to Figure 7 , Figure 7A schematic diagram of another module for performing feature fusion combined with attention according to an embodiment of the present disclosure is schematically shown.

[0122] like Figure 7 As shown, the module for performing feature fusion combined with attention can first perform feature fusion on the first feature data and the second feature data (i.e., the process shown in 701 in the figure) to obtain fused feature data (which can be referred to as first fused feature data). Then, on the one hand, based on the spatial attention mechanism and the previous image mask, the spatial attention weight corresponding to the target image can be calculated (i.e., the process shown in 702 in the figure), and the calculated spatial attention weight is multiplied with the first fused feature data (i.e., the process shown in 703 in the figure) to obtain the weighted first fused feature data; on the other hand, based on the channel attention mechanism and the first fused feature data, the channel attention weight corresponding to both the first feature data and the second feature data can be calculated (i.e., the process shown in 704 in the figure), and the first feature data and the second feature data are respectively multiplied with the calculated channel attention weight (i.e., the process shown in 705 in the figure) to obtain the weighted first feature data and the weighted second feature data. Finally, the previous image mask, the weighted first fused feature data, the weighted first feature data and the weighted second feature data may be subjected to feature fusion (i.e., the processing shown in 706 in the figure) to obtain the above-mentioned second fused feature data.

[0123] Usually, after the convolution layer performs convolution processing on the input image, the output data is still in the form of a matrix, and the number of output matrices is the number of convolution channels corresponding to the convolution layer. In this case, the same number of convolution channels can be set for the first convolution layer and the second convolution layer, and other model parameters of the first image processing model can be adjusted so that the data format of the first feature data output by the first convolution layer and the second feature data output by the second convolution layer are the same. For example, assuming that the first feature data is 32 64*64 matrices, the second feature data is also 32 64*64 matrices.

[0124] Continue to refer Figure 7, the module for feature fusion combined with attention calculates the channel attention weight based on the channel attention mechanism and the first fused feature data. Specifically, the first fused feature data can be firstly subjected to global average pooling (Global Average Pooling, GAP) to obtain pooled data; then, based on a linear rectified (Rectified Linear Unit, ReLU) function as an activation function corresponding to the first image processing model, the pooled data is mapped to the first mapping data; finally, based on another Sigmoid function as an activation function corresponding to the first image processing model, the first mapping data is further mapped to the second mapping data. At this point, the second mapping data can be determined as the channel attention weight.

[0125] In order to reduce the amount of calculation, after obtaining the above-mentioned pooled data, the pooled data can be firstly reduced in dimension based on a fully connected layer (FC) to obtain the pooled data after dimension reduction; then based on the above-mentioned linear rectification function, the pooled data after dimension reduction is mapped to the first mapping data; then based on another fully connected layer, the first mapping data is dimensionally upgraded to obtain the first mapping data after dimension upgrade, so that the dimension of the first mapping data after dimension upgrade is equal to the dimension of the pooled data; finally, based on the above-mentioned Sigmoid function, the first mapping data after dimension upgrade is further mapped to the second mapping data.

[0126] Continue to refer Figure 7 When the module for performing feature fusion combined with attention performs feature fusion on the preceding image mask, the weighted first fused feature data, the weighted first feature data and the weighted second feature data, it can specifically calculate the sum of the preceding image mask, the weighted first fused feature data, the weighted first feature data and the weighted second feature data, and determine the sum of the calculated feature data as the second fused feature data.

[0127] That is, assuming that X represents the first feature data, Y represents the second feature data, and Z represents the preceding image mask, then Figure 7 The above-mentioned second fusion feature data shown = Z+(Z / 255)*(X+Y)+X*Sigmoid(ReLu(GAP(X+Y)))+Y*Sigmoid(ReLu(GAP(X+Y))).

[0128] Or, in the case of reverse dimension reduction and dimension increase through two fully connected layers, such as Figure 7The second fused feature data shown above = Z+(Z / 255)*(X+Y)+X*Sigmoid(FC(ReLu(FC(GAP(X+Y)))))+Y*Sigmoid(FC(ReLu(FC(GAP(X+Y))))).

[0129] It should be noted that, for the above-mentioned first image processing model, Figure 2 , Figure 4 , Figure 6 The embodiments shown are all implemented based on the trained first image processing model.

[0130] In practical applications, the first image processing model obtained through knowledge distillation can be used as the trained first image processing model; alternatively, the first image processing model obtained through knowledge distillation can be further trained to optimize the model effect of the first image processing model, and the trained first image processing model can be used as the final first image processing model.

[0131] The training process for the first image processing model is described below.

[0132] In one embodiment shown, reference Figure 8 , Figure 8 A flowchart of a training method for a first image processing model according to an embodiment of the present disclosure is schematically shown.

[0133] The training method for the first image processing model may include the following steps:

[0134] Step 801: Acquire a labeled target image sample; wherein the labeled image sample is labeled with an image mask corresponding to the portrait area therein.

[0135] Step 802: Determine an image mask corresponding to the portrait area in a preceding image sample corresponding to the labeled target image sample.

[0136] Step 803: Input the image mask in the previous image sample and the labeled target image sample into the first image processing model, so that the first image processing model performs feature fusion on the image mask in the previous image sample and the labeled target image sample in combination with spatial attention, and predicts the image mask corresponding to the portrait area in the labeled target image sample from the labeled target image sample based on the fusion result.

[0137] Step 804: Through back propagation, based on the image mask predicted by the first image processing model from the labeled image samples and the labeled image masks in the labeled image samples, the model parameters of the first image processing model are adjusted to complete the training of the first image processing model.

[0138] In order to further train the first image processing model, firstly, a number of annotated images may be obtained as annotated samples (which may be referred to as target image samples) corresponding to the target image.

[0139] For any labeled target image sample, the labeled target image sample is labeled with an image mask corresponding to the portrait region therein. Subsequently, a previous image (which may be referred to as a previous image sample) corresponding to the labeled target image sample may be determined, and an image mask corresponding to the portrait region in the previous image sample may be obtained; wherein the image mask corresponding to the portrait region in the previous image sample may also be an image mask labeled for the previous image.

[0140] In the above case, the first image processing model may be supervised trained based on the several labeled target image samples and image masks corresponding to the portrait areas in the preceding image samples respectively corresponding to the several labeled target image samples.

[0141] Specifically, for any labeled target image sample, the image mask corresponding to the portrait area in the preceding image sample corresponding to the labeled target image sample, and the labeled target image sample can be input into the above-mentioned first image processing model, so that the first image processing model performs feature fusion combined with spatial attention (or feature fusion combined with spatial attention and channel attention) on the image mask corresponding to the portrait area in the preceding image sample and the labeled target image sample, and predicts the image mask corresponding to the portrait area in the labeled target image sample from the labeled target image sample based on the fusion result.

[0142] Finally, through back propagation, based on the labeled image masks in the above-mentioned several labeled target image samples and the image masks predicted by the above-mentioned first image processing model from these several labeled image samples, the model parameters of the first image processing model can be adjusted to complete the training of the first image processing model.

[0143] Furthermore, in one embodiment shown, the target image sample may be any frame image in any image sequence (referred to as an image sequence sample). The preceding image sample corresponding to the target image sample may be an image in the image sequence sample whose frame number of images spaced apart from the target image sample does not exceed a preset threshold value (referred to as a second threshold value). In this way, since any image whose interval does not exceed the second threshold value may be used as the preceding image sample, the first image processing model may be adapted to different degrees of variation of the preceding image mask, thereby further improving the accuracy of the target image mask predicted from the target image using the preceding image mask.

[0144] In practical applications, the second threshold may be preset by a technician, and the present disclosure does not limit this.

[0145] For example, assuming that a video has a total of n frames of images, and the target image sample is the mth frame of the video, if m is less than 30, then a frame of image can be randomly selected from the 0th frame to the 30th frame as the preceding image sample corresponding to the target image sample; if m is greater than n-30, then a frame of image can be randomly selected from the n-30th frame to the nth frame as the preceding image sample corresponding to the target image sample; if m is between 30 and n-30, then a frame of image can be randomly selected from the m-30th frame to the m+30th frame as the preceding image sample corresponding to the target image sample.

[0146] In another embodiment shown, the target image sample may be any image. Based on a preset image transformation relationship, the target image sample may be affine transformed, and the image obtained by the affine transformation may be used as a preceding image sample corresponding to the target image sample.

[0147] In practical applications, specific values ​​of parameters such as linear transformation parameters and translation parameters in the above image transformation relationship can be preset by technicians, and the present disclosure does not impose any restrictions on this.

[0148] When the above-mentioned preceding image mask is used as auxiliary information to predict the above-mentioned target image mask from the above-mentioned target image, if the accuracy of the preceding image mask is poor, the accuracy of the predicted target image mask will be affected. In this case, when the above-mentioned first image processing model is supervised trained, a part of the image masks corresponding to the portrait area in the preceding image samples corresponding to the above-mentioned several labeled target image samples can be set to 0 to simulate the situation where the accuracy of the preceding image mask is poor or there is no preceding image mask. In this way, the robustness of the above-mentioned first image processing model can be improved.

[0149] In an embodiment shown, when the first image processing model is supervisedly trained based on the aforementioned several labeled target image samples and the image masks corresponding to the portrait area in the preceding image samples corresponding to the aforementioned several labeled target image samples, a portion of the image masks may be selected from the plurality of image masks according to a preset ratio (which may be referred to as the first ratio), and the pixel value of each pixel point in each image mask in the selected portion of the image masks is set to 0. Subsequently, the first image processing model may be supervisedly trained based on the portion of the image masks set to 0, the remaining image masks not set to 0, and the aforementioned several labeled target image samples.

[0150] In practical applications, the first ratio may be preset by a technician, and the present disclosure does not limit this.

[0151] When predicting the target image mask from the target image, there may be a problem of missed prediction, that is, the predicted target image mask corresponds not only to the portrait area in the target image, but also to a part of the background area in the target image. For example, if the background area in the target image contains the back of other people, the target image mask predicted from the target image may be an image mask corresponding to both the portrait area and the background area where the back is located. In this case, when the first image processing model is supervised, background replacement can be performed on a part of the several labeled target image samples, so that the first image processing model can adapt to the situation where the same portrait is in the background corresponding to different scenes, so as to achieve the purpose of enriching the scene data. In this way, the possibility of missed prediction can be reduced.

[0152] In one embodiment shown, when the first image processing model is supervisedly trained based on the aforementioned several labeled target image samples and the image masks corresponding to the portrait area in the preceding image samples corresponding to the aforementioned several labeled target image samples, it is also possible to select some target image samples from the aforementioned several labeled target image samples according to a preset ratio (which may be referred to as a second ratio), and replace the image of the background area in each target image sample in the aforementioned several target image samples with any background image in the preset background image set. Subsequently, the first image processing model may be supervisedly trained based on the aforementioned several image masks, the aforementioned target image samples whose backgrounds are replaced, and the remaining target image samples whose backgrounds are not replaced.

[0153] In practical applications, the second ratio may be preset by a technician, and the background image set may also be preset by a technician, and the present disclosure does not limit this.

[0154] When the first image processing model is subjected to supervised training, it may not be possible to obtain enough labeled target image samples within a period of time. Therefore, in one embodiment shown, in addition to the supervised training of the first image processing model based on the aforementioned several labeled target image samples and the image masks corresponding to the portrait area in the preceding image samples corresponding to the several labeled target image samples, several unlabeled images may be obtained as the unlabeled target image samples corresponding to the target images, and the image masks corresponding to the portrait area in the preceding image samples corresponding to the unlabeled target image samples may be determined.

[0155] For any unlabeled target image sample, the image mask corresponding to the portrait area in the preceding image sample corresponding to the unlabeled target image sample and the unlabeled target image sample can be input into the above-mentioned first image processing model, so that the first image processing model performs feature fusion combining spatial attention (or feature fusion combining spatial attention and channel attention) on the image mask corresponding to the portrait area in the preceding image sample and the unlabeled target image sample, and predicts the image mask corresponding to the portrait area in the unlabeled target image sample from the unlabeled target image sample based on the fusion result. Subsequently, the unlabeled target image sample can be labeled based on the image mask predicted by the first image processing model from the unlabeled image sample, so that the unlabeled target image sample can be updated to a labeled target image sample, and the first image processing model can be supervisedly trained again based on the labeled target image sample and the image mask corresponding to the portrait area in the preceding image sample corresponding to the labeled target image sample. In this way, the workload of labeling the target image sample can also be reduced.

[0156] For the above-mentioned first image processing model, the larger the image sizes of the above-mentioned previous image mask and the above-mentioned target image input to the first image processing model are, the longer the calculation time of the first image processing model will be, that is, the longer it will take for the first image processing model to output the above-mentioned target image mask predicted from the target image.

[0157] In practical applications, the image size of an image can be reduced by downsampling the image, and the image size of the image can be increased by upsampling the image.

[0158] In the above case, when executing the above step 204, in order to reduce the calculation time of the above-mentioned first image processing model, the input image corresponding to the first image processing model can be upsampled; in order to ensure the consistency of the image size of the image mask corresponding to the portrait area predicted from the above-mentioned target image, the output image corresponding to the first image processing model can be downsampled.

[0159] In one embodiment shown, reference Fig. 9 , Fig. 9 The flowchart of another image prediction method according to an embodiment of the present disclosure is schematically shown.

[0160] The above image prediction method may include the following steps:

[0161] Step 901: down-sample the target image to obtain a down-sampled target image.

[0162] Step 902: Input the previous image mask and the downsampled target image into the first image processing model, so that the first image processing model performs feature fusion on the previous image mask and the downsampled target image in combination with spatial attention, and predicts a target image mask corresponding to the portrait area in the downsampled target image from the downsampled target image based on the fusion result.

[0163] Step 903: up-sample the target image mask to obtain an image mask corresponding to the portrait area in the target image.

[0164] For the preceding image corresponding to the above-mentioned target image, the preceding image can be first downsampled to obtain a downsampled preceding image, and then based on the above-mentioned first image processing model, an image mask corresponding to the portrait area in the downsampled preceding image can be predicted from the downsampled preceding image; wherein, the image mask can be used as the above-mentioned preceding image mask.

[0165] When the preceding image is determined, the preceding image mask can be further obtained. At this time, the target image can be downsampled to obtain a downsampled target image, and then the preceding image mask and the downsampled target image can be input into the first image processing model. In this case, the first image processing model can perform feature fusion of the preceding image mask and the downsampled target image in combination with spatial attention, and predict an image mask corresponding to the portrait area in the downsampled target image from the downsampled target image based on the fusion result; wherein the image mask can be used as the target image mask.

[0166] It should be noted that the image mask corresponding to the portrait area in a downsampled image can usually only be used to extract the image of the portrait area from the downsampled image, but cannot be used to extract the image of the portrait area from the original image that has not been sampled.

[0167] Therefore, after predicting the target image mask from the downsampled target image, the target image mask can be further upsampled to obtain an image mask corresponding to the portrait region in the target image. Specifically, the upsampling process for the target image mask can be used to restore the downsampled target image to the target image. Subsequently, the image mask corresponding to the portrait region in the target image can be used to extract the image of the portrait region from the target image.

[0168] Generally, an image is first downsampled and then upsampled. The restored image obtained may lose some image information compared to the original image. For example, the details of the restored image may be blurrier than the details of the original image. Therefore, when the target image mask is upsampled to obtain an upsampled image mask, the details of the upsampled image mask may be further refined.

[0169] In one embodiment shown, when the target image mask is upsampled to obtain an image mask corresponding to the portrait area in the target image, the target image mask can be first upsampled to obtain an upsampled image mask, and then the upsampled image mask can be divided into a number of image blocks based on a preset block rule, and the confidence corresponding to each image block is calculated. Then, the feature data contained in each image block (which can be called the first type of image block) whose confidence is less than a preset threshold (which can be called the third threshold) can be feature refined; since these feature data can characterize the details of the image mask, the obtained refined feature data can be determined as the refined first type of image block. Finally, based on the refined first type of image block and each image block (which can be called the second type of image block) whose confidence is greater than or equal to the third threshold, an image mask corresponding to the portrait area in the target image can be generated.

[0170] In practical applications, the third threshold may be preset by a technician, and the present disclosure does not limit this.

[0171] Furthermore, in one embodiment shown, a convolution result in matrix form corresponding to the above-mentioned upsampled image mask can be calculated based on a preset convolution layer for dividing the image into blocks (which may be called a block convolution layer), and each convolution value in the convolution result can be determined as the confidence of the image block in the upsampled image mask corresponding to the convolution value.

[0172] In practical applications, the specific values ​​of parameters such as the convolution kernel and convolution step corresponding to the above-mentioned block convolution layer can be pre-set by technical personnel according to actual needs, and the present disclosure does not impose any restrictions on this.

[0173] In another embodiment shown, feature extraction may be performed on the first category image blocks based on several preset convolutional layers (referred to as refined convolutional layers) for refining image features, and the extracted feature data may be determined as the refined first category image blocks.

[0174] In practical applications, the number of the above-mentioned refined convolutional layers, as well as the specific values ​​of parameters such as the convolution kernel and convolution step corresponding to each refined convolutional layer, can be pre-set by technical personnel according to actual needs, and the present disclosure does not impose any restrictions on this.

[0175] For the target image mask predicted from the target image, there may be edge jitter, that is, the edge of the predicted target image mask is not smooth enough. Therefore, in order to further optimize the target image mask, in one embodiment shown, reference is made to Fig.10 , Fig.10 A flowchart of a target image mask optimization method according to an embodiment of the present disclosure is schematically shown.

[0176] The target image mask optimization method may include the following steps:

[0177] Step 1001: extracting an edge image from the target image.

[0178] Step 1002: extracting an optical flow map corresponding to the target image based on a preset optical flow map extraction algorithm.

[0179] Step 1003: extracting an edge optical flow map from the optical flow map based on the edge image.

[0180] Step 1004: Based on preset weights, weighted processing is performed on the pixel value of each pixel point in the target image mask and the pixel value of each pixel point in the edge optical flow map to obtain weighted pixel values, and a weighted target image mask is generated based on the weighted pixel values.

[0181] For the above target image, an edge image can be extracted from the target image.

[0182] Specifically, in one embodiment shown, the target image may be first dilated to obtain a dilated image, and then the target image may be eroded to obtain an eroded image. Subsequently, the pixel value of each pixel in the dilated image may be subtracted from the pixel value of each pixel in the eroded image to obtain a pixel difference corresponding to each pixel, and a difference image may be generated based on the position of each pixel in the dilated image and the eroded image, and the pixel difference corresponding to the pixel, so that the generated difference image may be determined as the edge image.

[0183] In addition, an optical flow map corresponding to the target image may be extracted based on a preset optical flow map extraction algorithm.

[0184] In practical applications, the above optical flow map extraction algorithm may be a general optical flow map extraction algorithm, and the present disclosure does not limit this.

[0185] When the edge image and the optical flow map are obtained, the edge optical flow map can be extracted from the optical flow map based on the edge image. For example, for a pixel point whose pixel value in the edge image is not 0, the pixel value of the pixel point in the optical flow map can be retained; and for a pixel point whose pixel value in the edge image is 0, the pixel value of the pixel point in the optical flow map can be set to 0.

[0186] For the target image mask predicted by the first image processing model from the target image, the pixel values ​​of each pixel in the target image mask and the pixel values ​​of each pixel in the edge optical flow map can be weighted based on preset weights to obtain weighted pixel values ​​corresponding to each pixel, and based on the position of each pixel in the target image mask and the weighted pixel value corresponding to the pixel, a weighted target image mask is generated. Subsequently, the weighted target image mask can be used to extract an image of the portrait area from the target image.

[0187] Exemplary Media

[0188] After introducing the method of the exemplary embodiment of the present disclosure, next, refer to Fig.11 A medium for image processing according to an exemplary embodiment of the present disclosure will be described.

[0189] In this exemplary embodiment, the above method can be implemented by a program product, such as a portable compact disk read-only memory (CD-ROM) and including program code, and can be run on a device, such as a personal computer. However, the program product of the present disclosure is not limited thereto, and in this document, a readable storage medium can be any tangible medium containing or storing a program, which can be used by or in combination with an instruction execution system, an apparatus, or a device.

[0190] The program product may adopt any combination of one or more readable media, which may be readable signal media or readable storage media.

[0191] The readable storage medium may be, for example, but not limited to, an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, device or device, or any combination thereof. More specific examples of readable storage media (a non-exhaustive list) include: an electrical connection with one or more conductors, a portable disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination thereof.

[0192] The readable signal medium may include a data signal propagated in baseband or as part of a carrier wave, wherein the readable program code is carried. Such propagated data signals may take a variety of forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination of the above. The readable signal medium may also be any readable medium other than a readable storage medium, which may send, propagate, or transmit a program for use by or in conjunction with an instruction execution system, apparatus, or device.

[0193] The program code contained on the readable medium may be transmitted using any appropriate medium, including but not limited to wireless, wired, optical cable, RE, etc., or any suitable combination of the foregoing.

[0194] Program code for performing the operations of the present disclosure may be written in any combination of one or more programming languages, including object-oriented programming languages ​​such as Java, C++, etc., and conventional procedural programming languages ​​such as C or similar programming languages. The program code may be executed entirely on the user computing device, partially on the user computing device, partially on a remote computing device, or entirely on a remote computing device or server. In the case of a remote computing device, the remote computing device may be connected to the user computing device through any type of network, including a local area network (LAN) or a wide area network (WAN), or may be connected to an external computing device (e.g., using an Internet service provider to connect through the Internet).

[0195] Exemplary Devices

[0196] After introducing the medium of the exemplary embodiment of the present disclosure, next, reference is made to Fig.12 A device for image processing according to an exemplary embodiment of the present disclosure will be described.

[0197] The implementation process of the functions and effects of each module in the following device is specifically described in the implementation process of the corresponding steps in the above method, which will not be repeated here. For the device embodiment, since it basically corresponds to the method embodiment, the relevant parts can be referred to the partial description of the method embodiment.

[0198] refer to Fig.12 , Fig.12 An image processing device according to an embodiment of the present disclosure is schematically shown.

[0199] The above-mentioned image processing device may include:

[0200] An acquisition module 1201 is used to acquire a target image to be processed;

[0201] A determination module 1202 is used to determine a preceding image corresponding to the target image; wherein the preceding image is an image that is located before the target image in time sequence;

[0202] A first prediction module 1203 is used to predict a previous image mask corresponding to the portrait area in the previous image from the previous image based on the trained first image processing model;

[0203] The second prediction module 1204 is used to input the previous image mask and the target image into the first image processing model, so that the first image processing model performs feature fusion on the previous image mask and the target image in combination with spatial attention, and predicts a target image mask corresponding to the portrait area in the target image from the target image based on the fusion result.

[0204] Optionally, the model parameters of the first image processing model are model parameters migrated from the second image processing model through knowledge distillation; the number of convolutional layers in the first image processing model is less than the number of convolutional layers in the second image processing model.

[0205] Optionally, the target image is any frame image except the first frame in the image sequence; the preceding image is an image in the image sequence that is located before the target image and the number of frames between the preceding image and the target image is a preset first threshold.

[0206] Optionally, the preceding image is an image obtained by performing an affine transformation on the target image based on a preset image transformation relationship.

[0207] Optionally, the device further comprises:

[0208] A binarization module 1205 is used to perform binarization processing on the target image mask to obtain a corresponding binarization mask;

[0209] A first extraction module 1206, configured to extract an image of a portrait region from the target image based on the binary mask;

[0210] The fusion module 1207 is used to fuse the image of the portrait area with the background image to obtain a fused image.

[0211] Optionally, the first image processing model includes at least a first convolutional layer and a second convolutional layer; wherein the first convolutional layer is used to extract first feature data related to image details; and the second convolutional layer is used to extract second feature data related to image semantics;

[0212] The second prediction module 1204 is specifically used for:

[0213] The target image and the previous image mask are input into the first image processing model so that the first image processing model performs the following steps:

[0214] Acquire the first feature data extracted by the first convolution layer for the target image, and acquire the second feature data extracted by the second convolution layer for the target image;

[0215] Performing feature fusion on the first feature data and the second feature data to obtain first fused feature data;

[0216] Based on the spatial attention mechanism and the preceding image mask, calculating a spatial attention weight corresponding to the target image, and multiplying the spatial attention weight by the first fused feature data to obtain weighted first fused feature data;

[0217] The preceding image mask and the weighted first fused feature data are subjected to feature fusion to obtain second fused feature data, and feature extraction is further performed on the second fused feature data to obtain a target image mask corresponding to the portrait area in the target image.

[0218] Optionally, the calculating, based on the spatial attention mechanism and the preceding image mask, a spatial attention weight corresponding to the target image includes:

[0219] The pixel value of each pixel in the previous image mask is divided by 255 to obtain a corresponding normalized matrix, and the normalized matrix is ​​determined as the spatial attention weight corresponding to the target image.

[0220] Optionally, the first characteristic data and the second characteristic data have the same data format;

[0221] The first image processing model further performs the following steps:

[0222] Based on the channel attention mechanism and the first fused feature data, calculating the channel attention weights corresponding to both the first feature data and the second feature data, and multiplying the first feature data and the second feature data by the channel attention respectively to obtain weighted first feature data and weighted second feature data;

[0223] The step of performing feature fusion on the preceding image mask and the weighted first fused feature data to obtain second fused feature data includes:

[0224] The preceding image mask, the weighted first fused feature data, the weighted first feature data and the weighted second feature data are subjected to feature fusion to obtain second fused feature data.

[0225] Optionally, the calculating, based on the channel attention mechanism and the first fused feature data, a channel attention weight corresponding to both the first feature data and the second feature data includes:

[0226] Performing global average pooling on the first fused feature data to obtain pooled data;

[0227] Based on a linear rectification function, mapping the pooled data into first mapping data;

[0228] Based on the Sigmoid function, the first mapping data is further mapped into second mapping data, and the second mapping data is determined as a channel attention weight corresponding to both the first feature data and the second feature data.

[0229] Optionally, performing feature fusion on the feature data includes: calculating the sum of the feature data.

[0230] Optionally, the first image processing model is trained in the following manner:

[0231] Acquire a labeled target image sample; wherein the labeled image sample is labeled with an image mask corresponding to the portrait area therein;

[0232] Determine an image mask corresponding to a portrait area in a preceding image sample corresponding to the labeled target image sample;

[0233] Inputting the image mask in the preceding image sample and the labeled target image sample into the first image processing model, so that the first image processing model performs feature fusion combining spatial attention on the image mask in the preceding image sample and the labeled target image sample, and predicts an image mask corresponding to a portrait region in the labeled target image sample from the labeled target image sample based on the fusion result;

[0234] Through back propagation, based on the image mask predicted by the first image processing model from the labeled image samples and the labeled image mask in the labeled image samples, the model parameters of the first image processing model are adjusted to complete the training of the first image processing model.

[0235] Optionally, the target image sample is any frame image in an image sequence sample; the preceding image sample is an image in the image sequence sample whose frame number interval with the target image sample does not exceed a preset second threshold.

[0236] Optionally, the preceding image is an image obtained by performing an affine transformation on the target image based on a preset image transformation relationship.

[0237] Optionally, the first image processing model is also trained in the following manner:

[0238] According to a preset first ratio, a partial image mask is selected from image masks in preceding image samples corresponding to a number of labeled target image samples, and a pixel value of each pixel point in each image mask in the partial image mask is set to 0.

[0239] Optionally, the first image processing model is also trained in the following manner:

[0240] According to a preset second ratio, some target image samples are selected from a number of marked target image samples, and an image of a background area in each target image sample in the partial target image samples is replaced with any background image in a preset background image set.

[0241] Optionally, the first image processing model is also trained in the following manner:

[0242] Obtain unlabeled target image samples;

[0243] Determine an image mask corresponding to the portrait area in a preceding image sample corresponding to the unlabeled target image sample;

[0244] Inputting the image mask in the preceding image sample and the unlabeled target image sample into the first image processing model, so that the first image processing model performs feature fusion combining spatial attention on the image mask in the preceding image sample and the unlabeled target image sample, and predicts an image mask corresponding to a portrait region in the unlabeled target image sample from the unlabeled target image sample based on the fusion result;

[0245] Based on the image mask predicted by the first image processing model from the unlabeled image samples, the unlabeled target image samples are labeled to obtain labeled target image samples, and the first image processing model is trained based on the labeled target image samples.

[0246] Optionally, the preceding image mask is an image mask corresponding to a portrait area in the preceding image after downsampling, predicted from the preceding image after downsampling based on the first image processing model;

[0247] The second prediction module 1204 is specifically used for:

[0248] Downsampling the target image to obtain a downsampled target image;

[0249] Inputting the preceding image mask and the downsampled target image into the first image processing model, so that the first image processing model performs feature fusion combining spatial attention on the preceding image mask and the downsampled target image, and predicts a target image mask corresponding to a portrait region in the downsampled target image from the downsampled target image based on the fusion result;

[0250] The device also includes:

[0251] The up-sampling module 1208 is used to up-sample the target image mask to obtain an image mask corresponding to the portrait area in the target image.

[0252] Optionally, the up-sampling module 1208 is specifically configured to:

[0253] Upsampling the target image mask to obtain an upsampled image mask;

[0254] Based on a preset block rule, the upsampled image mask is divided into a plurality of image blocks, and a confidence level corresponding to each image block is calculated;

[0255] Performing feature refinement on the feature data contained in the first-category image blocks whose confidence is less than a preset third threshold value to obtain refined first-category image blocks;

[0256] Based on the refined first-category image blocks and the second-category image blocks whose confidence is greater than or equal to the third threshold, an image mask corresponding to the portrait area in the target image is generated.

[0257] Optionally, the up-sampling module 1208 is specifically configured to:

[0258] Based on a preset block convolution layer, a convolution result in matrix form corresponding to the upsampled image mask is calculated, and each convolution value in the convolution result is determined as a confidence of an image block in the upsampled image mask corresponding to the convolution value.

[0259] Optionally, the up-sampling module 1208 is specifically configured to:

[0260] Based on a plurality of preset refined convolutional layers, feature extraction is performed on the first-category image blocks whose confidence is less than a preset third threshold, and the extracted feature data is determined as the refined first-category image blocks.

[0261] Optionally, the device further comprises:

[0262] A second extraction module 1209, configured to extract an edge image from the target image;

[0263] A third extraction module 1210 is used to extract an optical flow map corresponding to the target image based on a preset optical flow map extraction algorithm;

[0264] A fourth extraction module 1211, configured to extract an edge optical flow map from the optical flow map based on the edge image;

[0265] The weighting module 1212 is used to perform weighted processing on the pixel value of each pixel point in the target image mask and the pixel value of each pixel point in the edge optical flow map based on preset weights to obtain weighted pixel values, and generate a weighted target image mask based on the weighted pixel values.

[0266] Optionally, the second extraction module 1209 is specifically used to:

[0267] Performing dilation processing on the target image to obtain a dilated image, and performing erosion processing on the target image to obtain an eroded image;

[0268] Calculate a pixel difference obtained by subtracting a pixel value of each pixel point in the eroded image from a pixel value of each pixel point in the expanded image, and generate a difference image based on the pixel difference;

[0269] The difference image is determined as an edge image.

[0270] Exemplary Computing Devices

[0271] After introducing the method, medium and apparatus of the exemplary embodiments of the present disclosure, next, reference is made to Fig.13 A computing device for image processing according to an exemplary embodiment of the present disclosure is described.

[0272] Fig.13The computing device 1300 shown is merely an example and should not bring any limitation to the functionality and scope of use of the embodiments of the present disclosure.

[0273] like Fig.11 As shown, the computing device 1300 is in the form of a general computing device. The components of the computing device 1300 may include but are not limited to: at least one processing unit 1301, at least one storage unit 1302, and a bus 1303 connecting different system components (including the processing unit 1301 and the storage unit 1302).

[0274] The bus 1303 includes a data bus, a control bus, and an address bus.

[0275] The storage unit 1302 may include a readable medium in the form of a volatile memory, such as a random access memory (RAM) 13021 and / or a cache memory 13022 , and may further include a readable medium in the form of a non-volatile memory, such as a read-only memory (ROM) 13023 .

[0276] The storage unit 1302 may also include a program / utility 13025 having a set (at least one) of program modules 13024, such program modules 13024 including but not limited to: an operating system, one or more application programs, other program modules, and program data, each of which or some combination may include an implementation of a network environment.

[0277] The computing device 1300 may also communicate with one or more external devices 1304 (eg, a keyboard, pointing device, etc.).

[0278] Such communication may be performed via input / output (I / O) interface 1305. In addition, computing device 1300 may also communicate with one or more networks (e.g., a local area network (LAN), a wide area network (WAN), and / or a public network, such as the Internet) via network adapter 1306. Figure 8 As shown, the network adapter 1306 communicates with other modules of the computing device 1300 via the bus 1303. It should be understood that, although not shown in the figure, other hardware and / or software modules can be used in conjunction with the computing device 1300, including but not limited to: microcode, device drivers, redundant processing units, external disk drive arrays, RAID systems, tape drives, and data backup storage systems.

[0279] It should be noted that, although several units / modules or sub-units / modules of the image processing device are mentioned in the above detailed description, such division is merely exemplary and not mandatory. In fact, according to an embodiment of the present disclosure, the features and functions of two or more units / modules described above may be embodied in one unit / module. Conversely, the features and functions of one unit / module described above may be further divided to be embodied by multiple units / modules.

[0280] In addition, although the operations of the disclosed method are described in a specific order in the drawings, this does not require or imply that the operations must be performed in this specific order, or that all the operations shown must be performed to achieve the desired results. Additionally or alternatively, some steps may be omitted, multiple steps may be combined into one step, and / or one step may be decomposed into multiple steps.

[0281] Although the spirit and principle of the present disclosure have been described with reference to several specific embodiments, it should be understood that the present disclosure is not limited to the disclosed specific embodiments, and the division of various aspects does not mean that the features in these aspects cannot be combined to benefit, and such division is only for the convenience of expression. The present disclosure is intended to cover various modifications and equivalent arrangements included in the spirit and scope of the attached claims.

Claims

1. An image processing method, the method comprising: Obtaining a target image to be processed; Determine a preceding image corresponding to the target image; wherein the preceding image is an image that precedes the target image in time sequence; Based on the trained first image processing model, predicting a preceding image mask corresponding to the portrait area in the preceding image from the preceding image; The target image and the previous image mask are input into the first image processing model so that the first image processing model performs the following steps: Extracting feature data from the target image; Based on the spatial attention mechanism and the preceding image mask, calculating a spatial attention weight corresponding to the target image, and multiplying the spatial attention weight by the feature data to obtain weighted feature data; The preceding image mask and the weighted feature data are subjected to feature fusion to obtain fused feature data, and feature extraction is further performed on the fused feature data to obtain a target image mask corresponding to the portrait region in the target image.

2. According to the method of claim 1, the model parameters of the first image processing model are model parameters migrated from the second image processing model through knowledge distillation; the number of convolutional layers in the first image processing model is less than the number of convolutional layers in the second image processing model.

3. According to the method of claim 1, the target image is any frame image in the image sequence except the first frame; the preceding image is an image in the image sequence that is located before the target image and the number of frames between the preceding image and the target image is a preset first threshold.

4. According to the method of claim 1, the preceding image is an image obtained by performing an affine transformation on the target image based on a preset image transformation relationship.

5. The method according to claim 1, further comprising: Binarizing the target image mask to obtain a corresponding binary mask; Based on the binary mask, extracting an image of a portrait area from the target image; The image of the portrait area is fused with the background image to obtain a fused image.

6. According to the method of claim 1, the first image processing model comprises at least a first convolutional layer and a second convolutional layer; wherein, The first convolutional layer is used to extract first feature data related to image details; The second convolutional layer is used to extract second feature data related to image semantics; The step of extracting feature data from the target image includes: Acquire the first feature data extracted by the first convolution layer for the target image, and acquire the second feature data extracted by the second convolution layer for the target image; The first feature data and the second feature data are subjected to feature fusion to obtain first fused feature data as feature data extracted from the target image.

7. The method according to claim 6, wherein the calculating the spatial attention weight corresponding to the target image based on the spatial attention mechanism and the previous image mask comprises: The pixel value of each pixel in the previous image mask is divided by 255 to obtain a corresponding normalized matrix, and the normalized matrix is ​​determined as the spatial attention weight corresponding to the target image.

8. The method according to claim 6, wherein the first characteristic data and the second characteristic data have the same data format; The first image processing model further performs the following steps: Based on the channel attention mechanism and the first fused feature data, calculating the channel attention weights corresponding to both the first feature data and the second feature data, and multiplying the first feature data and the second feature data by the channel attention respectively to obtain weighted first feature data and weighted second feature data; The step of performing feature fusion on the preceding image mask and the weighted first fused feature data to obtain second fused feature data includes: The preceding image mask, the weighted first fused feature data, the weighted first feature data and the weighted second feature data are subjected to feature fusion to obtain second fused feature data.

9. The method according to claim 8, wherein the step of calculating the channel attention weights corresponding to both the first feature data and the second feature data based on the channel attention mechanism and the first fused feature data comprises: Performing global average pooling on the first fused feature data to obtain pooled data; Based on a linear rectification function, mapping the pooled data into first mapping data; Based on the Sigmoid function, the first mapping data is further mapped into second mapping data, and the second mapping data is determined as a channel attention weight corresponding to both the first feature data and the second feature data.

10. According to any one of the methods of claims 6 to 9, feature fusion is performed on the feature data, comprising: Calculates the sum of feature data.

11. The method according to claim 1, wherein the first image processing model is trained in the following manner: Get the labeled target image sample; where, The labeled image sample is labeled with an image mask corresponding to the portrait area therein; Determine an image mask corresponding to a portrait area in a preceding image sample corresponding to the labeled target image sample; Inputting the image mask in the preceding image sample and the labeled target image sample into the first image processing model, so that the first image processing model performs feature fusion combining spatial attention on the image mask in the preceding image sample and the labeled target image sample, and predicts an image mask corresponding to a portrait region in the labeled target image sample from the labeled target image sample based on the fusion result; Through back propagation, based on the image mask predicted by the first image processing model from the labeled image samples and the labeled image mask in the labeled image samples, the model parameters of the first image processing model are adjusted to complete the training of the first image processing model.

12. According to the method of claim 11, the target image sample is any frame image in the image sequence samples; the preceding image sample is an image in the image sequence samples whose frame number interval with the target image sample does not exceed a preset second threshold.

13. According to the method of claim 11, the preceding image is an image obtained by performing an affine transformation on the target image based on a preset image transformation relationship.

14. The method according to claim 11, further training the first image processing model by: According to a preset first ratio, a partial image mask is selected from image masks in preceding image samples corresponding to a number of labeled target image samples, and a pixel value of each pixel point in each image mask in the partial image mask is set to 0.

15. The method according to claim 11, further training the first image processing model by: According to a preset second ratio, some target image samples are selected from a number of marked target image samples, and an image of a background area in each target image sample in the partial target image samples is replaced with any background image in a preset background image set.

16. The method according to claim 11, further training the first image processing model by: Obtain unlabeled target image samples; Determine an image mask corresponding to the portrait area in a preceding image sample corresponding to the unlabeled target image sample; Inputting the image mask in the preceding image sample and the unlabeled target image sample into the first image processing model, so that the first image processing model performs feature fusion combining spatial attention on the image mask in the preceding image sample and the unlabeled target image sample, and predicts an image mask corresponding to a portrait region in the unlabeled target image sample from the unlabeled target image sample based on the fusion result; Based on the image mask predicted by the first image processing model from the unlabeled image samples, the unlabeled target image samples are labeled to obtain labeled target image samples, and the first image processing model is trained based on the labeled target image samples.

17. The method according to claim 1, wherein the preceding image mask is an image mask corresponding to a portrait area in the preceding image after downsampling, predicted from the preceding image after downsampling based on the first image processing model; The step of inputting the preceding image mask and the target image into the first image processing model so that the first image processing model performs feature fusion combining spatial attention on the preceding image mask and the target image, and predicting a target image mask corresponding to a portrait area in the target image from the target image based on the fusion result, comprises: Downsampling the target image to obtain a downsampled target image; Inputting the preceding image mask and the downsampled target image into the first image processing model, so that the first image processing model performs feature fusion combining spatial attention on the preceding image mask and the downsampled target image, and predicts a target image mask corresponding to a portrait region in the downsampled target image from the downsampled target image based on the fusion result; The method further comprises: The target image mask is upsampled to obtain an image mask corresponding to the portrait area in the target image.

18. The method according to claim 17, wherein upsampling the target image mask to obtain an image mask corresponding to the portrait area in the target image comprises: Upsampling the target image mask to obtain an upsampled image mask; Based on a preset block rule, the upsampled image mask is divided into a plurality of image blocks, and a confidence level corresponding to each image block is calculated; Performing feature refinement on the feature data contained in the first-category image blocks whose confidence is less than a preset third threshold value to obtain refined first-category image blocks; Based on the refined first-category image blocks and the second-category image blocks whose confidence is greater than or equal to the third threshold, an image mask corresponding to the portrait area in the target image is generated.

19. The method according to claim 18, wherein the upsampled image mask is divided into a plurality of image blocks based on a preset block rule, and the confidence corresponding to each image block is calculated, comprising: Based on a preset block convolution layer, a convolution result in matrix form corresponding to the upsampled image mask is calculated, and each convolution value in the convolution result is determined as a confidence of an image block in the upsampled image mask corresponding to the convolution value.

20. The method according to claim 18, wherein the step of refining the feature data contained in the first-category image blocks whose confidence is less than a preset third threshold to obtain the refined first-category image blocks comprises: Based on a plurality of preset refined convolutional layers, feature extraction is performed on the first-category image blocks whose confidence is less than a preset third threshold, and the extracted feature data is determined as the refined first-category image blocks.

21. The method according to claim 1, further comprising: Extracting an edge image from the target image; Based on a preset optical flow map extraction algorithm, extracting an optical flow map corresponding to the target image; Extracting an edge optical flow map from the optical flow map based on the edge image; Based on preset weights, pixel values ​​of each pixel point in the target image mask and pixel values ​​of each pixel point in the edge optical flow map are weighted to obtain weighted pixel values, and a weighted target image mask is generated based on the weighted pixel values.

22. The method according to claim 21, wherein extracting an edge image from the target image comprises: Performing dilation processing on the target image to obtain a dilated image, and performing erosion processing on the target image to obtain an eroded image; Calculate a pixel difference obtained by subtracting a pixel value of each pixel point in the eroded image from a pixel value of each pixel point in the expanded image, and generate a difference image based on the pixel difference; The difference image is determined as an edge image.

23. An image processing device, the device comprising: An acquisition module, used for acquiring a target image to be processed; A determination module, used to determine a preceding image corresponding to the target image; wherein the preceding image is an image that precedes the target image in time sequence; A first prediction module, configured to predict, from the preceding image, a preceding image mask corresponding to a portrait region in the preceding image based on a trained first image processing model; The second prediction module is used to input the target image and the previous image mask into the first image processing model so that the first image processing model performs the following steps: Extracting feature data from the target image; Based on the spatial attention mechanism and the preceding image mask, calculating a spatial attention weight corresponding to the target image, and multiplying the spatial attention weight by the feature data to obtain weighted feature data; The preceding image mask and the weighted feature data are subjected to feature fusion to obtain fused feature data, and feature extraction is further performed on the fused feature data to obtain a target image mask corresponding to the portrait region in the target image.

24. According to the device according to claim 23, the model parameters of the first image processing model are model parameters migrated from the second image processing model through knowledge distillation; the number of convolutional layers in the first image processing model is less than the number of convolutional layers in the second image processing model.

25. According to the device of claim 23, the target image is any frame image in the image sequence except the first frame; the preceding image is an image in the image sequence that is located before the target image and the number of frames between the image and the target image is a preset first threshold.

26. According to the device of claim 23, the preceding image is an image obtained by performing an affine transformation on the target image based on a preset image transformation relationship.

27. The apparatus according to claim 23, further comprising: A binarization module, used for performing binarization processing on the target image mask to obtain a corresponding binarization mask; A first extraction module, used for extracting an image of a portrait area from the target image based on the binary mask; The fusion module is used to fuse the image of the portrait area with the background image to obtain a fused image.

28. The apparatus according to claim 23, wherein the first image processing model comprises at least a first convolutional layer and a second convolutional layer; wherein The first convolution layer is used to extract first feature data related to image details; the second convolution layer is used to extract second feature data related to image semantics; The second prediction module is specifically used for: The target image and the previous image mask are input into the first image processing model so that the first image processing model performs the following steps: Acquire the first feature data extracted by the first convolution layer for the target image, and acquire the second feature data extracted by the second convolution layer for the target image; The first feature data and the second feature data are subjected to feature fusion to obtain first fused feature data as feature data extracted from the target image.

29. The apparatus according to claim 28, wherein the step of calculating the spatial attention weight corresponding to the target image based on the spatial attention mechanism and the preceding image mask comprises: The pixel value of each pixel in the previous image mask is divided by 255 to obtain a corresponding normalized matrix, and the normalized matrix is ​​determined as the spatial attention weight corresponding to the target image.

30. The device according to claim 28, wherein the first characteristic data and the second characteristic data have the same data format; The first image processing model further performs the following steps: Based on the channel attention mechanism and the first fused feature data, calculating the channel attention weights corresponding to both the first feature data and the second feature data, and multiplying the first feature data and the second feature data by the channel attention respectively to obtain weighted first feature data and weighted second feature data; The step of performing feature fusion on the preceding image mask and the weighted first fused feature data to obtain second fused feature data includes: The preceding image mask, the weighted first fused feature data, the weighted first feature data and the weighted second feature data are subjected to feature fusion to obtain second fused feature data.

31. The apparatus according to claim 30, wherein the calculating, based on the channel attention mechanism and the first fused feature data, the channel attention weights corresponding to both the first feature data and the second feature data comprises: Performing global average pooling on the first fused feature data to obtain pooled data; Based on a linear rectification function, mapping the pooled data into first mapping data; Based on the Sigmoid function, the first mapping data is further mapped into second mapping data, and the second mapping data is determined as a channel attention weight corresponding to both the first feature data and the second feature data.

32. According to any one of claims 28 to 31, the device performs feature fusion on the feature data, comprising: Calculates the sum of feature data.

33. The apparatus according to claim 23, wherein the first image processing model is trained by: Get the labeled target image sample; where, The labeled image sample is labeled with an image mask corresponding to the portrait area therein; Determine an image mask corresponding to a portrait area in a preceding image sample corresponding to the labeled target image sample; Inputting the image mask in the preceding image sample and the labeled target image sample into the first image processing model, so that the first image processing model performs feature fusion combining spatial attention on the image mask in the preceding image sample and the labeled target image sample, and predicts an image mask corresponding to a portrait region in the labeled target image sample from the labeled target image sample based on the fusion result; Through back propagation, based on the image mask predicted by the first image processing model from the labeled image samples and the labeled image mask in the labeled image samples, the model parameters of the first image processing model are adjusted to complete the training of the first image processing model.

34. According to the device of claim 33, the target image sample is any frame image in the image sequence samples; the preceding image sample is an image in the image sequence samples whose frame number interval with the target image sample does not exceed a preset second threshold.

35. According to the device of claim 33, the preceding image is an image obtained by performing an affine transformation on the target image based on a preset image transformation relationship.

36. The apparatus according to claim 33, further training the first image processing model by: According to a preset first ratio, a partial image mask is selected from image masks in preceding image samples corresponding to a number of labeled target image samples, and a pixel value of each pixel point in each image mask in the partial image mask is set to 0.

37. The apparatus according to claim 33, further training the first image processing model by: According to a preset second ratio, some target image samples are selected from a number of marked target image samples, and an image of a background area in each target image sample in the partial target image samples is replaced with any background image in a preset background image set.

38. The apparatus according to claim 33, further training the first image processing model by: Obtain unlabeled target image samples; Determine an image mask corresponding to the portrait area in a preceding image sample corresponding to the unlabeled target image sample; Inputting the image mask in the preceding image sample and the unlabeled target image sample into the first image processing model, so that the first image processing model performs feature fusion combining spatial attention on the image mask in the preceding image sample and the unlabeled target image sample, and predicts an image mask corresponding to a portrait region in the unlabeled target image sample from the unlabeled target image sample based on the fusion result; Based on the image mask predicted by the first image processing model from the unlabeled image samples, the unlabeled target image samples are labeled to obtain labeled target image samples, and the first image processing model is trained based on the labeled target image samples.

39. The device according to claim 23, wherein the preceding image mask is an image mask corresponding to a portrait area in the preceding image after downsampling, predicted from the preceding image after downsampling based on the first image processing model; The second prediction module is specifically used for: Downsampling the target image to obtain a downsampled target image; Inputting the preceding image mask and the downsampled target image into the first image processing model, so that the first image processing model performs feature fusion combining spatial attention on the preceding image mask and the downsampled target image, and predicts a target image mask corresponding to a portrait region in the downsampled target image from the downsampled target image based on the fusion result; The device also includes: The up-sampling module is used to up-sample the target image mask to obtain an image mask corresponding to the portrait area in the target image.

40. The apparatus according to claim 39, wherein the up-sampling module is specifically configured to: Upsampling the target image mask to obtain an upsampled image mask; Based on a preset block rule, the upsampled image mask is divided into a plurality of image blocks, and a confidence level corresponding to each image block is calculated; Performing feature refinement on the feature data contained in the first-category image blocks whose confidence is less than a preset third threshold value to obtain refined first-category image blocks; Based on the refined first-category image blocks and the second-category image blocks whose confidence is greater than or equal to the third threshold, an image mask corresponding to the portrait area in the target image is generated.

41. The apparatus according to claim 40, wherein the up-sampling module is specifically configured to: Based on a preset block convolution layer, a convolution result in matrix form corresponding to the upsampled image mask is calculated, and each convolution value in the convolution result is determined as a confidence of an image block in the upsampled image mask corresponding to the convolution value.

42. The apparatus according to claim 40, wherein the up-sampling module is specifically configured to: Based on a plurality of preset refined convolutional layers, feature extraction is performed on the first-category image blocks whose confidence is less than a preset third threshold, and the extracted feature data is determined as the refined first-category image blocks.

43. The apparatus of claim 23, further comprising: A second extraction module, used to extract an edge image from the target image; A third extraction module, configured to extract an optical flow map corresponding to the target image based on a preset optical flow map extraction algorithm; a fourth extraction module, configured to extract an edge optical flow map from the optical flow map based on the edge image; The weighting module is used to perform weighted processing on the pixel value of each pixel point in the target image mask and the pixel value of each pixel point in the edge optical flow map based on a preset weight to obtain a weighted pixel value, and generate a weighted target image mask based on the weighted pixel value.

44. The apparatus according to claim 43, wherein the second extraction module is specifically configured to: Performing dilation processing on the target image to obtain a dilated image, and performing erosion processing on the target image to obtain an eroded image; Calculate a pixel difference obtained by subtracting a pixel value of each pixel point in the eroded image from a pixel value of each pixel point in the expanded image, and generate a difference image based on the pixel difference; The difference image is determined as an edge image.

45. A medium having a computer program stored thereon, wherein when the computer program is executed by a processor, the method according to any one of claims 1 to 22 is implemented.

46. ​​A computing device comprising: processor; a memory for storing a processor executable program; The processor implements the method according to any one of claims 1 to 22 by running the executable program.

Citation Information

Patent Citations

  • Human body posture real-time detection method and device, computer equipment and storage medium

    CN112560796A

  • Hyperspectral image super-resolution reconstruction method and device and electronic equipment

    CN113139902A