Video processing method, device, electronic device and storage medium

By performing image processing only on the part of the unobstructed area of ​​the target object in the video image in the video processing technology, the problem of occlusion affecting the processing effect is solved, and a better video image display effect is achieved.

CN116132732BActive Publication Date: 2025-06-06BEIJING DAJIA INTERNET INFORMATION TECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202310093799.2
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2023-01-30
Publication Date
2025-06-06
Estimated Expiration
2043-01-30

AI Technical Summary

Technical Problem

When the existing video processing technology processes the target object, when there is an occlusion, the occlusion will be processed together with the target object, resulting in poor display of the processed video image.

Method used

By acquiring the target object and the occlusion area in the video image, only the part of the target object is not occluded is image-processed to obtain the processed video image, thereby avoiding affecting the occlusion area.

Benefits of technology

The image processing of the target object is achieved without affecting the occlusion area, avoiding abnormal display of non-target objects, and improving the display effect of the processed video image.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN116132732B_ABST
    Figure CN116132732B_ABST
Patent Text Reader

Abstract

The present disclosure relates to a video processing method, device, electronic device and storage medium, wherein the video processing method comprises: obtaining a video image, wherein the video image comprises a target object to be processed and an occlusion area that occludes the target object; performing image processing on a portion of the target object that is not occluded by the occlusion area to obtain a processed video image. The video processing method, device, electronic device and storage medium disclosed in the present disclosure can solve the problem of poor display effect of processed video images caused by the presence of occlusions in front of the target object, and can realize image processing of the target object without affecting the occlusion area, thereby avoiding abnormal display of non-target objects in the occlusion area and improving the display effect of the processed video image.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present disclosure relates to the field of video processing technology, and in particular to a video processing method, device, electronic device and storage medium. Background Art

[0002] With the continuous development of video technology, it is possible to process target objects in video images such as live video and recorded video, for example, by adjusting the color and brightness of local objects, applying deformation, etc. Here, when processing the target object, there may be occluders that block the target object. For example, in a live video scene, when applying deformation to a person's face, there may be occluders such as headphones and human hands that block part of the face.

[0003] However, in related video processing technologies, when processing a video scene in which there is an occluder in front of the target object, the occluder will be processed together with the target object, which may cause deformation and distortion of non-target objects, display abnormalities, etc., resulting in poor display effect of the processed video image. Summary of the invention

[0004] The present disclosure provides a video processing method, device, electronic device and storage medium to at least solve the problem in the related art that there are occluders in front of the target object, resulting in poor display effect of the processed video image. The technical solution of the present disclosure is as follows:

[0005] According to a first aspect of an embodiment of the present disclosure, a video processing method is provided, comprising: acquiring a video image, wherein the video image includes a target object to be processed and an occlusion area that occludes the target object; performing image processing on a portion of the target object that is not occluded by the occlusion area to obtain a processed video image.

[0006] Optionally, performing image processing on the portion of the target object not obscured by the occlusion area to obtain a processed video image includes: determining a first target area in the video image, wherein the first target area is a region containing the target object; determining the occlusion area in the first target area; performing image processing on the portion of the target object in the first target area that is not obscured by the occlusion area to obtain a second target area; and obtaining a processed video image based on the second target area.

[0007] Optionally, performing image processing on a portion of the target object in the first target area that is not obscured by the occlusion area to obtain a second target area includes: performing image processing on a portion of the target object in the first target area that is not obscured by the occlusion area and the occlusion area to obtain a candidate second target area; and superimposing the unprocessed occlusion area on the processed occlusion area in the candidate second target area to obtain the second target area.

[0008] Optionally, before performing image processing on the portion of the target object in the first target area that is not obscured by the occlusion area, the video processing method further includes: removing the occlusion area from the first target area according to the position of the occlusion area in the first target area to obtain the first target area after the occlusion is removed.

[0009] Optionally, removing the occluded area from the first target area according to the position of the occluded area in the first target area to obtain the first target area after the occlusion is removed includes:

[0010] According to the position of the occluded area in the first target area, all pixel values ​​of the occluded area in the first target area are set to preset pixel values ​​to obtain the first target area after the occlusion is removed.

[0011] Optionally, performing image processing on the portion of the target object in the first target area that is not obscured by the occlusion area to obtain a second target area includes: in response to the image processing not satisfying a preset completion condition, performing image processing on the portion of the target object in the first target area that is not obscured by the occlusion area to obtain the second target area; in response to the image processing satisfying the preset completion condition, performing pixel completion on the portion of the target object in the first target area that is obscured by the occlusion area to obtain a completed target object, wherein the completion condition includes: the image processing belongs to a preset image processing type and / or the image processing includes removing the occlusion area from the video image; performing the image processing on the completed target object to obtain the second target area.

[0012] Optionally, obtaining a processed video image based on the second target area includes: superimposing the occluded area on the second target area according to the position of the occluded area in the video image to obtain the processed video image; or, displaying the entire second target area in the processed video image, and not displaying the occluded area in the processed video image.

[0013] Optionally, determining the first target area in the video image includes: determining a visible outline of the target object in the video image; patching the visible outline based on preset outline characteristics of the target object to obtain a patched outline of the target object; and determining an area enclosed by the patched outline as the first target area.

[0014] Optionally, the image processing includes performing at least one of the following operations globally or locally: deformation, pixel filtering, color adjustment and image superposition, and the video image is a video image collected in real time.

[0015] According to a second aspect of an embodiment of the present disclosure, a video processing device is provided, comprising: an acquisition unit, configured to acquire a video image, wherein the video image includes a target object to be processed and an occlusion area that occludes the target object; and a processing unit, configured to perform image processing on a portion of the target object that is not occluded by the occlusion area to obtain a processed video image.

[0016] Optionally, the processing unit includes: a first target area determination unit, configured to determine a first target area in the video image, wherein the first target area is an area containing the target object; an occlusion area determination unit, configured to determine the occlusion area in the first target area; a second target area determination unit, configured to perform the image processing on a portion of the target object in the first target area that is not occluded by the occlusion area to obtain a second target area; and a video image determination unit, configured to obtain a processed video image based on the second target area.

[0017] Optionally, the second target area determination unit is further configured to: perform image processing on the portion of the target object in the first target area that is not obscured by the occlusion area and the occlusion area to obtain a candidate second target area; and superimpose the unprocessed occlusion area on the processed occlusion area in the candidate second target area to obtain a second target area.

[0018] Optionally, the video processing device also includes a removal unit, which is configured to remove the occlusion area from the first target area according to the position of the occlusion area in the first target area before performing image processing on the part of the target object in the first target area that is not occluded by the occlusion area, so as to obtain the first target area after the occlusion is removed.

[0019] Optionally, the removal unit is further configured to: according to the position of the occluded area in the first target area, set all pixel values ​​of the occluded area in the first target area to preset pixel values ​​to obtain the first target area after occlusion is removed.

[0020] Optionally, the second target area determination unit is further configured to: in response to image processing not satisfying a preset completion condition, perform image processing on the portion of the target object in the first target area that is not obscured by the occlusion area to obtain a second target area; in response to the image processing satisfying a preset completion condition, perform pixel completion on the portion of the target object in the first target area that is obscured by the occlusion area to obtain a completed target object, wherein the completion condition includes: the image processing belongs to a preset image processing type and / or the image processing includes removing the occlusion area from the video image; and perform the image processing on the completed target object to obtain a second target area.

[0021] Optionally, the video image determination unit is further configured to: superimpose the occluded area onto the second target area according to the position of the occluded area in the video image to obtain the processed video image; or, display the entire second target area in the processed video image, and not display the occluded area in the processed video image.

[0022] Optionally, the first target area determination unit is further configured to: determine a visible outline of the target object in the video image; based on preset outline characteristics of the target object, patch the visible outline to obtain a patched outline of the target object; and determine the area enclosed by the patched outline as the first target area.

[0023] Optionally, the image processing includes performing at least one of the following operations globally or locally: deformation, pixel filtering, color adjustment and image superposition, and the video image is a video image collected in real time.

[0024] According to a third aspect of an embodiment of the present disclosure, an electronic device is provided, comprising: a processor; and a memory for storing instructions executable by the processor, wherein the processor executable instructions, when executed by the processor, cause the processor to execute the video processing method according to the exemplary embodiment of the present disclosure.

[0025] According to a fourth aspect of an embodiment of the present disclosure, a computer-readable storage medium is provided. When instructions in the computer-readable storage medium are executed by a processor of an electronic device, the electronic device is enabled to perform the video processing method according to the exemplary embodiment of the present disclosure.

[0026] According to a fifth aspect of an embodiment of the present disclosure, a computer program product is provided, wherein the computer program product comprises computer instructions, and when the computer instructions are executed by a processor, the video processing method according to the exemplary embodiment of the present disclosure is implemented.

[0027] The technical solution provided by the embodiments of the present disclosure brings at least the following beneficial effects:

[0028] A video image including a target object to be processed and an occluded area can be obtained, and a processed video image can be obtained by performing image processing on the portion of the target object that is not occluded by the occluded area. In this way, image processing of the target object can be achieved without affecting the occluded area, thereby avoiding abnormal display of non-target objects in the occluded area and improving the display effect of the processed video image.

[0029] It is to be understood that the foregoing general description and the following detailed description are exemplary and explanatory only and are not restrictive of the present disclosure. BRIEF DESCRIPTION OF THE DRAWINGS

[0030] The drawings herein are incorporated into and constitute a part of the specification, illustrate embodiments consistent with the present disclosure, and together with the description are used to explain the principles of the present disclosure, and do not constitute improper limitations on the present disclosure.

[0031] Figure 1 The diagram is a schematic diagram of an example implementation scenario of a video processing method according to an exemplary embodiment.

[0032] Figure 2 The figure is a flowchart of a video processing method according to an exemplary embodiment.

[0033] Figure 3 The present invention is a flowchart of processing a video image in a video processing method according to an exemplary embodiment.

[0034] Figure 4 The present invention is a flowchart showing a step of determining a first target area in a video processing method according to an exemplary embodiment.

[0035] Figure 5 The diagram is a schematic diagram showing a method for determining a first target area in a video processing method according to an exemplary embodiment.

[0036] Figure 6 The present invention is a flowchart of a step of obtaining a second target area in a video processing method according to an exemplary embodiment.

[0037] Figure 7 The flowchart of an example of a video processing method is shown according to an exemplary embodiment.

[0038] Fig. 8A is a schematic diagram showing a video image processed according to an existing video processing method.

[0039] Figure 8Bis a schematic diagram showing a video image processed by a video processing method according to an exemplary embodiment.

[0040] Fig. 9 The invention is a block diagram of a video processing device according to an exemplary embodiment.

[0041] Fig.10 It is a block diagram of an electronic device according to an exemplary embodiment. DETAILED DESCRIPTION

[0042] In order to enable ordinary persons in the art to better understand the technical solutions of the present disclosure, the technical solutions in the embodiments of the present disclosure will be clearly and completely described below in conjunction with the accompanying drawings.

[0043] It should be noted that the terms "first", "second", etc. in the specification and claims of the present disclosure and the above-mentioned drawings are used to distinguish similar objects, and are not necessarily used to describe a specific order or sequence. It should be understood that the data used in this way can be interchanged where appropriate, so that the embodiments of the present disclosure described herein can be implemented in an order other than those illustrated or described herein. The embodiments described in the following exemplary embodiments do not represent all embodiments consistent with the present disclosure. Instead, they are merely examples of devices and methods consistent with some aspects of the present disclosure as detailed in the appended claims.

[0044] It should be noted that the phrase "at least one of the items" in the present disclosure includes three types of parallel situations: "any one of the items", "a combination of any number of the items", and "all of the items". For example, "including at least one of A and B" includes the following three parallel situations: (1) including A; (2) including B; (3) including A and B. Another example is "executing at least one of step 1 and step 2" which means the following three parallel situations: (1) executing step 1; (2) executing step 2; (3) executing step 1 and step 2.

[0045] As mentioned above, in relevant video processing technologies, when there is an occlusion in front of the target object to be processed, image processing will be applied to both the occlusion and the target object, so that the occlusion as a non-target object will also be processed, resulting in poor display effect of the processed video image, affecting the viewing experience.

[0046] Taking real-time video broadcasting as an example, with the rapid development of the short video industry, the frequency of use of facial effects in videos, pictures, and live broadcasting scenarios is increasing. For example, in order to improve the viewing experience of the video or highlight the theme that the video wants to express, you can apply facial deformation, add attachments, and other video effects to the characters in the video, such as face slimming, makeup, etc.

[0047] However, in real-time live broadcast and online scenarios, some obstructions that appear in front of the person's face will cause distortion of the video effects. For example, in some food recommendation video live broadcast scenarios, after the user turns on the face deformation effect, in order to highlight the theme of the video, intrusive objects such as food and chopsticks may appear in front of the user's face, causing the deformation of intrusive objects such as food and chopsticks; for another example, in some video live broadcast scenarios that add color effects, when intrusive objects such as headphones and hands suddenly cover the user's face, the color effects originally applied to the face will become invalid.

[0048] For the above-mentioned scenes, in the relevant video processing technologies, when processing scenes with deformable target occluders or deformable area intruders, the interfering objects will not be detected and processed, so deformation distortion, abnormal display of non-target objects, etc. will occur. In particular, in real-time video scenes, when the user sets the intensity of the deformation effect to a large value, this abnormal deformation phenomenon is more obvious, thereby affecting the display effect of the video.

[0049] In view of the above problems, the following will provide a video processing method, a video processing device, an electronic device, a computer-readable storage medium and a computer program product according to an exemplary embodiment of the present disclosure with reference to the accompanying drawings, so as to realize a video processing and display method, which can provide a more robust and robust video special effects presentation and processing solution, and solve at least one of the problems such as abnormal deformation of intrusive objects and loss of special effects application of image processing.

[0050] Here, it should be noted that, although the context is described using a live video scene as an example, it should be understood that the video processing method according to the exemplary embodiment of the present disclosure is not limited to this scene, and it can be applied to any scene that requires video processing, for example, it can also be applied to post-production of shot or recorded videos, video frame special effects processing, and other scenes.

[0051] In addition, although the context describes the application of special effects to a person's face as an example, it should be understood that the role of the video processing method according to the exemplary embodiment of the present disclosure is not limited to this. For example, it can also be used to improve video quality, improve video display effects, and improve local video clarity during post-video processing. For example, improving the color or brightness of people or objects in video images, adding animation special effects to people or objects, repairing light and shadow defects in video shooting, etc., as well as any video processing scenarios that require distinguishing objects for processing.

[0052] According to an exemplary embodiment of the present disclosure, a video processing method is provided. The video processing method can be applied to any video processing scenario. Figure 1 An example implementation scenario of a video processing method according to an exemplary embodiment of the present disclosure is given.

[0053] like Figure 1 As shown, when a user uses a video application client at a user terminal (e.g., a mobile phone 111, a desktop computer 112, a tablet computer 113, etc.) to obtain a video to be processed (e.g., a real-time live video or a recorded video, etc.) from a server 130 or a video processing platform through a network 120, the server 130 or the video processing platform can send the video to user terminals such as a mobile phone 111, a desktop computer 112, and a tablet computer 113 through the network 120, and the user can process the received video through the video application client and display the processed video.

[0054] Specifically, in the above process, the user terminal can process the received video according to the video processing method of the exemplary embodiment of the present disclosure. Specifically, the user terminal can obtain a video image, which includes a target object to be processed and an occluded area that occludes the target object. The user terminal can perform image processing on the portion of the target object that is not blocked by the occluded area to obtain a processed video image. According to this method, image processing of the target object can be achieved without affecting the occluded area, thereby avoiding abnormal display of non-target objects in the occluded area, improving the display effect of the processed video image, and improving the viewing experience of the video viewers.

[0055] It should be noted that although the above description uses the user terminal as an example, it is only an example. The executor of the video processing method can be any electronic device, which may include, for example, smart phones, tablet computers, laptop computers, digital assistants, wearable devices, vehicle-mounted terminals, etc., which are physical devices serving as video processing hardware, and may also include software running on physical devices serving as video processing software.

[0056] In addition, despite the above Figure 1 As an example, the above video processing is performed by a user terminal. Figure 1 As shown, the user terminal can perform video processing on the video received from the server, or the user terminal can also perform real-time or post-processing on the video taken or recorded by itself, or, however, the exemplary embodiments of the present disclosure are not limited to this, and the user terminal can also send a request to the server and execute the above-mentioned video processing process through the server, and then send the processed video to the user terminal.

[0057] Figure 2It is a flow chart of a video processing method according to an exemplary embodiment. As described above, the video processing method according to the exemplary embodiment of the present disclosure can perform image processing on the part of the target object that is not blocked by the blocked area based on the acquired video image to obtain a processed video image. In this way, the image processing of the target object can be realized without affecting the blocked area, thereby avoiding abnormal display of non-target objects in the blocked area and improving the display effect of the processed video image.

[0058] like Figure 2 As shown, the video processing method according to an exemplary embodiment of the present disclosure may include the following steps:

[0059] In step S210, a video image may be acquired, wherein the video image includes a target object to be processed and an occlusion region that occludes the target object.

[0060] Here, the video image may be a video image captured in real time, such as a live video, but is not limited thereto, and the video image may also be a video image shot or recorded in advance.

[0061] The target object to be processed can be any object in the video image, such as but not limited to people, objects, background, etc., and can also be any local part of the video image, such as the central area, corner area, etc. of the video image. The target object can be a moving object in the video or a stationary object.

[0062] In step S220, image processing may be performed on the portion of the target object that is not blocked by the blocked area to obtain a processed video image.

[0063] As an example, Figure 3 As shown, step S220 may include the following steps:

[0064] In step S310, a first target area in a video image may be determined.

[0065] In this step, the first target area is an area including the target object. For example, in a video image, a portion of the target object may be exposed, while another portion may be blocked by an obstruction. As an example, the target object may be a person's face, and in a video image, a portion of the person's face may be blocked by headphones, a person's hand, etc. (e.g., the chin). Here, the first target area may be an area including the exposed portion of the target object and the blocked portion of the target object.

[0066] As an example, in a real-time video scenario, based on a pre-trained first detection model, real-time detection and contour recognition of the contour area of ​​a target object such as a face can be performed on a video image acquired in real time. Here, the first detection model can be but is not limited to a real-time detection and segmentation method based on a center network (CenterNet), a mask region-based convolutional neural network (Mask R-CNN), a center mask (CenterMask), a segmentation object by location (Segmenting Objects by Locations, SOLO) method, etc., or it can be based on image algorithms such as clustering and watershed. The present disclosure does not impose any special restrictions on the first detection model and its training method, and any other detection method can also be used to determine the first target area in the video image.

[0067] As an example, Figure 4 As shown, the first target area in the video image can be determined in the following manner:

[0068] In step S410, a visible outline of a target object in a video image is determined.

[0069] In this step, the visible outline of the target object in the video image may be identified first. Since the target object may have an obscured portion, the shape of the visible outline may not be the actual shape of the target object. Figure 5 As shown, when the target object is a human face, due to the presence of an occluder, the visible outline 510 is a discontinuous circle, which lacks a part of the arc.

[0070] In step S420, the visible contour is repaired based on the preset contour characteristics of the target object to obtain a repaired contour of the target object.

[0071] Here, contour characteristics corresponding to the target object may be preset. For example, when the target object is a human face, its contour shape may be circular or elliptical and have a continuous and smooth contour curve. Figure 5 For example, when the visible contour 510 is determined in step S410, the missing contour portion of the visible contour 510 can be repaired according to the preset contour characteristics to obtain the supplementary contour 520. In this way, the entire repaired contour formed by the visible contour 410 and the supplementary contour 520 can be obtained.

[0072] In step S430, the area enclosed by the repair outline is determined as the first target area.

[0073] In this step, it can be considered that the target object is located within the repair contour, and therefore, the area surrounded by the repair contour can be determined as the first target area.

[0074] In this way, the first target area determined by determining the visible outline of the target object and repairing it can include the actual outline shape of the target object, so that even if there are obstructions, image processing can be performed on the area where the entire target object is located, ensuring the integrity of the processed target object and improving the processing effect.

[0075] In step S320 , an occlusion area may be determined in the first target area.

[0076] In this step, the first target area may be detected to determine the area that blocks the target object. For example, non-target objects in the first target area or parts outside the target object may be identified.

[0077] As an example, based on a pre-trained second detection model, abnormal objects intruding into the face can be identified in real time, and the boundaries of the abnormal objects can be segmented. Here, the second detection model can be but is not limited to a segmentation method based on a U-Net method, a DeepLab method, a Fully Convolution Network (FCN) method, etc. The present disclosure does not impose any special restrictions on the second detection model and its training method, and any other detection method can also be used to determine the occlusion area of ​​the occluded target object. Alternatively, the second detection model can also be implemented in the same model as the first detection model.

[0078] In step S330, image processing may be performed on a portion of the target object in the first target area that is not blocked by the blocking area to obtain a second target area.

[0079] In this step, image processing includes but is not limited to performing at least one of the following operations globally or locally: deformation, pixel filtering, color adjustment, and image superposition.

[0080] In the first example, image processing of the portion of the target object not obscured by the occluded area can be achieved by performing image processing on the first target area as a whole. Specifically, image processing can be performed on the portion of the target object in the first target area not obscured by the occluded area and the occluded area to obtain a candidate second target area; the unprocessed occluded area is superimposed on the processed occluded area in the candidate second target area to obtain the second target area.

[0081] In this example, the first target area including the target object and the occluded area can be subjected to the desired image processing as a whole to achieve the image processing of the target object. In this process, the occluded area is also subjected to image processing. Therefore, the original unprocessed occluded area can be superimposed on the processed occluded area in the selected second target area to obtain the second target area, so that only the image processing of the first target area is shown in the second target area, while the occluded area still maintains its original form. In this way, the problem that the occluded area is also processed in the processing of the target object can be solved by superimposing the original unprocessed occluded area, and the processing of the target object can be quickly achieved while keeping the occluded area unchanged.

[0082] As an example, this example can be applied to scenarios where the area of ​​the target object is not enlarged (that is, the area of ​​the target object after processing is less than or equal to the area of ​​the target object before processing), for example, partially or globally reducing the target object, adding attachments to the target object, adjusting the color or brightness of the target object, etc. Taking the live video scene as an example, this example can be applied to but not limited to: special effects such as beautification and skin smoothing that do not cause deformation based on filtering algorithms or deep learning algorithms; global or local face-slimming special effects based on elastic deformation, facial feature point recognition, and deep learning algorithms; facial attachments generated based on the anatomical position relationship of the face, such as cat ears and beards. Based on the face after the above special effects are applied, the detected occluded area is displayed as the top layer, which will not affect the display of the occluded area (such as food, chopsticks, etc.).

[0083] In the second example, if Figure 6 As shown, in step S610, in response to the image processing satisfying the preset completion condition, the pixel completion can be performed on the portion of the target object in the first target area that is occluded by the occluded area to obtain the completed target object; in step S620, the image processing can be performed on the completed target object to obtain the second target area.

[0084] Here, the completion condition may include: the image processing belongs to a preset image processing type; and / or the image processing includes removing an occluded area from the video image.

[0085] As an example, the preset image processing types may include, but are not limited to, special effects processing such as beautification and skin smoothing that do not cause deformation based on filtering algorithms or deep learning algorithms; global or local face-thinning special effects processing based on elastic deformation, facial feature point recognition, and deep learning algorithms; and processing to generate facial attachments (such as cat ears, beards, etc.) based on the facial anatomical position relationship. However, this example is not limited to this and can be applied to scenarios where any image processing is performed on the target object. This example can be applied, but is not limited to, to live video scenarios.

[0086] Furthermore, in the case where the image processing includes removing an occluded area from a video image, the content of the occluded area may not be displayed in the video image finally obtained by the processing, but the entire target object may be fully shown.

[0087] In this example, based on the part of the target object that is blocked by the unblocked area, the part of the target object blocked by the blocked area can be pixel-completed, that is, the blocked area can be pixel-completed. After the target object is completed, the desired image processing can be performed on the entire completed target object to obtain a second target area. In this way, the unblocked part and the blocked part of the target object can be uniformly image-processed, thereby maintaining the integrity of the target object after image processing, providing more possibilities for subsequent processing.

[0088] As an example, pixel completion processing can be performed based on a pre-trained image generation model to complete the currently missing occluded area of ​​the target object, thereby obtaining a completed complete target object, such as a real-time image of the complete target object. Here, the image generation model can be but is not limited to pixel generation pre-training (Generative Pretraining from Pixels, IGPT), conformer based metric generative adversarial networks (conformer based metric, Generative Adversarial Networks, CM-GAN), etc. and other network structures of generative adversarial networks (Generative Adversarial Networks, GAN) structures. The present disclosure does not impose any special restrictions on the image generation model and its training method. In addition, in addition to the method of completing by the image generation model, any other pixel completion method can also be used to complete the pixel of the part of the target object that is occluded by the occluded area.

[0089] In addition, in response to the image processing not satisfying the preset completion condition, the completion operation may not be performed, and image processing is performed on a portion of the target object in the first target area that is not blocked by the blocked area to obtain a second target area.

[0090] The above describes a process of performing image processing on the portion of the target object in the first target area that is not obscured by the obscured area. In this process, the image processing described in step S330 can be directly performed on the first target area in the video image to obtain the second target area. Specifically, the above processing can be applied to the first target area including the portion of the target object that is not obscured by the obscured area and the obscured area, wherein no special processing is performed on the obscured area.

[0091] However, the exemplary embodiments of the present disclosure are not limited thereto, and the occluded area may also be removed before performing image processing. Specifically, before performing image processing on the portion of the target object in the first target area that is not occluded by the occluded area, the occluded area may be removed from the first target area according to the position of the occluded area in the first target area to obtain the first target area after the occlusion is removed. Based on this, the processing described in the above step S330 may be performed on the first target area after the occlusion is removed to obtain the second target area. In this way, after removing the occluded area, it may be easier to perform image processing on the first target area, reduce the amount of calculation caused by the existence of the occluded area, and improve processing efficiency.

[0092] As an example, according to the position of the occluded area in the first target area, all pixel values ​​of the occluded area in the first target area may be set to preset pixel values ​​to obtain the first target area after occlusion is removed.

[0093] Here, the preset pixel value may be, for example, 0, so that the first target area after occlusion includes the target object and the occluded area filled with black. In this way, the occluded areas can be unified into the same pixel value, which is convenient for the management and storage of pixel data in subsequent processing and improves the calculation speed. Although it is described here that the occluded area is removed by setting all the pixel values ​​of the occluded area to the same value, the exemplary embodiments of the present disclosure are not limited thereto, and for example, all pixels in the occluded area may also be removed to facilitate subsequent image processing.

[0094] As for whether to remove the occluded area, in the first example of step S330, the portion of the target object in the first target area that is not blocked by the occluded area and the original occluded area can be image processed to obtain a candidate second target area. Since in this example, the unprocessed original occluded area will be superimposed on the processed occluded area later, the processed occluded area obtained based on the original occluded area processing will be covered and will not be shown in the processed video image. The above process can also be performed after removing the occluded area. Specifically, the portion of the target object in the first target area that is not blocked by the occluded area and the removed occluded area can be image processed to obtain a candidate second target area.

[0095] In the second example of step S330, as shown in FIG. Figure 6 As described, the pixel complementation can be performed on the portion of the target object in the first target area that is blocked by the blocked area, that is, the pixel complementation can be performed on the original blocked area. In this example, since the original blocked area no longer exists after the complementation, what is displayed is the complemented area, the complementation can be performed after removing the blocked area, or the complementation can be performed directly without removing the blocked area.

[0096] In step S340, a processed video image may be obtained based on the second target area.

[0097] In this step, without completing the target object, the second target area can be directly displayed in the processed video image to obtain the final video image.

[0098] In addition, when completing the target object, in one example, the occluded area can be superimposed on the second target area according to the position of the occluded area in the video image to obtain a processed video image. Alternatively, in another example, the second target area can be displayed in its entirety in the processed video image, and the occluded area is not displayed in the processed video image. Compared with the superimposed occluded area in the above example, not displaying the occluded area in this example is more suitable for highly emphasizing the target object to which the image processing is applied, and the visual area of ​​the target object is not desired to be occluded by other distracting objects.

[0099] In this way, based on the completed target object, the occluded area can be selectively superimposed back to the video image or no longer displayed according to actual needs, thereby making image processing more flexible and suitable for more video processing requirements.

[0100] Refer to above Figure 2-Figure 6 The video processing method according to the exemplary embodiment of the present disclosure is described below, and the real-time video scene is used as a reference. Figure 7 An example of the video processing method is described.

[0101] like Figure 7 As shown, in step S701, real-time image data may be obtained. Here, the real-time image data may be, for example, but not limited to, real-time picture data in a live video scene or a video connection scene. The data may be initiated by a client, for example, and received by a server, and all calculation steps may be completed on the server, but it is not limited thereto, and all or part of the calculation steps may also be completed by the client itself.

[0102] In step S702, a first target area including a target object may be detected. The target object may be an object or area in an image where a special effect is expected to act, such as a human face, and the special effect may be, for example, a face-thinning special effect.

[0103] In step S703, an occlusion area in the first target area may be detected. The occlusion area may be, for example, an intrusion that covers the target object.

[0104] In step S704, the occluded area may be removed. Specifically, after the occluded area covering the target object is detected, the occluded area may be removed, for example, the occluded area may be removed from the original image or reset with the same pixel value.

[0105] In step S705, it can be determined whether to complete the target object. After the occluded area is cut out in the original image, there will inevitably be a vacant display in the original occluded area. Therefore, it is necessary to determine whether to complete the image of the target object, so as to restore the complete target object with the help of the algorithm.

[0106] If completion is required, the target object may be completed in step S706, and then step S707 is performed to perform image processing on the first target area. For example, for some effects that stretch the area after acting on the target object, the target object needs to be completed to fill the part after the mask area is deducted, so as to ensure that the final visual effect is complete.

[0107] If completion is not required, step S707 can be directly executed to perform image processing on the first target area. For example, for some applications that reduce the target object, it is possible to choose not to complete the cut-out area. For example, for a face-slimming effect, after cutting out the occluded area, the face-slimming effect is applied without completion. After application, the cut-out occluded area will be compressed along with the face-slimming effect. At this time, the original image of the covered area is added for combination, and the final visual effect can achieve that the covered area part completely covers the vacant position after the face-slimming effect.

[0108] In step S708, it can be determined whether to superimpose the occluded area. If it is necessary to superimpose the occluded area, then in step S709, the occluded area can be superimposed. If not, the processed video image can be directly obtained. Specifically, after applying special effects to the target object, it can be selected whether to re-superimpose the subtracted occluded area image layer. In the case where the target object is completed, it can be selected not to superimpose.

[0109] As described above, when the video processing method according to the exemplary embodiment of the present disclosure is applied to a real-time video scene, it can improve the user experience of the special effects application in the real-time video scene. For example, when a user (such as a host) turns on the face shaping special effect, the deformation of the objects in front of the face will occur. This obvious visual abnormality will divert the viewer's attention and is not conducive to the communication of the video content. The application of the video processing method according to the exemplary embodiment of the present disclosure can optimize this situation, so that the objects in the occluded area are no longer affected by the image processing of the target object and will not be abnormally deformed, thereby improving the video viewing experience, removing visual interference, allowing the viewer to pay more attention to the content expressed by the user, and also giving the user a better special effects experience.

[0110] Fig. 8A and Figure 8B Schematic diagrams respectively show a video image processed according to an existing video processing method and a video image processed according to a video processing method of an exemplary embodiment.

[0111] Fig. 8A and Figure 8B This is a comparison of the effect of horizontal reduction applied to the original image with the occluded area. Fig. 8A As shown in FIG. 1 , when the existing method is applied to perform processing, it can be clearly seen that the expected effect of only shrinking the target horizontally is obvious, while in the overlapping area of ​​the occluded area and the target object, the occluded area is also deformed, which is not expected. Figure 8B As shown, when the processing method of the embodiment of the present disclosure is used, the result of the final effect application is that only the target object changes, while the occluded area overlapping the target object remains the same as before.

[0112] Fig. 9 is a block diagram of a video processing device according to an exemplary embodiment. Fig. 9 The video processing device includes an acquisition unit 100 and a processing unit 200.

[0113] The acquisition unit 100 is configured to acquire a video image, wherein the video image includes a target object to be processed and an occlusion region that occludes the target object.

[0114] The processing unit 200 is configured to perform image processing on the portion of the target object that is not blocked by the blocked area to obtain a processed video image.

[0115] As an example, the processing unit 200 may include a first object region determining unit 210 , an occlusion region determining unit 220 , a second object region determining unit 230 , and a video image determining unit 240 .

[0116] The first target region determining unit 210 is configured to determine a first target region in a video image, wherein the first target region is a region containing a target object.

[0117] The occlusion region determining unit 220 is configured to determine an occlusion region in the first target region.

[0118] The second target region determining unit 230 is configured to perform image processing on a portion of the target object in the first target region that is not blocked by the blocking region to obtain a second target region.

[0119] The video image determination unit 240 is configured to obtain a processed video image based on the second target area.

[0120] As an example, the second target area determination unit 230 is also configured to: perform image processing on the portion of the target object in the first target area that is not obscured by the obscured area and the obscured area to obtain a candidate second target area; and superimpose the unprocessed obscured area on the processed obscured area in the candidate second target area to obtain a second target area.

[0121] As an example, the video processing device also includes a removal unit, which is configured to: before performing image processing on a portion of the target object in the first target area that is not obscured by the occlusion area, remove the occlusion area from the first target area according to the position of the occlusion area in the first target area to obtain the first target area after the occlusion is removed.

[0122] As an example, the removal unit is further configured to: according to the position of the occluded area in the first target area, set all pixel values ​​of the occluded area in the first target area to preset pixel values ​​to obtain the first target area after occlusion is removed.

[0123] As an example, the second target area determination unit 230 is also configured to: in response to the image processing not satisfying a preset completion condition, perform image processing on the portion of the target object in the first target area that is not obscured by the occluded area to obtain a second target area; in response to the image processing satisfying a preset completion condition, perform pixel completion on the portion of the target object in the first target area that is obscured by the occluded area to obtain a completed target object, wherein the completion condition includes: the image processing belongs to a preset image processing type and / or the image processing includes removing the occluded area from the video image; and perform image processing on the completed target object to obtain a second target area.

[0124] As an example, the video image determination unit 240 is also configured to: superimpose the occluded area onto the second target area according to the position of the occluded area in the video image to obtain a processed video image; or, display the entire second target area in the processed video image, and not display the occluded area in the processed video image.

[0125] As an example, the first target area determination unit 210 is also configured to: determine the visible outline of the target object in the video image; based on preset outline characteristics of the target object, patch the visible outline to obtain a patched outline of the target object; and determine the area enclosed by the patched outline as the first target area.

[0126] As an example, the image processing includes performing at least one of the following operations globally or locally: deformation, pixel filtering, color adjustment, and image overlay, and the video image is a video image acquired in real time.

[0127] Regarding the device in the above embodiment, the specific manner in which each unit performs the operation has been described in detail in the embodiment of the method, and will not be elaborated here.

[0128] Fig.10 FIG. 1 is a block diagram of an electronic device according to an exemplary embodiment. Fig.10 As shown, the electronic device 10 includes a processor 101 and a memory 102 for storing processor executable instructions. Here, when the processor executable instructions are executed by the processor, the processor is prompted to execute the video processing method as described in the above exemplary embodiment.

[0129] As an example, the electronic device 10 is not necessarily a single device, but may be any collection of devices or circuits that can execute the above instructions (or instruction sets) individually or in combination. The electronic device 10 may also be part of an integrated control system or system manager, or may be configured as an electronic device interconnected with a local or remote (e.g., via wireless transmission) interface.

[0130] In the electronic device 10, the processor 101 may include a central processing unit (CPU), a graphics processing unit (GPU), a programmable logic device, a dedicated processor system, a microcontroller or a microprocessor. As an example and not a limitation, the processor 101 may also include an analog processor, a digital processor, a microprocessor, a multi-core processor, a processor array, a network processor, etc.

[0131] The processor 101 may execute instructions or codes stored in the memory 102, wherein the memory 102 may also store data. Instructions and data may also be sent and received over a network via a network interface device, wherein the network interface device may employ any known transmission protocol.

[0132] The memory 102 may be integrated with the processor 101, for example, by placing RAM or flash memory within an integrated circuit microprocessor or the like. In addition, the memory 102 may include a separate device, such as an external disk drive, a storage array, or any other storage device that can be used by a database system. The memory 102 and the processor 101 may be operatively coupled, or may communicate with each other, such as through an I / O port, a network connection, etc., so that the processor 101 can read files stored in the memory 102.

[0133] In addition, the electronic device 10 may also include a video display (such as a liquid crystal display) and a user interaction interface (such as a keyboard, a mouse, a touch input device, etc.) All components of the electronic device 10 may be connected to each other via a bus and / or a network.

[0134] In an exemplary embodiment, a computer-readable storage medium may also be provided. When the instructions in the computer-readable storage medium are executed by a processor of an electronic device, the electronic device can perform the video processing method as described in the above exemplary embodiment. The computer-readable storage medium may be, for example, a memory including instructions. Optionally, the computer-readable storage medium may be: a read-only memory (ROM), a random access memory (RAM), a random access programmable read-only memory (PROM), an electrically erasable programmable read-only memory (EEPROM), a dynamic random access memory (DRAM), a static random access memory (SRAM), a flash memory, a non-volatile memory, a CD-ROM, a CD-R, a CD+R, a CD-RW, a CD+RW, a DVD-ROM, a DVD-R, a DVD+R, a DVD-RW, a DVD+RW, a DVD-RAM, a BD-ROM, a BD-R, a BD-R LTH, BD-RE, Blu-ray or optical disk storage, hard disk drive (HDD), solid state drive (SSD), card storage (such as, multimedia card, secure digital (SD) card or extreme digital (XD) card), magnetic tape, floppy disk, magneto-optical data storage device, optical data storage device, hard disk, solid state disk and any other device, any other device is configured to store computer programs and any associated data, data files and data structures in a non-transitory manner and provide the computer programs and any associated data, data files and data structures to a processor or computer so that the processor or computer can execute the computer program. The computer program in the above-mentioned computer-readable storage medium can be run in an environment deployed in a computer device such as a client, a host, an agent device, a server, etc. In addition, in one example, the computer program and any associated data, data files and data structures are distributed on a networked computer system, so that the computer program and any associated data, data files and data structures are stored, accessed and executed in a distributed manner by one or more processors or computers.

[0135] In an exemplary embodiment, a computer program product may also be provided. The computer program product includes computer instructions. When the computer instructions are executed by a processor, the video processing method as described in the exemplary embodiment above is implemented.

[0136] Those skilled in the art will readily appreciate other embodiments of the present disclosure after considering the specification and practicing the invention disclosed herein. The present disclosure is intended to cover any variations, uses or adaptations of the present disclosure that follow the general principles of the present disclosure and include common knowledge or customary techniques in the art that are not disclosed in the present disclosure. The description and examples are to be considered exemplary only, and the true scope and spirit of the present disclosure are indicated by the following claims.

[0137] It should be understood that the present disclosure is not limited to the exact structures that have been described above and shown in the drawings, and that various modifications and changes may be made without departing from the scope thereof. The scope of the present disclosure is limited only by the appended claims.

Claims

1. A video processing method, It is characterized in that The video processing method comprises: Acquire a video image, wherein the video image includes a target object to be processed and an occlusion area that occludes the target object; Performing image processing on the portion of the target object that is not blocked by the blocking area to obtain a processed video image, The performing image processing on the portion of the target object not blocked by the blocking area to obtain a processed video image includes: determining a first target area in the video image, wherein the first target area is an area containing the target object; determining the blocking area in the first target area; performing image processing on the portion of the target object in the first target area not blocked by the blocking area to obtain a second target area; and obtaining a processed video image based on the second target area. Before performing image processing on a portion of the target object in the first target area that is not blocked by the blocking area, the video processing method further comprises: removing the blocking area from the first target area according to a position of the blocking area in the first target area to obtain a first target area after the blocking is removed, wherein the blocking area is removed by removing all pixels in the blocking area or setting the pixels in the blocking area to the same pixel value, Among them, the image processing is performed on the part of the target object in the first target area that is not obscured by the occlusion area to obtain the second target area, including: in response to the image processing satisfying a preset completion condition, pixel completion is performed on the part of the target object in the first target area that is obscured by the occlusion area to obtain a completed target object, wherein the completion condition includes: the image processing will stretch the processed area after acting on the target object; and the image processing is performed on the completed target object to obtain the second target area.

2. The video processing method according to claim 1, It is characterized in that The performing image processing on a portion of the target object in the first target area that is not blocked by the blocking area to obtain a second target area comprises: Performing image processing on a portion of the target object in the first target area that is not blocked by the blocking area and the blocking area to obtain a candidate second target area; The unprocessed occlusion region is superimposed on the processed occlusion region in the candidate second target region to obtain the second target region.

3. The video processing method according to claim 1, It is characterized in that The removing the occluded area from the first target area according to the position of the occluded area in the first target area to obtain the first target area after the occlusion is removed includes: According to the position of the occluded area in the first target area, all pixel values ​​of the occluded area in the first target area are set to preset pixel values ​​to obtain the first target area after the occlusion is removed.

4. The video processing method according to claim 1, It is characterized in that The performing image processing on the portion of the target object in the first target area that is not blocked by the blocking area to obtain the second target area further includes: In response to the image processing not satisfying the completion condition, performing image processing on a portion of the target object in the first target area that is not blocked by the blocking area to obtain a second target area, Wherein, the completion condition also includes: the image processing includes removing the blocked area from the video image.

5. The video processing method according to claim 4, It is characterized in that The step of obtaining a processed video image based on the second target area includes: According to the position of the occluded area in the video image, the occluded area is superimposed on the second target area to obtain the processed video image; or, The second target area is fully displayed in the processed video image, and the blocked area is not displayed in the processed video image.

6. The video processing method according to claim 1, It is characterized in that The determining of the first target area in the video image includes: Determining a visible outline of the target object in the video image; Based on the preset contour characteristics of the target object, the visible contour is repaired to obtain a repaired contour of the target object; The area enclosed by the repair outline is determined as the first target area.

7. The video processing method according to claim 1, It is characterized in that The image processing includes performing at least one of the following operations globally or locally: deformation, pixel filtering, color adjustment and image superposition, and the video image is a video image collected in real time.

8. A video processing device, It is characterized in that The video processing device comprises: An acquisition unit is configured to acquire a video image, wherein the video image includes a target object to be processed and an occlusion area that occludes the target object; a processing unit configured to perform image processing on a portion of the target object that is not blocked by the blocking area to obtain a processed video image, The processing unit includes a first target area determination unit, an occlusion area determination unit, a second target area determination unit and a video image determination unit. The first target region determining unit is configured to determine a first target region in the video image, wherein the first target region is a region including the target object. The occlusion region determining unit is configured to determine the occlusion region in the first target region. The second target area determination unit is configured to perform image processing on a portion of the target object in the first target area that is not blocked by the blocking area to obtain a second target area. The video image determination unit is configured to obtain a processed video image based on the second target area, The video processing device further includes a removal unit, which is configured to: before performing image processing on the portion of the target object in the first target area that is not blocked by the blocking area, remove the blocking area from the first target area according to the position of the blocking area in the first target area to obtain a first target area after the blocking is removed, wherein the blocking area is removed by removing all pixels in the blocking area or setting the pixels in the blocking area to the same pixel value, Among them, the second target area determination unit is further configured to: in response to the image processing satisfying a preset completion condition, perform pixel completion on the portion of the target object in the first target area that is occluded by the occlusion area to obtain a completed target object, wherein the completion condition includes: the image processing will stretch the processed area after acting on the target object; and perform the image processing on the completed target object to obtain a second target area.

9. An electronic device, It is characterized in that The electronic device comprises: processor; a memory for storing instructions executable by the processor, When the processor executable instructions are executed by the processor, the processor is prompted to execute the video processing method according to any one of claims 1 to 7.

10. A computer-readable storage medium, It is characterized in that When the instructions in the computer-readable storage medium are executed by a processor of an electronic device, the electronic device is enabled to perform the video processing method according to any one of claims 1 to 7.

11. A computer program product, It is characterized in that The computer program product comprises computer instructions, and when the computer instructions are executed by a processor, the video processing method according to any one of claims 1 to 7 is implemented.

Citation Information

Patent Citations

  • Image processing method and device, electronic equipment and storage medium

    CN110929651A

  • Image processing method, device and equipment and computer storage medium

    CN113284041A