Portrait elimination method and device and storage medium
By identifying the similarity between the face features in the video frame and the specified face features, the target video frame is automatically determined, and the target portrait is eliminated using the image repair model, the problems of low efficiency and unstable quality of portrait elimination in the prior art are solved, and efficient and reliable portrait elimination effect is achieved.
Patent Information
- Application Number
- CN202411817275.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2024-12-10
- Publication Date
- 2025-05-16
AI Technical Summary
The prior art cannot realize the automatic elimination of designated portraits in videos, resulting in low video editing efficiency and unstable quality.
By identifying the similarity between face features in video frames and specified face features, the target video frame is automatically determined, and the target portrait is eliminated using the image repair model to achieve automated portrait elimination.
This greatly improves the efficiency of portrait elimination at video level, ensures the reliability and quality of the elimination process, and avoids the cumbersome and errors of manual frame-by-frame processing.
Smart Images

Figure CN120013793A_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of video processing technology, and in particular to a method, device and storage medium for removing a human portrait. Background Art
[0002] In the field of multimedia processing, especially in video editing and post-production, there is often a need to remove specified human images in a video.
[0003] Currently, the above requirements are mostly met through video editing software. Video editing software can provide a series of image editing tools, allowing users to view video content frame by frame and manually operate image editing tools to remove portraits in each frame, but does not support automatic removal of specified portraits. Summary of the invention
[0004] In a first aspect, an embodiment of the present application provides a method for removing a human portrait, comprising: In response to a designated operation on the video, determining designated facial features of a target portrait designated by the designated operation from the video; Determining a target video frame from the video frames based on the similarity between the facial features of a person included in the video frames in the video and the designated facial features; Obtaining a mask of the target portrait in the target video frame, inputting the target video frame and the mask of the target portrait into an image restoration model, and obtaining a restoration image output by the image restoration model, wherein the mask of the target portrait is used to specify a to-be-restored area of the target video frame, and the restoration image no longer includes the target portrait; Extracting a repaired area from the repaired image based on the mask of the target portrait, and extracting the target portrait from the target video frame based on the mask of the target portrait to obtain an image to be synthesized; The repaired area is superimposed on the image to be synthesized to obtain a repaired video frame after the target portrait is removed from the target video frame.
[0005] In some embodiments, determining the specified facial features of the target portrait specified by the specified operation from the video includes: Performing image segmentation on the video frames associated with the specified operation in the video to obtain masks of one or more first portraits in the associated video frames; Determine one or more first portrait detection frames according to the one or more first portrait masks; Determine a first portrait detection frame corresponding to the specified operation, and perform face detection on a region of interest sub-image corresponding to the first portrait detection frame corresponding to the specified operation; In response to detecting a designated face in the region of interest sub-image, features of the designated face are extracted as the designated face features, and a first portrait corresponding to a first portrait detection frame where the designated operation occurs is the target portrait.
[0006] In some embodiments, determining a target video frame from the video frames based on the similarity between the facial features of a person included in the video frame in the video and the specified facial features includes: Obtaining candidate facial features of each of the video frames; Determining that the video frame has a similarity between the candidate facial feature and the designated facial feature that is greater than a preset threshold, then using the video frame as the target video frame; Wherein, obtaining the mask of the target portrait in the target video frame includes: Acquire a candidate face detection frame corresponding to the candidate face feature whose similarity is greater than a preset threshold as a target face detection frame; perform image segmentation on the target video frame to obtain a mask of one or more second portraits in the target video frame; Based on the overlap relationship between the masks of the one or more second portraits and the target face detection frame, the mask of the target portrait is determined from the masks of the one or more second portraits.
[0007] In some embodiments, it also includes: Taking the target video frame as a starting point, performing at least one of forward video tracking and backward video tracking on the video to obtain a video frame in which the target portrait still exists in the video, and using the video frame in which the target portrait still exists as a tracking video frame; The tracking video frame is used as a new target video frame.
[0008] In some embodiments, taking the target video frame as a starting point, performing forward video tracking on the video, obtaining a video frame in which the target portrait is still present in the video, and using the video frame in which the target portrait is still present as a tracking video frame, includes: Taking any target video frame as a current frame, in response to a video frame preceding the current frame being a non-target video frame, judging whether the previous video frame contains the target human portrait based on the overlap relationship between the masks of the target human portrait in the current frame and the masks of the human portraits in the previous video frame; If it is determined that the previous video frame contains the target portrait, the previous video frame is used as the tracking video frame, and in response to the previous video frame being a non-transition frame, the previous video frame is used as the new current frame for the forward video tracking.
[0009] In some embodiments, it also includes: It is determined that the previous video frame does not contain the target portrait, or it is determined that the previous video frame contains the target portrait and the previous video frame is a transition frame, and the forward video tracking ends.
[0010] In some embodiments, taking the target video frame as a starting point, performing backward video tracking on the video to obtain a video frame in which the target portrait is still present in the video, and using the video frame in which the target portrait is still present as a tracking video frame includes: Taking any target video frame as the current frame, in response to a subsequent video frame of the current frame being a non-target video frame and a non-transition frame, judging whether the subsequent video frame contains the target portrait based on the overlap relationship between the masks of the target portrait in the current frame and the masks of each portrait in the subsequent video frame; If it is determined that the target portrait is included in the subsequent video frame, the subsequent video frame is used as the tracking video frame, and the subsequent video frame is used as the new current frame for the backward video tracking.
[0011] In some embodiments, it also includes: It is determined that a subsequent video frame of the current frame is the transition frame, or it is determined that a subsequent video frame of the current frame does not contain the target portrait, and the backward video tracking ends.
[0012] In some embodiments, it also includes: The target video frame in the video is replaced with the repaired video frame.
[0013] In a second aspect, an embodiment of the present application further provides an electronic device, comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein when the processor executes the program, a portrait removal method as described above is implemented.
[0014] In a third aspect, an embodiment of the present application further provides a non-transitory computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements any of the above-described portrait removal methods.
[0015] In a fourth aspect, an embodiment of the present application further provides a computer program product, including a computer program, which, when executed by a processor, implements any of the above-mentioned portrait removal methods. BRIEF DESCRIPTION OF THE DRAWINGS
[0016] In order to more clearly illustrate the technical solutions in the present application or related technologies, the following is a brief introduction to the drawings required for use in the embodiments or related technical descriptions. Obviously, the drawings described below are some embodiments of the present application. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying any creative work.
[0017] Figure 1 It is a schematic diagram of the structure of the terminal provided in the embodiment of the present application.
[0018] Figure 2 This is one of the flow charts of the portrait removal method provided in the embodiment of the present application.
[0019] Figure 3 It is a flowchart of the image restoration method provided in an embodiment of the present application.
[0020] Figure 4 It is a schematic diagram of the target portrait designation process provided in an embodiment of the present application.
[0021] Figure 5 Schematic diagram of a mask of a second portrait and a target face detection frame provided in an embodiment of the present application.
[0022] Figure 6 It is a flowchart of the method for forward video tracking provided in an embodiment of the present application.
[0023] Figure 7 It is a flowchart of the method for backward video tracking provided in an embodiment of the present application.
[0024] Figure 8 This is the second flowchart of the portrait removal method provided in the embodiment of the present application.
[0025] Fig. 9 It is a schematic diagram of the structure of an electronic device provided in an embodiment of the present application. DETAILED DESCRIPTION
[0026] In order to make the purpose, technical solutions and advantages of this application clearer, the technical solutions in this application will be clearly and completely described below in conjunction with the drawings in this application. Obviously, the described embodiments are part of the embodiments of this application, not all of them. Based on the embodiments in this application, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of this application.
[0027] The terms "first", "second", etc. in the specification and claims of the present application are used to distinguish similar objects, and are not used to describe a specific order or sequence. It should be understood that the terms used in this way are interchangeable under appropriate circumstances, so that the embodiments of the present application can be implemented in an order other than those illustrated or described here, and the objects distinguished by "first" and "second" are generally of the same type, and the number of objects is not limited. For example, the first object can be one or more. In addition, "and / or" in the specification and claims represents at least one of the connected objects, and the character " / " generally represents that the objects associated with each other are in an "or" relationship.
[0028] In the field of multimedia processing, especially in video editing and post-production, there is often a need to modify or remove specific content in a video. Among them, removing specific portraits in a video is a common and important application scenario. For example, in order to improve the viewing experience of the video or protect the privacy of others, the creator may remove irrelevant portraits of passers-by in the video, or in order to produce a video that meets the playback requirements, some portraits may need to be removed in the later stage.
[0029] Currently, this function is mainly achieved by video editing software. Video editing software usually provides a series of image editing tools, allowing users to view the video content frame by frame and manually remove the specified portrait in each frame. Although this method can meet the needs of portrait removal to a certain extent, its operation process is cumbersome and time-consuming, requiring users to have high image editing skills and patience. More importantly, since it requires manual frame-by-frame processing, this method cannot achieve automatic removal of specified portraits, which greatly limits the efficiency and flexibility of video editing.
[0030] In addition, as users' requirements for video editing speed and quality continue to increase, the current method of manually processing frames to remove specified people in videos is no longer able to meet the needs of practical applications. Especially when processing long videos, the workload of manual frame-by-frame processing is huge, which not only increases production costs, but may also cause unstable video editing quality due to human factors.
[0031] To this end, the present application provides a method for portrait removal, which determines the target video frame and performs portrait removal based on the similarity between the facial features of the faces contained in each video frame in the video and the specified facial features. In this process, the user only needs to specify the target portrait, and the determination of the target video frame and the portrait removal for the target video frame can be automatically realized, which greatly improves the efficiency of video-level portrait removal. In addition, in the process of portrait removal, the target video frame is repaired based on the image repair model to obtain a repaired image, and the image to be synthesized from the target video frame and the repaired image obtained by cutting out the target portrait from the repaired image are superimposed, so that the portrait removal of each target video frame can be achieved, which further ensures the reliability and quality of portrait removal.
[0032] The portrait removal method provided in the embodiment of the present application can be applied to electronic devices, specifically to mobile phones, tablet computers, wearable devices, vehicle-mounted devices, augmented reality (AR) / virtual reality (VR) devices, laptop computers, ultra-mobile personal computers (UMPC), netbooks, personal digital assistants (PDA) and other terminals. It can also be applied to databases, servers and service response systems based on terminal artificial intelligence. The embodiment of the present application does not impose any restrictions on the specific type of terminal.
[0033] For example, the terminal can be a station (STAION, ST) in a WLAN, a cellular phone, a cordless phone, a Session Initiation Protocol (SIP) phone, a Wireless Local Loop (WLL) station, a Personal Digital Assistant (PDA) device, a handheld device with wireless communication function, a computing device or other processing device connected to a wireless modem, a computer, a laptop computer, a handheld communication device, a handheld computing device, and / or other devices for communicating on a wireless system and a next-generation communication system, such as a mobile terminal in a 5G network, a mobile terminal in a future evolved Public Land Mobile Network (PLMN) or a mobile terminal in a future evolved Non-terrestrial Network (NTN), etc.
[0034] As an example but not limitation, when the terminal is a wearable device, the wearable device can also be a general term for wearable devices that are developed by applying wearable technology to intelligently design daily wearables, such as gloves, watches, AR (Augmented Reality) head-mounted display devices, VR (Virtual Reality) head-mounted display devices or MR (Mixed Reality) head-mounted display devices equipped with far-field communication modules and / or near-field communication modules.
[0035] In some embodiments, the terminal may be a Figure 1 The mobile phone 100 of the hardware structure shown in FIG. Figure 1 As shown, the mobile phone 100 may specifically include: a radio frequency (RF) circuit 110, a memory 120, an input unit 130, a display unit 140, a sensor 150, an audio circuit 160, a short-range wireless communication module 170, a processor 180, and a power supply 190. Those skilled in the art will appreciate that Figure 1 The structure of the mobile phone 100 shown in the figure does not constitute a limitation on the electronic device, and the electronic device may include more or fewer components than shown in the figure, or combine certain components, or arrange the components differently.
[0036] Combine the following Figure 1 A detailed introduction to the various components of the mobile phone: The RF circuit 110 can be used for receiving and sending signals during information transmission or calls. In particular, after receiving the downlink information of the base station, it is sent to the processor 180 for processing; in addition, the designed uplink data is sent to the base station. Generally, the RF circuit includes but is not limited to an antenna, at least one amplifier, a transceiver, a coupler, a low noise amplifier (Low Noise Amplifier, LNA), a duplexer, etc. In addition, the RF circuit 110 can also communicate with the network and other devices through wireless communication. The above-mentioned wireless communication can use any communication standard or protocol, and the wireless communication can include global system for mobile communications (global system for mobile communications, GSM), general packet radio service (general packet radio service, GPRS), code division multiple access (code division multiple access, CDMA), wideband code division multiple access (wideband code division multiple access, WCDMA), time division code division multiple access (time-division code division multiple access, TD-SCDMA), long term evolution (long term evolution, LTE), new radio (new radio, NR), GNSS, FM, low-orbit satellite connection and / or IR technology, etc. The GNSS may include a global positioning system (GPS), a global navigation satellite system (GLONASS), a Beidou navigation satellite system (BDS), a quasi-zenith satellite system (QZSS) and / or a satellite based augmentation system (SBAS), etc.
[0037] The memory 120 can be used to store software programs and modules. The processor 180 executes various functional applications and data processing of the mobile phone by running the software programs and modules stored in the memory 120. The memory 120 may mainly include a program storage area and a data storage area, wherein the program storage area may store an operating system, an application required for at least one function (such as a sound playback function, an image playback function, etc.), etc.; the data storage area may store data created according to the use of the mobile phone (such as pictures, audio data, phone books, etc.), etc. In addition, the memory 120 may include a high-speed random access memory, and may also include a non-volatile memory, such as at least one disk storage device, a flash memory device, or other volatile solid-state storage devices. Specifically, the memory 120 may store pictures taken by an electronic device or downloaded via a wireless network.
[0038] The input unit 130 can be used to receive input digital or character information, and to generate key signal input related to the user settings and function control of the mobile phone 100. Specifically, the input unit 130 may include a touch panel 131 and other input devices 132. The touch panel 131, also known as a touch screen, can collect the user's touch operation on or near it (such as the user's operation on the touch panel 131 or near the touch panel 131 using any suitable object or accessory such as a finger, stylus, etc.), and drive the corresponding connection device according to a pre-set program. Optionally, the touch panel 131 may include two parts: a touch detection device and a touch controller. Among them, the touch detection device detects the user's touch orientation, detects the signal brought by the touch operation, and transmits the signal to the touch controller; the touch controller receives the touch information from the touch detection device, converts it into the touch point coordinates, and then sends it to the processor 180, and can receive and execute the command sent by the processor 180. In addition, the touch panel 131 can be implemented in various types such as resistive, capacitive, infrared, and surface acoustic wave. In addition to the touch panel 131, the input unit 130 may also include other input devices 132. Specifically, other input devices 132 may include but are not limited to one or more of a physical keyboard, function keys (such as volume control keys, switch keys, etc.), a trackball, a mouse, a joystick, etc.
[0039] The display unit 140 can be used to display information input by the user or information provided to the user and various menus of the mobile phone. The display unit 140 may include a display panel 141. Optionally, the display panel 141 may be configured in the form of a liquid crystal display (LCD), a light emitting diode (LED), an organic light emitting diode (OLED), an active-matrix organic light emitting diode (AMOLED), etc. Further, the touch panel 131 may cover the display panel 141. When the touch panel 131 detects a touch operation on or near it, it is transmitted to the processor 180 to determine the type of touch event. Subsequently, the processor 180 provides a corresponding visual output on the display panel 141 according to the type of touch event. Although in Figure 1 In the embodiment, the touch panel 131 and the display panel 141 are used as two independent components to realize the input and output functions of the mobile phone, but in some embodiments, the touch panel 131 and the display panel 141 can be integrated to realize the input and output functions of the mobile phone.
[0040] The mobile phone 100 may also include at least one sensor 150, such as a light sensor, a motion sensor, and other sensors. Specifically, the light sensor may include an ambient light sensor and a proximity sensor, wherein the ambient light sensor may adjust the brightness of the display panel 141 according to the brightness of the ambient light, and the proximity sensor may turn off the display panel 141 and / or the backlight when the mobile phone is moved to the ear. As a type of motion sensor, the accelerometer sensor can detect the magnitude of acceleration in all directions (generally three axes), and can detect the magnitude and direction of gravity when stationary. It can be used for applications that identify the posture of the mobile phone (such as horizontal and vertical screen switching, related games, magnetometer posture calibration), vibration recognition related functions (such as pedometer, tapping), etc.; as for other sensors that can be configured in the mobile phone, such as gyroscopes, barometers, hygrometers, thermometers, infrared sensors, etc., they will not be repeated here.
[0041] The audio circuit 160, the speaker 161, and the microphone 162 can provide an audio interface between the user and the mobile phone. The audio circuit 160 can transmit the received audio data to the speaker 161 after converting the received audio data into an electrical signal, which is converted into a sound signal for output; on the other hand, the microphone 162 converts the collected sound signal into an electrical signal, which is received by the audio circuit 160 and converted into audio data, and then the audio data is output to the processor 180 for processing, and then sent to another electronic device through the RF circuit 110, or the audio data is output to the memory 120 for further processing.
[0042] Communication technologies such as Wi-Fi, Bluetooth, and Near Field Communication (NFC) belong to short-range wireless transmission technologies. The mobile phone can help users send and receive emails, browse web pages, access streaming media, etc. through the short-range wireless communication module 170, which provides users with wireless broadband Internet access. The above-mentioned short-range wireless communication module 170 may include a Wi-Fi chip, a Bluetooth chip, and an NFC chip. The Wi-Fi chip can realize the function of Wi-Fi Direct connection between the mobile phone 100 and other electronic devices, and can also make the mobile phone 100 work in an AP mode (Access Point mode) that can provide wireless access services and allow other wireless devices to access, or work in a STA mode (Station mode) that can connect to an AP but does not accept wireless devices to access, thereby establishing point-to-point communication between the mobile phone 100 and other Wi-Fi devices.
[0043] The processor 180 is the control center of the mobile phone. It uses various interfaces and lines to connect various parts of the entire mobile phone. By running or executing software programs and / or modules stored in the memory 120, and calling data stored in the memory 120, it executes various functions of the mobile phone and processes data, thereby monitoring the mobile phone as a whole. Optionally, the processor 180 may include one or more processing units; optionally, the processor 180 may include, for example, an application processor (application processor, AP), a modem processor, a graphics processor (graphics processing unit, GPU), an image signal processor (image signal processor, ISP), a controller, a video codec, a digital signal processor (digital signal processor, DSP), a baseband processor, and / or a neural-network processing unit (neural-network processing unit, NPU), etc. Among them, different processing units can be independent devices or integrated in one or more processors.
[0044] The mobile phone 100 also includes a power source 190 (such as a battery) for supplying power to various components. Preferably, the power source can be logically connected to the processor 180 through a power management system, so that the power management system can manage functions such as charging, discharging, and power consumption.
[0045] The mobile phone 100 may further include a camera. Optionally, the camera may be located at the front or rear of the mobile phone, which is not limited in the present embodiment.
[0046] Figure 2is one of the flow charts of the method for removing a portrait provided in an embodiment of the present application, Figure 2 The illustrated portrait removal method can be applied to the electronic device described above, and the method includes step 210, step 220, step 230, step 240, and step 250. The method flow steps are only a possible implementation of the present application.
[0047] Step 210: In response to a designated operation on the video, determine designated facial features of a target portrait designated by the designated operation from the video.
[0048] The video here refers to the video that needs to be removed from the portrait or the part of the video that needs to be removed from the portrait, such as a certain episode, a certain issue, or a certain season in a whole video collection. The video is composed of a series of continuous static images (frames), which are played at a certain rate to create a dynamic visual effect for the audience. In the embodiment of the present application, the static images that make up the video are recorded as video frames, that is, the video may include multiple continuous video frames.
[0049] The user can perform specified operations on the video. The instruction operation here is used to specify the target portrait to be eliminated in the video. The specified operation can be an operation of specifying or selecting a face or portrait as the target portrait on the video. The specified operation can be a user interaction operation in the form of clicking, box selection, gesture recognition, voice recognition, etc. For example, the user can implement the specified operation by clicking on a face on the screen on the video playback interface. For example, the user can select any video frame from the video and click on a portrait in the video frame to implement the specified operation.
[0050] It should be noted that a portrait refers to a description of a person in a video frame, and may include, for example, a full-body portrait, a half-body portrait, or other portraits. Some portraits may include the person's limbs, body, and face, some portraits may only include part of the limbs, body, and face, some portraits may only include the face, and some portraits may be portraits with the back turned. That is, a portrait may include the face or not (e.g., a portrait with the back turned), and the target portrait specified by the specified operation is usually a portrait including the face.
[0051] Accordingly, after receiving the designated operation, the electronic device can respond to the designated operation. The specific response method can be to obtain the designated position corresponding to the designated operation, and then determine the portrait at the position specified by the designated operation from the video as the target portrait, and extract the facial features of the target portrait as the designated facial features. For example, the coincidence relationship between the position specified by the designated operation and the area where each portrait in the video frame is located can be determined from the video frame where the designated operation occurs in the video, and then the portrait that coincides with the position corresponding to the designated operation is determined from the portraits in the video frame, and the portrait is used as the target portrait, that is, the portrait that the user specifies through the designated operation and needs to be eliminated. Subsequently, the facial features of the target portrait can be extracted from the area where the target portrait is located in the video frame, that is, the designated facial features are obtained. It can be understood that the designated facial features can be used as the identification of the target portrait.
[0052] Step 220: Determine a target video frame from the video frames based on the similarity between the facial features of the portrait contained in the video frames in the video and the designated facial features.
[0053] Specifically, after determining the specified facial features, portrait detection can be performed on each video frame in the video to obtain the portraits contained in each video frame, and facial features can be extracted for the portraits contained in each video frame to obtain the facial features of the portraits contained in each video frame.
[0054] On this basis, the similarity between the facial features of the portraits contained in each video frame and the designated facial features can be calculated. It can be understood that for any portrait in any video frame, the similarity between the facial features of the portrait and the designated facial features is used to reflect the probability that the portrait is the target portrait to be eliminated, that is, the higher the similarity, the more likely the portrait is the target portrait to be eliminated, and the lower the similarity, the less likely the portrait is the target portrait to be eliminated.
[0055] After obtaining the similarity between the facial features of the portraits respectively contained in each video frame and the designated facial features, the target video frame can be determined from each video frame based on this. Here, the target video frame is the video frame that needs to remove the portrait, that is, the video frame with the target portrait or target face. For example, for any video frame, if there is a portrait in the video frame whose similarity with the designated facial features exceeds the similarity threshold, it is determined that the target portrait exists in the video frame, and the video frame is used as the target video frame; if there is no portrait in the video frame whose similarity with the designated facial features exceeds the similarity threshold, it is determined that there is no target portrait in the video frame, and the video frame is not used as the target video frame.
[0056] Therefore, the target video frames that need to be removed from the portrait can be screened out from the video frames in the video according to the similarity between the facial features.
[0057] Optionally, before calculating the similarity between facial features, the face of the portrait in the video frame and the face of the target portrait can be aligned. That is, the position, angle, size, etc. of the face can be standardized with the help of the key points of the face. After completing the face alignment, the facial features of the portrait in the video frame after the face alignment can be extracted. Applying the facial features thus extracted to the similarity calculation between facial features can effectively improve the accuracy and robustness of the similarity calculation.
[0058] Step 230: obtain the mask of the target portrait in the target video frame, input the target video frame and the mask of the target portrait into the image restoration model, and obtain a restored image output by the image restoration model, wherein the mask of the target portrait is used to specify the area to be restored in the target video frame, and the restored image no longer includes the target portrait.
[0059] Specifically, after determining the target video frame, the mask of the target portrait in the target video frame can be obtained. Here, for the target video frame, the mask of the target portrait is the area occupied by the target portrait in the target video frame, that is, the image area in the target video frame where the portrait needs to be removed. The mask of the target portrait can be obtained by performing image segmentation on the target video frame, for example, by obtaining the mask of the target portrait through methods such as global threshold, edge detection, and region growing.
[0060] After obtaining the mask of the target portrait in the target video frame, the target video frame and the mask of the target portrait can be input into the image restoration model together. Here, the image restoration model is used to restore the image, and the mask of the target portrait is used to specify the restoration area. In an embodiment of the present application, for the case where the target video frame and the mask of the target portrait are input into the image restoration model together, the target video frame can be regarded as an image that needs to be restored, and the mask of the target portrait can be regarded as the area that needs to be restored in the target video frame.
[0061] Here, there may be a variety of ways to perform image restoration based on the image restoration model, which may be traditional implementation methods, such as a diffusion-based image restoration method, a texture synthesis-based image restoration method, a sparse representation-based image restoration method, etc., or may be a deep learning-based image restoration method, such as a codec-based image restoration method, a GAN (Generative Adversarial Networks)-based image restoration method, an RNN (Recurrent Neural Networks)-based image restoration method, etc.
[0062] Considering that in the actual application scenario of portrait removal, the main area of the target portrait to be removed usually accounts for a large proportion of the video frame, in the embodiment of the present application, a large model based on LaMa (large mask inpainting) can be used as an image restoration model to achieve image restoration. In the LaMa-based image restoration method, LaMa can use FFCs (fast Fourier convolutions) to increase the receptive field, and can use the occlusion area mask generation strategy during the training process, so it can perform high-quality restoration for situations where the area is large.
[0063] After the target video frame and the mask of the target portrait are input into the image restoration model, the image restoration model can perform image restoration on the target video frame. Here, performing image restoration on the target video frame can specifically be a process of treating the mask of the target portrait in the target video frame as the missing area in the target video frame and filling the missing area. The output of the image restoration model obtained in this way is the restored image, which is an image obtained by eliminating the target portrait in the target video frame and generating and filling the area where the eliminated target portrait is located. The restored image no longer contains the target portrait.
[0064] Step 240: extracting a restored area from the restored image based on the mask of the target portrait, and extracting the target portrait from the target video frame based on the mask of the target portrait to obtain an image to be synthesized.
[0065] Specifically, in the obtained repaired image, the repaired area can be cut out from the repaired image based on the mask of the target portrait. The repaired area here is the area in the repaired image that replaces the target portrait in the target video frame, which can be understood as the image area in the repaired image that is filled with the area where the target portrait is actually located.
[0066] In addition, the target person portrait can be removed from the target video frame based on the mask of the target person portrait to obtain the image to be synthesized. Here, the image to be synthesized is the target video frame without the target person portrait, and the area where the target person portrait originally existed in the image to be synthesized is reflected as a blank area.
[0067] Step 250: superimpose the repaired area on the image to be synthesized to obtain a repaired video frame after removing the target portrait from the target video frame.
[0068] Specifically, after respectively obtaining the repaired area and the image to be synthesized, the repaired area can be superimposed on the image to be synthesized, that is, the blank area in the image to be synthesized caused by the missing target portrait is covered by the repaired area, and the superposition method here can be bitmap superposition. In an embodiment of the present application, the image obtained by superimposing the repaired area on the image to be synthesized is recorded as a repaired video frame. It can be understood that the repaired video frame is a video frame obtained by removing the target portrait from the target video frame. Here, removing the target portrait includes cutting out the target portrait from the target video frame and superimposing the repaired area to the original position of the target portrait.
[0069] It should be noted that the repaired video frame obtained here is different from the repaired image obtained in step 230. Specifically, the repaired image is directly output by the image repair model. In the process of performing image repair on the target video frame based on the image repair model, the image repair model may also make slight adjustments to the target video frame outside the mask of the target portrait. The repaired image obtained in this way may have slight differences from the target video frame in the area outside the mask of the target portrait. For example, some areas outside the mask are processed at the pixel level and are difficult to find visually, but if it is enlarged to the pixel level, it can be clearly found that there is a difference. The repaired video frame is to superimpose the repaired area on the image to be synthesized. Since the image to be synthesized only removes the target portrait from the target video frame and does not adjust the area outside the target portrait, the area outside the mask of the repaired video frame is completely consistent with the target video frame. That is, compared with the repaired image, the repaired video frame completely retains the original appearance of the area outside the mask of the target portrait in the target video frame, and only eliminates and repairs the image at the mask of the target portrait.
[0070] For example, Figure 3 Schematic diagram of the process of the image restoration method provided in the embodiment of the present application. Figure 3As shown, the target video frame that needs to be removed from the portrait and the mask of the target portrait in the target video frame can be first obtained. Subsequently, the target video frame and the mask of the target portrait can be pre-processed respectively, and the pre-processing here can include a resize operation performed to adapt to the input of the image restoration model to achieve size adjustment. Then, the target video frame and the mask of the target portrait after pre-processing can be input into the image restoration model together, thereby obtaining the restored image output by the image restoration model. Subsequently, the restored image can be post-processed, and the post-processing here can include a reverse resize operation, so as to restore the restored image to a form consistent with the original target video frame size. Then, based on the mask of the target portrait, the restored area can be extracted from the restored image. In addition, based on the mask of the target portrait, the target portrait can be removed from the target video frame to obtain the image to be synthesized. Subsequently, the restored area can be superimposed on the image to be synthesized, thereby obtaining a restored video frame with the target portrait eliminated.
[0071] It is understood that the target portrait specified from a video or a portion of a video may be one or more, and the corresponding designated facial features may be one or more. Similarly, when there are multiple target portraits and designated facial features, the facial features of the portrait in each video frame may be compared with the multiple designated facial features to find out whether the target portrait exists, or whether there are multiple target portraits.
[0072] In the embodiment of the present application, based on the similarity between the facial features of the faces contained in each video frame in the video and the specified facial features, the target video frame is determined and the portrait is removed. In this process, the user only needs to specify the target portrait, and the determination of the target video frame and the portrait removal for the target video frame can be automatically realized, which greatly improves the efficiency of video-level portrait removal. In addition, in the process of portrait removal, the target video frame is repaired based on the image repair model to obtain a repaired image, and the image to be synthesized from the target video frame and the repaired image obtained by cutting out the target portrait from the repaired image are superimposed, so that the portrait removal of each target video frame can be achieved, further ensuring the reliability and quality of portrait removal.
[0073] It should be noted that each implementation method of the present application can be freely combined, the order can be changed, or it can be executed separately, and does not need to rely on or depend on a fixed execution order.
[0074] In some embodiments, in step 210, determining the specified facial features of the target portrait specified by the specified operation from the video includes: Performing image segmentation on the video frames associated with the specified operation in the video to obtain masks of one or more first portraits in the associated video frames; Determine one or more first portrait detection frames according to the one or more first portrait masks; Determine a first portrait detection frame corresponding to the specified operation, and perform face detection on a region of interest sub-image corresponding to the first portrait detection frame corresponding to the specified operation; In response to detecting a designated face in the region of interest sub-image, features of the designated face are extracted as the designated face features, and a first portrait corresponding to a first portrait detection frame where the designated operation occurs is the target portrait.
[0075] Specifically, in response to the specified operation, a video frame associated with the specified operation can be determined from the video frames. Here, the video frame associated with the specified operation, that is, the video frame where the specified operation occurs, can specifically be the video frame pointed to when the specified operation occurs, such as the video frame currently displayed on the screen or one of the multiple video frames currently displayed.
[0076] After determining the video frames associated with the specified operation, image segmentation can be performed on the associated video frames. Image segmentation is a technology that separates the target in the image from the background. Specifically, in the embodiment of the present application, it can be used to separate the first portrait as the target from the associated video frame. The result of the image segmentation obtained in this way can be a mask of each first portrait in the associated video frame in the associated video frame, and the mask of the first portrait indicates the position and shape of the first portrait in the associated video frame. The image segmentation here can be achieved through an image segmentation model, for example, it can be achieved through image segmentation models with structures such as FCN (Fully Convolutional Networks), U-Net, DeepLab, Mask R-CNN (Region-based Convolutional Neural Network), PSPNet (PyramidScene Parsing Network), SegNet, etc.
[0077] After completing the portrait segmentation, the masks of the first portraits in the associated video frames can be obtained, so that the first portrait detection frame can be determined based on the masks of the first portraits. For example, the minimum circumscribed rectangle of the masks of the first portraits can be matched to obtain the first portrait detection frame. Then, according to the location where the specified operation occurs and the location of the first portrait detection frame, the first portrait detection frame corresponding to the specified operation is obtained.
[0078] After determining the first portrait detection frame corresponding to the specified operation, face detection may be performed on the sub-image of interest corresponding to the first portrait detection frame.
[0079] Here, the sub-image of interest corresponding to the first portrait detection frame is the regional image inside the first portrait detection frame, and the sub-image of interest corresponding to the first portrait detection frame reflects the regional image of the first portrait corresponding to the first portrait detection frame. After determining the sub-image of interest, face detection can be performed on the sub-image of interest. Here, face detection is a technology for identifying and locating faces in an image. For example, face detection models such as MTCNN (Multi-task Convolutional Neural Network), RetinaFace, and CenterFace can be used to implement face detection for the sub-image of interest. The result of face detection can be whether the sub-image of interest contains a face, and if the sub-image of interest contains a face, it can also include the position of the face in the sub-image of interest. Here, the position of the face in the sub-image of interest can be shown in the form of a face detection frame (box frame). In addition, the result of face detection can also include the position of key points of the face, and the key points of the face here can be key points of organs such as eyes, nose, and mouth.
[0080] When it is determined through face detection that a face is included in the sub-graph of interest, that is, when a specified face is detected in the region of interest, the features of the specified face can be extracted from the location of the specified face in the sub-graph of interest as the specified face features. Here, the extraction of face features can be achieved by extracting face features of the face detection frame region in the sub-graph of interest through a deep learning feature extraction backbone network. For example, the deep learning feature extraction backbone network can be ResNet (Residual Networks), MobileNet (Mobile Network), ShuffleNet, EfficientNet, GhostNet, GoogLeNet, DenseNet (Dense Convolutional Network), etc., or the key point data output by the face detection model can be directly used as the face features.
[0081] In addition, when it is determined through face detection that a face is included in the sub-image of interest, it is also necessary to determine the overlap relationship between the mask of one or more first portraits and the first portrait detection frame where the specified face is located, based on the mask of one or more first portraits in the associated video frame and the first portrait detection frame where the specified face is located. The overlap relationship here can be expressed as whether it overlaps or not, and can also be further expressed as the overlap area, for example, it can be the number of pixels that overlap. It can be understood that each mask of the first portrait has an overlap relationship with the first portrait detection frame where the specified face is located. Assuming that the area intersection and union ratio is used as a representation of the overlap relationship, the area intersection and union ratio can range from 0 to 1.
[0082] It should be noted that the first portrait only refers to the definition used in the process of specifying the target portrait, and the first portrait is the portrait in the associated video frame in which the specified operation occurs during the process of specifying the target portrait. Portraits in other video frames (such as non-associated video frames) or portraits extracted from associated video frames in subsequent other processing can be referred to as second portraits in this article.
[0083] After obtaining the overlap relationship between the mask of one or more first portraits and the first portrait detection frame where the specified face is located, the target portrait can be determined from the one or more first portraits based on this. Here, the first portrait that completely overlaps with the first portrait detection frame can be used as the target portrait, or the first portrait with the largest overlap area can be used as the target portrait, and this embodiment of the application does not specifically limit this.
[0084] In some embodiments, Figure 4 Schematic diagram of the process of specifying a target portrait provided by the embodiment of the present application. Figure 4 As shown, step 210 can be implemented by the following steps: First, the user can arbitrarily select a video frame from the video.
[0085] Secondly, the user can click on the video frame to perform a specified operation. Accordingly, the electronic device obtains the clicked position in response to the specified operation.
[0086] Next, the electronic device can perform image segmentation based on the image segmentation model for the video frame selected by the user, that is, the video frame associated with the specified operation, thereby obtaining each subject and the category of each subject in the video frame, and thereby obtaining one or more first portraits in the video frame according to the category of the subject.
[0087] Subsequently, the electronic device can determine the first portrait at the click position from one or more first portraits in the video frame, further determine the first portrait detection frame of the clicked first portrait, and perform face detection on the sub-image of interest in the first portrait detection frame to determine whether a face can be detected.
[0088] If a face can be detected, feature extraction can be performed on the detected face to obtain the specified facial features of the target portrait and the corresponding target portrait, thereby completing the target portrait specification process.
[0089] If no face is detected, a prompt may be issued and it may be determined whether the user continues to select a portrait in the current video frame. If so, the specified operation may continue to be obtained in the current video frame; otherwise, the video frame selected by the user may be obtained again.
[0090] In some embodiments, in step 220, based on the similarity between the facial features of the portrait contained in the video frame in the video and the specified facial features, determining the target video frame from the video frame includes: Obtaining candidate facial features of each of the video frames; If it is determined that the similarity between the candidate facial feature and the designated facial feature in the video frame is greater than a preset threshold, the video frame is used as the target video frame.
[0091] Specifically, in the process of determining the target video frame from each video frame of the video, for any video frame in the video frame, the similarity between the corresponding candidate facial features of the second portrait contained in the video frame and the specified facial features can be calculated to determine whether to use the video frame as the target video frame.
[0092] Here, for any video frame, the portrait contained in the video frame can be recorded as the second portrait, and the candidate facial features of each second portrait are extracted, and then the candidate facial features of each second portrait are respectively calculated with the specified facial features for similarity, thereby obtaining the similarity between the candidate facial features of each second portrait and the specified facial features. The similarity here can be cosine similarity, Euclidean distance, etc., which is not specifically limited in the embodiment of the present application. After obtaining the similarity between the candidate facial features of each second portrait and the specified facial features, each similarity can be compared with a preset threshold. The preset threshold here is to determine the minimum value of the similarity when two faces belong to the same portrait. Therefore, when there is a similarity greater than the preset threshold, it can be determined that there is a face of the target portrait in the video frame, that is, the video frame is determined to be a target video frame; otherwise, it is determined that the video frame is not a target video frame.
[0093] For example, referring to the aforementioned content of this article, the target portrait determined is Y, and the corresponding designated facial feature is aa. Three candidate facial features {aa, ab, ac} are extracted from the second portrait I, the second portrait II and the second portrait III of a video frame S1 respectively. It can be seen that the candidate facial features of the second portrait I have the highest similarity with the designated facial features and are greater than the preset threshold. Then the video frame S1 is the target video frame, and the second portrait I is the target portrait Y.
[0094] In some embodiments, in step 230, obtaining a mask of the target person portrait in the target video frame includes: Acquire a candidate face detection frame corresponding to the candidate face feature whose similarity is greater than a preset threshold as a target face detection frame; Performing image segmentation on the target video frame to obtain masks of one or more second portraits in the target video frame; Based on the overlap relationship between the masks of the one or more second portraits and the target face detection frame, the mask of the target portrait is determined from the masks of the one or more second portraits.
[0095] Specifically, for the target video frame determined from the video, a candidate face detection frame of the candidate face features of the second portrait whose similarity is greater than a preset threshold in the target video frame can be obtained as the target face detection frame. Then, the target video frame can also be segmented to obtain a mask of each second portrait in the target video frame, that is, to obtain the regional position of each second portrait in the target video frame.
[0096] Subsequently, the overlap relationship between one or more second portrait masks and the target face detection frame in the target video frame can be determined. The overlap relationship here can be expressed as whether it overlaps or not, and can also be further expressed as the overlap area, for example, it can be the number of pixels that overlap. It can be understood that the mask of each second portrait needs to be judged to overlap with the target face detection frame, so that the mask of each second portrait has an overlap relationship with the target face detection frame where the specified face is located. After obtaining the overlap relationship between the mask of each second portrait and the target face detection frame where the facial features of the target portrait are located, the target portrait in the target video frame can be determined from each second portrait based on this. Here, the second portrait with the highest overlap area can be used as the target portrait, and this embodiment of the present application does not specifically limit this.
[0097] For example, Figure 5 is a schematic diagram of a mask of a second portrait and a target face detection frame provided in an embodiment of the present application. Figure 5 As shown, in video frame 50, Figure 5 The frame marked in red is the target face detection frame 51 determined according to the similarity with the specified facial features as mentioned above in this article. The target face detection frame 51 identifies a rectangular area of the face part of the second portrait. Figure 5 The mask 52 of the second person portrait is marked with a black frame. Figure 5 The second portrait detection frame 53 is marked with a blue frame. For the same second portrait, the difference between the second portrait mask 52 and the second portrait detection frame 53 is that the second portrait mask 52 describes the main shape and position of the portrait, which is usually irregular, and the second portrait mask 52 coincides with the edge of the portrait, reflecting the position and shape of the portrait, while the second portrait detection frame 53 is a rectangular frame circumscribed by the second portrait mask 52.
[0098] Therefore, when calculating the overlap relationship between the mask 52 of the second portrait and the target face detection frame 51, the number of pixels overlapped by the mask 52 of the second portrait and the target face detection frame 53 can be calculated. It can be understood that the higher the number of overlapping pixels, the higher the probability that the mask 52 of the second portrait and the target face detection frame 53 represent the same portrait.
[0099] In some embodiments, the method further comprises: Taking the target video frame as a starting point, performing at least one of forward video tracking and backward video tracking on the video to obtain a video frame in which the target portrait still exists in the video, and using the video frame in which the target portrait still exists as a tracking video frame; The tracking video frame is used as a new target video frame.
[0100] Specifically, the target video frames determined based on the similarity between facial features can reflect the part of the video frames that need to remove the portrait in the video, but there are still some video frames that cannot be identified based on the similarity between facial features because the portraits contained in them are seriously sideways or facing away. For such video frames, video tracking can be performed based on the detected target video frames, thereby detecting such video frames with target portraits.
[0101] In a specific implementation, any target video frame in the video can be used as a starting point to perform forward video tracking on the video, that is, the target portrait in the target video frame is used as the tracking target, and the target video frame is used as the starting point for tracking, and the target portrait is tracked to the video frame arranged before the target video frame in the video to determine whether there is still a target portrait in the video frame before the target video frame. It can be understood that for any video frame, when the video frame is not a target video frame, that is, the target portrait cannot be identified in the video frame based on the similarity between facial features, and the target portrait is identified based on video tracking after the target portrait recognition based on the similarity between facial features fails.
[0102] In addition, any target video frame in the video can be used as a starting point to perform backward video tracking on the video, that is, the target portrait in the target video frame is used as the tracking target, and the target video frame is used as the starting point for tracking. The target portrait is tracked to the video frames arranged after the target video frame in the video to determine whether the target portrait still exists in the video frames after the target video frame.
[0103] In the embodiment of the present application, a video frame containing a target person portrait obtained based on at least one of forward video tracking and backward video tracking can be recorded as a tracking video frame. It can be understood that a tracking video frame is a video frame containing a target person portrait obtained based on video tracking. Thus, the tracking video frame can be used as a target video frame, thereby achieving the elimination of the target person portrait for the tracking video frame.
[0104] In an embodiment of the present application, a tracking video frame in which a target portrait exists is determined by at least one of forward video tracking and backward video tracking, so that when a video frame containing a target portrait is determined from a video frame, video frames that cannot be identified based on face detection due to the portrait's severe side profile or back facing away can be included, thereby ensuring the comprehensiveness and reliability of eliminating the target portrait from the video.
[0105] In some embodiments, Figure 6 Schematic diagram of the method flow of forward video tracking provided by the embodiment of the present application. Figure 6 As shown, taking the target video frame as the starting point, forward video tracking is performed on the video to obtain a video frame in which the target portrait is still present in the video, and the video frame in which the target portrait is still present is used as the tracking video frame, including: Taking any target video frame as a current frame, in response to a video frame preceding the current frame being a non-target video frame, judging whether the previous video frame contains the target human portrait based on the overlap relationship between the masks of the target human portrait in the current frame and the masks of the human portraits in the previous video frame; If it is determined that the previous video frame contains the target portrait, the previous video frame is used as the tracking video frame, and in response to the previous video frame being a non-transition frame, the previous video frame is used as the new current frame for the forward video tracking.
[0106] Specifically, forward video tracking starting from the target video frame can be achieved through the following steps: First, any target video frame is used as the current frame in the forward video tracking process. The current frame here refers to the currently tracked video frame.
[0107] Secondly, it is determined whether the previous video frame of the current frame is the target video frame, that is, it is determined whether there is a target portrait in the previous video frame that can be detected based on the similarity between facial features. In the case where the previous video frame is the target video frame, that is, it has been determined that the previous video frame contains the target portrait based on the similarity between facial features, the forward video tracking process can be directly terminated. In addition, in the case where there is no previous video frame, that is, the current frame is the first video frame in the video, the forward video tracking process can also be directly terminated.
[0108] In the case where the previous video frame is not the target video frame, that is, there is no target portrait in the previous video frame that can be detected based on the similarity between facial features, image segmentation can be performed on the previous video frame to obtain the masks of all portraits in the previous video frame, and then the overlap relationship between the mask of the target portrait in the current frame and the mask of each portrait in the previous video frame is calculated. The overlap relationship here can be expressed as IOU (Intersection over Union), which is specifically the ratio of the intersection and union of the areas of two portraits. For example, the calculation formula for the intersection and union of portrait mask A and portrait mask B can be expressed as follows: After calculating the overlap relationship between the mask of the target portrait in the current frame and the mask of each portrait in the previous video frame, it can be determined whether the previous video frame contains the target portrait based on the overlap relationship. The specific judgment method can be to determine whether the maximum IOU value is greater than the IOU threshold. If the maximum IOU value is greater than the IOU threshold, it can be determined that the portrait in the previous video frame corresponding to the maximum IOU value is the target portrait; if the maximum IOU value is less than or equal to the IOU threshold, it is determined that the target portrait does not appear in the previous video frame, and the forward video tracking process can be terminated at this time. As an example, the IOU threshold here can be set in the interval [0.8, 1).
[0109] In the case where it is determined based on the coincidence relationship that the target portrait exists in the previous video frame, the previous video frame can be used as a tracking video frame.
[0110] In addition, for the situation that the target portrait exists in the previous video frame based on the coincidence relationship, it is also necessary to determine whether the previous video frame is a transition frame in the video. The transition frame here refers to a video frame that plays a transition role in the transition process from one scene to another in the video, specifically refers to the first frame in the rear lens video segment of two adjacent lens video segments. It can be understood that if the previous video frame is a transition frame, it means that the video has transitioned at the previous video frame, and the previous video frame is the first frame of the lens after the transition. The lens before the previous video frame is discontinuous with the previous video frame, and the forward video tracking process can be terminated at this time. If the previous video frame is not a transition frame, it is considered that the lens before the previous video frame is continuous with the previous video frame, and the previous video frame can be used as a new current frame to perform forward video tracking.
[0111] In some embodiments, the method further comprises: It is determined that the previous video frame does not contain the target portrait, or it is determined that the previous video frame contains the target portrait and the previous video frame is a transition frame, and the forward video tracking ends.
[0112] Specifically, during the forward video tracking process, if it is determined that the previous video frame does not contain the target portrait based on the overlap relationship between the mask of the target portrait in the current frame and the mask of each portrait in the previous video frame, the forward video tracking process can be terminated.
[0113] In addition, if it is determined that the previous video frame does not contain the target portrait based on the overlapping relationship between the masks of the target portrait in the current frame and the masks of each portrait in the previous video frame, but the previous video frame is a transition frame, it means that the shot before the previous video frame is discontinuous with the previous video frame, and it is meaningless to perform video tracking from the previous video frame forward, so the forward video tracking process can be terminated.
[0114] In some embodiments, the transition frame may be determined based on the following steps: Based on the difference between two adjacent video frames in the video, it is determined whether the latter video frame of the two adjacent video frames is a transition frame.
[0115] Specifically, when the difference between two adjacent video frames exceeds a difference threshold, it can be determined that the latter video frame is a transition frame; otherwise, it is determined that the latter video frame is not a transition frame.
[0116] In an embodiment of the present application, the difference between two adjacent video frames can be expressed as the difference between the two video frames in each color channel. For example, HSV decomposition can be performed on the two video frames respectively. Among them, H represents hue, S represents saturation, and V represents value. After completing the HSV decomposition, the differences between the two video frames in the H, S, and V channels can be compared respectively, and then the differences in each channel combined with the weights of each channel can be used to calculate the weighted average of the H, S, and V channels to measure the difference between the two video frames. The specific calculation method can be expressed as the following steps: First, two adjacent video frames in the video are taken, namely, video frame a and video frame b, where video frame b is the next frame of video frame a.
[0117] Secondly, HSV decomposition is performed on video frames a and b respectively to obtain H, S, and V single-channel images of video frames a and b.
[0118] Next, the differences between video frames a and b in H, S, and V channels are calculated using the following formulas, which are respectively , , : Where w is the number of columns of pixels in the video frame, h is the number of rows of pixels in the video frame, Represents the pixel value of a single-channel video frame a at column i and row j, represents the pixel value of the single-channel video frame b at column i and row j. It can be understood that and When representing the pixel value on the H channel, the calculated d is ;exist and When representing the pixel value on the S channel, the calculated d is ;exist and When representing the pixel value on the V channel, the calculated d is .
[0119] Then, the weight difference score of video frames a and b can be calculated by the following formula: : In the formula, , , They are the weights of the H, S, and V channels, respectively, and their value range is .
[0120] Finally, we can judge Is it greater than a difference threshold (for example, the difference threshold can be set to A constant in the interval), if it is greater than the threshold, it is considered that there is a transition between video frames a and b, and video frame b is used as the transition frame. Otherwise, it is considered that there is no transition between video frames a and b.
[0121] In some embodiments, Figure 7 Schematic diagram of the backward video tracking method provided by the embodiment of the present application. Figure 7 As shown, taking the target video frame as the starting point, performing backward video tracking on the video, obtaining a video frame in which the target portrait still exists in the video, and taking the video frame in which the target portrait still exists as the tracking video frame, includes: Taking any target video frame as the current frame, in response to a subsequent video frame of the current frame being a non-target video frame and a non-transition frame, judging whether the subsequent video frame contains the target portrait based on the overlap relationship between the masks of the target portrait in the current frame and the masks of each portrait in the subsequent video frame; If it is determined that the target portrait is included in the subsequent video frame, the subsequent video frame is used as the tracking video frame, and the subsequent video frame is used as the new current frame for the backward video tracking.
[0122] Specifically, backward video tracking starting from the target video frame can be achieved through the following steps: First, any target video frame is used as the current frame in the backward video tracking process. The current frame here refers to the currently tracked video frame.
[0123] Secondly, it is determined whether the next video frame of the current frame is the target video frame and whether the next video frame is a transition frame, that is, it is determined whether there is a target portrait that can be detected based on the similarity between facial features in the next video frame, and whether the next video frame is a transition frame in the video. For the case where the next video frame is the target video frame, that is, the case where the target portrait has been determined to be included in the next video frame based on the similarity between facial features, the backward video tracking process can be directly ended. For the case where the next video frame is a transition frame in the video, that is, the video has a transition at the next video frame, and the next video frame is the first frame of the lens after the transition, the current frame and the next video frame are not continuous on the lens, and the next video frame cannot be tracked based on the current frame at this time, and the backward video tracking process can be directly ended. In addition, for the case where there is no next video frame, that is, the current frame is the last video frame in the video, the backward video tracking process can also be directly ended.
[0124] If the next video frame is not the target video frame and is not a transition frame, the next video frame can be segmented to obtain all the portraits in the next video frame, and then the overlap relationship between the mask of the target portrait in the current frame and the mask of each portrait in the next video frame is calculated. The overlap relationship here can also be expressed as IOU.
[0125] After calculating the overlap relationship between the mask of the target portrait in the current frame and the mask of each portrait in the next video frame, it is possible to determine whether the next video frame contains the target portrait based on the overlap relationship. The specific judgment method can be to determine whether the maximum IOU value is greater than the IOU threshold. If the maximum IOU value is greater than the IOU threshold, it can be determined that the portrait in the next video frame corresponding to the maximum IOU value is the target portrait; if the maximum IOU value is less than or equal to the IOU threshold, it is determined that the target portrait does not appear in the next video frame, and the backward video tracking process can be ended at this time.
[0126] In the case where it is determined based on the coincidence relationship that there is a target portrait in the subsequent video frame, the subsequent video frame can be used as a tracking video frame, and the subsequent video frame can be used as a new current frame to perform backward video tracking.
[0127] In some embodiments, the method further comprises: It is determined that a subsequent video frame of the current frame is a transition frame, or it is determined that a subsequent video frame of the current frame does not contain the target portrait, and the backward video tracking ends.
[0128] Specifically, during the backward video tracking process, if the next video frame is a transition frame, that is, the video transitions at the next video frame, and the next video frame is the first frame of the shot after the transition, the current frame and the next video frame are not continuous on the shot. At this time, the next video frame cannot be tracked based on the current frame, and the backward video tracking process can be terminated directly. In addition, in the case where it is determined based on the overlap relationship that the target person portrait does not exist in the next video frame, the backward video tracking process can also be directly terminated.
[0129] In some embodiments, the method further comprises: The target video frame in the video is replaced with the repaired video frame.
[0130] Specifically, after removing the portrait from the target video frame, that is, after obtaining the repaired video frame of the target video frame, the repaired video frame can be used to replace the original target video frame in the video. As a result, there is no longer any video frame containing the target portrait in the video, and all target portraits in the video have been removed.
[0131] Specifically, in the process of replacing the target video frame with the repaired video frame, the audio and video frame sequences can be first extracted from the video, and the width, height and frame rate of the video frame can be obtained. Subsequently, in the video frame sequence, according to the frame index of the target video frame detected, the target video frame corresponding to each frame index can be replaced with the repaired video frame after the target portrait is eliminated, thereby obtaining a video frame sequence after the replacement. Then, based on the video frame sequence after the replacement, a video with the same width, height and frame rate as the original video can be synthesized, and the previously extracted audio can be added to the newly synthesized video, thereby obtaining a video after the target portrait is eliminated.
[0132] In some embodiments, Figure 8 FIG. 2 is a flow chart of the method for removing a person's portrait provided in the embodiment of the present application. Figure 8 As shown, a method for removing a portrait comprises the following steps: First, the user can specify the target portrait to be removed in the video through a specified operation. Accordingly, the electronic device can determine the target portrait specified by the specified operation from the video frame and extract the specified facial features of the target portrait.
[0133] Secondly, the electronic device can read the video and perform the following processing frame by frame for all video frames in the video.
[0134] Next, for each video frame in the video, the face contained in the video frame can be identified based on the face detection model, and the similarity between the facial features of the identified face and the specified facial features can be calculated, thereby achieving face matching, and then determining whether the video frame is a target video frame that needs to remove the portrait, and obtaining the mask of the target portrait in the target video frame.
[0135] Then, for each video frame in the video, it is determined whether the video frame is a transition frame.
[0136] Next, starting from the target video frame, the video frame containing the target portrait is further searched through video tracking and combined with the transition frame, and the video frame thus found is also used as the target video frame.
[0137] Then, portrait removal processing is performed on all target video frames to obtain repaired video frames corresponding to each target video frame.
[0138] Finally, the repaired video frame is used to replace the target video frame in the video, and the portrait removal process is completed.
[0139] In the embodiment of the present application, the target portrait specified by the specified operation can be determined from the video frame based on the face detection model, and the mask of the target portrait can be determined by the image segmentation model, and the portrait area specified by the target portrait mask can be generatively repaired by the image repair model, thereby realizing automatic portrait removal. In this process, the user only needs to perform the specified operation to realize automatic portrait removal at the video level with one click, which greatly improves the efficiency of portrait removal compared to the frame-by-frame manual processing method.
[0140] In addition, in case that facial feature matching cannot be performed to determine the target portrait in the video frame because the portrait is severely sideways or facing away, the video frame containing the target portrait can be determined through video tracking, thereby ensuring the comprehensiveness and reliability of video-level portrait removal.
[0141] Fig. 9 An example of a physical structure diagram of an electronic device is shown in FIG. Fig. 9 As shown, the electronic device may include: a processor 910, a communication interface 920, a memory 930 and a communication bus 940, wherein the processor 910, the communication interface 920 and the memory 930 communicate with each other through the communication bus 940. The processor 910 may call the logic instructions in the memory 930 to execute the portrait removal method, which includes: In response to a designated operation on the video, determining designated facial features of a target portrait designated by the designated operation from the video; Determining a target video frame from the video frames based on the similarity between the facial features of a person included in the video frames in the video and the designated facial features; Obtaining a mask of the target portrait in the target video frame, inputting the target video frame and the mask of the target portrait into an image restoration model, and obtaining a restoration image output by the image restoration model, wherein the mask of the target portrait is used to specify a to-be-restored area of the target video frame, and the restoration image no longer includes the target portrait; Extracting a repaired area from the repaired image based on the mask of the target portrait, and extracting the target portrait from the target video frame based on the mask of the target portrait to obtain an image to be synthesized; The repaired area is superimposed on the image to be synthesized to obtain a repaired video frame after the target portrait is removed from the target video frame.
[0142] In addition, the logic instructions in the above-mentioned memory 930 can be implemented in the form of a software functional unit and can be stored in a computer-readable storage medium when it is sold or used as an independent product. Based on this understanding, the technical solution of the present application can be essentially or partly embodied in the form of a software product that contributes to the relevant technology. The computer software product is stored in a storage medium, including several instructions to enable a computer device (which can be a personal computer, a server, or a network device, etc.) to perform all or part of the steps of the method described in each embodiment of the present application. The aforementioned storage medium includes: U disk, mobile hard disk, read-only memory (ROM, Read-Only Memory), random access memory (RAM, Random Access Memory), disk or optical disk, etc. Various media that can store program codes.
[0143] On the other hand, the present application also provides a computer program product, the computer program product includes a computer program, the computer program can be stored on a non-transitory computer-readable storage medium, when the computer program is executed by a processor, the computer can execute the portrait removal method provided by the above methods, the method includes: In response to a designated operation on the video, determining designated facial features of a target portrait designated by the designated operation from the video; Determining a target video frame from the video frames based on the similarity between the facial features of a person included in the video frames in the video and the designated facial features; Obtaining a mask of the target portrait in the target video frame, inputting the target video frame and the mask of the target portrait into an image restoration model, and obtaining a restoration image output by the image restoration model, wherein the mask of the target portrait is used to specify a to-be-restored area of the target video frame, and the restoration image no longer includes the target portrait; Extracting a repaired area from the repaired image based on the mask of the target portrait, and extracting the target portrait from the target video frame based on the mask of the target portrait to obtain an image to be synthesized; The repaired area is superimposed on the image to be synthesized to obtain a repaired video frame after the target portrait is removed from the target video frame.
[0144] On the other hand, the present application further provides a non-transitory computer-readable storage medium having a computer program stored thereon, and when the computer program is executed by a processor, the method for removing a human portrait provided by the above methods is implemented, and the method includes: In response to a designated operation on the video, determining designated facial features of a target portrait designated by the designated operation from the video; Determining a target video frame from the video frames based on the similarity between the facial features of a person included in the video frames in the video and the designated facial features; Obtaining a mask of the target portrait in the target video frame, inputting the target video frame and the mask of the target portrait into an image restoration model, and obtaining a restoration image output by the image restoration model, wherein the mask of the target portrait is used to specify a to-be-restored area of the target video frame, and the restoration image no longer includes the target portrait; Extracting a repaired area from the repaired image based on the mask of the target portrait, and extracting the target portrait from the target video frame based on the mask of the target portrait to obtain an image to be synthesized; The repaired area is superimposed on the image to be synthesized to obtain a repaired video frame after the target portrait is removed from the target video frame.
[0145] The device embodiments described above are merely illustrative, wherein the units described as separate components may or may not be physically separated, and the components displayed as units may or may not be physical units, that is, they may be located in one place, or they may be distributed on multiple network units. Some or all of the modules may be selected according to actual needs to achieve the purpose of the scheme of this embodiment. Ordinary technicians in this field can understand and implement it without paying creative labor.
[0146] Through the description of the above implementation methods, those skilled in the art can clearly understand that each implementation method can be implemented by means of software plus a necessary general hardware platform, and of course, can also be implemented by hardware. Based on this understanding, the above technical solution is essentially or the part that contributes to the relevant technology can be embodied in the form of a software product, and the computer software product can be stored in a computer-readable storage medium, such as ROM / RAM, a disk, an optical disk, etc., including a number of instructions for a computer device (which can be a personal computer, a server, or a network device, etc.) to execute the methods described in each embodiment or some parts of the embodiment.
[0147] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present application, rather than to limit it. Although the present application has been described in detail with reference to the aforementioned embodiments, those skilled in the art should understand that they can still modify the technical solutions described in the aforementioned embodiments, or make equivalent replacements for some of the technical features therein. However, these modifications or replacements do not deviate the essence of the corresponding technical solutions from the spirit and scope of the technical solutions of the embodiments of the present application.
Claims
1. A method for removing a human portrait, comprising: In response to a designated operation on the video, determining designated facial features of a target portrait designated by the designated operation from the video; Determining a target video frame from the video frames based on the similarity between the facial features of a person included in the video frames in the video and the designated facial features; Obtaining a mask of the target portrait in the target video frame, inputting the target video frame and the mask of the target portrait into an image restoration model, and obtaining a restoration image output by the image restoration model, wherein the mask of the target portrait is used to specify a to-be-restored area of the target video frame, and the restoration image no longer includes the target portrait; Extracting a repaired area from the repaired image based on the mask of the target portrait, and extracting the target portrait from the target video frame based on the mask of the target portrait to obtain an image to be synthesized; The repaired area is superimposed on the image to be synthesized to obtain a repaired video frame after the target portrait is removed from the target video frame.
2. The method for removing a human portrait according to claim 1, wherein: Determining, from the video, specified facial features of the target portrait specified by the specified operation, includes: Performing image segmentation on the video frames associated with the specified operation in the video to obtain masks of one or more first portraits in the associated video frames; Determine one or more first portrait detection frames according to the one or more first portrait masks; Determine a first portrait detection frame corresponding to the specified operation, and perform face detection on a region of interest sub-image corresponding to the first portrait detection frame corresponding to the specified operation; In response to detecting a designated face in the region of interest sub-image, features of the designated face are extracted as the designated face features, and a first portrait corresponding to a first portrait detection frame where the designated operation occurs is the target portrait.
3. The method for removing a human portrait according to claim 1, wherein: Determining a target video frame from the video frames based on similarity between facial features of a person included in a video frame in the video and the specified facial features includes: Obtaining candidate facial features of each of the video frames; Determining that the video frame has a similarity between the candidate facial feature and the designated facial feature that is greater than a preset threshold, then using the video frame as the target video frame; Wherein, obtaining the mask of the target portrait in the target video frame includes: Acquire a candidate face detection frame corresponding to the candidate face feature whose similarity is greater than a preset threshold as a target face detection frame; perform image segmentation on the target video frame to obtain a mask of one or more second portraits in the target video frame; Based on the overlap relationship between the masks of the one or more second portraits and the target face detection frame, the mask of the target portrait is determined from the masks of the one or more second portraits.
4. The method for removing a human portrait according to claim 1, further comprising: Taking the target video frame as a starting point, performing at least one of forward video tracking and backward video tracking on the video to obtain a video frame in which the target portrait still exists in the video, and using the video frame in which the target portrait still exists as a tracking video frame; The tracking video frame is used as a new target video frame.
5. The method for removing a human portrait according to claim 4, wherein: Taking the target video frame as a starting point, performing forward video tracking on the video to obtain a video frame in which the target portrait is still present in the video, and using the video frame in which the target portrait is still present as a tracking video frame, including: Taking any target video frame as a current frame, in response to a video frame preceding the current frame being a non-target video frame, judging whether the previous video frame contains the target human portrait based on the overlap relationship between the masks of the target human portrait in the current frame and the masks of the human portraits in the previous video frame; If it is determined that the previous video frame contains the target portrait, the previous video frame is used as the tracking video frame, and in response to the previous video frame being a non-transition frame, the previous video frame is used as the new current frame for the forward video tracking.
6. The method for removing a human portrait according to claim 5, further comprising: It is determined that the previous video frame does not contain the target portrait, or it is determined that the previous video frame contains the target portrait and the previous video frame is a transition frame, and the forward video tracking ends.
7. The method for removing a human portrait according to claim 4, wherein: Taking the target video frame as a starting point, performing backward video tracking on the video to obtain a video frame in which the target portrait still exists in the video, and using the video frame in which the target portrait still exists as a tracking video frame, including: Taking any target video frame as the current frame, in response to a subsequent video frame of the current frame being a non-target video frame and a non-transition frame, judging whether the subsequent video frame contains the target portrait based on the overlap relationship between the masks of the target portrait in the current frame and the masks of each portrait in the subsequent video frame; If it is determined that the target portrait is included in the subsequent video frame, the subsequent video frame is used as the tracking video frame, and the subsequent video frame is used as the new current frame for the backward video tracking.
8. The method for removing a human portrait according to claim 7, further comprising: It is determined that a subsequent video frame of the current frame is a transition frame, or it is determined that a subsequent video frame of the current frame does not contain the target portrait, and the backward video tracking ends.
9. The method for removing a human portrait according to any one of claims 1 to 8, further comprising: The target video frame in the video is replaced with the repaired video frame.
10. An electronic device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor implements the method for removing a human portrait as claimed in any one of claims 1 to 9 when executing the computer program.
11. A non-transitory computer-readable storage medium having a computer program stored thereon, wherein the computer program, when executed by a processor, implements the method for removing a human portrait as claimed in any one of claims 1 to 9.