Face image replacement data processing method and device based on sub-pixels
By using preset detection models in face image processing for refined key point detection and generative adversarial networks to replace face images, the problems of inefficient and poor results of face image processing in traditional technology are solved, and higher positioning accuracy and visual effects are achieved.
Patent Information
- Application Number
- CN202510123161.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-26
- Publication Date
- 2025-05-09
AI Technical Summary
In real-time live broadcast and post-video production, traditional film and television shooting coordination technology lags behind, resulting in inefficient face image processing when continuous frames are processed, and errors are prone to occur, such as blurred facial expressions generated on the side face and unclear boundaries.
The preset detection model performs pixel-level key point position detection and sub-pixel-level key point offset detection, accurately determines the position information of the face image area to be replaced, and uses a generative adversarial network to perform face image replacement processing.
It significantly improves the positioning accuracy of the face replacement area, improves the efficiency, visual effect and realism of face replacement.
Smart Images

Figure CN119963680A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of artificial intelligence technology, and more particularly to a method and device for processing sub-pixel facial image replacement data. Background Art
[0002] At present, with the development of artificial intelligence technology, facial image synthesis technology has gradually matured. However, in practical applications, such as face swapping in real-time live broadcast and post-video production, due to the lag of traditional film and television shooting coordination technology, the processing of continuous frames is inefficient and prone to errors, such as blurred facial expressions and unclear boundaries when generating side faces.
[0003] To address the above problems, no effective solution has been proposed yet. Summary of the invention
[0004] This specification provides a method and device for processing face image replacement data based on sub-pixel, by using the first module of the preset detection model to perform pixel-level key point position detection processing, obtain the first position information of the first key point in the first image frame, and then use the second module of the preset detection model to perform sub-pixel level key point offset detection processing on the first key point to obtain the first offset, and then accurately determine the position information of the first face image area to be replaced in the first image frame according to the first position information and the first offset. This step-by-step refinement from pixel to sub-pixel not only significantly improves the positioning accuracy of the face replacement area, but also significantly improves the efficiency, visual effect and realism of face replacement.
[0005] This specification provides a sub-pixel-based face image replacement data processing method, including:
[0006] Receiving a facial image replacement request for a first video, and in response to the facial image replacement request, determining a first image frame in the first video and a second image frame for performing facial image replacement processing on the first video; wherein the first image frame is an image frame in the first video where a first facial image area to be replaced exists;
[0007] Using a first module of a preset detection model, performing pixel-level key point position detection processing on the first image frame to obtain first position information of a first key point in the first image frame, and using a second module of the preset detection model, performing sub-pixel-level key point offset detection processing on the first key point according to the first position information to obtain a first offset;
[0008] Determine, according to the first position information and the first offset, position information of the first facial image region to be replaced in the first image frame;
[0009] Based on the position information of the first facial image area to be replaced in the first image frame and the second facial image area contained in the second image frame, the first facial image area contained in each of the first image frames is replaced to obtain a replaced first video.
[0010] In one embodiment, the replacing process is performed on the first face image region contained in each of the first image frames according to the position information of the first face image region to be replaced in the first image frame and the second face image region contained in the second image frame to obtain the replaced first video, including:
[0011] Determine the position information of the second facial image region in the second image frame by using the preset detection model;
[0012] Using a preset face image replacement model, according to the position information of the first face image area to be replaced in the first image frame, and the position information of the second face image area in the second image frame, the first face image area contained in each of the first image frames is replaced to obtain a replaced first image frame; wherein the preset face image replacement model is a model constructed according to a preset generative algorithm;
[0013] The replaced first video is determined according to the replaced first image frame.
[0014] In one embodiment, before using the preset facial image replacement model to replace the first facial image region contained in each of the first image frames according to the position information of the first facial image region to be replaced in the first image frame and the position information of the second facial image region in the second image frame to obtain the replaced first image frame, the method further includes:
[0015] Acquire sample data, the sample data including a first face image and a second face image;
[0016] Using the generator of the preset facial image replacement model, the first facial image is replaced by the second facial image to obtain a replaced first facial image;
[0017] Determining whether the preset face image replacement model has converged based on the first face image and the replaced first face image using the discriminator of the preset face image replacement model, and if it is determined that the preset face image replacement model has not converged, continuing to train the preset face image replacement model based on the sample data until the preset face image replacement model has converged;
[0018] The method of using a preset facial image replacement model to replace the first facial image region contained in each of the first image frames according to the position information of the first facial image region to be replaced in the first image frame and the position information of the second facial image region in the second image frame to obtain the replaced first image frame includes:
[0019] Using a generator of a preset facial image replacement model, replacement processing is performed on the first facial image area contained in each of the first image frames according to the position information of the first facial image area to be replaced in the first image frame, and the position information of the second facial image area in the second image frame, so as to obtain the replaced first image frame.
[0020] In one embodiment, determining the position information of the first facial image region to be replaced in the first image frame according to the first position information and the first offset includes:
[0021] Determine a first adjacent key point related to the first key point in the first image frame, and determine second position information of the first adjacent key point in the first image frame according to the first position information;
[0022] Using the second module of the preset detection model, and according to the second position information, performing sub-pixel level key point offset detection processing on the first adjacent key point to obtain a second offset;
[0023] The position information of the first facial image region to be replaced in the first image frame is determined according to the first position information, the second position information, the first offset, and the second offset.
[0024] In one embodiment, determining the first image frame in the first video includes:
[0025] According to the number of image frames contained in a preset time period in the first video and the image quality of the image frames contained in the first video, the first video is segmented into video frames to obtain a plurality of candidate image frames;
[0026] Using a preset face region detection model, detecting and processing the region where the first face image to be replaced is located in the candidate image frame, and determining a first region in each of the candidate image frames that contains the first face image to be replaced;
[0027] Using a preset face detection model, performing face detection processing on the first area to obtain position information of the first face image to be replaced in the first area;
[0028] According to the position information of the first facial image to be replaced in the first area, the pixel value corresponding to the first facial image to be replaced is set as the first pixel value, and the pixel value of the area other than the first facial image to be replaced in the candidate image frame is set as the second pixel value to obtain the first image frame.
[0029] In one embodiment, determining the second image frame for performing face image replacement processing on the first video includes:
[0030] Acquire a second video corresponding to the first video;
[0031] Determine a degree of matching between each image frame in the second video and the first image frame, and determine, based on the degree of matching, a second image frame in the second video corresponding to each of the first image frames.
[0032] In one embodiment, determining the replaced first video according to the replaced first image frame includes:
[0033] Acquire a third image frame and a fourth image frame in the first video that are adjacent to the replaced first image frame;
[0034] According to a preset size, resizing the replaced first image frame, the third image frame, and the fourth image frame respectively to obtain resized image frames;
[0035] Determining an initial transformation matrix and Euler angles according to the adjusted image frame;
[0036] According to the Euler angles and the difference between any two adjacent image frames in the replaced first image frame, the third image frame and the fourth image frame, adjusting the position information of the first facial image area to be replaced in the first image frame, and updating the initial transformation matrix according to the position information obtained by the adjustment process to obtain a target transformation matrix;
[0037] According to the target transformation matrix, the replaced first image frame is aligned to determine the replaced first video.
[0038] This specification provides a sub-pixel-based face image replacement data processing device, comprising:
[0039] A receiving response module, configured to receive a face image replacement request for a first video, and in response to the face image replacement request, determine a first image frame in the first video, and a second image frame for performing face image replacement processing on the first video; wherein the first image frame is any image frame in the first video;
[0040] a first information determination module, configured to perform pixel-level key point position detection processing on the first image frame using a first module of a preset detection model to obtain first position information of a first key point in the first image frame, and to perform sub-pixel-level key point offset detection processing on the first key point according to the first position information using a second module of the preset detection model to obtain a first offset;
[0041] a position information determining module, configured to determine position information of a first facial image region to be replaced in the first image frame according to the first position information and the first offset;
[0042] A facial image replacement processing module is used to perform replacement processing on the first facial image area contained in each of the first image frames based on the position information of the first facial image area to be replaced in the first image frame and the second facial image area contained in the second image frame, so as to determine the first video after replacement.
[0043] This specification also provides an electronic device, including a processor and a memory for storing processor executable instructions, wherein when the processor executes the instructions, a sub-pixel-based facial image replacement data processing method is implemented.
[0044] The present specification also provides a computer-readable storage medium on which computer instructions are stored, and when the instructions are executed, a sub-pixel-based facial image replacement data processing method is implemented.
[0045] A sub-pixel-based facial image replacement data processing method provided in the present specification receives a facial image replacement request for a first video, and in response to the facial image replacement request, determines a first image frame in the first video and a second image frame for performing facial image replacement processing on the first video; wherein the first image frame is an image frame in the first video where a first facial image region to be replaced exists; using a first module of a preset detection model, performing pixel-level key point position detection processing on the first image frame to obtain first position information of a first key point in the first image frame, and using a second module of the preset detection model, performing sub-pixel-level key point offset detection processing on the first key point according to the first position information to obtain a first offset; determining position information of the first facial image region to be replaced in the first image frame according to the first position information and the first offset; performing replacement processing on the first facial image region contained in each of the first image frames according to the position information of the first facial image region to be replaced in the first image frame and the second facial image region contained in the second image frame to obtain the replaced first video. In this way, by using the first module of the preset detection model to perform pixel-level key point position detection processing, the first position information of the first key point in the first image frame is obtained, and then using the second module of the preset detection model to perform sub-pixel-level key point offset detection processing on the first key point to obtain the first offset, and then according to the first position information and the first offset, the position information of the first face image area to be replaced in the first image frame is accurately determined. This step-by-step refinement from pixel to sub-pixel not only significantly improves the positioning accuracy of the face replacement area, but also significantly improves the efficiency, visual effect and realism of face replacement. BRIEF DESCRIPTION OF THE DRAWINGS
[0046] In order to more clearly illustrate the embodiments of this specification, the drawings required for use in the embodiments will be briefly introduced below. The drawings described below are only some embodiments recorded in this specification. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying creative work.
[0047] Figure 1 It is a flowchart of a sub-pixel-based face image replacement data processing method provided by an embodiment of this specification;
[0048] Figure 2 is a schematic diagram of a generative adversarial network structure provided by an embodiment of this specification;
[0049] Figure 3 It is a schematic diagram of the structure of an electronic device provided by an embodiment of this specification;
[0050] Figure 4 It is a schematic diagram of the structure of a sub-pixel-based face image replacement data processing device provided by an embodiment of the present specification;
[0051] Figure 5 It is a flowchart of another sub-pixel-based face image replacement data processing method provided by an embodiment of this specification;
[0052] Figure 6 It is a schematic diagram of a sub-pixel level face key point detection model structure provided by an embodiment of this specification;
[0053] Figure 7 This is a flowchart of a face swapping and frame sequence alignment algorithm provided by an embodiment of this specification;
[0054] Figure 8 is a schematic diagram of a visualization result of a face swapping process provided by an embodiment of the present specification;
[0055] Fig. 9 This is a schematic diagram of the face-swapping effect of a different model method on a human face at multiple side angles provided by an embodiment of this specification. DETAILED DESCRIPTION
[0056] In order to enable those skilled in the art to better understand the technical solutions in this specification, the technical solutions in the embodiments of this specification will be clearly and completely described below in conjunction with the drawings in the embodiments of this specification. Obviously, the described embodiments are only part of the embodiments of this specification, not all of the embodiments. Based on the embodiments in this specification, all other embodiments obtained by ordinary technicians in this field without creative work should fall within the scope of protection of this specification.
[0057] In China, with the development of artificial intelligence technology, face image synthesis technology has gradually matured, especially the application of adversarial neural networks has become mainstream. However, in practical applications, such as face swapping in real-time live broadcasts and post-production video production, due to the lag of traditional film and television shooting coordination technology, the efficiency of processing continuous frames is low and errors are prone to occur, such as blurred facial expressions and unclear boundaries when generating side faces.
[0058] Early face generation methods, such as the U-Net structure-based method proposed by P. Isola et al., achieved style transfer, but due to the irregularities in data set construction and the disorder and confusion in data processing, the generated images had inconsistent semantic information. In addition, the management complexity caused by data loss or human factors increased the difficulty of technical implementation.
[0059] The shortcomings of current technology mainly include the following: Distortion and deformation caused by special angle problems: When the face is at a special angle, the existing technology may not be able to accurately capture the features of the face, resulting in distortion and deformation during the face swap process. Difficulty in matching: Faces at special angles may make it difficult to extract and match facial features, affecting the final swap effect; Low face similarity problem: Difficulty in feature matching: When the similarity between two faces is low, face swap technology may find it difficult to accurately match and synthesize the features of the two faces, resulting in unrealistic synthesis results; Quality degradation: Face swaps with low similarity may result in a degradation in the quality of the synthesized image, lacking realism and authenticity. Face orientation problem: Difficulty in changing viewing angle: Changes in facial orientation will affect the accuracy of facial feature extraction and matching, and current technology may have difficulty processing faces of different orientations; Poor fusion effect: Changes in facial orientation may result in poor fusion of the synthesized image, making the synthesized result look unnatural or unrealistic.
[0060] In view of the root cause of the above problems, this manual uses the first module of the preset detection model to perform pixel-level key point position detection processing to obtain the first position information of the first key point in the first image frame, and then uses the second module of the preset detection model to perform sub-pixel-level key point offset detection processing on the first key point to obtain the first offset, and then accurately determine the position information of the first face image area to be replaced in the first image frame based on the first position information and the first offset. This step-by-step refinement from pixel to sub-pixel not only significantly improves the positioning accuracy of the face replacement area, but also significantly improves the efficiency, visual effect and realism of face replacement.
[0061] See also Figure 1 As shown, the embodiment of this specification provides a method for processing sub-pixel-based facial image replacement data, wherein the method is specifically applied to the server side. When implemented specifically, the method may include the following contents:
[0062] S101: receiving a facial image replacement request for a first video, and in response to the facial image replacement request, determining a first image frame in the first video and a second image frame for performing facial image replacement processing on the first video; wherein the first image frame is an image frame in the first video where a first facial image region to be replaced exists;
[0063] S102: using a first module of a preset detection model, performing pixel-level key point position detection processing on the first image frame to obtain first position information of a first key point in the first image frame, and using a second module of the preset detection model, performing sub-pixel-level key point offset detection processing on the first key point according to the first position information to obtain a first offset;
[0064] S103: Determine position information of a first facial image region to be replaced in the first image frame according to the first position information and the first offset;
[0065] S104: Based on the position information of the first facial image area to be replaced in the first image frame and the second facial image area included in the second image frame, replace the first facial image area included in each of the first image frames to obtain a replaced first video.
[0066] Among them, the above-mentioned image frame can be a still image frame in the first video, and the first video is composed of a series of continuous image frames. The above-mentioned sub-pixel can refer to the accuracy of coordinates or positions refined to a range smaller than a single pixel. Furthermore, there is a distance of 5.2 microns between two pixels, which can be regarded as connected at the macro level, but at the micro level, there are infinitely smaller things between them. We call this smaller thing a sub-pixel. The above-mentioned key point can refer to the position of a specific feature in the first image frame, such as the corners of the eyes, the tip of the nose, and the corners of the mouth of the face. The above-mentioned offset position detection processing can refer to further calculating the precise offset of the key point relative to the center of the pixel by analyzing the probability distribution of the local area of the heat map on the basis of the pixel-level detection results. The above-mentioned offset can refer to the distance between the precise position of the key point and the center of the pixel where it is located, usually a floating point number relative to the center of the pixel.
[0067] In some embodiments, the first module of the preset detection model is used to perform pixel-level key point position detection processing on the first image frame to obtain first position information of the first key point in the first image frame. When specifically implemented, it may include:
[0068] By processing the image frame through the first module of the preset detection model, the key point position can be detected at the pixel level, and the pixel coordinate information of the key point (such as x and y coordinates) can be output; among them, the detected key point coordinate set can be used to calculate the boundary range of the target area and generate a rectangular box containing the upper left corner coordinates and width and height information.
[0069] In some embodiments, using the second module of the preset detection model, according to the first position information, performing sub-pixel level key point offset detection processing on the first key point to obtain a first offset, the specific implementation may include:
[0070] After obtaining the pixel-level position information of the key points through the second module of the preset detection model, the sub-pixel level offset detection processing is further performed on the key point position. This process uses the feature information of the input image to calculate a more accurate offset within the pixel for each key point, thereby determining the more fine-grained position information of the key point. Among them, the first offset reflects the slight adjustment of the relative position of the key point in the pixel coordinate system, which is an important parameter for achieving high-precision key point positioning.
[0071] Furthermore, this offset detection is usually implemented through heat map regression, and the model is optimized based on a specific loss function to ensure that the predicted offset of the key point can be close to the actual offset value. Combining pixel-level position information and sub-pixel-level offset, a high-precision estimate of the key point position can be achieved.
[0072] In some embodiments, determining the position information of the first facial image region to be replaced in the first image frame according to the first position information and the first offset may include:
[0073] For example, when detecting the position of the corner of the eye in a face image, the first module of the preset detection model may locate the pixel block where the corner of the eye is located, such as the coordinates of the upper left corner of the rectangular box is (50,100). However, in order to achieve higher-precision positioning, the second module will further detect the offset of the corner of the eye in the pixel block based on heat map regression. For example, the offset is (0.3,0.6), which means that the corner of the eye is offset horizontally by 0.3 pixels and vertically by 0.6 pixels in the pixel block. Finally, combining the pixel-level coordinates and the offset, the exact position of the corner of the eye is (50.3,100.6).
[0074] Based on the above embodiment, by using the first module of the preset detection model, pixel-level key point position detection processing is performed to obtain the first position information of the first key point in the first image frame, and then using the second module of the preset detection model, sub-pixel-level key point offset detection processing is performed on the first key point to obtain the first offset, and then based on the first position information and the first offset, the position information of the first face image area to be replaced in the first image frame is accurately determined. This step-by-step refinement from pixel to sub-pixel not only significantly improves the positioning accuracy of the face replacement area, but also significantly improves the efficiency, visual effect and realism of face replacement.
[0075] In some embodiments, according to the position information of the first face image area to be replaced in the first image frame and the second face image area included in the second image frame, the first face image area included in each of the first image frames is replaced to obtain the replaced first video. When the method is specifically implemented, it may also include the following contents:
[0076] S1: using the preset detection model, determining the position information of the second face image area in the second image frame;
[0077] S2: using a preset face image replacement model, according to the position information of the first face image area to be replaced in the first image frame, and the position information of the second face image area in the second image frame, replacing the first face image area contained in each of the first image frames, to obtain a replaced first image frame; wherein the preset face image replacement model is a model constructed according to a preset generative algorithm;
[0078] S3: Determine the replaced first video according to the replaced first image frame.
[0079] Among them, the above-mentioned preset face image replacement model can be a model constructed using generative adversarial networks (GAN) or variational autoencoders (VAE).
[0080] In some embodiments, before using the preset facial image replacement model to replace the first facial image region contained in each of the first image frames according to the position information of the first facial image region to be replaced in the first image frame and the position information of the second facial image region in the second image frame to obtain the replaced first image frame, the method may further include the following contents when implemented:
[0081] S1: Acquire sample data, where the sample data includes a first face image and a second face image;
[0082] S2: using the generator of the preset facial image replacement model to replace the first facial image with the second facial image to obtain a replaced first facial image;
[0083] S3: using the discriminator of the preset face image replacement model, determining whether the preset face image replacement model has converged according to the first face image and the replaced first face image, and if it is determined that the preset face image replacement model has not converged, continuing to train the preset face image replacement model according to the sample data until the preset face image replacement model has converged;
[0084] The method of using a preset facial image replacement model to replace the first facial image region contained in each of the first image frames according to the position information of the first facial image region to be replaced in the first image frame and the position information of the second facial image region in the second image frame to obtain the replaced first image frame includes:
[0085] S4: Using a generator of a preset facial image replacement model, replace the first facial image area contained in each of the first image frames according to the position information of the first facial image area to be replaced in the first image frame and the position information of the second facial image area in the second image frame to obtain the replaced first image frame.
[0086] For details, see Figure 2 As shown (in order to avoid infringement of portrait rights, part of the face is blocked), taking the preset face image replacement model as a generative adversarial network model as an example, the generative adversarial network training trains two networks, namely the generator network and the discriminator network.
[0087] In this network structure, the input image is divided into a first face image and a second face image, both of which are obtained through the encoding structure of the generator to obtain image features, and then the intermediate representation layer stores the corresponding feature matrix, and then passes through two decoding structures to generate the image of the person to be replaced (i.e., the first face image) and the image of the person to be replaced (i.e., the second face image), respectively. Finally, the corresponding discriminator network is used to identify the authenticity between the generated image (i.e., the first face image after replacement) and the original image (i.e., the first face image). In this process, the generator generates the image and the discriminator identifies the image, and the balance between the two is constrained by the adversarial loss, wherein the adversarial loss constraint can be determined according to the following formula:
[0088]
[0089] in, is the adversarial loss constraint value, V(D,G) represents the adversarial objective function of the generator and the discriminator, D(x) is the output of the discriminator for the first face image, G(z) is the replaced first face image, and D(G(z))) is the output of the discriminator for the replaced first face image.
[0090] Furthermore, the detailed network layer description of the generative adversarial network model is as follows: the generative adversarial network structure is composed of a symmetrical U-Net, the encoding stage is downsampling feature extraction, and the decoding stage is upsampling feature recovery. The downsampling process is connected to the network layer with the same scale as the upsampling process, which is conducive to the construction of detailed features of the image during the upsampling process; the discriminator network structure obtains image features through continuous downsampling, and the probability of the image being true or false is determined by the softmax function at the end of the discriminator network.
[0091] Based on the above embodiments, the adversarial training mechanism of the generator and the discriminator is used to significantly improve the generation quality of replacement faces, ensure the naturalness and realism of the images, and optimize the efficiency and robustness of the training process through automated convergence judgment.
[0092] In some embodiments, the method of determining the position information of the first facial image region to be replaced in the first image frame according to the first position information and the first offset may further include the following when implemented:
[0093] S1: determining a first adjacent key point related to the first key point in the first image frame, and determining second position information of the first adjacent key point in the first image frame according to the first position information;
[0094] S2: using the second module of the preset detection model, and performing sub-pixel level key point offset detection processing on the first adjacent key point according to the second position information, to obtain a second offset;
[0095] S3: Determine the position information of the first facial image area to be replaced in the first image frame according to the first position information, the second position information, the first offset, and the second offset.
[0096] Based on the above embodiments, by introducing adjacent key points, the accuracy and stability of face image replacement can be significantly improved by comprehensively analyzing the position information of the target key point and its surrounding related key points. Through sub-pixel level offset detection processing, the information of adjacent key points can effectively capture local details and minor changes, reduce error accumulation, and enhance the resolution and accuracy of key point detection.
[0097] In some embodiments, the method for determining the first image frame in the first video may further include the following when implemented:
[0098] S1: performing video frame segmentation processing on the first video according to the number of image frames included in a preset time period in the first video and the image quality of the image frames included in the first video to obtain a plurality of candidate image frames;
[0099] S2: using a preset face region detection model, detecting and processing the region where the first face image to be replaced is located in the candidate image frame, and determining a first region in each of the candidate image frames that contains the first face image to be replaced;
[0100] S3: performing face detection processing on the first area using a preset face detection model to obtain position information of the first face image to be replaced in the first area;
[0101] S4: According to the position information of the first facial image to be replaced in the first area, the pixel value corresponding to the first facial image to be replaced is set as the first pixel value, and the pixel value of the area other than the first facial image to be replaced in the candidate image frame is set as the second pixel value, so as to obtain the first image frame.
[0102] The position information in the first area may be pixel coordinate position information of the first face image, specifically including the coordinates of the upper left corner and the width and height of the rectangular frame.
[0103] In some embodiments, according to the position information of the first facial image to be replaced in the first area, the pixel value corresponding to the first facial image to be replaced is set as the first pixel value, and the pixel value of the area other than the first facial image to be replaced in the candidate image frame is set as the second pixel value to obtain the first image frame. When specifically implemented, it may include:
[0104] Specifically, taking the masking processing of the face mask area as an example, a pixel-level segmentation network is used to detect and segment the face area from the face image, and a black and white mask image of the same size as the face image is generated, where black represents the non-face area and white represents the face area.
[0105] Based on the above embodiment, by distinguishing and allocating pixel values in this way, it is possible to ensure that the face replacement operation focuses on the target area and avoids interference with the background or irrelevant areas. At the same time, the target pixel area and the background area can be clearly defined, and the first face image to be replaced can be accurately positioned to avoid erroneous replacement or affecting the integrity of the background.
[0106] In some embodiments, the method of determining the second image frame for performing face image replacement processing on the first video may further include the following when the method is specifically implemented:
[0107] S1: Acquire a second video corresponding to the first video;
[0108] S2: Determine a matching degree between each image frame in the second video and the first image frame, and determine a second image frame in the second video corresponding to each first image frame based on the matching degree.
[0109] Among them, the above matching degree can be calculated in a variety of ways, including feature point-based matching (such as SIFT, SURF, ORB), deep learning feature extraction (such as using a pre-trained model to extract feature vectors and calculate similarity), or template matching-based methods (such as normalized cross-correlation coefficients).
[0110] Based on the above embodiment, by calculating the matching degree, the second image frame that is most similar to the first image frame can be effectively screened out to ensure the consistency and visual coherence between frames during the replacement process. At the same time, this matching strategy can adapt to the lighting changes and posture deviations in complex scenes, and improve the robustness of the matching.
[0111] In some embodiments, the method of determining the replaced first video according to the replaced first image frame may further include the following contents when the method is specifically implemented:
[0112] S1: Acquire a third image frame and a fourth image frame adjacent to the replaced first image frame in the first video;
[0113] S2: performing size adjustment processing on the replaced first image frame, the third image frame and the fourth image frame respectively according to a preset size to obtain adjusted image frames;
[0114] S3: Determine an initial transformation matrix and Euler angles according to the adjusted image frame;
[0115] S4: adjusting the position information of the first facial image area to be replaced in the first image frame according to the Euler angles and the difference between any two adjacent image frames in the replaced first image frame, the third image frame and the fourth image frame, and updating the initial transformation matrix according to the position information obtained by the adjustment process to obtain a target transformation matrix;
[0116] S5: performing alignment processing on the replaced first image frame according to the target transformation matrix to determine the replaced first video.
[0117] Based on the above embodiment, during the replacement process, the facial posture is accurately described based on the Euler angle, and the position information of the first facial image area to be replaced is dynamically adjusted in combination with the difference between any two adjacent frames in the first image frame, the third image frame, and the fourth image frame after replacement. The subtle changes between the image frames are captured by difference analysis, the regional position deviation is corrected, and the initial transformation matrix is gradually updated to generate the final target transformation matrix. This matrix not only reflects the latest spatial position of the facial area, but also ensures the overall coordination of posture changes and image replacement, providing support for generating high-quality replacement results.
[0118] As can be seen from the above, an embodiment of the present specification provides a sub-pixel-based face image replacement data processing method, which receives a face image replacement request for a first video, and in response to the face image replacement request, determines a first image frame in the first video, and a second image frame for performing face image replacement processing on the first video; wherein the first image frame is an image frame in which a first face image area to be replaced exists in the first video; using a first module of a preset detection model, performing pixel-level key point position detection processing on the first image frame to obtain first position information of a first key point in the first image frame, and using a second module of the preset detection model, performing sub-pixel-level key point offset detection processing on the first key point according to the first position information to obtain a first offset; determining the position information of the first face image area to be replaced in the first image frame according to the first position information and the first offset; performing replacement processing on the first face image area contained in each of the first image frames according to the position information of the first face image area to be replaced in the first image frame and the second face image area contained in the second image frame to obtain the replaced first video. In this way, by using the first module of the preset detection model to perform pixel-level key point position detection processing, the first position information of the first key point in the first image frame is obtained, and then using the second module of the preset detection model to perform sub-pixel-level key point offset detection processing on the first key point to obtain the first offset, and then according to the first position information and the first offset, the position information of the first face image area to be replaced in the first image frame is accurately determined. This step-by-step refinement from pixel to sub-pixel not only significantly improves the positioning accuracy of the face replacement area, but also significantly improves the efficiency, visual effect and realism of face replacement.
[0119] See also Figure 3 As shown, the embodiment of this specification also provides a specific electronic device, wherein the electronic device includes a network communication port 301, a processor 302 and a memory 303, and the above structures are connected through internal cables so that each structure can perform specific data interaction.
[0120] Among them, the network communication port 301 can be specifically used to receive a facial image replacement request for a first video, and in response to the facial image replacement request, determine a first image frame in the first video, and a second image frame for performing facial image replacement processing on the first video; wherein the first image frame is an image frame in the first video where a first facial image area to be replaced exists.
[0121] The processor 302 can be specifically used to use the first module of the preset detection model to perform pixel-level key point position detection processing on the first image frame to obtain first position information of the first key point in the first image frame, and use the second module of the preset detection model to perform sub-pixel level key point offset detection processing on the first key point according to the first position information to obtain a first offset; determine the position information of the first face image area to be replaced in the first image frame according to the first position information and the first offset; and replace the first face image area contained in each of the first image frames according to the position information of the first face image area to be replaced in the first image frame and the second face image area contained in the second image frame to obtain a replaced first video.
[0122] The memory 303 may be specifically used to store corresponding instruction programs.
[0123] Based on the above method, the relevant structural performance of the electronic device can be effectively utilized, the data processing speed of the electronic device can be improved, and the sub-pixel-based face image replacement data processing method can be efficiently implemented.
[0124] In this embodiment, the network communication port 301 can be a virtual port that is bound to different communication protocols so that different data can be sent or received. For example, the network communication port can be a port responsible for web data communication, a port responsible for FTP data communication, or a port responsible for email data communication. In addition, the network communication port can also be a physical communication interface or communication chip. For example, it can be a wireless mobile network communication chip, such as GSM, CDMA, etc.; it can also be a Wifi chip; it can also be a Bluetooth chip.
[0125] In this embodiment, the processor 302 may be implemented in any appropriate manner. For example, the processor may take the form of a microprocessor or processor and a computer-readable medium storing a computer-readable program code (such as software or firmware) executable by the (micro)processor, a logic gate, a switch, an application specific integrated circuit (ASIC), a programmable logic controller, and an embedded microcontroller, etc. This specification does not limit this.
[0126] In this embodiment, the memory 303 may include multiple levels. In a digital system, anything that can store binary data can be a memory; in an integrated circuit, a circuit with a storage function but no physical form is also called a memory, such as RAM, FIFO, etc.; in a system, a storage device with a physical form is also called a memory, such as a memory stick, TF card, etc.
[0127] The embodiment of the present specification also provides a computer-readable storage medium based on the above-mentioned sub-pixel-based facial image replacement data processing method, which receives a facial image replacement request for a first video, and in response to the facial image replacement request, determines a first image frame in the first video, and a second image frame for performing facial image replacement processing on the first video; wherein the first image frame is an image frame in the first video where a first facial image area to be replaced exists; using a first module of a preset detection model, performing pixel-level key point position detection processing on the first image frame to obtain first position information of a first key point in the first image frame, and using a second module of the preset detection model, performing sub-pixel-level key point offset detection processing on the first key point according to the first position information to obtain a first offset; determining the position information of the first facial image area to be replaced in the first image frame according to the first position information and the first offset; performing replacement processing on the first facial image area contained in each of the first image frames according to the position information of the first facial image area to be replaced in the first image frame and the second facial image area contained in the second image frame to obtain the replaced first video.
[0128] In this embodiment, the storage medium includes, but is not limited to, a random access memory (RAM), a read-only memory (ROM), a cache, a hard disk (HDD), or a memory card. The memory may be used to store computer program instructions. The network communication unit may be an interface for network connection communication set in accordance with the standard specified by the communication protocol.
[0129] In this embodiment, the functions and effects specifically implemented by the program instructions stored in the computer-readable storage medium can be explained in comparison with other implementations and will not be described in detail here.
[0130] See also Figure 4 At the software level, the embodiment of this specification also provides a sub-pixel-based face image replacement data processing device, which may specifically include the following structural modules:
[0131] A receiving response module 401 is used to receive a face image replacement request for a first video, and in response to the face image replacement request, determine a first image frame in the first video and a second image frame for performing face image replacement processing on the first video; wherein the first image frame is any image frame in the first video;
[0132] A first information determination module 402 is configured to perform pixel-level key point position detection processing on the first image frame using a first module of a preset detection model to obtain first position information of a first key point in the first image frame, and perform sub-pixel-level key point offset detection processing on the first key point according to the first position information using a second module of the preset detection model to obtain a first offset;
[0133] A position information determination module 403, configured to determine position information of a first facial image region to be replaced in the first image frame according to the first position information and the first offset;
[0134] The face image replacement processing module 404 is used to perform replacement processing on the first face image area contained in each of the first image frames according to the position information of the first face image area to be replaced in the first image frame and the second face image area contained in the second image frame, so as to determine the first video after replacement.
[0135] In some embodiments, the above-mentioned face image replacement processing module 404, when implemented, uses the preset detection model to determine the position information of the second face image area in the second image frame; the first image frame determination module is used to use the preset face image replacement model to replace the first face image area contained in each first image frame according to the position information of the first face image area to be replaced in the first image frame, and the position information of the second face image area in the second image frame, to obtain the replaced first image frame; wherein the preset face image replacement model is a model constructed according to a preset generative algorithm; the first video determination module is used to determine the replaced first video according to the replaced first image frame.
[0136] In some embodiments, before the above-mentioned first image frame determination module, during specific implementation, sample data is obtained, and the sample data includes a first face image and a second face image; using the generator of the preset face image replacement model, the first face image is replaced by the second face image to obtain the replaced first face image; using the discriminator of the preset face image replacement model, according to the first face image and the replaced first face image, it is determined whether the preset face image replacement model has converged, and if it is determined that the preset face image replacement model has not converged, the preset face image replacement model is continued to be trained according to the sample data until the preset face image replacement model is replaced. Model convergence; using the preset face image replacement model to replace the first face image area contained in each of the first image frames according to the position information of the first face image area to be replaced in the first image frame and the position information of the second face image area in the second image frame to obtain the replaced first image frame, including: using a generator of the preset face image replacement model to replace the first face image area contained in each of the first image frames according to the position information of the first face image area to be replaced in the first image frame and the position information of the second face image area in the second image frame to obtain the replaced first image frame.
[0137] In some embodiments, the above-mentioned position information determination module 403, when specifically implemented, determines a first adjacent key point related to the first key point in the first image frame, and determines the second position information of the first adjacent key point in the first image frame based on the first position information; utilizes the second module of the preset detection model to perform sub-pixel level key point offset detection processing on the first adjacent key point based on the second position information to obtain a second offset; determines the position information of the first face image area to be replaced in the first image frame based on the first position information, the second position information, the first offset and the second offset.
[0138] In some embodiments, the above-mentioned receiving response module 401, when specifically implemented, performs video frame segmentation processing on the first video according to the number of image frames contained in a preset time period in the first video and the image quality of the image frames contained in the first video to obtain multiple candidate image frames; uses a preset face area detection model to detect the area where the first face image to be replaced is located in the candidate image frame, and determines the first area in each of the candidate image frames that contains the first face image to be replaced; uses a preset face detection model to perform face detection processing on the first area to obtain position information of the first face image to be replaced in the first area; according to the position information of the first face image to be replaced in the first area, sets the pixel value corresponding to the first face image to be replaced to the first pixel value, and sets the pixel value of the area other than the first face image to be replaced in the candidate image frame to the second pixel value, to obtain the first image frame.
[0139] In some embodiments, in the above-mentioned receiving response module 401, a second image frame for performing facial image replacement processing on the first video is determined. In a specific embodiment, a second video corresponding to the first video is obtained; a matching degree between each image frame in the second video and the first image frame is determined, and based on the matching degree, a second image frame in the second video corresponding to each of the first image frames is determined.
[0140] In some embodiments, the above-mentioned first video determination module, when specifically implemented, obtains the third image frame and the fourth image frame adjacent to the replaced first image frame in the first video; according to the preset size, resizes the replaced first image frame, the third image frame and the fourth image frame respectively to obtain the adjusted image frame; determines the initial transformation matrix and Euler angles according to the adjusted image frame; adjusts the position information of the first face image area to be replaced in the first image frame according to the Euler angles and the difference between any two adjacent image frames in the replaced first image frame, the third image frame and the fourth image frame, and updates the initial transformation matrix according to the position information obtained by the adjustment process to obtain the target transformation matrix; aligns the replaced first image frame according to the target transformation matrix to determine the replaced first video.
[0141] It should be noted that the units, devices or modules described in the above embodiments can be implemented by computer chips or entities, or by products with certain functions. For the convenience of description, the above devices are described separately by functions divided into various modules. Of course, when implementing this specification, the functions of each module can be implemented in the same or more software and / or hardware, or the modules that implement the same function can be implemented by a combination of multiple sub-modules or sub-units. The device embodiments described above are only schematic. For example, the division of the units is only a logical function division. There may be other division methods in actual implementation, such as multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the mutual coupling or direct coupling or communication connection shown or discussed can be through some interfaces, indirect coupling or communication connection of devices or units, which can be electrical, mechanical or other forms.
[0142] As can be seen from the above, a sub-pixel-based face image replacement data processing device provided in the embodiment of this specification uses the first module of the preset detection model to perform pixel-level key point position detection processing to obtain the first position information of the first key point in the first image frame, and then uses the second module of the preset detection model to perform sub-pixel-level key point offset detection processing on the first key point to obtain the first offset, and then accurately determine the position information of the first face image area to be replaced in the first image frame based on the first position information and the first offset. This step-by-step refinement from pixel to sub-pixel not only significantly improves the positioning accuracy of the face replacement area, but also significantly improves the efficiency, visual effect and realism of face replacement.
[0143] In a specific scenario example, a sub-pixel-based facial image replacement data processing method and device provided in this specification can be applied. By using the first module of the preset detection model, pixel-level key point position detection processing is performed to obtain the first position information of the first key point in the first image frame, and then the second module of the preset detection model is used to perform sub-pixel-level key point offset detection processing on the first key point to obtain the first offset, and then accurately determine the position information of the first facial image area to be replaced in the first image frame based on the first position information and the first offset. This gradual refinement from pixel to sub-pixel not only significantly improves the positioning accuracy of the face replacement area, but also significantly improves the efficiency, visual effect and realism of the face replacement. The specific implementation process may include the following content.
[0144] See also Figure 5 As shown, the sub-pixel-based facial image replacement data processing method also includes the following steps:
[0145] S1: video frame segmentation;
[0146] In this step, the ffmpeg library and opencv library related video and image processing technologies are combined to realize the segmentation of image files from video files. The controllable parameters include the frequency of segmented image frames, that is, the number of image frames contained in one second; image quality, that is, the information entropy density contained in a segmented image, which is divided into lossy compression and lossless compression. In lossy compression, the greater the compression rate, the worse the image quality.
[0147] S2: face detection;
[0148] In this step, the application adopts open source to implement the s3fd algorithm based on the convolutional neural network model, which has a good effect on blurry face and small face detection. The model outputs the coordinates of the rectangular target box containing the face, specifically the coordinates of the upper left corner of the rectangular box and the width and height.
[0149] S3: face mask area mask;
[0150] This step detects and segments the facial area from the face image through a pixel-level segmentation network, and generates a black-and-white mask image of the same size as the face image, where black represents the non-facial area and white represents the facial area. The mask area mask model in this application adopts the open source BiSeNet segmentation model.
[0151] S4: Sub-pixel facial key point detection;
[0152] In this step, the present application uses a sub-pixel level face key point detection model, the structure of which is as follows Figure 6 shown.
[0153] In this network structure, the convolution layer continuously learns relevant features in the image. Compared with the coordinate point regression method, the detection head does not directly output the two-dimensional coordinates of N key points, but predicts the predicted score heat map of N key points, the key point position offset heat map and the C adjacent key point coordinate offset heat map, so as to achieve the prediction of sub-pixel key point position coordinates. Specifically, firstly, the model algorithm essentially embeds the heat map regression method into the key point coordinate regression method, and outputs the feature vector heat map at the head of the detection head. Secondly, in order to obtain the sub-pixel coordinates of the key points, the offset coordinates of the key points and a custom number of adjacent offset components are predicted respectively. At the model implementation level, the predictions of the two are separated and can be processed in parallel.
[0154] The principle of sub-pixel face key point detection includes the following: first, the convolutional neural network input accepts an RGB image containing a face, and the feature extraction layer is described by the ResNet network structure. The feature matrix is extracted, and its mathematical expression is shown in Formula 1:
[0155] F = ResNet(I) (1)
[0156] Where I is the input image and F is the feature matrix.
[0157] During the model training phase, the weights of each layer of the convolutional neural network are continuously updated with the iteration of the input image and the constraints of the loss function. The definition of the total loss function is shown in Formula 2:
[0158] L=L s +αL O +βL N (2)
[0159] Among them, L s is the prediction score loss, L O is the key point coordinate offset loss, L N is the loss of the offset component of the adjacent key points, α and β are hyperparameters, satisfying the relationship α+β=1, balancing the loss between the offset coordinates and the adjacent offset components.
[0160] Specifically, the prediction score loss is calculated as shown in Formula 3:
[0161]
[0162] Among them, s * is the true score value, s ′ is the predicted score, N is the number of key points, H F and W F Represent the height and width of the feature map respectively, i is the two-dimensional coordinate of the i-th key point; j and k are the height and width of the feature map respectively. F and width WF Any pixel position.
[0163] Specifically, the key point coordinate offset loss is calculated as shown in Formula 4:
[0164]
[0165] Among them, * is the actual offset value, o ′ is the predicted offset value, N is the number of key points, i is the two-dimensional coordinate of the i-th key point; j and k are the feature map height H F and width W F Any pixel position. l is the coordinate subscript value, the key point coordinate is a two-dimensional coordinate, l∈[1,2]
[0166] Specifically, the loss of the offset component of the adjacent key points is calculated as shown in Formula 5:
[0167]
[0168] Among them, n * is the actual adjacent offset value, n ′ To predict the neighboring offset value, N is the number of key points, C is the offset heat map of the key point coordinates with C neighbors, and m is the offset heat map of the mth neighboring key point coordinates.
[0169] When determining the loss value function, the score loss uses L2 loss, which has a better calculation effect for classification problems, while the coordinate loss uses L1 loss, which has a better calculation effect for regression problems. The feature extraction stage includes but is not limited to ResNet network structure, MobileNet network structure and VGG network structure.
[0170] S5: Generative adversarial network training;
[0171] In this step, two networks are trained, namely the generator network and the discriminator network. The specific network structure diagram is as follows: Figure 2 shown.
[0172] In this network structure, the input image is divided into a first face image and a second face image, both of which are obtained through the encoding structure of the generator to obtain image features, and then the intermediate representation layer stores the corresponding feature matrix, and then passes through two decoding structures to generate the image of the person to be replaced (i.e., the first face image) and the image of the person to be replaced (i.e., the second face image), respectively. Finally, the corresponding discriminator network is used to identify the authenticity between the generated image (i.e., the first face image after replacement) and the original image (i.e., the first face image). In this process, the generator generates the image and the discriminator identifies the image, and the balance between the two is constrained by the adversarial loss, wherein the adversarial loss constraint can be determined according to the following formula:
[0173]
[0174] in, is the adversarial loss constraint value, V(D,G) represents the adversarial objective function of the generator and the discriminator, D(x) is the output of the discriminator for the first face image, G(z) is the replaced first face image, and D(G(z))) is the output of the discriminator for the replaced first face image.
[0175] S6: face swapping and inter-frame sequence alignment;
[0176] In this step, the face swap model and the aligned sub-pixel coordinates of the key points of the face need to be combined to calculate the affine transformation matrix, and the relative position of the real face in the three-dimensional space is obtained (the relative position is expressed in Euler angles), so as to realize the alignment of the face swap and the inter-frame sequence. Figure 7 shown.
[0177] Specifically, according to the affine transformation principle of the image, the coordinates of the key points are sampled and the transformation matrix is calculated. The transformation matrix can ensure the consistency of the flatness and parallelism of the image content before and after the transformation. Through the transformation matrix, the rectangular area containing the face can be cut out from the frame image, and the cropped image can be arbitrarily enlarged or reduced by the scaling characteristics of the transformation matrix. Secondly, the orientation information of the face image in three-dimensional space, namely the Euler angle, needs to be calculated. In the continuous frame sequence alignment stage, three images are taken as a processing unit, and each image is scaled to three different scales using the Laplace pyramid. First, the transformation matrix is fine-tuned for the first time according to the images of different scales, and then the key points are fine-tuned with the difference between the Euler angle and the continuous frames. Finally, the transformation matrix is recalculated according to the fine-tuned key points; in the face-changing stage, the image containing the face area is cut out according to the transformation matrix calculated in the sequence frame alignment stage, and the synthetic face image is generated by the model and restored to the original sequence frame to obtain the final face-changing image frame. Then, the process is repeated in sequence until the end frame. The visualization result from video frame extraction to face synthesis is shown in the figure below. Figure 8 As shown (in order to avoid infringement of portrait rights, part of the face is blocked).
[0178] In some embodiments, in order to verify the effectiveness of the method proposed in this application, an experiment is designed to compare and analyze the FaceSwap face swapping algorithm. The experimental steps are as follows:
[0179] S1: Collect and prepare data. In order to clearly observe the difference of transformed faces, we choose to transform female faces into male faces. The training set contains 5238 continuous frames of male videos, the training set contains 5256 continuous frames of female videos, and the test set contains 1000 continuous frames of female videos. The image frames contain transformations of various perspectives and different facial expressions.
[0180] S2: Use the training set data to train the FaceSwap face-changing model and the face-changing model trained by the method proposed in this application respectively;
[0181] S3: Use the test set data to test the FaceSwap face-changing model and the face-changing model trained by the method proposed in this application;
[0182] S4: Analysis of the effects of two model methods on face swapping at different side face angles. Fig. 9 As shown (in order to avoid infringement of portrait rights, part of the face is blocked):
[0183] like Fig. 9 As shown in the visualization effect, when it comes to side faces and large side faces, the stability and image consistency of the method proposed in the present application are better than those of the FaceSwap method. The FaceSwap method has obvious signs of inconsistency on the edge of the face, especially when converting large side faces, the junction of the character's hair and face shows a completely inconsistent synthesis effect, while the method of the present application can maintain consistency very well.
[0184] Obviously, the embodiments described above are only some embodiments of the present application, rather than all embodiments. The preferred embodiments of the present application are given in the accompanying drawings, but they do not limit the patent scope of the present application. The present application can be implemented in many different forms. On the contrary, the purpose of providing these embodiments is to make the understanding of the disclosure of the present application more thorough and comprehensive. Although the present application is described in detail with reference to the aforementioned embodiments, for those skilled in the art, it is still possible to modify the technical solutions recorded in the aforementioned specific implementation methods, or to perform equivalent replacement of some of the technical features therein. Any equivalent structure made using the contents of the specification and drawings of this application, directly or indirectly used in other related technical fields, is similarly within the scope of patent protection of this application.
[0185] In some embodiments, by splitting the extraction of facial feature information into a facial mask module and a facial key point extraction module, the feature granularity is refined, which helps to extract facial feature information more accurately and provide a better data basis for subsequent model training. The training method of the generative adversarial network (GAN) is adopted, combined with the generator and discriminator networks, and the process of face synthesis is controlled through adversarial loss to improve the similarity between the generated face image and the real face, thereby ensuring the quality of the synthesized image. Face alignment based on sub-pixel features: Using sub-pixel features for face alignment during the face swapping process can achieve more precise facial feature matching and improve the fusion and visual effect of the synthesized face with the original image.
[0186] Based on the above embodiments, the following economic benefits or advantages can be brought about: Entertainment industry: Face swap technology can add fun and innovation to entertainment content such as film and television works, television programs, and advertisements, attract more audiences, and enhance the attractiveness and market competitiveness of the content, thereby bringing higher ratings and advertising benefits. Advertising marketing: Face swap technology can be used to create vivid and interesting advertisements, attract more consumers' attention, increase the click-through rate and conversion rate of advertisements, and increase brand exposure and sales. Virtual image endorsement: Face swap technology can allow virtual images such as stars and celebrities to endorse brands or products, saving endorsement fees and reducing marketing costs. At the same time, it can also avoid the risks that may exist in real image endorsements to a certain extent. Save production costs: In film and television production, face swap technology can reduce shooting costs and post-production costs, improve production efficiency, shorten production cycles, thereby saving production costs and improving production benefits. Personalized customization: Through face swap technology, individual users can customize their own unique images to meet personalized needs, promote the development of the personalized customization market, and bring more business opportunities to individual users and related companies.
[0187] Although the present specification provides method operation steps as described in the embodiments or flow charts, more or less operation steps may be included based on conventional or non-creative means. The order of steps listed in the embodiments is only one way of executing the order of many steps, and does not represent a unique execution order. When the device or client product in practice is executed, it can be executed in sequence or in parallel according to the method shown in the embodiments or the drawings (for example, a parallel processor or a multi-threaded processing environment, or even a distributed data processing environment). The term "include", "include" or any other variant thereof is intended to cover non-exclusive inclusion, so that the process, method, product or device including a series of elements includes not only those elements, but also includes other elements that are not explicitly listed, or also includes elements inherent to such a process, method, product or device. In the absence of more restrictions, it is not excluded that there are other identical or equivalent elements in the process, method, product or device including the elements. The first, second, etc. words are used to represent the name, and do not represent any particular order.
[0188] Those skilled in the art also know that, in addition to implementing the controller in a purely computer-readable program code, the controller can be made to implement the same function in the form of logic gates, switches, application-specific integrated circuits, programmable logic controllers, and embedded microcontrollers by logically programming the method steps. Therefore, such a controller can be considered as a hardware component, and the devices for implementing various functions included therein can also be considered as structures within the hardware component. Or even, the devices for implementing various functions can be considered as both software modules for implementing the method and structures within the hardware component.
[0189] Through the description of the above embodiments, it can be known that those skilled in the art can clearly understand that the present specification can be implemented by means of software plus a necessary general hardware platform. Based on such an understanding, the technical solution of the present specification can essentially be embodied in the form of a software product, which can be stored in a storage medium such as ROM / RAM, a disk, an optical disk, etc., and includes a number of instructions for enabling a computer device (which can be a personal computer, a mobile terminal, a server, or a network device, etc.) to execute the methods described in each embodiment of the present specification or some parts of the embodiments.
[0190] Although the present specification is described through embodiments, those skilled in the art will appreciate that there are many modifications and changes to the present specification without departing from the spirit of the present specification, and it is intended that the appended claims include these modifications and changes without departing from the spirit of the present specification.
Claims
1. A method for processing face image replacement data based on sub-pixel, characterized in that: include: Receiving a facial image replacement request for a first video, and in response to the facial image replacement request, determining a first image frame in the first video and a second image frame for performing facial image replacement processing on the first video; wherein the first image frame is an image frame in the first video where a first facial image area to be replaced exists; Using a first module of a preset detection model, performing pixel-level key point position detection processing on the first image frame to obtain first position information of a first key point in the first image frame, and using a second module of the preset detection model, performing sub-pixel-level key point offset detection processing on the first key point according to the first position information to obtain a first offset; Determine, according to the first position information and the first offset, position information of the first facial image region to be replaced in the first image frame; Based on the position information of the first facial image area to be replaced in the first image frame and the second facial image area contained in the second image frame, the first facial image area contained in each of the first image frames is replaced to obtain a replaced first video.
2. The method according to claim 1, characterized in that The replacing process is performed on the first face image region contained in each of the first image frames according to the position information of the first face image region to be replaced in the first image frame and the second face image region contained in the second image frame to obtain the replaced first video, including: Determine the position information of the second facial image region in the second image frame by using the preset detection model; Using a preset face image replacement model, according to the position information of the first face image area to be replaced in the first image frame, and the position information of the second face image area in the second image frame, the first face image area contained in each of the first image frames is replaced to obtain a replaced first image frame; wherein the preset face image replacement model is a model constructed according to a preset generative algorithm; The replaced first video is determined according to the replaced first image frame.
3. The method according to claim 2, characterized in that Before using the preset facial image replacement model to replace the first facial image region contained in each of the first image frames according to the position information of the first facial image region to be replaced in the first image frame and the position information of the second facial image region in the second image frame to obtain the replaced first image frame, the method further includes: Acquire sample data, the sample data including a first face image and a second face image; Using the generator of the preset facial image replacement model, the first facial image is replaced by the second facial image to obtain a replaced first facial image; Determining whether the preset face image replacement model has converged based on the first face image and the replaced first face image using the discriminator of the preset face image replacement model, and if it is determined that the preset face image replacement model has not converged, continuing to train the preset face image replacement model based on the sample data until the preset face image replacement model has converged; The method of using a preset facial image replacement model to replace the first facial image region contained in each of the first image frames according to the position information of the first facial image region to be replaced in the first image frame and the position information of the second facial image region in the second image frame to obtain the replaced first image frame includes: Using a generator of a preset facial image replacement model, replacement processing is performed on the first facial image area contained in each of the first image frames according to the position information of the first facial image area to be replaced in the first image frame, and the position information of the second facial image area in the second image frame, so as to obtain the replaced first image frame.
4. The method according to claim 1, characterized in that: Determining the position information of the first facial image region to be replaced in the first image frame according to the first position information and the first offset includes: Determine a first adjacent key point related to the first key point in the first image frame, and determine second position information of the first adjacent key point in the first image frame according to the first position information; Using the second module of the preset detection model, and according to the second position information, performing sub-pixel level key point offset detection processing on the first adjacent key point to obtain a second offset; The position information of the first facial image region to be replaced in the first image frame is determined according to the first position information, the second position information, the first offset, and the second offset.
5. The method according to claim 1, characterized in that The determining a first image frame in the first video includes: According to the number of image frames contained in a preset time period in the first video and the image quality of the image frames contained in the first video, the first video is segmented into video frames to obtain a plurality of candidate image frames; Using a preset face region detection model, detecting and processing the region where the first face image to be replaced is located in the candidate image frame, and determining a first region in each of the candidate image frames that contains the first face image to be replaced; Using a preset face detection model, performing face detection processing on the first area to obtain position information of the first face image to be replaced in the first area; According to the position information of the first facial image to be replaced in the first area, the pixel value corresponding to the first facial image to be replaced is set as the first pixel value, and the pixel value of the area other than the first facial image to be replaced in the candidate image frame is set as the second pixel value to obtain the first image frame.
6. The method according to claim 1, characterized in that The determining of a second image frame for performing face image replacement processing on the first video includes: Acquire a second video corresponding to the first video; Determine a degree of matching between each image frame in the second video and the first image frame, and determine, based on the degree of matching, a second image frame in the second video corresponding to each of the first image frames.
7. The method according to claim 2, characterized in that The step of determining the replaced first video according to the replaced first image frame includes: Acquire a third image frame and a fourth image frame in the first video that are adjacent to the replaced first image frame; According to a preset size, resizing the replaced first image frame, the third image frame, and the fourth image frame respectively to obtain resized image frames; Determining an initial transformation matrix and Euler angles according to the adjusted image frame; According to the Euler angles and the difference between any two adjacent image frames in the replaced first image frame, the third image frame and the fourth image frame, adjusting the position information of the first facial image area to be replaced in the first image frame, and updating the initial transformation matrix according to the position information obtained by the adjustment process to obtain a target transformation matrix; According to the target transformation matrix, the replaced first image frame is aligned to determine the replaced first video.
8. A sub-pixel-based facial image replacement data processing device, characterized in that: include: A receiving response module, configured to receive a face image replacement request for a first video, and in response to the face image replacement request, determine a first image frame in the first video, and a second image frame for performing face image replacement processing on the first video; wherein the first image frame is any image frame in the first video; a first information determination module, configured to perform pixel-level key point position detection processing on the first image frame using a first module of a preset detection model to obtain first position information of a first key point in the first image frame, and to perform sub-pixel-level key point offset detection processing on the first key point according to the first position information using a second module of the preset detection model to obtain a first offset; a position information determining module, configured to determine position information of a first facial image region to be replaced in the first image frame according to the first position information and the first offset; A facial image replacement processing module is used to perform replacement processing on the first facial image area contained in each of the first image frames based on the position information of the first facial image area to be replaced in the first image frame and the second facial image area contained in the second image frame, so as to determine the first video after replacement.
9. An electronic device, characterized in that: It comprises a processor and a memory for storing processor executable instructions, and when the processor executes the instructions, the steps of the sub-pixel based facial image replacement data processing method described in any one of claims 1 to 7 are implemented.
10. A computer-readable storage medium, characterized in that: Computer instructions are stored thereon, and when the instructions are executed by a processor, the steps of the sub-pixel-based facial image replacement data processing method described in any one of claims 1 to 7 are implemented.