Image Processing Method, Apparatus, Device, and Storage Medium
In the image processing method, the blur processing technology of motion feature parameter set and texture image is used to generate natural virtual face image frames, which solves the problem of face distortion and deformation of virtual face images when expression changes, and achieves smooth follow-up of face expressions.
Patent Information
- Application Number
- CN202010250999.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2020-04-01
- Publication Date
- 2025-05-30
- Estimated Expiration
- 2040-04-01
AI Technical Summary
In the prior art, when the virtual face image changes in expressions with the driving face image, it may cause the key points of the face to overlap, resulting in the deformation and distortion of the face in the virtual face image.
By determining the motion feature parameter set of each face key points based on the current driving face image frame and the adjacent historical driving face image frame, the motion feature adjustment parameter set is generated, and finally generating the currently driven face image frame to avoid face distortion and deformation.
It is realized that when the driven face image is adjusted according to the driving face image, it avoids the distortion and deformation of the human face image, and ensures the natural and smooth changes of the virtual face image.
Smart Images

Figure CN113496506B_ABST
Abstract
Description
Technical Field
[0001] Embodiments of the present invention relate to the technical field of image processing, and in particular, to an image processing method, apparatus, device, and storage medium. Background Art
[0002] With the development of society, electronic devices such as mobile phones and tablet computers have been widely used in aspects such as learning, entertainment, and work. Many electronic devices are equipped with cameras for operations such as taking pictures, recording videos, and live streaming. When the image data of the camera contains a human face, static virtual human face images can also be made to follow the facial expressions of the human face to achieve various action effects.
[0003] Currently, according to the position correspondence relationship between the human face key points in the driving human face image and the virtual human face key points in the virtual human face image, each virtual human face key point can be made to change accordingly following each human face key point, so that the virtual human face image follows the driving human face image to perform corresponding facial expression changes.
[0004] In the process of implementing the present invention, the inventors found that the prior art has the following defects: in the driving human face and the virtual human face corresponding to the driving human face, if the positions between some human face key points are relatively close, then when the driving human face undergoes a facial expression change, the relatively close virtual human face key points in the virtual human face image will overlap in position after corresponding adjustment, resulting in the distortion of the human face in the driven virtual human face image. Summary of the Invention
[0005] Embodiments of the present invention provide an image processing method, apparatus, device, and storage medium, which realize that when the driven human face image follows the driving human face image to perform corresponding human face expression adjustment, the distortion of the human face in the adjusted driven human face image can be avoided.
[0006] In a first aspect, embodiments of the present invention provide an image processing method, including:
[0007] Determine a set of motion feature parameters corresponding to each human face key point in the historical driving human face according to the current driving human face image frame and the adjacent historical driving human face image frame;
[0008] Map each motion feature parameter in the set of motion feature parameters to a texture image to obtain a motion texture image, and perform image blurring processing on the motion texture image;
[0009] Generate a set of motion feature adjustment parameters according to the blurred image, and generate the current driven human face image frame according to the set of motion feature adjustment parameters and the historical driven human face image frame.
[0010] Optionally, mapping each motion feature parameter in the motion feature parameter set to a texture image to obtain a motion texture image, including:
[0011] Obtaining the position correspondence between each face key point in the historical driving face image frame and the pixel points in the fragment shader of the texture image;
[0012] According to the position correspondence, mapping the motion feature parameters corresponding to each face key point to the texture image respectively to obtain a motion texture image.
[0013] Optionally, the motion feature parameters corresponding to the face key points include: horizontal motion change amount and vertical motion change amount;
[0014] According to the position correspondence, mapping the motion feature parameters corresponding to each face key point to the texture image respectively, including:
[0015] Obtaining a target horizontal motion change amount and a target vertical motion change amount corresponding to the target face key point currently being mapped;
[0016] According to the position correspondence, obtaining a target pixel point in the texture image that matches the target face key point;
[0017] Assigning the R channel of the target pixel point to the target horizontal motion change amount, the G channel to the target vertical motion change amount, and the B channel to 0.
[0018] Optionally, performing image blurring processing on the motion texture image, including:
[0019] Performing a Gaussian convolution operation on the motion texture image using a Graphics Processing Unit (GPU) to obtain a blurred image.
[0020] Optionally, generating a motion feature adjustment parameter set according to the blurred image, including:
[0021] According to the position correspondence, determining the pixel points corresponding to the face key points in the blurred image respectively;
[0022] Generating motion feature adjustment parameters corresponding to the face key points respectively according to the channel values of the R channel and the G channel of each pixel point;
[0023] Combining the motion feature adjustment parameters to obtain a motion feature adjustment parameter set.
[0024] Optionally, before determining the motion feature parameter set corresponding to each face key point in the historical driving face according to the current driving face image frame and the adjacent historical driving face image frame, it further includes:
[0025] Determine the target fixed points corresponding to the current driving face image frame, the historical driving face image frame, and the historical driven face image frame respectively;
[0026] According to the current driving face image frame and the adjacent historical driving face image frame, determine the set of motion feature parameters corresponding to each face key point in the historical driving face, including:
[0027] Generate a set of motion feature parameters according to each face key point in the current driving face image frame, each face key point in the historical driving face image frame, the target fixed point corresponding to the current driving face image frame, and the target fixed point corresponding to the historical driving face image frame.
[0028] Optionally, generate a set of motion feature parameters according to each face key point in the current driving face image frame, each face key point in the historical driving face image frame, the target fixed point corresponding to the current driving face image frame, and the target fixed point corresponding to the historical driving face image frame, including:
[0029] Obtain the first position vector between each face key point in the current driving face image frame and the corresponding target fixed point, and the second position vector between each face key point in the historical driving face image frame and the corresponding target fixed point;
[0030] Calculate the vector difference between each second position vector and the corresponding first position vector;
[0031] Calculate the product of each vector difference and the face scaling ratio to obtain the face key point acceleration vector matching each face key point in the historical driving face image frame;
[0032] Decompose each face key point acceleration vector along the horizontal and vertical directions to obtain the horizontal motion change amount and the vertical motion change amount corresponding to each face key point as the set of motion feature parameters.
[0033] Optionally, generate the current driven face image frame according to the set of motion feature adjustment parameters and the historical driven face image frame, including:
[0034] Establish a blank image matching the historical driven face image frame;
[0035] According to the set of motion feature adjustment parameters, determine the grid deformation method of each original grid in the historical driven face image frame; wherein, the historical driven face image frame is composed of multiple original grids divided by each face key point therein;
[0036] According to the grid deformation method, divide a plurality of target deformation grids corresponding to the original grids in the blank image;
[0037] According to the positional correspondence between the original mesh and the target deformed mesh, each pixel point in the original mesh is mapped to the corresponding target deformed mesh to obtain the current driven face image frame.
[0038] In a second aspect, an embodiment of the present invention further provides an image processing apparatus, including:
[0039] A motion feature parameter set generation module, configured to determine a motion feature parameter set corresponding to each face key point in the historical driven face according to the current driven face image frame and the adjacent historical driven face image frames;
[0040] An image blurring processing module, configured to map each motion feature parameter in the motion feature parameter set to a texture image to obtain a motion texture image, and perform image blurring processing on the motion texture image;
[0041] A face driving module, configured to generate a motion feature adjustment parameter set according to the blurred image, and generate the current driven face image frame according to the motion feature adjustment parameter set and the historical driven face image frame.
[0042] In a third aspect, an embodiment of the present invention further provides a device, including:
[0043] One or more processors;
[0044] A storage device, configured to store one or more programs,
[0045] When the one or more programs are executed by the one or more processors, the one or more processors implement the image processing method provided in any embodiment of the present invention.
[0046] In a fourth aspect, an embodiment of the present invention further provides a computer-readable storage medium, on which a computer program is stored, and when the program is executed by a processor, the image processing method provided in any embodiment of the present invention is implemented.
[0047] In the embodiment of the present invention, according to the current driven face image frame and the adjacent historical driven face image frames, a motion feature parameter set corresponding to each face key point in the historical driven face is determined. By mapping each motion feature parameter in the motion feature parameter set to a texture image, a motion texture image is obtained, and image blurring processing is performed on the motion texture image. Then, a motion feature adjustment parameter set is generated according to the blurred image, and the current driven face image frame is generated according to the motion feature adjustment parameter set and the historical driven face image frame, which solves the problem of face deformation and distortion in the driven virtual face image in the prior art, and realizes that when the driven face image follows the driving face image for corresponding face expression adjustment, the face in the adjusted driven face image can be avoided from being distorted and deformed. Description of the Drawings
[0048] Figure 1a It is a flowchart of an image processing method in Embodiment 1 of the present invention;
[0049] Figure 1b It is a schematic diagram of facial key points in Embodiment 1 of the present invention;
[0050] Figure 1c It is a schematic diagram of the motion feature parameters of facial key points in Embodiment 1 of the present invention;
[0051] Figure 2a It is a flowchart of an image processing method in Embodiment 2 of the present invention;
[0052] Figure 2b It is a schematic diagram of the original grid in a facial image in Embodiment 2 of the present invention;
[0053] Figure 2c It is a schematic diagram of the interpupillary distance of a facial image in Embodiment 2 of the present invention;
[0054] Figure 2d It is a schematic diagram of a facial positioning rectangle in Embodiment 2 of the present invention;
[0055] Figure 3 It is a schematic structural diagram of an image processing device in Embodiment 3 of the present invention;
[0056] Figure 4 It is a schematic structural diagram of a device in Embodiment 4 of the present invention. Detailed implementation manners
[0057] The present invention will be further described in detail below with reference to the accompanying drawings and embodiments. It can be understood that the specific embodiments described herein are only used to explain the present invention, rather than limiting the present invention. In addition, it should be noted that for the sake of description, only parts related to the present invention rather than all structures are shown in the accompanying drawings.
[0058] Embodiment 1
[0059] Figure 1a It is a flowchart of an image processing method in Embodiment 1 of the present invention. This embodiment is applicable to the situation where a historical driven face follows a driving face to perform real-time expression transformation. This method can be executed by an image processing device, which can be implemented by hardware and / or software and is generally integrated in a device providing image processing services. As Figure 1a shown, the method includes:
[0060] Step 110: Determine a set of motion feature parameters corresponding to each facial key point in the historical driving face according to the current driving face image frame and the adjacent historical driving face image frames.
[0061] In this embodiment, both the current driving face image frame and the adjacent historical driving face image frames are video frames intercepted from a live video. The historical driving face image frame is intercepted earlier in time and includes the face of the live streamer before the expression change, for example, the smiling face of the live streamer. The current driving face image frame is intercepted later in time and includes the face of the live streamer after the expression change, for example, the laughing face of the live streamer. Among them, the faces in the historical driving face image frame and the current driving face image frame can be the front face facing the camera or the side face facing the camera.
[0062] Optionally, by performing face detection on the driving face image frame, each face key point in the driving face image frame can be recognized, such as eyebrows, eyes, nose, mouth, facial contour, etc., as Figure 1b shown. Among them, the number of face key points can be set according to the actual situation.
[0063] In this embodiment, according to the coordinates of each face key point in the current driving face image frame and the coordinates of each face key point in the historical driving face image frame, the coordinate change amount of each face key point in the current driving face image frame relative to each face key point in the historical driving face image frame can be calculated, and then the motion feature parameter set corresponding to each face key point in the historical driving face can be determined.
[0064] Step 120: Map each motion feature parameter in the motion feature parameter set to the texture image to obtain a motion texture image, and perform image blurring processing on the motion texture image.
[0065] In this embodiment, in order to facilitate data smoothing of each motion feature parameter and enable the face image frame adjusted according to the motion feature parameter set to achieve a blurring effect, during the live video texture rendering process, in addition to processing the texture images of the current driving face image frame and the historical driving face image frames, a texture image for processing motion feature parameters can be bound at the same time. After mapping each motion feature parameter to this texture image, data smoothing of each motion feature parameter can be achieved by performing corresponding operations on this texture image. Among them, the texture image for processing each motion feature parameter can be understood as a picture with the same picture as the historical driving face image frame. Hereinafter, the texture image for processing motion feature parameters will be simply referred to as the texture image.
[0066] Optionally, mapping each motion feature parameter in the motion feature parameter set to the texture image to obtain a motion texture image may include: obtaining the position correspondence between each face key point in the historical driving face image frame and the pixel points in the fragment shader of the texture image; according to the position correspondence, mapping the motion feature parameter corresponding to each face key point to the texture image respectively to obtain a motion texture image.
[0067] In this embodiment, during the rendering of the live video texture, the vertex shader is used to process the rotation transformation, translation transformation, etc. of each vertex in the texture image, and the fragment shader is used to process the color calculation and filling of each pixel point in the texture image. Compared with processing the texture images of the current driving face image frame and the historical driving face image frames, the vertex shader for processing the texture image of the motion feature parameters remains unchanged. A position correspondence relationship is established between each pixel point in the fragment shader and each face key point in the historical driving face image frame. For example, the nose tip key point in the historical driving face image frame corresponds to the pixel point at the nose tip position in the fragment shader. According to the position correspondence relationship, the motion feature parameters corresponding to each face key point of the historical driving face image frame can be respectively mapped to the corresponding pixel points in the fragment shader to obtain a motion texture image.
[0068] Optionally, the motion feature parameters corresponding to the face key points include: the horizontal motion change amount and the vertical motion change amount; according to the position correspondence relationship, mapping the motion feature parameters corresponding to each face key point to the texture image respectively may include: obtaining the target horizontal motion change amount and the target vertical motion change amount corresponding to the target face key point currently mapped; obtaining the target pixel point matching the target face key point in the texture image according to the position correspondence relationship; assigning the R channel of the target pixel point to the target horizontal motion change amount, the G channel to the target vertical motion change amount, and the B channel to 0.
[0069] In this embodiment, in order to facilitate mapping the coordinate change amount corresponding to the face key point to different channels of the corresponding pixel points in the fragment shader, the motion feature parameters are set to include the horizontal motion change amount and the vertical motion change amount. Exemplarily, as Figure 1c shown, the horizontal motion change amount X corresponding to the face contour key point in the figure is 0.1, and the vertical motion change amount Y is 0.1.
[0070] In this embodiment, when mapping each motion feature parameter to the texture image, first determine the target face key point to be mapped currently, as well as the target horizontal motion change amount and the target vertical motion change amount corresponding to the target face key point, such as 0.2 and 0.15. Then, according to the position correspondence between each face key point and the pixel points in the fragment shader, find the target pixel points in the texture image that match the target face key point, assign the R channel of the target pixel point to the target horizontal motion change amount, such as 0.2, the G channel to the target vertical motion change amount, such as 0.15, and the B channel to 0, to complete the mapping of the motion feature parameter corresponding to the current target face key point. Then update the target face key point and repeat the above process to map the motion feature parameter corresponding to the updated target face key point until the mapping of the motion feature parameters corresponding to all face key points is completed, and a motion texture image is obtained.
[0071] Optionally, performing image blurring processing on the motion texture image may include: using a graphics processing unit (GPU) to perform a Gaussian convolution operation on the motion texture image to obtain a blurred image.
[0072] In this embodiment, in order to perform parallel processing on multiple texture images to achieve the effect of driving the historical driven face image frames in real time at high speed, the GPU is used to perform image blurring processing on the motion texture image. Blurring processing can be understood as taking the average value of the surrounding pixel points for each pixel point. Since the image is continuous, the closer the pixel points are, the closer their relationship is, and the farther the pixel points are, the more distant their relationship is. Therefore, the weight of the pixel points closer in distance can be set to be larger, and the weight of the pixel points farther in distance can be set to be smaller. Thus, when performing a Gaussian convolution operation on the motion texture image through the GPU, an image blurred according to the distance weight value can be obtained. Through blurring, the purpose of eliminating side face distortion and readjusting the motion feature parameters can be achieved.
[0073] Step 130: Generate a motion feature adjustment parameter set according to the blurred image, and generate the current driven face image frame according to the motion feature adjustment parameter set and the historical driven face image frames.
[0074] Optionally, generating a motion feature adjustment parameter set according to the blurred image may include: determining, according to the position correspondence, each pixel point corresponding to the face key points in the blurred image; generating motion feature adjustment parameters corresponding to each face key point according to the channel values of the R channel and the G channel of each pixel point; and combining the motion feature adjustment parameters to obtain a motion feature adjustment parameter set.
[0075] In this embodiment, after performing image blurring processing on the motion texture image through the GPU, according to the position correspondence relationship, each pixel point corresponding to the face key points is found from the blurred image, and the channel value of the R channel of each pixel point is used as the adjusted horizontal motion change amount of the face key point corresponding to this pixel point, and the channel value of the G channel is used as the adjusted vertical motion change amount of the face key point corresponding to this pixel point, so as to obtain the motion feature adjustment parameters corresponding to each face key point respectively, and a set of motion feature adjustment parameters is obtained by combining each motion feature adjustment parameter.
[0076] In this embodiment, after generating the set of motion feature adjustment parameters, each motion feature adjustment parameter in the set of motion feature adjustment parameters is respectively superimposed on the corresponding face key points in the historical driven face image frame, to obtain the coordinates of the face key points after real-time expression transformation following the current driven face image frame, and generate the current driven face image frame.
[0077] In the embodiment of the present invention, according to the current driven face image frame and the adjacent historical driven face image frame, a set of motion feature parameters corresponding to each face key point in the historical driven face is determined. By mapping each motion feature parameter in the set of motion feature parameters to the texture image, a motion texture image is obtained, and image blurring processing is performed on the motion texture image. Then, a set of motion feature adjustment parameters is generated according to the blurred image, and according to the set of motion feature adjustment parameters and the historical driven face image frame, the current driven face image frame is generated, which solves the problem of face distortion in the driven virtual face image in the prior art, and realizes that when the driven face image follows the driving face image for corresponding face expression adjustment, the face in the adjusted driven face image can be avoided from being distorted.
[0078] Embodiment 2
[0079] Figure 2a It is a flowchart of an image processing method in Embodiment 2 of the present invention. This embodiment can be combined with each optional solution in the above embodiments. Specifically, referring to Figure 2a , the method may include the following steps:
[0080] Step 210, in the current driven face image frame, the historical driven face image frame, and the historical driven face image frame, respectively obtain the face key points matching the face area.
[0081] In this embodiment, in order to enable the historical driven face in the historical driven face image frame to follow the corresponding action effect of the historical driving face in the historical driving face image frame, it is necessary to first obtain the face key points matching the face area in the current driving face image frame, the historical driving face image frame, and the historical driven face image frame, and then determine the corresponding relationship between the face key points in the historical driving face image frame and the current driving face image frame, as well as the corresponding relationship between the face key points in the historical driving face image frame and the historical driven face image frame. Only then can the coordinate changes of the face key points in the historical driven face image frame that match the expression changes of the driving face be found.
[0082] Step 220: Divide the historical driven face image frame into a plurality of original grids according to the face key points in the historical driven face image frame.
[0083] In this embodiment, using the face key points in the historical driven face image frame as at least some of the vertices of the grid, the historical driven face image frame is meshed and divided into two or more grids, as Figure 2b shown.
[0084] Step 230: Determine the target fixed points corresponding to the current driving face image frame, the historical driving face image frame, and the historical driven face image frame respectively.
[0085] Optionally, determining the target fixed points corresponding to the current driving face image frame, the historical driving face image frame, and the historical driven face image frame respectively may include: determining face positioning rectangles in the current driving face image frame, the historical driving face image frame, and the historical driven face image frame respectively; and obtaining the corner points at the same position in each face positioning rectangle as the target fixed points.
[0086] Optionally, determining the face positioning rectangles in the current driving face image frame, the historical driving face image frame, and the historical driven face image frame respectively may include: obtaining the eye distance and the nose tip key point of each face in the current driving face image frame, the historical driving face image frame, and the historical driven face image frame respectively; constructing the face positioning rectangles corresponding to the current driving face image frame, the historical driving face image frame, and the historical driven face image frame respectively with the product of the eye distance and the first ratio value as the length, the product of the eye distance and the second ratio value as the width, and the nose tip key point as the center point; where the eye distance and the nose tip key point are determined by the face key points in the corresponding face image frame.
[0087] Exemplarily, taking the current driving face image frame as an example, as Figure 2cAs shown, by obtaining the facial key point coordinates of the two eyes among the facial key points in the current driving facial image frame and calculating the difference, the eye distance E of the current driving facial image frame can be obtained. According to experience, the first proportional value can be set to 2, and the second proportional value can be set to 2.5, so as to obtain a facial positioning rectangle with a length of 2*E, a width of 2.5*E, and the center point being the nose tip key point S, as Figure 2d shown. Of course, the first proportional value and the second proportional value can also be set to other values.
[0088] Step 240: Generate a set of motion feature parameters according to the facial key points in the current driving facial image frame, the facial key points in the historical driving facial image frame, the target fixed points corresponding to the current driving facial image frame, and the target fixed points corresponding to the historical driving facial image frame.
[0089] Optionally, generating a set of motion feature parameters according to the facial key points in the current driving facial image frame, the facial key points in the historical driving facial image frame, the target fixed points corresponding to the current driving facial image frame, and the target fixed points corresponding to the historical driving facial image frame may include: obtaining the first position vectors between the facial key points in the current driving facial image frame and the corresponding target fixed points, and the second position vectors between the facial key points in the historical driving facial image frame and the corresponding target fixed points; calculating the vector differences between the second position vectors and the corresponding first position vectors; calculating the product of each vector difference and the facial scaling ratio to obtain the facial key point acceleration vectors matching the facial key points in the historical driving facial image frame; decomposing each facial key point acceleration vector along the horizontal and vertical directions to obtain the horizontal motion change amount and the vertical motion change amount corresponding to each facial key point as the set of motion feature parameters.
[0090] Exemplarily, as Figure 2d shown, assuming that both the current driving facial image frame and the historical driving facial image frame use the vertex A of the facial positioning rectangle as the target fixed point, obtaining the first position vector X n between each current driving facial key point X n and the target fixed point A, and the second position vector X n ' between each historical driving facial key point X n ' and the target fixed point A. By calculating X n A - X n 'A, the coordinate change amount of each facial key point corresponding to the historical driving face when changing from the historical driving facial image frame to the current driving facial image frame is obtained. Considering that the sizes of the historical driving face and the historical driven face are inconsistent, it is necessary to calculate according to the formula a n =(X n A - X n'A)*Q calculates the acceleration vectors of each facial key point so that the calculated coordinate changes of each facial key point are suitable for the size of the historical driven face, where Q is the face scaling ratio. Then, the acceleration vectors of each facial key point are decomposed along the x-axis direction and the y-axis direction respectively to obtain the horizontal motion change amount and the vertical motion change amount corresponding to each facial key point, which are used as the motion feature parameter set.
[0091] Optionally, before calculating the product of the difference of each vector and the face scaling ratio, it may further include: taking the quotient obtained by dividing the eye distance of the historical driven face image frame by the eye distance of the current driven face image frame as the face scaling ratio.
[0092] In this embodiment, since the historical driven face image frame is a static image, the eye distance of the historical driven face will not change. When the size of the driven face in front of the camera changes, in order to ensure that the calculated coordinate changes of each facial key point conform to the size of the historical driven face, the quotient of the eye distance of the historical driven face image frame and the eye distance of the current driven face image frame can be used as the face scaling ratio to scale the coordinate changes of each facial key point.
[0093] Step 250: Map each motion feature parameter in the motion feature parameter set to the texture image to obtain a motion texture image, and perform image blurring processing on the motion texture image.
[0094] In this embodiment, in order to facilitate data smoothing of each motion feature parameter and enable the facial image frame adjusted according to the motion feature parameter set to achieve a blurring effect, during the live video texture rendering process, in addition to processing the texture images of the current driven face image frame and the historical driven face image frame, a texture image for processing motion feature parameters can be bound at the same time. After mapping each motion feature parameter to this texture image, a Gaussian convolution operation is performed on this texture image through the GPU to achieve data smoothing of each motion feature parameter, so as to eliminate side face distortion and readjust the motion feature parameters.
[0095] Step 260: Generate a motion feature adjustment parameter set according to the blurred image, and generate the current driven face image frame according to the motion feature adjustment parameter set and the historical driven face image frame.
[0096] Optionally, adjusting the parameter set and the historical driven face image frames according to the motion features to generate the current driven face image frame may include: establishing a blank image matching the historical driven face image frame; adjusting the parameter set according to the motion features to determine the mesh deformation mode of each original mesh in the historical driven face image frame, where the historical driven face image frame is composed of multiple original meshes divided by each face key point; dividing a plurality of target deformation meshes corresponding to the original meshes in the blank image according to the mesh deformation mode; and mapping each pixel point in the original mesh to the corresponding target deformation mesh according to the position correspondence between the original mesh and the target deformation mesh to obtain the current driven face image frame.
[0097] In this embodiment, after determining each face key point of the adjusted historical driven face image frame according to the motion features to adjust the parameter set, a blank image with the same size as the historical driven face image frame may be established, so as to facilitate subsequent division of the target deformation meshes corresponding to the original meshes in the blank image, and after deforming and adjusting each original mesh in the historical driven face image frame, displaying the adjusted historical driven face in the blank image to obtain the current driven face image frame.
[0098] In this embodiment, in order to accelerate the processing speed of the historical driven face image frame, after dividing a plurality of target deformation meshes corresponding to the original meshes, directly according to the mapping relationship between the original mesh and the target deformation mesh, the pixel points in each original mesh are sequentially mapped to the corresponding target deformation mesh to obtain the current driven face image frame, without re-rendering the pixel points in the original mesh to the target deformation mesh.
[0099] In the embodiment of the present invention, according to the current driven face image frame and the adjacent historical driven face image frames, a motion feature parameter set corresponding to each face key point in the historical driven face is determined. By mapping each motion feature parameter in the motion feature parameter set to the texture image, a motion texture image is obtained, and the motion texture image is subjected to image blurring processing. Then, a motion feature adjustment parameter set is generated according to the blurred image, and the current driven face image frame is generated according to the motion feature adjustment parameter set and the historical driven face image frame, solving the problem of face deformation and distortion in the virtual face image after being driven in the prior art, and realizing that when the driven face image follows the driving face image for corresponding face expression adjustment, the face in the adjusted driven face image can be avoided from being distorted and deformed.
[0100] Embodiment III
[0101] Figure 3It is a schematic structural diagram of an image processing device in Embodiment 3 of the present invention. This embodiment is applicable to the situation where a historical driven face follows a driving face to perform real-time expression transformation. As Figure 3 shown, the image processing device includes:
[0102] A motion feature parameter set generation module 310, configured to determine a motion feature parameter set corresponding to each face key point in the historical driving face according to the current driving face image frame and the adjacent historical driving face image frames;
[0103] An image blurring processing module 320, configured to map each motion feature parameter in the motion feature parameter set to a texture image to obtain a motion texture image, and perform image blurring processing on the motion texture image;
[0104] A face driving module 330, configured to generate a motion feature adjustment parameter set according to the blurred image, and generate a current driven face image frame according to the motion feature adjustment parameter set and the historical driven face image frame.
[0105] In the embodiment of the present invention, according to the current driving face image frame and the adjacent historical driving face image frames, a motion feature parameter set corresponding to each face key point in the historical driving face is determined. By mapping each motion feature parameter in the motion feature parameter set to a texture image, a motion texture image is obtained, and image blurring processing is performed on the motion texture image. Then, a motion feature adjustment parameter set is generated according to the blurred image, and a current driven face image frame is generated according to the motion feature adjustment parameter set and the historical driven face image frame, which solves the problem of face deformation and distortion in the virtual face image after being driven in the prior art, and realizes that when the driven face image follows the driving face image to perform corresponding face expression adjustment, the face in the adjusted driven face image can be avoided from being distorted and deformed.
[0106] Optionally, the image blurring processing module 320 includes: a position correspondence obtaining unit, configured to obtain the position correspondence between each face key point in the historical driving face image frame and the pixel points in the fragment shader of the texture image; a mapping unit, configured to map the motion feature parameters corresponding to each face key point to the texture image respectively according to the position correspondence to obtain a motion texture image.
[0107] Optionally, the motion feature parameters corresponding to the facial key points include: horizontal motion change amount and vertical motion change amount; the mapping unit is specifically configured to: obtain a target horizontal motion change amount and a target vertical motion change amount corresponding to the target facial key point currently mapped; obtain a target pixel point matching the target facial key point in the texture image according to the position correspondence relationship; assign the R channel of the target pixel point to the target horizontal motion change amount, the G channel to the target vertical motion change amount, and the B channel to 0.
[0108] Optionally, the image blurring processing module 320 is specifically configured to: perform a Gaussian convolution operation on the motion texture image using a graphics processing unit (GPU) of an image processor to obtain a blurred image.
[0109] Optionally, the face driving module 330 is specifically configured to: determine, in the blurred image, pixel points corresponding to the facial key points respectively according to the position correspondence relationship; generate motion feature adjustment parameters corresponding to the facial key points respectively according to the channel values of the R channel and the G channel of each pixel point; and combine the motion feature adjustment parameters to obtain a motion feature adjustment parameter set.
[0110] Optionally, it further includes a target fixed point determination module, configured to: before determining a motion feature parameter set corresponding to each facial key point in the historical driven face according to the current driven face image frame and the adjacent historical driven face image frame, determine target fixed points corresponding to the current driven face image frame, the historical driven face image frame, and the historical driven face image frame respectively;
[0111] The motion feature parameter set generation module 310 is specifically configured to: generate a motion feature parameter set according to each facial key point in the current driven face image frame, each facial key point in the historical driven face image frame, the target fixed point corresponding to the current driven face image frame, and the target fixed point corresponding to the historical driven face image frame.
[0112] Optionally, the motion feature parameter set generation module 310 is specifically configured to: obtain a first position vector between each facial key point in the current driven face image frame and the corresponding target fixed point, and a second position vector between each facial key point in the historical driven face image frame and the corresponding target fixed point; calculate the vector difference between each second position vector and the corresponding first position vector; calculate the product of each vector difference and the face scaling ratio to obtain a facial key point acceleration vector matching each facial key point in the historical driven face image frame; decompose each facial key point acceleration vector along the horizontal direction and the vertical direction to obtain a horizontal motion change amount and a vertical motion change amount corresponding to each facial key point, as the motion feature parameter set.
[0113] Optionally, the face driving module 330 is specifically configured to: establish a blank image that matches the historical driven face image frame; adjust the parameter set according to the motion features to determine the mesh deformation mode of each original mesh in the historical driven face image frame, where the historical driven face image frame is composed of multiple original meshes divided by each face key point in it; divide multiple target deformed meshes corresponding to the original meshes in the blank image according to the mesh deformation mode; map each pixel point in the original mesh to the corresponding target deformed mesh according to the position correspondence between the original mesh and the target deformed mesh to obtain the current driven face image frame.
[0114] The image processing device provided by the embodiments of the present invention can execute the image processing method provided by any embodiment of the present invention, and has the corresponding functional modules and beneficial effects for executing the method.
[0115] Embodiment 4
[0116] Figure 4 It is a schematic structural diagram of a device disclosed in Embodiment 4 of the present invention. Figure 4 It shows a block diagram of an exemplary device 12 suitable for implementing the embodiments of the present invention. Figure 4 The shown device 12 is only an example and should not impose any limitation on the functions and usage scope of the embodiments of the present invention.
[0117] As Figure 4 shown, the device 12 is presented in the form of a general-purpose computing device. The components of the device 12 may include, but are not limited to: one or more processors or processing units 16, a system memory 28, and a bus 18 connecting different system components (including the system memory 28 and the processing unit 16).
[0118] The bus 18 represents one or more of several types of bus structures, including a memory bus or a memory controller, a peripheral bus, a graphics acceleration port, a processor, or a local bus using any of the multiple bus structures. For example, these architectures include, but are not limited to, Industry Standard Architecture (ISA) bus, Micro Channel Architecture (MAC) bus, Enhanced ISA bus, Video Electronics Standards Association (VESA) local bus, and Peripheral Component Interconnect (PCI) bus.
[0119] The device 12 typically includes a variety of computer system-readable media. These media can be any available media accessible by the device 12, including volatile and non-volatile media, removable and non-removable media.
[0120] System memory 28 may include computer system readable media in the form of volatile memory, such as random access memory (RAM) 30 and / or cache memory 32. Device 12 may further include other removable / non-removable, volatile / non-volatile computer system storage media. By way of example only, storage system 34 can be used for reading and writing on non-removable, non-volatile magnetic media ( Figure 4 not shown, typically referred to as a "hard disk drive"). Although Figure 4 not shown in the figure, a disk drive for reading and writing on removable non-volatile disks (such as "floppy disks") and an optical disk drive for reading and writing on removable non-volatile optical disks (such as CD-ROM, DVD-ROM or other optical media) can be provided. In these cases, each drive can be connected to bus 18 through one or more data media interfaces. Memory 28 may include at least one program product having a set (e.g., at least one) of program modules configured to perform the functions of the embodiments of the present invention.
[0121] A program / utility 40 having a set (at least one) of program modules 42 can be stored, for example, in memory 28. Such program modules 42 include, but are not limited to, an operating system, one or more application programs, other program modules, and program data. Each or some combination of these examples may include an implementation of a network environment. Program modules 42 generally perform the functions and / or methods in the embodiments described in the present invention.
[0122] Device 12 can also communicate with one or more external devices 14 (such as a keyboard, a pointing device, a display 24, etc.), and can also communicate with one or more devices that enable a user to interact with the device 12, and / or communicate with any device that enables the device 12 to communicate with one or more other computing devices (such as a network card, a modem, etc.). Such communication can be carried out through an input / output (I / O) interface 22. Moreover, device 12 can also communicate with one or more networks (such as a local area network (LAN), a wide area network (WAN), and / or a public network, such as the Internet) through a network adapter 20. As shown in the figure, network adapter 20 communicates with other modules of device 12 through bus 18. It should be understood that although not shown in the figure, other hardware and / or software modules can be used in combination with device 12, including but not limited to: microcode, device drivers, redundant processing units, external disk drive arrays, RAID systems, tape drives, and data backup storage systems, etc.
[0123] Processing unit 16 performs various functional applications and data processing by running programs stored in system memory 28, such as implementing the image processing method provided by the embodiments of the present invention.
[0124] That is to say, an image processing method is implemented, including: determining a set of motion feature parameters corresponding to each facial key point in the historical driven face according to the current driven face image frame and the adjacent historical driven face image frames; mapping each motion feature parameter in the set of motion feature parameters to a texture image to obtain a motion texture image, and performing image blurring processing on the motion texture image; generating a set of motion feature adjustment parameters according to the blurred image, and generating the current driven face image frame according to the set of motion feature adjustment parameters and the historical driven face image frames.
[0125] Embodiment 5
[0126] Embodiment 5 of the present invention also discloses a computer storage medium, on which a computer program is stored. When the program is executed by a processor, an image processing method is implemented, including: determining a set of motion feature parameters corresponding to each facial key point in the historical driven face according to the current driven face image frame and the adjacent historical driven face image frames; mapping each motion feature parameter in the set of motion feature parameters to a texture image to obtain a motion texture image, and performing image blurring processing on the motion texture image; generating a set of motion feature adjustment parameters according to the blurred image, and generating the current driven face image frame according to the set of motion feature adjustment parameters and the historical driven face image frames.
[0127] The computer storage medium of the embodiments of the present invention can adopt any combination of one or more computer-readable media. The computer-readable medium can be a computer-readable signal medium or a computer-readable storage medium. The computer-readable storage medium can be, for example, but not limited to, an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any combination of the above. More specific examples (non-exhaustive list) of the computer-readable storage medium include: an electrical connection having one or more wires, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the above. In this document, the computer-readable storage medium can be any tangible medium that contains or stores a program, which can be used by or in combination with an instruction execution system, apparatus, or device.
[0128] A computer-readable signal medium may include a data signal propagated in a baseband or as part of a carrier wave, which carries computer-readable program code. Such a propagated data signal may take various forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination of the foregoing. A computer-readable signal medium may also be any computer-readable medium other than a computer-readable storage medium, which can send, propagate, or transmit a program for use by or in connection with an instruction execution system, apparatus, or device.
[0129] The program code contained on a computer-readable medium may be transmitted using any appropriate medium, including but not limited to wireless, wire, optical fiber cable, RF, etc., or any suitable combination of the foregoing.
[0130] The computer program code for performing the operations of the present invention may be written in one or more programming languages or combinations thereof. The programming languages include object-oriented programming languages such as Java, Smalltalk, C++, and also include conventional procedural programming languages such as the "C" language or similar programming languages. The program code may be executed entirely on the user's computer, partially on the user's computer, executed as a stand-alone software package, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In the case of a remote computer, the remote computer may be connected to the user's computer through any type of network, including a local area network (LAN) or a wide area network (WAN), or may be connected to an external computer (e.g., through the Internet using an Internet service provider).
[0131] Note that the above is only the preferred embodiment of the present invention and the technical principles applied. Those skilled in the art will understand that the present invention is not limited to the specific embodiments described herein. Various obvious changes, re-adjustments, and substitutions can be made by those skilled in the art without departing from the protection scope of the present invention. Therefore, although the present invention has been described in detail through the above embodiments, the present invention is not limited to the above embodiments. Without departing from the concept of the present invention, more other equivalent embodiments may be included, and the scope of the present invention is determined by the scope of the appended claims.
Claims
1. An image processing method, characterized in that, comprising: determining a set of motion feature parameters corresponding to each face key point in the historical driven face according to the current driven face image frame and the adjacent historical driven face image frame; mapping each motion feature parameter in the set of motion feature parameters to a texture image to obtain a motion texture image, and performing image blurring processing on the motion texture image; generating a set of motion feature adjustment parameters according to the blurred image, and generating a current driven face image frame according to the set of motion feature adjustment parameters and the historical driven face image frame; The generating a current driven face image frame according to the set of motion feature adjustment parameters and the historical driven face image frame includes: establishing a blank image matching the historical driven face image frame; determining the grid deformation mode of each original grid in the historical driven face image frame according to the set of motion feature adjustment parameters; wherein, the historical driven face image frame is composed of a plurality of original grids divided by each face key point; dividing a plurality of target deformed grids corresponding to the original grids in the blank image according to the grid deformation mode; mapping each pixel point in the original grid to the corresponding target deformed grid according to the position correspondence between the original grid and the target deformed grid to obtain the current driven face image frame.
2. The method according to claim 1, characterized in that, mapping each motion feature parameter in the set of motion feature parameters to a texture image to obtain a motion texture image, including: obtaining the position correspondence between each face key point in the historical driven face image frame and the pixel points in the fragment shader of the texture image; mapping the motion feature parameter corresponding to each face key point to the texture image respectively according to the position correspondence to obtain the motion texture image.
3. The method according to claim 2, characterized in that, the motion feature parameters corresponding to the face key points include: horizontal motion change amount and vertical motion change amount; mapping the motion feature parameter corresponding to each face key point to the texture image respectively according to the position correspondence, including: obtaining a target horizontal motion change amount and a target vertical motion change amount corresponding to the currently mapped target face key point; obtaining a target pixel point matching the target face key point in the texture image according to the position correspondence; assigning the R channel of the target pixel point to the target horizontal motion change amount, the G channel to the target vertical motion change amount, and the B channel to 0.
4. The method according to claim 1, characterized in that, performing image blurring processing on the motion texture image, including: performing a Gaussian convolution operation on the motion texture image using a graphics processing unit (GPU) to obtain a blurred image.
5. The method according to claim 3, characterized in that, generating a set of motion feature adjustment parameters according to the blurred image, including: determining each pixel point corresponding to the face key points respectively in the blurred image according to the position correspondence; Generate motion feature adjustment parameters corresponding to each of the face key points according to the channel values of the R channel and the G channel of each pixel point; Combine the motion feature adjustment parameters to obtain the motion feature adjustment parameter set.
6. The method according to claim 1, wherein, before determining the motion feature parameter set corresponding to each face key point in the historical driven face according to the current driven face image frame and the adjacent historical driven face image frame, further comprising: Determine target fixed points corresponding to the current driven face image frame, the historical driven face image frame, and the historical driven face image frame respectively; Determining the motion feature parameter set corresponding to each face key point in the historical driven face according to the current driven face image frame and the adjacent historical driven face image frame includes: Generate a motion feature parameter set according to each face key point in the current driven face image frame, each face key point in the historical driven face image frame, the target fixed point corresponding to the current driven face image frame, and the target fixed point corresponding to the historical driven face image frame.
7. The method according to claim 6, wherein, Generating a motion feature parameter set according to each face key point in the current driven face image frame, each face key point in the historical driven face image frame, the target fixed point corresponding to the current driven face image frame, and the target fixed point corresponding to the historical driven face image frame includes: Obtain a first position vector between each face key point in the current driven face image frame and the corresponding target fixed point, and a second position vector between each face key point in the historical driven face image frame and the corresponding target fixed point; Calculate the vector difference between each of the second position vectors and the corresponding first position vector; Calculate the product of each vector difference and the face scaling ratio to obtain a face key point acceleration vector matching each face key point in the historical driven face image frame; Decompose each face key point acceleration vector along the horizontal and vertical directions to obtain the horizontal motion change amount and the vertical motion change amount corresponding to each face key point as the motion feature parameter set.
8. An image processing apparatus, wherein, comprising: A motion feature parameter set generation module, configured to determine a motion feature parameter set corresponding to each face key point in the historical driven face according to the current driven face image frame and the adjacent historical driven face image frame; An image blurring processing module, configured to map each motion feature parameter in the motion feature parameter set to a texture image to obtain a motion texture image, and perform image blurring processing on the motion texture image; A face driving module, configured to generate a motion feature adjustment parameter set according to the blurred image, and generate a current driven face image frame according to the motion feature adjustment parameter set and the historical driven face image frame; The face driving module is specifically configured to: establish a blank image that matches the historical driven face image frames; adjust a parameter set according to motion features to determine the mesh deformation modes of each original mesh in the historical driven face image frames, where the historical driven face image frames are composed of multiple original meshes divided by each face key point in the historical driven face image frames; divide a plurality of target deformation meshes corresponding to the original meshes in the blank image according to the mesh deformation modes. Map each pixel point in the original mesh to the corresponding target deformation mesh according to the position correspondence between the original mesh and the target deformation mesh to obtain the current driven face image frame.
9. An image processing device Characterized in that The device includes: One or more processors; A storage device for storing one or more programs, When the one or more programs are executed by the one or more processors, the one or more processors implement the image processing method according to any one of claims 1-7.
10. A computer-readable storage medium, on which a computer program is stored, Characterized in that When the program is executed by a processor, it implements the image processing method according to any one of claims 1-7.
Citation Information
Patent Citations
Image processing method and device, mobile terminal and storage medium
CN109788190A