An image generation method, apparatus and electronic device based on action transfer
By identifying key points and dividing the template object into meshes, and adjusting the key points of the reference object, motion transfer images with consistent poses are generated, solving the problem of high production costs for 2D character graffiti animation and realizing automatic animation generation.
Patent Information
- Application Number
- CN202211431643.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-11-15
- Publication Date
- 2025-10-31
- Estimated Expiration
- 2042-11-15
AI Technical Summary
Existing 2D character graffiti animation production costs are high, and traditional methods are time-consuming and labor-intensive.
By acquiring the image and key points of the template object, performing mask division and mesh transformation, and adjusting the key points of the reference object, motion transfer image generation is achieved.
It enables automatic animation generation, reducing production costs.
Smart Images

Figure CN115908858B_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of information technology, and in particular to an image generation method, apparatus and electronic device based on motion transfer. Background Technology
[0002] As living standards continue to improve, people are paying more and more attention to their spiritual and cultural lives. Animation, as a popular industry, is becoming increasingly mature. There are generally two methods for creating 2D character doodle animation: frame-by-frame animation, where each frame is hand-drawn and then pieced together; and skeletal animation, where the character is divided into different parts, each of which must be complete, and then the positions of different parts in each frame are manually adjusted to achieve the animation effect. However, regardless of the method, the production cost is very high. Summary of the Invention
[0003] The purpose of this application is to provide an image generation method, apparatus, and electronic device based on motion transfer, so as to reduce the production cost of animation. The specific technical solution is as follows:
[0004] In a first aspect of this application, an image generation method based on action transfer is provided, the method comprising:
[0005] Obtain the first image of the template object and multiple key points of the template object in the first image;
[0006] Based on the key points of each template object, a template object mask of the template object is determined in the first image;
[0007] The template object mask is meshed, and the key points of the template object associated with the vertices of each mesh are determined;
[0008] Obtain key points of multiple reference objects in the second image;
[0009] Based on the reference object key points, adjust the mapping position of each template object key point so that the relative positions between the reference object key points and the key points of the template object in the same location are the same.
[0010] For each vertex in each of the aforementioned meshes, the mapping position of the vertex is determined based on the mapping position of the key point of the template object associated with that vertex;
[0011] Based on the mapping positions of each vertex, an affine transformation is performed on each mesh of the template object mask to obtain an action transfer image of the template object whose pose is the same as that of the reference object in the second image.
[0012] In one possible implementation, the method further includes:
[0013] Obtain multi-frame motion transfer images of a template object whose pose is the same as that of the reference object in multiple frames of images, and obtain a set of motion transfer images;
[0014] Animation of the template object is generated according to the time sequence of the multi-frame images based on the motion migration image set.
[0015] In one possible implementation, acquiring multiple reference object key points in the second image includes:
[0016] Obtain a second image of the reference object sent by the client; perform keypoint detection on the reference object in the second image to obtain multiple reference object keypoints; or
[0017] Obtain key points of multiple reference objects in the second image from the reference objects sent by the client.
[0018] In one possible implementation, determining the template object mask of the template object in the first image based on each of the template object key points includes:
[0019] Based on the key points of each template object, the template object in the first image is segmented to obtain each mask component of the template object; wherein, each mask component includes at least one key point of the template object;
[0020] The various mask components are combined into a template object mask of the template object.
[0021] In one possible implementation, the template object is a person template object, and the segmentation of the template object in the first image based on the key points of each template object to obtain the mask components of the template object includes:
[0022] Based on the key points of each template object, the human character template object in the first image is segmented into a head, left hand, right hand, body, left foot, and right foot mask components;
[0023] The various mask components are merged into a template object mask for the character template object.
[0024] In one possible implementation, the template object is a humanoid template object, and the segmentation of the template object in the first image based on the key points of each template object to obtain the mask components of the template object includes:
[0025] Based on the key points of each template object, the anthropomorphic template object in the first image is segmented into a head, left hand, right hand, body, left foot, right foot, and tail mask components;
[0026] The various mask components are merged into a template object mask for the anthropomorphic template object.
[0027] In one possible implementation, determining the template object keypoints associated with the vertices of each of the meshes includes:
[0028] For the vertices of the mesh of the mask components other than the head mask components, determine the N template object keypoints that are closest to the vertex as the template object keypoints associated with the vertex, where N is an integer greater than 1;
[0029] For each vertex of the mesh of the head mask component, the template object keypoint that is closest to that vertex is determined as the template object keypoint associated with that vertex.
[0030] In one possible implementation, before determining the template object keypoints associated with the vertices of each of the meshes, the method further includes:
[0031] Expand the template object keypoints between at least one set of adjacent template object keypoints.
[0032] In one possible implementation, calculating the mapping position of each template object key point based on the reference object key points includes:
[0033] In each of the key points of the template object, determine the sub-key points whose positions need to be adjusted and the parent key point corresponding to each sub-key point;
[0034] For each sub-keypoint, calculate the first relative rotation angle of the sub-keypoint relative to its parent keypoint based on the reference keypoint.
[0035] Keeping the distance between the key points of each template object constant, and using the preset key points of the template objects as a reference, the mapping position of each key point of the template objects is calculated according to the first relative rotation angle.
[0036] In one possible implementation, the template object is a character template object, and the difference between the display angle of the reference object's key points and the display angle of the template object in the first image is less than a preset angle error. The key points of the template object include joint key points.
[0037] The step of determining, among the key points of each template object, the sub-key points whose positions need to be adjusted and the parent key point corresponding to each sub-key point includes:
[0038] Among the key points of each template object, select joint key points other than the neck, left shoulder, right shoulder, left hip, middle of hip, and right hip to obtain sub-key points;
[0039] For each sub-keypoint, among the joint keypoints that are closer to the middle keypoint of the buttock in terms of connection relationship, select the joint keypoint that is closest to the sub-keypoint in terms of connection relationship to obtain the parent keypoint of the sub-keypoint.
[0040] In one possible implementation, the template object is a humanoid template object, and the difference between the display angle of the reference object key points and the display angle of the template object in the first image is less than a preset angle error. The key points of the template object include joint key points.
[0041] The step of determining, among the key points of each template object, the sub-key points whose positions need to be adjusted and the parent key point corresponding to each sub-key point includes:
[0042] Among the key points of each template object, sub-key points are obtained from the joint key points other than the neck, left shoulder, right shoulder, left hip, middle of the hip, right hip, and sacrum.
[0043] For each sub-keypoint, among the joint keypoints that are closer to the middle keypoint of the buttock in terms of connection relationship, select the joint keypoint that is closest to the sub-keypoint in terms of connection relationship to obtain the parent keypoint of the sub-keypoint.
[0044] In one possible implementation, determining the mapping position of each vertex based on the mapping position of the keypoints of the template object associated with that vertex includes:
[0045] For each vertex, obtain the second relative rotation angle between the mapping position of the vertex and the associated key point of the vertex, the first relative rotation angle of the associated key point of the vertex, and the third relative rotation angle between the associated key point of the vertex and the parent key point of the associated key point of the vertex in the first image, wherein the associated key point of the vertex is the template object key point associated with the vertex.
[0046] The mapping position of the vertex is calculated based on the second relative rotation angle of the vertex, the first relative rotation angle of the associated key point of the vertex, and the third relative rotation angle of the vertex.
[0047] In one possible implementation, the step of performing an affine transformation on each mesh of the template object mask based on the mapping positions of each of the vertices to obtain an action transfer image of the template object whose pose is the same as that of the reference object in the second image includes:
[0048] For each grid, the affine transformation matrix of the grid is obtained based on the mapping positions of each vertex of the grid;
[0049] According to the corresponding affine transformation matrix, the points of each grid in the template object mask are subjected to affine transformation to obtain the motion transfer image of the template object with the same pose as the reference object in the second image.
[0050] In one possible implementation, the step of performing an affine transformation on the points of each grid in the template object mask according to the corresponding affine transformation matrix to obtain an action transfer image of the template object whose pose is the same as that of the reference object in the second image includes:
[0051] For each grid, the grid depth information is calculated based on the keypoint depth information of the keypoints associated with the vertices of that grid.
[0052] Following the order of decreasing mesh depth information, the meshes in the template object mask are selected sequentially. Using the corresponding affine transformation matrix, the points of the currently selected mesh are affinely transformed until all the meshes have been affinely transformed, resulting in an action transfer image of the template object whose pose is the same as that of the reference object in the second image.
[0053] In a second aspect of this application, an image generation method based on motion transfer is provided, the method being applied to a client, the method comprising:
[0054] Get the first image of the template object and the second image of the reference object;
[0055] Based on the first image and the second image, motion transfer data is generated, wherein the motion transfer data includes the first image and the second image, or the motion transfer data includes the first image, multiple template object key points of the template object in the first image, and multiple reference object key points of the reference object in the second image.
[0056] Send the motion transfer data to the server so that the server can generate a motion transfer image of a template object whose pose is the same as that of the reference object in the second image by using any of the methods described in the first aspect of the embodiments of this application;
[0057] Receive the motion migration image sent by the server.
[0058] In a third aspect of this application, an image generation apparatus based on motion transfer is provided, the apparatus comprising:
[0059] The first image key point acquisition module is used to acquire the first image of the template object and multiple key points of the template object in the first image.
[0060] A mask determination module is used to determine the template object mask of the template object in the first image based on the key points of each template object;
[0061] The mesh generation module is used to divide the template object mask into meshes and determine the key points of the template object associated with the vertices of each mesh.
[0062] The second image key point acquisition module is used to acquire multiple key points of the reference object in the second image.
[0063] The key point position adjustment module is used to adjust the mapping position of each key point of the template object according to the key points of the reference object, so that the relative positions between the key points of the reference object and the key points of the template object in the same part are the same.
[0064] The vertex position determination module is used to adjust the mapping position of each vertex in each of the meshes according to the mapping position of the key point of the template object associated with the vertex.
[0065] The affine transformation module is used to perform affine transformations on each mesh of the template object mask according to the mapping positions of each vertex, so as to obtain an action transfer image of the template object whose pose is the same as that of the reference object in the second image.
[0066] In one possible implementation, the device further includes:
[0067] An animation production module is used to acquire multi-frame motion transfer images of a template object whose posture is the same as that of a reference object in multiple frames of images, to obtain a motion transfer image set; and to generate an animation of the template object according to the time sequence of the multi-frame images based on the motion transfer image set.
[0068] In one possible implementation, the second image key point acquisition module includes:
[0069] The second image acquisition submodule is specifically used to acquire a second image of the reference object sent by the client; perform key point detection on the reference object in the second image to obtain multiple reference object key points; or
[0070] The reference object key point acquisition submodule is specifically used to acquire multiple reference object key points in the second image of the reference object sent by the client.
[0071] In one possible implementation, the mask determination module includes:
[0072] The template object segmentation submodule is specifically used to segment the template object in the first image based on the key points of each template object to obtain each mask component of the template object; wherein, each mask component includes at least one key point of the template object;
[0073] The mask merging submodule is specifically used to merge the various mask components into a template object mask of the template object.
[0074] In one possible implementation, the template object is a character template object, and the template object is divided into sub-modules, including:
[0075] The character template object segmentation unit is specifically used to segment the character template object in the first image into head, left hand, right hand, body, left foot, and right foot mask components based on the key points of each template object;
[0076] The character template object mask merging unit is specifically used to merge each of the mask components into a template object mask of the character template object.
[0077] In one possible implementation, the template object is a humanoid template object, and the template object is divided into sub-modules, including:
[0078] The anthropomorphic template object segmentation unit is specifically used to segment the anthropomorphic template object in the first image into a head, left hand, right hand, body, left foot, right foot, and tail mask component based on the key points of each template object;
[0079] The anthropomorphic template object mask merging unit is specifically used to merge each of the mask components into a template object mask of the anthropomorphic template object.
[0080] In one possible implementation, the mesh generation module includes:
[0081] Other mask partitioning submodules are specifically used to determine the N template object key points that are closest to the vertex of the mesh of the other mask components besides the head mask components as the template object key points associated with the vertex, where N is an integer greater than 1;
[0082] The head mask partitioning submodule is specifically used to determine the template object key point that is closest to the vertex of the mesh of the head mask components as the template object key point associated with the vertex.
[0083] In one possible implementation, the device further includes:
[0084] The keypoint expansion module is used to expand the keypoints of a template object between at least one set of adjacent keypoints.
[0085] In one possible implementation, the key point position adjustment module includes:
[0086] The parent key point determination sub-module is specifically used to determine, among the key points of each template object, each sub-key point whose position needs to be adjusted and the parent key point corresponding to each sub-key point;
[0087] The first relative rotation angle calculation submodule is specifically used to calculate the first relative rotation angle of each sub-key point relative to its parent key point, based on the reference object key point.
[0088] The key point position calculation submodule is specifically used to keep the distance between the key points of each template object unchanged, and to calculate the mapping position of each key point of each template object based on the preset key points of the template object and the first relative rotation angle.
[0089] In one possible implementation, the template object is a character template object, and the difference between the display angle of the reference object's key points and the display angle of the template object in the first image is less than a preset angle error. The key points of the template object include joint key points.
[0090] The parent keypoint determines the sub-module, including:
[0091] The joint key point selection unit is specifically used to select joint key points other than the neck, left shoulder, right shoulder, left hip, middle of hip, and right hip from the key points of each template object to obtain sub-key points;
[0092] The parent keypoint determination unit is specifically used to select the joint keypoint that is closest to the middle keypoint of the buttock in terms of connection relationship among the joint keypoints that are closer to the child keypoint in terms of connection relationship for each child keypoint, and obtain the parent keypoint of the child keypoint.
[0093] In one possible implementation, the template object is a humanoid template object, and the difference between the display angle of the reference object key points and the display angle of the template object in the first image is less than a preset angle error. The key points of the template object include joint key points.
[0094] The parent keypoint determines the sub-module, including:
[0095] The anthropomorphic joint key point selection unit is specifically used to obtain sub-key points from the joint key points other than the neck, left shoulder, right shoulder, left hip, middle of hip, right hip, and sacrum among the key points of each template object;
[0096] The anthropomorphic parent key point determination unit is specifically used to select the joint key point that is closest to the child key point in terms of connection relationship among the joint key points that are closer to the middle key point of the buttock in terms of connection relationship for each child key point, and obtain the parent key point of the child key point.
[0097] In one possible implementation, the vertex position determination module includes:
[0098] The vertex information acquisition submodule is specifically used to acquire, for each vertex, the second relative rotation angle between the mapping position of the vertex and the associated key point of the vertex, the first relative rotation angle of the associated key point of the vertex, and the third relative rotation angle between the associated key point of the vertex and the parent key point of the associated key point of the vertex in the first image, wherein the associated key point of the vertex is the template object key point associated with the vertex.
[0099] The vertex position calculation submodule is specifically used to calculate the mapped position of the vertex based on the second relative rotation angle of the vertex, the first relative rotation angle of the associated key point of the vertex, and the third relative rotation angle of the vertex.
[0100] In one possible implementation, the affine transformation module includes:
[0101] The affine transformation matrix acquisition submodule is specifically used to obtain the affine transformation matrix of each mesh based on the mapping positions of each vertex of the mesh.
[0102] The motion transfer submodule is specifically used to perform affine transformation on the points of each grid in the template object mask according to the corresponding affine transformation matrix, so as to obtain a motion transfer image of the template object with the same pose as the reference object in the second image.
[0103] In one possible implementation, the action transfer submodule includes:
[0104] The depth information calculation unit is specifically used to calculate the mesh depth information of each mesh based on the key point depth information of the key points associated with the vertices of that mesh.
[0105] The drawing unit is specifically used to sequentially select the meshes in the template object mask according to the mesh depth information from large to small, and use the corresponding affine transformation matrix to perform affine transformation on the points of the currently selected meshes until all the meshes have been affine transformed, so as to obtain the motion transfer image of the template object with the same pose as the reference object in the second image.
[0106] In a fourth aspect of this application, an image generation apparatus based on motion transfer is provided, the apparatus being applied to a client, the apparatus comprising:
[0107] The image acquisition module is used to acquire the first image of the template object and the second image of the reference object;
[0108] The motion migration data generation module is used to generate motion migration data based on the first image and the second image, wherein the motion migration data includes the first image and the second image, or the motion migration data includes the first image, multiple template object key points of the template object in the first image, and multiple reference object key points of the reference object in the second image.
[0109] The motion migration data sending module is used to send the motion migration data to the server so that the server can generate a motion migration image of a template object whose posture is the same as that of the reference object in the second image through the apparatus described in the third aspect of the present application.
[0110] In a fifth aspect of this application, an electronic device is provided, including a processor, a communication interface, a memory, and a communication bus, wherein the processor, the communication interface, and the memory communicate with each other through the communication bus;
[0111] Memory, used to store computer programs;
[0112] When a processor executes a program stored in memory, it implements the steps of the method described in any of the first aspects of the embodiments of this application.
[0113] In a sixth aspect of the present application, a computer-readable storage medium is provided, wherein a computer program is stored therein, and when the computer program is executed by a processor, it implements the steps of any of the methods described in the first aspect of the present application.
[0114] In another aspect of this application, a computer program product containing instructions is also provided, which, when run on a computer, causes the computer to execute any of the above-described motion-transfer-based image generation methods.
[0115] This application provides an image generation method, apparatus, and electronic device based on action transfer. The method involves: acquiring a first image of a template object and multiple template object keypoints in the first image; determining a template object mask in the first image based on each template object keypoint; dividing the template object mask into a mesh and determining the template object keypoints associated with the vertices of each mesh; acquiring multiple reference object keypoints of a reference object in a second image; adjusting the mapping positions of each template object keypoint according to the reference object keypoints so that the relative positions between the reference object keypoints and keypoints in the same location of the template object keypoints are the same; adjusting the mapping position of each vertex in each mesh according to the mapping position of the template object keypoints associated with that vertex; and performing an affine transformation on each mesh of the template object mask according to the mapping positions of each vertex to obtain an action transfer image of the template object with the same pose as the reference object in the second image. By applying the method of this application, key points of the template object can be identified and the mask can be meshed. Then, based on the key points of the reference object, the position transformation of the key points of the template object and the mesh vertices can be calculated to determine the final pose of the template object, thereby realizing the automatic generation of animation and ultimately reducing the animation production cost. Attached Figure Description
[0116] To more clearly illustrate the technical solutions in the embodiments of this application or the prior art, the accompanying drawings used in the description of the embodiments or the prior art will be briefly introduced below.
[0117] Figure 1 A flowchart illustrating a method for image generation based on motion migration applied to a server, as provided in this application embodiment.
[0118] Figure 2 This is a schematic diagram illustrating the determination of key points of a character template object, as provided in an embodiment of this application.
[0119] Figure 3 A flowchart detailing step S105 provided in the embodiments of this application is shown.
[0120] Figure 4 A flowchart illustrating a method for image generation based on action migration applied to a client, as provided in this application embodiment.
[0121] Figure 5 This is a schematic diagram of image generation based on action transfer provided in an embodiment of this application.
[0122] Figure 6 This is a schematic diagram of a device structure for image generation based on motion migration applied to a server, as provided in an embodiment of this application.
[0123] Figure 7 This is a schematic diagram of a key point position adjustment module provided in an embodiment of this application.
[0124] Figure 8 This is a schematic diagram of a device structure for image generation based on motion migration applied to a client, as provided in an embodiment of this application.
[0125] Figure 9 This is a schematic diagram of an electronic device applied to a server, as provided in an embodiment of this application.
[0126] Figure 10 This is a schematic diagram of a structure of an electronic device applied to a client, provided in an embodiment of this application. Detailed Implementation
[0127] The technical solutions in the embodiments of this application will now be described with reference to the accompanying drawings.
[0128] To reduce animation production costs, this application provides an image generation method based on motion transfer, comprising: acquiring a first image of a template object and multiple template object key points in the first image; determining a template object mask of the template object in the first image based on each template object key point; dividing the template object mask into a mesh and determining the template object key points associated with the vertices of each mesh; acquiring multiple reference object key points of a reference object in a second image; adjusting the mapping positions of each template object key point according to the reference object key points, so that the relative positions between the reference object key points and the key points of the template object in the same location are the same; for each vertex in each mesh, adjusting the mapping position of the vertex according to the mapping position of the template object key point associated with the vertex; and performing an affine transformation on each mesh of the template object mask according to the mapping positions of each vertex to obtain a motion transfer image of the template object with the same pose as the reference object in the second image. By applying the method of this application, key points of the template object can be identified and the mask can be meshed. Then, based on the key points of the reference object, the position transformation of the key points of the template object and the mesh vertices can be calculated to determine the final pose of the template object, thereby realizing the automatic generation of animation and ultimately reducing the animation production cost.
[0129] The following is a detailed explanation:
[0130] In a first aspect of this application, a method for image generation based on action transfer is provided, the method comprising: Figure 1 The steps shown are as follows:
[0131] Step S101: Obtain the first image of the template object and multiple key points of the template object in the first image.
[0132] The method described in this application can be implemented using a terminal device, which may be an electronic device such as a server, computer, or mobile phone.
[0133] The type of template object can be customized according to actual needs. For example, the template object can be a person, an animal, a humanoid cartoon character, or a humanoid 3D character. Obtaining multiple key points of the template objects in the first image can be achieved by performing key point detection on the template objects in the first image, thus obtaining multiple key points of the template objects. Alternatively, the key points of the template objects can also be sent from the client to the server. Key point detection can be implemented using relevant key point detection algorithms based on the type of the template object. For example, when the template object type is a person, a human key point detection algorithm can be used to detect the key points. In some scenarios, key points of the template objects can also be obtained through manual annotation based on manually input instructions.
[0134] Step S102: Based on the key points of each template object, determine the template object mask of the template object in the first image.
[0135] In practical applications, the range of the template object to be masked in the first image can be determined based on the key points of the template object. Then, a mask matrix can be generated by obtaining the pixel values of the image within the mask range. Based on the mask matrix, the template object in the first image can be pixel filtered to obtain the template object mask.
[0136] Step S103: Mesh the template object mask and determine the key points of the template object associated with the vertices of each mesh.
[0137] There are several ways to mesh the mask, such as using triangular meshes or quadrangular meshes. Each mesh has multiple vertices, which can be shared with other meshes. All meshes together form the template object mask. Each vertex is associated with at least one template object keypoint. The association method between vertices and template object keypoints can be customized according to the actual situation. For example, based on the distance from the vertex to the template object keypoint, one or more template object keypoints with the smallest distance from the vertex can be selected as the template object keypoints associated with that vertex; or the template object can be divided into multiple units, and the template object keypoints in the same unit as the vertex can be determined as the template object keypoints associated with that vertex.
[0138] Step S104: Obtain multiple reference object key points in the second image.
[0139] In practical applications, the key points of the reference object in the second image can be directly sent from the client to the server, or obtained by the server using a key point detection algorithm after acquiring the second image. The second image of the reference object can be obtained in various ways, such as by importing videos or images. The pose of the reference object in the second image is the pose that the template object needs to be converted into. Since the second image is obtained by capturing reference objects in a real scene (e.g., real people or real animals), the key points of the reference object can be obtained by recognizing the 3D (3-dimensional) pose of the reference object in the second image. The detection methods for reference object key points can refer to key point detection methods in related technologies. For example, pose detection can be performed using various video processing platforms, such as media pipes and Mocap (motion capture), to obtain the reference object key points.
[0140] Step S105: Based on the key points of the reference object, adjust the mapping position of the key points of each template object so that the relative positions of the key points of the same part in the reference object and the key points of the template object are the same.
[0141] The reference object and the template object can be of the same type, for example, both being humans; or the reference object can be a real human, and the template object can be an anthropomorphic human. Therefore, the mapping position of the template object's key points can be determined based on the position of the reference object's key points. For example, the mapping position of the template object's left elbow key point can be determined based on the position of the reference object's left elbow key point. Specifically, the second angle of the template object's left elbow key point relative to the left shoulder can be adjusted according to the first angle of the reference object's left elbow key point relative to the left shoulder, so that the first angle and the second angle are the same.
[0142] Step S106: For each vertex in each mesh, adjust the mapping position of the vertex according to the mapping position of the key point of the template object associated with the vertex.
[0143] Each vertex is associated with at least one keypoint of a template object. The mapping position of a vertex can be determined and adjusted based on the mapping positions of the keypoints of each template object associated with it, through methods such as weighted averaging and affine transformation.
[0144] Step S107: Perform affine transformation on each mesh of the template object mask according to the mapping position of each vertex to obtain the motion transfer image of the template object with the same pose as the reference object in the second image.
[0145] For each mesh, given the mapped positions of its vertices, each pixel within that mesh can be transformed using an affine transformation, thus completing the affine transformation of the mesh. After all meshes have undergone the affine transformation, the template object is drawn in the pose of the reference object. In practical applications, drawing can be performed on each mesh texture. Once all meshes are drawn, a motion transfer image of the template object with the same pose as the reference object in the second image is obtained.
[0146] By applying the method of this application embodiment, the template object mask can be meshed, and then the position transformation of the template object key points and mesh vertices can be calculated based on the key points of the reference object, thereby determining the final pose of the template object, realizing automatic animation generation, and ultimately reducing animation production costs.
[0147] The above embodiments provide a method for generating motion transition images of template objects from a single frame image. When there are multiple frames, the animation of the template object can be obtained based on the motion transition images of the template object in each frame. In one possible implementation, the method further includes:
[0148] Step 1: Obtain multi-frame motion transfer images of a template object whose pose is the same as that of the reference object in the multi-frame images, and obtain a set of motion transfer images.
[0149] The multi-frame images can be multiple video frames from the acquired video data of the reference object. In one example, the multi-frame object includes a second image. For each frame of the multi-frame images, a motion transfer image of the template object with the same pose as the reference object in that frame can be obtained. In practical applications, the motion transfer image generated by the template object based on the pose of the reference object in each frame can be obtained according to steps S101-S107 above, thereby obtaining the motion transfer image of the template object in the multi-frame images, which is called the motion transfer image set.
[0150] Step 2: Based on the motion transfer image set, generate the animation of the template object according to the time sequence of multiple frames.
[0151] After obtaining the painting collection, the animation of the template object is obtained by combining the motion transfer images of the template object with the temporal sequence of the multiple frames as the reference.
[0152] By applying the method of this application embodiment, motion transfer images corresponding to the pose of the template object under the reference object in multiple frames of images can be obtained to generate a motion transfer image set, and then an animation can be generated based on the motion transfer image set. This allows animation to be generated directly based on the reference object during animation production, reducing manual adjustments to the drawings and realizing automatic animation generation, thereby reducing animation production costs.
[0153] In one possible implementation, the step S104 described above, which involves acquiring multiple reference object key points in the second image, includes:
[0154] Step 1: Obtain the second image of the reference object sent by the client; perform keypoint detection on the reference object in the second image to obtain multiple reference object keypoints; or
[0155] Step 2: Obtain key points of multiple reference objects in the second image, as sent by the client.
[0156] By applying the method of the embodiments of this application, key points of a reference object can be obtained directly or indirectly. Based on the key points of the reference object, the mapping position of key points of a template object can be determined, and finally a motion transfer image with the same posture as the reference object can be obtained, thereby realizing the automatic generation of animation.
[0157] The template object mask is composed of mask components corresponding to each limb part of the template object. In one possible implementation, step S102 may include:
[0158] Step 1: Based on the key points of each template object, segment the template objects in the first image to obtain the mask components of each template object;
[0159] Step 2: Merge the various mask components into a template object mask for the template object.
[0160] Each mask component includes at least one template object key point; and the mask components can be divided according to the template object composition in the first image. For example, when the template object is a dog, the template object mask can be divided into mask components corresponding to the head, body, limbs and tail, etc., based on the body structure features of the dog.
[0161] By applying the method of the embodiments of this application, the mask of the template object is divided into masks of various parts, which can help determine the mapping position of the key points of the template object and make the obtained motion transfer image more accurate.
[0162] In one possible implementation, when the template object in the first image is a person template object, the segmentation step of the template object may include: based on the key points of each person template object, segmenting the person template object in the first image into head, left hand, right hand, body, left foot, and right foot mask components.
[0163] Among them, the key points of the character template object can be the key points of the joints that make up the character's skeleton, such as... Figure 2As shown, the key points of each character template object can be the joints corresponding to the nose (0), neck (1), right shoulder (2), right elbow (3), right wrist (4), left shoulder (5), left elbow (6), left wrist (7), middle of the buttocks (8), right buttock (9), right knee (10), right ankle (11), left buttock (12), left knee (13), and left ankle (14). Based on these key points, the character object can be divided into mask components for the head, left hand, right hand, body, left foot, and right foot. These mask components can be combined into a complete character template object mask.
[0164] In practical applications, the head mask can include the key points of the template object corresponding to the nose of the character; the left hand mask can include the key points of the template object corresponding to the left shoulder, left elbow, and left wrist of the character; the right hand mask can include the key points of the template object corresponding to the right shoulder, right elbow, and right wrist of the character; the body mask can include the key points of the template object corresponding to the middle of the neck and hip of the character; the left foot mask can include the key points of the template object corresponding to the left hip, left knee, and left ankle of the character; and the right foot mask can include the key points of the template object corresponding to the right hip, right knee, and right ankle of the character.
[0165] By applying the method of the embodiments of this application, the mask of the template object can be divided into multiple mask components, which can more accurately calculate the changing position of the template object, thereby improving the animation effect.
[0166] In addition, in animation, there is another form of anthropomorphizing non-human objects. Therefore, when the template object is an anthropomorphic template object, the above-mentioned template object segmentation steps may include: based on the key points of each anthropomorphic template object, segmenting the anthropomorphic template object in the first image into head, left hand, right hand, body, left foot, right foot and tail mask components.
[0167] Among them, anthropomorphic template objects include objects with human characteristics, such as those that can speak, have facial expressions, and perform human actions; anthropomorphic template objects can be animal or plant-based.
[0168] By applying the embodiments of this application, anthropomorphic animations can be automatically generated, thereby expanding the scope of automatically generated animations and reducing animation production costs.
[0169] Taking a mask component including a head mask component as an example, in one possible implementation, "determining the template object keypoints associated with the vertices of each of the meshes" in step S103 may include:
[0170] Step 1: For the vertices of the mesh of the mask components other than the head mask components, determine the N template object keypoints that are closest to the vertex as the template object keypoints associated with the vertex, where N is an integer greater than 1;
[0171] Step 2: For each vertex of the mesh that makes up the head mask, determine the template object keypoint that is closest to that vertex as the template object keypoint associated with that vertex.
[0172] In practical applications, since the head undergoes a rigid transformation, only one template object keypoint needs to be determined as the template object keypoint associated with the mesh vertex in the head mask. However, for other mask components besides the head mask, which undergo significant changes with character movement (such as affine transformations and tangential transformations), multiple template object keypoints need to be determined and associated with each vertex.
[0173] By applying the method of this application embodiment, key points of the associated template object corresponding to the mesh vertices can be set so that when the template object changes, the position transformation of the mesh vertices can be calculated based on the key points of the corresponding associated template object, thereby making the obtained transformed motion transfer image more accurate.
[0174] To further improve the painting effect of the template object, in one possible implementation, before determining the template object key points associated with the vertices of each mesh in step S103, the above method may further include: expanding the template object key points between at least one set of adjacent template object key points.
[0175] In practical applications, there are several ways to expand the key points of a template object. For example, you can randomly add key points to adjacent key points of the template object; you can also add key points by evenly distributing the distance between adjacent key points of the template object; or you can add key points to the template object by changing the posture of the action, adding more key points between adjacent key points of the template object where the action changes frequently.
[0176] By applying the method of this application embodiment, the key points of the template object can be made denser by expanding the key points of the template object, thereby making the motion transfer image calculated based on the key points of the template object more accurate.
[0177] The mapping positions of key points in the template object can be calculated using parent-child key points. In one possible implementation, step S105 may include, for example... Figure 3 The steps shown are as follows:
[0178] Step S301: In each template object key point, determine the sub-key points whose positions need to be adjusted and the parent key point corresponding to each sub-key point.
[0179] The key points whose positions need adjustment can be determined based on the specific structure of the template object. For example, when the template object is a dog, the key point corresponding to its head can be preset to remain unchanged, and the key points corresponding to the other parts are the child key points whose positions need adjustment. The correspondence between child key points and parent key points can be preset, and the principle is that the child key point rotates around the parent key point as an axis; in some examples, the same template object key point can be both a child key point and a parent key point, for example, with Figure 2 For example, template object key point 10 is a child key point of template object key point 9, and at the same time, template object key point 10 is the parent key point of template object key point 11.
[0180] Step S302: For each sub-keypoint, calculate the first relative rotation angle of the sub-keypoint relative to its parent keypoint based on the reference object keypoint.
[0181] The first relative rotation angle can be calculated based on the angle between the child keypoint and the parent keypoint. For example, taking a certain child keypoint as an example, first, using the parent keypoint of the child keypoint as the axis, obtain the angle α1 of the child keypoint relative to the parent keypoint. Figure 2 For example, the angle α1 of child keypoint 3 relative to parent keypoint 2 can be the angle between the vector pointing from keypoint 2 to keypoint 3 and a preset direction. Then, obtain the angle α2 of the child keypoint relative to the parent keypoint in the pose of the reference object, and then calculate the difference between the two angles to obtain the first relative rotation angle of the child keypoint relative to its parent keypoint in the pose of the reference object.
[0182] Step S303: Keep the distance between the key points of each template object unchanged, and use the preset key points of the template object as a reference to calculate the mapping position of each key point of the template object according to each first relative rotation angle.
[0183] In practical applications, the mapping position of each template object's key point can be calculated using the coordinate values of the key point, based on each first relative rotation angle.
[0184] X = x p +cos(θ1)·R
[0185] Y = y p +sin(θ1)·R
[0186] Where X represents the x-coordinate value of the sub-keypoint; Y represents the y-coordinate value of the sub-keypoint; x p This represents the x-coordinate of the parent keypoint corresponding to the child keypoint; yp θ1 represents the ordinate of the parent keypoint corresponding to the child keypoint; θ1 represents the first relative rotation angle of the child keypoint relative to the parent keypoint; R represents the distance from the child keypoint to the parent keypoint.
[0187] By applying the method of this application, the mapping position of the drawing key point can be calculated by determining the correspondence and coordinate values of the child key point and the parent key point, the first relative rotation angle, etc., so as to make the motion transfer image more accurate.
[0188] In some scenarios, to further improve the accuracy of automatically generated template objects, the display angles of the reference object in the second image and the template object in the first image can be made approximately the same, thereby reducing the situation where the automatically generated template object is severely deformed due to a large difference in their display angles. In one possible implementation, the difference between the display angle of the key points of the reference object and the display angle of the template object in the first image is less than a preset angle error, and the key points of the template object include joint key points. Step S301 may include:
[0189] Step 1: Among the key points of each template object, select the joint key points other than the neck, left shoulder, right shoulder, left hip, middle of the hip, and right hip to obtain sub-key points.
[0190] In practical applications, the preset angle error can be customized according to the actual situation. If the difference between the display angle of the key point of the reference object and the display angle of the template object in the first image is less than the preset angle error, it is assumed that the display angles of the reference object and the template object are basically the same. This ensures that the display angle of the template object in the animation production will not change significantly, and can reduce the situation where the automatically generated template object is severely deformed due to the large difference between the two display angles.
[0191] Since the reference object and the template object have basically the same display angle, it can be assumed that the joint key points related to the display angle do not change. Furthermore, considering that the proportion difference between the template object and the reference object is too large when making animations, it is necessary to fix the corresponding joint key points. For example, fix the relative rotation angle between the joint key points from the neck to the right shoulder, the joint key points from the neck to the left shoulder, the joint key points from the neck to the middle of the hip, the joint key points from the middle of the hip to the right hip, and the joint key points from the middle of the hip to the left hip.
[0192] Step 2: For each sub-keypoint, among the joint keypoints that are closer to the middle keypoint of the buttock in terms of connection relationship, select the joint keypoint that is closest to the sub-keypoint in terms of connection relationship, and obtain the parent keypoint of the sub-keypoint.
[0193] In practical applications, the rotation angles of other key points can be calculated in a tree-like structure, using the central key point in the buttocks as the center. For example... Figure 2 In the process of calculating the mapping position of keypoint 10, since keypoint 9 is keypoint 9 which is closer to the middle keypoint 8 of the hip, when keypoint 10 is a child keypoint, keypoint 9 is the parent keypoint corresponding to keypoint 10.
[0194] In one possible implementation, when the template object is a humanoid template object, the difference between the display angle of the reference object key points and the display angle of the template object in the first image can also be less than a preset angle error. The key points of the template object include humanoid joint key points.
[0195] Step S301 may include:
[0196] Step 1: Among the key points of each template object, select the joint key points other than the neck, left shoulder, right shoulder, left hip, middle of the hip, right hip, and sacrum to obtain sub-key points;
[0197] Step 2: For each sub-keypoint, among the joint keypoints that are closer to the middle keypoint of the buttock in terms of connection relationship, select the joint keypoint that is closest to the sub-keypoint in terms of connection relationship, and obtain the parent keypoint of the sub-keypoint.
[0198] By applying the method of this application embodiment, the relative rotation angle between some joint key points can be fixed, and the correspondence between parent key points and child key points can be determined, thereby facilitating the calculation of the position transformation of key points of template objects.
[0199] To facilitate the calculation of the mapping positions of each mesh vertex, in one possible implementation, step S106 may include:
[0200] Step 1: For each vertex, obtain the second relative rotation angle between the mapping position of the vertex and the associated key point of the vertex, the first relative rotation angle of the associated key point of the vertex, and the third relative rotation angle between the associated key point of the vertex and the parent key point of the associated key point of the vertex in the first image, wherein the associated key point of the vertex is the template object key point associated with the vertex.
[0201] Step 2: Calculate the mapped position of the vertex based on the second relative rotation angle of the vertex, the first relative rotation angle of the associated key point of the vertex, and the third relative rotation angle of the vertex.
[0202] In practical applications, the mapped position of a vertex can be calculated using the following formula:
[0203]
[0204]
[0205] Where x represents the x-coordinate of the vertex; y represents the y-coordinate of the vertex; N represents the set of keypoints associated with the vertex; x n This represents the x-coordinate value of the nth associated key node; y n θ represents the ordinate value of the nth associated key joint; θ2 represents the second relative rotation angle between the mapped positions of this vertex and its associated key joints; θ n This represents the first relative rotation angle between the vertex and the nth associated key point of the vertex; This represents the third relative rotation angle between the vertex and its associated keypoint in the first image, and between the vertex's parent keypoint and its associated keypoint; r represents the distance from the vertex to the associated keypoint; w n This represents the weight of the nth associated keypoint, where the weight is determined by the distance from the associated keypoint to the vertex. For example, the smaller the distance from the associated keypoint to the vertex, the greater the weight of the associated keypoint.
[0206] By applying the method of this application embodiment, the corresponding positions of the key points of the template object and the mesh vertices after the pose transformation can be calculated based on the relationship and coordinate values of the key points and mesh vertices of each template object, thereby obtaining the motion transfer image after the pose transformation, reducing the production cost of animation.
[0207] After obtaining the mapped positions of each mesh vertex, in one possible implementation, step S107 may include:
[0208] Step 1: For each grid, obtain the affine transformation matrix of the grid based on the mapping positions of each vertex of the grid;
[0209] Step 2: According to the corresponding affine transformation matrix, perform affine transformation on the points of each grid in the template object mask to obtain the motion transfer image of the template object with the same pose as the reference object in the second image.
[0210] Each grid contains multiple pixels. The affine transformation matrix of the grid can be calculated based on the position transformation of each pixel in the grid. The calculation of the affine transformation matrix can be referred to the relevant existing technology, and will not be described in detail in this application.
[0211] To achieve better animation effects, one possible implementation is to draw based on the depth information of the mesh. Specifically,
[0212] For each grid, the grid depth information can be calculated based on the keypoint depth information of the keypoints associated with the vertices of that grid. In descending order of grid depth information, the grids in the template object mask are selected sequentially, and the points of the currently selected grid are affinely transformed using the corresponding affine transformation matrix until all grids have been affinely transformed, resulting in an action transfer image of the template object whose pose is the same as that of the reference object in the second image.
[0213] The grid depth information represents the distance from each point in the image to the camera, and is usually represented in grayscale. For example, a two-dimensional function can be used to give the depth information of each coordinate point in grayscale, where the value of each pixel is the depth of that point.
[0214] By applying the method of this application embodiment, the final drawing can be obtained by performing affine transformation on each grid vertex, and the grid can be drawn in an orderly manner according to the depth information of each grid, so that the final motion transfer image has a more 3D effect.
[0215] In a second aspect of this application, an image generation method based on motion transfer is also provided, applied to a client. The method includes, for example: Figure 4 The steps shown are as follows:
[0216] Step S401: Obtain the first image of the template object and the second image of the reference object.
[0217] Step S402: Generate motion transfer data based on the first image and the second image.
[0218] The motion transfer data includes a first image and a second image, or the motion transfer data includes a first image, multiple template object key points in the first image, and multiple reference object key points in the second image.
[0219] Step S403: Send motion migration data to the server so that the server can generate a motion migration image of a template object whose pose is the same as that of the reference object in the second image by using any of the motion migration-based image generation methods applied to the server in this application embodiment.
[0220] Step S404: Receive the motion migration image sent by the server.
[0221] By applying the method of this application embodiment, motion transfer data can be sent to the server, which can then perform mesh division on the template object mask. Based on the key points of the reference object, the position transformation of the key points of the template object and the mesh vertices can be calculated, and the motion transfer image of the template object can be determined and output. This achieves automatic animation generation and ultimately reduces animation production costs.
[0222] The following example illustrates the action transfer-based image generation method of this application. Figure 5 As shown, Figure 5 The template object is a graffiti character, and its first image is of a character with both hands outstretched, such as... Figure 5 As shown in picture ②, its final posture is with both hands falling down, as... Figure 5 As shown in image ⑨, to obtain image ⑨, this application employs the following steps:
[0223] The final pose ① of the reference object in the video input is obtained, and pose recognition is performed on it to obtain image ⑦; in this example, the 3D pose of the person in each frame of the video is recognized by Mediepipe.
[0224] Pose recognition is performed on the graffiti figure in the first image to obtain the template object key points, resulting in image ②. The template object key points are then expanded to obtain image ⑤. In this example, three template object key points are added between adjacent template object key points to evenly distribute the distance between the template object key points in image ②. The graffiti pose recognition model trained by the detectron2 (object detection platform) framework identifies the 2D (2-dimensional) reference object key points of the graffiti figure.
[0225] The graffiti semantic segmentation model trained using the detectron2 framework identified the masks of six parts of the graffiti character: head, left hand, right hand, body, left foot, and right foot, resulting in image ③.
[0226] The various mask components in image ③ are merged into a complete character mask, and this mask is then meshed to obtain image ④; in this example, triangular meshing is used.
[0227] For each mesh vertex, determine the template object keypoints that are associated with it to obtain image ⑥; in this example, for each mesh vertex in the head mask, determine one template object keypoint that is associated with that vertex; for each mesh vertex in each mask other than the head mask, determine four template object keypoints that are associated with that vertex.
[0228] Based on the key points of the template object in the final pose in image ⑦, calculate the final transformation position of each mesh vertex to obtain image ⑧;
[0229] Finally, based on the depth information of the template object, all the triangular mesh textures in the mask are drawn in sequence to obtain the graffiti character ⑨ in the current frame.
[0230] In a third aspect of this application, an image generation apparatus based on motion transfer is provided, the apparatus comprising: Figure 6 As shown:
[0231] The first image key point acquisition module 601 is used to acquire a first image of a template object and multiple key points of the template object in the first image.
[0232] The mask determination module 602 is used to determine the template object mask of the template object in the first image based on the key points of each template object;
[0233] The mesh generation module 603 is used to perform mesh generation on the template object mask and determine the key points of the template object associated with the vertices of each mesh.
[0234] The second image key point acquisition module 604 is used to acquire multiple key points of the reference object in the second image.
[0235] The key point position adjustment module 605 is used to adjust the mapping position of each key point of the template object according to the key points of the reference object, so that the relative positions between the key points of the reference object and the key points of the template object in the same part are the same.
[0236] Vertex position determination module 606 is used to adjust the mapping position of each vertex in each of the meshes according to the mapping position of the key point of the template object associated with the vertex.
[0237] The affine transformation module 607 is used to perform affine transformation on each mesh of the template object mask according to the mapping position of each vertex, so as to obtain an action transfer image of the template object with the same pose as the reference object in the second image.
[0238] The apparatus of this application embodiment can automatically generate animation by identifying key points of the template object and dividing the mask into meshes, and then calculating the position transformation of the key points of the template object and the mesh vertices based on the key points of the reference object, thereby determining the final pose of the template object and reducing the cost of animation production.
[0239] In one possible implementation, the device further includes:
[0240] The animation production module is used to obtain multi-frame motion transfer images of a template object whose posture is the same as that of a reference object in multiple frames of images, and obtain a set of motion transfer images; based on the set of motion transfer images, the animation of the template object is generated according to the time sequence of the multi-frame images.
[0241] The apparatus of this application embodiment can generate a motion transfer image set by using multi-frame motion transfer images of template objects whose postures are the same as those of reference objects in multi-frame images. Animation can then be generated based on the drawing set. This allows animation to be generated directly based on reference objects during animation production, reducing manual adjustments to the drawings and achieving automatic animation generation, thereby reducing animation production costs.
[0242] In one possible implementation, the second image key point acquisition module 604 includes:
[0243] The second image acquisition submodule is specifically used to acquire a second image of the reference object sent by the client; perform key point detection on the reference object in the second image to obtain multiple reference object key points; or
[0244] The reference object key point acquisition submodule is specifically used to acquire multiple reference object key points in the second image of the reference object sent by the client.
[0245] In one possible implementation, the mask determination module 602 may include:
[0246] The template object segmentation submodule is specifically used to segment the template object in the first image based on the key points of each template object to obtain each mask component of the template object; wherein, each mask component includes at least one key point of the template object;
[0247] The Mask Merging submodule is specifically used to merge the various mask components into a template object mask of the template object.
[0248] In one possible implementation, the template object is a character template object, and the template object is divided into sub-modules, which may include:
[0249] The character template object segmentation unit is specifically used to segment the character template object in the first image into head, left hand, right hand, body, left foot, and right foot mask components based on the key points of each character template object;
[0250] The Character Template Object Mask Merging Unit is specifically used to merge the various mask components into a template object mask for the character template object.
[0251] In one possible implementation, the template object is a humanoid template object, and the template object is divided into sub-modules, including:
[0252] The anthropomorphic template object segmentation unit is specifically used to segment the anthropomorphic template object in the first image into a head, left hand, right hand, body, left foot, right foot, and tail mask component based on the key points of each template object;
[0253] The anthropomorphic template object mask merging unit is specifically used to merge each of the mask components into a template object mask of the anthropomorphic template object.
[0254] The apparatus of this application embodiment can divide the mask of the template object into multiple mask components, and can more accurately calculate the changing position of the template object, thereby improving the animation effect.
[0255] In one possible implementation, the mesh generation module 603 may include:
[0256] Other mask partitioning submodules are specifically used to determine the N template object key points that are closest to the vertex of the mesh of the other mask components besides the head mask components as the template object key points associated with the vertex, where N is an integer greater than 1;
[0257] The head mask partitioning submodule is specifically used to determine the template object keypoint that is closest to the vertex of the mesh that constitutes the head mask as the template object keypoint associated with that vertex.
[0258] The apparatus of this application embodiment can set key points of associated template objects corresponding to mesh vertices, so that when the template object changes, the position transformation of the mesh vertices can be calculated based on the corresponding key points of associated template objects, thereby making the posture of the transformed template object more accurate.
[0259] In one possible implementation, the apparatus may further include:
[0260] The keypoint expansion module is used to expand the keypoints of a template object between at least one set of adjacent keypoints.
[0261] The apparatus of this application embodiment can make the key points of the template object denser by expanding the key points of the template object, thereby making the pose of the template object calculated based on the key points of the template object more accurate.
[0262] In one possible implementation, the key point position adjustment module 605 may include, for example: Figure 7 As shown:
[0263] The parent key point determination submodule 701 is specifically used to determine the child key points whose positions need to be adjusted and the parent key point corresponding to each child key point in the key points of each template object.
[0264] The first relative rotation angle calculation submodule 702 is specifically used to calculate the first relative rotation angle of each sub-key point relative to its parent key point, based on the reference object key point.
[0265] The key point position calculation submodule 703 is specifically used to keep the distance between the key points of each template object unchanged, and to calculate the mapping position of each key point of the template object based on the preset key points of the template object and the first relative rotation angle.
[0266] Using the apparatus of this application, the mapping position of the drawing key point can be calculated by determining the correspondence and coordinate values of the child key point and the parent key point, the first relative rotation angle, etc., so that the posture of the template object after transformation is more accurate.
[0267] In one possible implementation, the difference between the display angle of the reference object's pose and the display angle of the template object in the first image is less than a preset angle error, and the key points of the template object include joint key points.
[0268] The parent keypoint determines the submodule 701, which may include:
[0269] The joint key point selection unit is specifically used to select joint key points other than the neck, left shoulder, right shoulder, left hip, middle of hip, and right hip from the key points of each template object to obtain sub-key points;
[0270] The parent keypoint determination unit is specifically used to select the joint keypoint that is closest to the middle keypoint of the buttock in terms of connection relationship among the joint keypoints that are closer to the child keypoint in terms of connection relationship for each child keypoint, and obtain the parent keypoint of the child keypoint.
[0271] In one possible implementation, the template object is a humanoid template object, and the difference between the display angle of the reference object key points and the display angle of the template object in the first image is less than a preset angle error. The key points of the template object include joint key points.
[0272] The parent keypoint determines the submodule 701, including:
[0273] The anthropomorphic joint key point selection unit is specifically used to select joint key points other than the neck, left shoulder, right shoulder, left hip, middle of the hip, right hip, and sacrum from the key points of each template object to obtain sub-key points;
[0274] The anthropomorphic parent key point determination unit is specifically used to select the joint key point that is closest to the child key point in terms of connection relationship among the joint key points that are closer to the middle key point of the buttock in terms of connection relationship for each child key point, and obtain the parent key point of the child key point.
[0275] The apparatus of this application embodiment can determine the relative rotation angle between some joint key points and the correspondence between parent key points and child key points, thereby facilitating the calculation of the position transformation of key points of template objects.
[0276] In one possible implementation, the vertex position determination module 606 may include:
[0277] The vertex information acquisition submodule is specifically used to acquire, for each vertex, the second relative rotation angle between the mapping position of the vertex and the associated key point of the vertex, the first relative rotation angle of the associated key point of the vertex, and the third relative rotation angle between the associated key point of the vertex and the parent key point of the associated key point of the vertex in the first image, wherein the associated key point of the vertex is the template object key point associated with the vertex.
[0278] The vertex position calculation submodule is specifically used to calculate the mapped position of the vertex based on the second relative rotation angle of the vertex, the first relative rotation angle of the associated key point of the vertex, and the third relative rotation angle of the vertex.
[0279] The apparatus of this application embodiment can calculate the corresponding positions of the key points and mesh vertices of the template object after the pose transformation based on the relationship and coordinate values of the key points and mesh vertices of each template object, thereby obtaining the pose of the template object after the pose transformation, reducing the production cost of animation.
[0280] In one possible implementation, the affine transformation module 607 may include:
[0281] The affine transformation matrix acquisition submodule is specifically used to obtain the affine transformation matrix of each mesh based on the mapping positions of each vertex of the mesh.
[0282] The motion transfer submodule is specifically used to perform affine transformation on the points of each grid in the template object mask according to the corresponding affine transformation matrix, so as to obtain the motion transfer image of the template object with the same pose as the reference object in the second image.
[0283] In one possible implementation, the action transition submodule includes:
[0284] The depth information calculation unit is specifically used to calculate the mesh depth information of each mesh based on the key point depth information of the key points associated with the vertices of that mesh.
[0285] The drawing unit is specifically used to select the meshes in the template object mask in descending order of mesh depth information, and use the corresponding affine transformation matrix to perform affine transformation on the points of the currently selected meshes until all meshes have been affine transformed, thereby obtaining the motion transfer image of the template object whose pose is the same as that of the reference object in the second image.
[0286] The apparatus of this application embodiment can obtain the final drawing by performing affine transformation on each grid vertex, and draw the grid in an orderly manner according to the depth information of each grid, so that the final drawing has a more 3D effect.
[0287] In a fourth aspect of this application, an image generation apparatus based on motion transfer is provided, such as... Figure 8 As shown, the device, applied to a client, includes:
[0288] Image acquisition module 801 is used to acquire a first image of a template object and a second image of a reference object;
[0289] The motion transfer data generation module 802 is used to generate motion transfer data based on the first image and the second image. The motion transfer data includes the first image and the second image, or the motion transfer data includes the first image, multiple template object key points of the template object in the first image, and multiple reference object key points of the reference object in the second image.
[0290] The motion migration data sending module 803 is used to send motion migration data to the server so that the server can generate a motion migration image of a template object whose posture is the same as that of the reference object in the second image through the apparatus of the third aspect of the present application.
[0291] The apparatus using the embodiments of this application can send motion transfer data to a server, which then performs mesh division on the template object mask. Based on the key points of the reference object, the position transformations of the key points of the template object and the mesh vertices are calculated, and the motion transfer image of the template object is determined and output. This achieves automatic animation generation and ultimately reduces animation production costs.
[0292] This application also provides an electronic device, such as... Figure 9 As shown, it includes a processor 901, a communication interface 902, a memory 903, and a communication bus 904, wherein the processor 901, the communication interface 902, and the memory 903 communicate with each other through the communication bus 904.
[0293] Memory 903 is used to store computer programs;
[0294] When processor 901 executes a program stored in memory 903, it performs the following steps:
[0295] Get the first image of the template object and multiple key points of the template object in the first image;
[0296] Based on the key points of each template object, the template object mask of the template object is determined in the first image;
[0297] Mesh the template object mask and determine the key points of the template object associated with the vertices of each mesh;
[0298] Obtain key points of multiple reference objects in the second image;
[0299] Based on the key points of the reference object, adjust the mapping position of the key points of each template object so that the relative positions of the key points in the same part of the reference object and the key points of the template object are the same.
[0300] For each vertex in each mesh, adjust the mapping position of the vertex according to the mapping position of the key points of the template object associated with that vertex;
[0301] Based on the mapping positions of each vertex, an affine transformation is performed on each mesh of the template object mask to obtain an action transfer image of the template object whose pose is the same as that of the reference object in the second image.
[0302] This application also provides an electronic device, such as... Figure 10 As shown, it includes a processor 1001, a communication interface 1002, a memory 1003, and a communication bus 1004, wherein the processor 1001, the communication interface 1002, and the memory 1003 communicate with each other through the communication bus 1004.
[0303] Memory 1003 is used to store computer programs;
[0304] When processor 1001 executes a program stored in memory 1003, it performs the following steps:
[0305] Get the first image of the template object and the second image of the reference object;
[0306] Based on the first image and the second image, motion transfer data is generated, wherein the motion transfer data includes the first image and the second image, or the motion transfer data includes the first image, multiple template object key points of the template object in the first image, and multiple reference object key points of the reference object in the second image.
[0307] Send motion transfer data to the server so that the server can generate a motion transfer image of a template object whose pose is the same as that of the reference object in the second image using any of the methods in the first aspect of the embodiments of this application;
[0308] Receive the motion migration image sent by the server.
[0309] The communication bus mentioned above can be a Peripheral Component Interconnect (PCI) bus or an Extended Industry Standard Architecture (EISA) bus, etc. This communication bus can be divided into address bus, data bus, control bus, etc. For ease of illustration, only one thick line is used to represent it in the diagram, but this does not mean that there is only one bus or one type of bus.
[0310] The communication interface is used for communication between the aforementioned terminal and other devices.
[0311] The memory may include random access memory (RAM) or non-volatile memory, such as at least one disk storage device. Optionally, the memory may also be at least one storage device located remotely from the aforementioned processor.
[0312] The processors mentioned above can be general-purpose processors, including central processing units (CPUs), network processors (NPs), etc.; they can also be digital signal processors (DSPs), application-specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs), or other programmable logic devices, discrete gate or transistor logic devices, or discrete hardware components.
[0313] In another embodiment provided in this application, a computer-readable storage medium is also provided, wherein a computer program is stored therein, and when the computer program is executed by a processor, it implements the image generation method based on action transfer described in any of the above embodiments.
[0314] In another embodiment provided in this application, a computer program product containing instructions is also provided, which, when run on a computer, causes the computer to execute any of the motion-transfer-based image generation methods described in the above embodiments.
[0315] In the above embodiments, implementation can be achieved entirely or partially through software, hardware, firmware, or any combination thereof. When implemented using software, it can be implemented entirely or partially in the form of a computer program product. The computer program product includes one or more computer instructions. When the computer program instructions are loaded and executed on a computer, all or part of the processes or functions described in the embodiments of this application are generated. The computer can be a general-purpose computer, a special-purpose computer, a computer network, or other programmable device. The computer instructions can be stored in a computer-readable storage medium or transmitted from one computer-readable storage medium to another. For example, the computer instructions can be transmitted from one website, computer, server, or data center to another website, computer, server, or data center via wired (e.g., coaxial cable, fiber optic, digital subscriber line (DSL)) or wireless (e.g., infrared, wireless, microwave, etc.) means. The computer-readable storage medium can be any available medium that a computer can access or a data storage device such as a server or data center that integrates one or more available media. The available medium can be a magnetic medium (e.g., floppy disk, hard disk, magnetic tape), an optical medium (e.g., DVD), or a semiconductor medium (e.g., solid state disk (SSD)).
[0316] It should be noted that, in this document, relational terms such as "first" and "second" are used only to distinguish one entity or operation from another, and do not necessarily require or imply any such actual relationship or order between these entities or operations. Furthermore, the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such a process, method, article, or apparatus. Without further limitations, an element defined by the phrase "comprising one..." does not exclude the presence of other identical elements in the process, method, article, or apparatus that includes said element.
[0317] The various embodiments in this specification are described in a related manner. Similar or identical parts between embodiments can be referred to mutually. Each embodiment focuses on describing the differences from other embodiments. In particular, the apparatus embodiments are basically similar to the method embodiments, so the description is relatively simple; relevant parts can be referred to the descriptions of the method embodiments.
[0318] The above description is merely a preferred embodiment of this application and is not intended to limit the scope of protection of this application. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of this application are included within the scope of protection of this application.
Claims
1. An image generation method based on action transfer, characterized in that, The method is applied to the server side, and the method includes: Obtain the first image of the template object and multiple key points of the template object in the first image; Based on the key points of each template object, a template object mask of the template object is determined in the first image; The template object mask is meshed, and the key points of the template object associated with the vertices of each mesh are determined; Obtain key points of multiple reference objects in the second image; Based on the reference object key points, adjust the mapping position of each template object key point so that the relative positions between the reference object key points and the key points of the template object in the same location are the same. For each vertex in each of the aforementioned meshes, adjust the mapping position of the vertex according to the mapping position of the key point of the template object associated with that vertex; Based on the mapping positions of each vertex, an affine transformation is performed on each mesh of the template object mask to obtain an action transfer image of the template object whose pose is the same as that of the reference object in the second image.
2. The method according to claim 1, characterized in that, The method further includes: Obtain multi-frame motion transfer images of a template object whose pose is the same as that of the reference object in multiple frames of images, and obtain a set of motion transfer images; Animation of the template object is generated according to the time sequence of the multi-frame images based on the motion migration image set.
3. The method according to claim 1, characterized in that, The acquisition of multiple reference object key points in the second image includes: Obtain a second image of the reference object sent by the client; perform keypoint detection on the reference object in the second image to obtain multiple reference object keypoints; or Obtain key points of multiple reference objects in the second image from the reference objects sent by the client.
4. The method according to claim 1, characterized in that, The step of determining the template object mask of the template object in the first image based on the key points of each template object includes: Based on the key points of each template object, the template object in the first image is segmented to obtain each mask component of the template object; wherein, each mask component includes at least one key point of the template object; The various mask components are combined into a template object mask of the template object.
5. The method according to claim 4, characterized in that, The template object is a character template object. Based on the key points of each template object, the template object in the first image is segmented to obtain the mask components of the template object, including: Based on the key points of each template object, the human character template object in the first image is segmented into a head, left hand, right hand, body, left foot, and right foot mask components; The various mask components are merged into a template object mask for the character template object.
6. The method according to claim 4, characterized in that, The template object is a humanoid template object. Based on the key points of each template object, the template object in the first image is segmented to obtain the mask components of the template object, including: Based on the key points of each template object, the anthropomorphic template object in the first image is segmented into a head, left hand, right hand, body, left foot, right foot, and tail mask components; The various mask components are merged into a template object mask for the anthropomorphic template object.
7. The method according to claim 4, characterized in that, The determination of the template object key points associated with the vertices of each of the meshes includes: For the vertices of the mesh of the mask components other than the head mask components, determine the N template object keypoints that are closest to the vertex as the template object keypoints associated with the vertex, where N is an integer greater than 1; For each vertex of the mesh of the head mask component, the template object keypoint that is closest to that vertex is determined as the template object keypoint associated with that vertex.
8. The method according to claim 7, characterized in that, Before determining the template object keypoints associated with the vertices of each of the meshes, the method further includes: Expand the template object keypoints between at least one set of adjacent template object keypoints.
9. The method according to claim 1, characterized in that, The step of adjusting the mapping position of each template object key point according to the reference object key points includes: In each of the key points of the template object, determine the sub-key points whose positions need to be adjusted and the parent key point corresponding to each sub-key point; For each sub-keypoint, calculate the first relative rotation angle of the sub-keypoint relative to its parent keypoint based on the reference keypoint. Keeping the distance between the key points of each template object constant, and using the preset key points of the template objects as a reference, the mapping position of each key point of the template objects is calculated according to the first relative rotation angle.
10. The method according to claim 9, characterized in that, The template object is a character template object. The difference between the display angle of the reference object's key points and the display angle of the template object in the first image is less than a preset angle error. The key points of the template object include joint key points. The step of determining, among the key points of each template object, the sub-key points whose positions need to be adjusted and the parent key point corresponding to each sub-key point includes: Among the key points of each template object, select joint key points other than the neck, left shoulder, right shoulder, left hip, middle of hip, and right hip to obtain sub-key points; For each sub-keypoint, among the joint keypoints that are closer to the middle keypoint of the buttock in terms of connection relationship, select the joint keypoint that is closest to the sub-keypoint in terms of connection relationship to obtain the parent keypoint of the sub-keypoint.
11. The method according to claim 9, characterized in that, The template object is a humanoid template object. The difference between the display angle of the reference object's key points and the display angle of the template object in the first image is less than a preset angle error. The key points of the template object include joint key points. The step of determining, among the key points of each template object, the sub-key points whose positions need to be adjusted and the parent key point corresponding to each sub-key point includes: Among the key points of each template object, select joint key points other than the neck, left shoulder, right shoulder, left hip, middle of hip, right hip, and sacrum to obtain sub-key points; For each sub-keypoint, among the joint keypoints that are closer to the middle keypoint of the buttock in terms of connection relationship, select the joint keypoint that is closest to the sub-keypoint in terms of connection relationship to obtain the parent keypoint of the sub-keypoint.
12. The method according to claim 9, characterized in that, For each vertex, determining the mapping position of the vertex based on the mapping position of the keypoints of the template object associated with that vertex includes: For each vertex, obtain the second relative rotation angle between the mapping position of the vertex and the associated key point of the vertex, the first relative rotation angle of the associated key point of the vertex, and the third relative rotation angle between the associated key point of the vertex and the parent key point of the associated key point of the vertex in the first image, wherein the associated key point of the vertex is the template object key point associated with the vertex. The mapping position of the vertex is calculated based on the second relative rotation angle of the vertex, the first relative rotation angle of the associated key point of the vertex, and the third relative rotation angle of the vertex.
13. The method according to claim 1, characterized in that, The step of performing an affine transformation on each mesh of the template object mask based on the mapping positions of each vertex to obtain an action transfer image of the template object whose pose is the same as that of the reference object in the second image includes: For each grid, the affine transformation matrix of the grid is obtained based on the mapping positions of each vertex of the grid; According to the corresponding affine transformation matrix, the points of each grid in the template object mask are subjected to affine transformation to obtain the motion transfer image of the template object with the same pose as the reference object in the second image.
14. The method according to claim 13, characterized in that, The step of performing an affine transformation on the points of each grid in the template object mask according to the corresponding affine transformation matrix to obtain an action transfer image of the template object whose pose is the same as that of the reference object in the second image includes: For each grid, the grid depth information is calculated based on the keypoint depth information of the keypoints associated with the vertices of that grid. Following the order of decreasing mesh depth information, the meshes in the template object mask are selected sequentially. Using the corresponding affine transformation matrix, the points of the currently selected mesh are affinely transformed until all the meshes have been affinely transformed, resulting in an action transfer image of the template object whose pose is the same as that of the reference object in the second image.
15. An image generation method based on action transfer, characterized in that, The method is applied to a client, and the method includes: Get the first image of the template object and the second image of the reference object; Based on the first image and the second image, motion transfer data is generated, wherein the motion transfer data includes the first image and the second image, or the motion transfer data includes the first image, multiple template object key points of the template object in the first image, and multiple reference object key points of the reference object in the second image. Send the motion transfer data to the server so that the server can generate a motion transfer image of a template object whose pose is the same as that of the reference object in the second image by using the method described in any one of claims 1-14; Receive the motion migration image sent by the server.
16. An image generation device based on motion transfer, characterized in that, The device is used on a server side, and the device includes: The first image key point acquisition module is used to acquire the first image of the template object and multiple key points of the template object in the first image. A mask determination module is used to determine the template object mask of the template object in the first image based on the key points of each template object; The mesh generation module is used to divide the template object mask into meshes and determine the key points of the template object associated with the vertices of each mesh. The second image key point acquisition module is used to acquire multiple key points of the reference object in the second image. The key point position adjustment module is used to adjust the mapping position of each key point of the template object according to the key points of the reference object, so that the relative positions between the key points of the reference object and the key points of the template object in the same part are the same. The vertex position determination module is used to adjust the mapping position of each vertex in each of the meshes according to the mapping position of the key point of the template object associated with the vertex. The affine transformation module is used to perform affine transformations on each mesh of the template object mask according to the mapping positions of each vertex, so as to obtain an action transfer image of the template object whose pose is the same as that of the reference object in the second image.
17. An image generation device based on motion transfer, characterized in that, The device is used on a client side, and the device includes: The image acquisition module is used to acquire the first image of the template object and the second image of the reference object; The motion migration data generation module is used to generate motion migration data based on the first image and the second image, wherein the motion migration data includes the first image and the second image, or the motion migration data includes the first image, multiple template object key points of the template object in the first image, and multiple reference object key points of the reference object in the second image. The motion migration data sending module is used to send the motion migration data to the server, so that the server can generate a motion migration image of a template object whose pose is the same as that of the reference object in the second image using the apparatus of claim 16. The motion migration image receiving module is used to receive the motion migration image sent by the server.
18. An electronic device, characterized in that, It includes a processor, a communication interface, a memory, and a communication bus, wherein the processor, the communication interface, and the memory communicate with each other through the communication bus; Memory, used to store computer programs; A processor, when executing a program stored in memory, implements the steps of the method described in any one of claims 1-15.
19. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores a computer program that, when executed by a processor, implements the steps of the method described in any one of claims 1-15.
Citation Information
Patent Citations
Three-dimensional animation generation method based on deep recurrent neural network algorithm
CN106971414A
Skeletal animation vertex correction method and model learning method, device and apparatus
CN113554736A