Motion video generation method and device, computer equipment and storage medium
By extracting and fusing two-dimensional and three-dimensional pose features in action videos and inputting video generation models, the problem of discontinuity and jitter during motion video generation in the prior art is solved, and the video generation quality is improved.
Patent Information
- Application Number
- CN202510495629.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-21
- Publication Date
- 2025-05-16
- Estimated Expiration
- 2045-04-21
AI Technical Summary
The motion video generated in the prior art is prone to problems of picture discontinuity and jitter.
By obtaining the target character image, target background and action-driven video, extracting the two-dimensional and three-dimensional pose features of each frame, fusing them, and inputting the video generation model to generate the target motion video.
This method reduces the probability of ambiguity action generation by simultaneously using three-dimensional motion information and two-dimensional motion information to guide actions, significantly improves the quality of video generation, and reduces picture discontinuity and jitter.
Smart Images

Figure CN120017917A_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of artificial intelligence technology, and in particular to a motion video generation method, device, computer equipment and storage medium. Background Art
[0002] With the development of artificial intelligence and image processing technology, it is now possible to generate a video by driving a character image through a specified target action sequence. This technology is relatively mature, such as Latent Image Animator (an image generator based on shallow coding). With the development of related technologies, the diffusion model is widely used, which has demonstrated powerful generation and editing capabilities in both image and video generation. Some researchers have also applied it to human posture migration, such as DreamPose, MagicDance, MagicAnimate, AnimateAnyone and MimicMotion, which are all algorithms for human posture migration based on diffusion models. These methods all use the pose information of the driving video as a condition and guide to drive the target image to be driven to generate the result video. The videos produced in the current mainstream framework are prone to discontinuity and jitter. Summary of the invention
[0003] The purpose of the present application is to solve at least one of the above-mentioned technical defects, especially the defect that the motion video generated in the prior art is prone to picture discontinuity and jitter.
[0004] In a first aspect, the present application provides a motion video generation method, comprising:
[0005] Obtain target person image, target background and action-driven video;
[0006] According to the action-driven video, the two-dimensional posture features and three-dimensional posture features corresponding to each frame are extracted respectively;
[0007] Fusing the two-dimensional posture feature and the three-dimensional posture feature to obtain a fused posture feature;
[0008] The target person image, target background and fused posture features are input into the video generation model to obtain the target motion video.
[0009] In one embodiment, the three-dimensional posture feature includes the rotation angles corresponding to the first number of joint points, and the two-dimensional posture feature and the three-dimensional posture feature are fused to obtain the fused posture feature, including:
[0010] The initial number of channels of the two-dimensional posture feature is expanded to a first number of times, and the two-dimensional posture feature is divided into sub-feature graphs corresponding to each joint point in the channel dimension, and a first mean and a first standard deviation corresponding to each channel of each sub-feature graph are determined;
[0011] Transform and map the three-dimensional posture feature to obtain the three-dimensional mapping posture feature, and determine the second mean and the second standard deviation corresponding to each joint point in the three-dimensional mapping posture feature; each row of the three-dimensional mapping posture feature corresponds to a joint point, each second number column corresponds to a rotation angle, and the second number is the ratio between the initial channel number and the number of types of rotation angles;
[0012] For any sub-feature map, a fused sub-feature map is obtained according to the corresponding first mean, first standard deviation, second mean and second standard deviation;
[0013] The fused posture feature is obtained according to each fused sub-feature map.
[0014] In one embodiment, obtaining a fused sub-feature map according to the corresponding first mean, first standard deviation, second mean, and second standard deviation includes:
[0015] The intermediate sub-feature map is obtained according to the first expression; the first expression is:
[0016]
[0017] in, is the intermediate sub-feature map, is the i-th sub-feature map, is the first mean, is the first standard deviation, is the second mean, is the second standard deviation;
[0018] The intermediate sub-feature map is weightedly summed with the corresponding sub-feature map to obtain a fused sub-feature map.
[0019] In one embodiment, the process of acquiring the target background includes:
[0020] Perform human body detection on the original background to obtain a human body area mask;
[0021] According to the human body region mask and the original background, the target background is obtained.
[0022] In one embodiment, before obtaining the target background according to the human body region mask and the original background, the method further includes:
[0023] The human body region mask is sequentially subjected to closing and dilation operations.
[0024] In one embodiment, the video generation model includes an encoder, a feature extraction module, a UNET network and a decoder, and inputs the target person image, the target background and the fusion posture features into the video generation model to obtain the target motion video, including:
[0025] Input the fused posture features into the downsampling layer of the UNET network;
[0026] Input the target person image into the encoder;
[0027] Add potential noise encoding to the target background and sum it with the output of the encoder; the summation result is input to the downsampling layer of the UNET network;
[0028] The target person image is input into the feature extraction module; the result of the feature extraction module is input into the downsampling layer and upsampling layer of the UNET network, the output of the upsampling layer of the UNET network is connected to the decoder, and the decoder is used to generate the target motion video according to the output of the upsampling layer of the UNET network.
[0029] In one embodiment, the training process of the video generation model includes:
[0030] According to the training action driving video, the training three-dimensional posture features are obtained;
[0031] The predicted three-dimensional posture features are obtained by using the video generation model to obtain the predicted motion video from the training action-driven video;
[0032] Obtaining a target loss value according to the predicted 3D posture features and the trained 3D posture features;
[0033] Tune the video generation model according to the target loss value.
[0034] In one embodiment, obtaining a target loss value according to the predicted three-dimensional posture feature and the training three-dimensional posture feature includes:
[0035] The target loss value is obtained according to the second expression, the predicted three-dimensional posture features and the trained three-dimensional posture features; the second expression is:
[0036]
[0037] in, is the target loss value, n is the number of joint points in the 3D posture feature, is the training 3D posture feature of the i-th joint point, is the predicted 3D posture feature of the i-th joint point.
[0038] In a second aspect, the present application provides a motion video generation device, comprising:
[0039] A data acquisition module is used to acquire the target person image, target background and action-driven video;
[0040] A posture feature extraction module is used to extract two-dimensional posture features and three-dimensional posture features corresponding to each frame according to the action-driven video;
[0041] A posture feature fusion module is used to fuse the two-dimensional posture feature and the three-dimensional posture feature to obtain a fused posture feature;
[0042] The video generation module is used to input the target person image, target background and fusion posture features into the video generation model to obtain the target motion video.
[0043] In a third aspect, the present application provides a computer device comprising one or more processors and a memory, wherein the memory stores computer-readable instructions, and when the computer-readable instructions are executed by the one or more processors, the steps of the motion video generation method in any of the above embodiments are performed.
[0044] In a fourth aspect, the present application provides a storage medium storing computer-readable instructions. When the computer-readable instructions are executed by one or more processors, the one or more processors execute the steps of the motion video generation method in any of the above embodiments.
[0045] It can be seen from the above technical solutions that the embodiments of the present application have the following advantages:
[0046] Based on the motion video generation method in this embodiment, the target person image, target background and action-driven video are first obtained, then two-dimensional and three-dimensional posture features are extracted from the action-driven video, and then these two features are fused to obtain fused posture features, and finally the target person image, target background and fused posture features are input into the video generation model to obtain the target motion video. This solution uses both three-dimensional motion information and two-dimensional motion information to guide the action, reduce the probability of ambiguous action generation, and greatly improve the quality of video generation. BRIEF DESCRIPTION OF THE DRAWINGS
[0047] In order to more clearly illustrate the embodiments of the present application or the technical solutions in the prior art, the drawings required for use in the embodiments or the description of the prior art will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present application. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying creative labor.
[0048] Figure 1 A schematic diagram of a flow chart of a motion video generation method provided by an embodiment of the present application;
[0049] Figure 2A schematic diagram of a process for obtaining fused posture features in one embodiment of the present application;
[0050] Figure 3 This is a schematic diagram of a process for obtaining fused posture features in one embodiment of the present application;
[0051] Figure 4 A schematic diagram of a process for obtaining a target background in one embodiment of the present application;
[0052] Figure 5 A schematic diagram of a process for obtaining a target background in one embodiment of the present application;
[0053] Figure 6 This is a schematic diagram of the structure of a video generation model in one embodiment of the present application;
[0054] Figure 7 A schematic diagram of the working process of a video generation model in one embodiment of the present application;
[0055] Figure 8 This is a schematic diagram of a process for training a video generation model in one embodiment of the present application;
[0056] Fig. 9 An internal structure diagram of a computer device provided for one embodiment of the present application. DETAILED DESCRIPTION
[0057] The following will be combined with the drawings in the embodiments of the present application to clearly and completely describe the technical solutions in the embodiments of the present application. Obviously, the described embodiments are only part of the embodiments of the present application, not all of the embodiments. Based on the embodiments in the present application, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of this application.
[0058] This application provides a motion video generation method, see Figure 1 , including steps S102 to S108.
[0059] S102, obtaining a target person image, a target background, and an action-driven video.
[0060] It can be understood that the target character image refers to the appearance characteristics of the character that the user expects to be presented in the generated video. It can be a character image captured from pictures or videos uploaded by the user. The character image image contains information such as the character's appearance, body shape, and clothing. This information will be used to shape the character image when the subsequent video is generated. The target background is the scene where the character is located after the video is generated. Its sources are diverse, such as specific scene pictures provided by the user, or pictures of natural scenery, indoor environments, etc. selected from the material library. Action-driven videos are videos that contain specific action sequences and are the key basis for guiding the target character to make corresponding actions.
[0061] The principle of this step is to provide basic materials for subsequent video generation. By obtaining these three types of data, the input conditions are provided for the human motion transfer algorithm based on the diffusion model, so as to integrate the target character image into the target background and make it perform the actions in the action-driven video. In the entire video generation process, the target character image, target background and action-driven video cooperate with each other. The target character image determines the main characteristics of the character in the video, the target background provides the scene environment for the character's actions, and the action-driven video gives the character dynamic behavior. The three together constitute the complete video generation elements. These three types of data can be specifically acquired by building a user interaction interface. For example, when developing a web application or mobile application, the user can select a locally stored full-body photo as the target character image, select a local picture or select a target background from the built-in background material library of the application, and upload a locally stored action-driven video. If the user needs to use a video or some images uploaded by himself as the target character image, the character image can be cut out by character detection and cutout, so that it is decoupled from the original background, so that the model can better generate videos based on a simple character image.
[0062] S104, extracting two-dimensional posture features and three-dimensional posture features corresponding to each frame according to the action-driven video.
[0063] It can be understood that two-dimensional posture features refer to features that describe the position of human joints and the relative relationship of limbs on a two-dimensional plane, which can intuitively show the shape of human body movements on the plane. Three-dimensional posture features are features that describe the position of human joints, limb rotation angles and spatial relative relationships in three-dimensional space, which can more accurately express the depth and three-dimensional sense of human body movements. Specifically, two-dimensional posture features can be extracted from human body images through mediapipe, which can detect multiple key joints of the human body, such as head, shoulder, elbow, wrist, hip, knee and ankle, and convert their position information into feature data. Three-dimensional posture features can be obtained from human body images with the help of PyMAF-X, which accurately estimates three-dimensional human body posture and shape information based on monocular images, and presents them in the form of posture parameters of the SMPL model, including the rotation angles of 24 key points of the human body. The relevant functions and interfaces of mediapipe and PyMAF-X can be directly called. Taking Python language as an example, after installing the corresponding library, use the human posture detection module of mediapipe, input each frame of the action-driven video, and obtain two-dimensional posture feature data, such as the coordinate value of each joint point. For the three-dimensional posture features, the frame image is input into the PyMAF-X model, and after calculation and processing by the model, the posture parameters of the SMPL model are output.
[0064] This step is to provide multi-dimensional posture information for subsequent action guidance. The two-dimensional posture features can quickly capture the general outline and basic posture changes of human body movements, and the three-dimensional posture features supplement the depth and space information. The combination of the two can more comprehensively describe human body movements. In the entire video generation process, they work closely with other steps. The extracted posture features provide a data basis for the feature fusion of step S106, provide accurate action guidance for the video generation model, and make the character movements in the generated video more in line with expectations.
[0065] S106, fusing the two-dimensional posture feature and the three-dimensional posture feature to obtain a fused posture feature.
[0066] It can be understood that the fusion of two-dimensional posture features and three-dimensional posture features is to combine the advantages of the two and provide more accurate and comprehensive action guidance information. In the present invention, the fusion process is based on specific algorithms and operations. First, the two-dimensional posture feature map is mapped to a higher-dimensional space, and the number of channels is increased to align it with the number of joint points in the three-dimensional posture feature. Then, the mean and standard deviation of the two-dimensional posture feature map and the three-dimensional posture feature vector are calculated respectively. The original features in the two-dimensional posture feature map are removed by these statistics, and the three-dimensional posture features are integrated into it. Finally, the fused posture feature map is obtained by weighted fusion.
[0067] This step is to make up for the shortcomings of using two-dimensional and three-dimensional posture features alone. Although two-dimensional posture features can reflect the plane information of the action, they lack depth information; although three-dimensional posture features can accurately describe spatial actions, they are not intuitive enough in some cases. Through fusion, the advantages of the two are combined to provide richer and more accurate action guidance for the video generation model, reduce the ambiguity of action results, and make the generated character actions more natural and reasonable.
[0068] S108, inputting the target person image, target background and fused posture features into a video generation model to obtain a target motion video.
[0069] It can be understood that the video generation model is a system built on the diffusion model, and its core components include UNet, encoder, decoder, etc. After the target person image, target background and fused posture features are input into the video review model, the diffusion model gradually adjusts the generated image through an iterative denoising process under the guidance of the fused posture features to generate a target action video in which the target person image moves in a similar action in the action-driven video under the target background.
[0070] Based on the motion video generation method in this embodiment, the target person image, target background and action-driven video are first obtained, then two-dimensional and three-dimensional posture features are extracted from the action-driven video, and then these two features are fused to obtain fused posture features, and finally the target person image, target background and fused posture features are input into the video generation model to obtain the target motion video. This solution uses both three-dimensional motion information and two-dimensional motion information to guide the action, reduce the probability of ambiguous action generation, and greatly improve the quality of video generation.
[0071] In one embodiment, the three-dimensional posture feature includes the rotation angles corresponding to the first number of joint points. If the three-dimensional posture feature is a Smpl model, the number of joint points is 24, and includes three different rotation angles. The two-dimensional posture feature and the three-dimensional posture feature are fused to obtain a fused posture feature. Figure 2 , including steps S202 to S208.
[0072] S202, expanding the initial number of channels of the two-dimensional posture feature to a first number of times, and dividing the two-dimensional posture feature into sub-feature graphs corresponding to each joint point in the channel dimension, and determining a first mean and a first standard deviation corresponding to each channel of each sub-feature graph.
[0073] It can be understood that the two-dimensional posture feature is the feature data that describes the posture of the human body on a two-dimensional plane, usually in the form of a feature map, which contains the position information of the joints of the human body on the plane. The initial number of channels is the initial number of channels of the two-dimensional posture feature map, which reflects the dimensional information of the feature map. The first number refers to the number of joints in the three-dimensional posture feature. The sub-feature map is obtained by dividing the two-dimensional posture feature after the number of channels is expanded on the channel dimension. Each sub-feature map corresponds to a joint point, which is used for the subsequent fusion of the three-dimensional posture features corresponding to the joint point. The first mean and the first standard deviation are the mean and standard deviation of the data on each channel of the sub-feature map, respectively, and are used for subsequent feature fusion operations. They can reflect the central tendency and discreteness of the sub-feature map data.
[0074] This step is to make the two-dimensional posture feature match the number of joint points of the three-dimensional posture feature in the channel dimension, so as to facilitate effective feature fusion in the future. By expanding the initial number of channels of the two-dimensional posture feature to the first number of times, the corresponding feature space can be allocated to each joint point in the channel dimension. After dividing the sub-feature graphs corresponding to each joint point, the first mean and the first standard deviation of each channel of each sub-feature graph are calculated. These statistics will be used to adjust the distribution of the sub-feature graphs in the subsequent fusion process, so that it can be better integrated with the three-dimensional posture feature, thereby fully combining the advantages of the two-dimensional and three-dimensional posture features, and finally laying the foundation for generating more accurate fused posture features.
[0075] Figure 3The upper part shows the operation of this step. For example, the number of joints is 24, the 2D posture features are represented by green blocks, and the initial number of channels is c. The first number is 24 (the figure shows that the feature map is divided into 24 parts according to the joints). After the number of channels is expanded to 24 times, it is divided into 24 sub-feature maps in the channel dimension (different color blocks represent different sub-feature maps). By mapping each sub-feature map (map kpi ) calculation, and the first mean value corresponding to each channel is obtained μ (map kpi ) and the first standard deviation σ(map kpi ) (i represents different sub-feature graphs, from 1 to 24). The purpose of this operation is to prepare for the subsequent fusion with the three-dimensional posture features, so that the two-dimensional posture features are adapted to the three-dimensional posture features in terms of structure and statistical characteristics.
[0076] S204, transforming and mapping the three-dimensional posture feature to obtain a three-dimensional mapping posture feature, and determining a second mean and a second standard deviation corresponding to each joint point in the three-dimensional mapping posture feature. Each row of the three-dimensional mapping posture feature corresponds to a joint point, and each second number of columns corresponds to a rotation angle, and the second number is the ratio between the initial channel number and the number of types of rotation angles.
[0077] It can be understood that the three-dimensional posture feature is the feature data that describes the human posture in three-dimensional space, and contains information such as the rotation angle of the human joints. Transformation and mapping are operations performed on the three-dimensional posture feature, the purpose of which is to convert it into a form suitable for fusion with the two-dimensional posture feature, and the result is the three-dimensional mapping posture feature. The second quantity is calculated based on the initial number of channels and the number of types of rotation angles, and is used to determine how many columns in the three-dimensional mapping posture feature correspond to a rotation angle. The second mean and the second standard deviation are the mean and standard deviation of the data corresponding to each joint point in the three-dimensional mapping posture feature, respectively, and are used for subsequent feature fusion operations, reflecting the central tendency and discreteness of the rotation angle data of each joint point.
[0078] The principle of this step is to preprocess the 3D posture features so that they match the sub-feature graph after the 2D posture features are divided in structure, which is convenient for subsequent fusion operations. Through transformation and mapping, the 3D posture features are converted into a specific matrix form, that is, each row corresponds to a joint point, and each second number of columns corresponds to a rotation angle. Such a structure enables the rotation angle information of each joint point to clearly correspond to the sub-feature graph in the 2D posture feature. The second mean and second standard deviation corresponding to each joint point are calculated in order to standardize the 3D mapping posture features in the subsequent fusion process, so that they are more consistent in distribution with the sub-feature graph of the 2D posture features, thereby achieving more effective feature fusion. In the entire video generation process, this step is one of the key links in feature fusion, providing appropriate 3D posture feature representation and statistical information for subsequent steps.
[0079] Figure 3 The lower part shows the operation of this step. The 3D posture feature is initially in the form of a first vector with 72 data. After transformation and mapping, it is converted into a second vector form with 24 rows (each row corresponds to a joint point), and each c / 3 column (assuming there are 3 rotation angles) corresponds to a rotation angle. Then the second mean corresponding to each joint point (vector) is calculated. μ (vector) and the second standard deviation σ(vector). Through such processing, the three-dimensional posture feature has the structure and statistical parameters that are fused with the two-dimensional posture feature sub-feature map.
[0080] S206: For any sub-feature graph, a fused sub-feature graph is obtained according to the corresponding first mean, first standard deviation, second mean and second standard deviation.
[0081] It can be understood that the fused sub-feature graph is the result of fusing the sub-feature graph of the two-dimensional posture feature with the feature graph obtained after processing the three-dimensional posture feature. It integrates the information of the two-dimensional and three-dimensional posture features for the subsequent construction of the fused posture feature. This step adjusts the sub-feature graph of the two-dimensional posture feature based on statistical information so that it can be better integrated with the three-dimensional posture feature. Through the first mean, the first standard deviation, the second mean and the second standard deviation, each sub-feature graph can be standardized and adjusted to remove the distribution influence of the original features in the sub-feature graph, and then the information of the three-dimensional posture feature is integrated into it. The purpose of this is to make the fused sub-feature graph retain the information about the position of the joint point in the two-dimensional posture feature, and integrate the information about the rotation angle of the joint point in the three-dimensional posture feature, so as to achieve the complementary advantages of the two-dimensional and three-dimensional posture features.
[0082] S208, obtaining fused posture features according to each fused sub-feature graph.
[0083] It can be understood that the fused posture feature is the final feature obtained by combining all the fused sub-feature graphs. It integrates the information of two-dimensional and three-dimensional posture features and is used for the subsequent video generation model to provide the model with more accurate and comprehensive action guidance information. The step integrates the fused sub-feature graphs to form a complete fused posture feature. Since each fused sub-feature graph corresponds to a joint point, combining them together can fully describe the posture information of the human body, including the position and rotation angle of the joint point. Such fused posture features can provide richer and more accurate action guidance for the video generation model, making the character movements in the generated video more natural and reasonable. In the entire video generation process, this step is the last step of feature fusion, which integrates the local fusion results obtained in the previous steps into an overall feature, providing a key input for subsequent video generation. Specifically, all sub-feature graphs can be added and averaged.
[0084] In one embodiment, obtaining a fused sub-feature map according to the corresponding first mean, first standard deviation, second mean, and second standard deviation includes:
[0085] The intermediate sub-feature map is obtained according to the first expression. The first expression is:
[0086]
[0087] in, is the intermediate sub-feature map, is the i-th sub-feature map, is the first mean, is the first standard deviation, is the second mean, is the second standard deviation. The intermediate sub-feature map is an intermediate product in the process of fusing the 2D posture feature sub-feature map with the 3D posture feature. It is obtained by processing the i-th sub-feature map through the first expression. It preliminarily fuses the statistical information of the 3D posture feature and prepares for the final formation of the fusion sub-feature map. In the first expression, each data point in the i-th sub-feature map is first subtracted from its mean, so that the center of the data moves to near 0 and then divided by its standard deviation, in order to scale the data distribution to a form with a standard deviation of 1. In this way, the standardization of the 2D posture feature sub-feature map is completed, the influence of its original data distribution is removed, and different sub-feature maps are comparable, which also facilitates the subsequent fusion operation with the 3D posture feature. At this time, multiply by the second standard deviation corresponding to the 3D posture feature. The role of this step is to scale the standardized data according to the discrete degree of the 3D posture feature. Because the range of change of the 3D posture feature may be different in different actions or scenes, by multiplying by the second standard deviation, the data range of the 2D sub-feature map can be adapted to the 3D posture feature. Finally, add the second mean corresponding to the i-th joint point of the three-dimensional posture feature. This step is to move the center of the processed data to the mean position of the i-th joint point in the three-dimensional posture feature. Through this series of operations, the statistical information of the three-dimensional posture feature (the distribution characteristics represented by the mean and standard deviation) is integrated into the standardized two-dimensional sub-feature map, and finally the intermediate sub-feature map is obtained. The intermediate sub-feature map obtained in this way not only retains some basic information of the two-dimensional posture feature, but also has the distribution characteristics of the three-dimensional posture feature, laying the foundation for further weighted summation with the original two-dimensional sub-feature map to generate a fused sub-feature map. After obtaining the intermediate sub-feature map, perform a weighted summation with the corresponding sub-feature map to obtain a fused sub-feature map. That is and The information of both is combined by weighted addition according to the set ratio, and finally a fused sub-feature map is obtained.
[0088] In one embodiment, see Figure 4 The target background acquisition process includes step S402 and step S404.
[0089] S402, performing human body detection on the original background to obtain a human body region mask.
[0090] It can be understood that the original background refers to the complete image information containing scenes and characters that is initially obtained in the motion video generation process. Human body detection is a technology in the field of computer vision that uses algorithms and models to identify and locate human bodies in images. The human body region mask is a special image representation that has the same size as the image cropped from the human body detection result. It marks the area where the human body is located in the original background in a binary form (usually represented by 0 and 1), where the part with a value of 1 corresponds to the human body area, and the part with a value of 0 corresponds to the non-human area. Figure 5 As shown in the figure, the leftmost group of images shows multiple frames of original background images, including dancers. The human detection link uses relevant algorithms to identify and locate the human body in the original background based on the shape, texture and other features of the human body, and determine the area where the dancer's body is located from the original background image containing the dancer. Then, the human body area mask is generated by human matting. Furthermore, because each person's appearance features are different, in order to prevent the mask obtained by human matting from carrying the human features in the original background sequence, it is necessary to perform some processing on the obtained mask sequence. First, the mask is closed to smooth the mask periphery and remove the edges and corners, and some appearance features of the characters in the original background are removed. Then, the mask that has been closed is expanded so that the human body generated by the subsequent drive can have enough space to be embedded in the specified scene. The original background here can be obtained by identifying the action-driven video, which can make the target character image move in the same background as the action-driven video, or it can be any background that the user wants to use. In addition, if it is a background image or video without a person, it can be directly input into the model as the target background.
[0091] S404, obtaining the target background according to the human body region mask and the original background.
[0092] It can be understood that the target background refers to an image that has been processed and the human body part in the original background has been removed. It will serve as the background environment of the characters in the subsequent motion video. Specifically, the human body region mask can be used as an index to operate the original background image. Specifically, according to the human body region marked in the mask, the corresponding human body part in the original background image is removed, replaced or otherwise processed, so as to obtain the target background that does not contain the original human body. In the entire video generation process, the role of this step is to provide a suitable background scene for the target character image, avoid the conflict between the human body in the original background and the target character image, and make the final generated motion video more natural and in line with expectations. As can be seen from the figure, through the "scaling & positioning" operation, the pixels corresponding to the human body (dancer) area in the original background image are removed or processed with the mask as a guide to obtain the "masked background", that is, the background picture of the dance room after the dancer is removed. This step obtains a clean and suitable background so that the target character image can be integrated into it to generate a motion video in the future, ensuring that the background and the character can be reasonably matched, and improving the quality and effect of video generation.
[0093] In one embodiment, see Figure 6 The video generation model includes an encoder (vae encoder), a feature extraction module (dinov2), a UNET network (DownSampling+UpSampling) and a decoder (vae decoder). The input of the encoder is the image of the target person, which is connected to the downsampling layer of the UNET network through an adder. The input of the feature extraction module is the image of the target person, and the output of the feature extraction module is connected to the downsampling layer and upsampling layer of the UNET network respectively. The output of the upsampling layer of the UNET network is connected to the decoder. Please refer to Figure 7 , input the target person image, target background and fused posture features into the video generation model to obtain the target motion video, including steps S702 to S708.
[0094] S702, inputting the fused posture features into the downsampling layer of the UNET network.
[0095] S704, input the target person image into the encoder.
[0096] S706, add potential noise coding to the target background and sum it with the output result of the encoder. The summation result is input to the downsampling layer of the UNET network.
[0097] S708, input the target person image into the feature extraction module. The result of the feature extraction module is input into the downsampling layer and upsampling layer of the UNET network. The output of the upsampling layer of the UNET network is connected to the decoder, and the decoder is used to generate the target motion video according to the output of the upsampling layer of the UNET network.
[0098] It can be understood that the downsampling layer of the UNET network undertakes the key task of feature extraction and abstraction of the input data. Through convolution and pooling operations, the spatial dimension of the data is gradually reduced, while the number of feature channels is increased, thereby obtaining high-level features in the data. The data input to the downsampling layer includes fused posture features, the target person image fused with the target background, and the target person image after additional feature extraction. Posture guidance is crucial in generating high-quality human action videos. The fused posture features integrate the two-dimensional and three-dimensional motion information of the human body, accurately depicting the state of human body motion at different times. Inputting the fused posture features into the downsampling layer of the UNET network can enable the model to play a guiding role in the denoising at the beginning of the latent space constructed by UNET. The fusion of the target person image and the target background can be achieved by superposition. The target background needs to add potential noise coding before superposition, and the model eliminates this part of the noise by diffusion in the UNET space. The target background is also a sequence of video frames with the same number of frames as the driving video sequence, while the target person image can be a single image. The target person image is encoded using an encoder to obtain its representation in the latent space. Then, the latent features of a single target person image are replicated along the time dimension to align with the features of each video frame. The copied target character image and the latent video frame corresponding to the target background are connected together along the channel dimension and then input into U-Net for diffusion. In addition, the target character image is input into the cross attention of the downsampling layer and upsampling layer of the UNET network respectively after feature extraction by the feature extraction module to control the output result. The result after diffusion denoising in the UNET network will be input into the decoder for decoding to obtain the target motion video. The encoder and decoder here can be a variational autoencoder and a variational autodecoder. The feature extraction module can be a module based on the DiNOv2 structure.
[0099] In one embodiment, see Figure 8 , the training process of the video generation model includes steps S802 to S808.
[0100] S802, obtaining training 3D posture features according to the training action driving video.
[0101] It can be understood that the training action-driven video, as the basic data of the training video generation model, is a video sequence containing a series of changes in human body action postures. It records the action performance of the human body at different times and provides rich information for the model to learn the human body action pattern. The training three-dimensional posture feature is extracted from the training action-driven video and is used to describe the characteristic information of the human body in three-dimensional space. The loss function used in the general diffusion model framework that drives the human body based on posture information is the pixel-level difference made on the potential coding layer as the loss, and then the parameters in the network model are updated through back propagation. The reasoning process of the diffusion model is a step-by-step iterative denoising process, and the mechanism in training is to randomly perform a denoising step for training. It has been observed that decoding the latent space features obtained from certain steps can obtain a picture with a human figure, and the posture of the human action in the image can be determined. Therefore, this embodiment mainly constructs a loss function based on the difference between the three-dimensional posture predicted in the middle of the training process and the three-dimensional data in the actual training data for training, which is more intuitive and effective than calculating the loss in the potential coding layer.
[0102] S804, obtaining predicted three-dimensional posture features based on the predicted motion video obtained by the video generation model for the training action-driven video.
[0103] It can be understood that the training action-driven video is a video sequence generated by the video generation model based on the training action-driven video. It is the model's simulation and reproduction of the action in the input video. The quality and accuracy of the predicted motion video reflects the performance of the video generation model. The predicted 3D posture feature is the feature information extracted from the predicted motion video that describes the posture of the human body in 3D space. Similar to the training 3D posture feature, it is also obtained by analyzing and processing the predicted motion video, and is used to evaluate the difference between the action generated by the model and the real action.
[0104] S806, obtaining a target loss value according to the predicted three-dimensional posture features and the trained three-dimensional posture features.
[0105] It can be understood that the target loss value is a value obtained by comparing the predicted 3D posture features and the trained 3D posture features, and it is used to measure the difference between the action generated by the video generation model and the real action. The smaller the target loss value, the closer the action generated by the model is to the real action, and the better the performance of the model. Specifically, the target loss value can be obtained according to the second expression, the predicted 3D posture features and the trained 3D posture features. The second expression is:
[0106]
[0107] in, is the target loss value, n is the number of joint points in the 3D posture feature, is the training 3D posture feature of the i-th joint point, is the predicted 3D posture feature of the i-th joint point. In the expression, the first part of the loss term is the sum of the absolute values of the differences between the training and predicted 3D posture features of each joint point, divided by the number of joint points, which measures the average absolute degree of the overall difference; the second part of the loss term is the sum of the squares of the differences between the training and predicted 3D posture features of each joint point, divided by the number of joint points, which highlights the impact of larger differences. The two parts are added together to obtain the final target loss value.
[0108] S808, adjusting the video generation model according to the target loss value.
[0109] It can be understood that the principle of this step is to use the back propagation algorithm and the optimizer to adjust the parameters of the video generation model according to the target loss value. In deep learning, the training process of the model is the process of constantly adjusting the model parameters to reduce the value of the loss function. After the target loss value is calculated, the gradient of the loss function relative to the model parameters is calculated by the back propagation algorithm. These gradients represent the rate of change of each parameter to the loss function. Then, the optimizer updates the parameters of the model according to these gradients and parameters such as the preset learning rate. By constantly repeating this process, the parameters of the model will gradually adjust to the state that minimizes the loss function value, thereby improving the performance of the model. In some embodiments, the UNET network part is mainly trained, and the encoder and decoder parts can use frozen pre-trained parameters.
[0110] The present application provides a motion video generation device, including a data acquisition module, a posture feature extraction module, a posture feature fusion module and a video generation module.
[0111] The data acquisition module is used to obtain the target person image, target background and action-driven video. The posture feature extraction module is used to extract the two-dimensional posture features and three-dimensional posture features corresponding to each frame according to the action-driven video. The posture feature fusion module is used to fuse the two-dimensional posture features and the three-dimensional posture features to obtain the fused posture features. The video generation module is used to input the target person image, target background and fused posture features into the video generation model to obtain the target motion video.
[0112] For the specific limitations of the motion video generation device, please refer to the limitations of the motion video generation method above, which will not be repeated here. The various modules in the above-mentioned motion video generation device can be implemented in whole or in part by software, hardware and a combination thereof. The above-mentioned modules can be embedded in or independent of the processor in the computer device in the form of hardware, or can be stored in the memory in the computer device in the form of software, so that the processor can call and execute the operations corresponding to the above modules. It should be noted that the division of modules in the embodiments of the present application is schematic and is only a logical function division. There may be other division methods in actual implementation.
[0113] The present application provides a computer device, including one or more processors and a memory, wherein the memory stores computer-readable instructions, and when the computer-readable instructions are executed by the one or more processors, the steps of the motion video generation method in any of the above embodiments are executed.
[0114] Indicatively, Fig. 9 As shown, Fig. 9 A schematic diagram of the internal structure of a computer device provided in an embodiment of the present application. Fig. 9 The computer device 900 includes a processing component 902, which further includes one or more processors, and a memory resource represented by a memory 901, for storing instructions executable by the processing component 902, such as an application. The application stored in the memory 901 may include one or more modules, each corresponding to a set of instructions. In addition, the processing component 902 is configured to execute instructions to perform the steps of the motion video generation method of any of the above embodiments.
[0115] The computer device 900 may further include a power supply component 903 configured to perform power management of the computer device 900 , a wired or wireless model interface 904 configured to connect the computer device 900 to the model, and an input / output (I / O) interface 905 .
[0116] The present application provides a storage medium storing computer-readable instructions. When the computer-readable instructions are executed by one or more processors, the one or more processors execute the steps of the motion video generation method in any of the above embodiments.
[0117] Finally, it should be noted that, in this article, relational terms such as first and second, etc. are only used to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any such actual relationship or order between these entities or operations. Moreover, the terms "include", "comprise" or any other variants thereof are intended to cover non-exclusive inclusion, so that a process, method, article or device including a series of elements includes not only those elements, but also other elements not explicitly listed, or also includes elements inherent to such process, method, article or device. In the absence of further restrictions, the elements defined by the sentence "comprise a ..." do not exclude the presence of other identical elements in the process, method, article or device including the elements.
[0118] The various embodiments in this specification are described in a progressive manner, and each embodiment focuses on the differences from other embodiments. The various embodiments can be combined as needed, and the same or similar parts can refer to each other.
[0119] The above description of the disclosed embodiments enables those skilled in the art to implement or use the present application. Various modifications to these embodiments will be apparent to those skilled in the art, and the general principles defined herein may be implemented in other embodiments without departing from the spirit or scope of the present application. Therefore, the present application will not be limited to the embodiments shown herein, but will conform to the widest scope consistent with the principles and novel features disclosed herein.
Claims
1. A motion video generation method, characterized in that: include: Obtain target person image, target background and action-driven video; Extracting two-dimensional posture features and three-dimensional posture features corresponding to each frame of the action-driven video respectively; Fusing the two-dimensional posture feature with the three-dimensional posture feature to obtain a fused posture feature; The target person image, the target background and the fused posture features are input into a video generation model to obtain a target motion video.
2. The motion video generation method according to claim 1, characterized in that: The three-dimensional posture feature includes rotation angles corresponding to the first number of joint points, and the two-dimensional posture feature and the three-dimensional posture feature are fused to obtain the fused posture feature, including: Expanding the initial number of channels of the two-dimensional posture feature to the first number of times, dividing the two-dimensional posture feature into sub-feature graphs corresponding to each of the joint points in the channel dimension, and determining a first mean and a first standard deviation corresponding to each channel of each of the sub-feature graphs; Transform and map the three-dimensional posture feature to obtain a three-dimensional mapping posture feature, and determine a second mean and a second standard deviation corresponding to each joint point in the three-dimensional mapping posture feature; each row of the three-dimensional mapping posture feature corresponds to one joint point, and each second number column corresponds to one rotation angle, and the second number is a ratio between the initial channel number and the number of types of the rotation angle; For any of the sub-feature graphs, obtaining a fused sub-feature graph according to the corresponding first mean, the first standard deviation, the second mean, and the second standard deviation; The fused posture feature is obtained according to each of the fused sub-feature graphs.
3. The motion video generation method according to claim 2, characterized in that: The obtaining a fused sub-feature map according to the corresponding first mean, the first standard deviation, the second mean, and the second standard deviation includes: The intermediate sub-feature graph is obtained according to the first expression; the first expression is: in, is the intermediate sub-feature map, is the i-th sub-feature map, is the first mean, is the first standard deviation, is the second mean, is the second standard deviation; The intermediate sub-feature map and the corresponding sub-feature map are weightedly summed to obtain the fused sub-feature map.
4. The motion video generation method according to claim 1, characterized in that: The target background acquisition process includes: Perform human body detection on the original background to obtain a human body area mask; The target background is obtained according to the human body region mask and the original background.
5. The motion video generation method according to claim 4, characterized in that: Before obtaining the target background according to the human body region mask and the original background, the method further includes: The human body region mask is sequentially subjected to a closing operation and a dilation operation.
6. The motion video generation method according to claim 1, characterized in that: The video generation model includes an encoder, a feature extraction module, a UNET network and a decoder. The target person image, the target background and the fusion posture feature are input into the video generation model to obtain the target motion video, including: Inputting the fused posture features into the downsampling layer of the UNET network; Inputting the target person image into the encoder; Adding potential noise coding to the target background and summing it with the output result of the encoder; the summation result is input into the downsampling layer of the UNET network; The target person image is input into the feature extraction module; the result of the feature extraction module is input into the downsampling layer and upsampling layer of the UNET network, the output of the upsampling layer of the UNET network is connected to the decoder, and the decoder is used to generate the target motion video according to the output of the upsampling layer of the UNET network.
7. The motion video generation method according to claim 6, characterized in that: The training process of the video generation model includes: According to the training action driving video, the training three-dimensional posture features are obtained; Obtain predicted three-dimensional posture features based on the predicted motion video obtained by the video generation model for the training action driving video; Obtaining a target loss value according to the predicted three-dimensional posture feature and the training three-dimensional posture feature; The video generation model is adjusted according to the target loss value.
8. The motion video generation method according to claim 7, characterized in that: The obtaining of a target loss value according to the predicted three-dimensional posture feature and the training three-dimensional posture feature comprises: The target loss value is obtained according to a second expression, the predicted three-dimensional posture feature and the training three-dimensional posture feature; the second expression is: in, is the target loss value, n is the number of joint points in the three-dimensional posture feature, is the training 3D posture feature of the i-th joint point, is the predicted three-dimensional posture feature of the i-th joint point.
9. A motion video generating device, characterized in that: include: A data acquisition module is used to acquire the target person image, target background and action-driven video; A posture feature extraction module, used to extract two-dimensional posture features and three-dimensional posture features corresponding to each frame according to the action-driven video; A posture feature fusion module, used for fusing the two-dimensional posture feature with the three-dimensional posture feature to obtain a fused posture feature; The video generation module is used to input the target person image, the target background and the fusion posture features into a video generation model to obtain a target motion video.
10. A computer device, characterized in that: The system comprises one or more processors and a memory, wherein the memory stores computer-readable instructions, and when the computer-readable instructions are executed by the one or more processors, the steps of the motion video generation method according to any one of claims 1 to 8 are executed.
11. A storage medium, characterized in that: The storage medium stores computer-readable instructions, and when the computer-readable instructions are executed by one or more processors, the one or more processors execute the steps of the motion video generation method according to any one of claims 1 to 8.
Citation Information
Patent Citations
Human motion posture migration method and device, control equipment and readable storage medium
CN114093033A
Three-dimensional human body posture migration method based on video time sequence information
CN115761801A
Image generation method and device based on action migration and electronic equipment
CN115908858A
Model training method based on motion capture
CN116091972A
Cited By
Video generation method and device based on diffusion model, equipment and medium
CN121967824A