Video generation method and apparatus, electronic device, and readable storage medium

By using semantic depth maps and key point maps to generate limb-driven videos in digital human limb-driven technology, the problem of low efficiency in existing technologies is solved, and realistic limb movement effects are generated efficiently.

CN119342246BActive Publication Date: 2025-11-25VIVO MOBILE COMM CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202411318250.X
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-09-20
Publication Date
2025-11-25
Estimated Expiration
2044-09-20

AI Technical Summary

Technical Problem

Existing digital human limb-driven technologies rely on specialized equipment and complex manual processing, resulting in low generation efficiency.

Method used

By determining M first semantic depth maps and M first keypoint maps based on M-frame pose reference video maps, and inputting them into the limb driving model to generate limb driving videos, the use of motion capture technology and hardware devices is avoided.

Benefits of technology

It improves the generation efficiency of digital human body-driven videos and achieves highly realistic, natural, and smooth motion effects.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119342246B_ABST
    Figure CN119342246B_ABST
Patent Text Reader

Abstract

The application discloses a video generation method and device, electronic equipment and readable storage medium, and belongs to the field of artificial intelligence. The method comprises the following steps: determining M first semantic depth maps and M first key point maps based on M posture reference video maps, the first semantic depth map being used for indicating depth information of at least one body part of a character in the posture reference video map, and the first key point map being used for indicating a body key point of the character in the posture reference video map; inputting a character reference map, the M first semantic depth maps and the M first key point maps into a limb driving model to output M first posture video maps; and generating a limb driving video corresponding to a character image map according to the M first posture video maps; wherein one first semantic depth map and one first key point map correspond to one posture reference video map, and the M first semantic depth maps and the M first key point maps correspond to each other in one-to-one manner.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application belongs to the field of artificial intelligence, and specifically relates to a video generation method, apparatus, electronic device, and readable storage medium. Background Technology

[0002] In recent years, with the rapid development of digital human technology, it is continuously driving innovation in multiple industries such as virtual assistants, online customer service, education, entertainment, and advertising, providing users with richer and more personalized experiences. Among them, digital human body motion-driven technology is an important branch.

[0003] Currently, digital human limb motion driving technology typically involves adding inertial and optical sensors to different joints of a real person to acquire real motion data, and then mapping it onto the skeletal system of a 3D model to drive the digital human to perform corresponding movements.

[0004] However, while the above methods can achieve realism and lifelikeness in digital human motion-driven animation, they rely on specialized equipment for data collection and processing. Furthermore, the digital human models to be driven require complex manual processing such as skeletal binding and weight adjustment, which is time-consuming and labor-intensive, increasing the production cycle and cost, thus resulting in low efficiency in generating digital human limb-driven videos. Summary of the Invention

[0005] The purpose of this application is to provide a video generation method, apparatus, electronic device, and readable storage medium that can improve the generation efficiency of digital human limb-driven videos.

[0006] In a first aspect, embodiments of this application provide a video generation method, which includes: determining M first semantic depth maps and M first keypoint maps based on M frames of pose reference video images, wherein the first semantic depth maps are used to indicate the depth information of at least one body part of a person in the pose reference video images, and the first keypoint maps are used to indicate the key points of the person's body in the pose reference video images; inputting a person image image, the M first semantic depth maps, and the M first keypoint maps into a limb driving model, and outputting M frames of first pose video images; generating a limb driving video corresponding to the person image image based on the M frames of first pose video images; wherein one frame of pose reference video image corresponds to one first semantic depth map and one first keypoint map, and the M first semantic depth maps and M first keypoint maps... Figure 1 One-to-one correspondence.

[0007] Secondly, embodiments of this application provide a video generation apparatus, comprising: a determining module, configured to determine M first semantic depth maps and M first keypoint maps based on M frames of pose reference video images, wherein the first semantic depth maps are used to indicate the depth information of at least one body part of a person in the pose reference video images, and the first keypoint maps are used to indicate key points of the person's body in the pose reference video images; a processing module, configured to input the person image image, the M first semantic depth maps determined by the determining module, and the M first keypoint maps determined by the determining module into a limb driving model, and output M frames of first pose video images; and a generation module, configured to generate a limb driving video corresponding to the person image image based on the M frames of first pose video images output by the processing module; wherein one frame of pose reference video image corresponds to one first semantic depth map and one first keypoint map, and the M first semantic depth maps and M first keypoint maps... Figure 1 One-to-one correspondence.

[0008] Thirdly, embodiments of this application provide an electronic device including a processor and a memory, the memory storing programs or instructions executable on the processor, the programs or instructions, when executed by the processor, implementing the steps of the method described in the first aspect.

[0009] Fourthly, embodiments of this application provide a readable storage medium on which a program or instructions are stored, which, when executed by a processor, implement the steps of the method described in the first aspect.

[0010] Fifthly, embodiments of this application provide a chip, the chip including a processor and a communication interface, the communication interface being coupled to the processor, the processor being used to run programs or instructions to implement the method as described in the first aspect.

[0011] In a sixth aspect, embodiments of this application provide a computer program product stored in a storage medium, which is executed by at least one processor to implement the method described in the first aspect.

[0012] In this embodiment, based on M frames of pose reference video images, M first semantic depth maps and M first keypoint maps are determined. The first semantic depth maps are used to indicate the depth information of at least one body part of the person in the pose reference video images, and the first keypoint maps are used to indicate the key points of the person's body in the pose reference video images. The person image, the M first semantic depth maps, and the M first keypoint maps are input into the limb driving model, and M frames of first pose video images are output. Based on the M frames of first pose video images, a limb driving video corresponding to the person image is generated. Each frame of the pose reference video image corresponds to one first semantic depth map and one first keypoint map, and the M first semantic depth maps and M first keypoint maps... Figure 1 One-to-one correspondence. In this scheme, by inputting a semantic depth map representing the depth information of at least one body part of a person in the pose reference video image, and a first keypoint map representing the key points of the person's body in the pose reference video image, into the limb driving model, the limb driving model can generate a limb driving video of the person in the character image image based on the limb driving details in the semantic depth map and the first keypoint map. This can achieve a limb driving video with free-moving movements without the need for motion capture technology and hardware devices, thereby improving the generation efficiency of digital human limb driving videos. Attached Figure Description

[0013] Figure 1 This is a schematic diagram of the video generation method provided in an embodiment of this application;

[0014] Figure 2 This is a schematic diagram of the video generation method provided in an embodiment of this application;

[0015] Figure 3 This is a schematic diagram of the posture extraction layer in the limb driving model provided in this application embodiment;

[0016] Figure 4 This is a schematic diagram of the structure of the video generation device provided in the embodiments of this application;

[0017] Figure 5 This is a schematic diagram of the structure of the video generation device provided in the embodiments of this application;

[0018] Figure 6 This is one of the hardware structure diagrams of an electronic device provided in the embodiments of this application;

[0019] Figure 7 This is a second schematic diagram of the hardware structure of an electronic device provided in an embodiment of this application. Detailed Implementation

[0020] The technical solutions of the embodiments of this application will be clearly described below with reference to the accompanying drawings. Obviously, the described embodiments are only some, not all, of the embodiments of this application. All other embodiments obtained by those skilled in the art based on the embodiments of this application are within the scope of protection of this application.

[0021] The terms "first," "second," etc., used in the specification and claims of this application are used to distinguish similar objects and not to describe a specific order or sequence. It should be understood that such terms can be used interchangeably where appropriate so that embodiments of this application can be implemented in orders other than those illustrated or described herein, and the objects distinguished by "first," "second," etc., are generally of the same class and the number of objects is not limited; for example, a first object can be one or more. Furthermore, in the specification and claims, "and / or" indicates at least one of the connected objects, and the character " / " generally indicates that the preceding and following objects are in an "or" relationship.

[0022] The terms "at least one," "at least one of," etc., used in the specification and claims of this application refer to any one, any two, or a combination of two or more of the included items. For example, "at least one of a, b, and c" can mean: "a," "b," "c," "a and b," "a and c," "b and c," and "a, b, and c," where a, b, and c can be single or multiple. Similarly, "at least two" refers to two or more items, and its meaning is similar to that of "at least one."

[0023] It should be noted that the video generation method provided in this application can be executed by electronic devices such as mobile phones, tablets, laptops, PDAs, and in-vehicle electronic devices. Some embodiments of this application use electronic devices as the executing entity to illustrate the video generation method provided in this application.

[0024] The video generation method, apparatus, electronic device, and readable storage medium provided in this application will be described in detail below with reference to the accompanying drawings and through specific embodiments and application scenarios.

[0025] This application provides a simple and easy-to-use technical solution for generating digital human limb-driven videos, which can achieve highly realistic, natural and smooth digital human limb movement-driven effects.

[0026] Specifically, the user can first select a limb-driven reference video and a human figure image. Then, based on the selected limb-driven reference video and human figure image, the electronic device accurately reproduces each action posture in the limb-driven reference video using the limb body model provided in this application embodiment, and finally generates a limb-driven video corresponding to the human figure image.

[0027] This application has a wide range of applications. On the one hand, it empowers new ways of personal creative expression, allowing users to easily create personalized action videos, such as dance videos, using only their own photos or photos of any person, greatly satisfying users' needs to share unique creative content on various social media platforms. On the other hand, in the field of game development, the technical solution provided by this application uses game character images as character images and combines different body-driven videos to generate the body movements of game characters, thereby optimizing the body driving of game characters, accelerating the design of game character action postures, and enhancing the richness and diversity of in-game character behavior.

[0028] The execution entity of the video generation method provided in this embodiment can be a video generation device, which can be an electronic device, or a control module or processing module within the electronic device. The following description uses an electronic device as an example to illustrate the technical solution provided in this application embodiment.

[0029] This application provides a video generation method. Figure 1 A flowchart illustrating a video generation method provided in an embodiment of this application is shown, which can be applied to electronic devices. Figure 1 As shown, the video generation method provided in this application embodiment may include the following steps 201 to 203.

[0030] Step 201: The electronic device determines M first semantic depth maps and M first key point maps based on the M-frame pose reference video map.

[0031] In some embodiments of this application, the first semantic depth map described above is used to indicate the depth information of at least one body part of a person in a pose reference video image.

[0032] In some embodiments of this application, the first key point map described above is used to indicate key points of the body of a person in a posture reference video image.

[0033] In some embodiments of this application, the aforementioned one-frame pose reference video graph corresponds to a semantic depth map and a first keypoint map.

[0034] In some embodiments of this application, M first semantic depth maps and M first keypoints Figure 1 One-to-one correspondence.

[0035] In some embodiments of this application, the electronic device acquires each video frame in the posture reference video, i.e., the aforementioned posture reference video image, and uses a human detection algorithm to detect human bounding boxes, expanding the height of the bounding box by 10% to increase the overall aspect ratio to 2:1. If no human is detected in a certain frame, the video segment is split into two parts, with that frame as the dividing point. After detection, the union of the human bounding boxes of all frames in a single video is taken to obtain a rectangular box that can cover the human area of ​​all frames, and the corresponding area is cropped out in each frame of the video to obtain M frames of posture reference video images.

[0036] In some embodiments of this application, the electronic device acquires a pose reference video and, based on the human model information contained in the human body model of the person in each frame of the pose reference video, renders and generates corresponding semantic maps and depth maps respectively. Then, the semantic map and depth map corresponding to a frame of the pose reference video are concatenated to obtain the aforementioned first semantic depth map. Simultaneously, based on the body key points of the person in each frame of the pose reference video, the aforementioned M first keypoint maps are generated.

[0037] Optionally, in some embodiments of this application, before step 201 above, the video generation method provided in this application further includes steps 301 and 302:

[0038] Step 301: The electronic device acquires the semantic map corresponding to the M-frame pose reference video map and the depth map corresponding to the M-frame pose reference video map.

[0039] In some embodiments of this application, the semantic map described above is used to characterize at least one body part of a person in a pose reference video image.

[0040] For example, the semantic map described above can use different colors to represent different body parts of the reference figure.

[0041] For example, hands can be represented in yellow, elbows in yellow-green, and shoulders in green.

[0042] In some embodiments of this application, the depth map described above is used to characterize the depth information of the body of a person in a pose reference video image.

[0043] For example, the depth map described above can represent the three-dimensional spatial position information of a person in a pose reference video image, allowing users to see the three-dimensional model of the person in the pose reference video image based on the depth map.

[0044] Step 302: The electronic device stitches together the semantic map and depth map corresponding to each frame of pose reference video according to the color channel to obtain M first semantic depth maps.

[0045] In some embodiments of this application, the electronic device can superimpose the color information corresponding to the same color channel in the semantic map and depth map of each frame of pose reference video map according to the color channel to obtain M first semantic depth maps.

[0046] It should be noted that when stitching images according to color channels, the height and width of the image remain unchanged.

[0047] Optionally, in some embodiments of this application, step 301 specifically includes steps 301a to 301d:

[0048] Step 301a: The electronic device constructs a first human body model based on the key points of the human body in the human figure image.

[0049] In some embodiments of this application, the electronic device obtains key points of the reference person's body using a key point detection algorithm, such as DWPose.

[0050] For example, the key points of the body of the figure in the above-mentioned figure image can be the key points of the body skeleton.

[0051] For example, the aforementioned key points of the skeletal structure include, but are not limited to, at least one of the following: key points of the hand bones, key points of the face bones, and key points of the joint bones.

[0052] In some embodiments of this application, the electronic device uses a 3D whole-body reconstruction algorithm (such as the ExPose algorithm) to generate a first human body model based on the above-mentioned key body points.

[0053] In some embodiments of this application, the first human body model described above includes first human body model parameters.

[0054] For example, the above-mentioned human body model reference includes, but is not limited to, at least one of the following: a first body shape parameter, a first body posture parameter, a first hand posture parameter, and a first facial expression parameter.

[0055] Step 301b: The electronic device constructs M second human body models based on the key points of the human body in each frame of the pose reference video image.

[0056] In some embodiments of this application, a second human body model corresponds to a frame of pose reference video image.

[0057] In some embodiments of this application, the electronic device uses a key point detection algorithm, such as DWPose, to obtain the body key points of the person in the video in each frame of the pose reference video.

[0058] In some embodiments of this application, the electronic device uses a 3D full-body reconstruction algorithm (such as the ExPose algorithm) to generate M second human body models based on the key points of the body of the person in the video in each frame of the pose reference video.

[0059] In some embodiments of this application, the second human body model described above includes second human body model parameters.

[0060] For example, the above-mentioned second human body model parameters include, but are not limited to, at least one of the following: second body shape parameters, second body posture parameters, second hand posture parameters, and second facial expression parameters.

[0061] Step 301c: The electronic device obtains M third human models based on the body shape parameters in the first human model and the posture parameters in the M second human models.

[0062] In some embodiments of this application, a third human body model corresponds to a frame of pose reference video image.

[0063] In some embodiments of this application, the aforementioned third human body model includes third human body model parameters.

[0064] In some embodiments of this application, the posture parameters in the aforementioned third human body model include, but are not limited to, at least one of the following: third body shape parameter, third body posture parameter, third hand posture parameter, and third facial expression parameter.

[0065] In some embodiments of this application, the electronic device merges the body shape parameters in the first human body model (i.e., the first body shape parameters) with the posture parameters of each of the M second human body models to obtain the M third human body model parameters, and then generates M third human body models based on the M third human body model parameters.

[0066] For example, the posture parameters of the second human body model mentioned above include at least one of the following: second body posture parameters, second hand posture parameters, and second facial expression parameters.

[0067] For example, the electronic device first uses a keypoint detection algorithm (such as DWPose) to detect the keypoints of the person's body in each video frame, and the keypoints of the person's body in the image. It should be noted that, compared with existing general skeletal keypoint detection algorithms, the method used in this application embodiment not only includes common human skeletal keypoints, but also adds hand keypoints and facial keypoints. This information will then be added to the driving model for fine control of local body positions.

[0068] Next, the electronic device reconstructs the SMPL-x human body model of the person based on the key points of the person's body in each video frame and the key points of the person's body in the image of the person, using a 3D full-body reconstruction algorithm (such as the ExPose algorithm), namely the first human body model and the second human body model mentioned above.

[0069] Because the human body model contains parameters such as body shape parameters, body posture parameters, hand posture parameters, and facial expression parameters, the human body model reconstructed by electronic devices can accurately capture the body shape, posture, hand movements, and facial expressions of a person.

[0070] In step 301d, the electronic device renders the M third human models respectively to obtain the semantic map corresponding to each frame of the pose reference video map and the depth map corresponding to each frame of the pose reference video map.

[0071] In some embodiments of this application, the electronic device renders the above M third human models by using rendering software or rendering processes and other algorithms with rendering functions to obtain a semantic map corresponding to each frame of pose reference video map and a depth map corresponding to each frame of pose reference video map.

[0072] For example, the electronic device uses rendering software, such as Blender, to render the semantic map and depth map of the human body model based on the established 3D human body model, that is, the third human body model corresponding to each frame of the pose reference video image mentioned above.

[0073] Thus, since the semantic map of the human body can provide information on the coverage area of ​​each joint as input, it helps to reduce the problem of limb distortion and improve the model generation effect; while the depth map can provide 3D spatial position information of the human body, which helps to improve the three-dimensionality of the person in the model generation result. Using these rendered result maps as input to the limb-driven model can generate more accurate limb-driven videos.

[0074] Step 202: The electronic device inputs the character image, M first semantic depth maps and M first key point maps into the limb driving model, and outputs M frames of first pose video images.

[0075] In some embodiments of this application, the appearance of the person in the above-mentioned character image is the same as that of the person in the above-mentioned first pose video image.

[0076] For example, the aforementioned character image can be selected by the user from the electronic device, i.e., uploaded by the user, or it can be a custom character image that is the default on the electronic device.

[0077] In some embodiments of this application, the electronic device uses the above-described limb-driven model to fuse the character image, M first semantic depth maps, and M first key point maps to generate M frames of first pose video images.

[0078] It is understandable that the pose in the first pose video image of the M frames is the same as the pose of the person in the M frame pose reference video image corresponding to the M first semantic depth maps and the M first key point maps, and is the same as the appearance of the person in the character image image.

[0079] Step 203: The electronic device generates a limb-driven video corresponding to the character image based on the first posture video image of frame M.

[0080] In some embodiments of this application, the electronic device stitches together M frames of first pose video images in a time sequence to generate a limb-driven video of the figure in the figure image image.

[0081] In some embodiments of this application, the electronic device can perform frame interpolation after stitching together the M-frame first posture video images in a time sequence, so that the resulting limb-driven video is smoother.

[0082] For example, to make the video frames in a limb-driven video appear more continuous, a video frame interpolation algorithm can be used to interpolate frames in the limb-driven video to increase the frame rate of the generated video. Additionally, background music can be extracted from a pose reference video and added to the limb-driven video to synthesize a complete video.

[0083] In the video generation method provided in this application embodiment, based on M frames of pose reference video images, M first semantic depth maps and M first keypoint maps are determined. The first semantic depth maps are used to indicate the depth information of at least one body part of a person in the pose reference video image, and the first keypoint maps are used to indicate the key points of the person's body in the pose reference video image. The person image image, the M first semantic depth maps, and the M first keypoint maps are input into a limb driving model, and M frames of first pose video images are output. Based on the M frames of first pose video images, a limb driving video corresponding to the person image image is generated. Each frame of the pose reference video image corresponds to one first semantic depth map and one first keypoint map, and the M first semantic depth maps and M first keypoint maps... Figure 1One-to-one correspondence. In this scheme, by inputting a semantic depth map representing the depth information of at least one body part of a person in the pose reference video image, and a first keypoint map representing the key points of the person's body in the pose reference video image, into the limb driving model, the limb driving model can generate a limb-driven video of the person's appearance in the character image based on the limb driving details in the semantic depth map and the first keypoint map. This can obtain a limb-driven video with free-moving movements without the need for motion capture technology and hardware devices, thereby improving the generation efficiency of digital human limb-driven videos.

[0084] Optionally, in some embodiments of this application, the limb-driven model described above includes an image encoder, a posture guidance module, a reference image fusion module, a diffusion module, and an image decoder.

[0085] For example, the image encoder model described above is used to encode images, that is, to extract image feature information from images.

[0086] For example, the posture guidance module described above is used to extract posture feature information from the posture reference video image.

[0087] For example, the above-mentioned pose guidance module includes at least two pose extraction layers, a first convolutional layer, and a second convolutional layer.

[0088] For example, the above image encoder model includes: a first convolution module, a second convolution module, and a third convolution module.

[0089] In one example, the first convolutional module includes a third convolutional layer, a fourth convolutional layer, and a fifth convolutional layer; the second convolutional module includes a third convolutional layer, a fourth convolutional layer, and a fifth convolutional layer; and the third convolutional module includes a third convolutional layer, a fourth convolutional layer, and a fifth convolutional layer.

[0090] In one example, the kernel of the third convolutional layer is 3×3, the stride is 2, and the number of output channels is 8; the kernel of the fourth convolutional layer is 3×3, the stride is 1, and the number of output channels is 8; the kernel of the fifth convolutional layer is 3×3, the stride is 1, and the number of output channels is 8.

[0091] For example, the ReferenceNet module described above is used to overlay the image feature information representing the shape of the reference person with the pose feature information in the pose reference video image.

[0092] For example, the above diffusion module is used to further fuse the pose feature information in detail.

[0093] For example, the image decoder described above is used to convert pose feature information into a final image.

[0094] In one example, the image decoder model described above includes: a first deconvolution module, a second deconvolution module, and a third deconvolution module.

[0095] In one example, the first deconvolution module includes a first deconvolution layer, a second deconvolution layer, and a third deconvolution layer; the second deconvolution module includes a first deconvolution layer, a second deconvolution layer, and a third deconvolution layer; and the third deconvolution module includes a first deconvolution layer, a second deconvolution layer, and a third deconvolution layer.

[0096] In one example, the first deconvolution layer has a 3×3 deconvolution kernel, a stride of 2, and 16 output channels; the second deconvolution layer has a 3×3 deconvolution kernel, a stride of 1, and 8 output channels; and the third deconvolution layer has a 3×3 deconvolution kernel, a stride of 1, and 8 output channels.

[0097] In one example, the image encoder, image decoder, and diffusion module described above can all reuse the module structure in the Singular Value Decomposition (SVD) model. For example, the image encoder can be a variational autoencoder, where the input to the image encoder is M consecutive video frames, and the image decoder can directly decode M pose-fused video frames.

[0098] Optionally, in some embodiments of this application, combined with Figure 1 ,like Figure 2 As shown, step 202 above, "The electronic device inputs the character image, M first semantic depth maps, and M first keypoint maps into the limb driving model, and outputs M frames of first pose video images," specifically includes steps 202a to 202f:

[0099] Step 202a: The electronic device inputs the character image, M first semantic depth maps and M first key point maps into the limb driving model.

[0100] In some embodiments of this application, the first semantic depth map is the result of superimposing the semantic map and depth map corresponding to the same video frame on the color channel. The input size of multiple first semantic depth maps is (B*M, 6, H, W), where B represents the batch size, i.e., the number of video segments, which is set to 1 in this embodiment, M represents the number of consecutive input video frames, and H and W represent the size of the input frames, which are usually set to 512.

[0101] Accordingly, the input size of the first key point map is (B*M,3,H,W).

[0102] Step 202b: The electronic device inputs the image of the person into the image encoder for feature extraction to obtain the first image feature information.

[0103] In some embodiments of this application, the aforementioned first image feature information is used to characterize the external features of the person in the figure image and the background features in the figure image.

[0104] For example, the physical characteristics of the figures in the above-mentioned character images may include, but are not limited to, at least one of the following: the figure of the figure in the character images, the clothing of the figure in the character images, and the hairstyle of the figure in the character images.

[0105] For example, the above-mentioned character image can be represented as [X,Y,Z], where X represents the number of color channels of the first image, and Y and Z represent the numerical values ​​of each pixel in the image.

[0106] For example, the aforementioned first image feature information can be represented in the form of a vector matrix.

[0107] Step 202c: The electronic device inputs M first semantic depth maps and M first key point maps into the pose guidance module for feature extraction to obtain first pose feature information.

[0108] In some embodiments of this application, the aforementioned first pose feature information is the pose feature information of M first semantic depth maps and the pose feature information of M first key point maps spliced ​​together.

[0109] Furthermore, in this embodiment of the application, step 202c specifically includes steps 202c1 to 202c3:

[0110] Step 202c1: The electronic device uses a first number of pose extraction layers and a first convolutional layer to extract pose feature information from M first semantic depth maps.

[0111] Step 202c2: The electronic device uses a second number of pose extraction layers and a second convolutional layer to extract pose feature information from M first keypoint maps.

[0112] In some embodiments of this application, the above-mentioned attitude extraction layer is a layer in the attitude guidance module.

[0113] In some embodiments of this application, the above-described pose extraction layer is used to extract pose feature information of people in an image.

[0114] In some embodiments of this application, the first quantity is greater than the second quantity.

[0115] For example, the first convolutional layer described above can be a convolutional layer with 3*3 convolutional kernels.

[0116] For example, given the different pixel density of the keypoint map and the semantic depth map mentioned earlier, this embodiment of the application can use one 3*3 convolution and two pose extraction layers (pose_module) to extract the pose feature information of the dense map, and one 3*3 convolution and one pose_module to extract the pose feature information of the sparse map. The dense map is the semantic depth map, and the sparse map is the keypoint map. The dense map has more underlying information, therefore more network layers are set, so that the fused dense and sparse action features are at closer feature levels.

[0117] In some embodiments of this application, the input to the pose guidance module includes a keypoint map and a semantic depth map. The significant differences in pixel distribution between these two pose guidance maps, namely the keypoint map and the semantic depth map, are fully considered. Specifically, the keypoint map, in its concise form, exhibits a relatively sparse pixel layout, while the semantic depth map presents richer and denser pixel details. Therefore, the pose guidance module employs different numbers of pose extraction layers for the keypoint map and the semantic depth map to extract their pose feature information respectively. For the relatively sparse keypoint map, a first number of pose extraction layers is used, while for the relatively dense semantic depth map, a second number of pose extraction layers are used. Therefore, the pose guidance module designed in this embodiment has the ability to efficiently encode and fuse features from pose guidance maps of different densities, thereby ensuring the precision and accuracy of pose control.

[0118] Step 202c3: The electronic device concatenates the pose feature information of the M first semantic depth maps and the pose feature information of the M first keypoint maps to obtain the first pose feature information.

[0119] In some embodiments of this application, the electronic device can further extract the pose feature information of the stitched M first semantic depth maps and M first key point maps, i.e., the aforementioned first pose feature information, through the pose extraction layer and convolutional layer.

[0120] For example, after the electronic device splices the pose feature information of the key point map and the semantic depth map together, it adds two pose_modules and one 1*1 convolution to further extract the fused pose feature information, namely the first pose feature information mentioned above.

[0121] In one example, the size of the pose feature vector corresponding to the output first pose feature information is (B*K, 320, H / 16, W / 16), and the specific network layer parameters are shown in Table 1.

[0122] The structure of the designed pose_module is as follows: Figure 3 As shown in Table 2, the specific network layer parameters are as follows.

[0123]

[0124]

[0125] Table 1 Parameters of Attitude Guidance Module

[0126]

[0127] Table 2 Parameters of the Attitude Extraction Layer

[0128] Step 202d: The electronic device inputs the first image feature information and the first pose feature information into the reference image fusion module for feature fusion to obtain the second pose feature information.

[0129] In some embodiments of this application, the above-described reference image fusion module includes at least one coding block and at least one zero convolutional layer.

[0130] Furthermore, in some embodiments of this application, step 202d specifically includes steps 202d1 to 202d3:

[0131] Step 202d1: The electronic device inputs the first image feature information into the reference image fusion module and extracts at least one shape feature information of the first image feature information through at least one coding block.

[0132] In some embodiments of this application, each of the at least one coding block described above is used to extract shape feature information of the corresponding level.

[0133] Step 202d2: The electronic device splices at least one shape feature information with each of the first posture feature information in the first posture feature information.

[0134] In some embodiments of this application, one shape feature information is spliced ​​together with M posture feature information in the first posture feature information.

[0135] It should be noted that the above M pose feature information refers to the pose feature information of the M semantic depth maps and M key point maps corresponding to the M frame pose reference video maps.

[0136] Step 202d3: The electronic device fuses the spliced ​​first pose feature information through at least one zero convolutional layer to obtain the second pose feature information.

[0137] In some embodiments of this application, the above-mentioned zero convolutional layer is used to fuse image feature information and pose feature information.

[0138] For example, assuming the size of the image of a person is B*C*H*W=1*3*512*512, the image of the person is first encoded into a latent space vector by the VAE encoder, that is, the image encoder mentioned above. The size of the feature vector corresponding to the first image feature information output is (B, 4, H / 8, W / 8).

[0139] Next, the feature vector corresponding to the first image feature information is input into ReferenceNet. In ReferenceNet, each coding block extracts shape feature information at different levels. For the shape feature information extracted by each coding block, it is first copied in the batch dimension, from (B,C,H,W) to (B*M,C,H,W). Then, the extracted first pose feature information is downsampled or upsampled to be scaled to the same shape size as the shape feature information of each layer, and then concatenated with the shape feature information. For example, after the first coding module is copied in the batch, the output is (B*M,320,H / 8,W / 8), and the size of the first pose feature information is (B*M,320,H / 16,W / 16). First, the first pose feature information is scaled to the size of (B*M,320,H / 8,W / 8) by the upsampling operation of dilated convolution, and then the two are concatenated together, with a size of (B*M,640,H / 8,W / 8).

[0140] Then, the features of the shape feature information and the first pose feature information are fused through a zero convolution layer to obtain the second pose feature information.

[0141] Step 202e: The electronic device inputs the second pose feature information into the diffusion module, uses preset noise, performs cross-attention mechanism and convolution calculation on the second pose feature information, and obtains the third pose feature information.

[0142] In some embodiments of this application, the preset noise can be a default setting of the electronic device or a user-selected setting.

[0143] For example, the preset noise mentioned above can be a random Gaussian noise matrix (GNM).

[0144] For example, the electronic device adds a random Gaussian noise matrix (GNM) to the second pose feature information to obtain the latent matrix of the image after noise addition. Then, the second pose feature information is combined with the output features of the corresponding spatial block in the image decoder through a jump connection. Finally, the pose feature information is converted into a pose video map to obtain the first pose video map of M frames.

[0145] Step 202f: The electronic device inputs the third posture feature information into the image decoder for image decoding and outputs M frames of the first posture video image.

[0146] In some embodiments of this application, the electronic device inputs third pose feature information into an image decoder, and then outputs a feature vector of the target dimension through a deconvolution module.

[0147] In some embodiments of this application, the electronic device converts the feature vector corresponding to the third posture feature information into an M-frame first posture video image for output.

[0148] In this way, by inputting a person image, M semantic depth maps, and M key point maps into the limb driving model, the electronic device outputs a posture video image of the reference person's shape and background. Without relying on motion capture technology and hardware devices, it can generate a limb driving video of the reference person, thereby improving the generation efficiency of the limb driving video of the reference person.

[0149] Optionally, in some embodiments of this application, the video generation method provided in this application further includes steps 401 and 402:

[0150] Step 401: The electronic device inputs the training samples into the limb driving model and outputs N frames of second pose video images.

[0151] In some embodiments of this application, the training samples include: N second semantic depth maps corresponding to N frames of pose training video images, N key point maps corresponding to N frames of pose training video images, and the first frame of pose training video image.

[0152] In some embodiments of this application, the posture training videos corresponding to the aforementioned N-frame posture training video images can be obtained by crawling a large number of individual videos of people dancing, exercising, etc., from social media platforms, and selecting videos with a duration between 0.5 minutes and 3 minutes. Additionally, videos that are too long are trimmed to a length of 2 minutes.

[0153] It is understood that the acquisition of the above training samples can refer to steps 301 and 302 above. The process of inputting the above training samples into the limb driving model and outputting N frames of second pose video images can refer to steps 202a to 202f above, and will not be repeated here.

[0154] Step 402: When the loss value between the N frames of the second posture video image and the N frames of the posture training video image is greater than or equal to the preset loss threshold, the electronic device trains the limb driving model based on the training loss function to obtain the trained limb driving model.

[0155] In some embodiments of this application, the preset loss threshold can be a default setting of the electronic device or a user setting. It is understood that, ideally, the preset loss threshold can be 0.

[0156] In some embodiments of this application, the training loss function is Loss = ||Input - Output||2.

[0157] In some embodiments of this application, the Input is N frames of pose training video images, and the Output is N frames of second pose video images output by the limb body model.

[0158] Understandably, the statement that "the loss value between the N frames of the second pose video image and the N frames of the pose training video image is greater than or equal to the preset loss threshold" indicates that the input image and the output image differ too much. Therefore, further training of the limb drive model is required.

[0159] For example, during model training, the network weights of the VAE's encoder and decoder modules reuse the weights of the SVD open-source model and remain fixed during training. This is to utilize the pre-trained video generation capabilities of the SVD open-source model. To adapt the model to the action-driven effect requirements of this scheme, the model weights of the pose_guider, referenceNet, and diffusion modules are first randomly initialized. Then, using the training data synthesized in step 102, the Adam optimization algorithm is used to iteratively update the model weights with the goal of minimizing the overall loss function. Through the updating of network weights, the limb-driven model can generate digital human action videos that maintain the shape of the reference image and have the same pose sequence as the action template.

[0160] It should be noted that the step of training the limb drive model can be performed before or after step 203 above, and this application does not impose any restrictions.

[0161] In this way, electronic devices improve the accuracy of the output image content of the limb driving model by continuously training the limb driving model according to the training loss function.

[0162] Each of the above-described method embodiments, or various possible implementations of each method embodiment, can be executed individually or in combination of any two or more. The specific implementation can be determined according to actual usage requirements, and this application does not impose any restrictions on this.

[0163] The video generation method provided in this application can be executed by an electronic device or a video generation apparatus. This application uses a video generation apparatus as an example to illustrate the video generation apparatus provided in this application.

[0164] Figure 4 A schematic diagram of a possible structure of the video generation apparatus involved in an embodiment of this application is shown. For example... Figure 4 As shown, the video generation device 700 may include: a determining module 701, a processing module 702, and a generation module 703;

[0165] The system comprises the following modules: a determining module 701, which determines M first semantic depth maps and M first keypoint maps based on M frames of pose reference video images; a first semantic depth map indicating the depth information of at least one body part of a person in the pose reference video image, and a first keypoint map indicating key points of the person's body in the pose reference video image; a processing module 702, which inputs the person image image, the M first semantic depth maps determined by the determining module 701, and the M first keypoint maps determined by the determining module 701 into a limb driving model, and outputs M frames of first pose video images; and a generating module 703, which generates a limb driving video corresponding to the person image image based on the M frames of first pose video images output by the processing module 702; wherein one frame of pose reference video image corresponds to one first semantic depth map and one first keypoint map, and the M first semantic depth maps and M first keypoint maps... Figure 1 One-to-one correspondence.

[0166] Optionally, in some embodiments of this application, combined with Figure 4 ,like Figure 5 As shown, the above-mentioned device 700 further includes: an acquisition module 704, used to acquire the semantic map corresponding to the M-frame posture reference video map and the depth map corresponding to the M-frame posture reference video map before the determination module 701 determines the M first semantic depth maps and the M first key point maps based on the M-frame posture reference video map; the above-mentioned processing module 702 is further used to stitch the semantic map and the depth map corresponding to each frame posture reference video map according to the color channel to obtain the M first semantic depth maps; wherein, the semantic map is used to represent at least one body part of the person in the posture reference video map, and the depth map is used to represent the depth information of the person's body in the posture reference video map.

[0167] Optionally, in some embodiments of this application, the above-mentioned acquisition module 704 is specifically used for:

[0168] Based on the key points of the human body referenced in the character image, a first human body model is constructed.

[0169] Based on the key points of the human body in the pose reference video image of each frame, construct M second human body models;

[0170] Based on the body shape parameters in the first human body model and the posture parameters in the M second human body models, M third human body models are obtained, and one third human body model corresponds to one frame of posture reference video image.

[0171] Render each of the M third-person human models to obtain the semantic map and the depth map corresponding to each frame of the pose reference video map.

[0172] Optionally, in some embodiments of this application, the above-mentioned limb driving model includes: an image encoder, a posture guidance module, a reference image fusion module, a diffusion module, and an image decoder;

[0173] The aforementioned processing module 702 is specifically used for:

[0174] The image of the person is input into the image encoder for feature extraction to obtain the first image feature information;

[0175] The M first semantic depth maps determined by the aforementioned determining module 701 and the M first key point maps determined by the aforementioned determining module 701 are input into the pose guidance module for feature extraction to obtain the first pose feature information.

[0176] The first image feature information and the first pose feature information are input into the reference image fusion module for feature fusion to obtain the second pose feature information;

[0177] The second pose feature information is input into the diffusion module, and a preset noise is used to perform cross-attention mechanism and convolution calculation on the second pose feature information to obtain the third pose feature information.

[0178] The third pose feature information is input into the image decoder for image decoding, and the M-frame first pose video image output by the above processing module 702 is output.

[0179] Optionally, in some embodiments of this application, the pose guidance module includes at least two pose extraction layers, a first convolutional layer, and a second convolutional layer; the processing module 702 is specifically used for:

[0180] Using a first number of pose extraction layers and a first convolutional layer, pose feature information of the M first semantic depth maps determined by the determination module 701 is extracted;

[0181] Using a second number of pose extraction layers and a second convolutional layer, the pose feature information of the M first keypoint maps determined by the determination module 701 is extracted;

[0182] The pose feature information of the M first semantic depth maps determined by the above-mentioned determining module 701 and the pose feature information of the M first key point maps determined by the above-mentioned determining module 701 are concatenated to obtain the first pose feature information.

[0183] The first quantity is greater than the second quantity.

[0184] Optionally, in some embodiments of this application, the above-mentioned reference image fusion module includes at least one coding block and at least one zero convolutional layer;

[0185] The aforementioned processing module 702 is specifically used for:

[0186] The first image feature information is input into the reference image fusion module, and at least one shape feature information of the first image feature information is extracted through at least one coding block;

[0187] At least one shape feature is concatenated with each of the first pose features in the first pose feature information;

[0188] The first pose feature information, after being stitched together, is fused through at least one zero convolutional layer to obtain the second pose feature information.

[0189] In the video generation apparatus provided in this application embodiment, based on M frames of pose reference video images, M first semantic depth maps and M first keypoint maps are determined. The first semantic depth maps are used to indicate the depth information of at least one body part of a person in the pose reference video images, and the first keypoint maps are used to indicate the key points of the person's body in the pose reference video images. The person image image, the M first semantic depth maps, and the M first keypoint maps are input into a limb driving model, and M frames of first pose video images are output. Based on the M frames of first pose video images, a limb driving video corresponding to the person image image is generated. Each frame of the pose reference video image corresponds to one first semantic depth map and one first keypoint map, and the M first semantic depth maps and M first keypoint maps... Figure 1 One-to-one correspondence. In this scheme, by inputting a semantic depth map representing the depth information of at least one body part of a person in the pose reference video image, and a first keypoint map representing the key points of the person's body in the pose reference video image, into the limb driving model, the limb driving model can generate a limb-driven video of the person's appearance in the character image based on the limb driving details in the semantic depth map and the first keypoint map. This can obtain a limb-driven video with free-moving movements without the need for motion capture technology and hardware devices, thereby improving the generation efficiency of digital human limb-driven videos.

[0190] The video generation device in this application embodiment can be an electronic device or a component within an electronic device, such as an integrated circuit or a chip. The electronic device can be a terminal or other devices besides a terminal. For example, the electronic device can be a mobile phone, tablet computer, laptop computer, PDA, in-vehicle electronic device, mobile internet device (MID), augmented reality (AR) / virtual reality (VR) device, robot, wearable device, ultra-mobile personal computer (UMPC), netbook, or personal digital assistant (PDA), etc. It can also be a server, network attached storage (NAS), personal computer (PC), television set (TV), ATM, or self-service machine, etc. This application embodiment does not specifically limit the device.

[0191] The video generation device in this application embodiment can be a device with an operating system. This operating system can be Android, iOS, or other possible operating systems; this application embodiment does not specifically limit the specific operating system used.

[0192] The video generation apparatus provided in this application embodiment can implement all the processes implemented in the video generation method embodiment, and will not be described again here to avoid repetition.

[0193] Optionally, such as Figure 6 As shown, this application embodiment also provides an electronic device 800, including a processor 801 and a memory 802. The memory 802 stores a program or instructions that can run on the processor 801. When the program or instructions are executed by the processor 801, they implement the various steps of the above-described video generation method embodiment and can achieve the same technical effect. To avoid repetition, they will not be described again here.

[0194] It should be noted that the electronic devices in the embodiments of this application include the mobile electronic devices and non-mobile electronic devices described above.

[0195] Figure 7 A schematic diagram of the hardware structure of an electronic device to implement an embodiment of this application.

[0196] The electronic device 100 includes, but is not limited to, components such as: radio frequency unit 101, network module 102, audio output unit 103, input unit 104, sensor 105, display unit 106, user input unit 107, interface unit 108, memory 109, and processor 110.

[0197] Those skilled in the art will understand that the electronic device 100 may also include a power supply (such as a battery) for supplying power to various components. The power supply may be logically connected to the processor 110 through a power management system, thereby enabling functions such as managing charging, discharging, and power consumption through the power management system. Figure 7 The electronic device structure shown does not constitute a limitation on the electronic device. The electronic device may include more or fewer components than shown, or combine certain components, or have different component arrangements, which will not be elaborated here.

[0198] The processor 110 is configured to determine M first semantic depth maps and M first keypoint maps based on M frames of pose reference video images. The first semantic depth maps indicate the depth information of at least one body part of a person in the pose reference video images, and the first keypoint maps indicate the key points of the person's body in the pose reference video images. The processor 110 is also configured to input the person image image, the M first semantic depth maps, and the M first keypoint maps into a limb driving model and output M frames of first pose video images. The processor 110 is also configured to generate a limb driving video corresponding to the person image image based on the M frames of first pose video images. Each frame of the pose reference video image corresponds to one first semantic depth map and one first keypoint map, and the M first semantic depth maps and M first keypoint maps... Figure 1 One-to-one correspondence.

[0199] Optionally, in some embodiments of this application, the processor 110 is further configured to obtain the semantic map and the depth map corresponding to the M frames of the pose reference video map before determining the M first semantic depth maps and the M first key point maps based on the M frames of pose reference video map; the processor 110 is further configured to stitch the semantic map and the depth map corresponding to each frame of the pose reference video map according to the color channel to obtain the M first semantic depth maps; wherein, the semantic map is used to represent at least one body part of the person in the pose reference video map, and the depth map is used to represent the depth information of the person's body in the pose reference video map.

[0200] Optionally, in some embodiments of this application, the processor 110 is specifically used for:

[0201] Based on the key points of the human body referenced in the character image, a first human body model is constructed.

[0202] Based on the key points of the human body in the pose reference video image of each frame, construct M second human body models;

[0203] Based on the body shape parameters in the first human body model and the posture parameters in the M second human body models, M third human body models are obtained, and one third human body model corresponds to one frame of posture reference video image.

[0204] Render each of the M third-person human models to obtain the semantic map and the depth map corresponding to each frame of the pose reference video map.

[0205] Optionally, in some embodiments of this application, the above-mentioned limb driving model includes: an image encoder, a posture guidance module, a reference image fusion module, a diffusion module, and an image decoder;

[0206] The aforementioned processor 110 is specifically used for:

[0207] The image of the person is input into the image encoder for feature extraction to obtain the first image feature information;

[0208] M first semantic depth maps and M first keypoint maps are input into the pose guidance module for feature extraction to obtain first pose feature information;

[0209] The first image feature information and the first pose feature information are input into the reference image fusion module for feature fusion to obtain the second pose feature information;

[0210] The second pose feature information is input into the diffusion module, and a preset noise is used to perform cross-attention mechanism and convolution calculation on the second pose feature information to obtain the third pose feature information.

[0211] The third pose feature information is input into the image decoder for image decoding, and M frames of the first pose video image are output.

[0212] Optionally, in some embodiments of this application, the above-mentioned pose guidance module includes at least two pose extraction layers, a first convolutional layer, and a second convolutional layer;

[0213] The aforementioned processor 110 is specifically used for:

[0214] Using a first number of pose extraction layers and a first convolutional layer, pose feature information of M first semantic depth maps is extracted;

[0215] A second number of pose extraction layers and a second convolutional layer are used to extract pose feature information from M first keypoint maps;

[0216] The pose feature information of M first semantic depth maps and the pose feature information of M first keypoint maps are concatenated to obtain the first pose feature information.

[0217] The first quantity is greater than the second quantity.

[0218] Optionally, in some embodiments of this application, the above-mentioned reference image fusion module includes at least one coding block and at least one zero convolutional layer;

[0219] The aforementioned processor 110 is specifically used for:

[0220] The first image feature information is input into the reference image fusion module, and at least one shape feature information of the first image feature information is extracted through at least one coding block;

[0221] At least one shape feature is concatenated with each of the first pose features in the first pose feature information;

[0222] The first pose feature information, after being stitched together, is fused through at least one zero convolutional layer to obtain the second pose feature information.

[0223] In the electronic device provided in this application embodiment, based on M frames of pose reference video images, M first semantic depth maps and M first keypoint maps are determined. The first semantic depth maps are used to indicate the depth information of at least one body part of a person in the pose reference video image, and the first keypoint maps are used to indicate the key points of the person's body in the pose reference video image. The person image image, the M first semantic depth maps, and the M first keypoint maps are input into a limb driving model, and M frames of first pose video images are output. Based on the M frames of first pose video images, a limb driving video corresponding to the person image image is generated. Each frame of the pose reference video image corresponds to one first semantic depth map and one first keypoint map, and the M first semantic depth maps and M first keypoint maps... Figure 1 One-to-one correspondence. In this scheme, by inputting a semantic depth map representing the depth information of at least one body part of a person in the pose reference video image, and a first keypoint map representing the key points of the person's body in the pose reference video image, into the limb driving model, the limb driving model can generate a limb-driven video of the person's appearance in the character image based on the limb driving details in the semantic depth map and the first keypoint map. This can obtain a limb-driven video with free-moving movements without the need for motion capture technology and hardware devices, thereby improving the generation efficiency of digital human limb-driven videos.

[0224] It should be understood that, in this embodiment, the input unit 104 may include a graphics processing unit (GPU) 1041 and a microphone 1042. The GPU 1041 processes image data of still images or videos obtained by an image capture device (such as a camera) in video capture mode or image capture mode. The display unit 106 may include a display panel 1061, which may be configured in the form of a liquid crystal display, an organic light-emitting diode, or the like. The user input unit 107 includes at least one of a touch panel 1071 and other input devices 1072. The touch panel 1071 is also called a touch screen. The touch panel 1071 may include a touch detection device and a touch controller. Other input devices 1072 may include, but are not limited to, a physical keyboard, function keys (such as volume control buttons, power buttons, etc.), a trackball, a mouse, and a joystick, which will not be described in detail here.

[0225] The memory 109 can be used to store software programs and various data. The memory 109 may primarily include a first storage area for storing programs or instructions and a second storage area for storing data. The first storage area may store the operating system, application programs or instructions required for at least one function (such as sound playback, image playback, etc.). Furthermore, the memory 109 may include volatile memory or non-volatile memory, or both. The non-volatile memory may be read-only memory (ROM), programmable read-only memory (PROM), erasable programmable read-only memory (EPROM), electrically erasable programmable read-only memory (EEPROM), or flash memory. Volatile memory can be random access memory (RAM), static random access memory (SRAM), dynamic random access memory (DRAM), synchronous dynamic random access memory (SDRAM), double data rate synchronous dynamic random access memory (DDRSDRAM), enhanced synchronous dynamic random access memory (ESDRAM), synchronous link dynamic random access memory (SLDRAM), and direct memory bus RAM (DRRAM). The memory 109 in the embodiments of this application includes, but is not limited to, these and any other suitable types of memory.

[0226] Processor 110 may include one or more processing units; optionally, processor 110 integrates an application processor and a modem processor, wherein the application processor mainly handles operations involving the operating system, user interface, and applications, and the modem processor mainly handles wireless communication signals, such as a baseband processor. It is understood that the aforementioned modem processor may also not be integrated into processor 110.

[0227] This application also provides a readable storage medium storing a program or instructions. When the program or instructions are executed by a processor, they implement the various processes of the above-described video generation method embodiments and achieve the same technical effect. To avoid repetition, they will not be described again here.

[0228] The processor is the processor in the electronic device described in the above embodiments. The readable storage medium includes computer-readable storage media, such as computer read-only memory (ROM), random access memory (RAM), magnetic disk, or optical disk.

[0229] This application embodiment also provides a chip, which includes a processor and a communication interface. The communication interface is coupled to the processor. The processor is used to run programs or instructions to implement the various processes of the above video generation method embodiments and can achieve the same technical effect. To avoid repetition, it will not be described again here.

[0230] It should be understood that the chip mentioned in the embodiments of this application may also be referred to as a system-on-a-chip, system chip, chip system, or system-on-a-chip, etc.

[0231] This application provides a computer program product, which is stored in a storage medium and executed by at least one processor to implement the various processes of the video generation method embodiments described above, and can achieve the same technical effect. To avoid repetition, it will not be described again here.

[0232] It should be noted that, in this document, the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such a process, method, article, or apparatus. Without further limitations, an element defined by the phrase "comprising one..." does not exclude the presence of other identical elements in the process, method, article, or apparatus that includes that element. Furthermore, it should be noted that the scope of the methods and apparatuses in the embodiments of this application is not limited to performing functions in the order shown or discussed, but may also include performing functions substantially simultaneously or in the reverse order, depending on the functions involved. For example, the described methods may be performed in a different order than described, and various steps may be added, omitted, or combined. Additionally, features described with reference to certain examples may be combined in other examples.

[0233] Through the above description of the embodiments, those skilled in the art can clearly understand that the methods of the above embodiments can be implemented by means of software plus necessary general-purpose hardware platforms. Of course, they can also be implemented by hardware, but in many cases the former is a better implementation method. Based on this understanding, the technical solution of this application, in essence, or the part that contributes to the prior art, can be embodied in the form of a computer software product. This computer software product is stored in a storage medium (such as ROM / RAM, magnetic disk, optical disk) and includes several instructions to cause a terminal (which may be a mobile phone, computer, server, or network device, etc.) to execute the methods described in the various embodiments of this application.

[0234] The embodiments of this application have been described above with reference to the accompanying drawings. However, this application is not limited to the specific embodiments described above. The specific embodiments described above are merely illustrative and not restrictive. Those skilled in the art can make many other forms under the guidance of this application without departing from the spirit and scope of the claims, and all of these forms are within the protection scope of this application.

Claims

1. A method of video generation, the method comprising: The method comprises: determining M first semantic depth maps and M first key point maps based on M frame posture reference video graphs, the first semantic depth map being used to indicate depth information of at least one body part of a character in the posture reference video graph, and the first key point map being used to indicate body key points of the character in the posture reference video graph; inputting the character image graph, the M first semantic depth maps and the M first key point maps into a limb driving model to output M frame first posture video graphs; generating a limb driving video corresponding to the character image graph according to the M frame first posture video graphs; wherein one frame of the posture reference video graph corresponds to one first semantic depth map and one first key point map, and the M first semantic depth maps and the M first key point maps correspond to each other in a one-to-one manner; before the determining of the M first semantic depth maps and the M first key point maps based on the M frame posture reference video graphs, the method further comprises: obtaining semantic graphs corresponding to the M frame posture reference video graphs and depth graphs corresponding to the M frame posture reference video graphs; splicing the semantic graphs and the depth graphs corresponding to each frame of the posture reference video graph according to a color channel to obtain the M first semantic depth maps; wherein the semantic graph is used to represent at least one body part of a character in the posture reference video graph, and the depth graph is used to represent depth information of the body of the character in the posture reference video graph; the obtaining of the semantic graphs corresponding to the M frame posture reference video graphs and the depth graphs corresponding to the M frame posture reference video graphs comprises: constructing a first human body model based on body key points of a character in a character image graph; constructing M second human body models based on the M frame posture reference video graphs, one second human body model being constructed based on body key points of a character in one frame of the posture reference video graph; obtaining M third human body models based on body shape parameters in the first human body model and posture parameters in the M second human body models, one third human body model corresponding to one frame of the posture reference video graph; respectively rendering the M third human body models to obtain a semantic graph corresponding to each frame of the M frame posture reference video graphs and the depth graph corresponding to each frame of the M frame posture reference video graphs.

2. The method of claim 1, wherein, The limb driving model comprises an image encoder, a posture guide module, a reference image fusion module, a diffusion module and an image decoder; the inputting of the character image graph, the M first semantic depth maps and the M first key point maps into the limb driving model to output the M frame first posture video graphs comprises: inputting the character image graph into the image encoder for feature extraction to obtain first image feature information; inputting the M first semantic depth maps and the M first key point maps into the posture guide module for feature extraction to obtain first posture feature information; inputting the first image feature information and the first posture feature information into the reference image fusion module for feature fusion to obtain second posture feature information; The second posture feature information is input into the diffusion module, preset noise is used, cross attention mechanism and convolution calculation are performed on the second posture feature information, and third posture feature information is obtained. The third posture feature information is input into the image decoder for image decoding, and the M first posture video graphs are output.

3. The method of claim 2, wherein, The posture guiding module comprises at least two posture extraction layers, a first convolution layer and a second convolution layer. The M first semantic depth graphs and the M first key point graphs are input into the posture guiding module for feature extraction, and first posture feature information is obtained, comprising: A first number of posture extraction layers and the first convolution layer are used to extract posture feature information of the M first semantic depth graphs; A second number of posture extraction layers and the second convolution layer are used to extract posture feature information of the M first key point graphs; The posture feature information of the M first semantic depth graphs and the posture feature information of the M first key point graphs are correspondingly spliced to obtain the first posture feature information. The first number is greater than the second number.

4. The method of claim 2, wherein, The reference image fusion module comprises at least one encoding block and at least one zero convolution layer. The first image feature information and the first posture feature information are input into the reference image fusion module for feature fusion to obtain second posture feature information, comprising: The first image feature information is input into the reference image fusion module, and at least one shape feature information of the first image feature information is extracted through the at least one encoding block; The at least one shape feature information is spliced with each of the first posture feature information; The spliced first posture feature information is fused through the at least one zero convolution layer to obtain the second posture feature information.

5. A video generation apparatus characterized by comprising: The video generation device comprises: A determination module configured to determine, based on M posture reference video graphs, M first semantic depth graphs and M first key point graphs, the first semantic depth graph being configured to indicate depth information of at least one body part of a character in the posture reference video graph, and the first key point graph being configured to indicate a body key point of the character in the posture reference video graph; A processing module configured to input a character image, the M first semantic depth graphs determined by the determination module and the M first key point graphs determined by the determination module into a limb driving model, and output M first posture video graphs; A generation module configured to generate a limb driving video corresponding to the character image according to the M first posture video graphs output by the processing module. Each posture reference video graph corresponds to one first semantic depth graph and one first key point graph, and the M first semantic depth graphs and the M first key point graphs correspond one by one. The device further comprises: An acquisition module configured to acquire semantic graphs corresponding to the M posture reference video graphs and depth graphs corresponding to the M posture reference video graphs before the determination module determines, based on the M posture reference video graphs, the M first semantic depth graphs and the M first key point graphs. The processing module is further configured to: splice the semantic map and the depth map corresponding to each frame of the posture reference video according to a color channel, to obtain M first semantic depth maps; The semantic map is used to represent at least one body part of a character in the posture reference video, and the depth map is used to represent depth information of a body of the character in the posture reference video. The obtaining module is specifically configured to: construct a first human body model based on body key points of the character in the character image; construct M second human body models based on the M frames of posture reference videos, and one second human body model is constructed based on body key points of the character in one frame of posture reference video; obtain M third human body models based on body shape parameters in the first human body model and posture parameters in the M second human body models, and one third human body model corresponds to one frame of posture reference video; render the M third human body models respectively, to obtain a semantic map corresponding to each frame of the M frames of posture reference videos and the depth map corresponding to each frame of the M frames of posture reference videos.

6. The apparatus of claim 5, wherein, The limb driving model comprises an image encoder, a posture guide module, a reference image fusion module, a diffusion module and an image decoder. The processing module is specifically configured to: input the character image into the image encoder to extract features, to obtain first image feature information; input the M first semantic depth maps determined by the determining module and the M first key point maps determined by the determining module into the posture guide module to extract features, to obtain first posture feature information; input the first image feature information and the first posture feature information into the reference image fusion module to fuse features, to obtain second posture feature information; input the second posture feature information into the diffusion module, adopt a preset noise, and perform cross-attention mechanism and convolution calculation on the second posture feature information, to obtain third posture feature information; input the third posture feature information into the image decoder to perform image decoding, and output the M frames of first posture video determined by the processing module.

7. The apparatus of claim 6, wherein, The posture guide module comprises at least two posture extraction layers, a first convolution layer and a second convolution layer. The processing module is specifically configured to: adopt a first number of posture extraction layers and a first convolution layer to extract posture feature information of the M first semantic depth maps determined by the determining module; adopt a second number of posture extraction layers and a second convolution layer to extract posture feature information of the M first key point maps determined by the determining module; splice the posture feature information of the M first semantic depth maps determined by the determining module and the posture feature information of the M first key point maps determined by the determining module, to obtain the first posture feature information; The first number is greater than the second number.

8. The apparatus of claim 6, wherein, The reference image fusion module comprises at least one encoding block and at least one zero convolution layer. The processing module is specifically configured to: inputting the first image feature information into the reference image fusion module, and extracting at least one contour feature information of the first image feature information through the at least one encoding block; splicing the at least one contour feature information and each of the first pose feature information; fusing the spliced first pose feature information through the at least one zero convolution layer to obtain the second pose feature information.

9. An electronic device, comprising: The device comprises a processor, a memory, and a program or instruction stored on the memory and executable on the processor, and the program or instruction is executed by the processor to implement the steps of the video generation method according to any one of claims 1 to 4.

10. A readable storage medium, characterized by, The readable storage medium stores a program or instruction, and the program or instruction is executed by the processor to implement the steps of the video generation method according to any one of claims 1 to 4.

Citation Information

Patent Citations

  • Video generation method and device

    CN110245638A

  • Image processing method and apparatus, and device

    WO2024109522A1