Video generation method, device, computer device and storage medium

By performing action-driven and video fusion on different parts of the human body image, the problem of low generation efficiency in the prior art is solved, and a rapid generation of natural and real virtual human videos is achieved.

CN114612595BActive Publication Date: 2025-07-18CHINA PING AN LIFE INSURANCE CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202210225839.X
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-03-07
Publication Date
2025-07-18
Estimated Expiration
2042-03-07

AI Technical Summary

Technical Problem

The existing virtual human video generation technology requires the construction of three-dimensional animations, which are cumbersome and cannot be adapted to other user images, resulting in low generation efficiency.

Method used

By activating different parts in the human body image according to the human body posture video, the first action video and the second action video are generated, and the first action video and the second action video are fused, so as to quickly generate the character action video.

Benefits of technology

The efficiency of human body action video generation is improved, and the characters' actions in the generated video are more natural and realistic, and are suitable for different user images.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114612595B_ABST
    Figure CN114612595B_ABST
Patent Text Reader

Abstract

This application relates to the field of artificial intelligence. By respectively performing action driving on the first part and the second part in the human body image according to the human body posture video, and performing fusion processing on the obtained first action video and the second action video, it realizes the rapid generation of a human action video using a single-frame human body image, with convenient operation and improved efficiency of human action video generation. It relates to a video generation method, device, computer device, and storage medium. The method includes: obtaining a human body image; performing action driving on the first part in the human body image according to the human body posture video to obtain a first action video corresponding to the human body image; performing action driving on the second part in the human body image according to the human body posture video to obtain a second action video corresponding to the human body image; fusing the first action video and the second action video to obtain a target human action video corresponding to the human body image. In addition, this application also relates to blockchain technology, and the human body image can be stored in the blockchain.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to artificial intelligence, and particularly to a video generation method, apparatus, computer device, and storage medium. Background Art

[0002] With the rapid growth of short video content consumption, quickly creating virtual human videos has become a typical demand. Existing virtual human video generation technologies generally generate based on 3D models such as 3Dmax. However, when using such 3D models to generate virtual human videos, it is necessary to construct 3D animations for each user's actions, which is cumbersome, time-consuming, and the constructed 3D animations cannot be adapted to the images of other users, reducing the efficiency of video generation.

[0003] Therefore, how to improve the efficiency of video generation has become an urgent problem to be solved. Summary of the Invention

[0004] This application provides a video generation method, apparatus, computer device, and storage medium. By respectively performing action driving on the first part and the second part in the human body image according to the human body posture video, and performing fusion processing on the obtained first action video and second action video, it realizes quickly generating a human action video using a single-frame human body image, improving the efficiency of human action video generation.

[0005] In a first aspect, this application provides a video generation method, the method comprising:

[0006] Obtaining a human body image of the video to be generated;

[0007] Performing action driving on the first part in the human body image according to a preset human body posture video to obtain a first action video corresponding to the human body image;

[0008] Performing action driving on the second part in the human body image according to the human body posture video to obtain a second action video corresponding to the human body image;

[0009] Fusing the first action video and the second action video to obtain a target human action video corresponding to the human body image.

[0010] In a second aspect, this application further provides a video generation apparatus, the apparatus comprising:

[0011] A human body image acquisition module, configured to obtain a human body image of the video to be generated;

[0012] A first action driving module, configured to perform action driving on the first part in the human body image according to a preset human body posture video to obtain a first action video corresponding to the human body image;

[0013] A second motion driving module, configured to perform motion driving on a second part in the human body image according to the human body posture video, so as to obtain a second motion video corresponding to the human body image;

[0014] A video fusion module, configured to fuse the first motion video and the second motion video to obtain a target human body motion video corresponding to the human body image.

[0015] In a third aspect, the present application further provides a computer device, which includes a memory and a processor;

[0016] The memory is used for storing a computer program;

[0017] The processor is configured to execute the computer program and implement the video generation method as described above when executing the computer program.

[0018] In a fourth aspect, the present application further provides a computer-readable storage medium, which stores a computer program, and when the computer program is executed by a processor, the processor is caused to implement the video generation method as described above.

[0019] The present application discloses a video generation method, device, computer device and storage medium. By performing motion driving on a first part in a human body image according to a preset human body posture video to obtain a first motion video corresponding to the human body image, motion driving on the head in the human body image can be realized, the motion details of the face are retained, and the situation of missing motion details is avoided; by performing motion driving on a second part in the human body image according to the human body posture video to obtain a second motion video corresponding to the human body image, motion driving on the limbs and torso in the human body image can be realized, so that the characters in the subsequent generated human body motion video are more natural and real; by fusing the first motion video and the second motion video to obtain a target human body motion video corresponding to the human body image, it is possible to quickly generate a human body motion video using a single-frame human body image, with convenient operation and improved efficiency of generating the human body motion video. BRIEF DESCRIPTION OF THE DRAWINGS

[0020] In order to more clearly illustrate the technical solutions of the embodiments of the present application, the following will briefly introduce the drawings required for the description of the embodiments. Obviously, the following drawings are some embodiments of the present application. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.

[0021] Figure 1 It is a schematic flowchart of a video generation method provided by an embodiment of the present application;

[0022] Figure 2It is a schematic diagram of generating a first action video provided by an embodiment of the present application;

[0023] Figure 3 It is a schematic flowchart of sub-steps of generating a first action video provided by an embodiment of the present application;

[0024] Figure 4 It is a schematic flowchart of sub-steps of generating a non-head action video provided by an embodiment of the present application;

[0025] Figure 5 It is a schematic block diagram of a video generation device provided by an embodiment of the present application;

[0026] Figure 6 It is a schematic block diagram of the structure of a computer device provided by an embodiment of the present application. Detailed implementation manners

[0027] Next, the technical solutions in the embodiments of the present application will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present application. Obviously, the described embodiments are part of the embodiments of the present application, rather than all the embodiments. Based on the embodiments in the present application, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of the present application.

[0028] The flowcharts shown in the accompanying drawings are only illustrative examples, and do not necessarily include all the contents and operations / steps, nor do they necessarily need to be executed in the described order. For example, some operations / steps can also be decomposed, combined, or partially merged. Therefore, the actual execution order may be changed according to the actual situation.

[0029] It should be understood that the terms used in this specification of the present application are only for the purpose of describing specific embodiments and are not intended to limit the present application. As used in this specification of the present application and the appended claims, unless the context clearly indicates otherwise, the singular forms "a", "an", and "the" are intended to include the plural forms.

[0030] It should also be understood that the term "and / or" used in this specification of the present application and the appended claims refers to any combination and all possible combinations of one or more of the associated listed items, and includes these combinations.

[0031] Embodiments of the present application provide a video generation method, apparatus, computer device, and storage medium. Among them, the video generation method can be applied to a server or a terminal. By respectively performing action driving on a first part and a second part in a human body image according to a human body posture video, and performing fusion processing on the obtained first action video and second action video, it is possible to quickly generate a human action video using a single-frame human body image, with convenient operation and improved efficiency of human action video generation.

[0032] Among them, the server can be an independent server or a cloud server that provides basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communications, middleware services, domain name services, security services, Content Delivery Network (CDN), and big data and artificial intelligence platforms. The terminal can be an electronic device such as a smart phone, a tablet computer, a notebook computer, and a desktop computer.

[0033] The following will describe in detail some embodiments of the present application with reference to the accompanying drawings. Without conflict, the following embodiments and the features in the embodiments can be combined with each other.

[0034] As Figure 1 shown, the video generation method includes steps S10 to S40.

[0035] Step S10, obtain a human body image of the video to be generated.

[0036] Exemplarily, an image uploaded or selected by a user can be determined as the human body image of the video to be generated. Among them, the human body image can be a full-body image, including the head, limbs, and torso.

[0037] To further ensure the privacy and security of the above human body image, the above human body image can be stored in a node of a blockchain. When video generation is to be performed, the human body image can be obtained from the blockchain node.

[0038] In the embodiments of the present application, by obtaining the human body image, it is possible to subsequently quickly generate a human action video using a single-frame human body image, without the need for the user to manually construct a 3D model, with convenient operation and improved efficiency of human action video generation.

[0039] Step S20, perform action driving on a first part in the human body image according to a preset human body posture video to obtain a first action video corresponding to the human body image.

[0040] Exemplarily, a preset human body pose video can be obtained from a local disk or a database. Among them, the human body pose video is used as a driving video to transfer the actions in the human body pose video to the human body image to generate a corresponding target human body action video. The person in the human body image and the person in the human body pose video can be the same person or different persons.

[0041] It should be noted that the human body pose video refers to a video including actions or poses. For example, a self-introduction video, a business explanation video, a course explanation video, and so on.

[0042] Exemplarily, the solution of the embodiment of the present application can be applied to the scenario of generating a virtual human explanation video. Of course, it can also be applied to other scenarios. For example, by obtaining a human body image provided by a user and a business explanation video recorded by the user, driving the head in the human body image according to the business explanation video to obtain a first action video corresponding to the human body image, and driving the limbs and torso in the human body image according to the business explanation video to obtain a second action video corresponding to the human body image; fusing the first action video and the second action video to obtain a virtual human business explanation video corresponding to the user.

[0043] It should be noted that in order to retain the action details in the generated human body action video, in the embodiment of the present application, the head in the human body image will be driven to obtain a first action video corresponding to the human body image, and the limbs and torso in the human body image will be driven to obtain a second action video corresponding to the human body image; then, the first action video and the second action video will be fused to obtain a target human body action video corresponding to the human body image.

[0044] By driving the first part in the human body image according to a preset human body pose video to obtain a first action video corresponding to the human body image, it is possible to drive the head in the human body image, retain the action details of the face, and avoid the situation of missing action details.

[0045] In some embodiments, driving the first part in the human body image according to a preset human body pose video to obtain a first action video corresponding to the human body image may include: determining a head region image corresponding to the human body image, and determining a head action video corresponding to the human body pose video; inputting the head action video and the head region image into an action driving model for action driving to obtain the first action video. The first part refers to the head.

[0046] Please refer to Figure 2 , Figure 2 which is a schematic diagram of generating a first action video provided by an embodiment of the present application. As Figure 2As shown, when determining the head region image corresponding to the human body image, face detection can be performed on the human body image according to a face detection algorithm to determine the face region in the human body image; the face region in the human body image is cropped, and the cropped image is determined as the head region image. When determining the head action video corresponding to the human body pose video, the video of the head region in the human body pose video can be cropped to obtain the head action video. Then, the head action video and the head region image are input into an action driving model for action driving to obtain a first action video.

[0047] Among them, the face detection algorithm can include but is not limited to a face detection algorithm based on histogram rough segmentation and singular value features, a face detection algorithm based on binary wavelet transform, a face detection algorithm based on AdaBoost algorithm, and a face detection algorithm based on facial binocular structure features, etc.

[0048] In the embodiment of the present application, the action driving model can be the First Order Motion Model. It should be noted that the First Order Motion Model is used to generate a target video according to the input source image and driving video. Among them, the protagonist in the target video is the source image, and the action in the target video is the action in the driving video. The First Order Motion Model includes a keypoint detector, a motion estimation module, and an image generation module. Among them, the keypoint detector is used to detect the keypoints in the image and the corresponding jaccobian matrix for each keypoint; the motion estimation module is used to generate a final transform map and an occlusion map based on the previous results; the image generation module is used to perform transformation and mask processing on the encoded source image according to the transform map and occlusion map, and then decode to generate the final result.

[0049] Please refer to Figure 3 , Figure 3 which is a schematic flowchart of the sub-steps for generating the first action video provided by the embodiment of the present application, and specifically may include the following steps S201 to step S204.

[0050] Step S201: Input the head action video and the head region image into the keypoint detector for processing to obtain first keypoint information corresponding to the head region image and second keypoint information corresponding to the head action video.

[0051] Exemplarily, the head motion video and the head region image can be input into the keypoint detector, and the keypoint detector outputs the first keypoint information corresponding to the head region image and the second keypoint information corresponding to the head motion video. Among them, the first keypoint information represents the mapping relationship from the reference frame to the head region image, and the second keypoint information represents the mapping relationship from the reference frame to the head motion video.

[0052] It can be understood that, in order to facilitate obtaining the mapping relationship between the head motion video and the head region image, a reference frame can be introduced so that the mapping relationship from the reference frame to the head region image and the mapping relationship from the reference frame to the head motion video can be independently estimated.

[0053] Exemplarily, the reference frame can be denoted as R; the first keypoint information can be denoted as T S←R ; the second keypoint information can be denoted as T D←R .

[0054] Step S202, determine the affine transformation matrix corresponding to the first keypoint information and the second keypoint information.

[0055] In some embodiments, determining the affine transformation matrix corresponding to the first keypoint information and the second keypoint information may include: taking the derivative of the first keypoint information to obtain the first derivative corresponding to the first keypoint information, and taking the derivative of the second keypoint information to obtain the second derivative corresponding to the second keypoint information; generating the affine transformation matrix from the ratio of the first derivative to the second derivative.

[0056] Exemplarily, the first derivative can be denoted as The second derivative can be denoted as Generating the affine transformation matrix from the ratio of the first derivative to the second derivative, the generated affine transformation matrix can be denoted as where p k is the keypoint position on the reference frame R.

[0057] Step S203, input the head region image, the first keypoint information, the second keypoint information, and the affine transformation matrix into the motion estimator for motion estimation processing to obtain the corresponding mapping relationship graph and occlusion graph.

[0058] Exemplarily, the head region image, the first keypoint information, the second keypoint information, and the affine transformation matrix can be input into the motion estimator for motion estimation processing, and the motion estimator outputs the mapping relationship graph and the occlusion graph. Among them, the specific motion estimation processing process is not limited herein.

[0059] Exemplarily, the mapping relationship diagram can be denoted as T S←D ; the occlusion mask can be denoted as O S←D . Among them, the mapping relationship diagram T S←D is obtained by the following formula:

[0060] T S←D (z)≈T S←R (p k )+J k (z - T D←R (p k ))

[0061] In the formula, z represents the key point in the head region image.

[0062] It should be noted that the mapping relationship diagram represents the mapping relationship of the key points in the head action video to the key points in the head region image. The occlusion mask represents which parts can be distorted by the head action video and which parts can be obtained by inpainting in the finally generated image.

[0063] Step S204, input the mapping relationship diagram, the occlusion mask, and the head region image into the image generator for image generation to obtain the first action video.

[0064] Exemplarily, the image generator can include an encoder and a decoder.

[0065] In some embodiments, inputting the mapping relationship diagram, the occlusion mask, and the head region image into the image generator for image generation to obtain the first action video may include: performing feature encoding on the head region image through the encoder to obtain an intermediate feature vector; performing an affine transformation on the intermediate feature vector according to the mapping relationship diagram to obtain an affine-transformed intermediate feature vector; multiplying the affine-transformed intermediate feature vector with the occlusion mask to obtain a feature vector map; and performing image reconstruction on the feature vector map through the decoder to obtain the first action video.

[0066] It should be noted that an affine transformation refers to a composition of a non-singular linear transformation followed by a translation transformation.

[0067] Exemplarily, the head region image can be feature-encoded through the encoder encoder to obtain a corresponding intermediate feature vector; then, an affine transformation is performed on the intermediate feature vector according to the mapping relationship diagram to obtain an affine-transformed intermediate feature vector, and the affine-transformed intermediate feature vector is multiplied with the occlusion mask to obtain a feature vector map; finally, the feature vector map is image-reconstructed through the decoder decoder to obtain the first action video.

[0068] It should be noted that by performing an affine transformation on the intermediate feature vector according to the mapping relationship diagram, the intermediate feature vector after the affine transformation can be obtained, which can load the mapping relationship between the key points in the head motion video and the key points in the head region image into the feature vector diagram, and then the action in the head motion video can be transferred to the first action video. By multiplying the intermediate feature vector after the affine transformation with the occlusion map, the feature vector map can be obtained, and the key points that need to be repaired during image reconstruction can be determined through the feature vector map.

[0069] By inputting the head motion video and the head region image into the action driving model for action driving, the first action video can be obtained conveniently and quickly without the user manually constructing a 3D model, and it can be applied to images of different user images, improving the efficiency of generating human action videos.

[0070] Step S30: Perform action driving on the second part in the human body image according to the human body pose video to obtain the second action video corresponding to the human body image.

[0071] It should be noted that in the embodiments of the present application, by performing head action driving and non-head action driving respectively, the action details of the head and non-head can be maximally retained, and the fineness of the target human action video can be effectively improved.

[0072] Exemplarily, the second part may include non-head regions, such as the limbs and torso of the human body; the second action video includes non-head action videos.

[0073] Please refer to Figure 4 , Figure 4 which is a schematic flowchart of a sub-step for generating a non-head action video provided by the embodiments of the present application, and specifically may include the following steps S301 to step S303.

[0074] Step S301: Clip the head region in the human body pose video to obtain the human body pose video after clipping the head region.

[0075] Exemplarily, the head region in the human body pose video can be clipped to obtain the human body pose video after clipping the head region.

[0076] Step S302: Clip the head region in the human body image to obtain the human body image after clipping the head region.

[0077] Exemplarily, the head region in the human body image can be clipped to obtain the human body image after clipping the head region. Of course, the regions where the limbs and torso of the human body image are located can also be cropped, and the cropped image can be determined as the human body image after clipping the head region.

[0078] Step S303: Input the human body posture video after head region shearing and the human body image after head region shearing into the action driving model for action driving to obtain the non-head action video.

[0079] Exemplarily, the human body posture video after head region shearing and the human body image after head region shearing can be input into the action driving model for action driving to obtain the non-head action video. Among them, the action driving model can be the FirstOrder Motion Model.

[0080] It can be understood that inputting the human body posture video after head region shearing and the human body image after head region shearing into the action driving model for action driving is equivalent to performing non-head action driving. In the embodiment of the present application, the specific process of performing non-head action driving is similar to the process of performing head action driving in the above embodiment. The specific process is as follows: Input the human body posture video after head region shearing and the human body image after head region shearing into the key point detector for processing to obtain the third key point information corresponding to the human body posture video after head region shearing and the fourth key point information corresponding to the human body image after head region shearing; determine the affine transformation matrix corresponding to the third key point information and the fourth key point information; input the human body image after head region shearing, the third key point information, the fourth key point information, and the affine transformation matrix into the motion estimator for motion estimation processing to obtain the corresponding mapping relationship diagram and occlusion diagram; input the mapping relationship diagram, the occlusion diagram, and the human body image after head region shearing into the image generator for image generation to obtain the non-head action video.

[0081] By inputting the human body posture video after head region shearing and the human body image after head region shearing into the action driving model for action driving, not only can the non-head action video be obtained conveniently and quickly, but also the limbs and torso in the human body image can be driven for actions, making the characters in the subsequent generated human body action video more natural and real.

[0082] Step S40: Fuse the first action video and the second action video to obtain the target human body action video corresponding to the human body image.

[0083] In some embodiments, fusing the first action video and the second action video to obtain the target human body action video corresponding to the human body image may include: aligning the images of the first action video and the second action video, and splicing each pair of aligned images to obtain the target human body action video.

[0084] It should be noted that image alignment means aligning each frame of the images in the first action video with each frame of the images in the second action video.

[0085] Exemplarily, when aligning the images of the first action video and the second action video, each image in the first action video can be numbered, and each image in the second action video can be numbered; then, the first image and the second image with the same number are aligned to obtain multiple pairs of images, where the first image is an image in the first action video and the second image is an image in the second action video; finally, each pair of aligned images is stitched to obtain the target human action video.

[0086] Exemplarily, an image stitching and fusion algorithm can be used to stitch each pair of images to obtain the target human action video. Among them, the image stitching and fusion algorithm can include, but is not limited to, the image stitching algorithm based on SURF, the image fusion algorithm based on gradient pyramid decomposition, and so on.

[0087] By fusing the first action video and the second action video to obtain the target human action video corresponding to the human image, it is possible to quickly generate a human action video using a single-frame human image, which is convenient to operate and improves the efficiency of generating the human action video.

[0088] The video generation method provided in the above embodiments can drive the action of the first part in the human image according to the preset human pose video to obtain the first action video corresponding to the human image, which can realize driving the action of the head in the human image, retain the action details of the face, and avoid the situation of missing action details; by inputting the head action video and the head region image into the action driving model for action driving, the first action video can be obtained conveniently and quickly without the user manually constructing a 3D model, and it can be applied to images of different user images, improving the efficiency of generating the human action video; by inputting the human pose video after cropping the head region and the human image after cropping the head region into the action driving model for action driving, not only can the non-head action video be obtained conveniently and quickly, but also the action of the limbs and torso in the human image can be driven, making the characters in the subsequent generated human action video more natural and real; by fusing the first action video and the second action video to obtain the target human action video corresponding to the human image, it is possible to quickly generate a human action video using a single-frame human image, which is convenient to operate and improves the efficiency of generating the human action video.

[0089] Please refer to Figure 5 , Figure 5 FIG. is a schematic block diagram of a video generation device 1000 provided by an embodiment of the present application. The video generation device is used to execute the foregoing video generation method. Among them, the video generation device can be configured in a server or a terminal.

[0090] As shown in Figure 5As shown in the figure, the video generation device 1000 includes: a human body image acquisition module 1001, a first motion driving module 1002, a second motion driving module 1003, and a video fusion module 1004.

[0091] The human body image acquisition module 1001 is configured to acquire a human body image of the video to be generated.

[0092] The first motion driving module 1002 is configured to drive the motion of the first part in the human body image according to a preset human body posture video, and obtain a first motion video corresponding to the human body image.

[0093] The second motion driving module 1003 is configured to drive the motion of the second part in the human body image according to the human body posture video, and obtain a second motion video corresponding to the human body image.

[0094] The video fusion module 1004 is configured to fuse the first motion video and the second motion video to obtain a target human body motion video corresponding to the human body image.

[0095] It should be noted that those skilled in the art can clearly understand that for the convenience and brevity of description, the specific working processes of the above-described device and each module can refer to the corresponding processes in the foregoing method embodiments, and will not be elaborated herein.

[0096] The above device can be implemented in the form of a computer program, and the computer program can run on a computer device as shown in Figure 6 the figure.

[0097] Please refer to Figure 6 , Figure 6 which is a schematic block diagram of the structure of a computer device provided in an embodiment of the present application.

[0098] Please refer to Figure 6 , the computer device includes a processor and a memory connected through a system bus. Among them, the memory may include a storage medium and an internal memory. The storage medium may be a non-volatile storage medium or a volatile storage medium.

[0099] The processor is configured to provide computing and control capabilities to support the operation of the entire computer device.

[0100] The internal memory provides an environment for the operation of the computer program in the non-volatile storage medium. When the computer program is executed by the processor, the processor can execute any video generation method.

[0101] It should be understood that the processor may be a Central Processing Unit (CPU), and the processor may also be other general-purpose processors, Digital Signal Processors (DSPs), Application Specific Integrated Circuits (ASICs), Field-Programmable Gate Arrays (FPGAs), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. Among them, the general-purpose processor may be a microprocessor or the processor may also be any conventional processor, etc.

[0102] Among them, in one embodiment, the processor is used to run a computer program stored in the memory to implement the following steps:

[0103] Obtain a human body image of the video to be generated; perform action driving on the first part in the human body image according to a preset human body posture video to obtain a first action video corresponding to the human body image; perform action driving on the second part in the human body image according to the human body posture video to obtain a second action video corresponding to the human body image; fuse the first action video and the second action video to obtain a target human body action video corresponding to the human body image.

[0104] In one embodiment, the first part includes the head; when the processor implements performing action driving on the first part in the human body image according to a preset human body posture video to obtain a first action video corresponding to the human body image, it is used to implement:

[0105] Determine the head region image corresponding to the human body image and determine the head action video corresponding to the human body posture video; input the head action video and the head region image into an action driving model for action driving to obtain the first action video.

[0106] In one embodiment, the action driving model includes a key point detector, a motion estimator, and an image generator; when the processor implements inputting the head action video and the head region image into the action driving model for action driving to obtain the first action video, it is used to implement:

[0107] Input the head motion video and the head region image into the key point detector for processing to obtain the first key point information corresponding to the head region image and the second key point information corresponding to the head motion video; determine the affine transformation matrix corresponding to the first key point information and the second key point information; input the head region image, the first key point information, the second key point information, and the affine transformation matrix into the motion estimator for motion estimation processing to obtain the corresponding mapping relationship graph and occlusion graph; input the mapping relationship graph, the occlusion graph, and the head region image into the image generator for image generation to obtain the first action video.

[0108] In one embodiment, when the processor implements determining the affine transformation matrix corresponding to the first key point information and the second key point information, it is used to implement:

[0109] Derive the first key point information to obtain the first derivative corresponding to the first key point information, and derive the second key point information to obtain the second derivative corresponding to the second key point information; generate the affine transformation matrix from the ratio of the first derivative and the second derivative.

[0110] In one embodiment, the image generator includes an encoder and a decoder; when the processor implements inputting the mapping relationship graph, the occlusion graph, and the head region image into the image generator for image generation to obtain the first action video, it is used to implement:

[0111] Perform feature encoding on the head region image through the encoder to obtain an intermediate feature vector; perform an affine transformation on the intermediate feature vector according to the mapping relationship graph to obtain an affine-transformed intermediate feature vector; perform a dot product of the affine-transformed intermediate feature vector and the occlusion graph to obtain a feature vector graph; perform image reconstruction on the feature vector graph through the decoder to obtain the first action video.

[0112] In one embodiment, the second part includes a non-head region, and the second action video includes a non-head action video; when the processor implements driving the action of the second part in the human body image according to the human body posture video to obtain the second action video corresponding to the human body image, it is used to implement:

[0113] Clip the head region in the human body posture video to obtain the human body posture video after clipping the head region; clip the head region in the human body image to obtain the human body image after clipping the head region; input the human body posture video after clipping the head region and the human body image after clipping the head region into an action driving model for action driving to obtain the non-head action video.

[0114] In one embodiment, when the processor implements the fusion of the first action video and the second action video to obtain the target human action video corresponding to the human body image, it is used to implement:

[0115] Perform image alignment on the first action video and the second action video, and perform image stitching on each pair of aligned images to obtain the target human action video.

[0116] An embodiment of the present application also provides a computer-readable storage medium. The computer-readable storage medium stores a computer program, and the computer program includes program instructions. The processor executes the program instructions to implement any video generation method provided by the embodiments of the present application.

[0117] For example, when the program is loaded by the processor, the following steps can be executed:

[0118] Obtain a human body image of the video to be generated; perform action driving on the first part of the human body image according to a preset human body posture video to obtain a first action video corresponding to the human body image; perform action driving on the second part of the human body image according to the human body posture video to obtain a second action video corresponding to the human body image; fuse the first action video and the second action video to obtain a target human action video corresponding to the human body image.

[0119] Among them, the computer-readable storage medium may be an internal storage unit of the computer device described in the foregoing embodiment, such as the hard disk or memory of the computer device. The computer-readable storage medium may also be an external storage device of the computer device, such as a plug-in hard disk equipped on the computer device, a SmartMedia Card (SMC), a Secure Digital Card (SD Card), a Flash Card, etc.

[0120] Further, the computer-readable storage medium may mainly include a storage program area and a storage data area. Among them, the storage program area may store an operating system, application programs required for at least one function, etc.; the storage data area may store data created according to the use of blockchain nodes.

[0121] The blockchain referred to in this application is a new application mode of computer technologies such as distributed data storage, peer-to-peer transmission, consensus mechanism, and encryption algorithms. Blockchain, in essence, is a decentralized database, a series of data blocks generated by using cryptographic methods. Each data block contains information on a batch of network transactions, which is used to verify the validity of the information (anti-counterfeiting) and generate the next block. The blockchain can include the blockchain underlying platform, the platform product service layer, and the application service layer, etc.

[0122] As mentioned above, it is only the specific implementation mode of this application, but the protection scope of this application is not limited thereto. Any person skilled in the art in the technical field disclosed in this application can easily think of various equivalent modifications or substitutions, and these modifications or substitutions should be covered within the protection scope of this application. Therefore, the protection scope of this application should be subject to the protection scope of the claims.

Claims

1. A video generation method, characterized in that, Including: Obtain a human body image of the video to be generated; Perform action driving on the first part in the human body image according to a preset human body posture video to obtain a first action video corresponding to the human body image; Perform action driving on the second part in the human body image according to the human body posture video to obtain a second action video corresponding to the human body image; Fuse the first action video and the second action video to obtain a target human body action video corresponding to the human body image; The first part includes the head; the performing action driving on the first part in the human body image according to a preset human body posture video to obtain a first action video corresponding to the human body image includes: determining a head region image corresponding to the human body image, and determining a head action video corresponding to the human body posture video; inputting the head action video and the head region image into an action driving model for action driving to obtain the first action video; The second part includes a non-head region, and the second action video includes a non-head action video; the performing action driving on the second part in the human body image according to the human body posture video to obtain a second action video corresponding to the human body image includes: shearing the head region in the human body posture video to obtain a human body posture video after head region shearing; shearing the head region in the human body image to obtain a human body image after head region shearing; inputting the human body posture video after head region shearing and the human body image after head region shearing into an action driving model for action driving to obtain the non-head action video; The fusing the first action video and the second action video to obtain a target human body action video corresponding to the human body image includes: performing image alignment on the first action video and the second action video, and performing image stitching on each pair of aligned images to obtain the target human body action video.

2. The video generation method according to claim 1, wherein The action driving model includes a key point detector, a motion estimator, and an image generator; The inputting the head action video and the head region image into the action driving model for action driving to obtain the first action video includes: Inputting the head action video and the head region image into the key point detector for processing to obtain first key point information corresponding to the head region image and second key point information corresponding to the head action video; Determine an affine transformation matrix corresponding to the first key point information and the second key point information; Inputting the head region image, the first key point information, the second key point information, and the affine transformation matrix into the motion estimator for motion estimation processing to obtain a corresponding mapping relationship graph and an occlusion graph; Inputting the mapping relationship graph, the occlusion graph, and the head region image into the image generator for image generation to obtain the first action video.

3. The video generation method according to claim 2, wherein The determining an affine transformation matrix corresponding to the first key point information and the second key point information includes: Derive the first key point information to obtain the first derivative corresponding to the first key point information, and derive the second key point information to obtain the second derivative corresponding to the second key point information; Generate the affine transformation matrix from the ratio of the first derivative to the second derivative.

4. The video generation method according to claim 2, wherein The image generator includes an encoder and a decoder; the step of inputting the mapping relationship graph, the occlusion graph, and the head region image into the image generator to generate an image and obtaining the first action video includes: Feature-encode the head region image through the encoder to obtain an intermediate feature vector; Perform an affine transformation on the intermediate feature vector according to the mapping relationship graph to obtain an affine-transformed intermediate feature vector; Perform a dot product of the affine-transformed intermediate feature vector and the occlusion graph to obtain a feature vector graph; Reconstruct an image from the feature vector graph through the decoder to obtain the first action video.

5. A video generation device, characterized in that, For executing the video generation method according to any one of claims 1 to 4, the video generation device includes: A human body image acquisition module, configured to acquire a human body image of a video to be generated; A first action driving module, configured to drive an action of a first part in the human body image according to a preset human body posture video to obtain a first action video corresponding to the human body image; A second action driving module, configured to drive an action of a second part in the human body image according to the human body posture video to obtain a second action video corresponding to the human body image; A video fusion module, configured to fuse the first action video and the second action video to obtain a target human body action video corresponding to the human body image.

6. A computer device, characterized in that, The computer device includes a memory and a processor; The memory is used to store a computer program; The processor is configured to execute the computer program and, when executing the computer program, implement the video generation method according to any one of claims 1 to 4.

7. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, the processor is caused to implement the video generation method according to any one of claims 1 to 4.

Citation Information

Patent Citations

  • Sign language video generation method, sign language video translation method, sign language video customer service method and device and readable medium

    CN113835522A

  • Character image video generation method and device, computer equipment and storage medium

    CN113920230A