Video generation method and device, electronic equipment and storage medium

By generating a three-dimensional 3D mannequin and combining pose sequence and background replacement technology, the problem of not being able to generate complex action videos and background replacement in the prior art is solved, and personalized video generation with high authenticity and diversity is achieved.

CN120378647APending Publication Date: 2025-07-25BEIJING X RING TECHNOLOGY CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202410675455.7
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2024-05-28
Publication Date
2025-07-25

AI Technical Summary

Technical Problem

The prior art cannot generate video content of any action sequence, especially videos of complex actions, and background replacement cannot be performed.

Method used

By receiving the human body image and human body posture map sent by the terminal device that includes a face but does not include a background, a three-dimensional 3D human body model is generated, and a target video is generated according to the pose sequence of the reference video, including background replacement and shadow rendering.

Benefits of technology

It improves the authenticity and accuracy of the video, increases the creativity and diversity of the video, and realizes personalized customized video content.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120378647A_ABST
    Figure CN120378647A_ABST
Patent Text Reader

Abstract

The invention provides a video generation method and device, electronic equipment and a storage medium, and belongs to the technical field of image processing. The method comprises the following steps: receiving a first human body image, a human body posture graph and a first posture sequence of a reference video sent by terminal equipment, wherein the first human body image comprises a human face but does not comprise a background; acquiring a 3D human body model according to the first human body image and the human body posture graph; generating a second posture sequence of the 3D human body model according to the first posture sequence; and performing video synthesis according to the second attitude sequence to generate a target video, and sending the target video to the terminal device. Therefore, according to the scheme, the target video is generated according to the first human body image, the human body posture graph and the first posture sequence, and the authenticity and accuracy of the video can be improved. By generating the posture sequence from the complex actions, creativity and diversity of the video can be increased, and personalized customization of video content is realized.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present disclosure relates to the field of image processing technologies, and in particular, to a video generation method, apparatus, electronic device, and storage medium. Background Art

[0002] With the development and application of computer technologies and artificial intelligence, the technology for generating videos from images has become increasingly mature. However, existing technologies often perform pose transfer based on video frame information and are unable to generate video content with arbitrary action sequences. Summary of the Invention

[0003] The present disclosure provides a video generation method, apparatus, electronic device, computer-readable storage medium, and computer program product to at least solve the problems of being unable to handle complex actions and unable to perform background replacement. The technical solutions of the present disclosure are as follows:

[0004] According to a first aspect of an embodiment of the present disclosure, there is provided a video generation method, including: receiving a first human body image, a human body pose map, and a first pose sequence of a reference video sent by a terminal device, where the first human body image is a human body image including a face but not including a background; obtaining a three-dimensional (3D) human body model according to the first human body image and the human body pose map; generating a second pose sequence of the 3D human body model according to the first pose sequence; performing video synthesis according to the second pose sequence to generate a target video, and sending the target video to the terminal device.

[0005] According to a first aspect of an embodiment of the present disclosure, there is provided another video generation method, including: obtaining a first human body image, where the first human body image includes a face but not including a background; determining a human body pose map of the human body in the first human body image; obtaining a reference video, and extracting a first pose sequence from the reference video; sending the first human body image, the human body pose map, and the first pose sequence to a server; receiving the target video sent by the server, where the target video is obtained according to the first human body image, the human body pose map, and the first pose sequence.

[0006] According to a second aspect of an embodiment of the present disclosure, there is provided a video generation apparatus, including: a receiving module, configured to receive a first human body image, a human body pose map, and a first pose sequence of a reference video sent by a terminal device, where the first human body image is a human body image including a face but not including a background; an obtaining module, configured to obtain a three-dimensional (3D) human body model according to the first human body image and the human body pose map; a first generating module, configured to generate a second pose sequence of the 3D human body model according to the first pose sequence; a second generating module, configured to perform video synthesis according to the second pose sequence to generate a target video, and send the target video to the terminal device.

[0007] According to a second aspect of the embodiments of the present disclosure, there is provided another video generation device, including: a first acquisition module, configured to acquire a first human body image, where the first human body image includes a human face but does not include a background; a determination module, configured to determine a human body posture map of the human body in the first human body image; a second acquisition module, configured to acquire a reference video and extract a first posture sequence from the reference video; a sending module, configured to send the first human body image, the human body posture map, and the first posture sequence to a server; and a receiving module, configured to receive a target video sent by the server, where the target video is obtained according to the first human body image, the human body posture map, and the first posture sequence.

[0008] According to a third aspect of the embodiments of the present disclosure, there is provided an electronic device, including a processor; a memory for storing processor-executable instructions; wherein, the processor is configured to implement the steps of the method described in the first aspect of the embodiments of the present disclosure.

[0009] According to a fourth aspect of the embodiments of the present disclosure, there is provided a computer-readable storage medium, on which computer program instructions are stored, and when the program instructions are executed by a processor, the steps of the method described in the first aspect of the embodiments of the present disclosure are implemented.

[0010] According to a fifth aspect of the embodiments of the present disclosure, there is provided a computer program product, including a computer program, characterized in that when the computer program is executed by a processor of an electronic device, the steps of the method described in the first aspect of the embodiments of the present disclosure are implemented.

[0011] The technical solutions provided by the embodiments of the present disclosure at least bring the following beneficial effects: By acquiring a first human body image, a human body posture map, and a first posture sequence, and according to the first human body image and the human body posture map, a 3D human body model can be generated. Furthermore, based on the 3D human body model and the first posture sequence, a second posture sequence can be obtained, and according to the second posture sequence, a target video can be generated to improve the authenticity and accuracy of the video. Generating a second posture sequence of the 3D human body model according to the first posture sequence can ensure that the target video maintains consistency with the reference video in terms of posture. The embodiments of the present disclosure can generate a posture sequence for complex actions. Through various unique and complex actions, the creativity and diversity of the video can be increased, and personalized customized video content can be realized.

[0012] It should be understood that the above general description and the following detailed description are only exemplary and explanatory, and cannot limit the present disclosure. BRIEF DESCRIPTION OF THE DRAWINGS

[0013] The accompanying drawings herein are incorporated into the specification and constitute a part of the specification, showing embodiments consistent with the present disclosure, and are used together with the specification to explain the principles of the present disclosure, and do not constitute an improper limitation of the present disclosure.

[0014] Figure 1 is a flowchart of a video generation method shown according to an exemplary embodiment.

[0015] Figure 2 is a flowchart of a video generation method shown according to another exemplary embodiment.

[0016] Figure 3 is a flowchart of a video generation method shown according to another exemplary embodiment.

[0017] Figure 4 is a flowchart of a video generation method shown according to another exemplary embodiment.

[0018] Figure 5 is an interaction diagram of a video generation method shown according to an exemplary embodiment.

[0019] Figure 6 is a flowchart of video generation shown according to an exemplary embodiment.

[0020] Figure 7 is a block diagram of a video generation device shown according to an exemplary embodiment.

[0021] Figure 8 is a block diagram of a video generation device shown according to another exemplary embodiment.

[0022] Figure 9 is a block diagram of an electronic device shown according to an exemplary embodiment. Detailed implementation manners

[0023] To enable those of ordinary skill in the art to better understand the technical solutions of the present disclosure, the technical solutions in the embodiments of the present disclosure will be clearly and completely described below with reference to the accompanying drawings.

[0024] It should be noted that the terms "first", "second", etc. in the specification and claims of the present disclosure and the above-mentioned drawings are used to distinguish similar objects, and do not necessarily have to be used to describe a specific order or sequence. It should be understood that such data can be interchanged under appropriate circumstances so that the embodiments of the present disclosure described herein can be implemented in an order other than those illustrated or described herein. The embodiments described in the following exemplary embodiments do not represent all embodiments consistent with the present disclosure. On the contrary, they are merely examples of devices and methods consistent with some aspects of the present disclosure as detailed in the appended claims.

[0025] In the technical solutions of the present disclosure, the acquisition, storage, use, processing, etc. of data all comply with the provisions of relevant laws and regulations.

[0026] It should be noted that the video generation method provided by the embodiments of the present disclosure can be applied to sub - fields such as image segmentation, pose estimation, pose transfer, and view synthesis.

[0027] Image segmentation: Image segmentation refers to the process of dividing an image into multiple non - overlapping regions or pixel sets. Its goal is to assign each pixel in the image to different categories or objects, so as to achieve semantic understanding and region recognition of the image. Image segmentation has a wide range of applications in the field of computer vision, including object detection, image analysis, image editing, and robot vision. By segmenting the image, the target regions of interest can be extracted, and then higher - level image analysis and understanding can be achieved.

[0028] Pose estimation: Human pose estimation is a fundamental and challenging task in the field of computer vision. Human pose estimation is crucial for describing human poses, human behaviors, etc. Many computer vision tasks are based on the human pose estimation task, including action recognition, action detection, etc. In recent years, with the development of deep learning technology, especially with the proposal of the convolutional neural network algorithm, an implicit human pose estimation model can be established through the powerful fitting ability and feature extraction ability of the neural network, greatly reducing the threshold of human pose estimation and improving the accuracy of human pose estimation at the same time, which also makes human pose estimation develop rapidly.

[0029] Pose transfer: Human pose transfer refers to the process of annotating the positions of human joint points in pictures or videos and making optimal connections for the joint points. Human pose transfer can be widely applied in fields such as action recognition, human - computer interaction, and intelligent tracking, and has become one of the research hotspots in the field of computer vision. In recent years, scholars at home and abroad have proposed a large number of human pose transfer methods and given some publicly available human pose transfer datasets. Therefore, using deep - learning - based methods for human pose transfer has become the main research direction. Deep - learning - based human pose transfer methods can fully extract image information, use deep neural networks to extract more accurate features, and obtain the connections between all human joint points. This method does not require local detectors and manually designed features, does not require designing the topological structure of the model and the interaction between joint points, and uses the outstanding performance of the neural network in positioning and classification to obtain more accurate human joint point positions.

[0030] View Synthesis: The novel view synthesis task refers to rendering a target image corresponding to a target pose given a source image, a source pose, and a target pose. Novel view synthesis has a wide range of applications in fields such as 3D reconstruction, augmented reality (AR) / virtual reality (VR).

[0031] The following describes the video generation method and apparatus according to embodiments of the present disclosure with reference to the accompanying drawings.

[0032] Figure 1 is a flowchart of a video generation method shown according to an exemplary embodiment. As Figure 1 shown, the video generation method according to embodiments of the present disclosure includes the following steps:

[0033] S101: Receive a first human body image, a human body pose map, and a first pose sequence of a reference video sent by a terminal device, where the first human body image is a human body image including a face but not including a background.

[0034] It should be noted that the execution subject of the video generation method according to embodiments of the present disclosure is an electronic device, and the electronic device may be a server. Optionally, the server includes, but is not limited to, a web server, an application server, may also be a server of a distributed system, or a server combined with a blockchain, etc. The video generation method according to embodiments of the present disclosure may be executed by the video generation apparatus according to embodiments of the present disclosure, and the video generation apparatus according to embodiments of the present disclosure may be configured in any electronic device to execute the video generation method according to embodiments of the present disclosure.

[0035] In some implementations, the server can perform video synthesis based on the image information and action information by receiving the image information and action information sent by the terminal device. Optionally, the image information includes a first human body image and a human body pose map, where the first human body image is a human body image including a face but not including a background. The action information includes a first pose sequence of a reference video, where the first pose sequence includes multiple pose frames.

[0036] S102: Obtain a three-dimensional (3D) human body model based on the first human body image and the human body pose map.

[0037] In some implementations, the server can perform 3D digital human modeling based on the first human body image and the human body pose map to obtain a three-dimensional (3D) human body model. Optionally, based on the first human body image and the human body pose map, the key points of the human body and the positions of the key points are determined, and then a 3D human body modeling framework is used to generate a 3D human body model. The key points can be the head, neck, hands, feet, etc.

[0038] S103. Generate a second pose sequence of the 3D human body model according to the first pose sequence.

[0039] In some implementations, the first pose sequence includes multiple pose frames of the human body, and the pose frames are arranged in chronological order to form the first pose sequence. The key points of the human body and the action information of the key points can be obtained from the pose frames, and the key points of the 3D human body model are obtained. The key points of the human body are matched with the key points of the 3D human body model, and the action information of the matched key points is matched to the 3D human body model to obtain the second pose sequence. Among them, the order of the pose frames in the second pose sequence is the same as that of the pose frames in the first pose sequence.

[0040] Optionally, the 3D human body model can be controlled to move according to the action information of the key points to obtain the second pose sequence.

[0041] S104. Perform video synthesis according to the second pose sequence to generate a target video, and send the target video to the terminal device.

[0042] In some implementations, the second pose sequence can be used to drive the 3D human body model to obtain a target video, and the target video is sent to the terminal device. Optionally, according to the pose frames in the second pose sequence, the pose of the 3D human body model is updated frame by frame to ensure that the pose of the 3D human body model matches the pose in the second pose sequence, so that the animation effect in the target video is smoother.

[0043] The video generation method provided by the embodiments of the present disclosure obtains the first human body image, the human body pose map, and the first pose sequence, and based on the first human body image and the human body pose map, a 3D human body model can be generated. Then, based on the 3D human body model and the first pose sequence, a second pose sequence is obtained, and a target video can be generated according to the second pose sequence to improve the authenticity and accuracy of the video. Generating a second pose sequence of the 3D human body model according to the first pose sequence can ensure that the target video maintains consistency with the reference video in terms of pose. The embodiments of the present disclosure can generate a pose sequence from complex actions. Through various unique and complex actions, the creativity and diversity of the video can be increased, and personalized customized video content can be realized.

[0044] Figure 2is a flowchart of a video generation method shown according to an exemplary embodiment. As Figure 2 shown, the video generation method of the embodiments of the present disclosure includes the following steps:

[0045] S201, receive a first human body image, a human body pose map, and a first pose sequence of a reference video sent by a terminal device, where the first human body image is a human body image that includes a human face but does not include a background.

[0046] For the relevant content of step S201, reference may be made to the above embodiments, which will not be elaborated here.

[0047] S202, perform 3D digital human modeling based on the first human body image and the human body pose map to generate a 3D human skinned multi-person linear SMPL model.

[0048] In some implementations, the body shape information of the human body can be obtained from the first human body image, and the key points of the human body, such as the head and feet of the human body, can be obtained from the human body pose map. Then, based on the body shape information and key points of the human body, the skinned multi-person linear (SMPL) model of the human body is fitted to obtain a 3D human SMPL model.

[0049] S203, diffuse the first human body image to generate a human back texture image.

[0050] S204, extract human texture information from the first human body image and the human back texture image, and based on the human texture information, perform texture rendering on the 3D human SMPL model to obtain a 3D human model.

[0051] In some implementations, to ensure the consistency and coherence of the 3D human model, the human back texture image can be determined, and based on the human back texture image and the human body image, texture rendering is performed on the 3D human SMPL model to obtain a 3D human model.

[0052] Optionally, based on the first human body image, a diffusion algorithm can be used to generate the first human body image into a human back texture image.

[0053] Further, feature extraction can be performed on the first human body image and the human back texture image to obtain human texture information, and then the human texture information is rendered onto the 3D human SMPL model to obtain a 3D human model.

[0054] Optionally, the key points of the 3D human SMPL model can be matched with the key points of the human texture information, and the matched human texture information is rendered onto the 3D human SMPL model to obtain a 3D human model.

[0055] S205. Generate a second pose sequence of the 3D human model according to the first pose sequence.

[0056] S206. Perform video synthesis according to the second pose sequence to generate a target video, and send the target video to the terminal device.

[0057] For the relevant content of steps S205 - S206, reference can be made to the above embodiments and will not be elaborated here.

[0058] The video generation method provided by the embodiments of the present disclosure obtains a first human body image, a human body pose map, and a first pose sequence, generates a 3D human body SMPL model according to the first human body image and the human body pose map, renders to obtain a 3D human model based on the human body texture information, further obtains a second pose sequence based on the 3D human model and the first pose sequence, and can generate a target video according to the second pose sequence to improve the authenticity and accuracy of the video. Generating a second pose sequence of the 3D human model according to the first pose sequence can ensure the consistency of the target video with the reference video in terms of pose. The embodiments of the present disclosure can generate a pose sequence for complex actions. Through various unique and complex actions, the creativity and diversity of the video can be increased, and personalized customized video content can be realized.

[0059] Figure 3 is a flowchart of a video generation method shown according to an exemplary embodiment. As Figure 3 shown, the video generation method of the embodiments of the present disclosure includes the following steps:

[0060] S301. Receive the first human body image, the human body pose map, and the first pose sequence of the reference video sent by the terminal device, where the first human body image is a human body image including the face but not the background.

[0061] S302. Obtain a three-dimensional 3D human model according to the first human body image and the human body pose map.

[0062] For the relevant content of steps S301 - S302, reference can be made to the above embodiments and will not be elaborated here.

[0063] S303. Sequentially extract the key points and the actions of the key points in the first pose frame.

[0064] In some implementations, the first pose sequence includes multiple first pose frames, and the multiple first pose frames are arranged in order to obtain the first pose sequence. The key points and the actions of the key points of each frame can be extracted from the first pose frames according to the order of the first pose frames.

[0065] S304. Drive the 3D human model to perform actions on the matching key points based on the key points and the actions of the key points to obtain the second pose frame of the 3D human model.

[0066] In some implementations, for any first pose frame, by obtaining the key points of the 3D human model and matching the key points of the 3D human model with the key points in the first pose frame, the matching key points on the 3D human model are determined, and then the 3D human model is driven to act according to the actions of the matching key points, and the second pose frame of the 3D human model corresponding to the first pose frame is obtained.

[0067] S305. Combine the second pose frames according to the matching relationship between the first pose frames and the second pose frames and the order of the first pose frames to generate a second pose sequence.

[0068] Optionally, by determining the second pose frame matched by each first pose frame and combining the second pose frames in the order of the first pose frames, a second pose sequence can be obtained. For example, the first pose frames in the first pose sequence are: Frame A, Frame B, Frame C, and the second pose frames matched by the first pose frames are: Frame 2, Frame 3, Frame 1 respectively, and the order of the first pose frames is Frame C, Frame A, Frame B, then the second pose sequence is: Frame 1, Frame 2, Frame 3.

[0069] S306. Receive the background image sent by the terminal device or select a background image locally.

[0070] In some implementations, in order to obtain the scene information of the target video, the background image sent by the terminal device can be received, where the background image sent by the terminal device can be a background image taken by the user himself. An image can also be selected locally as the background image.

[0071] S307. Synthesize a video according to the background image and the second pose sequence to generate a target video.

[0072] In some implementations, in order to make the effect of the target video more realistic, the light source information can be obtained from the background image, and the shadow information of the 3D human model can be determined according to the light source information, and then the shadow information, the second pose sequence and the background image are fused to obtain the target video.

[0073] Optionally, the background image can be analyzed based on a neural network to extract the light source point information from the background image, and then according to the light source point information, the 3D human model in the second pose sequence is shadow-rendered and background-combined with the background image to generate a target video.

[0074] In some implementations, the shadow area can be determined based on the light source information and the position information of the human body in the second pose sequence, and then combined with the shadow area, the 3D human model in the second pose sequence is shadow-rendered to obtain the target video.

[0075] Optionally, the light source position can be extracted from the light source point information, and the human body position of the 3D human body model in each second pose frame of the second pose sequence can be extracted. Further, for each second pose frame, according to the light source position and the human body position in the second pose frame, the relative position relationship between the light source and the 3D human body model can be determined. Furthermore, according to the relative position relationship and the height information of the 3D human body model, the shadow area in the second pose frame can be determined, and shadow rendering can be performed in the shadow area.

[0076] S308. Send the target video to the terminal device.

[0077] For the relevant content of step S308, reference can be made to the above embodiments and will not be elaborated here.

[0078] The video generation method provided by the embodiments of the present disclosure obtains a first human body image, a human body pose map, and a first pose sequence. According to the first human body image and the human body pose map, a 3D human body model can be generated. Furthermore, according to the key points and actions of the key points in the first pose frame of the first pose sequence, combined with the 3D human body model, a second pose sequence can be obtained. And according to the second pose sequence and the background image, a target video can be generated to achieve random replacement of the background. Determine the shadow area according to the background image and perform shadow rendering on the 3D human body model to improve the authenticity of the target video. Generating a second pose sequence of the 3D human body model according to the first pose sequence can ensure that the target video maintains consistency with the reference video in terms of pose. The embodiments of the present disclosure can generate a pose sequence for complex actions. Through various unique and complex actions, the creativity and diversity of the video can be increased, and personalized customized video content can be realized.

[0079] Figure 4 is a flowchart of a video generation method shown according to an exemplary embodiment. As Figure 4 shown, the video generation method of the embodiments of the present disclosure includes the following steps:

[0080] S401. Obtain a first human body image, where the first human body image includes a human face but does not include a background.

[0081] It should be noted that the execution subject of the video generation method of the embodiments of the present disclosure is an electronic device, and the electronic device can be a terminal device. The video generation method of the embodiments of the present disclosure can be executed by the video generation device of the embodiments of the present disclosure. The video generation device of the embodiments of the present disclosure can be configured in any electronic device to execute the video generation method of the embodiments of the present disclosure.

[0082] In some implementations, a second human body image including a human face can be collected, and background recognition can be performed on the second human body image to remove the background from the second human body image, obtaining a first human body image. Optionally, the terminal device can call a portrait segmentation algorithm to remove the background of the second human body image, obtaining a first human body image.

[0083] S402. Determine the human body pose map of the human body in the first human body image.

[0084] In some implementations, the terminal device can call a pose estimation algorithm to perform pose estimation on the first human body image to determine the human body pose map of the human body in the first human body image. Optionally, the terminal device can also perform pose estimation on the second human body image to determine the human body pose map. That is to say, the human body pose map can be extracted from the first human body image. Or, the human body pose map can be extracted from the second human body image.

[0085] S403. Obtain a reference video and extract a first pose sequence from the reference video.

[0086] In some implementations, the user can shoot an action video based on the terminal device and use this video as the reference video. Further, the terminal device calls a pose estimation algorithm to extract the pose of each frame from the reference video to obtain a first pose frame, and then arranges the first pose frames in chronological order to form a first pose sequence.

[0087] S404. Send the first human body image, the human body pose map, and the first pose sequence to the server.

[0088] In some implementations, after obtaining the first human body image, the human body pose map, and the first pose sequence, the terminal device sends them to the server, and the server generates a target video based on the first human body image, the human body pose map, and the first pose sequence.

[0089] Optionally, the terminal device can also collect a background image and send the background image to the server. The background image is used to generate the background and shadows in the target video to enrich the content of the target video and make the target video more realistic.

[0090] S405. Receive the target video sent by the server, where the target video is obtained based on the first human body image, the human body pose map, and the first pose sequence.

[0091] In some implementations, the terminal device receives the target video sent by the server and displays the target video to the user, where the target video is obtained by the server based on the first human body image, the human body pose map, and the first pose sequence.

[0092] The video generation method provided by an embodiment of the present disclosure can obtain a high-fidelity and accurate video by acquiring a first human body image and a human body pose map from a second human body image, obtaining a first pose sequence according to a reference video, sending the first human body image, the human body pose map, and the first pose sequence to a server, and receiving a target video from the server.

[0093] Figure 5 is an interaction diagram of a video generation method shown according to an exemplary embodiment. As Figure 5 shown, the video generation method of the embodiment of the present disclosure includes the following steps:

[0094] S501. The terminal device acquires a first human body image and determines the human body pose map of the human body in the first human body image.

[0095] S502. The terminal device acquires a reference video and extracts a first pose sequence from the reference video.

[0096] S503. The terminal device sends the first human body image, the human body pose map, and the first pose sequence to the server.

[0097] S504. The server acquires a 3D human body model according to the first human body image and the human body pose map.

[0098] S505. The server generates a second pose sequence of the 3D human body model according to the first pose sequence.

[0099] S506. The server performs video synthesis according to the second pose sequence to generate a target video.

[0100] S507. The server sends the target video to the terminal device.

[0101] The video generation method provided by an embodiment of the present disclosure can generate a 3D human body model by acquiring a first human body image, a human body pose map, and a first pose sequence, and according to the first human body image and the human body pose map. Furthermore, based on the 3D human body model and the first pose sequence, a second pose sequence can be obtained, and a target video can be generated according to the second pose sequence to improve the fidelity and accuracy of the video. Generating the second pose sequence of the 3D human body model according to the first pose sequence can ensure that the target video maintains consistency with the reference video in terms of pose. The embodiment of the present disclosure can generate a pose sequence for complex actions. Through various unique and complex actions, the creativity and diversity of the video can be increased, realizing personalized customization of video content.

[0102] As Figure 6 shown is a schematic flowchart of video generation. Figure 6 It includes the terminal device side and the server side. Among them, the terminal device side is the mobile phone side.

[0103] Mobile side:

[0104] 1. The user uses the mobile phone to take a full-body photo as the second human body image I human , and the mobile side invokes the human portrait segmentation algorithm to remove the background in I human , and obtains the first human body image, denoted as

[0105] 2. The mobile side invokes the pose estimation algorithm to process the second human body image I human , and obtains the human body pose map

[0106] 3. The user uses the mobile phone to record an action video as the reference video, and the mobile side invokes the pose estimation algorithm to extract the pose information of the human body in each frame of the reference video, and obtains the first pose sequence V pose ;

[0107] 4. The user takes a photo as the background image I background ;

[0108] 5. Send the first human body image the human body pose map the first pose sequence v pose and the background image I background to the server side.

[0109] Server side:

[0110] 1. Perform 3D digital human modeling according to the first human body image and the human body pose map , and obtain the 3D human body SMPL model SMPL human ;

[0111] 2. Diffuse the first human body image to generate the human back texture image

[0112] 3. Extract the human body texture information from the first human body image and the human back texture image , and based on the human body texture information, perform texture rendering on the 3D human body SMPL model SMPL human , and obtain the 3D human body model

[0113] 4. Use the first pose sequence V pose to drive the 3D human body model to obtain the second pose sequence

[0114] 5. Use the neural network to analyze the background image I background , and obtain the light source point information Pilluminant ;

[0115] 6. Use the light source point information P illuminant and the second pose sequence to calculate the shadow area;

[0116] 7. Fuse the second pose sequence the shadow area and the background image I background to obtain the target video and send the target video to the mobile phone side.

[0117] Figure 7 is a block diagram of a video generation device shown according to an exemplary embodiment. Refer to Figure 7 , the video generation device 700 of the embodiments of the present disclosure includes: a receiving module 701, an obtaining module 702, a first generating module 703, and a second generating module 704.

[0118] The receiving module 701 is configured to receive the first human body image, the human body pose map, and the first pose sequence of the reference video sent by the terminal device, where the first human body image is a human body image including a face but not including a background;

[0119] The obtaining module 702 is configured to obtain a three-dimensional 3D human body model according to the first human body image and the human body pose map;

[0120] The first generating module 703 is configured to generate a second pose sequence of the 3D human body model according to the first pose sequence;

[0121] The second generating module 704 is configured to perform video synthesis according to the second pose sequence to generate a target video and send the target video to the terminal device.

[0122] In an embodiment of the present disclosure, the obtaining module 702 is further configured to: perform 3D digital human modeling according to the first human body image and the human body pose map to generate a 3D human skin multi-person linear SMPL model; diffuse the first human body image to generate a human back texture image; extract human texture information from the first human body image and the human back texture image, and perform texture rendering on the 3D human body SMPL model based on the human texture information to obtain the 3D human body model.

[0123] In one embodiment of the present disclosure, the first generation module 703 is further configured to: sequentially extract key points and actions of the key points in the first pose frame; drive the matching key points on the 3D human body model to perform actions based on the key points and the actions of the key points, so as to obtain a second pose frame of the 3D human body model; and combine the second pose frames according to the matching relationship between the first pose frame and the second pose frame and the order of the first pose frame to generate the second pose sequence.

[0124] In one embodiment of the present disclosure, the second generation module 704 is further configured to: receive a background image sent by the terminal device or select a background image locally; and perform video synthesis according to the background image and the second pose sequence to generate the target video.

[0125] In one embodiment of the present disclosure, the second generation module 704 is further configured to: extract light source point information from the background image; perform shadow rendering on the 3D human body model in the second pose sequence according to the light source point information, and perform background synthesis with the background image to generate the target video.

[0126] In one embodiment of the present disclosure, the second generation module 704 is further configured to: extract the light source position from the light source point information; extract the human body position of the 3D human body model in each second pose frame of the second pose sequence; for each second pose frame, determine the relative position relationship between the light source and the 3D human body model according to the light source position and the human body position in the second pose frame; and determine the shadow area in the second pose frame according to the relative position relationship and the height information of the 3D human body model, and perform shadow rendering in the shadow area.

[0127] Regarding the device in the above embodiments, the specific manners in which each module performs operations have been described in detail in the embodiments related to the method, and will not be elaborated here.

[0128] The video generation device provided by the embodiments of the present disclosure can generate a 3D human body model by obtaining a first human body image, a human body pose map, and a first pose sequence, and then obtain a second pose sequence based on the 3D human body model and the first pose sequence, and generate a target video according to the second pose sequence, so as to improve the authenticity and accuracy of the video. Generating a second pose sequence of the 3D human body model according to the first pose sequence can ensure the consistency of the target video with the reference video in terms of pose. The embodiments of the present disclosure can generate a pose sequence for complex actions, and through various unique and complex actions, can increase the creativity and diversity of the video and achieve personalized customization of video content.

[0129] Figure 8 is a block diagram of a video generation device shown according to an exemplary embodiment. Referring to Figure 8 , the video generation device 800 according to the embodiments of the present disclosure includes: a first acquisition module 801, a determination module 802, a second acquisition module 803, a sending module 804, and a receiving module 805.

[0130] The first acquisition module 801 is configured to acquire a first human body image, where the first human body image includes a human face but does not include a background;

[0131] The determination module 802 is configured to determine a human body posture map of the human body in the first human body image;

[0132] The second acquisition module 803 is configured to acquire a reference video and extract a first posture sequence from the reference video;

[0133] The sending module 804 is configured to send the first human body image, the human body posture map, and the first posture sequence to the server;

[0134] The receiving module 805 is configured to receive a target video sent by the server, where the target video is obtained according to the first human body image, the human body posture map, and the first posture sequence.

[0135] In an embodiment of the present disclosure, the first acquisition module 801 is further configured to: collect a second human body image including the human face, perform background recognition on the second human body image, and remove the background from the second human body image to obtain the first human body image.

[0136] In an embodiment of the present disclosure, the determination module 802 is further configured to: extract the human body posture map from the first human body image; or extract the human body posture map from the second human body image.

[0137] In an embodiment of the present disclosure, the device further includes: collecting a background image and sending the background image to the server, where the background image is used to generate the background and shadows in the target video.

[0138] Regarding the device in the above embodiments, the specific manners in which each module performs operations have been described in detail in the embodiments related to the method, and will not be elaborated herein.

[0139] The video generation device provided by the embodiments of the present disclosure can obtain a video with high authenticity and accuracy by acquiring a first human body image and a human body posture map from a second human body image, obtaining a first posture sequence according to a reference video, sending the first human body image, the human body posture map, and the first posture sequence to the server, and receiving the target video from the server.

[0140] Figure 9 is a block diagram of an electronic device shown according to an exemplary embodiment.

[0141] As Figure 9 shown, the above-mentioned electronic device 900 includes:

[0142] a memory 901 and a processor 902, a bus 903 connecting different components (including the memory 901 and the processor 902), the memory 901 stores a computer program, and when the processor 902 executes the program, the video generation method described in the embodiments of the present disclosure is implemented.

[0143] The bus 903 represents one or more of several types of bus structures, including a memory bus or a memory controller, a peripheral bus, a graphics acceleration port, a processor, or a local bus using any bus structure in a variety of bus structures. For example, these architectures include, but are not limited to, Industry Standard Architecture (ISA) bus, Micro Channel Architecture (MAC) bus, Enhanced ISA bus, Video Electronics Standards Association (VESA) local bus, and Peripheral Component Interconnect (PCI) bus.

[0144] The electronic device 900 typically includes a variety of electronic device-readable media. These media can be any available media that can be accessed by the electronic device 900, including volatile and non-volatile media, removable and non-removable media.

[0145] The memory 901 may further include a computer system-readable medium in the form of volatile memory, such as random access memory (RAM) 904 and / or cache memory 905. The electronic device 900 may further include other removable / non-removable, volatile / non-volatile computer system storage media. By way of example only, a storage system 906 can be used for reading and writing non-removable, non-volatile magnetic media ( Figure 9 not shown, commonly referred to as a "hard disk drive"). Although Figure 9 not shown in the figure, a disk drive for reading and writing removable non-volatile disks (such as "floppy disks") and an optical disk drive for reading and writing removable non-volatile optical disks (such as CD-ROM, DVD-ROM, or other optical media) can be provided. In these cases, each drive can be connected to the bus 903 through one or more data media interfaces. The memory 901 may include at least one program product having a set (e.g., at least one) of program modules configured to perform the functions of the embodiments of the present disclosure.

[0146] A program / utilities 908 having a set (at least one) of program modules 907 can be stored in, for example, a memory 901. Such program modules 907 include, but are not limited to, an operating system, one or more application programs, other program modules, and program data. Each or some combination of these examples may include an implementation of a network environment. The program modules 907 generally execute the functions and / or methods in the embodiments described in the present disclosure.

[0147] The electronic device 900 can also communicate with one or more external devices 909 (such as a keyboard, a pointing device, a display 991, etc.), and can also communicate with one or more devices that enable a user to interact with the electronic device 900, and / or communicate with any device that enables the electronic device 900 to communicate with one or more other computing devices (such as a network card, a modem, etc.). Such communication can be carried out through an input / output (I / O) interface 992. And, the electronic device 900 can also communicate with one or more networks (such as a local area network (LAN), a wide area network (WAN), and / or a public network, such as the Internet) through a network adapter 993. As Figure 9 shown, the network adapter 993 communicates with other modules of the electronic device 900 through a bus 903. It should be understood that although not shown in the figure, other hardware and / or software modules can be used in conjunction with the electronic device 900, including but not limited to: microcode, device drivers, redundant processing units, external disk drive arrays, RAID systems, tape drives, and data backup storage systems, etc.

[0148] The processor 902 executes various functional applications and data processing by running programs stored in the memory 901.

[0149] It should be noted that for the implementation process and technical principle of the electronic device in this embodiment, refer to the foregoing explanation of the video generation method in the embodiments of the present disclosure, and details are not described herein again.

[0150] To implement the above embodiments, the present disclosure also proposes a computer-readable storage medium having computer program instructions stored thereon. When the program instructions are executed by a processor, the steps of the video generation method provided by the present disclosure are implemented.

[0151] Optionally, the computer-readable storage medium can be a ROM, a random access memory (RAM), a CD-ROM, a magnetic tape, a floppy disk, and an optical data storage device, etc.

[0152] To implement the above embodiments, the present disclosure also provides a computer program product, including a computer program, characterized in that when the computer program is executed by a processor of an electronic device, the video generation method as described above is implemented.

[0153] Other embodiments of the present disclosure will be readily apparent to those skilled in the art upon consideration of the specification and practice of the invention disclosed herein. The present disclosure is intended to cover any variations, uses, or adaptations of the present disclosure that follow the general principles of the present disclosure and include known common general knowledge or conventional technical means in the technical field not disclosed by the present disclosure. The specification and examples are only to be considered as exemplary, and the true scope and spirit of the present disclosure are pointed out by the following claims.

[0154] It should be understood that the present disclosure is not limited to the exact structures described above and shown in the drawings, and various modifications and changes can be made without departing from its scope. The scope of the present disclosure is only limited by the appended claims.

Claims

1. A video generation method, characterized in that, The method includes: Receiving a first human body image, a human body posture map, and a first posture sequence of a reference video sent by a terminal device, where the first human body image is a human body image including a human face but not including a background; Obtaining a three-dimensional (3D) human body model according to the first human body image and the human body posture map; Generating a second posture sequence of the 3D human body model according to the first posture sequence; Performing video synthesis according to the second posture sequence to generate a target video, and sending the target video to the terminal device.

2. The method according to claim 1, wherein The obtaining a three-dimensional (3D) human body model according to the first human body image and the human body posture map includes: Performing 3D digital human modeling according to the first human body image and the human body posture map to generate a 3D human skin multi-person linear SMPL model; Diffusing the first human body image to generate a human back texture image; Extracting human texture information from the first human body image and the human back texture image, and performing texture rendering on the 3D human body SMPL model based on the human texture information to obtain the 3D human body model.

3. The method according to claim 1, characterized in that, The first posture sequence includes a plurality of first posture frames arranged in sequence. The generating a second posture sequence of the 3D human body model according to the first posture sequence includes: Sequentially extracting key points and actions of the key points in the first posture frames; Driving the matching key points on the 3D human body model to perform actions based on the key points and the actions of the key points to obtain second posture frames of the 3D human body model; Combining the second posture frames according to the matching relationship between the first posture frames and the second posture frames and the order of the first posture frames to generate the second posture sequence.

4. The method according to any one of claims 1-3, characterized in that, The performing video synthesis according to the second posture sequence to generate a target video includes: Receiving a background image sent by the terminal device or selecting a background image locally; Performing video synthesis according to the background image and the second posture sequence to generate the target video.

5. The method according to claim 4, wherein The performing video synthesis according to the background image and the second posture sequence to generate the target video includes: Extracting light source point information from the background image; Performing shadow rendering on the 3D human body model in the second posture sequence according to the light source point information, and performing background synthesis with the background image to generate the target video.

6. The method according to claim 5, characterized in that, The performing shadow rendering on the 3D human body model in the second posture sequence according to the light source point information includes: Extracting the light source position from the light source point information; Extracting the human body position of the 3D human body model in each second posture frame of the second posture sequence; For each second posture frame, determining the relative position relationship between the light source and the 3D human body model according to the light source position and the human body position in the second posture frame; Determining the shadow area in the second posture frame according to the relative position relationship and the height information of the 3D human body model, and performing shadow rendering in the shadow area.

7. A video generation method, characterized in that, The method includes: Obtaining a first human body image, where the first human body image includes a human face but not including a background; Determine the human body posture map of the human body in the first human body image; Obtain a reference video, and extract a first posture sequence from the reference video; Send the first human body image, the human body posture map, and the first posture sequence to the server; Receive the target video sent by the server, where the target video is obtained according to the first human body image, the human body posture map, and the first posture sequence.

8. The method according to claim 7, wherein The obtaining of the first human body image includes: Collect a second human body image including the face, perform background recognition on the second human body image, and remove the background from the second human body image to obtain the first human body image.

9. The method according to claim 8, wherein The determining of the human body posture map of the human body in the first human body image includes: Extract the human body posture map from the first human body image; or, Extract the human body posture map from the second human body image.

10. The method according to any one of claims 7-9, characterized in that The method further includes: Collect a background image and send the background image to the server, where the background image is used to generate the background and shadows in the target video.

11. A video generation device, characterized in that, The device includes: A receiving module, configured to receive the first human body image, the human body posture map, and the first posture sequence of the reference video sent by the terminal device, where the first human body image is a human body image including the face but not including the background; An obtaining module, configured to obtain a three-dimensional (3D) human body model according to the first human body image and the human body posture map; A first generating module, configured to generate a second posture sequence of the 3D human body model according to the first posture sequence; A second generating module, configured to perform video synthesis according to the second posture sequence to generate a target video, and send the target video to the terminal device.

12. A video generation device, characterized in that, The device includes: A first obtaining module, configured to obtain a first human body image, where the first human body image includes the face but not including the background; A determining module, configured to determine the human body posture map of the human body in the first human body image; A second obtaining module, configured to obtain a reference video and extract a first posture sequence from the reference video; A sending module, configured to send the first human body image, the human body posture map, and the first posture sequence to the server; A receiving module, configured to receive the target video sent by the server, where the target video is obtained according to the first human body image, the human body posture map, and the first posture sequence.

13. An electronic device, characterized in that, Includes: A processor; A memory for storing processor-executable instructions; Wherein, the processor is configured to: Implement the steps of the method according to any one of claims 1-10.

14. A computer-readable storage medium having computer program instructions stored thereon, characterized in that, When the program instructions are executed by the processor, the steps of the method according to any one of claims 1-10 are implemented.