Video generation method, electronic equipment and computer readable storage medium
By identifying portraits in the initial video data and processing trajectory information, generating action data and applying them to virtual characters, the problem of high motion capture cost is solved and efficient video generation is achieved.
Patent Information
- Application Number
- CN202311649093.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2023-12-04
- Publication Date
- 2025-06-13
AI Technical Summary
In the prior art, motion capture is required by wearing a motion capture device for human body, resulting in high cost of motion capture.
By identifying the portraits in the initial video data, the trajectory information corresponding to each marking point in the portrait is obtained, and action data is generated based on the trajectory information, applied to virtual characters, and replica video data is generated without wearing a motion capture device.
It reduces the cost of motion capture, improves the efficiency of generating video data, and realizes video generation without real-person capture of actions.
Smart Images

Figure CN120147583A_ABST
Abstract
Description
Technical Field
[0001] This application belongs to the technical field of terminal devices, and particularly relates to a video generation method, an electronic device, and a computer-readable storage medium. Background Art
[0002] With the continuous development of virtual reality technology, by using motion capture technology to identify and capture the actions and expressions of the human body, and displaying the corresponding actions and expressions of the human body through virtual characters, the similarity between virtual characters and real characters can be improved.
[0003] In the related art, a motion capture device can capture human actions and transmit the captured data to a terminal device. Correspondingly, the terminal device can drive a pre-set virtual character according to the captured data, so that the actions of the virtual character are similar to the human actions.
[0004] However, in practical applications, it is necessary to wear a motion capture device on the human body to complete motion capture, which will cause the problem of high cost of motion capture. Summary of the Invention
[0005] This application provides a video generation method, an electronic device, and a computer-readable storage medium, which solve the problem in the prior art that it is necessary to wear a motion capture device on the human body to complete motion capture, resulting in a high cost of motion capture.
[0006] To achieve the above object, this application adopts the following technical solutions:
[0007] In a first aspect, an embodiment of this application provides a video generation method, and the method includes:
[0008] Identify the human figure in the initial video data to obtain the trajectory information corresponding to each marked point in the human figure;
[0009] Generate action data according to the trajectory information, where the action data is used to indicate the actions performed by the virtual character;
[0010] Apply the action data to the virtual character to generate replicated video data.
[0011] Optionally, the identifying the human figure in the initial video data to obtain the trajectory information corresponding to each marked point in the human figure includes:
[0012] Identify the human figure in the initial video data to determine each joint point of the human figure;
[0013] Establish a marked point corresponding to each joint point;
[0014] For each of the marked points, record the position change information and angle change information of the marked point, and combine them to obtain the trajectory information of the marked point.
[0015] Optionally, before generating the action data according to the trajectory information, the method further includes:
[0016] Determine the ground information where the portrait is located and the shooting information corresponding to the portrait according to the portrait in the initial video data, where the shooting information includes: shooting coordinates and shooting angles;
[0017] The generating action data according to the trajectory information includes:
[0018] Control the movement of the pre-set virtual skeleton according to the trajectory information to obtain the movement data of the virtual skeleton;
[0019] Generate the action data through the movement trajectory of the virtual skeleton according to the ground information, shooting information and pre-set algorithm.
[0020] Optionally, the applying the action data to the virtual character to generate the replicated video data includes:
[0021] Redirect at least one virtual character according to the action data;
[0022] For each of the virtual characters, adjust the body proportion of the virtual character according to the ratio between the bones at all levels of the virtual character;
[0023] Generate the replicated video data according to the adjusted virtual character.
[0024] Optionally, before generating the replicated video data according to the adjusted virtual character, the method further includes:
[0025] Optimize the action data according to a pre-set filtering algorithm.
[0026] Optionally, the applying the action data to the virtual character to generate the replicated video data includes:
[0027] Adjust the angle of the virtual character according to the shooting information;
[0028] Combine the adjusted virtual character with the background image to generate the replicated video data.
[0029] Optionally, before combining the adjusted virtual character with the background image to generate the replicated video data, the method further includes:
[0030] Determine the light information for irradiating the human figure according to the initial video data;
[0031] Render the virtual character according to the light information.
[0032] Optionally, the combining the adjusted virtual character with the background image to generate the replicated video data includes:
[0033] Perform de - human - figure processing on the initial video data, and fill the position occupied by the human figure to obtain the background image;
[0034] Combine the adjusted virtual character with the background image to generate the replicated video data.
[0035] In a second aspect, an embodiment of the present application provides a video generation device, and the device includes:
[0036] An identification module, configured to identify the human figure in the initial video data to obtain the trajectory information corresponding to each marker point in the human figure;
[0037] A first generation module, configured to generate action data according to the trajectory information, where the action data is used to indicate the actions performed by the virtual character;
[0038] A second generation module, configured to apply the action data to the virtual character to generate replicated video data.
[0039] Optionally, the identification module is specifically configured to identify the human figure in the initial video data, determine each joint point of the human figure; establish a marker point corresponding to each joint point; for each marker point, record the position change information and the angle change information of the marker point, and combine them to obtain the trajectory information of the marker point.
[0040] Optionally, the device further includes:
[0041] A first determination module, configured to determine the ground information where the human figure is located and the shooting information corresponding to the human figure according to the human figure in the initial video data, where the shooting information includes: shooting coordinates and shooting angles;
[0042] The first generation module is specifically configured to control the movement of a pre - set virtual skeleton according to the trajectory information to obtain the movement data of the virtual skeleton; and generate the action data through the movement trajectory of the virtual skeleton according to the ground information, the shooting information, and a pre - set algorithm.
[0043] Optionally, the second generation module is specifically configured to redirect at least one virtual character according to the action data; for each virtual character, adjust the body proportion of the virtual character according to the proportion between the bones at all levels of the virtual character; and generate the replicated video data according to the adjusted virtual character.
[0044] Optionally, the apparatus further includes:
[0045] An optimization module, configured to optimize the action data according to a preset filtering algorithm.
[0046] Optionally, the second generation module is specifically configured to adjust the angle of the virtual character according to the shooting information; combine the adjusted virtual character with a background image to generate the replicated video data.
[0047] Optionally, the apparatus further includes:
[0048] A second determination module, configured to determine light information for irradiating the portrait according to the initial video data;
[0049] A rendering module, configured to render the virtual character according to the light information.
[0050] Optionally, the second generation module is specifically configured to perform portrait removal processing on the initial video data, fill the position occupied by the portrait, and obtain the background image; combine the adjusted virtual character with the background image to generate the replicated video data.
[0051] In a third aspect, an embodiment of the present application provides an electronic device, including: a memory and a processor, where the memory is used to store a computer program; the processor is configured to execute the method described in the first aspect or any implementation manner of the first aspect when calling the computer program.
[0052] In a fourth aspect, an embodiment of the present application provides a computer-readable storage medium, on which a computer program is stored, and when the computer program is executed by a processor, the method described in the first aspect or any implementation manner of the first aspect is implemented.
[0053] A video generation method provided by an embodiment of the present application identifies a portrait in initial video data to obtain trajectory information corresponding to each marker point in the portrait, then generates action data according to the trajectory information, where the action data is used to indicate the actions performed by a virtual character, and finally applies the action data to the virtual character to generate replicated video data, without the need to capture actions by a real person, which can reduce the cost of action capture and improve the efficiency of generating video data. Description of the Drawings
[0054] Figure 1A Schematic diagram of the video generation scenario involved in a video generation method proposed in an embodiment of this application;
[0055] Figure 1B Schematic diagram of an initial video data proposed in an embodiment of this application;
[0056] Figure 1C Schematic diagram of each marker point obtained based on a human figure proposed in an embodiment of this application;
[0057] Figure 1D Schematic diagram of the trajectory of each marker point proposed in an embodiment of this application;
[0058] Figure 1E Schematic flow chart of restoring action data proposed in an embodiment of this application;
[0059] Figure 1F Schematic diagram of replicated video data proposed in an embodiment of this application;
[0060] Figure 2 Schematic flow chart of a video generation method provided in an embodiment of this application;
[0061] Figure 3 Schematic diagram of an operation interface provided in an embodiment of this application;
[0062] Figure 4 Schematic flow chart of generating replicated video data provided in an embodiment of this application;
[0063] Figure 5 Schematic diagram of a certain video frame in the initial video data provided in an embodiment of this application;
[0064] Figure 6 Schematic diagram of a certain video frame in the initial video data after removing the human figure provided in an embodiment of this application;
[0065] Figure 7 Schematic diagram of combining a virtual character with a background image provided in an embodiment of this application;
[0066] Figure 8 Structure block diagram of a video generation device provided in an embodiment of this application;
[0067] Figure 9 Structure schematic diagram of an electronic device provided in an embodiment of this application. Detailed implementation manners
[0068] In the following description, specific details such as specific system architectures and technologies are presented for the purpose of illustration rather than limitation, in order to provide a thorough understanding of the embodiments of the present application. However, those skilled in the art should clearly understand that the present application can also be implemented in other embodiments without these specific details. In other cases, detailed descriptions of well-known human recognition technologies, human erasure technologies, video completion algorithms, and electronic devices are omitted to avoid unnecessary details from hindering the description of the present application.
[0069] The terms used in the following embodiments are only for the purpose of describing specific embodiments and are not intended to limit the present application. As used in the specification and appended claims of the present application, the singular forms "a", "the", "above-mentioned", and "this" are also intended to include the forms such as "one or more", unless there is a clear contrary indication in the context.
[0070] With the continuous development of virtual reality technology, by using motion capture technology to recognize and capture the actions and expressions of the human body, and displaying the corresponding actions and expressions of the human body through virtual characters, the similarity between virtual characters and real characters can be improved.
[0071] In the related art, a motion capture device can capture human actions and transmit the captured data to a terminal device. Correspondingly, the terminal device can render a pre-set virtual character according to the captured data, so that the actions of the virtual character are similar to the human actions.
[0072] However, in practical applications, the human body needs to wear a motion capture device in advance so that the motion capture can be completed through the motion capture device, which requires a relatively long time and high cost to complete the motion capture, and thus will cause the problems of high cost and low efficiency of motion capture.
[0073] Therefore, the embodiments of the present application propose a video generation method, which captures the actions of the characters in the video data and drives the virtual characters according to the captured actions, and can highly restore the movement trajectory of the characters in the video without wearing any motion capture device, thereby reducing the cost of motion capture and improving the efficiency of motion capture.
[0074] See Figure 1A , Figure 1A FIG. is a schematic diagram of a video generation scenario related to a video generation method proposed by an embodiment of the present application. The video generation field may include: a terminal device 110 and a server 120.
[0075] Among them, the terminal device 110 is connected to the server 120 through a link.
[0076] During the process of video generation, the terminal device 110 can transmit initial video data including a human portrait to the server 120 according to an operation triggered by the user. Correspondingly, the server 120 can receive the initial video data and generate replicated video data based on the initial video data.
[0077] The human portrait is used to represent the image corresponding to the photographed user in the initial video data.
[0078] See Figures 1B to 1F , Figure 1B which is a schematic diagram of an initial video data proposed in an embodiment of the present application, Figure 1C which is a schematic diagram of each landmark point obtained based on the human portrait proposed in an embodiment of the present application, Figure 1D which is a schematic diagram of the trajectory of each landmark point proposed in an embodiment of the present application, Figure 1E which is a schematic diagram of a process for restoring action data proposed in an embodiment of the present application, Figure 1F which is a schematic diagram of replicated video data proposed in an embodiment of the present application.
[0079] After receiving the replication instruction and the initial video data sent by the terminal device 110, the server 120 can identify the human portrait in the initial video data according to the replication instruction, and determine each joint point included in the human portrait, so as to mark each joint point, and further record the movement trajectory of each joint point to obtain trajectory information.
[0080] After that, the server 120 can control the virtual skeleton to perform actions through the trajectory information, and generate action data for instructing the actions of the virtual character in combination with a pre-set algorithm. Correspondingly, the server 120 can redirect the action data to the virtual character, that is, apply the action data to the virtual character, so that the virtual character can perform the same actions as the human portrait in the initial video data, thereby generating replicated video data composed of the virtual character.
[0081] The replication instruction can be generated according to an operation triggered by the user. For example, when the terminal device is playing the initial video data, it can detect a replication action triggered by the user for the person in a certain frame image of the initial video data, then the terminal device can generate a replication instruction for replicating the actions of the person according to the replication action.
[0082] Moreover, during the process of generating the replication instruction, the terminal device can also trigger other operations on the person in the frame image according to the replication operation triggered by the user. For example, the terminal device can scale the size of the person, adjust the perspective such as the position and posture of the person, or trigger other operations on the person. The embodiments of the present application do not make specific limitations thereto.
[0083] It should be noted that the above takes the server 120 generating replicated video data as an example for illustration. In actual applications, the terminal device 110 can also generate replicated video data. In this case, the terminal device 110 does not need to send a replication instruction and the initial video data to the server 120. The terminal device 110 can process the initial video data according to the replication instruction to obtain the replicated video data. The embodiments of the present application do not specifically limit the execution entity for generating the replicated video data.
[0084] The following takes the terminal device as an example to introduce in detail the process of the terminal device generating replicated video data.
[0085] Figure 2 It is a schematic flowchart of a video generation method provided by an embodiment of the present application. As an example but not a limitation, refer to Figure 2 This method includes:
[0086] Step 201: Identify the human figures in the initial video data to obtain the trajectory information corresponding to each marked point in the human figures.
[0087] The terminal device can detect the operation triggered by the user and generate a replication instruction according to the detected operation, so as to determine the initial video data according to the replication instruction and determine the relevant information of the virtual character required in the generated replicated video data. In subsequent steps, the terminal device can generate the replicated video data by combining the initial video data with the virtual character.
[0088] For example, as Figure 3 shown, the terminal device can display an operation interface to the user and obtain the storage path corresponding to the initial video data according to the operation triggered by the user. Moreover, the terminal device also provides options such as the number and type of virtual characters to the user in this operation interface. When the terminal device detects the determined operation triggered by the user, the terminal device can generate a replication instruction according to the information in the operation interface and process the initial video data according to the replication instruction.
[0089] Optionally, after the terminal device obtains the initial video data according to the storage path in the replication instruction, the terminal device can first identify the human figures in the initial video data, determine each joint point of the human figure, and then establish a marked point corresponding to each joint point. For each marked point, the terminal device can record the position change information and angle change information of the marked point and combine them to obtain the trajectory information of the marked point.
[0090] Specifically, the terminal device can identify the human figures in the initial video data according to a pre-set algorithm, determine each joint point included in the human figure according to the actions of the human figure, so as to mark each joint point and generate a marked point corresponding to each joint point.
[0091] After that, the terminal device can establish a coordinate system for the scene in the initial video data to determine the coordinate information corresponding to each marker point. Moreover, the terminal device can analyze each video frame of the initial video data to determine the coordinate information corresponding to each marker point in each video frame. Thus, based on the multiple coordinate information corresponding to each marker point, the position change information and angle change information of each joint point can be determined, and the trajectory information corresponding to each marker point can be obtained.
[0092] For example, if the terminal device performs human body recognition on the initial video data and obtains the elbow joint of the human body, the terminal device can establish a marker point corresponding to the elbow joint, record the positions of the elbow joint in each video frame of the initial video data, and the angle change information between the upper arm and the forearm connected to the elbow joint. Thus, based on the respective position change information and angle change information, the trajectory information of the elbow joint can be generated.
[0093] Step 202: Generate action data according to the trajectory information.
[0094] Among them, the action data is used to indicate the actions performed by the virtual character.
[0095] After obtaining the trajectory information of each marker point, the terminal device can restore the actions of each joint point of the human body according to the trajectory information, and thus obtain the action data through the combination of the actions of each joint point, so that in the subsequent steps, the terminal device can control the virtual character according to the action data.
[0096] Optionally, the terminal device can control the movement of the pre-set virtual skeleton according to the trajectory information to obtain the movement data of the virtual skeleton, and then generate the action data through the movement trajectory of the virtual skeleton according to the ground information, shooting information, and pre-set algorithm.
[0097] Among them, the shooting information may include: shooting coordinates and shooting angle. The initial video data is obtained by shooting the human body with a shooting device. In the coordinate system established based on the scene of the initial video data, the shooting coordinates are the position where the shooting device is located, and the shooting angle is used to represent the included angles between the shooting device and each coordinate axis in the coordinate system.
[0098] Similarly, the ground information is used to represent the plane where the human figure is located in the scene of the initial video data.
[0099] It should be noted that the terminal device can analyze the plane where the human figure is located according to the human figure in the initial video data to determine the ground information where the human figure is located. Moreover, the terminal device can also analyze the human figure in the initial video data to determine the shooting coordinates and shooting angle corresponding to the shooting device.
[0100] Correspondingly, after obtaining the trajectory information, the terminal device can establish the correspondence between each marked point in the trajectory information and each joint point in the pre-set virtual skeleton, and then control the virtual skeleton to move according to the position change information and angle change information of each marked point, and record the posture of the virtual skeleton to obtain the motion data of the virtual skeleton.
[0101] After that, the terminal device can correct the motion data of the virtual skeleton through the pre-set algorithms related to rigid body dynamics and contact response respectively, and combine the obtained shooting information and ground information to obtain the action data.
[0102] Step 203: Apply the action data to the virtual character to generate the replicated video data.
[0103] Corresponding to step 202, after obtaining the action data, the terminal device can redirect the action data to the pre-set virtual character, and the virtual character executes the action corresponding to the action data, so that the replicated video data can be generated based on the moving virtual character.
[0104] During the process of generating the replicated video data according to the virtual character, the terminal device can further restore the environment where the virtual character is located according to the initial video data to improve the similarity of the replicated video data.
[0105] Therefore, as Figure 4 shown, step 203 may include the following sub-steps:
[0106] Step 203a: Redirect at least one virtual character according to the action data.
[0107] Corresponding to step 201, the terminal device can determine the initial video data according to the generated replication instruction, and the terminal device can also determine information such as the number and type of virtual characters according to the replication instruction. Correspondingly, after obtaining the action data, the terminal device can apply the action data to at least one virtual character that matches the replication instruction according to the replication instruction, and complete the redirection of the virtual character according to the action data, so that the virtual character can execute the corresponding action according to the action data.
[0108] It should be noted that in practical applications, different virtual characters may have different body proportions. Therefore, the terminal device can execute step 203a to adjust the body proportion of each virtual character for different virtual characters, so as to improve the beauty of each action completed by the virtual character.
[0109] Step 203b: For each virtual character, adjust the body proportion of the virtual character according to the proportion between the bones at all levels of the virtual character.
[0110] After redirecting the virtual character through the action data, the terminal device can adjust the body proportion of the virtual character based on the actions performed by the virtual character and in combination with the size ratios between the bones at all levels, so as to beautify the actions of the virtual character.
[0111] For example, the terminal device can obtain the sizes corresponding to the upper arm of the parent bone and the lower arm of the child bone, calculate the size ratio between the upper arm and the lower arm, and then adjust the sizes of the upper arm and the lower arm respectively in combination with the action data, so as to complete the adjustment of the body proportion of the virtual character.
[0112] Step 203c: Optimize the action data according to a preset filtering algorithm.
[0113] After the terminal device restores the action data, it can further optimize the action data, so that when the virtual character moves according to the action data, it can complete each action continuously and smoothly, and avoid abnormalities in the movement of each joint of the virtual character.
[0114] Specifically, the terminal device can model the movement process of the virtual character according to dynamics and kinematics, so as to maximize the physical coherence of the movement process of the virtual character as the objective function of the model, thereby optimizing the action data, so that when the virtual character moves according to the optimized action data, the movement chain formed by the bones at all levels of the virtual character can conform to the Newton-Euler formula.
[0115] For example, the terminal device can complete the modeling based on the physical coherence of kinematics according to factors such as the prior probability distribution of human motion data, joint angle end, and motion constraints. Similarly, when the terminal device models the physical coherence of dynamics, it can control the joint speed and acceleration of the virtual character to conform to Newton's second law, or it can also control the limbs of the virtual character to conform to Newton's third law and tribology laws when in contact with the ground (such as collision). The embodiments of the present application do not make specific limitations on the modeling method of the terminal device.
[0116] Step 203d: Determine the light information for irradiating the portrait according to the initial video data, and render the virtual character according to the light information.
[0117] In order to further restore the lighting environment of the portrait in the initial video data, the terminal device can extract the environmental scene of the initial video data according to a certain video frame of the initial video data, and determine information such as the color, normal, depth, roughness, and illumination intensity of the environmental scene, so as to obtain the light information of the environmental scene.
[0118] After that, the terminal device can determine the light information, irradiate the virtual character in the same environment, and obtain the light and shadow effect of the virtual character when irradiated by light, so as to complete the rendering of the virtual character and realize the further restoration of the portrait.
[0119] Step 203e: Generate replicated video data according to the adjusted virtual character.
[0120] After the terminal device adjusts and restores the virtual character, the terminal device can record the process of the virtual character's movement, so as to obtain the replicated video data generated by the virtual character, realizing the replication of the portrait in the initial video data through the virtual character.
[0121] Optionally, the terminal device can also adjust the angle of the virtual character according to the shooting information, so that the recorded virtual character and the portrait in the initial video data are in the same perspective. Correspondingly, the terminal device can combine the adjusted virtual character with the background image to generate replicated video data.
[0122] Among them, the background image can be pre-set by the user or the background image in the initial video data. Of course, the background image can be an image or pre-set video data, and the embodiments of the present application do not make specific limitations on the background image.
[0123] Moreover, if the background image is pre-set by the user, the background image can be a two-dimensional plane image or a three-dimensional space scene; if the background image is the background image in the initial video data, the background image can be a two-dimensional plane image obtained based on the initial video data, and the embodiments of the present application do not make specific limitations on this.
[0124] See Figure 5 、 Figure 6 and Figure 7 , Figure 5 which is a schematic diagram of a certain video frame in the initial video data provided by the embodiments of the present application, Figure 6 which is a schematic diagram of a certain video frame in the initial video data after removing the portrait provided by the embodiments of the present application, Figure 7 which is a schematic diagram of combining the virtual character with the background image provided by the embodiments of the present application.
[0125] See Figure 5 and Figure 6 When the terminal device obtains the background image corresponding to the initial video data, the terminal device can first perform portrait removal processing on a certain frame of the initial video data (as shown in Figure 5 ), that is, the terminal device can use the portrait erasure technology on a certain frame of image to form a background image including a blank area. After that, the terminal device can fill the blank area according to the pixels near the blank area to obtain the background image (as shown inFigure 6 as shown). Similarly, the terminal device can adopt a video completion algorithm to combine the generated background image with the adjusted virtual character, thereby generating replicated video data (such as Figure 7 as shown).
[0126] It should be noted that in practical applications, the terminal device can also obtain the background image corresponding to the initial video data in other ways. For example, the terminal device can use a portrait erasure technique on multiple video frames in the initial video data to obtain multiple background images including blank areas, and then combine the background images including blank areas to complete the filling of the blank areas. Correspondingly, if the combined background image still has blank areas, the terminal device can fill the blank areas according to the pixels near the blank areas, thereby obtaining the background image.
[0127] Among them, the multiple video frames can be consecutive video frames, adjacent video frames, or video frames with similar background images. The embodiments of the present application do not specifically limit the positions of the multiple video frames in the initial video data.
[0128] Of course, the terminal device can also extract the background image in other ways, and the embodiments of the present application do not specifically limit this.
[0129] In addition, it should be noted that in practical applications, the terminal device can execute step 203c and step 203d, or execute any one of step 203c and step 203d, or not execute step 203c and step 203d. The embodiments of the present application do not specifically limit this.
[0130] Furthermore, if the terminal device executes step 203c and step 203d, the terminal device can execute step 203c first and then step 203d; or execute step 203d first and then step 203c; or execute step 203c and step 203d simultaneously. The embodiments of the present application do not specifically limit the execution order of step 203c and step 203d by the terminal device.
[0131] In summary, a video generation method proposed by the embodiments of the present application identifies the portrait in the initial video data to obtain the trajectory information corresponding to each marker point in the portrait, and then generates action data according to the trajectory information. The action data is used to indicate the actions performed by the virtual character, and finally the action data is applied to the virtual character to generate replicated video data, without the need to capture actions by real people, which can reduce the cost of action capture and improve the efficiency of generating video data.
[0132] It should be understood that the sequence numbers of the steps in the above embodiments do not imply the order of execution. The order of execution of each process should be determined according to its function and internal logic, and should not constitute any limitation to the implementation process of the embodiments of the present application.
[0133] Corresponding to the data screening method described in the above embodiments, Figure 8 FIG. 3 is a structural block diagram of a video generation device provided by an embodiment of the present application. For the convenience of description, only the parts related to the embodiments of the present application are shown.
[0134] See Figure 8 , the device includes:
[0135] An identification module 801, configured to identify a person in the initial video data, and obtain trajectory information corresponding to each marker point in the person;
[0136] A first generation module 802, configured to generate action data according to the trajectory information, where the action data is used to indicate an action performed by a virtual character;
[0137] A second generation module 803, configured to apply the action data to the virtual character to generate replicated video data.
[0138] Optionally, the identification module 801 is specifically configured to identify a person in the initial video data, determine each joint point of the person; establish a marker point corresponding to each joint point; for each marker point, record the position change information and angle change information of the marker point, and combine them to obtain the trajectory information of the marker point.
[0139] Optionally, the device further includes:
[0140] A first determination module 804, configured to determine ground information where the person is located and shooting information corresponding to the person according to the person in the initial video data, where the shooting information includes: shooting coordinates and shooting angle;
[0141] The first generation module 802 is specifically configured to control the movement of a pre-set virtual skeleton according to the trajectory information to obtain movement data of the virtual skeleton; and generate the action data through the movement trajectory of the virtual skeleton according to the ground information, shooting information and a pre-set algorithm.
[0142] Optionally, the second generation module 803 is specifically configured to redirect at least one virtual character according to the action data; for each virtual character, adjust the body proportion of the virtual character according to the ratio between the bones at all levels of the virtual character; and generate the replicated video data according to the adjusted virtual character.
[0143] Optionally, the device further includes:
[0144] An optimization module 805 for optimizing the action data according to a preset filtering algorithm.
[0145] Optionally, the second generation module 803 is specifically configured to adjust the angle of the virtual character according to the shooting information; combine the adjusted virtual character with the background image to generate the replicated video data.
[0146] Optionally, the apparatus further includes:
[0147] A second determination module 806 for determining the light information for irradiating the portrait according to the initial video data;
[0148] A rendering module 807 for rendering the virtual character according to the light information.
[0149] Optionally, the second generation module 803 is specifically configured to perform a de-portrait processing on the initial video data, fill the position occupied by the portrait, obtain the background image; combine the adjusted virtual character with the background image to generate the replicated video data.
[0150] In summary, a video generation apparatus proposed in an embodiment of the present application identifies the portrait in the initial video data to obtain the trajectory information corresponding to each marker point in the portrait, then generates action data according to the trajectory information, and the action data is used to indicate the actions performed by the virtual character. Finally, the action data is applied to the virtual character to generate replicated video data, without the need to capture actions by real people, which can reduce the cost of action capture and improve the efficiency of generating video data.
[0151] Based on the same inventive concept, an embodiment of the present application further provides a terminal device. Figure 9 The structural schematic diagram of the electronic device provided in an embodiment of the present application is as Figure 9 shown. The electronic device provided in this embodiment includes: a memory 91 and a processor 92. The memory 91 is used to store a computer program 93; the processor 92 is used to execute the method described in the above method embodiment when calling the computer program 93.
[0152] The electronic device provided in this embodiment can execute the above method embodiment, and its implementation principle and technical effects are similar, and will not be elaborated here.
[0153] An embodiment of the present application further provides a computer-readable storage medium, on which a computer program is stored, and when the computer program is executed by a processor, the method described in the above method embodiment is implemented.
[0154] The embodiments of the present application also provide a computer program product. When the computer program product runs on an electronic device, it enables the electronic device to execute the method described in the above method embodiments when executed.
[0155] If the above integrated unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on such an understanding, to implement all or part of the processes in the above method embodiments of the present application, a computer program can be used to instruct the relevant hardware to complete. The computer program can be stored in a computer-readable storage medium. When the computer program is executed by a processor, the steps of the above various method embodiments can be implemented. Among them, the computer program includes computer program code, and the computer program code can be in the form of source code, object code, executable file or some intermediate form, etc. The computer-readable storage medium can at least include: any entity or device that can carry the computer program code to the photographing device / terminal device, recording medium, computer memory, read-only memory (ROM), random access memory (RAM), electrical carrier signal, telecommunication signal, and software distribution medium. For example, a USB flash drive, a mobile hard disk, a magnetic disk or an optical disc, etc. In some jurisdictions, according to legislation and patent practice, the computer-readable medium cannot be an electrical carrier signal and a telecommunication signal.
[0156] In the above embodiments, the descriptions of the various embodiments have their own emphases. For the parts not detailed or recorded in a certain embodiment, reference can be made to the relevant descriptions of other embodiments.
[0157] Those of ordinary skill in the art can realize that the units and algorithm steps of the examples described in combination with the embodiments disclosed herein can be implemented by electronic hardware, or a combination of computer software and electronic hardware. Whether these functions are executed in a hardware or software manner depends on the specific application and design constraints of the technical solution. Professional technicians can use different methods to implement the described functions for each specific application, but such implementation should not be considered to exceed the scope of the present application.
[0158] In the embodiments provided in this application, it should be understood that the disclosed device / equipment and method can be implemented in other ways. For example, the device / equipment embodiments described above are merely illustrative. For example, the division of the modules or units is only a logical function division. In actual implementation, there may be other division methods. For example, multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the displayed or discussed coupling or direct coupling or communication connection to each other can be through some interfaces. The indirect coupling or communication connection of the device or unit can be in electrical, mechanical or other forms.
[0159] It should be understood that when used in the specification and appended claims of this application, the term "comprising" indicates the presence of the described features, wholes, steps, operations, elements and / or components, but does not exclude the presence or addition of one or more other features, wholes, steps, operations, elements, components and / or their combinations.
[0160] It should also be understood that the term "and / or" as used in the specification and appended claims of this application refers to any combination and all possible combinations of one or more of the associated listed items, and includes these combinations.
[0161] As used in the specification and appended claims of this application, the term "if" can be interpreted as "when" or "once" or "in response to determining" or "in response to detecting" depending on the context. Similarly, the phrase "if determined" or "if [the described condition or event] is detected" can be interpreted as meaning "once determined" or "in response to determining" or "once [the described condition or event] is detected" or "in response to detecting [the described condition or event]" depending on the context.
[0162] In addition, in the description of the specification and appended claims of this application, the terms "first", "second", "third", etc. are only used for distinguishing descriptions and cannot be understood as indicating or implying relative importance.
[0163] The reference to "one embodiment" or "some embodiments" or the like described in the specification of this application means that a specific feature, structure or characteristic described in connection with the embodiment is included in one or more embodiments of this application. Thus, the statements "in one embodiment", "in some embodiments", "in other some embodiments", "in still other embodiments", etc. that appear in different places in this specification do not necessarily all refer to the same embodiment, but mean "one or more but not all embodiments", unless otherwise specifically emphasized in other ways. The terms "comprising", "including", "having" and their variants all mean "including but not limited to", unless otherwise specifically emphasized in other ways.
[0164] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present application, rather than to limit them; although the present application has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that they can still modify the technical solutions described in the foregoing embodiments, or perform equivalent replacements on some or all of the technical features; and these modifications or replacements do not cause the essence of the corresponding technical solutions to deviate from the scope of the technical solutions of the various embodiments of the present application.
Claims
1. A video generation method, characterized in that, the method includes: identifying the human figure in the initial video data to obtain the trajectory information corresponding to each marked point in the human figure; generating action data according to the trajectory information, where the action data is used to indicate the actions performed by the virtual character; applying the action data to the virtual character to generate replicated video data.
2. The method according to claim 1, characterized in that, the identifying the human figure in the initial video data to obtain the trajectory information corresponding to each marked point in the human figure includes: identifying the human figure in the initial video data to determine each joint point of the human figure; establishing a marked point corresponding to each joint point; for each marked point, recording the position change information and angle change information of the marked point, and combining them to obtain the trajectory information of the marked point.
3. The method according to claim 1, characterized in that, before generating the action data according to the trajectory information, the method further includes: determining the ground information where the human figure is located and the shooting information corresponding to the human figure according to the human figure in the initial video data, where the shooting information includes: shooting coordinates and shooting angle; the generating the action data according to the trajectory information includes: controlling the movement of pre-set virtual bones according to the trajectory information to obtain the movement data of the virtual bones; generating the action data through the movement trajectory of the virtual bones according to the ground information, shooting information and pre-set algorithm.
4. The method according to claim 1, characterized in that, the applying the action data to the virtual character to generate replicated video data includes: redirecting at least one virtual character according to the action data; for each virtual character, adjusting the body proportion of the virtual character according to the proportion between the bones at all levels of the virtual character; generating the replicated video data according to the adjusted virtual character.
5. The method according to claim 4, characterized in that, before generating the replicated video data according to the adjusted virtual character, the method further includes: optimizing the action data according to a pre-set filtering algorithm.
6. The method according to any one of claims 1 to 4, characterized in that, the applying the action data to the virtual character to generate replicated video data includes: adjusting the angle of the virtual character according to the shooting information; combining the adjusted virtual character with the background image to generate the replicated video data.
7. The method according to claim 6, characterized in that, before combining the adjusted virtual character with the background image to generate the replicated video data, the method further includes: determining the light information for irradiating the human figure according to the initial video data; rendering the virtual character according to the light information.
8. The method according to claim 6, characterized in that, the combining the adjusted virtual character with the background image to generate the replicated video data includes: Perform portrait removal processing on the initial video data, and fill the positions occupied by the portrait to obtain the background image; Combine the adjusted virtual character with the background image to generate the replicated video data.
9. An electronic device, characterized in that, comprising: A memory and a processor, the memory is used to store a computer program; the processor is used to execute the method according to any one of claims 1-8 when calling the computer program.
10. A computer-readable storage medium, on which a computer program is stored, characterized in that, When the computer program is executed by a processor, the method according to any one of claims 1-8 is implemented.
Citation Information
Cited By
E-commerce video intelligent duplicating method and system
CN121967832A