Method and apparatus for rendering virtual human interactive video

By pre-rendering virtual human body movement videos and lip position information, and only rendering lip videos in real time to synthesize virtual human interactive videos, the problem of high computational load in virtual human interactive videos is solved, application costs are reduced, and the application scope is expanded.

CN116016986BActive Publication Date: 2026-03-31SHANGHAI GAUDIAN INTELLIGENT TECHNOLOGY GROUP CO LTD +1
View PDF 3 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2023-01-09
Publication Date
2026-03-31

AI Technical Summary

Technical Problem

The real-time rendering computation of existing virtual human interactive videos is computationally intensive, resulting in high application costs and limiting their application scope and scenarios.

Method used

By pre-rendering virtual human body movement videos and lip position information, only the lip video is rendered in real time, and then combined with the pre-rendered body movement videos to create a virtual human interactive video.

Benefits of technology

It significantly reduces the computational load of real-time rendering of interactive virtual human videos, lowers application costs, and expands its application scope and scenarios.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN116016986B_ABST
    Figure CN116016986B_ABST
Patent Text Reader

Abstract

The application provides a rendering method and device for virtual human interactive video, and the method comprises the following steps: obtaining a voice to be broadcast; selecting limb action data matched with the voice to be broadcast in a limb action video library as target limb action data; the target limb action data comprises a limb action video pre-rendered based on a virtual human limb action, lip position information and lip posture information of the limb action video; rendering a lip video according to the lip posture information and the voice to be broadcast; and fusing the lip video and the limb action video based on the lip position information to obtain a virtual human interactive video for outputting the voice to be broadcast. According to the scheme, only the lip video needs to be rendered in real time, the lip video and the pre-rendered limb action video can be combined into a completed virtual human interactive video, and the calculation amount required for rendering the virtual human interactive video in real time is significantly reduced.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of virtual human interaction technology, and in particular to a rendering method and apparatus for virtual human interactive videos. Background Technology

[0002] Virtual human interactive video refers to a video with the following characteristics: the video screen displays a three-dimensional human model, and when the video is played, the human model's limbs (including limbs and torso) and lips change in accordance with the voice output in the video.

[0003] During human-computer interaction, devices can play virtual human interactive videos to simulate real human speech, improving the interactive experience. Therefore, virtual human interactive videos are increasingly being used in scenarios such as smart shopping guides, smart navigation systems, smart front desks, and mobile assistants.

[0004] In the above scenarios, the voice output by the virtual human interactive video usually carries a lot of real-time information that changes constantly, such as time, weather, stock, business status, personal information, etc. Therefore, the virtual human interactive video must be rendered in real time, that is, the virtual human interactive video needs to be generated and output in a short time after obtaining user input.

[0005] However, real-time rendering of videos containing 3D character models requires a significant amount of computation. Especially with advancements in modeling technology, the precision of 3D character models is increasing, leading to a surge in computational demands for rendering the corresponding videos. This issue greatly increases the application cost of virtual human interactive videos, limiting the scope and application scenarios of this technology. Summary of the Invention

[0006] To address the shortcomings of the prior art, the present invention provides a rendering method and apparatus for virtual human interactive videos, thereby reducing the computational load required for real-time rendering of virtual human interactive videos.

[0007] The first aspect of this application provides a method for rendering virtual human interactive videos, including:

[0008] Obtain the speech to be played; and select the body movement data that matches the speech to be played from the body movement video library as the target body movement data; wherein, the target body movement data includes a body movement video pre-rendered based on the body movement of a virtual human, and the lip position information and lip posture information of the body movement video.

[0009] Render the lip video based on the lip posture information and the speech to be played;

[0010] Based on the lip position information, the lip video and the body movement video are fused to obtain a virtual human interactive video for outputting the speech to be broadcast.

[0011] Optionally, rendering the lip video based on the lip pose information and the speech to be played includes:

[0012] Synchronize the timeline of the body movement video and the timeline of the audio to be played.

[0013] For each audio frame in the speech to be played, lip pose data of the corresponding motion video frame is obtained from the lip pose information, and a lip video frame corresponding to the audio frame is rendered based on the audio frame and the lip pose data; wherein, the motion video frame refers to the video frame of the limb motion video; and the lip video frame refers to the video frame that makes up the lip video.

[0014] Optionally, the step of fusing the lip video and the body movement video based on the lip position information to obtain the virtual human interactive video for outputting the speech to be broadcast includes:

[0015] For each of the aforementioned action video frames, the lip position data of the action video frame is obtained from the lip position information, and the lip video frame corresponding to the action video frame is superimposed on the position indicated by the lip position data in the action video frame to obtain the interactive video frame corresponding to the action video frame; wherein, a series of consecutive interactive video frames constitute the virtual human interactive video.

[0016] Optionally, selecting body movement data from the body movement video library that matches the speech to be broadcast as the target body movement data includes:

[0017] Identify the target body movements that match the voice content to be broadcast;

[0018] Select the limb movement data corresponding to the target limb movement from the limb movement video library as the target limb movement data.

[0019] Optionally, the body movement data includes facial expression tags, which represent virtual facial expressions when rendering the body movement video;

[0020] The step of selecting body movement data from the body movement video library that corresponds to the target body movement as the target body movement data includes:

[0021] The body movement data in the body movement video library that corresponds to the target body movement and has the same facial expression tag as the speech to be played is selected as the target body movement data; wherein, the facial expression tag of the speech to be played is determined according to the speech content of the speech to be played.

[0022] A second aspect of this application provides a rendering apparatus for virtual human interactive video, comprising:

[0023] The acquisition unit is used to acquire the speech to be played; and select the body movement data that matches the speech to be played from the body movement video library as the target body movement data; wherein, the target body movement data includes a body movement video pre-rendered based on the body movement of a virtual human, and the lip position information and lip posture information of the body movement video.

[0024] A rendering unit is used to render a lip video based on the lip pose information and the speech to be played.

[0025] The fusion unit is used to fuse the lip video and the body movement video based on the lip position information to obtain a virtual human interactive video for outputting the speech to be broadcast.

[0026] Optionally, when the rendering unit renders the lip video based on the lip pose information and the speech to be played, it is specifically used for:

[0027] Synchronize the timeline of the body movement video and the timeline of the audio to be played.

[0028] For each audio frame in the speech to be played, lip posture data of the corresponding motion video frame is obtained from the lip posture information, and a lip video frame corresponding to the audio frame is synthesized based on the audio frame and the lip posture data; wherein, the motion video frame refers to the video frame of the limb motion video; and the lip video frame refers to the video frame that makes up the lip video.

[0029] Optionally, when the fusion unit fuses the lip video and the body movement video based on the lip position information to obtain the virtual human interactive video for outputting the speech to be broadcast, it is specifically used for:

[0030] For each of the aforementioned action video frames, the lip position data of the action video frame is obtained from the lip position information, and the lip video frame corresponding to the action video frame is superimposed on the position indicated by the lip position data in the action video frame to obtain the interactive video frame corresponding to the action video frame; wherein, a series of consecutive interactive video frames constitute the virtual human interactive video.

[0031] Optionally, when the acquisition unit selects body movement data from the body movement video library that matches the speech to be played as the target body movement data, it is specifically used for:

[0032] Identify the target body movements that match the voice content to be broadcast;

[0033] Select the limb movement data corresponding to the target limb movement from the limb movement video library as the target limb movement data.

[0034] Optionally, the body movement data includes facial expression tags, which represent virtual facial expressions when rendering the body movement video;

[0035] When the acquisition unit selects limb movement data corresponding to the target limb movement from the limb movement video library as the target limb movement data, it is specifically used for:

[0036] The body movement data in the body movement video library that corresponds to the target body movement and has the same facial expression tag as the speech to be played is selected as the target body movement data; wherein, the facial expression tag of the speech to be played is determined according to the speech content of the speech to be played.

[0037] This application provides a method and apparatus for rendering virtual human interactive videos. The method includes: obtaining speech to be played; selecting body movement data matching the speech from a body movement video library as target body movement data; the target body movement data includes a pre-rendered body movement video based on the virtual human's body movements, lip position information and lip pose information of the body movement video; rendering a lip video based on the lip pose information and the speech to be played; and fusing the lip video and the body movement video based on the lip position information to obtain a virtual human interactive video for outputting the speech to be played. This solution only requires real-time rendering of the lip video to synthesize the lip video and the pre-rendered body movement video into a complete virtual human interactive video, significantly reducing the computational load required for real-time rendering of virtual human interactive videos. Attached Figure Description

[0038] To more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are only embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on the provided drawings without creative effort.

[0039] Figure 1 A flowchart illustrating a method for rendering virtual human interactive video provided in an embodiment of this application;

[0040] Figure 2 A schematic diagram of lip pose data provided in an embodiment of this application;

[0041] Figure 3 A schematic diagram of lip position data provided in an embodiment of this application;

[0042] Figure 4 This is a schematic diagram of the structure of a rendering device for virtual human interactive video provided in an embodiment of this application. Detailed Implementation

[0043] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.

[0044] This application provides a method for rendering virtual human interactive videos. Please refer to [link to relevant documentation]. Figure 1 Here is a flowchart of the method, which may include the following steps.

[0045] S101, obtain the voice to be played; and select the body movement data that matches the voice to be played from the body movement video library as the target body movement data.

[0046] The target limb movement data includes pre-rendered limb movement videos based on virtual human limb movements, as well as lip position and lip pose information from the limb movement videos.

[0047] The method for obtaining the voice message to be played is as follows:

[0048] First, the text to be synthesized is determined based on the input user command. The form of the user command is not limited; it can be in the form of voice or text.

[0049] For example, after a user asks a question, the terminal device captures the user's voice while speaking, identifies the voice as the user's voice command, and then uses voice recognition technology to identify the question asked by the user. Next, it finds the preset answer text based on the question, or uses artificial intelligence algorithms to generate the corresponding answer text based on the question, and identifies the answer text as the text to be synthesized corresponding to the user's command.

[0050] Once the text to be synthesized is obtained, speech synthesis technology can be used to generate the speech to be read aloud. Speech synthesis is a technology that generates artificial speech based on specific text through mechanical and electronic methods. This technology can convert text information generated by computer devices or externally input into fluent, understandable spoken Chinese output. The method provided in this embodiment can utilize any existing speech synthesis technology or speech synthesizer to synthesize the speech to be read aloud.

[0051] The following is an explanation of the body movement video library.

[0052] The body movement video library includes multiple body movement data sets, each corresponding to a specific body movement. A body movement data set may include the following: a body movement video of a 3D human model (hereinafter referred to as a virtual human) performing a specific body movement; lip position information describing the location of the lips in the body movement video; and lip posture information describing the lip pose in the body movement video.

[0053] In one alternative embodiment, the virtual human can maintain a default expression and silent lips in all body movement videos. The default expression can be a smiling expression or other expressions, without limitation.

[0054] Taking a smile as the default expression, when building a body movement video library, we can utilize the complete texture maps of the virtual human, 2D or 3D scene textures, and related prop textures to render the process of a virtual human with a smiling expression and silent lips performing a body movement. This yields a video segment corresponding to that body movement. Simultaneously, during rendering, the lip position data and lip pose data of each video frame are recorded in real time. After rendering, the set of lip position data from all video frames in the body movement video represents the lip position information corresponding to that body movement video, and the set of lip pose data from all video frames represents the lip pose information corresponding to that body movement video. Thus, we can obtain a set of body movement data corresponding to a single body movement. By repeating the above process for each preset body movement that the virtual human can perform, we can obtain the body movement data corresponding to each body movement, thereby creating a body movement video library.

[0055] In another alternative embodiment, the virtual human can maintain a silent lip movement in different body movement videos and have different facial expressions. For example, the virtual human may display a smiling expression in some body movement videos and a serious expression in others.

[0056] In this scenario, when building a body movement video library, it's necessary to pre-specify the facial expression tags corresponding to different body movement videos. When rendering the body movement videos, the virtual human is rendered according to the expressions indicated by the facial expression tags. After rendering, the facial expression tags need to be added to the corresponding body movement data. In other words, when the virtual human's expressions differ across different body movement videos, the body movement data can include the body movement video, lip position information, lip pose information, and facial expression tags representing the virtual human's expressions in the video.

[0057] Exemplarily, assume that when rendering a limb movement video corresponding to a certain limb movement, the expression of the virtual human in the video is specified as "serious". Then, during the rendering process, control the virtual human to present a serious expression. After rendering the limb movement video, determine the limb movement video, the corresponding lip position information and lip pose information, and the "serious" label as a piece of limb movement data corresponding to the limb movement.

[0058] Please refer to Figure 2 , the lip pose of the virtual human is defined as the zenith angle and azimuth angle of its normal direction with respect to the global coordinate system. Specifically, the lip region is defined as the part inside the outer edge of the orbicularis oris muscle, and its normal is perpendicular to the plane formed by the outer edge of the orbicularis oris muscle, facing the direction where the lips are facing. Use the spherical coordinate system to define its direction. The angle between the normal direction and the z-axis is the zenith angle (as Figure 2 shown, denoted as angle A), and the angle between the projection on the x-y plane and the x-axis is the azimuth angle (as Figure 2 shown, denoted as angle B). Among them, angle A controls the elevation and depression of the lips. When A = π / 2, the virtual human is in a平视 state. When A < π / 2, the lips are in a仰头 state. When A > π / 2, the virtual human is in a低头 state. Angle B controls the orientation of the lips. When B = 0, it is facing directly forward. When 0 < B < π, it is facing the left side of the virtual human. When π < B < 2π, it is facing the right side of the virtual human.

[0059] Therefore, the lip pose data of a video frame in the limb movement video can be the angle of the zenith angle A and the angle of the azimuth angle B in this video frame. Exemplarily, the lip pose data of a video frame can be: angle A is π / 4, and angle B is π / 3.

[0060] Please refer to Figure 3 , the lip position of the virtual human is defined as the global three-dimensional coordinates of the center point of the outer edge of the virtual human's orbicularis oris muscle. Correspondingly, the lip position data of a video frame in the limb movement video can be the coordinate data of the center point of the outer edge of the orbicularis oris muscle in this video frame, that is, the (x, y, z) coordinates of this point.

[0061] Optionally, the process of selecting the limb movement data in the limb movement video library that matches the to-be-broadcast speech as the target limb movement data in step S101 may include:

[0062] First, determine the target limb movement that matches the speech content of the to-be-broadcast speech;

[0063] Then, select the limb movement data corresponding to the target limb movement in the limb movement video library as the target limb movement data.

[0064] The matching relationship between voice content and body movements can be preset. For example, when the voice content is to instruct the user to click on a certain part of the device, the matching body movement can be to point to the direction of that part. When the voice content is to greet the user, the matching body movement can be to nod and wave.

[0065] It should be noted that a piece of audio to be played can involve multiple aspects of audio content, and therefore can be matched with one or more target body movements. Correspondingly, the target body movement data corresponding to a piece of audio to be played can be one set of data for one movement, or multiple sets of data for multiple movements.

[0066] Once the target limb movement is identified, its data can be found in the limb movement video library. In the video containing the target limb movement data, the virtual human performs the same movements as the target limb movement.

[0067] As mentioned above, in some embodiments, the virtual human can have different expressions in different body movement videos. In this case, the body movement data includes expression tags, which represent the virtual expressions when rendering the body movement videos.

[0068] In this case, when selecting target limb movement data, it is necessary to consider both limb movement and facial expression tags. In other words, selecting limb movement data corresponding to the target limb movement from the limb movement video library as the target limb movement data can include:

[0069] The target body movement data is selected from the body movement video library that corresponds to the target body movement and has the same facial expression label as the speech to be played; the facial expression label of the speech to be played is determined according to the speech content.

[0070] S102 renders the lip video based on lip posture information and the speech to be played.

[0071] The limb movement videos involved in steps S102 and S103 refer to the limb movement videos included in the target limb movement data selected in S101. Similarly, the lip position information and lip posture information involved in steps S102 and S103 refer to the lip position information and lip posture information included in the target limb movement data selected in S101.

[0072] Optionally, the execution process of step S102 may include:

[0073] A1 synchronizes the timeline of the body movement video with the timeline of the audio to be played.

[0074] A2: For each audio frame in the speech to be played, obtain the lip posture data of the corresponding motion video frame from the lip posture information, and render the lip video frame corresponding to the audio frame based on the audio frame and the lip posture data; where motion video frame refers to the video frame of the limb motion video; lip video frame refers to the video frame that makes up the lip video.

[0075] The specific execution process of step A1 is as follows:

[0076] If only one set of target body movement data is determined, then only one body movement video needs to be synchronized in A1. In this case, the start time of the body movement video is set as the start time of the voice to be played, and the end time of the body movement video is set as the end time of the voice to be played.

[0077] After completing the above settings, if the duration of the body movement video is the same as the duration of the voice to be played, the timeline synchronization is successful, and step A1 ends. If the duration of the body movement video is longer than the duration of the voice to be played, the duration of the body movement video can be shortened by speeding up the video playback speed or deleting several video frames from the video to make the two durations the same. If the duration of the body movement video is shorter than the duration of the voice to be played, the duration of the body movement video can be extended by slowing down the video playback speed, copying several video frames from the video and inserting the copied video frames into the original video to make the two durations the same. After the two durations are the same, the timeline synchronization is successful, and step A1 ends.

[0078] If multiple target limb movement data are determined in S101, then multiple limb movement videos need to be synchronized in A1. In this case, the playback order of the multiple limb movement videos is first determined. For example, based on the voice content, it can be determined that action 1 is executed first and action 2 is executed later. Then, the limb movement video corresponding to action 1 is played first and the limb movement video corresponding to action 2 is played later. Then, the multiple limb movement videos are spliced ​​together in the order of playback to obtain a spliced ​​video.

[0079] Once you have the spliced ​​video, you can synchronize the timeline of the spliced ​​video with the timeline of the audio to be played, just like you would with a single video. For details, please refer to the previous text.

[0080] In step A2, for each audio frame in the speech to be played, the lip pose data of the audio frame and the corresponding motion video frame can be input into the real-time lip synthesizer for lip rendering to obtain the lip video frame corresponding to the audio frame.

[0081] The real-time lip shape synthesizer is used to calculate the lip shape when uttering a specific speech under a specific lip posture, based on the speech signal and lip posture data.

[0082] S103, based on lip position information, fuses lip video and body movement video to obtain virtual human interactive video for outputting the speech to be broadcast.

[0083] Optionally, the execution process of step S103 may include:

[0084] For each action video frame, the lip position data of the action video frame is obtained from the lip position information, and the corresponding lip video frame is superimposed on the position indicated by the lip position data in the action video frame to obtain the interactive video frame corresponding to the action video frame; among them, multiple consecutive interactive video frames constitute the virtual human interactive video.

[0085] In S102, the timeline of the body movement video is synchronized with the timeline of the speech to be played. Therefore, each audio frame of the speech to be played corresponds to an action video frame in the body movement video, and the audio frames of the speech to be played and the action video frames of the body movement video correspond one-to-one.

[0086] Each lip video frame in the lip video is synthesized from the speech to be played, and there is a one-to-one correspondence between the audio frames of the speech and the lip video frames. Therefore, a one-to-one correspondence can be established between each lip video frame in the lip video and each movement video frame in the body movement video.

[0087] Based on the above correspondence, in step S103, when rendering the lip video, each lip video frame and its corresponding motion video frame can be superimposed and fused in real time based on the lip position data of the motion video frame to obtain an interactive video frame. The video composed of multiple consecutive interactive video frames is the interactive video used to output the speech to be played.

[0088] The aforementioned interactive videos can be played on terminal devices that support virtual human interactive videos.

[0089] The method provided in this embodiment can be executed by a server in the cloud. By executing the above method, the server generates interactive video in real time and pushes the interactive video to the terminal device for display in the form of a video stream.

[0090] The method provided in this embodiment can also be executed by a terminal device. The terminal device can download a body movement video library from the server to its local machine in advance, and then execute the above method based on the downloaded body movement video library to generate and play interactive videos locally in real time.

[0091] This application provides a method for rendering virtual human interactive videos. The method includes: obtaining speech to be played; selecting body movement data matching the speech from a body movement video library as target body movement data; the target body movement data includes a pre-rendered body movement video based on the virtual human's body movements, lip position information and lip pose information of the body movement video; rendering a lip video based on the lip pose information and the speech to be played; and fusing the lip video and the body movement video based on the lip position information to obtain a virtual human interactive video for outputting the speech to be played. This solution only requires real-time rendering of the lip video to synthesize the lip video and the pre-rendered body movement video into a complete virtual human interactive video, significantly reducing the computational load required for real-time rendering of virtual human interactive videos.

[0092] According to the rendering method for virtual human interactive videos provided in the embodiments of this application, this application also provides a rendering apparatus for virtual human interactive videos. Please refer to [link to relevant documentation]. Figure 4 This is a schematic diagram of the structure of the device, which may include the following units.

[0093] The acquisition unit 401 is used to acquire the speech to be broadcast; and select the body movement data that matches the speech to be broadcast from the body movement video library as the target body movement data; wherein, the target body movement data includes a body movement video pre-rendered based on the body movement of a virtual human, and the lip position information and lip posture information of the body movement video.

[0094] Rendering unit 402 is used to render lip video based on lip pose information and the speech to be played;

[0095] The fusion unit 403 is used to fuse lip video and body movement video based on lip position information to obtain a virtual human interactive video for outputting the speech to be broadcast.

[0096] Optionally, when rendering the lip video based on the lip pose information and the speech to be played, the rendering unit 402 is specifically used for:

[0097] Synchronize the timeline of the body movement video with the timeline of the audio to be broadcast;

[0098] For each audio frame in the speech to be played, lip pose data of the corresponding motion video frame is obtained from the lip pose information, and the lip video frame corresponding to the audio frame is synthesized based on the audio frame and the lip pose data; where motion video frame refers to the video frame of limb motion video; lip video frame refers to the video frame that makes up the lip video.

[0099] Optionally, when the fusion unit 403 fuses lip video and body movement video based on lip position information to obtain a virtual human interactive video for outputting the speech to be played, it is specifically used for:

[0100] For each action video frame, the lip position data of the action video frame is obtained from the lip position information, and the corresponding lip video frame is superimposed on the position indicated by the lip position data in the action video frame to obtain the interactive video frame corresponding to the action video frame; among them, multiple consecutive interactive video frames constitute the virtual human interactive video.

[0101] Optionally, when the acquisition unit 401 selects body movement data from the body movement video library that matches the speech to be broadcast as the target body movement data, it is specifically used for:

[0102] Identify the target body movements that match the audio content to be broadcast;

[0103] Select the limb movement data corresponding to the target limb movement from the limb movement video library as the target limb movement data.

[0104] Optionally, the body movement data includes facial expression tags, which represent virtual facial expressions when rendering body movement videos;

[0105] When the acquisition unit 401 selects the limb movement data corresponding to the target limb movement from the limb movement video library as the target limb movement data, it is specifically used for:

[0106] The target body movement data is selected from the body movement video library that corresponds to the target body movement and has the same facial expression label as the speech to be played; the facial expression label of the speech to be played is determined according to the speech content.

[0107] The specific working principle of the rendering device for virtual human interactive video provided in this embodiment can be found in the relevant steps of the rendering method for virtual human interactive video provided in this application embodiment, and will not be repeated here.

[0108] This application provides a rendering apparatus for virtual human interactive videos. The apparatus includes: an acquisition unit 401 that acquires the speech to be played; and selects body movement data matching the speech from a body movement video library as target body movement data; the target body movement data includes a pre-rendered body movement video based on the virtual human's body movements, lip position information, and lip posture information of the body movement video; a rendering unit 402 that renders the lip video according to the lip posture information and the speech to be played; and a fusion unit 403 that fuses the lip video and the body movement video based on the lip position information to obtain a virtual human interactive video for outputting the speech to be played. This solution only requires real-time rendering of the lip video to synthesize the lip video and the pre-rendered body movement video into a complete virtual human interactive video, significantly reducing the computational load required for real-time rendering of virtual human interactive videos.

[0109] Finally, it should be noted that in this document, relational terms such as "first" and "second" are used only to distinguish one entity or operation from another, and do not necessarily require or imply any such actual relationship or order between these entities or operations. Furthermore, the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such a process, method, article, or apparatus. Without further limitations, an element defined by the phrase "comprising one..." does not exclude the presence of other identical elements in the process, method, article, or apparatus that includes said element.

[0110] It should be noted that the concepts of "first" and "second" mentioned in this invention are only used to distinguish different devices, modules or units, and are not used to limit the order of functions performed by these devices, modules or units or their interdependencies.

[0111] Those skilled in the art will be able to implement or use this application. Various modifications to these embodiments will be readily apparent to those skilled in the art, and the general principles defined herein may be implemented in other embodiments without departing from the spirit or scope of this application. Therefore, this application is not to be limited to the embodiments shown herein, but is to be accorded the widest scope consistent with the principles and novel features disclosed herein.

Claims

1. A method for rendering virtual human interactive videos, characterized in that, The method comprises: obtaining a to-be-broadcast voice; and selecting the body movement data in the body movement video library that matches the to-be-broadcast voice as target body movement data; wherein the target body movement data comprises a body movement video pre-rendered based on a virtual person, lip position information and lip posture information of the body movement video; the lip posture information is used to indicate the orientation of the lips of the virtual person; wherein the lip position is defined as the global three-dimensional coordinates of the center point of the outer edge of the orbicularis oris of the virtual person, and the lip posture is defined as the zenith angle and the azimuth angle formed by the normal direction of the lips and the coordinate axes of the global three-dimensional coordinates; the zenith angle is the included angle between the normal direction of the lips and the Z axis, and the azimuth angle is the included angle between the projection of the normal direction on the X-Y plane and the X axis; rendering a lip video according to the lip posture information and the to-be-broadcast voice; the lip video at least comprises a lip shape corresponding to the corresponding lip posture in each lip video frame; based on the lip position information, fusing the lip video and the body movement video to obtain a virtual person interactive video for outputting the to-be-broadcast voice.

2. The method of claim 1, wherein, The method according to the lip posture information and the to-be-broadcast voice, rendering a lip video, comprises: synchronizing the time axis of the body movement video and the time axis of the to-be-broadcast voice; for each audio frame in the to-be-broadcast voice, obtaining the lip posture data of the action video frame corresponding to the audio frame from the lip posture information, and rendering the lip video frame corresponding to the audio frame according to the audio frame and the lip posture data; wherein the action video frame refers to the video frame of the body movement video; the lip video frame refers to the video frame constituting the lip video.

3. The method of claim 2, wherein, The method based on the lip position information, fusing the lip video and the body movement video to obtain a virtual person interactive video for outputting the to-be-broadcast voice, comprises: for each action video frame, obtaining the lip position data of the action video frame from the lip position information, and superimposing the lip video frame corresponding to the action video frame in the position indicated by the lip position data to obtain the interactive video frame corresponding to the action video frame; wherein a plurality of continuous interactive video frames constitute the virtual person interactive video.

4. The method of claim 1, wherein, The method of selecting the body movement data in the body movement video library that matches the to-be-broadcast voice as target body movement data comprises: determining a target body movement that matches the voice content of the to-be-broadcast voice; selecting the body movement data in the body movement video library that corresponds to the target body movement as target body movement data.

5. The method of claim 4, wherein, The body movement data comprises an expression label, and the expression label represents a virtual expression when rendering the body movement video; The method of selecting the body movement data in the body movement video library that corresponds to the target body movement as target body movement data comprises: The limb action data corresponding to the target limb action and having the same expression label as the to-be-broadcast voice in the limb action video library is selected as target limb action data, wherein the expression label of the to-be-broadcast voice is determined according to the voice content of the to-be-broadcast voice.

6. A rendering apparatus of a virtual human interactive video, characterized in that, Comprise: An acquisition unit is configured to obtain a to-be-broadcast voice; And select the limb action data matching the to-be-broadcast voice in the limb action video library as target limb action data; wherein the target limb action data comprises a limb action video pre-rendered based on a virtual person, and lip position information and lip posture information of the limb action video; the lip posture information is used to indicate the orientation of the lips of the virtual person; wherein the lip position is defined as the global three-dimensional coordinates of the center point of the outer edge of the orbicularis oris muscle of the virtual person, and the lip posture is defined as the zenith angle and the azimuth angle formed by the normal direction of the lips and the coordinate axis of the global three-dimensional coordinates; the zenith angle is the included angle between the normal direction of the lips and the Z axis, and the azimuth angle is the included angle between the projection of the normal direction on the X-Y plane and the X axis; A rendering unit is configured to render a lip video according to the lip posture information and the to-be-broadcast voice; the lip video at least comprises a lip shape corresponding to a corresponding lip posture in each lip video frame; A fusion unit is configured to fuse the lip video and the limb action video based on the lip position information to obtain a virtual person interactive video for outputting the to-be-broadcast voice.

7. The apparatus of claim 6, wherein, When the rendering unit renders a lip video according to the lip posture information and the to-be-broadcast voice, it is specifically configured to: Synchronize the time axis of the limb action video and the time axis of the to-be-broadcast voice; For each audio frame in the to-be-broadcast voice, obtain the lip posture data of the action video frame corresponding to the audio frame from the lip posture information, and synthesize the lip video frame corresponding to the audio frame according to the audio frame and the lip posture data; wherein the action video frame refers to the video frame of the limb action video; the lip video frame refers to the video frame constituting the lip video.

8. The apparatus of claim 7, wherein, When the fusion unit fuses the lip video and the limb action video based on the lip position information to obtain a virtual person interactive video for outputting the to-be-broadcast voice, it is specifically configured to: For each action video frame, obtain the lip position data of the action video frame from the lip position information, and superimpose the lip video frame corresponding to the action video frame in the position indicated by the lip position data in the action video frame to obtain the interactive video frame corresponding to the action video frame; wherein a plurality of continuous interactive video frames constitute the virtual person interactive video.

9. The apparatus of claim 6, wherein, When the acquisition unit selects the limb action data matching the to-be-broadcast voice in the limb action video library as target limb action data, it is specifically configured to: Determine a target limb action matching the voice content of the to-be-broadcast voice; Select the limb action data corresponding to the target limb action in the limb action video library as target limb action data.

10. The apparatus of claim 9, wherein, The limb action data comprises an expression label, and the expression label represents a virtual expression when the limb action video is rendered. When the acquisition unit selects limb action data corresponding to the target limb action in the limb action video library as target limb action data, the acquisition unit is specifically configured to: Select limb action data corresponding to the target limb action and having the same expression label as the to-be-broadcast voice in the limb action video library as the target limb action data, wherein the expression label of the to-be-broadcast voice is determined according to the voice content of the to-be-broadcast voice.

Citation Information

Patent Citations

  • Virtual image video generation method and device, electronic equipment and readable storage medium

    CN114866807A

  • Virtual character video generation method and device, computer equipment and storage medium

    CN114998489A

  • Virtual portrait video generation method and device

    CN115052197A