Whole-body digital human video generation method and system based on graph retrieval
Through the full-body digital human video generation method based on graph retrieval, the problem of digital human lack of body movement generation in the prior art is solved, and gesture action video generation matching audio is realized, which improves the immersion of digital humans in the display effect.
Patent Information
- Application Number
- CN202510225818.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-27
- Publication Date
- 2025-06-06
AI Technical Summary
In the prior art, the digital human solution is only aimed at the generation of facial expressions and lip shapes, and lacks the generation of body movements, resulting in the generated digital human effect being relatively weak in display and cannot bring an immersive viewing experience to the audience.
Through the full-body digital human video generation method based on the graph search, audio data is obtained and audio sequence vectorization is performed, the database of gesture actions is retrieved to match the vector module of audio data, and the audio data is embedded in the digital human video using the video frame sequence, thereby generating gesture action video matching the audio data.
It realizes adding gestures to audio matching for digital people, making them more vivid and realistic in news broadcasts and other scenarios, bringing users a more friendly and immersive viewing and reading experience.
Smart Images

Figure CN120104834A_ABST
Abstract
Description
Technical Field
[0001] The invention relates to a method and system for generating full-body digital human video based on graph retrieval. Background Art
[0002] The current multimodal video digital human mainly generates facial portraits and lip synchronization. For example, wav2lip focuses on achieving accurate lip synchronization, so that the lip movements of the characters in the generated video are highly matched with the speech content in the audio; EchoMimic generates highly realistic audio-visual synchronized portrait videos through audio and facial landmarks. These digital human solutions only focus on the generation of facial expressions and lip shapes, and lack the generation of body movements. The digital humans generated based on this are relatively weak in display effect and cannot provide an immersive viewing experience for the audience. Summary of the invention
[0003] The technical problem to be solved by the present invention is to overcome the defect that the digital human solutions in the prior art only focus on the generation of facial expressions and lip shapes, lack the generation of body movements, and the digital humans generated based on this have a relatively weak display effect. A method and system for generating full-body digital human videos based on graph retrieval is provided for adding gestures matching audio to multi-model video digital humans, making them more vivid and realistic in news broadcasts. Compared with digital human videos without this function, the method and system can provide users with a more friendly and immersive viewing and reading experience.
[0004] This application solves the above technical problems through the following technical solutions:
[0005] A method for generating a full-body digital human video based on graph retrieval, wherein the method comprises:
[0006] Get audio data;
[0007] Performing audio sequence vectorization on the audio data to obtain a vector module of the audio data;
[0008] For each vector module, a video frame sequence of a gesture action matching the vector module is retrieved from a database of gesture actions, wherein the database includes a video map of a full-body digital human;
[0009] The audio data is embedded into the digital human video using the video frame sequence to generate a full-body digital human video matching the audio data.
[0010] Preferably, the method for generating a full-body digital human video comprises:
[0011] Get the role video material;
[0012] Predict the action of each frame in the video material and generate nodes of the topology graph;
[0013] Adding topological graph edges to nodes based on the similarity of the character’s gestures between frames;
[0014] A video frame sequence of a character's gesture action is generated according to the nodes of the topology graph and the edges of the topology graph, and the database includes the video frame sequence.
[0015] Preferably, the method for generating a full-body digital human video comprises:
[0016] Predicting audio and lip shape using the full-body digital human video;
[0017] The prediction result is used to generate a full-body digital human video with lip movements.
[0018] Preferably, the step of generating a full-body digital human video with lip movements using the prediction result comprises:
[0019] Get lip shape video material;
[0020] A node that predicts the motion of each frame in the video material and generates a lip graph;
[0021] Adding topological graph edges to nodes based on the similarity of lip movements between frames;
[0022] Generate a lip-shaped video frame sequence according to the nodes of the lip-shaped graph and the edges of the topological graph;
[0023] The lip-shaped video frame sequence and the vector module are used to embed the lip-shaped video into the digital human video to generate a full-body digital human video with the lip-shaped video.
[0024] Preferably, the step of generating a full-body digital human video with lip movements using the prediction result comprises:
[0025] The next node position of the lip is obtained by using the video frame sequence of the character's gesture action;
[0026] Obtain the intensity of the gesture action at the next node according to the position of the next node;
[0027] The edge of the optimal lip shape topology graph is matched according to the strength, and the lip position of the full-body digital human video is adjusted according to the next node position.
[0028] Preferably, obtaining the intensity of the gesture action at the next node according to the position of the next node includes:
[0029] Connect the node position and the image frame of the gesture action at the next node;
[0030] Get the pixel distance that the arm moves in the image under the target number in the image frame after connection;
[0031] The intensity is obtained according to the score of the pixel distance and the score of the specific action.
[0032] Preferably, the method for generating a full-body digital human video comprises:
[0033] Get the node pairings other than the edges of the graph in the video frame sequence;
[0034] Generate supplementary video material of node pairing by using the character video material and the gesture action of node pairing;
[0035] The supplementary video material is added to the video frame sequence using the edges of the topological graph of paired nodes.
[0036] The present invention also provides a full-body digital human video generation system based on graph retrieval, which is characterized in that the full-body digital human video generation system based on graph retrieval is used to implement the full-body digital human video generation method based on graph retrieval as described above.
[0037] On the basis of being in accordance with the common sense in the art, the above-mentioned preferred conditions can be arbitrarily combined to obtain the preferred embodiments of the present invention.
[0038] The positive and progressive effects of the present invention are:
[0039] The present invention adds gestures matching audio to multi-model video digital humans, making them appear more vivid and realistic in news broadcasts. Compared with digital human videos without this function, the present invention can provide users with a more friendly and immersive viewing and reading experience. BRIEF DESCRIPTION OF THE DRAWINGS
[0040] Figure 1 Schematic diagram of the structure of the full-body digital human video generation system according to Embodiment 1 of the present invention.
[0041] Figure 2 This is a flow chart of a method for generating a full-body digital human video according to Embodiment 1 of the present invention. DETAILED DESCRIPTION
[0042] The present invention is further described below by way of examples, but the present invention is not limited to the scope of the examples.
[0043] Example 1
[0044] This embodiment provides a full-body digital human video generation system based on graph retrieval.
[0045] The full-body digital human video generation system includes a processing terminal and a server. The processing terminal can be a desktop computer, a mobile phone or a tablet computer.
[0046] The server can be an independent physical server, or a server cluster or distributed system composed of multiple physical servers. It can also be a cloud server that provides basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communications, middleware services, domain name services, security services, CDN (Content Delivery Network), as well as big data and artificial intelligence platforms.
[0047] The full-body digital human video generation system is used for:
[0048] Get audio data;
[0049] Performing audio sequence vectorization on the audio data to obtain a vector module of the audio data;
[0050] For each vector module, a video frame sequence of a gesture action matching the vector module is retrieved from a database of gesture actions, wherein the database includes a video map of a full-body digital human;
[0051] The audio data is embedded into the digital human video using the video frame sequence to generate a full-body digital human video matching the audio data.
[0052] Specifically, the processing terminal is used to obtain audio data and upload it to the server.
[0053] The server is used to:
[0054] Get audio data;
[0055] Performing audio sequence vectorization on the audio data to obtain a vector module of the audio data;
[0056] For each vector module, a video frame sequence of a gesture action matching the vector module is retrieved from a database of gesture actions, wherein the database includes a video map of a full-body digital human;
[0057] The audio data is embedded into the digital human video using the video frame sequence to generate a full-body digital human video matching the audio data.
[0058] The video frame sequence includes nodes of a graph. In this embodiment, the nodes include multiple action nodes of a continuous action. In this embodiment, the nodes of the graph may be nodes in a topological graph.
[0059] The nodes of these graphs can be connected to each other, and the nodes with connection relationships mark the edges of the graphs.
[0060] Furthermore, the server is used for:
[0061] Get the role video material;
[0062] Predict the action of each frame in the video material and generate nodes of the topology graph;
[0063] Adding topological graph edges to nodes based on the similarity of the character’s gestures between frames;
[0064] A video frame sequence of a character's gesture action is generated according to the nodes of the topology graph and the edges of the topology graph, and the database includes the video frame sequence.
[0065] The server is used to: predict audio and lip shape using the full-body digital human video;
[0066] The prediction result is used to generate a full-body digital human video with lip movements.
[0067] The server is used to:
[0068] Get lip shape video material;
[0069] A node that predicts the lip movements of each frame in the video material and generates a lip shape graph;
[0070] Adding topological graph edges to nodes based on the similarity of lip movements between frames;
[0071] Generate a lip-shaped video frame sequence according to the nodes of the lip-shaped graph and the edges of the topological graph;
[0072] The lip-shaped video frame sequence and the vector module are used to embed the lip-shaped video into the digital human video to generate a full-body digital human video with the lip-shaped video.
[0073] The server is used to:
[0074] The next node position of the lip is obtained by using the video frame sequence of the character's gesture action;
[0075] Obtain the intensity of the gesture action at the next node according to the position of the next node;
[0076] The edge of the optimal lip shape topology graph is matched according to the strength, and the lip position of the full-body digital human video is adjusted according to the next node position.
[0077] The server is used to:
[0078] Connect the node position and the image frame of the gesture action at the next node;
[0079] Get the pixel distance that the arm moves in the image under the target number in the image frame after connection;
[0080] The intensity is obtained according to the score of the pixel distance and the score of the specific action.
[0081] The server is used to:
[0082] Get the node pairings other than the edges of the graph in the video frame sequence;
[0083] Generate supplementary video material of node pairing by using the character video material and the gesture action of node pairing;
[0084] The supplementary video material is added to the video frame sequence using the edges of the topological graph of paired nodes.
[0085] See also Figure 1 ,The server includes a gesture action graph structure generation module, a joint embedding module, a gesture action retrieval module, and an audio and lip synchronization module:
[0086] Gesture action graph structure generation module: Based on the input character video material, the action of each frame of the video is first predicted to generate the nodes of the graph. The prediction is to identify the posture, meaning, and topological graph nodes, and then add the edges of the graph based on the similarity between frames to finally generate a gesture action frame graph sequence.
[0087] Joint embedding module: vectorize the input audio sequence and jointly embed the video frames in the gesture action graph.
[0088] Gesture action retrieval module: retrieves gesture action graphs based on the vectorized audio sequence, generates a video frame sequence with gestures, and then synthesizes the video.
[0089] Audio and lip synchronization module: Predict audio and lip shape based on the video generated in the previous step, and finally generate a multimodal digital human video with gestures.
[0090] Using the above system, this embodiment also provides a method for generating a full-body digital human video based on graph retrieval, comprising:
[0091] Step 100: Obtain role video material;
[0092] Step 101: predict the action of each frame in the video material and generate nodes of a topological graph;
[0093] Step 102: adding edges of a topological graph to nodes based on the similarity of the gestures of the characters between frames;
[0094] Step 103: Generate a video frame sequence of the character's gesture action according to the nodes and edges of the topological graph, and the database includes the video frame sequence.
[0095] Step 104: Acquire audio data;
[0096] Step 105: vectorize the audio data into an audio sequence to obtain a vector module of the audio data;
[0097] Step 106: for each vector module, searching a database of gesture actions for a video frame sequence of gesture actions that matches the vector module, wherein the database includes a video image of a full-body digital human;
[0098] Step 107: embed the audio data into the digital human video using the video frame sequence to generate a full-body digital human video matching the audio data.
[0099] Specifically, the method for generating a full-body digital human video includes:
[0100] Predicting audio and lip shape using the full-body digital human video;
[0101] The prediction result is used to generate a full-body digital human video with lip movements.
[0102] The step of generating a full-body digital human video with lip movements using the prediction result includes:
[0103] Get lip shape video material;
[0104] A node that predicts the motion of each frame in the video material and generates a lip graph;
[0105] Adding topological graph edges to nodes based on the similarity of lip movements between frames;
[0106] Generate a lip-shaped video frame sequence according to the nodes of the lip-shaped graph and the edges of the topological graph;
[0107] The lip-shaped video frame sequence and the vector module are used to embed the lip-shaped video into the digital human video to generate a full-body digital human video with the lip-shaped video.
[0108] The step of generating a full-body digital human video with lip movements using the prediction result includes:
[0109] The next node position of the lip is obtained by using the video frame sequence of the character's gesture action;
[0110] Obtain the intensity of the gesture action at the next node according to the position of the next node;
[0111] The edge of the optimal lip shape topology graph is matched according to the strength, and the lip position of the full-body digital human video is adjusted according to the next node position.
[0112] Wherein, obtaining the intensity of the gesture action at the next node according to the position of the next node includes:
[0113] Connect the node position and the image frame of the gesture action at the next node;
[0114] Get the pixel distance that the arm moves in the image under the target number in the image frame after connection;
[0115] The intensity is obtained according to the score of the pixel distance and the score of the specific action.
[0116] Wherein, the method for generating a full-body digital human video comprises:
[0117] Get the node pairings other than the edges of the graph in the video frame sequence;
[0118] Generate supplementary video material of node pairing by using the character video material and the gesture action of node pairing;
[0119] The supplementary video material is added to the video frame sequence using the edges of the topological graph of paired nodes.
[0120] Although the specific embodiments of the present invention are described above, those skilled in the art should understand that these are only examples, and the protection scope of the present invention is defined by the appended claims. Those skilled in the art may make various changes or modifications to these embodiments without departing from the principles and essence of the present invention, but these changes and modifications all fall within the protection scope of the present invention.
Claims
1. A method for generating full-body digital human video based on graph retrieval, characterized in that: The method for generating a full-body digital human video comprises: Get audio data; Performing audio sequence vectorization on the audio data to obtain a vector module of the audio data; For each vector module, a video frame sequence of a gesture action matching the vector module is retrieved from a database of gesture actions, wherein the database includes a video map of a full-body digital human; The audio data is embedded into the digital human video using the video frame sequence to generate a full-body digital human video matching the audio data.
2. The method for generating full-body digital human video based on graph retrieval as claimed in claim 1, characterized in that: The method for generating a full-body digital human video comprises: Get the role video material; Predict the action of each frame in the video material and generate nodes of the topology graph; Adding topological graph edges to nodes based on the similarity of the character’s gestures between frames; A video frame sequence of a character's gesture action is generated according to the nodes of the topology graph and the edges of the topology graph, and the database includes the video frame sequence.
3. The method for generating full-body digital human video based on graph retrieval as claimed in claim 1, characterized in that: The method for generating a full-body digital human video comprises: Predicting audio and lip shape using the full-body digital human video; The prediction result is used to generate a full-body digital human video with lip movements.
4. The method for generating full-body digital human video based on graph retrieval as claimed in claim 3, characterized in that: The method of generating a full-body digital human video with lip movements using the prediction result includes: Get lip shape video material; A node that predicts the motion of each frame in the video material and generates a lip graph; Adding topological graph edges to nodes based on the similarity of lip movements between frames; Generate a lip-shaped video frame sequence according to the nodes of the lip-shaped graph and the edges of the topological graph; The lip-shaped video frame sequence and the vector module are used to embed the lip-shaped video into the digital human video to generate a full-body digital human video with the lip-shaped video.
5. The method for generating full-body digital human video based on graph retrieval as claimed in claim 4, characterized in that: The method of generating a full-body digital human video with lip movements using the prediction result includes: The next node position of the lip is obtained by using the video frame sequence of the character's gesture action; Obtain the intensity of the gesture action at the next node according to the position of the next node; The edge of the optimal lip shape topology graph is matched according to the strength, and the lip position of the full-body digital human video is adjusted according to the next node position.
6. The method for generating full-body digital human video based on graph retrieval as claimed in claim 5, characterized in that: The obtaining the intensity of the gesture action at the next node according to the position of the next node includes: Connect the node position and the image frame of the gesture action at the next node; Get the pixel distance that the arm moves in the image under the target number in the image frame after connection; The intensity is obtained according to the score of the pixel distance and the score of the specific action.
7. The method for generating full-body digital human video based on graph retrieval according to claim 2 or 4, characterized in that: The method for generating a full-body digital human video comprises: Get the node pairings other than the edges of the graph in the video frame sequence; Generate supplementary video material of node pairing by using the character video material and the gesture action of node pairing; The supplementary video material is added to the video frame sequence using the edges of the topological graph of paired nodes.
8. A full-body digital human video generation system based on graph retrieval, characterized in that: The full-body digital human video generation system based on graph retrieval is used to implement the full-body digital human video generation method based on graph retrieval as described in claims 1 to 7.