A digital human generation and interaction method, device and medium
By constructing a general digital human avatar and reference motion video library, and combining personalized motion video generation with voice-driven technology, the time-consuming and labor-intensive problems of digital human avatar generation and motion driving have been solved, achieving high-quality, real-time digital human interaction.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2026-01-20
- Publication Date
- 2026-04-10
AI Technical Summary
Existing digital human technologies suffer from time-consuming, labor-intensive, and costly image generation and motion-driven capabilities. Furthermore, the lack of semantic relevance in motion performance makes it difficult to achieve rapid deployment and high-quality personalized interaction.
Construct a general digital human image and reference action video library, generate a personalized action video library through image and voice wake-up information, and realize digital human interaction by combining real-time synchronization of voice, action and subtitles.
It enables rapid and high-quality generation of personalized digital human motion videos, enhancing the realism, smoothness, and real-time responsiveness of digital human interaction.
Smart Images

Figure CN121547664B_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application generally relates to the technical field of digital human, and particularly relates to a digital human generation and interaction method, device and medium. BACKGROUND
[0002] With the rapid development of metaverse, virtual reality and human-computer interaction technology, digital human, as a key carrier connecting the real world and digital space, has shown broad application prospects in many fields such as education, finance, customer service, entertainment and media. An ideal digital human system should be able to quickly generate realistic virtual images and understand user intentions for natural, smooth and expressive real-time interaction.
[0003] However, the existing digital human technology system still has bottlenecks. First, in terms of digital human image, existing solutions mostly rely on professional artists to manually model and bind bones for specific characters, which is time-consuming, labor-intensive and costly. Or rely on real models for recording, which requires high recording environment and is time-consuming and labor-intensive. These limitations have seriously restricted the large-scale production and rapid deployment of digital humans. Second, in terms of motion generation and driving, existing technologies either rely on predefined fixed animation libraries for random playback, resulting in a lack of semantic relevance in motion performance; or generate motions based on voice through algorithms, but it is difficult to balance the naturalness, diversity and real-time performance of the motion. SUMMARY
[0004] In view of the above-mentioned defects or shortcomings in the prior art, it is desirable to provide a digital human generation and interaction method, device and medium with strong universality.
[0005] In a first aspect, the present application provides a digital human generation and interaction method, comprising the following steps:
[0006] constructing a universal digital human image and a reference motion video library matched therewith; the reference motion video library at least includes a plurality of reference motion video segments, and the reference motion video segments are video segments of a plurality of motions displayed by the universal digital human image;
[0007] generating an individualized motion video library corresponding to a target digital human image based on an image of the target digital human image provided by a user and the reference motion video segments; the individualized motion video library at least includes a plurality of target motion video segments, and the target motion video segments are video segments of the plurality of motions displayed by the target digital human image;
[0008] generating a corresponding reply text and a reply voice corresponding to the reply text in response to voice wake-up information issued by a user;
[0009] generating a plurality of reply motion video segments based on the reply voice and the individualized motion video library;
[0010] The reply voice, the reply action video segment and the reply text are time-aligned, and the reply text is bound as subtitles to obtain a synchronous digital human interaction video with real-time synchronization of voice, action and subtitles.
[0011] According to the technical scheme provided in the present application, a general digital human image and a reference action video library matched therewith are constructed, specifically including the following steps:
[0012] A digital human template prompt word is obtained, and the digital human template prompt word is input into a text-to-image model to obtain a digital human template image;
[0013] An action behavior description prompt word is obtained, and the digital human template image and the action behavior description prompt word are input into a video generation model to obtain a general digital human image and a reference action video library matched therewith.
[0014] According to the technical scheme provided in the present application, based on an image of a target digital human image provided by a user and the reference action video segment, a personalized action video library corresponding to the target digital human image is generated, specifically including the following steps:
[0015] The image of the target digital human image is input into a graph-to-graph large model for processing to generate a front full-body image of the target digital human; the front full-body image of the target digital human includes a complete front body shape of the target digital human;
[0016] Based on action migration technology, the action exhibited by the general digital human image in the reference action video library is migrated to the target digital human image to generate a plurality of target action video segments, thereby obtaining a personalized action video library corresponding to the target digital human image.
[0017] According to the technical scheme provided in the present application, based on action migration technology, the action exhibited by the general digital human image in the reference action video library is migrated to the target digital human image to generate a plurality of target action video segments, specifically including the following steps:
[0018] The action exhibited by the general digital human image in the reference action video library is replaced by the target digital human image to generate an initial personalized action video segment;
[0019] Based on the image of the target digital human image, face replacement and frame processing are performed on the initial personalized action video segment to obtain a repaired personalized action video segment;
[0020] Frame interpolation processing is performed on the repaired personalized action video segment to generate a frame-interconnected personalized action video segment with consistent first and last frames;
[0021] Super-resolution reconstruction is performed on all the inter-connection personalized action video segments to obtain a target action video segment.
[0022] According to the technical scheme provided in the present application, in response to the voice wake-up information issued by the user, the corresponding reply text and the reply voice corresponding to the reply text are generated, which specifically includes the following steps:
[0023] The voice wake-up information issued by the user is received, and the voice recognition model is activated after recognizing the voice wake-up word carried by the voice wake-up information;
[0024] The voice interaction information issued by the user is converted into text through the voice recognition model;
[0025] The text is input into the large language model to output the corresponding reply text, and at the same time, the reply voice corresponding to each sentence of the reply text is output.
[0026] According to the technical scheme provided in the present application, based on the reply voice and the personalized action video library, a plurality of reply action video segments are generated, which specifically includes the following steps:
[0027] A target action video segment matched with the reply voice is obtained from the personalized action video library as an interactive action video segment;
[0028] The reply voice is cut into a plurality of audio segments, and the audio segments are matched and bound with the video frames of the interactive action video segment to obtain the bound audio segments and video frames;
[0029] The bound audio segments and video frames are input into a lip synchronization model to generate corresponding lip action frames;
[0030] The lip action frames are fused with the interactive action video segment to obtain a reply action video segment.
[0031] According to the technical scheme provided in the present application, the method further includes the following steps:
[0032] The reply action video segment is played at a preset video frame rate, and at the same time, the reply action video segment to be played is cached to a buffer queue according to a preset rule.
[0033] According to the technical scheme provided in the present application, the reply action video segment to be played is cached to a buffer queue according to a preset rule, which specifically includes the following steps:
[0034] S1, determining whether the frame number of the reply action video segment to be played temporarily stored in the buffer queue exceeds a preset buffer threshold value; if the frame number of the reply action video segment to be played temporarily stored in the buffer queue does not exceed the preset buffer threshold value, performing a write operation of storing the reply action video segment to be played into the buffer queue;
[0035] S2, if the frame number of the reply action video segment to be played temporarily stored in the buffer queue exceeds the preset buffer threshold value, pausing the write operation and the reply action video segment synthesis process and entering an interval waiting stage;
[0036] S3, after the interval waiting stage ends, S1-S2 are re-executed until the user exits.
[0037] In a second aspect, the present application provides an electronic device, comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor implements the steps of the digital human generation and interaction method as described above when executing the computer program.
[0038] In a third aspect, the present application provides a computer readable storage medium, which stores a computer program, wherein the computer program is executed by a processor to implement the steps of the digital human generation and interaction method as described above.
[0039] From the above technical solution, the present application has at least the following beneficial effects:
[0040] The present application provides a digital human generation and interaction method, which comprises constructing a general digital human image and a reference action video library matched therewith; the reference action video library at least comprises a plurality of reference action video segments, and the reference action video segment is a video segment of a plurality of actions displayed by the general digital human image; based on an image of a target digital human image provided by a user and a reference action video segment, a personalized action video library corresponding to the target digital human image is generated; the personalized action video library at least comprises a plurality of target action video segments, and the target action video segment is a video segment of a plurality of actions displayed by the target digital human image; in response to voice wake-up information issued by the user, a corresponding reply text and a reply voice corresponding to the reply text are generated; based on the reply voice and the personalized action video library, a plurality of reply action video segments are generated; the reply voice, the reply action video segment, and the reply text are time-aligned, and then the reply text is bound as a subtitle to obtain a synchronous digital human interaction video in which voice, action, and subtitle are real-time synchronized.
[0041] The application lays a reusable basic resource for personalized digital people generation by constructing a general digital person image and a matched reference action video library, and then combines the image and the reference action video segment of the target digital person image provided by the user to migrate the action in the reference action video library to the target digital person image to generate a personalized action video library, thereby solving the problems of great customization difficulty and high cost of traditional personalized digital people through the general action resource, realizing fast and high-quality personalized digital people action video generation, and in the digital people interaction stage, through voice driving and semantic understanding and combining action scheduling and multi-modal time alignment technology, ensuring that the most matched action video segment is selected from the pre-made high-quality reference action video library as the target action video segment, and the voice, lip shape, body movement and subtitles are synchronously fused in real time and smoothly, significantly enhancing the realism, smoothness and real-time response capability of the digital people interaction. BRIEF DESCRIPTION OF DRAWINGS
[0042] Other features, objects and advantages of the application will become more apparent from the following detailed description of non-limiting embodiments made with reference to the accompanying drawings.
[0043] Figure 1 Flowchart of the method for generating and interacting with digital people.
[0044] Figure 2 Example diagram generated for reply text and action label.
[0045] Figure 3 Example diagram for pre-processing of reply voice.
[0046] Figure 4 Example diagram for audio segment corresponding video frame and lip shape synchronization.
[0047] Figure 5 Example diagram for video frame synthesis.
[0048] Figure 6 Example diagram for audio and video playing.
[0049] Figure 7 Example diagram of general digital person image.
[0050] Figure 8 Example diagram of waving action of general digital person image.
[0051] Figure 9 Example diagram of face image in the image of target digital person image.
[0052] Figure 10 Example diagram of front full-body image of target digital person.
[0053] Figure 11 Example diagram of waving action of target digital person image.
[0054] Figure 12 Example diagram of a face region extracted before face replacement.
[0055] Figure 13 Example diagram of a face region after face replacement.
[0056] Figure 14 Example diagram of a first frame of an action video segment with consistent first and last frames.
[0057] Figure 15 Example diagram of a hand waving action corresponding to an action video segment with consistent first and last frames.
[0058] Figure 16 Example diagram of a last frame of an action video segment with consistent first and last frames.
[0059] Figure 17 Example diagram of a structure of an electronic device.
[0060] Figure 18 Example diagram of digital human generation and interaction.
[0061] Figure 19 Example diagram of digital human template and reference action video production.
[0062] Figure 20 Example diagram of a specific digital human image production.
[0063] Figure 21 Example diagram of a specific digital human application.
[0064] Figure 22 Example diagram of digital human reply text and action tag generation.
[0065] Figure 23 Example diagram of speech preprocessing for reply.
[0066] Figure 24 Example diagram of audio segment corresponding video frame and lip shape synchronization.
[0067] Figure 25 Example diagram of video frame synthesis.
[0068] Figure 26 Example diagram of audio and video playback.
[0069] Reference numerals in the figure: 500, electronic device; 501, CPU; 502, ROM; 503, RAM; 504, bus; 505, I / O interface; 506, input part; 507, output part; 508, storage part; 509, communication part; 510, drive; 511, detachable medium. DETAILED DESCRIPTION
[0070] The application will be described in further detail below with reference to the drawings and embodiments. It can be understood that the specific embodiments described herein are only used to explain the related application, and not to limit the application. In addition, it should be noted that only parts related to the application are shown in the drawings for ease of description.
[0071] It should be noted that the embodiments in the present application and the features in the embodiments can be combined with each other without conflict. The present application will be described in detail below with reference to the drawings and embodiments.
[0072] As shown in the drawings, Figure 1 The present application provides a digital human generation and interaction method, comprising the following steps:
[0073] S100, constructing a general digital human image and a reference action video library matched therewith; the reference action video library at least includes a plurality of reference action video segments, and the reference action video segments are video segments of a plurality of actions displayed by the general digital human image.
[0074] Here, the general digital human image is an image benchmark template generated by a prompt word driving and a text-to-image model.
[0075] Further, constructing the general digital human image and the reference action video library matched therewith specifically comprises the following steps:
[0076] Obtaining a digital human template prompt word, and inputting the digital human template prompt word into a text-to-image model to obtain a digital human template image;
[0077] Obtaining an action behavior description prompt word, and inputting the digital human template image and the action behavior description prompt word into a video generation model to obtain the general digital human image and the reference action video library matched therewith.
[0078] It should be noted that the digital human template prompt word includes basic identity features, appearance detail features, clothing style features, background and picture quality features, and posture features. The basic identity features refer to basic attributes such as age range, gender, and regional style of the digital human. The appearance detail features refer to facial expressions, skin texture, hair style, etc. The clothing style features refer to clothing types, colors, styles, and need to consider universality and scene adaptability. The background and picture quality features refer to background styles avoiding complex background interference of the subject; at the same time, the picture quality requirements are clear, providing a high-quality image basis for subsequent action video generation. The posture features refer to the basic standing posture. The text-to-image model is, for example, Stable Diffusion, MidJourney, Qianwen Image, Flux2, Nano Banana, Hunyuan Image, etc.
[0079] Exemplarily, the digital human template prompt word is a 5-30 year old Asian female digital human, with a kind and gentle face, a natural smile, clear and flawless skin texture, brown medium-length hair with slight curls, soft hair quality, wearing a simple black business suit, standing naturally with arms hanging down, and the background is pure white, super-realistic style, 8K resolution, photo-level texture, clear and sharp picture, and bright eyes.
[0080] As Figure 7 shown, the designed digital human template prompt word is input into the text-to-image model, and the text-to-image model generates a corresponding digital human template image based on the text description of the digital human template prompt word through a deep learning algorithm. After generation, preliminary verification is required. First, it is confirmed whether the digital human template image completely conforms to the prompt word description and whether the image is free of obvious defects. If there is a deviation, the prompt word details are adjusted, and then the digital human template image is regenerated until a digital human template image that meets the generalization and high-quality requirements is obtained. This image serves as the fixed main character for all subsequent action video generation, ensuring that all reference actions are generated based on the same image, achieving strong binding of actions and images.
[0081] The action behavior description prompt word includes default actions and specific actions. The default action is the basic dynamic action without specific interaction instructions, which is used to avoid the static stiffness of the digital human image. For example, the prompt word corresponding to natural standing breathing is that the chest slightly rises and falls with breathing in a natural standing state, with a frequency of 18-20 times per minute, the arms naturally hang down without extra movements, and the head does not have obvious shaking. The prompt word corresponding to slight nodding is that the head slowly nods down with the cervical spine as the axis, with an amplitude of about 10-15 degrees, and returns to normal after 0.5 seconds, with smooth and uninterrupted movements.
[0082] The specific action is a functional action that meets the daily interaction scenario. For example, the prompt word corresponding to right hand waving to greet is that from the natural standing posture, the right hand slowly rises with the shoulder joint as the axis, the elbow bending angle is about 120 degrees, the palm naturally opens, the fingers are close together, the wrist drives the palm to swing along the horizontal direction to the outside of the body, the swing amplitude is 15-20 degrees, a total of 2-3 swings, each swing lasts for 0.8-1 second, the left arm remains naturally hanging down, the head slightly turns to the waving direction (amplitude does not exceed 10 degrees), and the right hand slowly falls back to the initial position after the action is completed.
[0083] As Figure 8As shown, the arrow in the figure is the motion trajectory of the hand, and the action behavior description prompt word corresponding to the figure is "right hand waving hello". Starting from the default standing posture, the right hand is slowly raised with the shoulder joint as the axis, the elbow is slightly bent (the bending angle is about 120 degrees), the palm is naturally opened (the five fingers are close together, and the fingertips are forward), the wrist drives the palm to make a light swing in the horizontal direction to the outside (away from the body side), the swing amplitude is controlled within 15-20 degrees, and the swing times are 2-3 times (each swing lasts for 0.8-1 second), and during the action process, the left arm remains in the default state of falling, and the head can be accompanied by a slight turning in the waving direction (the turning amplitude is not more than 10 degrees), and after completion, the right hand slowly falls back to the default position.
[0084] Here, the video generation model is, for example, Veo3, Wan, and Meta Video.
[0085] The digital human template image and the single action behavior description prompt word are input as joint input data into the video generation model. The video generation model first identifies the subject in the digital human template image, which includes the contour, skeletal structure, and initial posture. Then, based on the description content of the action behavior description prompt word, the video generation model adds a dynamic frame sequence that meets the action details to the subject through a time sequence generation algorithm. Finally, a continuous video is synthesized, and the video length is, for example, 3-5 seconds. This process needs to generate a corresponding reference video for each action to ensure that each action has an independent video carrier.
[0086] According to the above generation process, after generating all the reference videos of single actions, first, the action integrity and fluency of each reference video are checked by manual or automatic tools. If there is a problem, the prompt word of the corresponding action needs to be adjusted, and the video generation model needs to be re-input to generate until all videos meet the standard. For example, if the waving action does not complete the falling back, the prompt word of the corresponding action needs to be adjusted, and the content of the prompt word needs to be supplemented with the action that the right hand needs to completely fall back to the natural falling position after the action is completed. Second, all the reference videos that pass the audit are unified into the same format, resolution, and frame rate to avoid compatibility problems in subsequent personalized action migration due to format differences. The format is, for example, MP4 format, the resolution is, for example, 1920x1080 pixels, and the frame rate is, for example, 30fps. Finally, the videos are classified according to the action type, a clear directory structure is established to form the final reference action video library matching the general digital human image, as shown in Table 1, and each video in the library has a unified digital human template image as the subject and can be directly reused.
[0087] Table 1 Reference action video library
[0088]
[0089] S101, based on the image of the target digital human figure provided by the user and the reference action video segment, generate a personalized action video library corresponding to the target digital human figure; the personalized action video library at least includes a plurality of target action video segments, and the target action video segment is a video segment for the target digital human figure to demonstrate a plurality of actions.
[0090] Among them, based on the image of the target digital human figure provided by the user and the reference action video segment, generate a personalized action video library corresponding to the target digital human figure, specifically including the following steps:
[0091] The image of the target digital human figure is input into the image-to-image model for processing to generate a front full-body image of the target digital human; the front full-body image of the target digital human includes the complete front body shape of the target digital human;
[0092] Based on the action migration technology, the actions demonstrated by the general digital human figure in the reference action video library are migrated to the target digital human figure to generate a plurality of target action video segments, and a personalized action video library corresponding to the target digital human figure is obtained.
[0093] It should be noted that the image of the target digital human figure is, for example, Figure 9 As shown in the figure, the image of the target digital human figure (the core is the front face) is input into the image-to-image model, and the image-to-image model is, for example, Qwen-Image-Edit, NanoBanana image-to-image model, Flux Context, and can be image-to-image model. The image-to-image model is a commonly used model in the art, and its operation process is prior art, which will not be repeated here. The generated front full-body image of the target digital human is as shown in Figure 10 Specifically, the front face of the image of the target digital human figure is input into the image-to-image model, and the image-to-image model demonstrates the facial features, skin color, hairstyle, etc. based on the information of the front face, and completes the complete body shape details, and the body shape is from the complete body shape template generated by the image-to-image model. Finally, the front full-body image of the target digital human is obtained, which contains the complete front body shape of the target digital human. Here, the body shape details include trunk proportion, limb posture, basic clothing style adaptation, etc.
[0094] Further, based on the action migration technology, the actions demonstrated by the general digital human figure in the reference action video library are migrated to the target digital human figure to generate a plurality of target action video segments, specifically including the following steps:
[0095] Replace the actions demonstrated by the general digital human figure in the reference action video library with the target digital human figure to generate an initial personalized action video segment;
[0096] Based on the image of the target digital human figure, the initial personalized action video segment is replaced and processed to obtain a repaired personalized action video segment;
[0097] frame interpolation processing is performed on the personalized action video segment after repair to generate an inter-frame connected personalized action video segment with consistent first and last frames;
[0098] All the inter-frame connected personalized action video segments are subjected to super-resolution reconstruction to obtain a target action video segment.
[0099] The action migration technology can adopt a deep learning-based action cloning / migration model or use a graph-to-video model using ControlNet technology, such as Mimicmotion, Wan2.1 Vace, Wan2.2-Animate, etc. Specifically, the deep learning-based action cloning / migration model first analyzes the reference action video frame by frame, extracts the skeletal key points of the general digital human image in each frame, such as shoulder joints, elbow joints, hip joints, knee joints, etc. through a pose estimation algorithm (such as OpenPose); secondly, the skeletal key points of the target digital human are labeled, and a skeletal framework matching the target digital human is established according to the body size ratio of the target digital human; then, the skeletal motion trajectory of the reference action is mapped to the skeletal framework of the target digital human in proportion, for example, the right hand in the reference action is lifted by 30 degrees, the model will calculate the corresponding actual lifting angle and motion distance according to the shoulder width and arm length of the target digital human, to ensure that the action amplitude matches the body size of the target digital human, and to avoid small target digital humans making too large actions and large target digital humans making too small actions; finally, based on the mapped skeletal trajectory, a frame-by-frame dynamic picture is generated for the target digital human full-body image, which keeps the head features (facial features, hairstyle), body size, and clothing unchanged, and makes the torso and limbs move according to the mapped trajectory, i.e. the initial personalized action video segment, as shown in Figure 11 As an example of waving hands, the action logic in the video library is completely consistent with the reference action video library, but the execution subject changes from the general digital human image to the target digital human image, realizing the preliminary combination of action reuse and image customization.
[0100] Here, a face swapping model can be used to replace the face in the initial personalized action video segment, such as inswapper, blendswap, simswap, Uniface, etc. Specifically, as shown in Figure 12 A complete facial region is extracted from the front full-body image of the target digital human image provided by the user as a facial reference template, including forehead, eyebrows, eyes, nose, mouth, and chin, without extra background, to ensure that the details of the facial features, skin color, and light and shadow effects of the template are completely consistent with the target digital human image. Then, the initial personalized action video segment is analyzed frame by frame, and the facial region in each frame is located through a face detection algorithm, as shown in Figure 13As shown, the face detection algorithm is used to align the face reference template with the face region in the video frame of the initial personalized action video segment, and then the two faces are input into the face swapping model to output a face picture whose face features are the same as the face reference template and whose facial expressions, shadows, etc. are the same as those in the video frame. Finally, the face picture output by the face swapping model is pasted back to the original video frame, and a mask fusion algorithm is used to ensure a natural edge transition. Here, the face detection algorithm is, for example, YOLO-Face, RetinaFace, etc. After the face swapping process is completed, the blurred skin texture, eye details, and hairstyle edges are repaired, optimized, and corrected by the high-definition face repair model to ensure that the facial details of each frame are clear and real, consistent with the target digital human image, i.e., to obtain the repaired personalized action video segment. Here, the high-definition face repair model is, for example, GFPGAN, CodeFormer.
[0101] As shown, the face detection algorithm is used to align the face reference template with the face region in the video frame of the initial personalized action video segment, and then the two faces are input into the face swapping model to output a face picture whose face features are the same as the face reference template and whose facial expressions, shadows, etc. are the same as those in the video frame. Finally, the face picture output by the face swapping model is pasted back to the original video frame, and a mask fusion algorithm is used to ensure a natural edge transition. Here, the face detection algorithm is, for example, YOLO-Face, RetinaFace, etc. After the face swapping process is completed, the blurred skin texture, eye details, and hairstyle edges are repaired, optimized, and corrected by the high-definition face repair model to ensure that the facial details of each frame are clear and real, consistent with the target digital human image, i.e., to obtain the repaired personalized action video segment. Here, the high-definition face repair model is, for example, GFPGAN, CodeFormer. Figure 14 and Figure 16 As shown, the face detection algorithm is used to align the face reference template with the face region in the video frame of the initial personalized action video segment, and then the two faces are input into the face swapping model to output a face picture whose face features are the same as the face reference template and whose facial expressions, shadows, etc. are the same as those in the video frame. Finally, the face picture output by the face swapping model is pasted back to the original video frame, and a mask fusion algorithm is used to ensure a natural edge transition. Here, the face detection algorithm is, for example, YOLO-Face, RetinaFace, etc. After the face swapping process is completed, the blurred skin texture, eye details, and hairstyle edges are repaired, optimized, and corrected by the high-definition face repair model to ensure that the facial details of each frame are clear and real, consistent with the target digital human image, i.e., to obtain the repaired personalized action video segment. Here, the high-definition face repair model is, for example, GFPGAN, CodeFormer. Figure 15 shows the state of a certain intermediate frame of the waving action.
[0102] As shown, the face detection algorithm is used to align the face reference template with the face region in the video frame of the initial personalized action video segment, and then the two faces are input into the face swapping model to output a face picture whose face features are the same as the face reference template and whose facial expressions, shadows, etc. are the same as those in the video frame. Finally, the face picture output by the face swapping model is pasted back to the original video frame, and a mask fusion algorithm is used to ensure a natural edge transition. Here, the face detection algorithm is, for example, YOLO-Face, RetinaFace, etc. After the face swapping process is completed, the blurred skin texture, eye details, and hairstyle edges are repaired, optimized, and corrected by the high-definition face repair model to ensure that the facial details of each frame are clear and real, consistent with the target digital human image, i.e., to obtain the repaired personalized action video segment. Here, the high-definition face repair model is, for example, GFPGAN, CodeFormer.
[0103] S102, in response to the voice wake-up information issued by the user, generating a corresponding reply text and a reply voice corresponding to the reply text.
[0104] Here, in response to the voice wake-up information issued by the user, the corresponding reply text and the reply voice corresponding to the reply text are generated, which specifically includes the following steps:
[0105] Receiving the voice wake-up information issued by the user, and activating the speech recognition model after recognizing the voice wake-up word carried by the voice wake-up information;
[0106] The speech recognition model converts the user's voice interaction information into text.
[0107] Input the text into the large language model and output the corresponding reply text. At the same time, output the reply voice corresponding to each reply text.
[0108] It's important to note that speech recognition models such as SensVoice, Whisper, and SeACoParaformer, when they recognize a wake-up word in the wake-up message, are activated and switch to working mode. They then recognize the user's voice interaction and convert it into text. This text is then input into a Large Language Model (LLM). The LLM first analyzes the core intent of the text, while also recognizing its emotional tone and interaction scenario, laying the semantic foundation for subsequent responses and action label selection. Based on this semantic understanding, it generates a natural language response that matches the scenario and the user's intent, ensuring the response fully addresses the need and guides subsequent interactions. According to the semantics, tone, and interaction purpose of the response text, it selects the most matching label from a predefined list of action labels as the action instruction. For example, if the response text is a greeting, such as "Hello!", it matches a wave; if it's confirmation / acknowledgment, it matches a slight nod; if it emphasizes key information, it matches emphasizing with both hands; if the response text has no clear emotional or action tendency, it matches natural standing and breathing.
[0109] For example, such as Figure 2 As shown, the voice wake-up word is "hello." After being recognized by the speech recognition model, the model is activated. The speech recognition model then converts the voice interaction information into text and inputs it into the large language model, generating a matching digital human response and action type. The response text is then preprocessed, including word segmentation and sentence breaking, while matching corresponding predefined action tags. Each text segment, along with its corresponding predefined action tag, is stored in the large model's response queue for subsequent text-to-speech synthesis, until the entire response is complete.
[0110] Here, if the large language model is streaming, the response text preprocessing process does not need to wait for the large model to finish responding. Thus, while the large language model is responding, the already generated response text can be simultaneously converted into speech to reduce response latency.
[0111] While the large language model outputs the response text, a text-to-speech (TTS) synthesis task is initiated in parallel to convert the response text into corresponding response speech. The speech synthesis process involves obtaining text content from the large model's response queue, inputting it into the TTS model to generate response speech with natural intonation and rhythm.
[0112] S103, generating a plurality of reply action video segments based on the reply voice and the personalized action video library, specifically comprising the following steps:
[0113] obtaining a target action video segment matching the reply voice from the personalized action video library as the interactive action video segment;
[0114] cutting the reply voice into a plurality of audio segments, and binding the audio segments and the video frames of the interactive action video segment to obtain the bound audio segments and video frames;
[0115] inputting the bound audio segments and video frames into a lip synchronization model to generate corresponding lip action frames;
[0116] fusing the lip action frames and the interactive action video segment to obtain the reply action video segment.
[0117] It should be noted that, as shown in Figure 3 the action instruction corresponding to the reply voice is obtained, a target action video segment consistent with the action instruction is searched in the personalized action video library, and is taken as the interactive action video segment matching the reply voice; then, the reply voice is cut into a plurality of audio segments according to the time length of the voice, and then one video frame corresponding to each audio segment is matched in the interactive action video segment, so that the audio and video at each time are synchronized. The lip synchronization model is, for example, Wav2Lip, LipGAN, MuseTalk, etc. As shown in Figure 4 the bound audio segments and video frames are processed by the lip synchronization model to obtain corresponding lip action frames. The lip action frames are superimposed on the transparent face area of the body frame according to the position of the face bounding box, ensuring that the size and angle of the lip action frames completely match the body frame, avoiding proportion imbalance or angle misplacement of the lips and the body; the fused video frame is adjusted in quality to ensure that the resolution, brightness and contrast of the face area and the body area are consistent, avoiding visual fragmentation due to quality differences.
[0118] Moreover, after the bound audio segments and video frames are input into the lip synchronization model to generate corresponding lip action frames, if the current audio segment is a silent segment, the lip synchronization processing is skipped, and the corresponding original face frame in the target action video segment is directly used as the lip action frame, so as to reduce the calculation power consumption in the lip synchronization process. Here, the silent segment refers to generating a corresponding silent audio segment of the same length when judging that the audio segment information queue is empty.
[0119] S104, time sequence alignment of the reply voice, the reply action video segment and the reply text, and binding the reply text as subtitles to obtain a synchronous digital human interactive video with real-time synchronization of voice, action and subtitles.
[0120] The process involves aligning the timestamps of the response voice, response action video clips, and response text, using the response text as subtitles to obtain a synchronized digital human interactive video, and storing the synchronized digital human interactive video in a buffer queue for playback.
[0121] like Figure 5 As shown, the buffer queue is a temporary storage container based on memory or disk. It is designed with corresponding storage units for three types of resources: response voice, response action video segments, and response text, ensuring that resources can be stored in chronological order and retrieved in chronological order during playback. The synchronized digital human interactive video has synchronized video frames, each corresponding to a response voice, response action video segment, and response text. The buffer queue is shown in Table 2, where the queue index increases sequentially according to the resource generation order.
[0122] Table 2 Buffer Queue
[0123]
[0124] like Figure 6 As shown, the generated response video segments are batch-written into a buffer queue in ascending order of timestamp. When a response voice read from the buffer queue is marked as the first frame, its corresponding response text is immediately displayed as a subtitle on the playback interface. When a response voice is marked as the last frame, the current subtitle is stopped after the segment is played. This mechanism ensures that the subtitle content is strictly synchronized with the start and end times of the response voice, and is aligned with the lip movements and actions in time, thereby enhancing the immersiveness and comprehensibility of the interaction.
[0125] This method manages asynchronous tasks and data streams during the interaction process through a multi-queue collaboration mechanism. Specifically, it receives user voice streams in real time, performs wake-word detection and speech recognition; stores response text and corresponding action commands generated by a large language model; receives response text and asynchronously generates corresponding response speech; schedules and preloads interactive action video segments from a personalized action video library based on action commands; and receives and temporarily stores response speech, response action video segments, and response text that have undergone lip-syncing and frame fusion, sorted by timestamps, for sequential reading by the playback module. The queues collaborate through event-driven mechanisms and state machines to ensure a smooth, low-latency end-to-end connection from voice input to multimodal video output.
[0126] Furthermore, the method also includes the following steps:
[0127] The response video segment is played at a preset video frame rate, and the response video segment to be played is cached in the buffer queue according to preset rules.
[0128] Here, the preset video frame rate is, for example, 25 frames per second.
[0129] The reply action video segment to be played is cached into the buffer queue according to a preset rule, and the method comprises the following steps:
[0130] S1, determining whether the number of frames of the reply action video segment to be played temporarily stored in the buffer queue exceeds a preset buffer threshold value; if the number of frames of the reply action video segment to be played temporarily stored in the buffer queue does not exceed the preset buffer threshold value, performing a write operation of storing the reply action video segment to be played into the buffer queue;
[0131] S2, if the number of frames of the reply action video segment to be played temporarily stored in the buffer queue exceeds the preset buffer threshold value, pausing the write operation and the reply action video segment synthesis process and entering an interval waiting stage;
[0132] S3, after the interval waiting stage ends, re-executing S1-S2 until the user exits.
[0133] Here, the preset buffer threshold value can be set according to actual needs.
[0134] The above-mentioned playing and writing logic is mainly that the write operation is dynamically adjusted according to the load state of the buffer queue, and the write operation is normally performed when it is not overloaded, and the write operation is paused when it is overloaded, so as to avoid queue congestion through interval waiting and guarantee the smoothness of interaction, that is, if the number of frames of the reply action video segment to be played temporarily stored in the buffer queue does not exceed the preset buffer threshold value, it is indicated that the number of frames of the reply action video segment to be played can be added to the tail of the buffer queue, and if the number of frames of the reply action video segment to be played temporarily stored in the buffer queue exceeds the preset buffer threshold value, it is indicated that the current buffer queue is overloaded, and the number of frames of the reply action video segment to be played cannot be directly added to the buffer queue, at this time, the interval waiting stage is entered, that is, the write operation is performed after waiting for a waiting interval time.
[0135] The calculation formula of the waiting interval time is:
[0136] ;
[0137] In the formula, is the waiting interval time, is the remaining number of frames of the reply action video segment to be played temporarily stored in the buffer queue, is the preset buffer threshold value, is the number of frames per second of the currently played reply action video segment.
[0138] In order to facilitate understanding, the digital human generation and interaction method in the present application is briefly introduced as follows.
[0139] As Figure 18As shown, steps S11-S13 show that by prompting word-driven text-to-image, video generation model, making digital human template and reference action video segment, forming a general digital human image and reference action video library, laying a reusable foundation for personalized customization; then taking the image of the target digital human image provided by the user as input, through text-to-image, action migration, face optimization and other steps, generating a specific digital human image and exclusive personalized action video library; finally, the specific digital human is applied in the corresponding business scenario, that is, through voice wake-up trigger interaction, text generation, speech synthesis, action scheduling, lip synchronization and multi-modal synthesis, outputting the digital human interaction video with real-time synchronization of voice, action and subtitles, realizing the end-to-end closed loop from image generation to natural interaction.
[0140] As shown in Figure 19 Steps S14-S17 show that the digital human template prompt words containing basic identity, appearance, clothing, background and posture features are designed, which are input into the text-to-image model to generate digital human template images that meet the general requirements. After generation, it needs to be verified whether it completely matches the prompt word description. If there is deviation, adjust the prompt word and regenerate; design action behavior description prompt words for default actions (such as natural standing and breathing) and specific actions (such as right hand waving and greeting); input the digital human template image and the single action behavior description prompt word into the video generation model to generate independent reference videos for each action, with a video duration of 3-5 seconds; manually or automatically review the generated reference videos to determine the completeness and smoothness of the action, and adjust the prompt words for unqualified videos and regenerate them; unify the format, resolution and frame rate of all reference videos that pass the review, and classify and organize them by action type to form a standardized reference action video library.
[0141] As shown in Figure 20 Steps S18-S23 show that the front face of the target digital human provided by the user is obtained (the face features are required to be complete, unobstructed and front angle), which is input into the text-to-image large model. The model completes the body details such as the torso and limbs based on the face features to generate a front full-body image of the target digital human containing complete front body shape; input the front full-body image of the target digital human and the reference action video library into the action migration model, which extracts the skeletal motion trajectory and timing parameters of the reference action, maps the trajectory according to the body shape proportion of the target digital human, and generates an initial personalized action video; cut the face area in the target person's face image as a reference template, and perform face replacement on the initial personalized action video frame by frame, and then optimize the skin texture, eye details, etc. through the face high-definition repair model to obtain the repaired personalized action video; input the repaired video into the frame interpolation model to generate a video with consistent posture at the beginning and end, ensuring smooth action switching; eliminate video noise and artifacts through the super-resolution reconstruction model to optimize the image quality, and finally form a high-quality specific digital human action video library.
[0142] As shown in Figure 21 S24-S32 show the start of the digital human reply text and action label generation process, convert user voice to text and generate corresponding action instructions, and push to the large model reply queue; then enter the reply voice preprocessing process, synthesize the reply text into voice and cut it into attribute audio segments, and store it in the audio segment information queue; then through the audio segment corresponding video frame and lip shape synchronization process, bind the audio segment and the action video frame, generate the lip shape aligned face frame, and push it to the audio-face video frame queue; then through the video frame synthesis process, fuse the face frame and the body action video to generate the reply action video segment and store it in the buffer queue; finally, through the playback process, retrieve resources from the buffer queue, synchronize audio and video playback and display subtitles, and complete human-computer interaction.
[0143] As shown in Figure 22 S33-S37 show that the user issues voice wake-up information, and the voice wake-up model returns the answer after recognizing the wake-up word, and activates the voice recognition model; the voice recognition model receives the user's subsequent voice interaction information and converts it into text; the converted text is input into the large language model, the model analyzes the text intent and scene, generates natural language reply text, and matches the predefined action label corresponding to the reply semantics and tone (such as the "wave" label matched by the greeting reply); During generation, the end punctuation of the sentence is detected in real time, and the content segment before the punctuation is intercepted, and the remaining reply is continuously generated until the entire reply is completed; the intercepted reply text segment and the corresponding action label are associated one by one, and pushed to the large model reply queue in order, waiting for subsequent processing.
[0144] As shown in Figure 23 S38-S43 show that the system cyclically detects whether the large model reply queue is empty, and if it is empty, it waits for a specified time and then detects again; if it is not empty, it goes to the next step; take out a reply content and the corresponding action label from the large model reply queue; input the reply content into the speech synthesis model, the model performs preprocessing such as word segmentation, sentence segmentation, and prosody marking on the text, and generates a reply voice with natural intonation and rhythm; cut the reply voice into continuous fixed-length audio segments according to the video frame rate; finally, add attribute information to each audio segment, including whether the segment is the first or last frame of audio, the corresponding reply content, and the associated action label, and then push it to the audio segment information queue in order.
[0145] As shown in Figure 24As shown, steps S44-S58 show whether the remaining time length of the audio and video in the audio and video buffer queue is greater than the buffer playing threshold, if so, the waiting interval is dynamically calculated according to the queue, and after waiting, it is re-detected; if not, the next step is entered; whether the audio segment information queue is empty, if so, a mute audio segment equal to the audio segment time length is generated as a processing object; if not, an audio segment and attribute information are taken out from the queue; whether the audio segment is the first frame of audio, if so, it is further detected whether the number of unprocessed frames of the current action video is less than the action switching threshold, if so, the action label in the attribute information is set as a to-be-played action; whether the current action video is processed, if so, whether there is a to-be-played action, if so, the corresponding action video is set as the current processing video, if not, the default action video is enabled; the first unprocessed face part is taken out from the current processing action video, and is bound to the audio segment-video frame pair being processed; the binding pair is input into the lip synchronization model to generate a face video frame with aligned lips and audio; the audio segment, attribute information and generated face video frame are pushed to the audio-face video frame queue.
[0146] As shown in Figure 25 steps S59-S63, the system continuously detects whether the audio-face video frame queue is empty, if empty, it waits for a specified time length and re-detects; if not empty, the next step is entered; the audio segment, corresponding attribute information and face video frame with aligned lips are taken out from the queue; the face video frame is superimposed and fused with the body part in the high-quality specific digital person specified action video according to the face bounding box position, ensuring the size, angle and quality of the face and body consistent, avoiding visual fragmentation; a high-quality specific digital person specified action video frame with synchronized voice and lip shape is generated; finally, the synthesized synchronous video frame, corresponding audio segment and attribute information are pushed to the audio and video buffer queue, preparing for subsequent playing.
[0147] As shown in Figure 26 steps S64-S71, the system detects whether the audio and video buffer queue is empty, if empty, it waits for a specified time length and re-detects; if not empty, the next step is entered; the synchronous video frame, audio segment and corresponding attribute information are taken out from the audio and video buffer queue; whether the audio segment is the first frame of audio, if so, the reply content in the attribute information is taken as the subtitle, and the display on the playing interface is started immediately; if not, the current subtitle display state is maintained; whether the audio segment is the last frame of audio, if so, the current subtitle display is stopped after the segment is played; if not, it continues to display; the taken audio segment and high-quality specific digital person specified action video frame are played synchronously on the playing interface, ensuring the time sequence of voice, picture and subtitle consistent, completing the interactive output.
[0148] This application provides an electronic device, including a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor executes the computer program to implement the steps of the digital human generation and interaction method as described in the above embodiments.
[0149] Among them, such as Figure 17 As shown, the electronic device 500 includes a CPU 501, which can perform various appropriate actions and processes according to a program stored in ROM 502 or a program loaded from storage section 508 into RAM 503.
[0150] RAM 503 also stores various programs and data required for system operation. CPU 501, ROM 502, and RAM 503 are interconnected via bus 504. I / O interface 505 is also connected to bus 504.
[0151] The following components are connected to I / O interface 505: an input section 506 including a keyboard, mouse, etc.; an output section 507 including a cathode ray tube (CRT), liquid crystal display (LCD), etc., and speakers, etc.; a storage section 508 including a hard disk, etc.; and a communication section 509 including a network interface card such as a LAN card, modem, etc. The communication section 509 performs communication processing via a network such as the Internet. A drive 510 is also connected to I / O interface 505 as needed. A removable medium 511, such as a disk, optical disk, magneto-optical disk, semiconductor memory, etc., is installed on drive 510 as needed so that computer programs read from it can be installed into storage section 508 as needed.
[0152] Specifically, according to embodiments of this application, the above reference flow Figure 1 The described process can be implemented as a computer software program.
[0153] For example, this application includes a computer program product comprising a computer program carried on a computer-readable medium, the computer program containing program code for performing the methods shown in the flowchart. In such an embodiment, the computer program can be downloaded and installed from a network via communication section 509, and / or installed from removable medium 511. When the computer program is executed by CPU 501, it performs the functions defined in the system of this application.
[0154] It should be noted that the computer-readable medium can be a computer-readable signal medium or a computer-readable storage medium or any combination thereof. The computer-readable storage medium can be, for example, but not limited to, an electronic, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any suitable combination of the foregoing. More specific examples of a computer-readable storage medium can include, but are not limited to, the following: an electrical connection having one or more wires, a portable computer diskette, a hard disk, a RAM (Random Access Memory), a ROM (Read-Only Memory), an erasable programmable ROM (EPROM or Flash memory), an optical fiber, a portable compact disc read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the foregoing.
[0155] In this application, a computer readable storage medium can be any tangible medium that can retain, store, or maintain the program for use by or in connection with an instruction execution system, apparatus, or device. In this application, a computer readable signal medium can include a computer readable program code, propagated by any means, including but not limited to wireless, wire line, optical fiber, RF, etc. The computer readable medium discussed in this application can be transitory or non-transitory computer readable medium.
[0156] The flow diagrams and the block diagrams in the drawings are illustrations of the possible architectures, functions, and operations of systems, methods, and computer program products according to various embodiments of the present application. In this regard, each block in the flow diagrams or block diagrams can represent a module, segment, or portion of code, which comprises one or more executable instructions for implementing the specified logical functions. It should also be noted that in some alternative implementations, the functions noted in the blocks can occur out of the order noted in the figures. For example, two blocks shown in succession may, in fact, be executed substantially concurrently or the blocks may sometimes be executed in the reverse order, depending upon the functionality involved.
[0157] It should also be noted that in some alternative implementations, the functions noted in the blocks can occur out of the order noted in the figures. For example, two blocks shown in succession may, in fact, be executed substantially concurrently or the blocks may sometimes be executed in the reverse order, depending upon the functionality involved. It will also be noted that each block of the block diagrams and / or flowchart illustrations, and combinations thereof, can be implemented by special purpose hardware-based systems that perform the specified functions or operations, or combinations of special purpose hardware and computer instructions.
[0158] The units described in the embodiments of the present application can be implemented in the form of software, or can be implemented in the form of hardware, and the described units can also be arranged in a processor. In some cases, the names of the units do not constitute a limitation on the units themselves. The described units or modules can also be arranged in a processor.
[0159] The present application also provides a computer readable storage medium, which can be included in the electronic device described in the above embodiments, or can exist independently without being assembled into the electronic device. The above computer readable medium carries one or more programs, which, when executed by the electronic device, enable the electronic device to implement the digital human generation and interaction method provided in the above embodiments.
[0160] The above description is merely preferred embodiments of the present application and a description of the principles of the technology used. Those skilled in the art should understand that the scope of the application disclosed in the present application is not limited to the technical solutions formed by the specific combinations of the above technical features, and also covers other technical solutions formed by any combination of the above technical features or equivalent features without departing from the inventive concept. For example, the above features are replaced with each other to form technical solutions with similar functions disclosed in the present application (but not limited to).
Claims
1. A method for generating and interacting with digital humans, characterized in that, Includes the following steps: Construct a universal digital human avatar and a reference motion video library that matches it; The reference action video library includes at least a number of reference action video segments, which are video segments of multiple actions displayed by the general digital human image; Based on the image of the target digital human provided by the user and the reference action video segment, a personalized action video library corresponding to the target digital human is generated; the personalized action video library includes at least a number of target action video segments, which are video segments in which the target digital human displays the multiple actions. In response to a user's voice wake-up message, a corresponding reply text and a corresponding reply voice are generated; Based on the response voice and the personalized action video library, multiple response action video segments are generated; The response voice, the response action video segment, and the response text are time-aligned, and then the response text is bound as a subtitle to obtain a synchronized digital human interactive video with real-time synchronization of voice, action, and subtitle; Constructing a universal digital human avatar and a matching reference motion video library includes the following steps: Obtain digital human template prompts and input the digital human template prompts into the text-based image model to obtain digital human template images; Obtain action behavior description prompts, and input the digital human template image and the action behavior description prompts into the video generation model to obtain a general digital human image and a reference action video library that matches it; Based on the image of the target digital human provided by the user and the reference action video clips, a personalized action video library corresponding to the target digital human is generated, specifically including the following steps: The image of the target digital human is input into the image-generated image-large model for processing to generate a full-body frontal image of the target digital human; the full-body frontal image of the target digital human includes the complete frontal body shape of the target digital human. Based on motion transfer technology, the motion displayed by the general digital human image in the reference motion video library is transferred to the target digital human image to generate multiple target motion video segments and obtain a personalized motion video library corresponding to the target digital human image. Based on motion transfer technology, the motion displayed by the general digital human image in the reference motion video library is transferred to the target digital human image to generate multiple target motion video segments, specifically including the following steps: Replace the actions displayed by the general digital human image in the reference action video library with the target digital human image to generate an initial personalized action video segment; Based on the image of the target digital human, the initial personalized action video segment is subjected to face replacement and frame processing to obtain the repaired personalized action video segment. The repaired personalized action video segment is subjected to frame interpolation processing to generate a frame-connected personalized action video segment with consistent first and last frames. Super-resolution reconstruction is performed on all the inter-frame stitched personalized motion video segments to obtain the target motion video segment.
2. The digital human generation and interaction method according to claim 1, characterized in that, In response to a user's voice wake-up message, a corresponding reply text and a corresponding reply voice are generated, specifically including the following steps: Receive voice wake-up information from the user, and activate the voice recognition model after recognizing the voice wake-up word carried in the voice wake-up information; The speech recognition model converts the user's voice interaction information into text. The text is input into the large language model, and the corresponding reply text is output. At the same time, the reply voice corresponding to each reply text is output.
3. The digital human generation and interaction method according to claim 1, characterized in that, Based on the response voice and the personalized action video library, multiple response action video segments are generated, specifically including the following steps: Obtain the target action video segment that matches the reply voice from the personalized action video library as the interactive action video segment; The response voice is cut into multiple audio segments, and the audio segments are matched and bound to the video frames of the interactive action video segment to obtain the bound audio segments and video frames; The bound audio clips and video frames are input into the lip-sync model to generate corresponding lip motion frames; The lip movement frame is fused with the interactive action video segment to obtain the response action video segment.
4. The digital human generation and interaction method according to claim 3, characterized in that, The method further includes the following steps: The response action video segment is played at a preset video frame rate, and the response action video segment to be played is cached in a buffer queue according to a preset rule.
5. The digital human generation and interaction method according to claim 4, characterized in that, The video segment of the response action to be played is cached into the buffer queue according to preset rules, specifically including the following steps: S1. Determine whether the number of frames of the response action video segment to be played temporarily stored in the buffer queue exceeds the preset buffer threshold; if the number of frames of the response action video segment to be played temporarily stored in the buffer queue does not exceed the preset buffer threshold, then perform a write operation to store the response action video segment to be played into the buffer queue. S2. If the number of frames of the response action video segment to be played temporarily stored in the buffer queue exceeds the preset buffer threshold, the writing operation and the response action video segment synthesis process are paused and the interval waiting stage is entered. S3. After the interval waiting phase ends, S1-S2 are executed again until the user exits.
6. An electronic device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that, When the processor executes the computer program, it implements the steps of the digital human generation and interaction method as described in any one of claims 1-5.
7. A computer-readable storage medium storing a computer program, characterized in that, When the computer program is executed by a processor, it implements the steps of the digital human generation and interaction method as described in any one of claims 1-5.
Citation Information
Patent Citations
Virtual image video generation method and device, electronic equipment and storage medium
CN116645455A
Digital human video generation method and device, equipment, medium and product
CN120751201A