Intelligent dialogue method and system based on digital human
Through large-scale model generation reply, TTS model conversion and audio-driven lip technology, combined with video generation, streaming and display, the problem of insufficient implementation of the complete process of intelligent dialogue systems in the existing technology is solved, and natural and anthropomorphic dialogue interaction and high user experience are achieved.
Patent Information
- Application Number
- CN202510294687.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-13
- Publication Date
- 2025-06-27
AI Technical Summary
The prior art has shortcomings in integrating artificial intelligence, big models, TTS technology and audio-driven lip technology to build a complete intelligent dialogue system, especially in implementing a complete process from input to video output.
It provides an intelligent dialogue method based on digital people, through big model generation reply, TTS model conversion and audio-driven lip typing technology, combining video generation, streaming and display to realize natural and anthropomorphic dialogue interaction between users and digital people.
It realizes a more natural and anthropomorphic dialogue interaction, improves the user experience, and has high flexibility and scalability. It is suitable for a variety of customer service scenarios, supports image customization and sound customization, and improves the authenticity and immersion of users' interactions with digital people.
Smart Images

Figure CN120216646A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical fields of artificial intelligence and multimedia processing, and specifically provides an intelligent dialogue method and system based on a digital human. Background Art
[0002] With the development of artificial intelligence technology, especially the emergence of large models, the ability of machines to understand and generate human language has made a qualitative leap. At the same time, the progress of TTS technology and audio-driven lip-sync technology has also made it possible to create more realistic and vivid virtual characters.
[0003] However, there are still deficiencies in the prior art in integrating these advanced technologies, especially in building an intelligent dialogue system that can achieve a complete process from input reception to video output. Summary of the Invention
[0004] The present invention aims at the above deficiencies of the prior art and provides a practical intelligent dialogue method based on a digital human.
[0005] A further technical task of the present invention is to provide a reasonably designed, safe and applicable intelligent dialogue system based on a digital human.
[0006] The technical solution adopted by the present invention to solve its technical problems is as follows:
[0007] The intelligent dialogue method based on a digital human has the following steps:
[0008] S1. Input stage, the user inputs a question or instruction through text or voice;
[0009] S2. The processing stage includes large model generating a reply, TTS model conversion and audio-driven lip-sync;
[0010] S3. The output stage includes video generation, streaming and display.
[0011] Further, in step S2, in the large model generating a reply, the question or instruction input by the user is input into the large model to generate a corresponding reply text;
[0012] In the TTS model conversion, the generated reply text is input into the TTS model and converted into audio.
[0013] Further, in the audio-driven lip-sync, the generated audio is input into the audio-driven lip-sync model MuseTalk to generate the lip animation of the digital human.
[0014] Further, in step S3, in the video generation, the generated lip animation is combined with other actions and expressions of the digital human to generate a complete video;
[0015] Streaming is to push the generated video to the SRS server;
[0016] Display means that the user pulls the video stream through the client to watch the dialogue video of the digital human.
[0017] Furthermore, the pre-trained large model Baichuan14B is adopted in the large model to generate responses;
[0018] The advanced TTS model Cosyvoice is adopted in the TTS model conversion;
[0019] The MuseTalk model is used for audio-driven lip synchronization;
[0020] The SRS server is adopted in streaming to push the video stream;
[0021] In display, the user pulls the video stream through the client to achieve real-time display of the digital human dialogue video.
[0022] Based on the intelligent dialogue system of the digital human, first, in the input stage, the user inputs questions or instructions through text or voice;
[0023] Then, the processing stage includes large model response generation, TTS model conversion, and audio-driven lip synchronization;
[0024] Finally, the output stage includes video generation, streaming, and display.
[0025] Furthermore, in the large model to generate responses, the questions or instructions input by the user are input into the large model to generate corresponding response texts;
[0026] In the TTS model conversion, the generated response text is input into the TTS model to be converted into audio.
[0027] Furthermore, in the audio-driven lip synchronization, the generated audio is input into the MuseTalk model of audio-driven lip synchronization to generate the lip animation of the digital human.
[0028] Furthermore, in the video generation, the generated lip animation is combined with other actions and expressions of the digital human to generate a complete video;
[0029] Streaming is to push the generated video to the SRS server;
[0030] Display means that the user pulls the video stream through the client to watch the dialogue video of the digital human.
[0031] Furthermore, the pre-trained large model Baichuan14B is adopted in the large model to generate responses;
[0032] The advanced TTS model Cosyvoice is adopted in the TTS model conversion;
[0033] The audio-driven lip-sync uses the MuseTalk model;
[0034] During the streaming, an SRS server is used to push the video stream.
[0035] During the display, the video stream is pulled through the client to achieve the real-time display of the digital human dialogue video.
[0036] Compared with the prior art, the intelligent dialogue method and system based on digital humans of the present invention have the following outstanding beneficial effects:
[0037] The intelligent dialogue method based on digital humans of the present invention generates responses through a large model, combines the TTS model and the model for audio-driven lip-sync, realizes more natural and anthropomorphic dialogue interactions, and improves the user experience. At the same time, it has high flexibility and scalability, and is applicable to various customer service scenarios. And it supports image customization and voice customization to provide customized services for customers. Description of the Drawings
[0038] In order to more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the following will briefly introduce the drawings required for use in the description of the embodiments or the prior art. Obviously, the drawings in the following description are some embodiments of the present invention. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.
[0039] Attached Figure 1 is a schematic flowchart of an intelligent dialogue method based on digital humans. Detailed Embodiments
[0040] In order to enable those skilled in the art to better understand the solutions of the present invention, the following will further elaborate on the present invention in combination with specific embodiments. Obviously, the described embodiments are only some embodiments of the present invention, rather than all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts fall within the scope of protection of the present invention.
[0041] The following gives a best embodiment:
[0042] As Figure 1 shown, the intelligent dialogue method based on digital humans in this embodiment has the following steps:
[0043] S1. Input stage, the user inputs questions or instructions through text or voice;
[0044] S2. The processing stage includes generating responses by a large model, TTS model conversion, and audio-driven lip-sync;
[0045] Among them, in the large model to generate a response, the question or instruction input by the user is input into the large model to generate the corresponding response text;
[0046] In the TTS model conversion, the generated response text is input into the TTS model and converted into audio.
[0047] In the audio-driven lip sync, the generated audio is input into the MuseTalk model of the audio-driven lip sync to generate the lip sync animation of the digital human.
[0048] S3. The output stage includes video generation, streaming, and display;
[0049] In video generation, the generated lip sync animation is combined with other actions and expressions of the digital human to generate a complete video;
[0050] Streaming is to stream the generated video to the SRS server;
[0051] Display means that the user pulls the video stream through the client to watch the conversation video of the digital human.
[0052] In this embodiment, the pre-trained large model Baichuan14B is adopted in the large model to generate a response;
[0053] The advanced TTS model Cosyvoice is adopted in the TTS model conversion;
[0054] The MuseTalk model is adopted for audio-driven lip sync;
[0055] The SRS server is adopted for video stream pushing during streaming;
[0056] During display, the video stream is pulled through the client to realize the real-time display of the digital human conversation video.
[0057] Based on the above method, for the intelligent conversation system based on digital human in this embodiment, first, in the input stage, the user inputs a question or instruction through text or voice;
[0058] Then, the processing stage includes large model to generate a response, TTS model conversion, and audio-driven lip sync;
[0059] Finally, the output stage includes video generation, streaming, and display.
[0060] Among them, in the large model to generate a response, the question or instruction input by the user is input into the large model to generate the corresponding response text;
[0061] In the TTS model conversion, the generated response text is input into the TTS model and converted into audio.
[0062] In audio-driven lip-syncing, the generated audio is input into the MuseTalk model of audio-driven lip-syncing to generate the lip animation of the digital human.
[0063] In video generation, the generated lip animation is combined with other actions and expressions of the digital human to generate a complete video;
[0064] Streaming is to stream the generated video to the SRS server;
[0065] Display means that the user pulls the video stream through the client to watch the conversation video of the digital human.
[0066] In the generation of responses by the large model, the pre-trained large model Baichuan14B is adopted;
[0067] In the TTS model conversion, the advanced TTS model Cosyvoice is adopted;
[0068] MuseTalk model is adopted for audio-driven lip-syncing;
[0069] SRS server is adopted in streaming for the push of the video stream;
[0070] In display, the video stream is pulled through the client to realize the real-time display of the conversation video of the digital human.
[0071] The intelligent conversation method based on digital human of the present invention generates responses through a large model, combines the TTS model and the model of audio-driven lip-syncing, realizes more natural and anthropomorphic conversation interaction, and improves the user experience. At the same time, the system has high flexibility and scalability, and is applicable to various customer service scenarios. And it supports image customization and voice customization, providing customized services for customers. It improves the realism and immersion when the user interacts with the digital human, and is especially suitable for occasions that require efficient customer service and personalized experience.
[0072] The above specific embodiments are only specific cases of the present invention. The patent protection scope of the present invention includes but is not limited to the above specific embodiments. Any technical solution that conforms to the technical solutions recorded in the above specific embodiments of the present invention and any appropriate changes or substitutions made by those of ordinary skill in the art shall fall within the patent protection scope of the present invention.
[0073] Although the embodiments of the present invention have been shown and described, it will be understood by those of ordinary skill in the art that various changes, modifications, substitutions and variations can be made to these embodiments without departing from the principles and spirit of the present invention. The scope of the present invention is defined by the appended claims and their equivalents.
Claims
1. An intelligent dialogue method based on digital human, characterized in that: The steps are as follows: S1, input stage, the user inputs questions or instructions through text or voice; S2, the processing stage includes large model generation response, TTS model conversion and audio-driven lip shape; S3, the output stage includes video generation, streaming and display.
2. The intelligent dialogue method based on digital human according to claim 1, characterized in that: In step S2, when the big model generates a reply, the question or instruction input by the user is input into the big model to generate a corresponding reply text; In the TTS model conversion, the generated reply text is input into the TTS model and converted into audio.
3. The intelligent dialogue method based on digital human according to claim 2, characterized in that: The audio driven lip-syncing process inputs the generated audio into the audio driven lip-syncing model MuseTalk to generate the lip-syncing animation of the digital human.
4. The intelligent dialogue method based on digital human according to claim 3 is characterized in that: In step S3, the generated lip animation is combined with other movements and expressions of the digital human in the video generation to generate a complete video; Push streaming is to push the generated video to the SRS server; It shows that the user pulls the video stream through the client and watches the conversation video of the digital person.
5. The intelligent dialogue method based on digital human according to claim 4 is characterized in that: The pre-trained large model Baichuan14B is used in the large model generation response; The advanced TTS model Cosyvoice is used in TTS model conversion; The audio-driven lip-sync uses the MuseTalk model; The SRS server is used to push the video stream during streaming; During the display, the video stream is pulled through the client to realize the real-time display of the digital human conversation video.
6. The intelligent dialogue system based on digital human is characterized by: First, in the input stage, the user enters questions or instructions through text or voice; Then, the processing stage includes large model generation response, TTS model conversion and audio-driven lip shape; Finally, the output stage includes video generation, streaming, and display.
7. The intelligent dialogue system based on digital human according to claim 6 is characterized in that: When the big model generates a reply, the question or instruction entered by the user is input into the big model to generate the corresponding reply text; In the TTS model conversion, the generated reply text is input into the TTS model and converted into audio.
8. The intelligent dialogue system based on digital human according to claim 7 is characterized in that: The audio driven lip-syncing process inputs the generated audio into the audio driven lip-syncing model MuseTalk to generate the lip-syncing animation of the digital human.
9. The intelligent dialogue system based on digital human according to claim 8, characterized in that: In the video generation, the generated lip animation is combined with other movements and expressions of the digital human to generate a complete video; Push streaming is to push the generated video to the SRS server; It shows that the user pulls the video stream through the client and watches the conversation video of the digital person.
10. The intelligent dialogue system based on digital human according to claim 9, characterized in that: The pre-trained large model Baichuan14B is used in the large model generation response; The advanced TTS model Cosyvoice is used in TTS model conversion; The audio-driven lip-sync uses the MuseTalk model; The SRS server is used to push the video stream during streaming; During the display, the video stream is pulled through the client to realize the real-time display of the digital human conversation video.
Citation Information
Cited By
Automatic lecturer video generation method based on AI speech synthesis and animation driving
CN120897102A
Method for automatically generating lecturer video based on AI speech synthesis and animation driving
CN120897102B
AI ancient poetry multi-round spoken language dialogue method and device and electronic equipment
CN121681767A
Interaction method based on lightweight modular three-dimensional digital human
CN121861175A