Digital human voice interaction optimization method and system
By using pre-trained models and load balancing algorithms in the digital human voice interaction system, the problems of latency, context understanding and resource utilization efficiency are solved, and a more natural and smooth voice interaction experience is achieved.
Patent Information
- Application Number
- CN202510085183.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-20
- Publication Date
- 2025-05-06
AI Technical Summary
The existing digital human voice interaction system has delay problems, limited context understanding, insufficient emotional expression and inefficient resource utilization, resulting in poor user experience.
Real-time speech stream is obtained through the pre-trained speech recognition model, combined with the large language model and the speech synthesis model, dynamically adjust the recognition strategy and insert the tone words, and optimize resource utilization using the load balancing algorithm.
It reduces users' perception of delay, improves system response speed and dialogue comprehension capabilities, makes voice interaction more natural and smooth, and improves user experience.
Smart Images

Figure CN119943045A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of digital human technology, and in particular to a digital human voice interaction optimization method and system. Background Art
[0002] With the continuous advancement of digital human voice interaction technology, this technology has been widely used in multiple scenarios such as customer service, education, smart home, medical care, etc.
[0003] Existing digital human voice interaction systems usually include the following core modules: automatic speech recognition (ASR), natural language processing (NLP), and text-to-speech synthesis (TTS). The typical process of these systems is to convert the user's voice input into text, generate answers based on the text content, and convert the answer text into voice output. However, the existing technology still has the following shortcomings: Delay issue: Delays may occur during speech recognition and synthesis, affecting user experience; Limited context understanding: The ability to understand multiple rounds of dialogue is weak, and it is difficult to maintain a consistent dialogue state; Insufficient emotional expression: Lack of natural emotional expression makes the interaction seem mechanical; Inefficient resource utilization: Failure to fully utilize hardware resources such as multi-core CPUs, resulting in low processing efficiency; Duplicate calculation: There is no effective caching mechanism for processed voice or text segments, resulting in unnecessary duplicate calculations.
[0004] Therefore, there is an urgent need for a digital human voice interaction optimization method and system that can reduce the user's perception of delay, improve the system's response speed, make the voice interaction more natural and smooth, and improve the user experience. Summary of the invention
[0005] In order to solve the above technical problems, the present invention provides a digital human voice interaction optimization method and system, which can reduce the user's perception of delay, improve the system's response speed, make the voice interaction more natural and smooth, and improve the user experience.
[0006] The present invention provides a method for optimizing digital human voice interaction, comprising the following steps: S1. Obtain the user's real-time voice stream through a pre-trained voice recognition model to obtain a text recognition result of the real-time voice stream; S2, input the text recognition results into the pre-trained large language model to generate the answer text; S3, synthesizing the answer text into an answer speech stream through a pre-trained speech synthesis model; S4, judging whether it is necessary to add modal particles according to the recognition delay of the speech recognition model; if so, proceeding to S5, otherwise proceeding to S6; S5, selecting a target modal particle through a pre-trained context-aware model according to the context of the current conversation, and inserting the target modal particle into the front end of the answering voice stream, obtaining an updated answering voice stream, and entering S6; wherein the context of the current conversation includes the text recognition result of the real-time voice stream and the historical record of the system response; S6. Play the answering voice stream in real time through the audio output module to realize the voice interaction of the digital human.
[0007] Furthermore, in S1, the real-time voice stream of the user is obtained through the pre-trained voice recognition model, and the text recognition result of the real-time voice stream is obtained, including: S11, obtaining a real-time voice stream of a user, and dividing the real-time voice stream into a plurality of voice sub-segments; S12, inputting each speech sub-segment into a pre-trained speech recognition model to obtain a text recognition result corresponding to each speech sub-segment; S13, inputting the text recognition results corresponding to each speech sub-segment obtained into the pre-trained context-aware model, and dynamically adjusting the recognition strategy of subsequent speech sub-segments; S14. When all speech sub-segments are recognized, the text recognition results corresponding to all speech sub-segments are concatenated to obtain the text recognition results of the real-time speech stream.
[0008] Furthermore, in S2, the text recognition result is input into the pre-trained large language model to generate the answer text including: S21. Build a knowledge graph and obtain entity embedding vectors and relationship embedding vectors through the knowledge graph embedding algorithm; S22, inputting the text recognition result of the real-time voice stream into the pre-trained knowledge graph retrieval module, and retrieving the entity embedding vector and the relationship embedding vector that match the text recognition result of the real-time voice stream; S23. Input the text recognition results of the real-time voice stream, the matching entity embedding vector and the matching relationship embedding vector into the pre-trained large language model to generate a response text.
[0009] Furthermore, in S3, the answer text is synthesized into an answer voice stream through a pre-trained speech synthesis model, including: S31, dividing the answer text into a plurality of text sub-segments; S32, inputting each text sub-segment into a pre-trained speech synthesis model to obtain a speech waveform corresponding to each text sub-segment; S33, inputting the speech waveform corresponding to each obtained text sub-segment into a pre-trained deep learning model, and dynamically adjusting the synthesis strategy of subsequent text sub-segments; S34. When all text sub-segments are synthesized, the speech waveforms corresponding to all text sub-segments are concatenated and encoded to obtain a response speech stream.
[0010] Further, in S4, judging whether to add an interjection word according to the recognition delay of the speech recognition model includes: The system monitors the recognition delay of the speech recognition model in real time. If the recognition delay is greater than the preset delay threshold, an interjection needs to be added; otherwise, an interjection does not need to be added.
[0011] Furthermore, in S5, a target modal particle is selected through a pre-trained context-aware model according to the context of the current conversation, and the target modal particle is inserted into the front end of the answer voice stream, and the updated answer voice stream includes: S51, analyzing the context of the current conversation through a pre-trained context-aware model; wherein the context of the current conversation includes a text recognition result of the real-time voice stream and a historical record of system responses; S52, selecting a target modal particle from a predefined modal word library according to the context of the current conversation; S53, inputting the target modal particle and the answer speech stream into the pre-trained deep learning model to obtain the speech waveform of the target modal particle; S54, encode the speech waveform of the target modal particle and insert it into the front end of the answering speech stream to obtain an updated answering speech stream.
[0012] Furthermore, when the real-time voice stream of the user is acquired through the pre-trained voice recognition model, the text recognition result of the real-time voice stream is obtained, and the answer text is synthesized into the answer voice stream through the pre-trained voice synthesis model, it also includes: Acquire a real-time voice processing task; wherein the real-time voice processing task refers to a processing operation performed on each group of continuous frames in a voice stream, and the voice stream includes a real-time voice stream or an answer voice stream; The load balancing algorithm monitors the load of each CPU core in real time, and distributes the real-time voice processing tasks to each CPU core according to the load of each CPU core; Cache the processed voice and text segments to the local server.
[0013] The present invention also provides a digital human voice interaction optimization system, which is used to execute any one of the digital human voice interaction optimization methods described above, and the system includes the following modules: The speech recognition module is used to obtain the user's real-time speech stream through a pre-trained speech recognition model and obtain the text recognition result of the real-time speech stream; The answer generation module is connected to the speech recognition module and is used to input the text recognition results into the pre-trained large language model to generate the answer text; A speech stream synthesis module, connected to the answer generation module, is used to synthesize the answer text into an answer speech stream through a pre-trained speech synthesis model; The modal particle judgment module is connected to the speech recognition module and the speech stream synthesis module, and is used to judge whether a modal particle needs to be added according to the recognition delay of the speech recognition model; if so, it enters the modal particle insertion module, otherwise it enters the audio output module; The modal particle insertion module is connected to the modal particle judgment module, and is used to select the target modal particle through the pre-trained context-aware model according to the context of the current conversation, and insert the target modal particle into the front end of the answer voice stream to obtain the updated answer voice stream, and enter the audio output module; wherein the context of the current conversation includes the text recognition result of the real-time voice stream and the historical record of the system response; The audio output module is connected to the modal particle judgment module and the modal particle insertion module, and is used to play the answer voice stream in real time through the audio output module to realize the voice interaction of the digital human.
[0014] The embodiments of the present invention have the following technical effects: The present invention ensures low latency and efficient processing by adopting block processing and incremental update methods. It is suitable for application scenarios such as speaking and recognizing, improves response speed, and combines with knowledge graph embedding algorithm to enhance the ability of dialogue understanding and generation, making answers more accurate. At the same time, the recognition strategy of subsequent voice sub-segments is dynamically adjusted through the context-aware model to improve recognition accuracy and coherence; appropriate modal particles are intelligently selected and inserted according to the recognition delay of the voice recognition model, thereby improving the naturalness and fluency of the interaction; the load of the system is monitored in real time through the load balancing algorithm, and the task allocation strategy is dynamically adjusted to ensure that the system can still maintain an efficient response speed under high load conditions, effectively utilize all CPU cores, and cache processed data to reduce repeated calculations, thereby improving the response speed and resource utilization of the system; this solution can reduce the user's perception of delay, improve the response speed of the system, make voice interaction more natural and smooth, and enhance the user experience. BRIEF DESCRIPTION OF THE DRAWINGS
[0015] In order to more clearly illustrate the specific implementation methods of the present invention or the technical solutions in the prior art, the drawings required for use in the specific implementation methods or the description of the prior art will be briefly introduced below. Obviously, the drawings described below are some implementation methods of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying creative work.
[0016] Figure 1is a flow chart of a digital human voice interaction optimization method provided by an embodiment of the present invention; Figure 2 It is a structural diagram of a digital human voice interaction optimization system provided by an embodiment of the present invention. DETAILED DESCRIPTION
[0017] In order to make the purpose, technical solution and advantages of the present invention clearer, the technical solution of the present invention will be described clearly and completely below. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without creative work belong to the scope of protection of the present invention.
[0018] The present invention proposes a digital human voice interaction optimization method. Figure 1 is a flowchart of a digital human voice interaction optimization method provided by an embodiment of the present invention, see Figure 1 , including: S1. Obtain the user's real-time voice stream through a pre-trained voice recognition model to obtain the text recognition result of the real-time voice stream.
[0019] Specifically include: S11. Acquire a user's real-time voice stream, and divide the real-time voice stream into several voice sub-segments.
[0020] In some embodiments, the real-time voice stream may be segmented according to actual conditions. For example, the real-time voice stream is segmented into a number of voice sub-segments with a duration of 50 ms.
[0021] S12. Input each speech sub-segment into a pre-trained speech recognition model to obtain a text recognition result corresponding to each speech sub-segment.
[0022] In some embodiments, context analysis and response generation are immediately performed based on the text recognition results corresponding to the acquired speech sub-segments, that is, speech-to-text is directly converted based on the currently acquired data and the previously acquired data, and as new data is acquired, the processing results are gradually updated to achieve instant response.
[0023] S13, inputting the text recognition results corresponding to each speech sub-segment into the pre-trained context-aware model, and dynamically adjusting the recognition strategy of subsequent speech sub-segments.
[0024] S14. When all speech sub-segments are recognized, the text recognition results corresponding to all speech sub-segments are concatenated to obtain the text recognition results of the real-time speech stream.
[0025] S2. Input the text recognition results into the pre-trained large language model to generate the answer text.
[0026] Specifically include: S21. Build a knowledge graph and obtain entity embedding vectors and relationship embedding vectors through the knowledge graph embedding algorithm.
[0027] In some embodiments, a pre-built knowledge graph can be constructed by extracting entities and relationships from text data sources such as expertise in the target field, literature, historical data, etc. through a large model relying on semantic understanding.
[0028] In some embodiments, the knowledge graph embedding algorithm may adopt TransE, TransR or DistMult, etc. Through these methods, entities and relationships in the knowledge graph are represented as vectors, which facilitates subsequent vector calculation and reasoning.
[0029] S22. Input the text recognition result of the real-time voice stream into the pre-trained knowledge graph retrieval module to retrieve the entity embedding vector and the relationship embedding vector that match the text recognition result of the real-time voice stream.
[0030] In some embodiments, the knowledge graph retrieval module retrieves entity embedding vectors and relationship embedding vectors that match the text recognition results of the real-time voice stream. The purpose is to improve the relevance and accuracy of the answer by identifying the entities and relationships that are most relevant to the user query, thereby generating a more appropriate answer. When the user's expression method changes, as long as the text recognition results point to the same entity or relationship, the system can still recognize and respond correctly. This improves the flexibility of the system, enabling it to cope with a variety of different questioning styles. In multiple rounds of conversations, the system can infer the user's intentions based on previously mentioned entities, and then provide more relevant follow-up information to maintain the consistency and coherence of the conversation.
[0031] S23. Input the text recognition results of the real-time voice stream, the matching entity embedding vector and the matching relationship embedding vector into the pre-trained large language model to generate a response text.
[0032] In some embodiments, the retrieved knowledge graph information is integrated into the generation process, thereby enhancing the model's ability to understand and generate domain-specific knowledge and obtaining more accurate answer text.
[0033] S3. The answer text is synthesized into an answer speech stream through a pre-trained speech synthesis model.
[0034] Specifically include: S31. Segment the answer text into a number of text sub-segments.
[0035] In some embodiments, the answer text may be segmented according to actual conditions. For example, the answer text is segmented into a number of text sub-segments with a length of 10 characters.
[0036] S32: Input each text sub-segment into a pre-trained speech synthesis model to obtain a speech waveform corresponding to each text sub-segment.
[0037] S33. Input the speech waveform corresponding to each text sub-segment obtained into a pre-trained deep learning model, and dynamically adjust the synthesis strategy of subsequent text sub-segments.
[0038] In some embodiments, the deep learning model may adopt the Sovits V2 model.
[0039] S34. When all text sub-segments are synthesized, the speech waveforms corresponding to all text sub-segments are concatenated and encoded to obtain a response speech stream.
[0040] In some embodiments, when a real-time voice stream of a user is acquired through a pre-trained voice recognition model, a text recognition result of the real-time voice stream is obtained, and the answer text is synthesized into an answer voice stream through a pre-trained voice synthesis model, it also includes: Get real-time speech processing tasks.
[0041] The real-time voice processing task refers to the processing operation performed on each group of continuous frames in the voice stream, and the voice stream includes the real-time voice stream or the answer voice stream.
[0042] The load balancing algorithm monitors the load of each CPU core in real time, and distributes the real-time voice processing tasks to each CPU core according to the load of each CPU core.
[0043] In some embodiments, the load balancing algorithm includes polling, minimum number of connections, response time and other methods. The appropriate algorithm can be selected according to the actual situation. In this embodiment, the polling method is used to monitor the load of each CPU core in real time, and dynamically adjust the task allocation strategy to ensure load balancing of each CPU core.
[0044] Cache the processed voice and text segments to the local server.
[0045] S4. Determine whether it is necessary to add modal particles according to the recognition delay of the speech recognition model; if so, proceed to S5; otherwise, proceed to S6.
[0046] In some embodiments, the system monitors the recognition delay of the speech recognition model in real time. If the recognition delay is greater than a preset delay threshold, an interjection needs to be added, otherwise, no interjection needs to be added; wherein, the preset delay threshold can be set according to actual conditions, for example, 200ms.
[0047] S5. Select the target modal particle through the pre-trained context-aware model according to the context of the current conversation, and insert the target modal particle into the front end of the answer voice stream to obtain an updated answer voice stream, and enter S6.
[0048] The context of the current conversation includes the text recognition result of the real-time voice stream and the historical record of the system response. The historical record of the system response includes the response delay information of the system and the like.
[0049] Specifically include: S51. Analyze the context of the current conversation through a pre-trained context-aware model.
[0050] S52. Select a target modal particle from a predefined modal particle library according to the context of the current conversation.
[0051] In some embodiments, the modal particles include words such as "ah", "oh", "um", etc. Real-world use cases can be collected from various sources such as social media, literary works, and conversation records, and annotated and classified through natural language processing technology to build a predefined modal word library. According to the context of the current conversation, the most appropriate modal particle is selected from the predefined modal word library as the target modal particle.
[0052] S53. Input the target modal particle and the answer speech stream into the pre-trained deep learning model to obtain the speech waveform of the target modal particle.
[0053] S54, encode the speech waveform of the target modal particle and insert it into the front end of the answering speech stream to obtain an updated answering speech stream.
[0054] S6. Play the answering voice stream in real time through the audio output module to realize the voice interaction of the digital human.
[0055] The present invention ensures low latency and efficient processing by adopting block processing and incremental update methods. It is suitable for application scenarios such as speaking and recognizing, improves response speed, and combines with knowledge graph embedding algorithm to enhance the ability of dialogue understanding and generation, making answers more accurate. At the same time, the recognition strategy of subsequent voice sub-segments is dynamically adjusted through the context-aware model to improve recognition accuracy and coherence; appropriate modal particles are intelligently selected and inserted according to the recognition delay of the voice recognition model, thereby improving the naturalness and fluency of the interaction; the load of the system is monitored in real time through the load balancing algorithm, and the task allocation strategy is dynamically adjusted to ensure that the system can still maintain an efficient response speed under high load conditions, effectively utilize all CPU cores, and cache processed data to reduce repeated calculations, thereby improving the response speed and resource utilization of the system; this solution can reduce the user's perception of delay, improve the response speed of the system, make voice interaction more natural and smooth, and enhance the user experience.
[0056] Figure 2 is a structural diagram of a digital human voice interaction optimization system provided by an embodiment of the present invention, the system is used to execute a digital human voice interaction optimization method described in the above embodiment, such as Figure 2 As shown, the system includes the following modules: The speech recognition module is used to obtain the user's real-time speech stream through a pre-trained speech recognition model and obtain the text recognition result of the real-time speech stream; The answer generation module is connected to the speech recognition module and is used to input the text recognition results into the pre-trained large language model to generate the answer text; A speech stream synthesis module, connected to the answer generation module, is used to synthesize the answer text into an answer speech stream through a pre-trained speech synthesis model; The modal particle judgment module is connected to the speech recognition module and the speech stream synthesis module, and is used to judge whether a modal particle needs to be added according to the recognition delay of the speech recognition model; if so, it enters the modal particle insertion module, otherwise it enters the audio output module; The modal particle insertion module is connected to the modal particle judgment module, and is used to select the target modal particle through the pre-trained context-aware model according to the context of the current conversation, and insert the target modal particle into the front end of the answer voice stream to obtain the updated answer voice stream, and enter the audio output module; wherein the context of the current conversation includes the text recognition result of the real-time voice stream and the historical record of the system response; The audio output module is connected to the modal particle judgment module and the modal particle insertion module, and is used to play the answer voice stream in real time through the audio output module to realize the voice interaction of the digital human.
[0057] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit it. Although the present invention has been described in detail with reference to the aforementioned embodiments, those skilled in the art should understand that they can still modify the technical solutions described in the aforementioned embodiments, or replace some or all of the technical features therein by equivalents. However, these modifications or replacements do not deviate the essence of the corresponding technical solutions from the technical solutions of the embodiments of the present invention.
Claims
1. A method for optimizing digital human voice interaction, characterized in that: The steps include: S1. Obtain the user's real-time voice stream through a pre-trained voice recognition model to obtain a text recognition result of the real-time voice stream; S2, inputting the text recognition result into a pre-trained large language model to generate a response text; S3, synthesizing the answer text into an answer speech stream through a pre-trained speech synthesis model; S4, judging whether it is necessary to add an interjection according to the recognition delay of the speech recognition model; If yes, go to S5, otherwise go to S6; S5, selecting a target modal particle through a pre-trained context-aware model according to the context of the current conversation, and inserting the target modal particle into the front end of the answering voice stream to obtain an updated answering voice stream, and entering S6; wherein the context of the current conversation includes the text recognition result of the real-time voice stream and the historical record of the system response; S6. Play the answering voice stream in real time through the audio output module to realize the voice interaction of the digital human.
2. A digital human voice interaction optimization method according to claim 1, characterized in that: In S1, the real-time voice stream of the user is obtained through the pre-trained voice recognition model, and the text recognition result of the real-time voice stream is obtained, including: S11, obtaining a real-time voice stream of a user, and dividing the real-time voice stream into a plurality of voice sub-segments; S12, inputting each speech sub-segment into a pre-trained speech recognition model to obtain a text recognition result corresponding to each speech sub-segment; S13, inputting the text recognition results corresponding to each speech sub-segment obtained into the pre-trained context-aware model, and dynamically adjusting the recognition strategy of subsequent speech sub-segments; S14. When all speech sub-segments are recognized, the text recognition results corresponding to all speech sub-segments are concatenated to obtain the text recognition results of the real-time speech stream.
3. A digital human voice interaction optimization method according to claim 1, characterized in that: In S2, the text recognition result is input into a pre-trained large language model to generate a response text, including: S21. Build a knowledge graph and obtain entity embedding vectors and relationship embedding vectors through the knowledge graph embedding algorithm; S22, inputting the text recognition result of the real-time voice stream into a pre-trained knowledge graph retrieval module, and retrieving entity embedding vectors and relationship embedding vectors that match the text recognition result of the real-time voice stream; S23, inputting the text recognition result of the real-time voice stream, the matching entity embedding vector and the matching relationship embedding vector into a pre-trained large language model to generate a response text.
4. A digital human voice interaction optimization method according to claim 1, characterized in that: In S3, synthesizing the answer text into an answer voice stream through a pre-trained speech synthesis model includes: S31, dividing the answer text into a plurality of text sub-segments; S32, inputting each text sub-segment into a pre-trained speech synthesis model to obtain a speech waveform corresponding to each text sub-segment; S33, inputting the speech waveform corresponding to each obtained text sub-segment into a pre-trained deep learning model, and dynamically adjusting the synthesis strategy of subsequent text sub-segments; S34. When all text sub-segments are synthesized, the speech waveforms corresponding to all text sub-segments are concatenated and encoded to obtain a response speech stream.
5. The method for optimizing digital human voice interaction according to claim 1, characterized in that: In S4, judging whether it is necessary to add an interjection according to the recognition delay of the speech recognition model includes: The system monitors the recognition delay of the speech recognition model in real time. If the recognition delay is greater than a preset delay threshold, an interjection needs to be added; otherwise, an interjection does not need to be added.
6. A digital human voice interaction optimization method according to claim 1, characterized in that: In S5, a target modal particle is selected through a pre-trained context-aware model according to the context of the current conversation, and the target modal particle is inserted into the front end of the answer voice stream, and the updated answer voice stream is obtained, including: S51, analyzing the context of the current conversation through a pre-trained context-aware model; wherein the context of the current conversation includes a text recognition result of the real-time voice stream and a historical record of system response; S52, selecting a target modal particle from a predefined modal word library according to the context of the current conversation; S53, inputting the target modal particle and the answer speech stream into a pre-trained deep learning model to obtain a speech waveform of the target modal particle; S54, encoding the speech waveform of the target modal particle and inserting it into the front end of the answering speech stream to obtain an updated answering speech stream.
7. A digital human voice interaction optimization method according to claim 1, characterized in that: When the user's real-time voice stream is acquired through a pre-trained voice recognition model, a text recognition result of the real-time voice stream is obtained, and the answer text is synthesized into an answer voice stream through a pre-trained voice synthesis model, it also includes: Acquire a real-time voice processing task; wherein the real-time voice processing task refers to a processing operation performed on each group of continuous frames in a voice stream, wherein the voice stream includes the real-time voice stream or the answer voice stream; The load balancing algorithm monitors the load of each CPU core in real time, and distributes the real-time voice processing tasks to each CPU core according to the load of each CPU core; Cache the processed voice and text segments to the local server.
8. A digital human voice interaction optimization system, used to implement a digital human voice interaction optimization method as described in any one of claims 1 to 7, characterized in that: The system includes the following modules: The speech recognition module is used to obtain the user's real-time speech stream through a pre-trained speech recognition model and obtain the text recognition result of the real-time speech stream; An answer generation module, connected to the speech recognition module, is used to input the text recognition result into a pre-trained large language model to generate an answer text; A speech stream synthesis module, connected to the answer generation module, for synthesizing the answer text into an answer speech stream through a pre-trained speech synthesis model; An interjection judgment module, connected to the speech recognition module and the speech stream synthesis module, for judging whether an interjection needs to be added according to the recognition delay of the speech recognition model; If yes, it goes to the modal particle insertion module, otherwise it goes to the audio output module; An interjection word insertion module is connected to the interjection word judgment module, and is used to select a target interjection word through a pre-trained context-aware model according to the context of the current conversation, and insert the target interjection word into the front end of the answer voice stream to obtain an updated answer voice stream, and enter the audio output module; wherein the context of the current conversation includes the text recognition result of the real-time voice stream and the historical record of the system response; The audio output module is connected to the modal particle judgment module and the modal particle insertion module, and is used to play the answer voice stream in real time through the audio output module to realize the voice interaction of the digital human.
Citation Information
Cited By
Digital human intelligent interaction method and system based on deep learning
CN120708611A
Digital human explanation and display control method based on voice recognition driving
CN121191518A
Digital human explanation and display control method based on voice recognition driving
CN121191518B