Fusion intelligent voice interaction system based on multithreading and dynamic context understanding

By designing a converged intelligent voice interaction system based on multi-threading and dynamic context understanding, the problems of low speech recognition accuracy, difficulty in distinguishing multiple people, and relying on network connection in the prior art are solved, and efficient, personalized and diversified voice interaction services are achieved.

CN120126482APending Publication Date: 2025-06-10BEIJING UNION UNIVERSITY
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
CN202510309596.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-17
Publication Date
2025-06-10

AI Technical Summary

Technical Problem

The existing speech recognition technology has low recognition accuracy in complex environments, making it difficult to distinguish between multiple people speaking at the same time. The large-model dialogue system relies on text input, which limits application scenarios. The speech synthesis technology is difficult to ensure emotional expression and personalized services. Most voice interaction systems rely on network connections, limiting offline use.

Method used

A converged intelligent voice interaction system based on multi-threading and dynamic context understanding is designed, including voice input, recognition, intelligent dialogue, voice synthesis and playback modules, and multi-threading voice generation and page updates are used to provide personalized and efficient voice interaction services in combination with offline and online modes.

Benefits of technology

It improves the accuracy and fluency of voice interaction, enhances the robustness and real-timeness of the system, meets the diverse needs in different scenarios, provides personalized voice synthesis services, and supports offline use.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120126482A_ABST
    Figure CN120126482A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of voice intelligent interaction, in particular to a fusion intelligent voice interaction system based on multithreading and dynamic context understanding, which comprises a voice input module, a voice recognition module, an intelligent dialogue module, a voice synthesis module and a voice playing module, and is characterized in that the voice input module is used for inputting voice information by a user; the voice recognition module is used for recognizing the voice information to obtain text information; the intelligent dialogue module is used for processing the character information to obtain answer information; the voice synthesis module is used for performing voice synthesis processing on the answer information to obtain a voice answer; and the voice playing module is used for storing and playing the voice answer. The voice interaction experience of the user can be improved, and diversified requirements of the user in different scenes can be met.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of voice intelligent interaction, and particularly to a fusion intelligent voice interaction system based on multi-threading and dynamic context understanding. Background Art

[0002] As the first threshold of human-computer interaction, the accuracy and robustness of speech recognition technology directly determine the quality of the user experience. Although significant progress has been made in speech recognition technology in recent years with the introduction of advanced algorithms such as deep learning, many problems still emerge in practical applications. First of all, the recognition accuracy needs to be improved. Especially in complex environments, such as noisy streets, crowded public places, etc., noise interference often leads to an increase in the recognition error rate. Secondly, the speech recognition system is sensitive to the accents, speaking speeds, and emotional changes of speakers. The differences between different users make it difficult for the system to treat everyone equally. Moreover, in the scenario where multiple people speak simultaneously, how to effectively distinguish and accurately recognize the speech of each person is a major difficulty in current technology.

[0003] The large model dialogue system, which complements speech recognition technology, has mastered rich language knowledge and dialogue strategies through pre-training on large-scale text data, and can generate logically coherent and linguistically fluent dialogue content. The emergence of large model dialogue systems such as kimi marks the leap of machine dialogue from a simple question-and-answer mode to complex interaction scenarios. However, current large model dialogue systems mostly rely on text input, which greatly limits the breadth of their application scenarios. In order to achieve a more natural and fluent human-machine dialogue, it is necessary to explore the path of deeply integrating the large model dialogue system with speech recognition technology. Combining speech recognition with the large model dialogue system means that the machine needs to process audio signals and text information simultaneously to achieve end-to-end voice interaction. This process not only requires high-precision speech recognition as the front-end input, but also requires the large model dialogue system to have strong semantic understanding and generation capabilities. In addition, how to ensure the efficient and accurate data flow between the two, and avoid information loss or misunderstanding, is also an issue that needs to be focused on during the technology integration process.

[0004] In a complete voice interaction experience, speech synthesis technology plays an indispensable role. The development of speech synthesis technology has made the speech generated by machines increasingly close to human natural speech, bringing a more natural and fluent interaction experience to users. However, speech synthesis technology is not perfect. While pursuing high naturalness, how to ensure that the emotional expression, intonation changes of the generated speech match the text content is a major problem in current research. In addition, different users have different preferences for speech styles. How to provide personalized speech synthesis services for users is also an important direction for future technology development.

[0005] With the popularization of Internet of Things devices and the widespread application of mobile applications, more and more users hope to enjoy convenient voice interaction services even in a network-free environment. However, most of the current voice interaction systems in the market rely on network connections, which to a certain extent limits their usage scenarios. Therefore, it is particularly important to develop a voice interaction system that supports the offline mode. The implementation of an offline voice interaction system faces many technical challenges. First of all, how to store and efficiently run large voice recognition, dialogue processing, and speech synthesis models on local devices is a key issue. This requires researchers to continuously optimize the model structure and compress model parameters to achieve the deployment of lightweight models. Secondly, how to ensure the real-time performance and accuracy of the system in a network-free environment is also a huge challenge. This requires careful design and tuning of the model to ensure its robustness in various complex environments.

[0006] In multi-turn conversations, the ability of a voice interaction system to understand and remember context is crucial for enhancing the user experience. However, many current voice interaction systems still have deficiencies in this regard, and users may need to repeat providing information because the system cannot accurately remember previous conversation content or user preferences. In addition, for many new users, how to use a voice interaction system efficiently is still a challenge. The system needs to provide clear and easy-to-understand user guides and educational materials to help users get started quickly and fully utilize the potential of the system. At the same time, the system should also have the ability to self-learn and be able to continuously optimize the interaction process and prompt information according to the user's usage situation. Therefore, in view of the above technical defects, this application proposes a fusion intelligent voice interaction system based on multi-threading and dynamic context understanding. Summary of the Invention

[0007] The object of the present invention is to provide a fusion intelligent voice interaction system based on multi-threading and dynamic context understanding, which can enhance the user voice interaction experience and meet the diverse needs of users in different scenarios.

[0008] To achieve the above object, the present invention provides the following solutions:

[0009] A fusion intelligent voice interaction system based on multi-threading and dynamic context understanding, comprising: a voice input module, a voice recognition module, an intelligent dialogue module, a speech synthesis module, and a voice playback module. Among them, the voice input module is used for users to input voice information; the voice recognition module is used to recognize the voice information to obtain text information; the intelligent dialogue module is used to process the text information to obtain an answer message; the speech synthesis module is used to perform speech synthesis processing on the answer message to obtain a voice answer; the voice playback module is used to store and play the voice answer.

[0010] Optionally, the voice input module uses an audio data acquisition interface, and the voice playback module uses a voice playback interface. Both the audio data acquisition interface and the voice playback interface are provided with a storage function for storing question-and-answer interaction records.

[0011] Optionally, the voice recognition module uses a voice recognition technology interface to preprocess the voice information and perform feature extraction and recognition based on the preprocessed voice information.

[0012] Optionally, the voice recognition module includes a preprocessing unit, an online recognition unit, and an offline recognition unit. Among them, the preprocessing unit is used to perform adaptive noise suppression and echo cancellation preprocessing on the voice information; the online recognition unit is used to provide an online recognition mode and use a Baidu recognizer to recognize the preprocessed voice information to obtain an online recognition result; the offline recognition unit is used to provide an offline recognition mode and use a Whisper recognizer to recognize the preprocessed voice information to obtain an offline recognition result.

[0013] Optionally, the intelligent dialogue module uses an intelligent model dialogue interface to input the text information into a large language model, generate answer information, and continuously save it.

[0014] Optionally, the intelligent dialogue module includes an online dialogue unit and an offline dialogue unit. Among them, the online dialogue unit is used to provide an online dialogue mode and use a kimi large language model to process the text information, generate online answer information, and save it as a continuous dialogue; the offline dialogue unit is used to provide an offline dialogue mode and use a local large language model to process the text information, generate offline answer information, and save it as a continuous dialogue.

[0015] Optionally, when the intelligent dialogue module generates answer information, it also combines the historical dialogue record and the context information of the current dialogue.

[0016] Optionally, the voice synthesis module uses a voice synthesis technology interface to synthesize the answer information into a voice form to generate a voice answer.

[0017] Optionally, the voice synthesis module includes an online synthesis unit and an offline synthesis unit. Among them, the online synthesis unit is used to provide an online synthesis mode and use a moonshot voice synthesis model to synthesize the answer information into an online voice answer; the offline synthesis unit is used to provide an offline synthesis mode and use a ChatTTS voice synthesis model to synthesize the answer information into an offline voice answer.

[0018] Optionally, a multi-threaded long text speech generation algorithm is adopted in the speech synthesis module. The multi-threaded long text speech generation algorithm is used to split the answer information of the long text into several short texts, perform speech processing on each short text, and then splice them to obtain a complete speech answer.

[0019] The beneficial effects of the present invention are as follows:

[0020] The present invention designs a complete system interaction logic to ensure seamless connection of the three stages of speech recognition, dialogue generation, and speech synthesis. By optimizing the data processing flow and algorithm parameters, the system response speed and interaction fluency are improved; at the same time, a multi-threaded method is adopted, with the change of the page itself written as one thread and the speech generation written as another thread, ensuring that the threads of the two do not conflict and preventing problems such as thread jams and no page changes during speech generation, resulting in users not getting feedback; using a large model for dialogue, with context understanding and memory capabilities, it can continuously track the user's dialogue content and generate more coherent and natural dialogue responses; offline versions and online versions are set for all three stages of speech recognition, large model dialogue, and speech synthesis, and conversion between the online version and the offline version can be performed, which can meet the different needs of users and facilitate future access to devices for offline use. BRIEF DESCRIPTION OF THE DRAWINGS

[0021] In order to more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the following will briefly introduce the drawings required for use in the embodiments. Obviously, the drawings in the following description are only some embodiments of the present invention. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.

[0022] Figure 1 Structural diagram of the fusion intelligent voice interaction system based on multi-threading and dynamic context understanding according to an embodiment of the present invention;

[0023] Figure 2 Flowchart of the working process of the fusion intelligent voice interaction system based on multi-threading and dynamic context understanding according to an embodiment of the present invention;

[0024] Figure 3 Flowchart of the multi-threaded long text speech generation algorithm according to an embodiment of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0025] The following will clearly and completely describe the technical solutions in the embodiments of the present invention with reference to the drawings in the embodiments of the present invention. Obviously, the described embodiments are only some embodiments of the present invention, rather than all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of the present invention.

[0026] To make the above objects, features, and advantages of the present invention more obvious and understandable, the present invention will be further described in detail below with reference to the accompanying drawings and specific embodiments.

[0027] This embodiment provides a fusion intelligent voice interaction system based on multi-threading and dynamic context understanding, as Figure 1 , Figure 2 shown, including: a voice input module, a speech recognition module, an intelligent dialogue module, a speech synthesis module, and a voice playback module. Among them, the voice input module is used for the user to input voice information; the speech recognition module is used to recognize the voice information to obtain text information; the intelligent dialogue module is used to process the text information to obtain an answer message; the speech synthesis module is used to perform speech synthesis processing on the answer message to obtain a voice answer; the voice playback module is used to store and play the voice answer.

[0028] Specifically, this embodiment integrates a large model dialogue with context understanding and memory capabilities to ensure that the system can continuously track the user's conversation content, generate more coherent and natural dialogue responses, and make the human-computer dialogue closer to real communication scenarios. In addition, two working modes, an offline version and an online version, are provided for the three core stages of speech recognition, large model dialogue, and speech synthesis. Users can choose a suitable version according to the actual scenario and requirements. This flexible configuration method not only broadens the application scenarios of the system but also provides convenience for the access of future offline devices, further enhancing the practicality and popularity of the system.

[0029] Furthermore, the voice input module adopts an audio data acquisition interface, and the voice playback module adopts a voice playback interface. Both the audio data acquisition interface and the voice playback interface are provided with a storage function for storing question-and-answer interaction records.

[0030] Specifically, this embodiment builds an acquisition interface for audio data, allowing users to interact with the system through audio input devices such as microphones, and designs an interface to save the input voice files to ensure the security and traceability of the data.

[0031] Build a voice playback interface responsible for playing the audio file generated in the speech synthesis stage to the user, thus completing the closed-loop of the entire voice interaction.

[0032] Furthermore, the speech recognition module adopts a speech recognition technology interface for preprocessing the voice information and performing feature extraction and recognition based on the preprocessed voice information.

[0033] Among them, the speech recognition module includes a preprocessing unit, an online recognition unit, and an offline recognition unit. Among them, the preprocessing unit is used to perform adaptive noise suppression and echo cancellation preprocessing on the speech information; the online recognition unit is used to provide an online recognition mode, and use Baidu recognizer to recognize the preprocessed speech information to obtain an online recognition result; the offline recognition unit is used to provide an offline recognition mode, and use Whisper recognizer to recognize the preprocessed speech information to obtain an offline recognition result.

[0034] Specifically, in this embodiment, a speech recognition technology interface is built to accurately convert the collected audio data into text. The interface supports preprocessing technologies such as adaptive noise suppression and echo cancellation, and has the ability to optimize feature extraction, etc. to improve the recognition accuracy and robustness.

[0035] Build an online speech recognition interface and connect to a callable network demand interface, specifically the online version of Baidu Speech API; build an offline speech recognition interface and deploy a local model, specifically the offline version of the Whisper model.

[0036] Furthermore, the intelligent dialogue module uses an intelligent model dialogue interface to input the text information into a large language model, generate response information and save it continuously.

[0037] Among them, the intelligent dialogue module includes an online dialogue unit and an offline dialogue unit. Among them, the online dialogue unit is used to provide an online dialogue mode, and use the kimi large language model to process the text information, generate online response information and save it as a continuous dialogue; the offline dialogue unit is used to provide an offline dialogue mode, and use a local large language model to process the text information, generate offline response information and save it as a continuous dialogue.

[0038] When the intelligent dialogue module generates response information, it also combines the historical dialogue record and the context information of the current dialogue.

[0039] Specifically, in this embodiment, an intelligent dialogue interface is built to connect the data transmitted by speech recognition and conduct a dialogue.

[0040] Build an online intelligent dialogue interface, connect to the api of the online version of kimi, and through a large amount of pre-trained text data, achieve highly natural and fluent dialogue generation. Design the dialogue management logic to ensure that the system can accurately understand the user's intention and give appropriate responses, and implement the context memory function to enable the system to maintain coherence in multi-round dialogues.

[0041] Build an offline intelligent dialogue interface and deploy an offline version of the local large language model.

[0042] In this embodiment, the content input by the system each time will be passed into the sequence model, and the sequence model will be considered each time when answering to implement the context understanding and memory functions. The system maintains an internal state to represent the user's conversation history and current context. When generating a response, the intelligent conversation interface will refer to this state to ensure the coherence and personalization of the conversation.

[0043] Further, the speech synthesis module adopts a speech synthesis technology interface, which is used to synthesize the answer information into a voice form to generate a voice answer.

[0044] Among them, the speech synthesis module includes an online synthesis unit and an offline synthesis unit. The online synthesis unit is used to provide an online synthesis mode and synthesize the answer information into an online voice answer by using the moonshot speech synthesis model. The offline synthesis unit is used to provide an offline synthesis mode and synthesize the answer information into an offline voice answer by using the ChatTTS speech synthesis model.

[0045] The speech synthesis module adopts a multi-threaded long text speech generation algorithm, which is used to split the long text answer information into several short texts, perform speech processing on each short text, and then splice them to obtain a complete voice answer.

[0046] Specifically, in this embodiment, a speech synthesis technology interface is built to connect the data transmitted by the intelligent conversation and perform speech synthesis.

[0047] Build an offline speech synthesis interface and connect it to the offline ChatTTS model, which can convert the text conversation result into a realistic voice output. During this process, the speech synthesis algorithm automatically adjusts the prosody, speech rate and tone color of the voice according to the emotional tendency, grammatical structure and context information of the text to ensure that the synthesized voice is highly consistent with the original text content and conforms to the human auditory habit.

[0048] Build an online speech synthesis interface and connect it to the api of the online version of moonshot, which can connect to the network cloud to generate the voice of the text.

[0049] The process of the multi-threaded long text speech generation algorithm is as Figure 3 shown. In this embodiment, the speech synthesis is achieved by splitting the long text into several clauses, allocating independent threads to each clause for speech generation. After all threads complete the tasks, the generated speech segments are merged through a synchronization mechanism, and finally a complete speech file is output.

[0050] This embodiment adopts a global variable mechanism to achieve data transfer and status synchronization among modules. Each module runs independently and interacts through a shared data structure, avoiding direct dependencies and improving the maintainability and scalability of the system. At the same time, a real-time page update mechanism is implemented. First, a message communication framework based on the publish-subscribe pattern is built. In this framework, the front end listens for events from the back end. When the data in the back-end system changes, the back end pushes a message about the data change to the front end, and then calls the front-end interface update function. At this time, the front end will re-fetch the data from the back end and render the interface. Each component of the page does not directly operate on the back-end interface or global variables, but self-updates by listening for status changes and obtaining the latest data, realizing the decoupling and hierarchical design of interface components. For example, during the operation of the speech recognition module, the real-time recognized speech text will push a message to the front end. At this time, the front end obtains the speech text and updates the originally stored speech text, and then the interface will be refreshed and display the latest text content. This ensures the consistency between the user interface and the back-end state. Separate the front-end display from the back-end processing. The front-end thread is responsible for the interaction of the user interface, and the back-end thread is responsible for complex computing tasks (such as speech recognition, dialogue generation, etc.). The interaction between the two is limited to data transmission, so as to achieve low-latency communication between the front end and the back end, improving the system response speed and user experience.

[0051] In this embodiment, conversion options for the online version and the offline version are set for each of the three modules of speech recognition, intelligent dialogue, and speech synthesis. Through the user interface, the user can separately set the online version or the offline version for each module. When a module is called, the mark of the module will be checked, and then the online version or the offline version will be selected for invocation according to the key value of the mark. For the online version, the API will be called, and for the offline version, the local model will be called to ensure stable services can be provided in different network environments or meet the needs of different users.

[0052] The following provides the design process of the integrated intelligent voice interaction system based on multi-threading and dynamic context understanding in this embodiment, including the following steps:

[0053] S1. Design and build the overall system framework to ensure seamless docking of each module and achieve efficient data flow and processing. The framework needs to support both online and offline working modes to meet the usage requirements in different scenarios.

[0054] S2. Build an acquisition interface for audio data, allowing users to interact with the system through audio input devices such as microphones. Design an interface to save the input voice files to ensure the security and traceability of the data.

[0055] Specifically, in terms of saving voice files in this embodiment, there is the following logic: Detect whether there is a saved voice file locally. If so, the saved voice file will be deleted and replaced to update the audio acquisition data.

[0056] S3. Build a speech recognition technology interface to accurately convert the collected audio data into text. The interface supports preprocessing technologies such as adaptive noise suppression and echo cancellation, and has the ability to optimize feature extraction, etc. to improve the recognition accuracy and robustness.

[0057] S3.1 Build an online speech recognition module and connect to the callable network demand interface;

[0058] S3.2 Build an offline speech recognition module and deploy the local model;

[0059] S4. Build an intelligent model dialogue interface to connect and conduct conversations with the data transmitted from speech recognition;

[0060] S4.1 Build an online intelligent model dialogue interface. Through a large amount of pre-trained text data, achieve highly natural and fluent dialogue generation. Design the dialogue management logic to ensure that the system can accurately understand the user's intention and give appropriate responses. Implement the context memory function to enable the system to maintain coherence in multi-round conversations.

[0061] S4.2 Build an offline intelligent model dialogue module and deploy the local model;

[0062] S5. Build a speech synthesis technology interface, which connects and conducts speech synthesis with the data transmitted from the intelligent model dialogue;

[0063] S5.1 Build an offline speech synthesis module that can convert the text dialogue result into a realistic voice output. During this process, the speech synthesis algorithm automatically adjusts the rhythm, speed, and tone of the voice according to the emotional tendency, grammatical structure, and context information of the text to ensure that the synthesized voice is highly consistent with the original text content and conforms to the human auditory habit.

[0064] S5.2 Build an online speech synthesis interface that can connect to the network cloud to generate the voice of the text;

[0065] S6. Build a voice playback interface, which is responsible for playing the audio file generated in step S5 (speech synthesis stage) to the user, thus completing the closed-loop of the entire voice interaction;

[0066] S7. The user submits voice commands or queries through the integrated voice input interface. This interface uses high-precision microphone array technology to capture the original audio signal and convert it into a digital audio stream. Subsequently, the audio stream is passed to the speech recognition module, which performs processing steps such as feature extraction, noise suppression, and decoding using speech recognition to convert the audio stream into a readable text command or question. After receiving the text output by the speech recognition module, the intelligent dialogue module uses natural language processing (NLP) technology for semantic understanding, intent recognition, and context management to generate a response text with clear logic and fluent language. The generated response text is then sent to the speech synthesis module to convert the text into a natural and fluent voice output.

[0067] The following provides the working process of the integrated intelligent voice interaction system based on multi-threading and dynamic context understanding in this embodiment, including the following:

[0068] As Figure 1 , when the user presses the voice input key and performs voice input, the speech recognition interface will be called. The speech recognition interface has access to the online version of Baidu Speech API and the offline version of the Whisper model, and the user can choose freely. Before calling this interface, the microphone is first called to save the collected audio locally, and then the audio file will be recognized according to the model selected by the user. Among them, the attribute brecognizer is used to initialize the online version API call function, wrecognizer is used to initialize the offline version Whisper model call function, and then the attribute text is used to receive the text result recognized by the speech recognition model. Therefore, for the user's need to use their own speech recognition model, when replacing, only ensure that the method signature of the new module is consistent with the existing module. For example, the speech recognition module has a wrecognize or brecognize method for initializing the function, and the user's model only needs to use this method for initialization. The system will pass the speech recognition result to the next module, so only use the self.text attribute to receive the text result output by the user's speech recognition model.

[0069] After speech recognition is completed, the text result of speech recognition is given to the large language model to make an answer. The large language model interface has access to the online version of Kimi's API and the offline version of the local model. After the system determines the selection of the large language model, it will use the corresponding model to make an answer. Among them, the attribute dialog is used to initialize the online version API call function, udialog is used to initialize the offline version model, and then the result attribute is used to store the result returned by the large language model. Therefore, for the user's need to use their own large language model, just pass the initialization parameters of their own model into dialog or udialog and use result to store the result.

[0070] After the large language model obtains the result, it will pass it to the speech synthesis model. The speech synthesis interface of this system has been connected to the online version of the moonshot API and the offline version of the ChatTTS model. After the system determines the use of the speech synthesis model, it calls the corresponding model call function and uses the wavs attribute to initialize the model call function. Inside the call function, the generated voice file will be stored locally as "audio.mp3". Then the system will retrieve the local audio file named "audio.mp3" and use the built-in playsound function of python to play the audio file. Therefore, for users who want to use their own speech synthesis model, they only need to name the voice generated by the model as "audio.mp3" and store it locally.

[0071] The speech recognition module, intelligent dialogue module, and speech synthesis module in this system all use the factory mode to create module instances. When you need to switch modules, you only need to change the implementation in the factory method. For example, when creating a speech recognizer, decide whether to return the Baidu recognizer or the Whisper recognizer based on the parameters passed in. Then, the dependencies between modules are passed through the constructor and parameters, rather than directly instantiating the specific implementation inside the module. This means that when replacing your own module, you only need to pass in the initialization parameters and the text content generated by the model, which ensures the convenience of replacing modules without causing code errors due to module replacement.

[0072] In addition, inside the large language model calling function, a self.history is defined to store the results of the current round of dialogue. When the large language model gives the dialogue result, the result of the current round of dialogue will be recorded in self.history to achieve state update. In subsequent dialogues, the historical records stored in self.history will be traversed first to achieve the function of memorizing the dialogue.

[0073] The embodiments described above are only descriptions of the preferred embodiments of the present invention and are not intended to limit the scope of the present invention. Without departing from the design spirit of the present invention, various modifications and improvements made to the technical solutions of the present invention by ordinary technicians in this field should all fall within the protection scope determined by the claims of the present invention.

Claims

1. A fusion intelligent voice interaction system based on multi-threading and dynamic context understanding, characterized by: include: A voice input module, a voice recognition module, an intelligent dialogue module, a voice synthesis module, and a voice playback module, wherein the voice input module is used for the user to input voice information; The speech recognition module is used to recognize the speech information and obtain text information; the intelligent dialogue module is used to process the text information and obtain answer information; the speech synthesis module is used to perform speech synthesis processing on the answer information and obtain a speech answer; the speech playback module is used to store and play the speech answer.

2. The fusion intelligent voice interaction system based on multi-threading and dynamic context understanding according to claim 1 is characterized in that: The voice input module adopts an audio data acquisition interface, and the voice playback module adopts a voice playback interface. Both the audio data acquisition interface and the voice playback interface are provided with a storage function for storing question-and-answer interaction records.

3. The fusion intelligent voice interaction system based on multi-threading and dynamic context understanding according to claim 1 is characterized in that: The speech recognition module uses a speech recognition technology interface to preprocess the speech information and perform feature extraction and recognition based on the preprocessed speech information.

4. The fusion intelligent voice interaction system based on multi-threading and dynamic context understanding according to claim 3 is characterized in that: The speech recognition module includes a preprocessing unit, an online recognition unit and an offline recognition unit, wherein the preprocessing unit is used to perform adaptive noise suppression and echo elimination preprocessing on the speech information; the online recognition unit is used to provide an online recognition mode, use the Baidu recognizer to recognize the preprocessed speech information, and obtain an online recognition result; the offline recognition unit is used to provide an offline recognition mode, use the Whisper recognizer to recognize the preprocessed speech information, and obtain an offline recognition result.

5. The fusion intelligent voice interaction system based on multi-threading and dynamic context understanding according to claim 1 is characterized in that: The intelligent dialogue module adopts an intelligent model dialogue interface to input the text information into a large language model, generate answer information and save it continuously.

6. The fusion intelligent voice interaction system based on multithreading and dynamic context understanding according to claim 5 is characterized in that: The intelligent dialogue module includes an online dialogue unit and an offline dialogue unit, wherein the online dialogue unit is used to provide an online dialogue mode, use the kimi large language model to process the text information, generate online answer information and save it as a continuous dialogue; the offline dialogue unit is used to provide an offline dialogue mode, use the local large language model to process the text information, generate offline answer information and save it as a continuous dialogue.

7. The fusion intelligent voice interaction system based on multi-threading and dynamic context understanding according to claim 6 is characterized in that: The intelligent dialogue module also combines historical dialogue records with context information of the current dialogue when generating answer information.

8. The fusion intelligent voice interaction system based on multithreading and dynamic context understanding according to claim 1 is characterized in that: The speech synthesis module adopts a speech synthesis technology interface to synthesize the answer information into a speech form to generate a speech answer.

9. The fusion intelligent voice interaction system based on multithreading and dynamic context understanding according to claim 8 is characterized in that: The speech synthesis module includes an online synthesis unit and an offline synthesis unit, wherein the online synthesis unit is used to provide an online synthesis mode, and adopts a moonshot speech synthesis model to synthesize the answer information into an online speech answer; the offline synthesis unit is used to provide an offline synthesis mode, and adopts a ChatTTS speech synthesis model to synthesize the answer information into an offline speech answer.

10. The fusion intelligent voice interaction system based on multi-threading and dynamic context understanding according to claim 9 is characterized in that: The speech synthesis module adopts a multi-threaded long text speech generation algorithm, which is used to divide the answer information of the long text into several short texts, perform speech processing on each short text, and then splice them to obtain a complete speech answer.

Citation Information

Cited By

  • Event processing method and system based on large language model

    CN121260185A