Real-time voice interaction method and system, computer equipment and storage medium

By deploying the voice processing module on the local device and adopting streaming processing mechanism and asynchronous multi-threading, the existing voice interaction system's slow response speed and poor emotional tone processing are solved, and an efficient and natural voice interaction experience is achieved.

CN119993150AInactive Publication Date: 2025-05-13SHENZHEN ZHUOYUE ZHIYUN TECHNOLOGY CO LTD
View PDF 5 Cites 0 Cited by

Patent Information

Application Number
CN202510249092.5
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-04
Publication Date
2025-05-13
Estimated Expiration
Not applicable · inactive patent

AI Technical Summary

Technical Problem

Existing voice interaction systems are slow to respond when handling complex conversations, which is difficult to meet real-time interaction needs, and have limitations in processing emotions and intonation in speech, and cannot fully simulate human voice expression.

Method used

By deploying speech recognition modules, language processing modules and voice conversion modules on local devices, the text information is processed in segments using streaming processing mechanisms, dynamically adjusting the parameters of text segments, and using asynchronous multi-threading to perform audio generation tasks, optimizing workflows to improve response speed and interaction fluency.

Benefits of technology

It realizes the reduction of data transmission delay, improves response speed, enhances the accuracy of speech recognition and the naturalness of interaction, reduces the dependence on computing resources, and enables the system to operate efficiently on resource-constrained devices.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119993150A_ABST
    Figure CN119993150A_ABST
Patent Text Reader

Abstract

The invention provides a real-time voice interaction method and system, computer equipment and a storage medium. The real-time voice interaction method comprises the following steps: acquiring a voice signal input by voice input equipment; performing primary processing on the voice signal; converting and processing the signal through the voice recognition module; segmenting the text information through a streaming processing mechanism, and transmitting the segmented text information to a language processing module; reply information is generated through a language processing module according to the text segment, and parameters of the text segment are dynamically adjusted; the reply information is sent to the voice conversion module; and the voice conversion module converts the reply information into a synthesized voice signal in real time, and sends the synthesized voice signal to the loudspeaker for playing. According to the invention, the voice recognition module, the language processing module and the voice conversion module are deployed on the local equipment, the delay of data transmission is reduced, the response speed is improved, and the voice processing module containing the language model is arranged, so that the system can adapt to different interaction scenes. Feedback is quickly obtained through a streaming processing mechanism, and text segments are dynamically adjusted to improve emotion and context processing.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of intelligent interaction technology, and in particular to a real-time voice interaction method, system, computer equipment and storage medium. Background Art

[0002] With the rapid development of large language models (LLMs), voice interaction systems are increasingly used in various fields, such as intelligent assistants, customer service robots, and educational tools. These voice interaction systems can interact with users efficiently through natural language processing technology. Real-time performance is one of the key requirements of voice interaction systems, especially in scenarios such as online customer service and medical consultation that require instant feedback.

[0003] Most current voice interaction systems rely on traditional speech recognition and natural language understanding technologies. However, these technologies are prone to latency issues when processing complex conversations. Traditional voice interaction systems usually use a cascaded model of speech recognition (ASR), natural language processing (NLP), and text-to-speech synthesis (TTS). During this process, data sets need to be transferred between multiple modules, resulting in slow response speed, high latency, and difficulty in meeting the needs of real-time interaction.

[0004] On the other hand, existing technologies still have limitations in processing emotions and intonation in speech, and cannot fully simulate human speech expression, resulting in unnatural speech interaction. Moreover, when processing complex speech tasks, such as multi-round conversations and emotion recognition, the performance is not ideal, and it is difficult to meet the interaction needs of diverse scenarios. In addition, the demand for computing resources is high, resulting in high resource consumption, which limits its application on resource-constrained devices. Summary of the invention

[0005] The present invention aims to solve the problems in the above-mentioned prior art of slow response speed, poor emotion and intonation processing, and difficulty in meeting the interaction needs of diverse scenarios, and provides a real-time voice interaction method, system, computer device and storage medium.

[0006] The present invention provides a real-time voice interaction method, comprising the following steps:

[0007] Acquire a voice signal input by a preset voice input device;

[0008] Performing preliminary processing on the speech signal to obtain a processed signal;

[0009] Converting the processed signal through a preset speech recognition module to obtain text information;

[0010] The text information is segmented through a preset streaming processing mechanism to obtain a plurality of text segments, and transmitted to a preset language processing module; wherein the language processing module includes at least one language processing model;

[0011] Generate reply information according to the text segment through the language processing module, and dynamically adjust the parameters of the text segment;

[0012] Sending the reply information to a preset voice conversion module;

[0013] The reply information is converted into a synthetic voice signal in real time by the voice conversion module and sent to a preset speaker for playing.

[0014] Furthermore, the step of segmenting the text information through a preset streaming processing mechanism to obtain a plurality of text segments and transmitting the segments to a preset language processing module includes:

[0015] The text segments are transmitted one by one to the language processing module to generate reply information in real time and transmitted to the voice conversion module.

[0016] Furthermore, the step of generating reply information from the text segment through the language processing module and dynamically adjusting the parameters of the text segment includes:

[0017] Monitoring the length and processing speed of the text segment to obtain monitoring results;

[0018] The length and processing speed of the text segment are adjusted according to the monitoring result.

[0019] Furthermore, the step of monitoring the length and processing speed of the text segment and obtaining the monitoring result includes:

[0020] Semantic pause points of the text segment are identified, and the processing speed is adjusted to adjust the pause time.

[0021] Further, the processing signal includes a previous processing signal and a current processing signal; the text segment includes a previous text segment and a current text segment;

[0022] When the language model receives the text information converted from the last processed signal, the speech recognition module converts the current processed signal;

[0023] When the speech conversion module receives the reply information generated by the previous text segment, the language processing module generates the reply information of the current text segment.

[0024] Further, the reply information includes the last reply information and the current reply information; the last reply information and the current reply information are respectively converted into output voice signals by the voice conversion module;

[0025] The output speech signal is synthesized into the synthesized speech signal.

[0026] Furthermore, the text segment also includes a next text segment; the reply information includes next reply information; when the output voice signal is synthesized into the synthesized voice signal, the language processing module generates the next reply information according to the next text segment.

[0027] Further, the step of converting the previous reply information and the current reply information into output voice signals respectively by the voice conversion module includes:

[0028] The audio generation task is asynchronously performed by the voice conversion module to simultaneously obtain the output voice signals of the previous reply information and the current reply information.

[0029] Furthermore, before the step of converting the reply information into a synthetic voice signal in real time by the voice conversion module and sending it to a preset speaker for playing, the step includes:

[0030] The synthesized speech signal is optimized; wherein the optimization includes but is not limited to adjusting the signal-to-noise ratio, intonation and pronunciation.

[0031] The present invention also provides a real-time voice interaction system, comprising a voice input device, a local device and a speaker, wherein the voice input device is used to input voice signals; the local device comprises a voice recognition module, a language processing module and a voice conversion module, the voice recognition module is used to convert and process signals to obtain text information; the language processing module is used to generate reply information for each text segment; the voice conversion module is used to convert each of the text segments into an output voice signal in real time; and the speaker is used to play the reply information.

[0032] The present invention also provides a computer device, including a memory and a processor, wherein a computer program is stored in the memory, and the processor executes the computer program to implement the steps in any one of the above methods.

[0033] The present invention also provides a computer-readable storage medium on which a computer program is stored. When the computer program is executed by a processor, the steps of any of the above methods are implemented.

[0034] The present invention provides a real-time voice interaction method, system, computer device and storage medium, which have the following beneficial effects:

[0035] This application can reduce the delay of data transmission, improve the response speed, improve data privacy protection and reduce potential security risks by deploying speech recognition module, language processing module and speech conversion module on the local device. It is also equipped with a speech processing module containing a language model, which can understand and process complex conversations to adapt to different contexts and interaction scenarios, recognize and process the input speech signal, and improve the accuracy of speech recognition. It adopts a modular parallel processing architecture, uses asynchronous multi-threading for audio generation tasks, optimizes the workflow, and can avoid long waiting times due to thread congestion. It can continue to perform other tasks while waiting for reply information generation and audio generation, ensuring that each module can run independently and efficiently, generate voice feedback in time, and enhance the interactive experience. At the same time, it can reduce the dependence on computing resources, so that the system can run efficiently on resource-constrained devices, and broaden the application scenarios.

[0036] This application uses a streaming processing mechanism to process text information in real time and in segments, ensuring that users can quickly get feedback after issuing voice commands, improving the real-time nature of voice interaction and user satisfaction. It can also dynamically adjust the length and processing speed of text segments based on user interaction to improve processing of emotions and context. BRIEF DESCRIPTION OF THE DRAWINGS

[0037] Figure 1 A schematic diagram of the method steps of a real-time voice interaction method in the present invention;

[0038] Figure 2 A structural block diagram of a real-time voice interaction system in the present invention;

[0039] Figure 3 A structural block diagram of a computer device of the present invention;

[0040] Figure 4 The figure is a schematic diagram of the steps of an embodiment of a real-time voice interaction method of the present invention.

[0041] Description of the symbols: voice input device 10, local device 20, speaker 30, voice recognition module 21, language processing module 22, voice conversion module 23. DETAILED DESCRIPTION

[0042] It should be understood that the specific embodiments described herein are only used to explain the present invention, and are not used to limit the present invention.

[0043] The following will be combined with the drawings in the embodiments of the present invention to clearly and completely describe the technical solutions in the embodiments of the present invention. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of the present invention.

[0044] To improve real-time performance, many solutions are dedicated to optimizing each link of speech recognition, dialogue management and response generation to reduce the response time of the system.

[0045] Reference Figure 1 , is a real-time voice interaction method in one embodiment of the present invention, comprising:

[0046] S1, obtaining a voice signal input by a preset voice input device 10;

[0047] S2, preliminarily processing the speech signal to obtain a processed signal;

[0048] S3, converting and processing the signal through the preset speech recognition module 21 to obtain text information;

[0049] S4, segmenting the text information through a preset streaming processing mechanism to obtain a plurality of text segments, and transmitting the segments to a preset language processing module 22; wherein the language processing module 22 includes at least one language processing model;

[0050] S5, generating reply information according to the text segment through the language processing module 22, and dynamically adjusting the parameters of the text segment;

[0051] S6, sending the reply information to the preset voice conversion module 23;

[0052] S7, the reply information is converted into a synthetic voice signal in real time by the voice conversion module 23, and sent to the preset speaker 30 for playing.

[0053] In the above steps, the user first performs voice input through the voice input device 10 built into the local device 20. The local device 20 is set in a way that it can quickly receive the voice signal after the user inputs the voice, thereby improving the fluency of the voice interaction. Figure 4As shown, in a specific embodiment, the voice input device 10 is a microphone. Then, the voice signal input by the user is obtained, and noise reduction and enhancement processing are performed to obtain a processed signal. Noise reduction of the voice signal can remove environmental noise, enhance voice clarity, and improve the clarity and accuracy of voice input. Make the subsequent voice recognition and generation process more accurate, thereby improving the overall interaction effect. Then the processed signal is sent to the locally deployed voice recognition module 21, and the locally deployed method can efficiently convert the processed signal into text information in a low-latency environment. In a specific embodiment, the voice recognition module 21 is ASR, Automatic Speech Recognition. Then, the text information is segmented through a preset streaming processing mechanism and transmitted to the language processing module 22; specifically, the streaming processing mechanism is to segment the text information, and generate and output the reply information of the segmented text information in real time, which is different from the method of waiting for the entire reply information to be generated and then transmitted at one time. Then the language processing module 22 generates the reply information according to the text segment and dynamically adjusts the parameters of the text segment. Specifically, a language model is provided in the language processing module 22. Preferably, the language model is LLM, Large Language Model, which can generate reply information for multiple text segments respectively. In a specific embodiment, the parameters of the text segment include the length and processing speed of the text segment. After the reply information of the text segment is generated, the length and processing speed of the reply information are adjusted. Then the reply information is sent to the voice conversion module 23, and finally the reply information is converted into a synthetic voice signal by the voice conversion module 23, and the synthetic voice signal is played through the speaker 30, so that the user can hear the response of the system in time, and the interaction fluency is improved. In a specific embodiment, the voice conversion module 23 is TTS, Text To Speech, which can convert the reply information into a synthetic voice signal. By capturing the voice and performing noise reduction processing on the voice, the accuracy of voice recognition is significantly improved. At the same time, the language processing module 22 is used to generate a natural and fluent voice response, which further improves the naturalness and comfort of the interaction.

[0054] In one embodiment, the step of segmenting the text information by a preset streaming processing mechanism to obtain a plurality of text segments and transmitting the segments to the preset language processing module 22 includes:

[0055] The text segments are transmitted one by one to the language processing module 22 to generate reply information in real time and transmitted to the voice conversion module 23.

[0056] In this embodiment, after each text segment is generated, it is transmitted to the preset voice conversion module 23. Specifically, after each text segment is generated, the current text segment is transmitted to the language processing module 22, and the reply information is generated in real time. Then the next text segment is transmitted, and the reply information is generated again. That is, the transmission is performed in the form of serial transmission.

[0057] In one embodiment, the step of generating reply information from a text segment by the language processing module 22 and dynamically adjusting the parameters of the text segment includes:

[0058] Monitor the length and processing speed of the text segment and obtain monitoring results;

[0059] Adjust the length and processing speed of the text segment based on the monitoring results.

[0060] In this embodiment, when the parameters of the text segment are dynamically adjusted, the length and processing speed of the text segment are first monitored and the monitoring results are obtained. Then the length and processing speed of the text segment are adjusted according to the monitoring results. In a specific embodiment, the language model provided in the language processing module 22 captures the intonation of the voice signal input by the user, wherein the intonation can be lyrical or cheerful. When the output reply message requires a slow and lyrical intonation, the processing speed is reduced to reduce the speed of voice output. When the output reply message requires a cheerful intonation, the processing speed is increased to increase the voice output speed. It has a real-time feedback monitoring function, and can dynamically adjust the text segment length and processing speed according to the user interaction situation, ensuring the synchronization between the synthesized voice signal synthesized by the voice conversion module 23 and the generated by the language processing module 22, and improving the coherence and naturalness of user interaction.

[0061] In one embodiment, the step of monitoring the length and processing speed of the text segment and obtaining the monitoring result includes:

[0062] Identify semantic pause points in a text segment and adjust processing speed to adjust the pause time.

[0063] In this embodiment, semantic pause points in the text segment are identified, such as a period and a comma, and the processing speed is adjusted to adjust the pause time to make the speech output more natural. In a specific embodiment, the pause time of a period is 1 second and the pause time of a comma is 0.5 seconds. When a period or a comma is identified, the processing speed is adjusted to adjust the pause time.

[0064] In one embodiment, the processed signal includes a previous processed signal and a current processed signal; the text segment includes a previous text segment and a current text segment;

[0065] When the language model receives text information converted from the previous processed signal, the speech recognition module 21 converts the current processed signal;

[0066] When the speech conversion module 23 receives the reply information generated by the previous text segment, the language processing module 22 generates the reply information of the current text segment.

[0067] In this embodiment, the last processed signal and the current processed signal are sequentially processed signals. When the language model receives the last processed signal, the speech recognition model immediately converts the current processed signal, and the language model generates reply information according to the last processed signal at the same time, so as to form a parallel processing architecture of multiple modules to ensure that each text segment can be processed quickly.

[0068] The previous text segment and the current text segment are text segments in a segmentation order. When the speech conversion module 23 receives the reply information generated by the previous text segment, it immediately generates the reply information of the current text segment. That is, when the language processing module 22 generates the reply information according to the current text segment, the speech conversion module 23 synthesizes the speech signal according to the reply information generated by the previous text segment, ensuring the synchronization of speech synthesis and reply information generation, and improving the coherence and naturalness of user interaction.

[0069] In one embodiment, the reply information includes the previous reply information and the current reply information;

[0070] The previous reply information and the current reply information are respectively converted into output voice signals by the voice conversion module 23;

[0071] The output speech signal is synthesized into a synthesized speech signal.

[0072] In this embodiment, the previous reply information and the current reply information are reply information generated in sequence. The synthesized voice signal is audio. The voice conversion module 23 converts the previous reply information and the current reply information into output voice signals respectively. Then, the output voice signals generated by the previous reply information and the current reply information are synthesized into a synthesized voice signal. The output voice signal is synthesized in a manner that the audio can be continuously played in the speaker 30, making the voice interaction smoother.

[0073] In one embodiment, the text segment further includes the next text segment; and the reply information includes the next reply information. When the output voice signal is synthesized into a synthesized voice signal, the language processing module 22 generates the next reply information according to the next text segment.

[0074] In this embodiment, the current text segment and the next text segment are text segments after the segmentation is completed in sequence. When synthesizing the output voice signal into a synthesized voice signal, the language processing module 22 generates the next reply information according to the next text segment. That is, the generation of the reply information is synchronized with the synthesis of the synthesized voice signal.

[0075] More specifically, the speech recognition module 21, the language processing module 22 and the speech conversion module 23 adopt a modular parallel processing architecture, and perform audio generation tasks in a multi-threaded manner to avoid long waits due to thread congestion, which affects the response speed, so that the main program can continue to perform other tasks while waiting for time-consuming tasks such as reply information generation and audio generation, ensuring that each module can run independently and efficiently, reducing the delay perceived by the user, and making the interaction between speech input, processing and output smoother. Among them, multi-threading refers to running multiple threads simultaneously in the same program, and each thread has its own execution task.

[0076] In one embodiment, the step of converting the previous reply information and the current reply information into output voice signals respectively by the voice conversion module 23 includes:

[0077] The audio generation task is asynchronously performed by the voice conversion module 23 to simultaneously obtain the output voice signals of the previous reply information and the current reply information.

[0078] In this embodiment, when the voice conversion module 23 performs the audio generation task, an asynchronous execution mode is adopted; wherein, asynchronous means that the task execution may not be performed in sequence, multiple tasks may be performed simultaneously or the task execution may be staggered. At the same time, the previous reply information and the current reply information are converted into output voice signals, which can greatly improve the response speed.

[0079] In a specific embodiment, the voice conversion module 23 is TTS, Text To Speech. The asynchronous execution of the audio generation task is performed by asynchronous programming. The reply information is processed in parallel by asynchronous processing, making the interaction between voice input, processing and output smoother.

[0080] In one embodiment, before the step of converting the reply information into a synthesized voice signal in real time by the voice conversion module 23 and sending it to the preset speaker 30 for playing, the following steps are included:

[0081] The synthesized speech signal is optimized; wherein the optimization includes but is not limited to adjusting the signal-to-noise ratio, intonation and pronunciation.

[0082] In this embodiment, before the synthesized speech signal is played, the synthesized speech signal is optimized. Specifically, the signal-to-noise ratio of the synthesized speech signal is adjusted to reduce noise interference. The intonation is adjusted to make the speech structure clear and avoid the phenomenon of breaking or prolonging. The pronunciation is adjusted, such as open vowels and nasal sounds, so that the speech played by the speaker 30 is clearer and more coherent.

[0083] In summary, in the specific implementation, first obtain the voice signal input by the voice input device 10, then perform preliminary processing on the voice signal to obtain a processed signal; convert the processed signal through the voice recognition module 21 to obtain text information; segment the text information through the streaming processing mechanism to obtain multiple text segments, and transmit them to the preset language processing module 22; during this period, the text segments are transmitted to the language processing module 22 one by one to generate reply information in real time, and transmit it to the voice conversion module 23. During this period, the language processing module 22 generates reply information according to the text segment, and dynamically adjusts the parameters of the text segment; then monitors the length and processing speed of the text segment to obtain the monitoring result; specifically, identifies the semantic pause point of the text segment, and adjusts the processing speed to adjust the pause time. Then adjust the length and processing speed of the text segment according to the monitoring result. Then send the reply information to the voice conversion module 23. The voice conversion module 23 converts the previous reply information and the current reply information into output voice signals respectively; that is, by asynchronously executing the audio generation task, the output voice signals of the previous reply information and the current reply information are obtained at the same time, and then the output voice signals are synthesized into a synthetic voice signal. When the output voice signal is synthesized into a synthesized voice signal, the language processing module 22 generates the next reply information according to the next text segment, then optimizes the synthesized voice signal, and finally converts the reply information into a synthesized voice signal in real time through the voice conversion module 23, and sends it to the speaker 30 for playback.

[0084] Reference Figure 2 A real-time speech interaction system for improving a large language model includes a speech input device 10, a local device 20 and a speaker 30. The speech input device 10 is used to input a speech signal; the local device 20 includes a speech recognition module 21, a language processing module 22 and a speech conversion module 23. The speech recognition module 21 is used to convert and process the signal to obtain text information; the language processing module 22 is used to generate reply information for each text segment; the speech conversion module 23 is used to convert each text segment into an output speech signal in real time; and the speaker 30 is used to play the reply information.

[0085] In this embodiment, a voice input device 10 for inputting voice signals, a local device 20 and a speaker 30 are included; in a specific embodiment, the voice input device 10 is a microphone. The local device 20 can be a device used in a local network or system, such as a computer or a mobile phone. The local device 20 includes a voice recognition module 21 for converting and processing signals, a language processing module 22 for generating reply information from text segments respectively, and a voice conversion module 23 for converting each text segment into an output voice signal in real time. Furthermore, the language processing module 22 includes at least one language model, such as a large language model. By processing the text segments through the language processing module 22, the accuracy of language recognition can be improved.

[0086] Specifically, integrating the speech recognition module 21, the language processing module 22, and the speech conversion module 23 on the same device can reduce the delay of data transmission, improve the performance of the system in multiple rounds of dialogue, optimize the processing capability of complex tasks, enhance the overall performance of the system, and further improve the response speed and interaction fluency of the system. It not only realizes efficient real-time speech interaction, but also ensures data privacy and efficient operation of the system through localized deployment and modular parallel design. The system ensures that the speech recognition module 21, the language processing module 22, and the speech conversion module 23 run in parallel through asynchronous processing, which can significantly reduce the delay perceived by the user, ensure that speech feedback can be generated efficiently and timely, and realize a real-time speech interaction experience.

[0087] Reference Figure 3 In the embodiment of the present application, a computer device is also provided. The computer device may be a server, and its internal structure may be as follows: Figure 3 As shown. The computer device includes a processor, a memory, a network interface and a database. Among them, the processor of the computer device is used to provide computing and control capabilities. The memory of the computer device includes a non-volatile storage medium and an internal memory. The non-volatile storage medium stores an operating system, a computer program and a database. The internal memory is the operating system, computer program and database in the non-volatile storage medium. The internal memory provides an environment for the operation of the operating system and computer program in the non-volatile storage medium. The database of the computer device is used to store data such as templates, tables, preset fields, etc. The network interface of the computer device is used to communicate with an external terminal through a network connection. When the computer program is executed by the processor, a real-time voice interaction method is implemented, including the following steps:

[0088] Acquire a voice signal input by a preset voice input device 10;

[0089] Performing preliminary processing on the speech signal to obtain a processed signal;

[0090] The signal is converted and processed by a preset speech recognition module 21 to obtain text information;

[0091] The text information is segmented through a preset streaming processing mechanism to obtain a plurality of text segments, and transmitted to a preset language processing module 22; wherein the language processing module 22 includes at least one language processing module 22;

[0092] Generate reply information according to the text segment through the language processing module 22, and dynamically adjust the parameters of the text segment;

[0093] Sending the reply information to the preset voice conversion module 23;

[0094] The reply information is converted into a synthesized voice signal in real time by the voice conversion module 23 and sent to the preset speaker 30 for playing.

[0095] In one embodiment, the step of segmenting the text information by a preset streaming processing mechanism to obtain a plurality of text segments and transmitting the segments to the preset language processing module 22 includes:

[0096] The text segments are transmitted one by one to the language processing module 22 to generate reply information in real time and transmitted to the voice conversion module 23.

[0097] In one embodiment, the step of generating reply information from a text segment by the language processing module 22 and dynamically adjusting the parameters of the text segment includes:

[0098] Monitor the length and processing speed of the text segment and obtain monitoring results;

[0099] Adjust the length and processing speed of the text segment based on the monitoring results.

[0100] In one embodiment, the step of monitoring the length and processing speed of the text segment and obtaining the monitoring result includes:

[0101] Identify semantic pause points in a text segment and adjust processing speed to adjust the pause time.

[0102] In one embodiment, the processed signal includes a previous processed signal and a current processed signal; the text segment includes a previous text segment and a current text segment;

[0103] When the language model receives text information converted from the previous processed signal, the speech recognition module 21 converts the current processed signal;

[0104] When the speech conversion module 23 receives the reply information generated by the previous text segment, the language processing module 22 generates the reply information of the current text segment.

[0105] In one embodiment, the reply information includes the previous reply information and the current reply information;

[0106] The previous reply information and the current reply information are respectively converted into output voice signals by the voice conversion module 23;

[0107] The output speech signal is synthesized into a synthesized speech signal.

[0108] In one embodiment, the text segment further includes the next text segment; and the reply information includes the next reply information. When the output voice signal is synthesized into a synthesized voice signal, the language processing module 22 generates the next reply information according to the next text segment.

[0109] In one embodiment, the step of converting the previous reply information and the current reply information into output voice signals respectively by the voice conversion module 23 includes:

[0110] The audio generation task is asynchronously performed by the voice conversion module 23 to simultaneously obtain the output voice signals of the previous reply information and the current reply information.

[0111] In one embodiment, before the step of converting the reply information into a synthesized voice signal in real time by the voice conversion module 23 and sending it to the preset speaker 30 for playing, the following steps are included:

[0112] The synthesized speech signal is optimized; wherein the optimization includes but is not limited to adjusting the signal-to-noise ratio, intonation and pronunciation.

[0113] Those skilled in the art will understand that Figure 3 The structure shown in is merely a block diagram of a portion of the structure related to the present application solution and does not constitute a limitation on the computer device to which the present application solution is applied.

[0114] An embodiment of the present application further provides a computer storage medium on which a computer program is stored. When the computer program is executed by a processor, a real-time voice interaction method is implemented, including the following steps:

[0115] Acquire a voice signal input by a preset voice input device 10;

[0116] Performing preliminary processing on the speech signal to obtain a processed signal;

[0117] The signal is converted and processed by a preset speech recognition module 21 to obtain text information;

[0118] The text information is segmented through a preset streaming processing mechanism to obtain a plurality of text segments, and transmitted to a preset language processing module 22; wherein the language processing module 22 includes at least one language processing module 22;

[0119] Generate reply information according to the text segment through the language processing module 22, and dynamically adjust the parameters of the text segment;

[0120] Sending the reply information to the preset voice conversion module 23;

[0121] The reply information is converted into a synthesized voice signal in real time by the voice conversion module 23 and sent to the preset speaker 30 for playing.

[0122] In one embodiment, the step of segmenting the text information by a preset streaming processing mechanism to obtain a plurality of text segments and transmitting the segments to the preset language processing module 22 includes:

[0123] The text segments are transmitted one by one to the language processing module 22 to generate reply information in real time and transmitted to the voice conversion module 23.

[0124] In one embodiment, the step of generating reply information from a text segment by the language processing module 22 and dynamically adjusting the parameters of the text segment includes:

[0125] Monitor the length and processing speed of the text segment and obtain monitoring results;

[0126] Adjust the length and processing speed of the text segment based on the monitoring results.

[0127] In one embodiment, the step of monitoring the length and processing speed of the text segment and obtaining the monitoring result includes:

[0128] Identify semantic pause points in a text segment and adjust processing speed to adjust the pause duration.

[0129] In one embodiment, the processed signal includes a previous processed signal and a current processed signal; the text segment includes a previous text segment and a current text segment;

[0130] When the language model receives text information converted from the previous processed signal, the speech recognition module 21 converts the current processed signal;

[0131] When the speech conversion module 23 receives the reply information generated by the previous text segment, the language processing module 22 generates the reply information of the current text segment.

[0132] In one embodiment, the reply information includes the previous reply information and the current reply information;

[0133] The previous reply information and the current reply information are respectively converted into output voice signals by the voice conversion module 23;

[0134] The output speech signal is synthesized into a synthesized speech signal.

[0135] In one embodiment, the text segment further includes the next text segment; and the reply information includes the next reply information. When the output voice signal is synthesized into a synthesized voice signal, the language processing module 22 generates the next reply information according to the next text segment.

[0136] In one embodiment, the step of converting the previous reply information and the current reply information into output voice signals respectively by the voice conversion module 23 includes:

[0137] The audio generation task is asynchronously performed by the voice conversion module 23 to simultaneously obtain the output voice signals of the previous reply information and the current reply information.

[0138] In one embodiment, before the step of converting the reply information into a synthesized voice signal in real time by the voice conversion module 23 and sending it to the preset speaker 30 for playing, the following steps are included:

[0139] The synthesized speech signal is optimized; wherein the optimization includes but is not limited to adjusting the signal-to-noise ratio, intonation and pronunciation.

[0140] In summary, a real-time voice interaction method, system, computer device and storage medium are provided in the embodiments of the present application. The present application can reduce the delay of data transmission, improve the response speed, and improve data privacy protection and reduce potential security risks by deploying a voice recognition module 21, a language processing module 22 and a voice conversion module 23 on a local device 20. A voice processing module including a language model is provided, which can understand and process complex conversations to adapt to different contexts and interaction scenarios, and can recognize and process the input voice signal, thereby improving the accuracy of voice recognition. A modular parallel processing architecture is adopted, and asynchronous multithreading is used to perform audio generation tasks, optimize the workflow, avoid long waiting times due to thread congestion, and continue to perform other tasks while waiting for reply information generation and audio generation, ensuring that each module can run independently and efficiently, generate voice feedback in time, and enhance the interactive experience. At the same time, it can reduce the dependence on computing resources, so that the system can run efficiently on resource-constrained devices, and broaden the application scenarios. The present application can process text information in real time and in segments through a streaming processing mechanism, ensuring that users can quickly obtain feedback after issuing voice commands, and improving the real-time and user satisfaction of voice interaction. It can also dynamically adjust the length and processing speed of text segments based on user interactions to improve the processing of emotions and context.

[0141] Those of ordinary skill in the art can understand that all or part of the processes in the above-mentioned embodiment methods can be completed by instructing related hardware through a computer program, and the computer program can be stored in a non-volatile computer-readable storage medium. When the computer program is executed, it can include the processes of the embodiments of the above-mentioned methods. Among them, any reference to memory, storage, database or other media provided in this application and used in the embodiments may include non-volatile and / or volatile memory. Non-volatile memory may include read-only memory (ROM), programmable ROM (PROM), electrically programmable ROM (EPROM), electrically erasable programmable ROM (EEPROM) or flash memory. Volatile memory may include random access memory (RAM) or external cache memory. By way of illustration and not limitation, RAM is available in a variety of forms, such as static RAM (SRAM), dynamic RAM (DRAM), synchronous DRAM (SDRAM), double data rate SDRAM (SSRSDRAM), enhanced SDRAM (ESDRAM), synchronous link (Synchlink) DRAM (SLDRAM), memory bus (Rambus) direct RAM (RDRAM), direct memory bus dynamic RAM (DRDRAM), and memory bus dynamic RAM (RDRAM), etc.

[0142] It should be noted that, in this article, the terms "include", "comprises" or any other variations thereof are intended to cover non-exclusive inclusion, so that a process, device, article or method including a series of elements includes not only those elements, but also includes other elements not explicitly listed, or also includes elements inherent to such process, device, article or method. In the absence of further restrictions, an element defined by the sentence "includes a ..." does not exclude the presence of other identical elements in the process, device, article or method including the element.

[0143] The above description is only a preferred embodiment of the present application, and does not limit the patent scope of the present application. Any equivalent structure or equivalent process transformation made using the contents of the present application specification and drawings, or directly or indirectly used in other related technical fields, are also included in the patent protection scope of the present application.

Claims

1. A real-time voice interaction method, characterized in that: The following steps are involved: Obtaining a voice signal input by a preset voice input device; Performing preliminary processing on the speech signal to obtain a processed signal; Converting the processed signal through a preset speech recognition module to obtain text information; The text information is segmented through a preset streaming processing mechanism to obtain a plurality of text segments, and transmitted to a preset language processing module; wherein the language processing module includes at least one language processing model; Generate reply information according to the text segment through the language processing module, and dynamically adjust the parameters of the text segment; Sending the reply information to a preset voice conversion module; The reply information is converted into a synthetic voice signal in real time by the voice conversion module and sent to a preset speaker for playing.

2. The real-time voice interaction method according to claim 1, characterized in that: The step of segmenting the text information by a preset streaming processing mechanism to obtain a plurality of text segments and transmitting the segments to a preset language processing module includes: The text segments are transmitted one by one to the language processing module to generate reply information in real time and transmitted to the voice conversion module.

3. The real-time voice interaction method according to claim 1, characterized in that: The step of generating reply information from the text segment by the language processing module and dynamically adjusting the parameters of the text segment includes: Monitoring the length and processing speed of the text segment to obtain monitoring results; The length and processing speed of the text segment are adjusted according to the monitoring result.

4. The real-time voice interaction method according to claim 3, characterized in that: The step of monitoring the length and processing speed of the text segment and obtaining the monitoring result includes: Semantic pause points of the text segment are identified, and the processing speed is adjusted to adjust the pause time.

5. The real-time voice interaction method according to claim 1, characterized in that: The processing signal includes a previous processing signal and a current processing signal; the text segment includes a previous text segment and a current text segment; When the language model receives the text information converted from the last processed signal, the speech recognition module converts the current processed signal; When the speech conversion module receives the reply information generated by the previous text segment, the language processing module generates the reply information of the current text segment.

6. The real-time voice interaction method according to claim 5, characterized in that: The reply information includes the previous reply information and the current reply information; The previous reply information and the current reply information are respectively converted into output voice signals by the voice conversion module; The output speech signal is synthesized into the synthesized speech signal.

7. The real-time voice interaction method according to claim 6, characterized in that: The text segment also includes a next text segment; the reply information includes a next reply information; When the output voice signal is synthesized into the synthesized voice signal, the language processing module generates the next reply information according to the next text segment.

8. The real-time voice interaction method according to claim 6, characterized in that: The step of converting the previous reply information and the current reply information into output voice signals respectively by the voice conversion module includes: The audio generation task is asynchronously performed by the voice conversion module to simultaneously obtain the output voice signals of the previous reply information and the current reply information.

9. The real-time voice interaction method according to claim 1, characterized in that: Before the step of converting the reply information into a synthetic voice signal in real time by the voice conversion module and sending it to a preset speaker for playing, the step includes: The synthesized speech signal is optimized; wherein the optimization includes but is not limited to adjusting the signal-to-noise ratio, intonation and pronunciation.

10. A real-time voice interaction system, characterized in that: It includes a voice input device, a local device and a speaker, wherein the voice input device is used to input voice signals; the local device includes a voice recognition module, a language processing module and a voice conversion module, the voice recognition module is used to convert and process signals to obtain text information; the language processing module is used to generate reply information from text segments respectively; the voice conversion module is used to convert each of the text segments into an output voice signal in real time; and the speaker is used to play the reply information.

11. A computer device comprising a memory and a processor, wherein a computer program is stored in the memory, wherein: When the processor executes the computer program, the steps of a real-time voice interaction method according to any one of claims 1 to 9 are implemented.

12. A computer storage medium storing a computer program, characterized in that: When the computer program is executed by a processor, the steps of a real-time voice interaction method according to any one of claims 1 to 9 are implemented.

Citation Information

Patent Citations

  • Incremental semantic processing method

    CN111862980A

  • Intelligent glasses streaming voice dialogue interaction system and method based on large language model

    CN119091878A

  • Audio interaction method and device of generative language model based on dialogue control module, medium, program product and terminal

    CN119107952A

  • Voice conversion device, voice conversion system, and computer program product

    US20200279550A1

  • Voice response processing method and apparatus based on artificial intelligence, device, and medium

    WO2021169615A1