A Construction Method for a Real-Time Streaming Voice Intelligent Q&A Service System

By decoupling ASR, LLM and TTS services and building an asynchronous streaming framework, the low latency, high accuracy and scalability of voice Q&A systems are achieved, solving the shortcomings of existing systems in open Q&A and audio output latency.

CN119719438BActive Publication Date: 2025-06-10ROCK AI
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510220327.8
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-02-27
Publication Date
2025-06-10
Estimated Expiration
2045-02-27

AI Technical Summary

Technical Problem

When the existing voice Q&A system handles complex and changeable open Q&A, the NLU model has limited generalization capabilities, resulting in insufficient answer accuracy and diversity. At the same time, the autoregressive reasoning mechanism of the large language model leads to a long audio output delay, affecting the user experience.

Method used

By decoupling ASR, LLM and TTS services into independent modular services, an asynchronous streaming processing framework is built, and the "think while talking" strategy is adopted to achieve parallel promotion of text generation and speech synthesis. The dynamic sentence splitter splitter splits the text output by the LLM in real time and uses an asynchronous thread pool to submit it to the TTS service in parallel to generate an audio stream.

Benefits of technology

It significantly improves the real-time and flexibility of the system, reduces the audio stream return delay, improves the system throughput and resource utilization efficiency, and forms a low-latency, high-accuracy, and scalable intelligent voice interaction system.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119719438B_ABST
    Figure CN119719438B_ABST
Patent Text Reader

Abstract

The present application discloses a method for constructing a real-time streaming voice intelligent question-answering service system, including: receiving input voice data; invoking an independent speech recognition service to convert the voice data into input text; constructing a prompt for a large language model based on the input text, initiating an LLM streaming request to an independent large language model service, and obtaining a streaming text answer generated by the LLM in real time; performing real-time segmentation on the streaming text answer through a dynamic sentence segmenter to generate multiple clauses; parallelly invoking an independent speech synthesis service for each clause to convert the text into audio data blocks; combining the audio data blocks in the generation order into streaming audio data and returning it to the client for playback in real time. The present invention significantly improves the real-time performance and flexibility of the voice question-answering system by decoupling the ASR, LLM, and TTS services and combining an asynchronous streaming framework with the "thinking and speaking while" strategy.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of artificial intelligence technology, and particularly relates to a method for constructing a real-time streaming voice intelligent question-answering service system. Background Art

[0002] With the rapid development of artificial intelligence technology, voice question-answering systems have become one of the core applications of human-computer interaction and are widely used in fields such as smart homes, customer service, and mobile terminals. Traditional voice question-answering systems usually rely on the following technologies:

[0003] 1. Automatic Speech Recognition (ASR): Converts the user's voice input into text data.

[0004] 2. Natural Language Processing (NLP): Understands and processes the converted text data to generate corresponding answers. NLP technology includes Natural Language Understanding (NLU) and Natural Language Generation (NLG).

[0005] 3. Speech synthesis technology: Converts the generated text answer into voice output and returns it to the user.

[0006] The typical process of a traditional voice question-answering system is as Figure 1 shown. The voice input is first converted into text by an ASR model, then the question-answering result is obtained through an NLU model, and finally the voice output is generated through a Text-to-Speech (TTS) model. Such a solution has high real-time performance due to its small model size and serial process, but its drawback is that the generalization ability of the NLU model is limited and it is only applicable to a single scenario. Different scenarios require training different NLU models, so it is difficult for such a solution to adapt to complex and changing open-ended question-answering requirements.

[0007] To achieve more intelligent and generalized voice question-answering, the industry uses large language models (LLMs) to replace the NLU model for question understanding and answering. The use of LLMs can make the question-answering system more intelligent and personalized, enabling cross-scenario general question-answering and significantly improving the accuracy and diversity of answers. However, the autoregressive inference mechanism of LLMs results in a long time delay in obtaining a complete answer, and the TTS conversion cannot be started until the complete text is generated, leading to a significant increase in the overall delay of audio output and seriously affecting the fluency of the user experience.

[0008] Some solutions attempt to adopt multimodal large models, such asFigure 2 As shown directly using audio as the input of the large model and streaming to generate the response audio, but the single timbre generated by it, the strong coupling between the model and the audio codec module, and the high fine-tuning cost limit the practical application flexibility of this technology.

[0009] In addition, in the existing technologies, ASR, NLU / LLM, and TTS services mostly adopt a tightly coupled architecture. Replacing or upgrading the model requires reconstructing the entire system, making it difficult to adapt to the differentiated requirements of different business scenarios for model performance and resource occupancy.

[0010] Therefore, how to achieve low-latency and high-concurrency streaming voice interaction while ensuring the intelligence of question answering, and improve the modularity and scalability of the system architecture has become a technical problem to be solved urgently in this field. Summary of the Invention

[0011] The purpose of the present invention is to provide a construction method for a real-time streaming voice intelligent question answering service system to solve the problems raised in the above technical background.

[0012] To achieve the above purpose, the present invention adopts the following technical solutions:

[0013] The present application provides a construction method for a real-time streaming voice intelligent question answering service system, including the following steps:

[0014] Step S1: Receive the voice data input by the user;

[0015] Step S2: Invoke an independent automatic speech recognition service (ASR Server) to convert the voice data into input text;

[0016] Step S3: Based on the input text, construct a prompt for a large language model (LLM), initiate an LLM streaming request to an independent large language model service (LLM Server), and obtain the streaming text answer generated by the LLM in real time;

[0017] Step S4: Real-time segment the streaming text answer through a dynamic sentence segmenter to generate multiple clauses;

[0018] Step S5: Parallelly invoke an independent text-to-speech service (TTS Server) for each clause to convert the text into audio data chunks (chunks);

[0019] Step S6: Combine the audio data chunks into streaming audio data in the order of generation and return it to the client for playback in real time;

[0020] Among them, the ASR Server, LLM Server, and TTS Server are independent and decoupled services, and the LLM Server and TTS Server support streaming output.

[0021] In a preferred embodiment, the working rules of the dynamic sentence splitter include:

[0022] Split the streaming text generated by the LLM according to punctuation marks or semantic boundaries;

[0023] Limit the maximum length of the first clause, and if it exceeds the threshold, force truncation;

[0024] If the clause length is less than the minimum length threshold, merge it with the subsequent clause and then output.

[0025] In a more preferred embodiment, the dynamic sentence splitter performs the following operations after each iteration of text generation by the LLM:

[0026] Record the termination index position of the current text traversal to avoid repeated processing of the segmented text;

[0027] Merge the clauses that do not reach the minimum length threshold into the subsequent text.

[0028] In a preferred embodiment, the step S5 includes:

[0029] Create an independent TTS request thread for each clause, and control the concurrency number through a thread pool;

[0030] Each TTS request thread stores the generated audio data blocks into the asynchronous audio queue (audio_chunks-Queue) in sequence;

[0031] Real-time detect the asynchronous audio queue, extract the audio data blocks in the enqueue order and return them to the client.

[0032] In a more preferred embodiment, the step S5 further includes a termination control step:

[0033] When a client termination request is received, trigger the termination event stop-event to terminate the current LLM streaming request, dynamic sentence splitting, and all TTS request threads.

[0034] In a more preferred embodiment, the execution process of each clause's TTS request thread in the step S5 includes:

[0035] Create an independent asynchronous queue audio-chunk-j for the clause sj, and add audio-chunk-j to the asynchronous audio queue (audio_chunks-Queue);

[0036] Send a streaming conversion request for clause sj to the TTS Server, and store the returned audio data chunks in audio-chunk-j in order;

[0037] When the audio stream conversion of clause sj is completed, insert the clause end identifier None into audio-chunk-j;

[0038] After each audio data chunk is stored in audio-chunk-j, detect the termination event stop-event. If the termination event stop-event is triggered, terminate the current TTS request thread.

[0039] In a preferred embodiment, in step S3, the request process of the LLM Server includes:

[0040] Construct a prompt containing retrieval-augmented generation (RAG) technology based on the business scenario;

[0041] Receive the text fragments generated by the LLM in real time through an asynchronous streaming generator (async llm stream generator).

[0042] In a preferred embodiment, in step S6, the implementation method of returning the streaming audio data includes:

[0043] (a) During the process of the LLM generating text, detect in real time whether there are audio data chunks in the asynchronous audio queue. If there are, return them in the enqueue order;

[0044] (b) After the LLM generation is completed, insert the task end identifier end into the asynchronous audio queue, and continue to process the remaining audio data chunks in the asynchronous audio queue and return them in the enqueue order until the task end identifier end is detected or the asynchronous audio queue is emptied.

[0045] In a more preferred embodiment, step (a) specifically includes the following steps:

[0046] Step a.1: Before each iteration of the LLM streaming text generation, create an empty temporary storage variable audio_chunks to cache the audio data of the current sub-queue being processed;

[0047] Step a.2: After each iteration of the LLM generating text, detect whether the asynchronous audio queue is non-empty:

[0048] If the asynchronous audio queue is non-empty and audio_chunks is empty, dequeue the first sub-queue audio-chunk-j from the asynchronous audio queue and assign it to audio_chunks;

[0049] Step a.3: Determine whether audio_chunks is non-empty and contains unreturned audio data chunks:

[0050] If the condition is satisfied, extract all audio data chunks in audio_chunks in sequence and return them to the client in real time;

[0051] If the extracted audio data chunks are empty, reset audio_chunks to be empty and retrieve the next sub-queue from the asynchronous audio queue again;

[0052] If audio_chunks is empty or there is no extractable data, continue the next LLM iteration.

[0053] Further, the step (b) specifically includes the following steps:

[0054] Step b.1: After the LLM generation ends, continuously detect the termination event (stop-event) and the status of the asynchronous audio queue:

[0055] If the termination event is triggered, directly exit the processing flow;

[0056] If audio_chunks is empty, dequeue a new sub-queue audio-chunk-j from the asynchronous audio queue and assign it to audio_chunks;

[0057] If the asynchronous audio queue is empty, asynchronously wait for a preset time and then detect again;

[0058] Step b.2: When audio_chunks is the end identifier end, terminate the processing flow;

[0059] Step b.3: Extract audio data chunks from audio_chunks in sequence and return them to the client, and at the same time loop to execute:

[0060] If the audio data chunks are empty, reset audio_chunks to be empty and detect the queue again;

[0061] If audio_chunks is empty, asynchronously wait and then continue to extract;

[0062] Immediately interrupt the extraction operation after detecting the termination event.

[0063] In a preferred embodiment, the method is implemented through an asynchronous framework, including:

[0064] Use FastAPI to build an asynchronous service interface;

[0065] Design the ASR request, LLM streaming generation, and TTS request as asynchronous non-blocking operations.

[0066] Compared with the prior art, the technical solution of the present invention has the following beneficial effects:

[0067] The present invention discloses a method for constructing a real-time streaming voice intelligent question-answering service system. By decoupling automatic speech recognition (ASR), large language model (LLM), and text-to-speech (TTS) services into independent modular services, an asynchronous streaming processing framework is constructed, and the "thinking while speaking" strategy is adopted to achieve parallel advancement of text generation and speech synthesis, significantly improving the real-time performance and flexibility of the system: By using a dynamic sentence splitter to split the text streamed out by the LLM in real time, clauses are generated based on punctuation marks, semantic boundaries, and the first sentence length limit rules, and the clauses are submitted to the TTS service in parallel using an asynchronous thread pool to generate an audio stream, breaking the traditional serial processing mode.

[0068] The present invention also combines hierarchical management of an asynchronous audio queue (audio_chunks-Queue) and the design of the FastAPI asynchronous framework to achieve non-blocking operations for ASR requests, LLM generation, and TTS conversion, improving the system throughput and resource utilization efficiency; at the same time, integrating retrieval-augmented generation (RAG) technology to optimize the LLM input prompt, enhancing the accuracy and scenario adaptability of open-domain question answering, and ensuring the coherence of the voice output through the fault tolerance mechanism (minimum length merging, termination index recording) of the dynamic splitter, finally forming a low-latency, high-accuracy, and scalable intelligent voice interaction system, suitable for high-concurrency and strong real-time scenarios such as smart homes and real-time customer service. BRIEF DESCRIPTION OF THE DRAWINGS

[0069] In order to more clearly illustrate the specific embodiments of the present invention or the technical solutions in the prior art, the following will briefly introduce the drawings required for use in the description of the specific embodiments or the prior art. Obviously, the drawings in the following description are some embodiments of the present invention. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.

[0070] Figure 1 It is a schematic diagram of the technical solution of a traditional voice question-answering system;

[0071] Figure 2 It is an architecture diagram of using a multimodal large model for voice question answering in the prior art;

[0072] Figure 3 It is the intelligent question-answering service architecture diagram adopted in the present application;

[0073] Figure 4 It is the streaming intelligent voice question-answering service framework diagram of the present application;

[0074] Figure 5It is a framework diagram of the streaming voice intelligent question-answering service system provided by this application;

[0075] Figure 6 It is a schematic flowchart of step S4-3 in the embodiment of this application;

[0076] Figure 7 It is a schematic flowchart of step S5 in the embodiment of this application. Detailed implementation manners

[0077] In order to make the above and other features and advantages of the present invention clearer, the present invention will be further described below with reference to the accompanying drawings. It should be understood that the specific embodiments given herein are for the purpose of explaining to those skilled in the art and are merely exemplary, not restrictive.

[0078] In addition, the terms "first" and "second" are only used for descriptive purposes and cannot be understood as indicating or implying relative importance or implicitly specifying the quantity of the indicated technical features. Thus, the features defined with "first" and "second" may explicitly or implicitly include at least one of such features. In the description of the present invention, "a plurality" means at least two, such as two, three, etc., unless otherwise specifically defined.

[0079] The present invention designs an intelligent question-answering system using a modular service architecture, decoupling speech recognition (ASR), large language model (LLM), and text-to-speech (TTS) services into independent service modules, supporting flexible selection of different model combinations for each module, while retaining high concurrency characteristics and being compatible with existing model services. The service architecture is as Figure 3 shown. To improve the real-time performance of audio responses, based on the streaming generation characteristics of LLM and TTS, the present invention adopts a "thinking and speaking simultaneously" strategy. Through the parallel design of LLM service and TTS service, while the LLM generates text, part of the content is dynamically segmented and submitted to the TTS service for parallel generation of audio streams, significantly reducing the end-to-end delay. Further, the retrieval-augmented generation (RAG, Retrieval-Augmented Generation) technology can be integrated into the LLM service to enhance the construction of prompts with domain knowledge, improving the accuracy and scenario adaptability of question answering.

[0080] Embodiment:

[0081] The streaming intelligent voice question-answering service system (Streaming Intelligent Voice Question and Answer, SIVQA) constructed by the present invention is based on three decoupled independent service modules of ASR, LLM, and TTS, where LLM and TTS support streaming output, and the services do not interfere with each other and have high throughput. The SIVQA service can be constructed using any service that meets the above conditions.

[0082] The streaming voice intelligent question - answering service system (SIVQA) provided by the present invention is as follows Figure 4 shown. Its input is voice data, and the output is streaming audio data which is returned to the client. The client can receive the voice data and play it directly while receiving it. After receiving the voice input, the SIVQA service calls an independent automatic speech recognition service (ASR Server) to convert the voice data into input text. Based on the input text, prompts for a large - language model (LLM) can be constructed according to different business scenarios, and a streaming request is sent to an independent large - language model service (LLM Server). The LLM service can choose a general large - model or an enhanced model integrated with RAG technology to generate text answers in a streaming manner. In order to obtain the voice answer to the question faster, the present invention designs a dynamic sentence splitter. According to the results returned by the LLM in a streaming manner, the entire answer text of the LLM is split into multiple clauses in real - time. The dynamic sentence splitter splits the LLM answer text based on rules (such as punctuation marks, semantic boundaries). At the same time, to shorten the latency of the first audio return, the maximum length of the first clause is restricted (for example: 10 Chinese characters or 10 English words), thereby reducing the latency of the language streaming output. According to the "thinking while speaking" strategy, when the LLM outputs an answer in a streaming manner, the answer text is dynamically split. For each split clause sj, a request is sent in parallel to an independent text - to - speech service (TTS Server) to convert the corresponding text into voice data. At the same time, the SIVQA service will detect the return situation of the TTS requests in parallel, and in the order of the TTS requests, combine the audio results returned by each request into audio data chunks one by one, and stream each segment of audio data to the client in real - time.

[0083] Specifically, the implementation process of the construction method of the real - time streaming voice intelligent question - answering service system provided by the present invention is as follows.

[0084] The present invention uses FastAPI to build an asynchronous service interface. According to its asynchronous characteristics, the blocking parts in the system such as automatic speech recognition (ASR) requests, large - language model (LLM) streaming generation, and text - to - speech (TTS) requests are designed as asynchronous non - blocking operations through multi - threading and asynchronous functions, so as to improve the concurrency of the service system and build an asynchronous streaming processing framework.

[0085] The design of the entire asynchronous service is as follows Figure 5 shown, and the specific implementation steps are as follows:

[0086] 1. Service initialization stage:

[0087] S1: Create a TTS request thread pool (TTS post ThreadPool) to control the number of running child threads and the concurrency of TTS concurrent requests.

[0088] S2: Create a service route and a voice API interface to define the asynchronous endpoints for voice input and audio stream output.

[0089] 2. Service running stage:

[0090] Step 1: Receive the voice data (audio-data) input by the user.

[0091] Step 2: Asynchronously call an independent speech recognition service (ASR Server) to convert the received voice data into input text.

[0092] Step 3: Based on the input text, construct a prompt for the large language model (LLM), initiate an LLM streaming request to an independent large language model service (LLM Server), and obtain the streaming text answer generated by the LLM in real time through an asynchronous streaming generator (async llm streamgenerator).

[0093] Step 4: Real-time segment the streaming text generated by the LLM through a dynamic sentence segmenter to generate multiple clauses. The segmentation rules include:

[0094] Segment according to punctuation marks or semantic boundaries;

[0095] Limit the maximum length (L1) of the first clause and truncate it forcibly when it exceeds;

[0096] If the clause length is less than the minimum threshold (L2), merge it with the subsequent clauses and then output.

[0097] Step 5: For each clause, asynchronously call an independent text-to-speech service (TTS Server) in parallel to convert the text into audio data chunks (chunks). Specifically, it includes:

[0098] Create an independent TTS request thread for each clause and control the concurrency through the thread pool;

[0099] Each TTS request thread stores the generated audio data chunks into an asynchronous audio queue (audio_chunks-Queue) in order;

[0100] Real-time detect the asynchronous audio queue, extract the audio data chunks in the enqueue order and return them to the client.

[0101] Step 6: Combine the generated audio data chunks into streaming audio data in the generation order and return it to the client for playback in real time.

[0102] Among them, the ASR Server, LLM Server, and TTS Server are independent and decoupled services, and the LLM Server and TTS Server support streaming output.

[0103] The following is a detailed introduction:

[0104] S1: Receive the voice input data audio-data.

[0105] S2: Send audio-data to the ASR service in an asynchronous request manner to avoid blocking the service interface and wait for it to be converted into text text.

[0106] S3: Concatenate the converted text text into an LLM prompt (prompt), create a termination event (stop-event) to interrupt the child thread, initialize an asynchronous audio queue (audio_chunks-Queue) to store the audio data chunks returned by TTS, and create a dynamic sentence splitter for sentence splitting.

[0107] S4: Create a streaming asynchronous LLM request, and the specific steps are as follows:

[0108] S4-1: Send an LLM service request using the prompt to obtain an asynchronous streaming generator (async llmstream generator).

[0109] S4-2: Iterate over the asynchronous streaming generator to continuously obtain the streaming text output by the LLM. At the same time, perform dynamic sentence splitting based on the streaming text output by the LLM and convert the split sub-clauses into audio streams. The specific sub-steps are as follows:

[0110] S4-2-1: Each time the asynchronous streaming generator is iterated, the text generated by the LLM gradually becomes longer. At this time, perform a dynamic sentence split to split out the newly generated sub-clause sj (0 < j <= n); if no complete sub-clause is split out, return an empty string.

[0111] S4-2-2: After each iteration, check whether the termination event (stop-event) is triggered. If it is triggered, terminate the iteration and exit.

[0112] S4-2-3: For each complete sub-clause sj generated, immediately create a child thread task TTS-post-j and submit it to the thread pool (TTS ThreadPool) for execution to convert the text of the sub-clause sj into audio data.

[0113] S4-2-3-1: In the TTS-post-j task, first create an independent asynchronous queue audio-chunk-j for clause sj to store the audio data returned by the request, and add it to the asynchronous audio queue (audio_chunks-Queue).

[0114] S4-2-3-2: In the TTS-post-j task, use clause sj to send a streaming conversion request to the TTS Server, and sequentially store the returned audio data chunks into audio-chunk-j;

[0115] S4-2-3-3: After the conversion of the audio streaming data returned in step S4-2-3-2 is completed, insert the end identifier None into audio-chunk-j to mark the completion of the generation of this segment of audio.

[0116] S4-2-3-4: After each audio data chunk is stored in audio-chunk-j, detect whether the termination event (stop-event) is triggered. If triggered, terminate all operations of the current TTS request thread.

[0117] S4-3: During each iteration of the asynchronous streaming generator, simultaneously detect whether there are audio data chunks in the asynchronous audio queue (audio_chunks-Queue). If there are, return the audio data chunks to the service requester in the enqueue order.

[0118] S4-4: After iterating through the asynchronous streaming generator, insert the end identifier end into the asynchronous audio queue (audio_chunks-Queue) to mark the completion of the LLM text generation task.

[0119] S5: After the above step S4 is iterated, if there are still unreturned audio data chunks in the asynchronous audio queue (audio_chunks-Queue), continue to return each audio data chunk (chunk) in the queue in the enqueue order.

[0120] S6: Construct an asynchronous termination background function. When the client request triggers termination, activate this background function. At this time, the termination event (stop-event) is triggered, and all subtask threads in the current request are terminated.

[0121] In the above content, steps S4-3 and S5 together complete the orderly return of all audio data chunks (chunk). However, due to different stages, their working methods are different.

[0122] (1) Step S4-3:

[0123] In the iterative asynchronous streaming generator stage, each iteration of the text takes a certain amount of time. During this period, the audio_chunks-Queue may receive some audio data chunks. Therefore, some audio data chunks will be returned in an orderly manner at this time, without interrupting the LLM iteration. The detailed steps are as follows (see Figure 6 ):

[0124] S4-3-0: Create an empty variable audio_chunks before iteration to orderly retrieve the audio-chunk-j queue.

[0125] S4-3-1: After each iteration, check if the length of the audio_chunks-Queue is greater than 0. If it is, there is a text-to-audio task. If audio_chunks is empty, dequeue an audio-chunk from the audio_chunks-Queue and assign it to audio_chunks.

[0126] S4-3-2: Check if audio_chunks is empty and if the length of the audio_chunks queue is greater than 0.

[0127] S4-3-2-1: If audio_chunks is not empty and the length of the audio_chunks queue is greater than 0, pop all the audio data chunk data in the current audio_chunks queue in order and return it to the client, then proceed to the next iteration.

[0128] If the popped audio data chunk is a None value, it means the text-to-audio task for this segment is completed. Assign None to audio_chunks. In the next iteration, in step S4-3-1, dequeue an audio-chunk from the audio_chunks-Queue again and assign it to audio_chunks.

[0129] S4-3-2-2: If audio_chunks is empty or the length of the audio_chunks queue is 0, continue to the next iteration.

[0130] (2) Step S5:

[0131] At this stage, the LLM has completed the generation of all texts. At this time, all clauses have completed the creation of the TTS-post-j task. S4-3 may have returned some audio data chunks, and there may still be some audio data in audio_chunks that has not been arranged and returned. Therefore, in this stage, the unreturned audio data will continue to be arranged in order of enqueue and returned to the client. The detailed steps are as follows (see Figure 7):

[0132] S5-1: Detect whether the stop-event is triggered. If triggered, directly exit step S5. Otherwise, continue.

[0133] S5-2: Judge whether audio_chunks is None:

[0134] S5-2-1: If audio_chunks is None, dequeue an audio-chunk from the audio_chunks-Queue and assign it to audio_chunks. Otherwise, start executing from step S5-2-4.

[0135] S5-2-2: If the audio_chunks-Queue is empty in the previous step, asynchronously sleep or wait for 10 milliseconds, and repeat from step S5-1 again. Otherwise, continue to execute downward.

[0136] S5-2-3: If audio_chunks is the character 'end', directly exit step S5. Otherwise, continue.

[0137] S5-2-4: Orderly retrieve all audio data chunks from audio_chunks and return them to the client. Details are as follows:

[0138] S5-2-4-1: Detect whether the stop-event is triggered. If triggered, directly exit step S5-2-4. Otherwise, continue.

[0139] S5-2-4-2: Judge whether the audio_chunks queue is empty: If it is empty, asynchronously sleep or wait for 10 milliseconds, and repeat from step S5-2-4-1 again; otherwise, continue to execute downward;

[0140] S5-2-4-3: Pop an audio data chunk from audio_chunks;

[0141] S5-2-4-4: If the popped audio data chunk is None, it means that the audio data in audio_chunks has been retrieved, set audio_chunks to None, and then start executing from step S5-1; otherwise, return the audio data chunk to the client;

[0142] S5-2-4-5: Continue to execute from S5-2-4-1 until the audio data chunk is None, and then break out of the loop of S5-2-4.

[0143] The text generated by the LLM contains punctuation marks, so the dynamic sentence splitter uses common Chinese and English punctuation marks as the basis for sentence division.

[0144] When segmenting sentences, the text generated by LLM is first preprocessed, and each Chinese character and English word is segmented as a separate character. After each traversal of the characters, the subscript index position idx of the last character is recorded. The next time the text generated by LLM is traversed, it starts from the idx character index position to avoid repeated traversal of the previously generated text content; at the same time, in order to avoid the high latency of the first audio data return, the maximum length of the first clause is limited to L1, and it is directly truncated when it exceeds L1. In addition, the minimum length of each clause is also limited to L2. If the clause length is less than L2, it is merged with the next sentence and returned.

[0145] In summary, the present invention significantly improves the real-time performance and flexibility of the voice question-answering system by decoupling ASR, LLM and TTS services, combining the asynchronous streaming framework with the "think while speaking" strategy. The dynamic sentence segmenter and parallel TTS processing effectively reduce the audio stream return delay, and the asynchronous queue management ensures the orderliness of the data and the system throughput.

[0146] The specific embodiments of the present invention are described in detail above, but they are only examples, and the present invention is not limited to the specific embodiments described above. For those skilled in the art, any equivalent modifications and substitutions made to the present invention are also within the scope of the present invention. Therefore, the equalization changes and modifications made without departing from the spirit and scope of the present invention should be included in the scope of the present invention.

Claims

1. A method for constructing a real-time streaming voice intelligent question-answering service system, characterized in that: The steps include: Step S1: receiving voice data input by a user; Step S2: calling an independent speech recognition service ASR Server to convert the speech data into input text; Step S3: construct a prompt of the large language model LLM based on the input text, initiate an LLM streaming request to the independent large language model service LLM Server, and obtain the streaming text answer generated by the LLM in real time; Step S4: segmenting the streaming text answer in real time by a dynamic sentence segmenter to generate multiple clauses; Step S5: calling an independent speech synthesis service TTS Server for each clause in parallel to convert the text into an audio data block; Step S6: combining the audio data blocks into streaming audio data in the order of generation, and returning the streaming audio data to the client for playback in real time; The ASR Server, LLM Server and TTS Server are independent services that are decoupled from each other, and the LLM Server and TTS Server support streaming output; The step S5 comprises: Create an independent TTS request thread for each clause, and control the number of concurrent requests through the thread pool; Each TTS request thread stores the generated audio data blocks in the asynchronous audio queue in sequence; Detect the asynchronous audio queue in real time, extract the audio data blocks in the order of entering the queue and return them to the client; In step S6, the implementation method of returning the streaming audio data includes: (a) During the text generation process of LLM, it is detected in real time whether there is an audio data block in the asynchronous audio queue. If so, it is returned in the order in which it was queued; (b) After the LLM generation is completed, insert the task end identifier end into the asynchronous audio queue, and continue to process the remaining audio data blocks in the asynchronous audio queue and return them in the order in which they were queued until the task end identifier end is detected or the asynchronous audio queue is cleared; Wherein, the step (a) specifically comprises the following steps: Step a.1: Before each iteration of LLM streaming text generation, create an empty temporary storage variable audio_chunks to cache the audio data of the currently processed subqueue; Step a.2: After each iteration of LLM generates text, check whether the asynchronous audio queue is non-empty: If the asynchronous audio queue is not empty and audio_chunks is empty, dequeue the first subqueue audio-chunk-j from the asynchronous audio queue and assign it to audio_chunks; Step a.3: Determine whether audio_chunks is not empty and contains unreturned audio data chunks: If the conditions are met, all audio data blocks in audio_chunks are extracted in order and returned to the client in real time; If the extracted audio data block is empty, reset audio_chunks to empty and get the next subqueue from the asynchronous audio queue again; If audio_chunks is empty or there is no extractable data, continue to the next LLM iteration.

2. The method according to claim 1, characterized in that The working rules of the dynamic sentence segmenter include: Segment the streaming text generated by LLM according to punctuation or semantic boundaries; Limit the maximum length of the first clause. If it exceeds the threshold, it will be truncated. If the length of a clause is less than the minimum length threshold, it is merged with the subsequent clauses and output.

3. The method according to claim 1, characterized in that The step S5 also includes a termination control step: When a termination request is received from the client, the termination event stop-event is triggered to terminate the current LLM streaming request, dynamic sentence segmentation and all TTS request threads.

4. The method according to claim 1, characterized in that: The TTS request thread execution process of each clause in step S5 includes: Create an independent asynchronous queue audio-chunk-j for clause sj, and add audio-chunk-j to the asynchronous audio queue; Send a streaming conversion request for clause sj to the TTS Server, and store the returned audio data blocks in order into audio-chunk-j; When the audio stream conversion of clause sj is completed, insert the clause end identifier None into audio-chunk-j; Each time an audio data block is stored in audio-chunk-j, a stop-event is detected. If the stop-event is triggered, the current TTS request thread is terminated.

5. The method according to claim 1, characterized in that In step S3, the request process of the LLM Server includes: Build prompts containing search enhancement generation technology based on business scenarios; Receive LLM-generated text snippets in real time via an asynchronous streaming generator.

6. The method according to claim 1, characterized in that The step (b) specifically comprises the following steps: Step b.1: After LLM generation is completed, the stop-event and asynchronous audio queue status are detected cyclically: If the stop-event is triggered, the processing flow will be exited directly; If audio_chunks is empty, dequeue a new subqueue audio-chunk-j from the asynchronous audio queue and assign it to audio_chunks; If the asynchronous audio queue is empty, wait for the preset time and then re-detect; Step b.2: When audio_chunks is the end identifier end, the processing flow is terminated; Step b.3: Extract audio data blocks from audio_chunks in order and return them to the client, while looping: If the audio data block is empty, reset audio_chunks to empty and recheck the queue; If audio_chunks is empty, continue extracting after asynchronous waiting; The extraction operation is interrupted immediately after the stop-event is detected.

7. The method according to claim 1, characterized in that The method is implemented through an asynchronous framework and includes: Use FastAPI to build asynchronous service interfaces; ASR requests, LLM streaming generation, and TTS requests are designed as asynchronous non-blocking operations.

Citation Information

Patent Citations

  • 3D virtual human real-time interaction system and implementation method

    CN118552672A

  • Multi-language 3D digital human interaction method adopting GPT

    CN119376586A