Voice broadcasting methods, computer equipment and readable storage media
Patent Information
- Application Number
- CN202611176952.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2026-08-05
- Publication Date
- 2026-09-01
AI Technical Summary
现有方案通常采用终端人声活动检测触发打断或云端语义裁决触发打断,但二者难以兼顾实时性与准确性;快慢双模型方案中,慢模型仅作为离线精修,未与在线打断机制动态耦合;用户插话触发的打断与模型自我纠正触发的修正走两套独立逻辑,易产生时序冲突与工程冗余
[0014] Unlike related technologies, this application sends the recognized text obtained from streaming speech recognition to a first model and a second model in parallel. The first model generates and plays the response text first to reduce response latency, while the second model performs deep understanding in the background and produces correction text to ensure response accuracy. During the playback process, the terminal synchronizes the playback progress to the cloud in real time, so that the cloud can know the played position and the remaining content when making a decision. When the terminal detects that the voice or the cloud determines that the correction text and the response text are inconsistent, an interruption is triggered. The terminal immediately pauses the playback and records the progress. The cloud combines the recognized text with the current playback progress to determine whether it is a valid interruption. If it is a valid interruption, a new playback audio is generated based on the new response text or correction text and continued at the sentence boundary. If it is an invalid interruption, it seamlessly resumes from the preserved playback progress position, thereby improving the accuracy of interruption judgment and the playback quality while ensuring low response latency.
Smart Images

Figure CN122676804A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of speech processing technology, and in particular to a speech broadcasting method, computer device, and readable storage medium. Background Technology
[0002] In full-duplex voice interaction systems, users can interrupt voice broadcasts at any time. Achieving accurate interruption decisions while maintaining low response latency is a major challenge currently facing voice broadcasting. Existing solutions typically use terminal voice activity detection to trigger interruptions or cloud-based semantic adjudication to trigger interruptions, but both struggle to balance real-time performance and accuracy. In fast-slow dual-model solutions, the slow model is only used for offline refinement and is not dynamically coupled with the online interruption mechanism. User-triggered interruptions and model self-correction triggers follow two independent logics, easily leading to timing conflicts and engineering redundancy. Therefore, existing voice broadcasting solutions suffer from technical problems such as difficulty in balancing real-time interruption accuracy, lack of broadcast progress information in interruption adjudication, and the independence of user-triggered interruptions and model self-correction logic, which are prone to conflict. Summary of the Invention
[0003] Therefore, it is necessary to provide a voice broadcasting method, computer device, and readable storage medium that can improve the reliability of voice broadcasting and achieve low response latency, in order to address the above-mentioned technical problems.
[0004] To address the aforementioned technical issues, firstly, a voice broadcasting method is provided, which includes: The cloud responds to the audio collected by the terminal, performs streaming speech recognition on the audio to obtain the recognized text, and sends the recognized text to the first model and the second model in parallel; the first model outputs the response text corresponding to the recognized text, and the second model outputs the correction text corresponding to the recognized text; The cloud generates audio based on the reply text, transmits the audio to the terminal for playback, and the terminal synchronizes the playback progress to the cloud in real time. In response to an interruption event, the terminal pauses the broadcast and records the current broadcast progress; the interruption event is triggered by the terminal detecting voice and / or by the cloud determining that the correction text and the reply text are inconsistent; The cloud-based system combines the recognized text with the current broadcast progress to determine whether an interruption is valid. If the interruption is valid, the cloud generates a new audio message based on the new response text and / or correction text and transmits it to the terminal for playback. The new response text is obtained by processing the speech. If the interruption is not valid, the terminal continues to play from the current playback progress.
[0005] In one embodiment, the cloud performs streaming speech recognition on the audio to obtain the recognized text, which includes: The cloud segment the audio according to a preset frame length, extracts features from the segmented audio, and obtains acoustic features. The acoustic features are input into the streaming acoustic encoder for acoustic modeling to obtain the acoustic coding results; The acoustic coding results are decoded frame by frame, and the temporary recognition result of the current frame is determined and output based on the decoding result obtained after decoding the current frame or by combining the decoding results obtained after decoding the current frame and the historical frames. When the change in the temporary recognition result of multiple consecutive frames is less than a preset change threshold, the update is stopped and the current temporary recognition result is used as the final recognition result; The final recognition result is used to generate the recognition text.
[0006] In one embodiment, sending the recognized text to the first model and the second model in parallel includes: The cloud will combine the recognized text with historical dialogue records to generate a model input sequence, which will then be input into the first and second models in parallel. The first model predicts the next word word by word based on the input sequence. After each word is generated, it is appended to the end of the sequence. The next word is generated based on the updated input sequence until the termination condition is met. All generated words are then combined into the response text and output. The second model predicts the next word word by word based on the input sequence. After each word is generated, it is appended to the end of the sequence. The next word is generated based on the updated input sequence until the termination condition is met. All generated words are combined into candidate text. The target candidate text is determined based on the result of word-by-word evaluation of the candidate text and / or the result of retrieval based on key information in the candidate text. The target candidate text is output as the correction text. The parameters of the first model are different from those of the second model.
[0007] In one embodiment, the interruption event includes a first interruption event and a second interruption event. In response to triggering the interruption event, the terminal pauses the broadcast and records the current broadcast progress, including: Upon detecting voice, the terminal triggers the first interruption event, pausing the broadcast and recording the current broadcast progress, and / or... When the cloud detects a discrepancy in the keywords or semantic similarity between the correction text and the reply text, it triggers a second interruption event, causing the terminal to pause the broadcast and record the current broadcast progress.
[0008] In one embodiment, the cloud-based system combines the identified text with the current playback progress to determine whether an interruption is a valid interruption, including: The cloud platform determines the remaining unbroadcast amount and the broadcast ratio of the current response content based on the current broadcast progress. When the remaining unbroadcast count is less than the first threshold, the interruption event is determined to be an invalid interruption. When the proportion of broadcasts is less than the second threshold, the interruption event is determined to be a valid interruption. When the remaining unbroadcast quantity is not less than the first threshold and the broadcast ratio is not less than the second threshold, the cloud calculates the target score by combining the intent recognition result, the echo word recognition result and the speech recognition stability. When the target score is greater than the third threshold, the interruption event is determined to be a valid interruption. When the target score is less than the third threshold, the interruption event is determined to be an invalid interruption.
[0009] In one embodiment, the cloud generates new broadcast audio based on the new reply text and / or correction text and transmits it to the terminal for playback, including: When the interruption event is the first interruption event and is determined to be a valid interruption, the cloud performs streaming speech recognition on the audio frame corresponding to the speech to obtain the interruption recognition text; the interruption recognition text is input into the dialogue model, and the dialogue model generates new response content based on the interruption recognition text and / or historical dialogue records; the new response content is divided into multiple audio segments according to a preset segmentation strategy, each audio segment is sequentially processed for speech synthesis, and the synthesized new broadcast audio is transmitted to the terminal for playback; When the interruption event is the second interruption event and is determined to be a valid interruption, the cloud divides the correction text into multiple audio segments according to the preset segmentation strategy, performs speech synthesis on each audio segment in sequence, and transmits the synthesized new broadcast audio to the terminal for playback.
[0010] In one embodiment, transmitting the synthesized new broadcast audio to the terminal for playback includes: Determine the pause point for the voice broadcast based on the current broadcast progress; Determine whether the text corresponding to the paused audio broadcast is a sentence boundary; If it is not a sentence boundary, then start from the pause position and continue playing the voice broadcast that was not finished before the pause until the target sentence boundary is reached and then stop playing, and discard the voice broadcast that was not finished after the target sentence boundary. The cloud generates a new audio message based on the new response content and / or corrected text and transmits it to the terminal for playback. The terminal starts playing the synthesized new audio message at the boundary of the target sentence.
[0011] In one embodiment, the process before the terminal plays the broadcast audio includes: The terminal receives and plays the pre-response filler phrase. Upon receiving the first audio segment of the broadcast audio, it continues the playback of the first audio segment of the broadcast audio after the pre-response filler phrase, based on the current playback progress of the pre-response filler phrase and the timestamp of the broadcast audio.
[0012] To address the aforementioned technical problems, a second aspect provides a computer device, including a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor executes the computer program to perform the following steps: the processor executes the computer program to perform the steps of the method described in the first aspect.
[0013] To address the aforementioned technical problems, a third aspect provides a computer-readable storage medium storing a computer program thereon, characterized in that the computer program, when executed by a processor, implements the steps of the method described in the first aspect.
[0014] Unlike related technologies, this application sends the recognized text obtained from streaming speech recognition to a first model and a second model in parallel. The first model generates and plays the response text first to reduce response latency, while the second model performs deep understanding in the background and produces correction text to ensure response accuracy. During the playback process, the terminal synchronizes the playback progress to the cloud in real time, so that the cloud can know the played position and the remaining content when making a decision. When the terminal detects that the voice or the cloud determines that the correction text and the response text are inconsistent, an interruption is triggered. The terminal immediately pauses the playback and records the progress. The cloud combines the recognized text with the current playback progress to determine whether it is a valid interruption. If it is a valid interruption, a new playback audio is generated based on the new response text or correction text and continued at the sentence boundary. If it is an invalid interruption, it seamlessly resumes from the preserved playback progress position, thereby improving the accuracy of interruption judgment and the playback quality while ensuring low response latency. Attached Figure Description
[0015] To more clearly illustrate the embodiments of this application, the accompanying drawings used in the embodiments will be briefly introduced below. Obviously, the drawings described below are only some embodiments of this application. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0016] Figure 1 This is a flowchart illustrating a voice broadcasting method in one embodiment; Figure 2 This is a schematic diagram of the structure of a voice broadcasting system in one embodiment; Figure 3 This is a flowchart illustrating the voice broadcasting method in another embodiment; Figure 4 This is a schematic diagram of the external interruption processing timing in one embodiment; Figure 5 This is a schematic diagram of the internal interruption processing timing in one embodiment; Figure 6 This is a flowchart illustrating the voice broadcasting method in yet another embodiment; Figure 7This is an internal structural diagram of a computer device in one embodiment. Detailed Implementation
[0017] To make the objectives, technical solutions, and advantages of this application clearer, the following detailed description is provided in conjunction with the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are merely illustrative and not intended to limit the scope of this application.
[0018] Full-duplex voice agents typically consist of multiple cascaded components, forming a complete dialogue chain from hearing to speaking: User audio is captured via microphone; Voice Activity Detection (VAD) is performed on the audio to determine the presence of user voice and its endpoints; Automatic Speech Recognition (ASR) is then performed on the audio to obtain recognized text; a dialogue model generates response content based on this; and finally, Text-to-Speech (TTS) is used to obtain synthesized audio, which is then played back to the user. Due to the trade-off between computing power and latency, these components are often deployed in an edge-cloud collaborative manner: the terminal device closer to the user handles audio capture, voice activity detection, and playback, while the cloud, with its concentrated computing power, handles streaming speech recognition, dialogue generation, and streaming speech synthesis. Audio and signaling are transmitted between the two via a real-time communication channel.
[0019] Regarding the implementation of a full-duplex experience where users can interrupt at any time, one existing approach is as follows: once the terminal detects user voice activity, it immediately mutes the current playback and sends the uplink audio to the server. The server then performs semantic judgment based on the recognized text to determine whether the interruption constitutes a genuine interruption, and sends a control command back to the terminal after the determination. Another existing approach is to obtain the offset of the current audio playback when an interruption occurs, stop playback based on this offset, clear the playback queue, and then restart a new playback.
[0020] In terms of the balance between speed and quality in dialogue generation, existing technologies also employ a so-called fast-slow dual-model approach: a fast model with a faster response first produces a draft response and outputs it promptly, while a slower model with a slower reasoning and deeper understanding refines the draft, thereby improving the response quality while maintaining a certain response speed.
[0021] However, the inventors noted in practice that the aforementioned existing technologies have at least three shortcomings. First, while the terminal's immediate silencing upon detecting voice activity is a rapid response, it is easily triggered by user agreement, coughing, or even ambient noise, causing unnecessary interruptions. While cloud-based semantic judgment is more reliable, it inevitably introduces round-trip latency between uplink audio transmission and the feedback of the judgment result, making it difficult to achieve both real-time performance and accuracy. Second, although the aforementioned offset-based scheme can determine where playback has stopped, it only uses it for stopping and clearing the queue. It does not incorporate information about the current position of the response content and the remaining amount of un-played content into the judgment of whether to interrupt. This results in the decision lacking crucial information about how much content has been played, making it difficult to make a more appropriate choice. Third, in the existing implementation, user interruption of the model and model correction of its own spoken content are separated into two unrelated sets of logic: the former is triggered by user interruption and goes through the interruption processing link, while the latter is triggered by the model's correction of its own output and goes through another correction link. The two operate independently, which not only causes duplication of engineering construction, but also hinders each other in terms of timing and is prone to conflict.
[0022] To address the aforementioned technical problems, in one embodiment, such as Figure 1 As shown, a voice broadcasting method is provided, which includes the following steps: Step 101: Upon receiving the audio collected by the terminal, the cloud performs streaming speech recognition on the audio to obtain the recognized text, and sends the recognized text to the first model and the second model in parallel; the first model outputs the response text corresponding to the recognized text, and the second model outputs the correction text corresponding to the recognized text.
[0023] Specifically, after the cloud receives the audio collected by the terminal, it segments the audio according to a preset frame length, extracts features from the segmented audio features to obtain acoustic features; inputs the acoustic features into a streaming acoustic encoder for acoustic modeling to obtain acoustic coding results; decodes the acoustic coding results frame by frame, and determines the temporary recognition result of the current frame based on the decoding result obtained after decoding the current frame or by combining the decoding results obtained after decoding the current frame and historical frames; when the change in the temporary recognition results of multiple consecutive frames is less than a preset change threshold, the update stops and the current temporary recognition result is taken as the final recognition result; and the recognized text is generated based on the final recognition result.
[0024] The cloud-based system combines the recognized text with historical dialogue records to generate a model input sequence, which is then input into the first and second models in parallel. The first model predicts the next word word by word based on the model input sequence, appending each generated word to the end of the sequence. It continues to generate the next word based on the updated model input sequence until a termination condition is met, and then combines all generated words into a response text and outputs it. The second model predicts the next word word by word based on the model input sequence, appending each generated word to the end of the sequence. It continues to generate the next word based on the updated model input sequence until a termination condition is met, and then combines all generated words into candidate text. The target candidate text is determined based on the results of word-by-word evaluation of the candidate text and / or based on the results of retrieving key information from the candidate text, and the target candidate text is output as the correction text. The parameters of the first and second models are different.
[0025] The voice broadcasting method of this application is applied to a voice broadcasting system, which includes a terminal and a cloud, such as... Figure 2 As shown, the terminal (100) and the cloud (200) are connected via a real-time communication channel (300).
[0026] Among them, the terminal can be a terminal device, which can be an embedded voice terminal equipped with a microphone and speaker, including lightweight terminal devices such as in-vehicle voice interaction terminals, smart speakers, robot voice interaction terminals, mobile phone voice assistant clients, and tablets. The terminal has low computing power and only carries low-latency lightweight tasks such as human voice acquisition, local playback, soft stop control, and progress reporting. The cloud (200) is a voice service backend deployed on a server cluster. It consists of multiple high-performance servers with sufficient computing power to centrally carry out computing power-intensive processing tasks such as streaming speech recognition, parallel reasoning of dual dialogue models, streaming speech synthesis, and unified interruption adjudication. The real-time communication channel (300) adopts the WebRTC bidirectional real-time transmission protocol. The audio media stream is transmitted through the audio track, and the broadcast progress and interruption control signaling are transmitted through an independent and reliable data channel. It is compatible with various network environments such as WiFi, 4G, and 5G, ensuring that the signaling is orderly and without packet loss.
[0027] The terminal device includes an audio acquisition module (110), a human voice activity detection module (120), a playback and soft stop module (130), a broadcast progress and position reporting module (140), and a pre-response module (150); the cloud (200) includes a streaming speech recognition module (210), a parallel model module (220), a streaming speech synthesis module (230), and a unified interruption adjudication module (240), wherein the parallel model module (220) has a built-in fast model (221) and a thinking model (222).
[0028] In one example, the real-time communication channel (300) is implemented using LiveKit based on WebRTC: user audio and synthesized audio are transmitted via the audio track, while control information such as soft stop commands, sentence boundary stop commands, resume playback commands, and playback progress position are reliably and orderly transmitted via the data channel in structured messages (e.g., JSON format, with a defined subject identifier such as barge_control). Of course, the real-time communication channel (300) can also be replaced by other channels with real-time audio and reliable signaling capabilities.
[0029] Specifically, when a user initiates voice input, the terminal's audio acquisition module captures the audio of the user's voice input and uploads it to the cloud via a real-time communication channel. This audio can be in the form of an encoded, continuously uploaded audio data sequence, also known as an audio stream.
[0030] After receiving the audio stream from the terminal, the cloud performs streaming speech recognition. Streaming speech recognition refers to receiving, recognizing, and outputting audio as it is being spoken, continuously correcting the recognition method as subsequent audio arrives. Specifically, the cloud first divides the audio stream into consecutive audio frames according to a preset frame length, such as 20ms to 40ms, and extracts acoustic features from each frame, such as log-Mel filter bank features. Then, the extracted acoustic features are input into a streaming acoustic encoder for acoustic modeling. This encoder uses a causal or block structure, relying only on the context of historical frames and a limited number of future frames to ensure accurate recognition. While improving accuracy, latency is reduced. After encoding, the decoder decodes the acoustic coding results frame by frame. When processing the current frame, if it is the first frame, the first temporary recognition result is output based on the decoding result of the current frame. If it is not the first frame, the previous temporary recognition result is updated and output by combining the decoding results of the current frame and the historical frames. As audio is continuously input, the temporary recognition result is continuously corrected. When the change in the temporary recognition result of multiple consecutive frames is less than the preset change threshold, it indicates that the recognition result has become stable. At this time, the update stops and the current temporary recognition result is output as the final recognition result, which is the recognized text.
[0031] After obtaining the identified text, the cloud simultaneously sends it to both the first and second models for processing. The first model is a fast-response model, with a relatively small model size and fewer inference steps, optimizing for first-character delay to quickly generate response text. The second model is a deep inference model, with a relatively large model size and more inference steps, enabling more thorough inference in the background to generate more accurate correction text. Specifically, after obtaining the identified text, the cloud combines it with historical dialogue records to construct the model input sequence, and then inputs the same input sequence into both the first and second models. Each model performs autoregressive word-by-word prediction based on the input sequence: each predicted word is appended to the end of the sequence, and the next word is predicted based on the updated sequence until a termination condition is met, such as predicting a stop marker or reaching the maximum generation length. The first model directly outputs all generated word combinations as the response text; the second model first generates candidate text, and then optimizes the candidate text by evaluating its logical consistency word by word from the end (re-predicting when logical contradictions or semantic incoherence are found) and / or by performing external knowledge retrieval based on key information in the candidate text (comparing the retrieval results with the candidate content, and correcting with the retrieval results when there is inconsistency), and outputs the optimized text as the correction text.
[0032] Through the parallel processing described above, the first model quickly generates the response text and passes it to subsequent steps for broadcast, so that users can hear the response as soon as possible; the second model performs deep reasoning in the background, and the correction text it produces is used for internal interruption events that may be triggered later, thereby ensuring both low response latency and accuracy of the response content.
[0033] The first model in this application prioritizes low first-word latency as its primary optimization objective, employing a language model with a relatively small parameter scale. In one example, the first model uses a model with approximately 500 million parameters, which can be compressed from a larger-scale model through knowledge distillation, quantization, or pruning, enabling it to achieve fast inference and streaming word-by-word output while maintaining basic response quality. The first model can employ a Transformer decoder architecture, with relatively small values set for the number of layers, hidden dimensions, and attention heads. For training, the first model can use a second or larger teacher model as its instructor, acquiring basic capabilities through knowledge distillation, and then aligning with the dialogue task through instruction fine-tuning (SFT), enabling it to produce fluent and acceptable immediate responses even on a small scale; alternatively, it can be obtained by directly performing instruction fine-tuning on a general-purpose small model.
[0034] The second model in this application aims for high accuracy and completeness in responses, employing a language model with a larger parameter scale and multi-step reasoning capabilities. In one example, the second model can utilize a model with approximately 70 billion parameters, possessing the ability for thought chain reasoning, retrieval enhancement, or tool invocation, enabling more thorough reasoning in the background. The second model can also employ a Transformer decoder architecture, with a larger number of layers, hidden dimensions, and attention heads compared to the first model. In terms of training, the second model can be pre-trained on a large-scale corpus, followed by instruction fine-tuning (SFT) and preference alignment (such as RLHF, DPO, etc.) to improve reasoning accuracy and response quality.
[0035] Specifically, after acquiring the recognized text in the cloud, the recognized text is combined with historical dialogue records to construct a model input sequence. This same input sequence is then fed into both the first and second models. The first and second models run in parallel in the background without blocking each other. The first model aims for a delayed first-word response; once it produces the initial segment of the response, it is immediately handed over to the speech synthesis module for segmented synthesis and downstream playback, ensuring the user hears the response as quickly as possible. Simultaneously, the second model performs deep reasoning on the same user statement in the background without occupying the output audio link, thus not affecting the playback already started by the first model.
[0036] Step 102: The cloud generates a broadcast audio based on the reply text, transmits the broadcast audio to the terminal for broadcast, and the terminal synchronizes the broadcast progress to the cloud in real time.
[0037] After obtaining the response text output by the first model from the cloud, the response text is input into the streaming speech synthesis module for speech synthesis. The streaming speech synthesis module segments the response text according to a preset segmentation strategy to reduce the latency of the first packet. The segmentation granularity can be determined by one or a combination of the following methods: segmentation by phrases or clauses, such as segmentation by secondary punctuation marks such as commas, pauses, or natural pauses, with one phrase or clause synthesized as one segment; segmentation by duration, such as each segment corresponding to approximately 200-500 milliseconds of synthesized audio; segmentation by the number of characters or words, such as consuming one segment after accumulating a set number of characters or words. The first segment can be smaller in granularity to produce sound as quickly as possible, and subsequent segments can be appropriately enlarged to improve synthesis efficiency and naturalness.
[0038] To ensure clean pauses at sentence boundaries during effective interruptions, the streaming speech synthesis module aligns sentence boundaries during segment generation. A segment split must occur at each sentence boundary to ensure the sentence boundary falls on the segment boundary. When a segment is about to cross a sentence boundary, it is prematurely terminated at the boundary to prevent it from spanning two sentences. Thus, each sentence boundary corresponds to the end position of a segment. After the speech synthesis module synthesizes each audio segment, the cloud transmits it to the terminal via a downlink real-time communication channel. Upon receiving each audio segment, the terminal adds it to the playback queue and plays it sequentially, achieving a streaming playback effect that simultaneously receives and plays the audio.
[0039] During the playback of synthesized audio, the terminal determines the current playback progress position in real time and synchronizes this position information to the cloud. The playback progress position is used to represent the position of the played content in the complete response content, and its granularity can be represented by at least one of the following: the sequence number of the played sentence (which sentence has been played so far), the sequence number of the word (which word has been played so far), or the milliseconds played (the duration of the played audio).
[0040] The reporting of the playback progress position can be triggered at fixed intervals, such as every 100-200 milliseconds; a report can be triggered after each speech synthesis segment is played; and an additional report can be made at the moment an interruption event is detected to improve the alignment accuracy of the interruption time. The terminal sends the playback progress position to the cloud through the data channel of the real-time communication channel for the cloud to use in subsequent interruption decisions.
[0041] Through the aforementioned streaming speech synthesis and transmission methods, the cloud can synthesize and transmit simultaneously, while the terminal can receive and play simultaneously, effectively reducing the initial packet latency and enabling users to hear responses as quickly as possible. Simultaneously, by aligning sentence boundaries with segment boundaries, when an effective interruption occurs, the terminal only needs to play the current segment (which terminates precisely at the sentence boundary) to smoothly stop, discarding the buffer of segments to be played after the sentence boundary, avoiding the auditory discomfort caused by abrupt cuts in the middle of words. The terminal synchronously reports the playback progress position to the cloud in real time, allowing the cloud to know crucial information such as where it has progressed and how much remains during the decision-making process, thus enabling more appropriate interruption decisions.
[0042] Step 103: In response to the interruption event, the terminal pauses the broadcast and records the current broadcast progress; the interruption event is triggered by the terminal detecting voice and / or by the cloud determining that the correction text and the reply text are inconsistent.
[0043] Interruption events include a first interruption event and a second interruption event. The first interruption event is an external interruption event, and the second interruption event is an internal interruption event. In response to triggering the first interruption event, the terminal pauses the broadcast and records the current broadcast progress, including: the terminal responds to detecting voice, triggers the first interruption event, and the terminal pauses the broadcast and records the current broadcast progress; and / or, the cloud responds to detecting that the correction text and the reply text have inconsistent keywords or semantic similarity, triggers the second interruption event, and the terminal pauses the broadcast and records the current broadcast progress.
[0044] During the playback of the synthesized audio, the voice activity detection module continuously detects voice activity in the collected audio. When the terminal detects voice activity, it triggers the first interruption event.
[0045] Specifically, the terminal collects ambient sound signals through a microphone, and the human activity detection module analyzes the collected audio in real time to determine whether there is voice activity. When voice activity is detected, the terminal immediately triggers the first interruption event. To reduce false triggers, the human activity detection module can use at least one of the following methods for preliminary filtering: echo cancellation, energy threshold, and minimum duration threshold, to avoid false triggers caused by device-generated echoes, ambient noise, or brief sounds.
[0046] When the first interruption event is triggered, the playback and soft-stop module immediately performs a soft stop on the current playback, that is, pausing the playback while retaining the playback buffer and the current playback progress position. The soft stop is an immediate response action of the terminal, without waiting for the cloud's decision, thus ensuring the real-time nature of the interruption response.
[0047] After the second model (thinking model) completes reasoning and produces corrected text, the cloud compares the corrected text with the response text already generated by the first model (fast model). When the cloud determines that the corrected text and the response text are inconsistent and meet preset conditions, a second interruption event is triggered.
[0048] The cloud-based methods for determining inconsistencies include at least one of the following: Structured difference detection: The cloud extracts key semantic elements from the response and correction texts, such as key entities (e.g., names of people, places, organizations, things), numerical values (e.g., amounts, quantities), time (e.g., dates, moments), and negative semantics (affirmative / negative polarity), and compares them item by item. If any element conflicts, the two are determined to be inconsistent. Semantic similarity detection: The cloud maps the response and correction texts to semantic vectors using a vector representation model and calculates their cosine similarity. If the similarity is below a set threshold, the two are determined to be semantically inconsistent.
[0049] When the cloud determines that the correction text and the reply text are inconsistent, and the current reply content has not yet been fully read, a second interruption event is triggered.
[0050] Please see Figure 3 In this application, the first interruption event and the second interruption event are normalized into an interruption event object (the source identifier is set to internal, and the source of the continuation content is set to the thinking model correction text), and sent to the unified interruption adjudication module for processing.
[0051] The interruption event object in this application is a unified data structure that normalizes external and internal interruption events. In one example, its fields include: source identifier (e.g., source, value external or internal) to indicate whether the event is external or internal; trigger time (e.g., timestamp, in milliseconds); associated broadcast progress position (e.g., play_pos, which can be composed of sentence number, character offset, elapsed_ms, etc.); continuation content source (e.g., content_origin, value is new user reply or thought model correction text); and, if necessary, a confidence field can be attached to record the credibility of the ruling basis.
[0052] When the second interruption event is triggered, the cloud sends a soft stop command to the terminal via the real-time communication channel. The playback and soft stop module (130) executes the soft stop, that is, it pauses the broadcast while retaining the playback buffer and the current broadcast progress position. This soft stop operation is exactly the same as the soft stop triggered by the first interruption event, thus ensuring that external interruptions and internal interruptions follow the same soft stop mechanism.
[0053] Regardless of whether the interruption is triggered by the first or second interruption event, the terminal executes the soft stop in the same way: the playback and soft stop module pauses reading audio data from the playback queue for playback, freezes the playback pointer, retains the locally decoded audio buffer (i.e., jitter buffer), and records the playback progress position corresponding to the pause time. The playback progress position can be represented by at least one of the following: the sequence number of the played sentence, the sequence number of the word, or the milliseconds played. The recorded position information is used for seamless recovery in the event of an invalid interruption or to determine the start point for continuation in the event of a valid interruption.
[0054] By executing a soft stop in real-time at the terminal, this application ensures the real-time nature of the interruption response. When a user interrupts or model correction triggers an interruption, the broadcast can be paused immediately without any residual broadcasting. By unifying external interruptions (user interruptions) and internal interruptions (model corrections) with the same soft stop mechanism, both types of interruption events execute the same pause operation at the terminal, laying the foundation for subsequent unified adjudication. Simultaneously, by overlaying echo cancellation, energy thresholds, and minimum duration thresholds in the human voice activity detection stage, false triggers caused by environmental noise and echoes are effectively reduced. Combined with cloud-based semantic adjudication, this forms a dual-gatekeeping system from the terminal to the cloud.
[0055] Step 104: The cloud combines the recognized text with the current broadcast progress to determine whether the interruption event is a valid interruption.
[0056] The cloud platform determines the remaining unplayed content and the proportion of already played content based on the current playback progress. When the remaining unplayed content is less than a first threshold, the interruption event is deemed invalid. When the proportion of already played content is less than a second threshold, the interruption event is deemed valid. When the remaining unplayed content is not less than the first threshold and the proportion of already played content is not less than the second threshold, the cloud platform calculates the target score by combining the intent recognition result, the echo word recognition result, and the speech recognition stability. When the target score is greater than a third threshold, the interruption event is deemed valid. When the target score is less than the third threshold, the interruption event is deemed invalid.
[0057] The remaining unplayed content refers to the length of text or audio duration in the current response that has not yet been played to the user. When the remaining unplayed content is less than a first threshold, it indicates that the response is nearing completion. Interrupting at this point is not very meaningful, as the user is about to hear the complete response. If this is considered a valid interruption and new content is introduced, it would result in missing information or an abrupt listening experience. Therefore, when the remaining unplayed content is less than the first threshold, the cloud directly determines the interruption as invalid, and the terminal seamlessly resumes playback from the paused position, completing the current response. In one example, the first threshold could be: the remaining unplayed content is less than a sentence, or the remaining unplayed characters are less than a certain number of characters, such as approximately 6-10 characters, or the remaining unplayed duration is less than approximately 300-500 milliseconds.
[0058] The "already broadcast ratio" refers to the proportion of the current reply content that has already been broadcast to the user, out of the total reply content. It indicates how much has been said. When the already broadcast ratio is less than a second threshold, it means the reply content has just begun to be broadcast and the user has not yet received substantial information. In this case, user interruption or system correction triggering an interruption results in minimal information loss for the user, making it more reasonable to allow interruption and switching to new content. Therefore, when the already broadcast ratio is less than the second threshold, the cloud directly determines the interruption event as valid. In one example, the second threshold could be an already broadcast ratio of less than 15%, or a broadcast duration of less than approximately 800 milliseconds.
[0059] When the remaining unplayed content is not less than the first threshold (i.e., there is a significant amount of remaining content) and the proportion of played content is not less than the second threshold (i.e., the proportion of played content has reached a certain level), the interruption event is in the gray area between being neither about to finish playing nor just starting. At this point, a clear judgment cannot be made based solely on the playback progress. The cloud enters a comprehensive judgment stage, combining the intent recognition results, the echo word recognition results, and the stability of speech recognition to calculate the target score. Based on the score result, it is determined whether the interruption is valid.
[0060] Intent recognition refers to the semantic analysis of the speech recognition text of a user interruption to determine whether the interruption constitutes a new instruction or a new question. If it is identified as constituting a new instruction or a new question, it indicates that the user has a clear intention to interrupt, and is likely to be judged as a valid interruption. Intent recognition can use a lightweight classification model to classify the recognized text and output the intent category and corresponding confidence score.
[0061] Concurrency word recognition refers to keyword matching of the user's interrupted speech text to determine if it contains concurrency words such as "um," "yes," "good," or "oh." If identified as a concurrency word, it indicates that the user is merely expressing agreement or agreement and has no real intention to interrupt, which is tended to be considered an invalid interruption. Concurrency word recognition can be achieved through deterministic string matching using a preset list of concurrency words.
[0062] Speech recognition stability refers to the convergence of streaming speech recognition results from provisional results to the final result, used to measure the reliability of the currently recognized text. Streaming speech recognition continuously outputs provisional results before the audio finishes and continuously corrects them as subsequent audio arrives. When provisional results change frequently, it indicates that the user is still speaking or the recognition result is not yet reliable; in this case, it should be judged as invalid interruption or the judgment should be postponed. When provisional results tend to be stable and the amount of change is small, it indicates that the user has finished speaking and the recognition result is reliable; in this case, a normal judgment can be made based on the recognized text. Speech recognition stability can be quantified by at least one of the following: the rate of change of edit distance between adjacent provisional results, the number of tail word flips, or the duration of retention of the same recognition prefix.
[0063] When the broadcast progress falls within the gray zone between the two thresholds mentioned above, the second-level comprehensive judgment is initiated, which involves weighted fusion of the other three signals within the gray zone. This is only activated when the first-level judgment falls into the gray zone. The priority of the three signals is, for example, intent recognition > echo word recognition > streaming recognition stability. Intent recognition refers to identifying whether the user's interruption constitutes a new instruction or a new question. If it does, there is a strong tendency to interrupt effectively, as it most directly reflects whether the user truly intends to interrupt. Echo word recognition refers to identifying whether the interruption is an echo word. If it is, there is a tendency to interrupt ineffectively. Streaming recognition stability indicates that a low stability level suggests that the recognition is not yet reliable, which may reduce the confidence level of this decision or postpone the judgment.
[0064] The target score can be calculated using a weighted summation method, as shown in the example below: Score = w1 × Intention Score + w2 × (1 - Conformity Score) + w3 × Stability Score; In this context, Score represents the target score, w1 represents the weight of intent recognition, w2 represents the weight of echoing word recognition, and w3 represents the weight of speech recognition stability, with w1>w2>w3 (reflecting the above priority). Intent recognition has the highest weight because whether the user constitutes a new instruction or a new question best reflects the true intention of interruption. Echoing word recognition is secondary, used to filter out echoing interruptions without interruption intent. Speech recognition stability serves as an auxiliary factor, reducing the credibility of the decision or delaying the decision when the recognition result is unreliable. When the target score is greater than the third threshold, it is determined to be a valid interruption; when the target score is less than the third threshold, it is determined to be an invalid interruption. When the target score exceeds the set threshold, it is determined to be a valid interruption; otherwise, it is determined to be an invalid interruption. The above weights and thresholds can be configured according to the scenario and can be dynamically adjusted based on the measured network round-trip latency. For example, when the network round-trip latency is high, the initial threshold can be appropriately increased to compensate for the delay in decision feedback. This application is not limited to the specific weights, thresholds, and priorities mentioned above.
[0065] By employing a multi-layered adjudication strategy, the accuracy of interruption judgments is improved while ensuring decision-making efficiency. The first layer, based on the broadcast progress, ensures that responses nearing completion are not interrupted and allows for flexible switching between responses just beginning to be spoken, resulting in rapid and highly certain decision-making. The second layer, in the gray area, combines intent recognition, echoing word recognition, and speech recognition stability for comprehensive judgment. It utilizes semantic information to filter out false triggers such as echoing words and noise, while also avoiding misjudgments caused by unstable recognition results. Terminal soft-stopping prioritizes real-time performance, while cloud-based layered adjudication prioritizes accuracy; the two work together to achieve a balance between real-time performance and accuracy.
[0066] Step 105: If the interruption is valid, the cloud generates a new audio message based on the new response text and / or correction text and transmits it to the terminal for playback. The new response text is obtained by processing the speech. If the interruption is not valid, the terminal continues to play from the current playback progress.
[0067] When the interruption event is the first interruption event and is determined to be a valid interruption, the cloud performs streaming speech recognition on the audio frame corresponding to the speech to obtain the interruption recognition text; the interruption recognition text is input into the dialogue model, and the dialogue model generates new response content based on the interruption recognition text and / or historical dialogue records; the new response content is divided into multiple audio segments according to a preset segmentation strategy, each audio segment is sequentially processed for speech synthesis, and the synthesized new broadcast audio is transmitted to the terminal for playback; When the interruption event is the second interruption event and is determined to be a valid interruption, the cloud divides the correction text into multiple audio segments according to the preset segmentation strategy, performs speech synthesis on each audio segment in sequence, and transmits the synthesized new broadcast audio to the terminal for playback.
[0068] When the interruption event is determined to be invalid, it means that the interruption triggered by the user's interruption or model correction should not take effect. For example, the user may have simply echoed the message, environmental noise may have triggered the interruption, or the response may have been close to completion. In this case, the terminal will continue broadcasting from the current progress.
[0069] Specifically, the terminal unpauses the playback, obtains the audio position information corresponding to the current playback progress, and starts reading the unplayed audio data from the locally decoded audio buffer starting from that position. It then sequentially reads subsequent unplayed audio segments and plays the read audio data continuously. Because the terminal retains the playback buffer and playback progress position during the soft pause, there is no need to request audio data from the cloud again or re-decode it upon resumption; playback can resume directly from the frozen position, achieving seamless recovery.
[0070] When an interruption event is determined to be a valid interruption, the cloud generates new audio messages using different methods depending on the type of the interruption event: When an interruption is the first interruption and is deemed a valid interruption, it indicates that the user has a genuine intention to interrupt, such as asking a new question or giving a new instruction. At this point, the cloud needs to recognize and understand the user's interrupted speech. Specifically, the cloud inputs the audio frame corresponding to the user's interruption into a streaming speech recognition module for speech recognition, obtaining the interruption recognition text. Then, the interruption recognition text is input into the dialogue model. The dialogue model performs semantic understanding and reasoning based on the interruption recognition text and / or historical dialogue records—that is, the user's previous dialogue with the system—to generate a new response to the user's interruption.
[0071] When the interruption event is the second interruption event and is determined to be a valid interruption, it indicates that the response content generated by the first model contains errors or inaccuracies, and the correction result of the second model should be used as the standard. In this case, the cloud directly retrieves the corrected text already generated by the second model as the new broadcast content, without needing to call the dialogue model again to generate it.
[0072] Regardless of whether the new broadcast content comes from a new response text generated by the dialogue model or from the correction text of the second model, the cloud inputs it into the streaming speech synthesis module for speech synthesis. The streaming speech synthesis module divides the input text into multiple audio segments according to a preset segmentation strategy, performs speech synthesis on each audio segment in sequence, and transmits the synthesized broadcast audio segments to the terminal for playback. The specific implementation of the preset segmentation strategy has been described in detail in step 102 and will not be repeated here.
[0073] In scenarios requiring effective interruption, to ensure a smooth listening experience for the user, the cloud needs to control the terminal to switch from the original playback to the new playback at appropriate sentence boundaries, rather than abruptly cutting off playback at any point. Specifically, the cloud determines the pause position based on the current playback progress reported by the terminal, i.e., the position where the soft stop occurs. Then, the cloud judges the playback progress of the current sentence at the pause position, i.e., whether the pause position falls exactly on the sentence boundary. If the pause position is exactly on the sentence boundary, the terminal does not need to continue playing the original content and can directly resume the new playback audio at the current position; if the pause position is not on the sentence boundary, the terminal starts from the pause position and continues playing the currently unfinished sentence until it reaches the target sentence boundary and stops playing, discarding all unplayed segments after the sentence boundary, i.e., the unplayed content after the original response will no longer be played. Here, the sentence boundary is determined by punctuation marks or semantic pauses, such as period, question mark, exclamation mark, or other sentence-ending punctuation, or semantically complete pauses.
[0074] While the terminal plays the current sentence up to the sentence boundary, the cloud generates a new audio message based on the new response or correction text and transmits it to the terminal in parallel. When the terminal plays to the target sentence boundary and stops the original playback, it immediately starts playing the new audio message synthesized and transmitted from the cloud, achieving a smooth transition between the original and new playback at the sentence boundary.
[0075] This avoids the auditory discomfort caused by abrupt cuts in the middle of words and improves the naturalness of the interruption experience; it also enables seamless resumption from the pause position when an invalid interruption occurs, avoiding unnecessary broadcast interruptions and improving the smoothness of the interaction.
[0076] This application also includes a method whereby, before the terminal plays the broadcast audio, a pre-response filler phrase is acquired and played. Upon receiving the first audio segment of the broadcast audio, the terminal, based on the current playback progress of the pre-response filler phrase and the timestamp of the broadcast audio, continues the playback of the first audio segment of the broadcast audio after the pre-response filler phrase.
[0077] In practical applications, after the terminal collects and uploads the user's audio, but before receiving the first audio segment of the broadcast audio from the cloud, there is a waiting time. This time includes the processing time of cloud-based streaming speech recognition, the inference time for the first model to generate the response text, and the processing time and network transmission time of streaming speech synthesis. If the terminal remains silent during this period, the user will perceive a significant response delay, affecting the interactive experience. To solve this problem, after the terminal sends the user's audio but before receiving the first audio segment of the broadcast audio from the cloud, it acquires and plays pre-response filler phrases. Pre-response filler phrases are short interjections or conjunctions not strongly bound to the context, such as "um," "let me think about it," "okay," etc., used to fill the auditory gap during the waiting period, thus shortening the perceived response time for the user.
[0078] When the terminal receives the first audio segment of the broadcast audio from the cloud, it connects the first audio segment of the broadcast audio to the pre-response filler phrase based on the current playback progress of the pre-response filler phrase and the timestamp of the broadcast audio. Specifically, the terminal obtains the playback time position of the pre-response filler phrase and the start timestamp of the first audio segment of the broadcast audio. Based on the correspondence between the two on the timeline, the terminal seamlessly connects the first audio segment of the broadcast audio to the current playback progress of the pre-response filler phrase, allowing the user to smoothly transition from the filler phrase to the formal response without any stuttering, overlap, or jumps. In other embodiments, the pre-response filler phrase can be taken from the terminal's preset corpus or a very short starting segment provided by a fast model; this application does not impose any restrictions on this.
[0079] By playing pre-response filler text during the waiting period and smoothly transitioning to the main audio based on timestamp alignment, the perceived response time is shortened and the smoothness of the interaction is improved. The seamless transition between the pre-response filler text and the main audio further enhances the naturalness of the listening experience and eliminates the discomfort caused by the waiting period.
[0080] In this application, when a soft stop is executed, the playback and soft stop modules freeze the playback pointer and retain the locally decoded audio buffer (i.e., jitter buffer). When an invalid interruption is determined to be a seamless recovery, playback resumes from the frozen pointer without re-decoding, thus avoiding stuttering and repetition at the resumption point. When a valid interruption is determined to be a stop at the sentence boundary, the buffer to be played after the sentence boundary is discarded. Control messages transmitted via the data channel can be aligned and sorted according to their sequence number or timestamp to tolerate out-of-order and jitter. If the terminal device does not receive the cloud's adjudication result within a preset delay (e.g., 200 milliseconds), it can first handle the situation locally using a conservative strategy, such as continuing to maintain a soft stop or automatically seamless recovery, and then correcting it after the adjudication arrives to avoid prolonged suspension. For echoes, coughs, and environmental noise, echo cancellation, energy thresholds, and minimum duration thresholds can be superimposed at the human voice activity detection module for preliminary filtering, forming a dual-gatekeeping system from end to cloud with the cloud semantic adjudication, further reducing false triggering.
[0081] In a specific implementation, please refer to Figure 4 In this embodiment, for external interruption events triggered by user-initiated interruptions during broadcasting, two processing branches are distinguished: valid interruptions and invalid interruptions.
[0082] Valid interruption: During the playback of synthesized audio on the terminal device, a user raises a new question. The voice activity detection module identifies valid voice activity, triggering an external interruption event. The playback and soft-stop module immediately executes a soft-stop operation, pausing or lowering the volume of the currently playing audio. Simultaneously, the playback progress position reporting module records and stores the current playback progress position. The terminal device uploads the user's interruption audio and the current playback progress position to the cloud synchronously via a real-time communication channel. The cloud-based unified interruption adjudication module adjudicates the interruption: first, it analyzes the interruption content through the streaming speech recognition module to identify the user's intent and determine whether the user has raised a new instruction or a new question; simultaneously, it combines this with the synchronously uploaded playback progress to confirm that the current playback is in the middle, the proportion of played content is higher than the starting threshold, and the remaining playback content is greater than the ending threshold, thus comprehensively determining it as a valid interruption. The cloud sends out effective interruption control signals through a real-time communication channel; the playback and soft stop module controls the audio playback to stop the original broadcast after reaching the boundary of a complete semantic sentence. The cloud generates a reply text for the user's new question and completes streaming speech synthesis and sends out audio segments. The terminal continues to play the new reply audio at the sentence boundary.
[0083] Invalid interruption: During the terminal device's broadcast, the user only utters agreeing interjections. The voice activity detection module still detects the speech and triggers an external interruption event. The playback and soft-stop module executes a soft stop, and the playback progress position reporting module retains the playback progress. The interrupted audio and progress information are simultaneously uploaded to the cloud. The unified interruption adjudication module performs agreeing word matching and recognition on the interrupted text, determining it to be agreeing content without actual interactive intent; or, after reading the playback progress, confirming that the remaining un-played content is less than the end threshold, both scenarios are determined to be invalid interruptions. The cloud sends a resume playback control signal, and the playback and soft-stop module reads the cached playback progress position, seamlessly continuing playback of the original synthesized audio from the breakpoint.
[0084] Please see Figure 5 This embodiment describes an automatically triggered internal interruption event when there is a discrepancy between the corrected output of the thinking model and the content broadcast by the fast model. This includes: soft stop, progress retention, hierarchical adjudication, and sentence boundary continuation. It reuses the same processing flow as external interruptions, differing only in the event source identifier. An exemplary interaction flow: The user inquires about the time of the event. The fast model prioritizes inference and generates text, for example, approximately 3:00 PM. This text is segmented and distributed to the terminal via the streaming speech synthesis module. The playback and soft stop module begins broadcasting, and the broadcast progress reporting module continuously and periodically uploads the playback position. The synchronously running thinking model performs deep multi-step inference based on the same user speech recognition text, outputting accurate corrected text: it should be 3:15 PM. The cloud compares the text broadcast by the fast model with the corrected text from the thinking model. If a content discrepancy is found and the preset triggering conditions are met, an internal interruption event is generated.
[0085] Among them, the preset conditions for determining inconsistency include two types of judgment logic, and the internal interruption is triggered when either one is met: entity comparison: the two texts conflict on core entities such as key time, numerical value, personal name, and negative semantics; semantic similarity comparison: the cosine similarity of the two texts after vectorization is lower than the preset threshold.
[0086] After generating an internal interruption event, the event is uniformly encapsulated into a standardized object and sent to the unified interruption adjudication module, using the complete layered adjudication logic of the external interruption. After the adjudication is valid, the original broadcast stops at the current sentence boundary and the audio segment of the text synthesis corrected by the thinking model is played.
[0087] Please see Figure 6In this application, after the terminal collects the user's audio, it uploads it to the cloud via a real-time communication channel. Simultaneously, the terminal performs human voice activity detection on the audio. The cloud performs streaming speech recognition on the uploaded audio to obtain recognized text, and uses at least two parallel models to process the same user statement. One fast model is used to generate response content in a timely manner to reduce response latency, and a thinking model is used to perform in-depth understanding of the response content and make corrective judgments after the fast model. The cloud performs streaming speech synthesis on the response content to obtain synthesized audio, which is sent to the terminal for streaming playback via a downlink channel. During playback, the terminal synchronizes the current playback progress position to the cloud in real time. This position represents the position of the played content in the response content. This application establishes a unified interruption control process, treating external and internal interruption events as the same type of interruption event. External interruption events occur when the terminal detects a user interrupting during playback, while internal interruption events occur when the thinking model makes a correction judgment on the response content generated by the rapid model and at least partially played. In response to any interruption event, the terminal performs a soft stop on the current playback, that is, locally pauses or reduces the current playback while retaining the playback progress position. The cloud combines the semantics of the recognized text with the playback progress position synchronized with the terminal to adjudicate the interruption event. If the adjudication is a valid interruption, the current playback stops at the sentence boundary of the response content and continues with new playback content. This new content is either a response generated in response to the user's interruption or content corrected by the thinking model. If the adjudication is an invalid interruption, the terminal seamlessly resumes the current playback based on the retained playback progress position. Through the above methods, this application achieves a balance between real-time performance and accuracy by combining terminal soft pause with cloud-based broadcast progress positioning and semantic adjudication, significantly reducing false interruptions. Furthermore, the terminal's broadcast progress position is synchronized to the cloud in real-time and directly participates in adjudication, enabling the cloud to make more rational decisions based on information such as where the broadcast has progressed and how much remains. Simultaneously, external interruptions from user interjections and internal interruptions from thinking model corrections are unified into a single interruption event, reusing the same soft pause, adjudication, and sentence boundary continuation mechanism, eliminating redundant construction and temporal conflicts between the two sets of logics for user interruption and self-correction. The rapid model speaks first, the thinking model corrects online in the background, and the terminal pre-response fills in the waiting gap, ensuring that response quality and low latency are simultaneously achieved.
[0088] It should be understood that, although Figure 1 , Figures 3-6 The steps in the flowchart are shown sequentially as indicated by the arrows, but these steps are not necessarily executed in the order indicated by the arrows. Unless otherwise specified herein, there is no strict order in which these steps are executed, and they can be performed in other orders. Figure 1 , Figures 3-6At least some of the steps in the process may include multiple sub-steps or multiple stages. These sub-steps or stages are not necessarily completed at the same time, but can be executed at different times. The execution order of these sub-steps or stages is not necessarily sequential, but can be executed in turn or alternately with other steps or at least some of the sub-steps or stages of other steps.
[0089] In one embodiment, this application also provides a computer program product, which includes a computer program stored on a non-transitory computer-readable storage medium. The computer program includes program instructions, and when the program instructions are executed by a computer, the computer is able to execute the voice broadcasting methods provided by the above methods.
[0090] In one embodiment, a computer device is provided, which may be a server, and its internal structure diagram may be as follows: Figure 7 As shown, the computer device includes a processor, memory, network interface, and database connected via a system bus. The processor provides computational and control capabilities. The memory includes non-volatile storage media and internal memory. The non-volatile storage media stores the operating system, computer programs, and database. The internal memory provides an environment for the operation of the operating system and computer programs in the non-volatile storage media. The database stores data used in the voice broadcasting method. The network interface communicates with external terminals via a network connection. When the computer program is executed by the processor, it implements a voice broadcasting method.
[0091] Those skilled in the art will understand that Figure 7 The structure shown is merely a block diagram of a portion of the structure related to the present application and does not constitute a limitation on the computer device to which the present application is applied. Specific computer devices may include more or fewer components than those shown in the figure, or combine certain components, or have different component arrangements.
[0092] Those skilled in the art will understand that all or part of the processes in the methods of the above embodiments can be implemented by a computer program instructing related hardware. The computer program can be stored in a non-volatile computer-readable storage medium. When executed, the computer program can include the processes of the embodiments of the above methods. Any references to memory, storage, databases, or other media used in the embodiments provided in this application can include non-volatile and / or volatile memory. Non-volatile memory may include read-only memory (ROM), programmable ROM (PROM), electrically programmable ROM (EPROM), electrically erasable programmable ROM (EEPROM), or flash memory. Volatile memory may include random access memory (RAM) or external cache memory. By way of illustration and not limitation, RAM is available in a variety of forms, such as static RAM (SRAM), dynamic RAM (DRAM), synchronous DRAM (SDRAM), dual data rate SDRAM (DDRSDRAM), enhanced SDRAM (ESDRAM), synchronous link DRAM (SLDRAM), RAMbus direct RAM (RDRAM), direct memory bus dynamic RAM (DRDRAM), and memory bus dynamic RAM (RDRAM), etc.
[0093] The technical features of the above embodiments can be combined in any way. For the sake of brevity, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, they should be considered to be within the scope of this specification.
[0094] The above embodiments merely illustrate several implementation methods of this application, and while the descriptions are relatively specific and detailed, they should not be construed as limiting the scope of the invention patent. It should be noted that those skilled in the art can make various modifications and improvements without departing from the concept of this application, and these all fall within the protection scope of this application. Therefore, the protection scope of this patent application should be determined by the appended claims.
Claims
1. A voice broadcasting method, characterized in that, The method includes: The cloud responds to the audio collected by the terminal, performs streaming speech recognition on the audio to obtain the recognized text, and sends the recognized text to the first model and the second model in parallel; the first model outputs the response text corresponding to the recognized text, and the second model outputs the correction text corresponding to the recognized text. The cloud platform generates a broadcast audio based on the reply text, transmits the broadcast audio to the terminal for broadcast, and the terminal synchronizes the broadcast progress to the cloud platform in real time. In response to an interruption event, the terminal pauses broadcasting and records the current broadcasting progress; the interruption event is triggered by the terminal detecting voice and / or by the cloud determining that the correction text and the reply text are inconsistent; The cloud platform combines the identified text with the current broadcast progress to determine whether the interruption event is a valid interruption. If the interruption is valid, the cloud generates a new audio message based on the new response text and / or the correction text and transmits it to the terminal for playback. The new response text is obtained by processing the speech. If the interruption is not valid, the terminal continues broadcasting from the current progress.
2. The method according to claim 1, characterized in that, The cloud-based streaming speech recognition process generates the following text: The cloud segment the audio according to a preset frame length, and extracts features from the segmented audio to obtain acoustic features; The acoustic features are input into the streaming acoustic encoder for acoustic modeling to obtain the acoustic coding results; The acoustic coding results are decoded frame by frame, and the temporary recognition result of the current frame is determined and output based on the decoding result obtained after decoding the current frame or by combining the decoding results obtained after decoding the current frame and the historical frames. When the change in the temporary recognition result of multiple consecutive frames is less than a preset change threshold, the update is stopped and the current temporary recognition result is used as the final recognition result; The final recognition result is used to generate the recognition text.
3. The method according to claim 1, characterized in that, The step of sending the recognized text to the first model and the second model in parallel includes: The cloud platform combines the recognized text with historical dialogue records to generate a model input sequence, which is then input into the first model and the second model in parallel. The first model predicts the next word word by word based on the input sequence of the model. After each word is generated, the generated word is appended to the end of the sequence. The next word is generated based on the updated input sequence of the model until the termination condition is met. All generated words are combined into a response text and output. The second model predicts the next word word by word based on the input sequence of the model. After each word is generated, it is appended to the end of the sequence. The next word is generated based on the updated input sequence of the model until the termination condition is met. All generated words are combined into candidate text. The target candidate text is determined based on the result of word-by-word evaluation of the candidate text and / or the result of retrieval based on key information in the candidate text. The target candidate text is output as the correction text. The parameters of the first model are different from those of the second model.
4. The method according to claim 1, characterized in that, The interruption event includes a first interruption event and a second interruption event. The step of pausing the broadcast and recording the current broadcast progress in response to triggering the interruption event includes: In response to detecting voice, the terminal triggers a first interruption event, pausing the broadcast and recording the current broadcast progress, and / or... In response to the cloud's detection that the keywords or semantic similarity between the corrected text and the reply text are inconsistent, a second interruption event is triggered, and the terminal pauses the broadcast and records the current broadcast progress.
5. The method according to claim 4, characterized in that, The cloud-based system, by combining the identified text with the current playback progress, determines whether the interruption event is a valid interruption by: The cloud platform determines the remaining unbroadcast amount and the broadcast ratio of the current response content based on the current broadcast progress. When the remaining unbroadcast amount is less than the first threshold, the interruption event is determined to be an invalid interruption; When the broadcast ratio is less than the second threshold, the interruption event is determined to be a valid interruption. When the remaining unbroadcast quantity is not less than the first threshold and the broadcast ratio is not less than the second threshold, the cloud calculates the target score by combining the intent recognition result, the echo word recognition result and the speech recognition stability. When the target score is greater than the third threshold, the interruption event is determined to be a valid interruption. When the target score is less than the third threshold, the interruption event is determined to be an invalid interruption.
6. The method according to claim 4, characterized in that, The cloud-based generation of new broadcast audio based on the new reply text and / or the correction text, and its transmission to the terminal for playback, includes: When the interruption event is the first interruption event and is determined to be a valid interruption, the cloud performs streaming speech recognition on the audio frame corresponding to the speech to obtain the interruption recognition text; the interruption recognition text is input into the dialogue model, and the dialogue model generates new response content based on the interruption recognition text and / or historical dialogue records; the new response content is divided into multiple audio segments according to a preset segmentation strategy, each audio segment is sequentially processed by speech synthesis, and the synthesized new broadcast audio is transmitted to the terminal for playback; When the interruption event is the second interruption event and is determined to be a valid interruption, the cloud divides the correction text into multiple audio segments according to a preset segmentation strategy, performs speech synthesis on each audio segment in sequence, and transmits the synthesized new broadcast audio to the terminal for playback.
7. The method according to claim 6, characterized in that, The step of transmitting the synthesized new broadcast audio to the terminal for playback includes: The pause position for the voice broadcast is determined based on the current broadcast progress; Determine whether the text corresponding to the voice broadcast at the pause position is a sentence boundary; If it is not a sentence boundary, then start from the pause position and continue playing the voice broadcast that was not finished before the pause until the target sentence boundary is reached and then stop playing, and discard the voice broadcast that was not finished after the target sentence boundary. The cloud platform generates a new broadcast audio based on the new response content and / or the corrected text and transmits it to the terminal for playback. The terminal then begins playing the synthesized new broadcast audio at the boundary of the target sentence.
8. The method according to claim 1, characterized in that, Before the terminal plays the broadcast audio, the following steps are included: The terminal receives the first audio segment of the broadcast audio and, based on the current playback progress of the pre-response filler and the timestamp of the broadcast audio, continues the playback of the first audio segment of the broadcast audio after the pre-response filler.
9. A computer device, comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that, When the processor executes the computer program, it implements the steps of the method according to any one of claims 1 to 8.
10. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by a processor, it implements the steps of the method as described in any one of claims 1 to 8.