Voice signal truncation method and apparatus, and storage medium
Patent Information
- Application Number
- CN202610797046.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2026-06-03
- Publication Date
- 2026-09-29
AI Technical Summary
相关技术中,VAD通常依赖于声学特征,如能量、频谱熵,无法理解语言内容,且采用固定的监听等待时长,导致语音信号容易出现早切或晚切
本说明书实施例提供的语音信号截断方法中,在对用户的语音信号进行监听的过程中,对持续监听到的语音信号进行实时语音识别,得到目标语音识别结果,目标语音是被结果包括已经识别到的目标文本片段以及目标文本片段对应的目标词性标签序列;基于多个预设的语音判停策略分别对目标语音识别结果进行判停处理,若存在目标语音判停策略判定目标语音识别结果为完整指令,将目标语音判停策略对应的目标后置静音时长作为后续等待语音输入的最大限制时长,其中,每个预设的语音判停策略对应的后置静音时长均不相同;若在目标后置静音时长内未检测到新的语音信号输入,对语音信号进行截断。本方案中,通过设置多个预设的语音判停策略,当用户语音信号对应的目标语音识别结果命中任一语音判停策略后,则将后置静音时长调整为该语音判停策略所对应的后置静音时长,即,语音判停过程中的后置静音时长是可变的,而不是固定不变的,从而避免固定的后置静音时长导致长时间的监听等待耗时,有效缩短延时,实现语音信号的快速判停。同时,预设的语音判停策略通过判定目标语音识别结果是否为完整指令来对语音信号进行截断,能够有效避免在指令未完毕时就进行截断,大大降低了语音信号的早切发生。
Smart Images

Figure CN122842575A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of voice interaction, and in particular to a method, apparatus and storage medium for truncating voice signals. Background Technology
[0002] With the rapid development of human-computer interaction technology, voice interaction has received increasing attention and application. Voice Activity Detection (VAD), used to identify the start and end points of valid human voices in a continuous audio stream, is a crucial step in voice front-end processing. In related technologies, VAD typically relies on acoustic features such as energy and spectral entropy, but cannot interpret the language content. Furthermore, it uses a fixed listening wait time, leading to premature or late cut-off of the voice signal. Premature cut-off occurs when the user has not finished speaking, causing instruction truncation; late cut-off occurs when the user has finished speaking but the system continues listening, increasing response latency. Summary of the Invention
[0003] This invention provides a method, apparatus, and storage medium for truncating voice signals, so as to achieve more accurate truncating of voice signals.
[0004] In a first aspect, embodiments of the present invention provide a method for truncating a voice signal, comprising: During the process of monitoring the user's voice signal, real-time speech recognition is performed on the continuously monitored voice signal to obtain the target speech recognition result. The target speech recognition result includes the recognized target text segment and the target part-of-speech tag sequence corresponding to the target text segment. The target speech recognition result is processed for stopping based on multiple preset speech stopping strategies. If a target speech stopping strategy determines that the target speech recognition result is a complete instruction, the target post-silence duration corresponding to the target speech stopping strategy is used as the maximum limit duration for subsequent waiting for speech input. The post-silence duration corresponding to each preset speech stopping strategy is different. If no new voice signal input is detected within the target's post-silence duration, the voice signal is truncated.
[0005] In some implementations, the plurality of preset voice stop strategies include a first voice stop strategy based on instruction template for complete instruction recognition, a second voice stop strategy based on part-of-speech tag sequence for complete instruction recognition, and a third voice stop strategy based on semantics for complete instruction recognition.
[0006] In some implementations, the post-mute durations corresponding to the first voice stop strategy, the second voice stop strategy, and the third voice stop strategy increase sequentially.
[0007] In some implementations, the step of performing stop processing on the target speech recognition result based on multiple preset speech stop strategies includes: According to the preset priority order of the multiple preset voice pausing strategies, each preset voice pausing strategy sequentially performs pausing processing on the target voice recognition result.
[0008] In some implementations, the step of performing stop processing on the target speech recognition result based on multiple preset speech stop strategies includes: During the process of pausing the target speech recognition result based on the first speech pausing strategy, the target text segment is matched with a preset instruction template library; If a target instruction template that matches the target text fragment exists in the preset instruction template library, the target field corresponding to the target slot is extracted from the target text fragment based on the target slot corresponding to the target instruction template; If the length of the target field is greater than or equal to the preset length, and the target field contains the preset keyword corresponding to the target slot, then the target speech recognition result is determined to be a complete instruction, and the first speech pausing strategy is used as the target speech pausing strategy.
[0009] In some implementations, the step of performing stop processing on the target speech recognition result based on multiple preset speech stop strategies includes: During the process of pausing the target speech recognition result based on the second speech pausing strategy, the syntactic integrity of the target part-of-speech tag sequence is determined. If the target part-of-speech tag sequence is syntactically complete, the target speech recognition result is determined to be a complete instruction, and the second speech pausing strategy is used as the target speech pausing strategy.
[0010] In some implementations, the step of determining the syntactic integrity of the target part-of-speech tag sequence includes: If the last two consecutive part-of-speech tags of the target part-of-speech tag sequence are verbs and nouns, or verbs and pronouns, the target part-of-speech tag sequence is determined to be a complete imperative sentence syntax; If the first three parts of speech tags of the target part of speech tag sequence are noun, verb and noun in sequence, or the first three parts of speech tags are noun, verb and pronoun in sequence, the target part of speech tag sequence is determined to be a complete subject-verb-object sentence; If the last two consecutive part-of-speech tags of the target part-of-speech tag sequence are a preposition and a noun respectively, the target part-of-speech tag sequence is determined to be a complete prepositional complement syntax.
[0011] In some implementations, the step of performing stop processing on the target speech recognition result based on multiple preset speech stop strategies includes: In the process of pausing the target speech recognition result based on the third speech pausing strategy, a preset semantic recognition model is used to process the target text segment to obtain a first score that characterizes the semantic integrity of the target text segment. Based on the voice text segment from the previous round of voice interaction, the correlation between the target text segment and the voice text segment is determined to obtain a second score; Based on the first score and the second score, a target score is obtained. When the target score is greater than a preset score, the target speech recognition result is determined to be a complete instruction, and the third speech pausing strategy is used as the target speech pausing strategy.
[0012] In some implementations, if there is no target speech pausing strategy to determine that the target speech recognition result is a complete instruction, the method further includes: Identify the target intent of the target speech recognition result. If the target intent is a navigation intent, extend the maximum time limit for subsequent waiting for speech input to the first time limit. If the user's historical average command duration is greater than the preset duration, the maximum time limit for waiting for subsequent voice input will be extended to the second duration. Determine the target confidence level of the target speech recognition result. If the target confidence level is less than the preset confidence level, extend the maximum time limit for subsequent waiting for speech input to a third time limit.
[0013] Secondly, embodiments of the present invention provide a voice signal truncation device, the device comprising: The speech recognition module is used to perform real-time speech recognition on continuously monitored speech signals during the process of monitoring the user's speech signals, and to obtain the target speech recognition result. The target speech recognition result includes the recognized target text segment and the target part-of-speech tag sequence corresponding to the target text segment. The first processing module is used to perform stop processing on the target speech recognition result based on multiple preset speech stop strategies. If a target speech stop strategy determines that the target speech recognition result is a complete instruction, the target post-silence duration corresponding to the target speech stop strategy is used as the maximum limit duration for subsequent waiting for speech input. The post-silence duration corresponding to each preset speech stop strategy is different. The second processing module is used to truncate the voice signal when no new voice signal input is detected within the target post-silence duration.
[0014] Thirdly, embodiments of the present invention provide a voice signal truncation device, including a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor executes the program to implement the steps in the above-described voice signal truncation method.
[0015] Fourthly, embodiments of the present invention provide a computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements the steps in the above-described speech signal truncation method.
[0016] The above-described one or more technical solutions in the embodiments of this application have at least the following technical effects: In the speech signal truncation method provided in the embodiments of this specification, during the process of monitoring the user's speech signal, real-time speech recognition is performed on the continuously monitored speech signal to obtain the target speech recognition result. The target speech is the result including the already recognized target text segment and the target part-of-speech tag sequence corresponding to the target text segment. Based on multiple preset speech cessation strategies, the target speech recognition result is processed for cessation. If a target speech cessation strategy determines that the target speech recognition result is a complete instruction, the target post-silence duration corresponding to the target speech cessation strategy is used as the maximum limit duration for subsequent waiting for speech input. The post-silence duration corresponding to each preset speech cessation strategy is different. If no new speech signal input is detected within the target post-silence duration, the speech signal is truncated. In this solution, multiple preset voice stop strategies are set. When the target speech recognition result corresponding to the user's voice signal matches any of these strategies, the post-silence duration is adjusted to the duration corresponding to that strategy. This means the post-silence duration during the voice stop process is variable, not fixed, thus avoiding prolonged listening and waiting time caused by a fixed post-silence duration, effectively shortening latency and achieving rapid voice signal stoppage. Simultaneously, the preset voice stop strategies truncate the voice signal by determining whether the target speech recognition result is a complete instruction, effectively preventing premature truncation and significantly reducing the occurrence of premature voice signal cutoff. Attached Figure Description
[0017] Figure 1 A flowchart illustrating a speech signal truncation method provided in an embodiment of this specification; Figure 2 This is a schematic diagram of the functional modules of a voice signal truncation device provided in the embodiments of this specification; Figure 3 This is a schematic diagram of a voice signal truncation device provided in an embodiment of this specification. Detailed Implementation
[0018] The overall technical solution of this specification embodiment is as follows: During the process of monitoring the user's voice signal, real-time speech recognition is performed on the continuously monitored voice signal to obtain the target speech recognition result. The target speech recognition result includes the recognized target text segment and the target part-of-speech tag sequence corresponding to the target text segment. Based on multiple preset speech pausing strategies, the target speech recognition result is processed for pausing. If a target speech pausing strategy determines that the target speech recognition result is a complete instruction, the target post-silence duration corresponding to the target speech pausing strategy is used as the maximum limit duration for subsequent waiting for voice input. The post-silence duration corresponding to each preset speech pausing strategy is different. If no new voice signal input is detected within the target post-silence duration, the voice signal is truncated.
[0019] The solution in this embodiment sets multiple preset voice cessation strategies. When the target speech recognition result corresponding to the user's speech signal matches any of these strategies, the post-silence duration is adjusted to the duration corresponding to that strategy. In other words, the post-silence duration during the voice cessation process is variable, not fixed, thus avoiding prolonged listening and waiting time caused by a fixed post-silence duration, effectively shortening latency and achieving rapid cessation of the speech signal. Simultaneously, the preset voice cessation strategies truncate the speech signal by determining whether the target speech recognition result is a complete instruction, effectively preventing truncation before the instruction is complete and significantly reducing premature cutoff of the speech signal.
[0020] To better understand the above technical solutions, the technical solutions of the embodiments of this specification will be described in detail below with reference to the accompanying drawings and specific embodiments. It should be understood that the embodiments of this specification and the specific features in the embodiments are detailed descriptions of the technical solutions of the embodiments of this specification, rather than limitations on the technical solutions of this specification. Unless otherwise specified, the embodiments of this specification and the technical features in the embodiments can be combined with each other.
[0021] First, it should be clarified that the term "and / or" in this article is merely a description of the relationship between related objects, indicating that three relationships can exist. For example, A and / or B can represent: A existing alone, A and B existing simultaneously, or B existing alone. Additionally, the character " / " in this article generally indicates that the preceding and following related objects have an "or" relationship.
[0022] This specification provides an embodiment of a method for truncating voice signals, such as... Figure 1 As shown, the method includes the following steps: Step S101: During the process of monitoring the user's voice signal, real-time speech recognition is performed on the continuously monitored voice signal to obtain the target speech recognition result. The target speech recognition result includes the identified target text segment and the target part-of-speech tag sequence corresponding to the target text segment. Step S102: Based on multiple preset voice stop judgment strategies, the target voice recognition result is processed for stop judgment. If a target voice stop judgment strategy determines that the target voice recognition result is a complete instruction, the target post-silence duration corresponding to the target voice stop judgment strategy is used as the maximum limit duration for subsequent waiting voice input. The post-silence duration corresponding to each preset voice stop judgment strategy is different. Step S103: If no new voice signal input is detected within the target post-silence duration, the voice signal is truncated.
[0023] The solutions described in this specification can be applied to terminal devices capable of voice control, such as vehicles, user smart devices, and home appliances. They can also be applied to cloud servers connected to the terminal devices, or to systems consisting of the terminal devices and cloud servers; no limitation is made here. In some embodiments, considering that voice data uploading to the cloud for processing may experience delays due to network congestion, execution on the terminal may be selected as a limited option. For ease of explanation, the method provided in this specification embodiment will be described using a vehicle as an example.
[0024] In step S101, the user can interact with the vehicle via voice control. For example, if the user wants to plan a route, they can say "Navigate to XX Street" to control the vehicle to automatically plan the navigation route. Alternatively, the user can say "Open the window" to control the window to open. In some embodiments, the user's voice signal can be detected in various ways, such as by processing the detected audio stream using the Web RTC VAD lightweight voice activity detection module to distinguish between user voice and noise and identify the starting point of the user's voice signal. Of course, other methods can also be used to monitor the voice signal, which are not limited here.
[0025] After listening to the user's voice signal, real-time speech recognition is performed to obtain the target speech recognition result, which may include the target text fragment and the target part-of-speech tag sequence. The target text fragment can be obtained through a preset speech recognition model. In some embodiments, the speech recognition model can be a Hidden Markov Model, a Transformer speech model, etc., which is not limited here. When training the preset speech recognition model, the training data may include a set of speech signals, with each speech signal labeled with corresponding text. Based on the training data, the initial speech recognition model is iteratively trained, and the model parameters are adjusted until a preset number of iterations is reached or the accuracy of the model output exceeds a threshold, thus completing the model training.
[0026] The user's voice signal is input into a preset speech recognition model, which processes the voice signal and outputs the corresponding target text segment. Since the user's voice signal is continuously input as a continuous audio stream, the preset speech recognition model also processes incremental audio data in real time and outputs the recognized text segment incrementally.
[0027] In the embodiments of this specification, after obtaining the target text segment corresponding to the speech signal, part-of-speech tagging can be performed on the target text segment to determine the part-of-speech (e.g., noun, verb, adjective, etc.) of each word in the target text segment and assign a corresponding part-of-speech tag to each word. In some embodiments, a part-of-speech tagging model can be used to process the target text segment. During the training of the part-of-speech tagging model, the training data may include a set of texts that have already been segmented into words. For each text, there is a corresponding part-of-speech tag for each word segment in the text. The initial part-of-speech tagging model is iteratively trained using the training data, and the model parameters are adjusted until a preset number of iterations is reached or the accuracy of the model output is greater than a threshold, thus completing the model training. By processing the target text segment using the trained part-of-speech tagging model, the corresponding target part-of-speech tag sequence can be output.
[0028] In some embodiments, in order to minimize the duration of voice interaction, the part-of-speech tagging model can be a lightweight BiLSTM-CRF model that can support Chinese part-of-speech tags (including more than 20 part-of-speech categories such as nouns, verbs, prepositions, and pronouns).
[0029] In step S102, multiple preset voice pausing strategies are pre-set, each with its own pausing rules to determine whether the target speech recognition result is a complete instruction. In this embodiment, the multiple preset voice pausing strategies can cover various types of speech instructions, such as pausing strategies for high-frequency instructions, pausing strategies for fixed syntax instructions, and pausing strategies for complex speech patterns. The post-silence duration is dynamically adjusted for different speech instructions. The post-silence duration refers to the critical duration for continuously detecting silence after the speech instruction ends. When the cumulative silence duration reaches this critical duration, the current speech segment is determined to have ended, triggering segment segmentation and truncation.
[0030] In some embodiments, a plurality of preset voice stop strategies include a first voice stop strategy based on instruction template for complete instruction recognition, a second voice stop strategy based on part-of-speech tag sequence for complete instruction recognition, and a third voice stop strategy based on semantics for complete instruction recognition.
[0031] The duration of the silence following each voice stop strategy can be set according to actual needs. In some embodiments, the duration of the silence following the first, second, and third voice stop strategies increases sequentially. For example, the duration of the silence following the first voice stop strategy is 300ms, the duration of the silence following the second voice stop strategy is 600ms, and the duration of the silence following the third voice stop strategy is 1200ms. Of course, the duration of the silence following each voice stop strategy can also be set to other durations, which is not limited here.
[0032] In step S102, after obtaining the target speech recognition result, multiple preset speech cessation strategies can process the target speech recognition result simultaneously, or they can judge the target speech recognition result sequentially according to a predetermined order. If the target speech recognition result matches any speech cessation strategy, then that speech cessation strategy is taken as the target speech cessation strategy, and the corresponding target post-silence duration is taken as the maximum limit duration for subsequent waiting for speech input.
[0033] To better understand the various voice stop strategies, the processing rules for the first, second, and third voice stop strategies are explained below.
[0034] I. Strategy for stopping the first voice prompt During the process of pausing the target speech recognition result based on the first speech pausing strategy, the target text segment is matched with a preset instruction template library. If a target instruction template that successfully matches the target text segment exists in the preset instruction template library, the target field corresponding to the target slot is extracted from the target text segment based on the target slot corresponding to the target instruction template. If the length of the target field is greater than or equal to a preset length, and the target field contains a preset keyword corresponding to the target slot, then the target speech recognition result is determined to be a complete instruction, and the first speech pausing strategy is used as the target speech pausing strategy.
[0035] Specifically, the first voice stop strategy is an ultra-fast stop channel. In the embodiments of this specification, a preset command template library is pre-built for the first voice stop strategy. The preset command template library consists of statistically obtained high-frequency command templates, such as "Navigate to...", "Call...", "Play...", etc. Each command template in the preset command template library has its own corresponding slot. For example, the command template "Navigate to..." corresponds to a location slot, the command template "Call..." corresponds to a contact slot, and the command template "Play..." corresponds to a song slot.
[0036] Each slot can have its own preset keywords. For example, for the location slots mentioned above, the preset keywords include city, district, road, street, station, building, and neighborhood; for the contact slots mentioned above, the preset keywords include telephone and mobile phone; and for the song slots mentioned above, the preset keywords include song and music. It should be noted that the preset keywords for each slot can be set according to the actual situation, and there is no limitation here.
[0037] In some embodiments, if a target text fragment matches a target instruction template, the target field corresponding to the target slot of the target instruction template is extracted from the target text fragment. It should be noted that if the target field is too short, the user's voice instruction may not be finished; truncating it directly would result in premature cutoff of the voice instruction. Therefore, to avoid premature cutoff, in this embodiment, it can be first determined whether the length of the target field is greater than or equal to a preset length. The preset length can be set according to actual needs, for example, a preset length of 2.
[0038] If the length of the target field is greater than or equal to the preset length, the target field is then matched with the preset keyword corresponding to the target slot. If the target field contains the preset keyword, the target speech recognition result is determined to be a complete instruction, and the first speech pausing strategy is used as the target speech pausing strategy. In some embodiments, the preset keyword can be limited to the end of the sentence; that is, if the target field ends with the preset keyword, the target speech recognition result is determined to be a complete instruction; otherwise, the target speech recognition result is considered not to be a complete instruction.
[0039] It should be noted that if there is no instruction template matching the target text fragment, or the length of the target field is less than the preset length, or the target field does not include the corresponding preset keyword, it indicates that the target speech recognition field has not matched the first speech pausing strategy.
[0040] In the embodiments of this specification, if the target speech recognition field matches the first speech cessation strategy, the post-silence duration is set to the post-silence duration corresponding to the first speech cessation strategy. It should be noted that during the speech activity detection process, an initial post-silence duration, such as 2 seconds, can be set for each interactive command. That is, in the initial stage of speech signal monitoring, a 2-second post-silence duration is used to wait for the speech signal input. When the target speech recognition result of the speech signal matches the first speech cessation strategy, the post-silence duration is adjusted from 2 seconds to the post-silence duration corresponding to the first speech cessation strategy, such as 300ms, thereby enabling rapid cessation of the speech signal.
[0041] II. Strategy for stopping the second voice prompt During the process of pausing the target speech recognition result based on the second speech pausing strategy, the syntactic integrity of the target part-of-speech tag sequence is determined; if the target part-of-speech tag sequence is syntactically complete, the target speech recognition result is determined to be a complete instruction, and the second speech pausing strategy is used as the target speech pausing strategy.
[0042] Specifically, the second speech cessation strategy is to determine whether the syntactic structure of the target text segment is complete based on the target part-of-speech tag sequence corresponding to the currently identified target text segment during the speech process and before the end of the speech is detected.
[0043] In some embodiments, a sequence of part-of-speech tags for complete syntactic structures can be pre-set, such as the part-of-speech tag sequence corresponding to the syntactic structure of a complete imperative sentence, or the part-of-speech tag sequence corresponding to a complete subject-verb-object syntactic structure. The target part-of-speech tag sequence is matched against the pre-set sequence of part-of-speech tags for complete syntactic structures. If the target part-of-speech tag sequence matches any of the pre-set sequences of part-of-speech tags for complete syntactic structures, the target speech recognition result is determined to be a complete instruction, and the target speech recognition structure has matched the second speech pausing strategy. The second speech pausing strategy is then adopted as the target speech pausing strategy.
[0044] To better understand the determination of syntactic completeness based on target part-of-speech tag sequences, the following examples illustrate the matching process for several complete syntactic sequences: If the last two consecutive part-of-speech tags of the target part-of-speech tag sequence are verbs and nouns, or verbs and pronouns, the target part-of-speech tag sequence is determined to be a complete imperative sentence; if the first three part-of-speech tags of the target part-of-speech tag sequence are nouns, verbs, and nouns, or nouns, verbs, and pronouns, the target part-of-speech tag sequence is determined to be a complete subject-verb-object sentence; if the last two consecutive part-of-speech tags of the target part-of-speech tag sequence are prepositions and nouns, the target part-of-speech tag sequence is determined to be a complete prepositional complement sentence.
[0045] Specifically, when users use voice control, they often issue commands using imperative sentences, such as "open the car window," "adjust the air conditioning temperature," or "play music." For a complete imperative sentence, the corresponding part-of-speech tag sequence is either verb + noun or verb + pronoun. Therefore, if the last two consecutive part-of-speech tags of the target part-of-speech tag sequence are verb + noun or verb + pronoun, the target speech recognition result is considered a complete imperative sentence, and the syntactic completeness of the target speech recognition result is considered.
[0046] For complete subject-verb-object syntax, such as "voice assistant opens the car window," "seat massage starts," and "ambient light changes color," a complete subject-verb-object structure usually expresses a complete intention. Therefore, if the target part-of-speech tag sequence matches the part-of-speech tag sequence of a subject-verb-object syntax, the target speech recognition result is considered a complete syntactic structure. In some embodiments, the part-of-speech tag sequence corresponding to the subject-verb-object syntax can be noun + verb + noun, or noun + verb + pronoun. If the first three part-of-speech tags of the target part-of-speech tag sequence are noun, verb, and noun / pronoun, respectively, then the target speech recognition result is a complete subject-verb-object syntactic structure.
[0047] For prepositional complement syntax, for example, if a user's voice command in the previous interaction was "Drive the car to school," but the user needs to change the destination, and therefore outputs the voice command "No, to the company," then "to the company" corresponds to prepositional complement syntax, where the preposition is followed by an object to provide supplementary information. Prepositional complement syntax can be used alone as a command, or it can be understood in conjunction with the preceding context. Therefore, prepositional complement syntax can also indicate that the user's command is complete. In some embodiments, the part-of-speech tag sequence corresponding to prepositional complement syntax can be preposition + noun, or it can include preposition + pronoun, etc. If the last two consecutive part-of-speech tags of the target part-of-speech tag sequence are a preposition and a noun, respectively, then the target speech recognition result is the complete syntactic structure of the prepositional complement sentence.
[0048] Of course, in addition to the syntactic structures mentioned above, other sentence structures can also be included, which will not be listed here. For each sentence structure, the corresponding standard part-of-speech tag sequence can also be set according to actual needs, without limitation here. When the target part-of-speech tag sequence matches the part-of-speech tag sequence of any complete syntactic structure, it is determined that the target speech recognition structure has hit the second speech pausing strategy, and the second speech pausing strategy is used as the target speech pausing strategy, and the post-silence duration is set to the post-silence duration corresponding to the second speech pausing strategy.
[0049] III. Strategy for stopping third-party voice input In the process of pausing the target speech recognition result based on the third speech pausing strategy, a preset semantic recognition model is used to process the target text segment to obtain a first score that characterizes the semantic integrity of the target text segment; based on the speech text segment of the previous round of speech interaction, the correlation between the target text segment and the speech text segment is determined to obtain a second score; based on the first score and the second score, a target score is obtained, and when the target score is greater than a preset score, the target speech recognition result is determined to be a complete instruction, and the third speech pausing strategy is used as the target speech pausing strategy.
[0050] Specifically, the third voice cessation strategy is a precise cessation channel, capable of accurately recognizing user voice commands to precisely truncate speech. In the embodiments of this specification, the completeness of the target speech recognition result can be determined by considering two factors: the semantic completeness of the target speech recognition result itself, and the correlation between the target speech recognition result and the previous round of speech text segment, i.e., whether the previous round of interaction can help understand the target speech recognition result.
[0051] In some embodiments, the target speech recognition result can be output using a pre-trained semantic evaluation model. The type of pre-trained semantic evaluation model can be selected according to actual needs. For example, to compress the model size and processing time, a distilled version of the TinyBERT model can be selected. During the training of the semantic evaluation model, the training data can include a set of speech text collected from real speech interactions. For each speech text, the corresponding label can be semantically complete or semantically incomplete. The initial semantic evaluation model is iteratively trained using the training data, and the model parameters are adjusted until a preset number of iterations is reached or the accuracy of the model output exceeds a threshold, thus completing the model training. The target text segment from the target speech recognition result is input into the trained semantic evaluation model, which outputs a first score characterizing the semantic completeness of the target text segment, where the first score S... semantic The value range is [0,1].
[0052] In some embodiments, the user's voice signal often omits some content, but this omission can be filled in by the previous round of dialogue, thus identifying the user's interactive intent. For example, if the user's previous interaction command was "turn up the air conditioner temperature," and the currently listened-to target text fragment is "turn it up a little more," the target text fragment actually contains the slot data of the object being operated on. Through the previous interaction, it can be known that the air conditioner is being adjusted; therefore, the target text fragment can be completed, i.e., "turn the air conditioner temperature up a little more." If the semantics of the target text fragment can be completed by the voice text fragment from the previous round of voice interaction, then the target text fragment and the previous voice text fragment have a strong correlation. If the semantics of the target text fragment cannot be completed by the previous voice text fragment, or if the target text fragment is completely unrelated to the previous voice interaction and consists of two independent commands, then the target text fragment and the previous voice text fragment have no correlation.
[0053] In some embodiments, a corresponding second score S can be set to determine whether the target text segment is related to the previous round of speech text segment. context For example, if the target text segment is related to the audio text segment from the previous round, meaning the audio text segment from the previous round can help understand the target text segment, then the corresponding second score S context The score is 0.9. If the target text segment is unrelated to the previous round's audio text segment, or if the previous round's audio text segment cannot help in understanding the target text segment, the corresponding second score S is... contextThe score is 0.3. Of course, the score of the second score in each case can be set according to actual needs. In addition, the second score can also be obtained in other ways, such as by calculating the similarity between the previous round's speech text segment and the target text segment. These will not be listed here.
[0054] Furthermore, the first score and the second score can be weighted. In some embodiments, the weights of the first score and the second score can be set. For example, the weight of the first score is 0.7 and the weight of the second score is 0.3, then the final target score S total It can be obtained through the following formula: S total =0.7×S semantic +0.3×S context Of course, the weights of the first and second ratings can be set according to actual needs, and there are no restrictions here.
[0055] When the target score is greater than the preset score, the target speech recognition result is determined to be a complete instruction. The target speech recognition result hits the third speech stop strategy, and the third speech stop strategy is used as the target stop strategy. The post-silence duration is set to the post-silence duration corresponding to the third speech stop strategy. The preset score can be set according to actual needs. For example, the preset score can be 0.75, 0.8, etc.
[0056] It should be noted that when calculating the second score, you can examine the correlation between the target text segment and the previous round's audio text segment, or you can examine the correlation between the target text segment and the audio text segments from the previous K rounds. The value of K can be set according to actual needs. For example, K can be 2 or 3, which means that the understanding of the currently monitored audio command is based on the audio interactions from several previous rounds.
[0057] In some embodiments, in addition to the first and second scores mentioned above, other semantically relevant factors can be added. For example, a third score can be determined based on the syntactic completeness of the target text fragment. The third score can be assigned a corresponding value depending on the situation, such as a score of 1 for syntactic completeness and 0.2 for syntactic incompleteness. Alternatively, a syntactic completeness score can be output using a pre-trained model; this is not limited here. Therefore, when calculating the target score, the first, second, and third scores can be weighted and fused, for example: S total =0.4×S syntax +0.4×S semantic +0.2×S context , of which S syntax This is the third score used to characterize the syntactic integrity of the target text fragment.
[0058] In the embodiments of this specification, after obtaining the target speech recognition result, the first speech pausing strategy, the second speech pausing strategy, and the third speech pausing strategy can process the target speech recognition result simultaneously. If the target speech recognition result matches one of the speech pausing strategies, the corresponding post-mute duration is adjusted. In some embodiments, a priority order of each preset speech pausing strategy can also be set. Then, when processing the target speech recognition result for pausing, each preset speech pausing strategy is sequentially processed according to the preset priority order of the multiple preset speech pausing strategies.
[0059] Specifically, the preset priority order can be set according to actual needs. For example, the preset priority order is to process the target speech recognition result first through the first speech pausing strategy, then through the second speech pausing strategy, and finally through the third speech pausing strategy.
[0060] To facilitate understanding, the following example illustrates the processing flow of target speech recognition results by executing the first, second, and third speech cessation strategies sequentially in chronological order.
[0061] When a user's voice signal is detected, real-time speech recognition is performed on the continuously detected voice signal. When the voice signal lasts for a certain period of time (e.g., 200ms), it is determined as the voice start point T0. Starting from the voice start point, the recognition result is incrementally output in a streaming manner for the continuously detected voice signal. For example: T0+150ms: Output "guide"; T0+250ms: Output "Navigate to"; T0+350ms: Output "Navigate to North"; T0+450ms: Output "Navigate to Beijing"; At the same time, when outputting the target text fragment, the part-of-speech tag of each word is output to obtain the target part-of-speech tag sequence (e.g., "navigation"-VERB, "to"-PREP, "Beijing"-NOUN).
[0062] Upon obtaining the target speech recognition result, the first speech cessation strategy is activated. In some embodiments, the first speech cessation strategy can be activated after a first preset duration following the speech start point T0. That is, upon receiving a new target text segment, the following checks are performed: determining whether the target text segment matches any template in the preset instruction template library; if a target instruction template is matched, determining whether the length of the slot field extracted from the target text segment is greater than or equal to a preset length, and whether the corresponding position of the extracted slot field contains a keyword; if so, performing silence detection to determine whether no new speech signal input has been detected within the subsequent silence duration corresponding to the first speech cessation strategy; if all of the above are satisfied, then the speech signal is truncated. If any of the above is not satisfied, then listening continues, and the second speech cessation strategy is entered.
[0063] In some embodiments, a second speech cessation strategy may be activated if the first speech cessation strategy is not triggered and after a second preset duration following the speech start point T0. That is, based on the target part-of-speech tag sequence, it is determined whether the syntactic structure of the target text segment is complete. If it is complete, silence detection is entered to determine whether no new speech signal input is detected within the subsequent silence duration corresponding to the second speech cessation strategy. If so, the speech signal is truncated; otherwise, the third speech cessation strategy is entered.
[0064] In some embodiments, a third voice pausing strategy may be activated after a third preset duration following the voice start point T0, when the second voice pausing strategy is not triggered; that is, the S of the target text segment is calculated. total In S total If the score exceeds the preset threshold, a silence detection is initiated to determine if no new voice signal input has been detected within the subsequent silence duration corresponding to the third voice cessation strategy. If so, the voice signal is truncated. If the threshold is not met, the cessation strategy is exited and monitoring ends when the duration exceeds the limit.
[0065] In the above process, the first preset duration, the second preset duration, and the third preset duration can be set according to actual needs, and there is no limitation here. Among them, the first preset duration is less than the second preset duration, and the second preset duration is less than the third preset duration.
[0066] In the implementation of this manual, whether multiple voice pause strategies are executed in parallel or sequentially, a fallback mechanism for voice monitoring is set up to avoid endless continuous monitoring. This mechanism dynamically adjusts the monitoring duration of the subsequent silence based on different factors. Specifically, it includes: identifying the target intent of the target voice recognition result; if the target intent is a navigation intent, extending the maximum limit for subsequent voice input waiting time to a first duration; if the user's historical average command duration is greater than a preset duration, extending the maximum limit for subsequent voice input waiting time to a second duration; and determining the target confidence level of the target voice recognition result; if the target confidence level is less than a preset confidence level, extending the maximum limit for subsequent voice input waiting time to a third duration.
[0067] Specifically, when performing voice pause judgment, an initial post-mute duration can be set, which is the maximum time limit for waiting for voice input. For example, the initial post-mute duration is 2.0s.
[0068] In some embodiments, the target intent can be realized through rule matching or a preset intent analysis model, which is not limited here. When the target intent is a navigation intent, the initial post-mute duration can be increased to a first duration. The first duration can be set according to actual needs, for example, the first duration is 3s, 3.5s, etc.
[0069] In some embodiments, for each user, the duration of voice commands executed each time the user performs voice control can be recorded. By statistically analyzing the historical voice command durations, the user's historical average command duration can be obtained, and this historical average command duration can be linked to the user. When a user's historical average command duration exceeds a preset duration, the initial back-mute duration can be extended to a second duration. The preset duration can be set according to actual needs, such as 2.5s, 3s, etc. The second duration can be determined as follows: min(5.0, avg_len + 0.7s), where avg_len is the historical average command duration. That is, compare the magnitudes of 5.0s and avg_len + 0.7s, and use the smaller of the two as the second duration.
[0070] In some embodiments, when processing the speech signal using a preset speech recognition model, the target confidence level of the target speech recognition result can be output simultaneously. Of course, the target confidence level can also be determined in other ways, which are not limited here. If the target confidence level of the target speech recognition result is less than the preset confidence level, the initial post-silence duration can be extended to a third duration. The preset confidence level can be set according to actual needs, for example, a preset confidence level of 0.6, 0.65, etc. The third duration can also be set according to actual needs; for example, the third duration can be increased by 0.5 seconds based on the initial post-silence duration.
[0071] Additionally, to avoid endless monitoring, a maximum range for the back-mute duration can be set. In some implementations, the maximum back-mute duration is always limited to between 1.5 seconds and 5 seconds. If the duration of the voice conversation exceeds the current maximum back-mute duration, the monitoring is forcibly terminated.
[0072] It should be noted that the above-mentioned fallback mechanism can be executed when the preset voice stop strategy is not hit. If any voice stop strategy is hit, subsequent monitoring will be performed according to the post-silence duration of the hit voice stop strategy.
[0073] In the embodiments of this specification, if truncation of the voice signal is triggered, the timestamp corresponding to the truncated voice signal is recorded, such as T. end and [T0, T] end The corresponding audio segment is sent to the NLU (Natural Language Understanding) module for further processing.
[0074] As can be seen, the above-mentioned fallback mechanism can ensure that long commands (such as commands for complex locations or multiple actions) are not truncated, and at the same time prevent invalid voice from waiting indefinitely.
[0075] In some embodiments, the various preset voice determination strategies and fallback mechanisms in the embodiments of this specification will disable the listening permissions of other strategies and mechanisms once any strategy or mechanism triggers voice truncation, so as to ensure that only one breakpoint is adopted.
[0076] In summary, the method provided in this specification achieves low latency and high accuracy in voice signal cessation by setting multiple preset voice cessation strategies and dynamically adjusting the duration of post-silence. Specifically, the first and second semantic cessation strategies, based on lightweight instruction templates and part-of-speech tag sequences, can meet the requirements for rapid response to high-frequency instructions and instructions with specific syntactic structures. Since the first and second semantic cessation strategies consume very little computational power, and most user voice instructions can be matched by both strategies, the computational cost of the system is significantly reduced. Furthermore, for multi-turn dialogue cessation and complex dialogue, the third semantic cessation strategy comprehensively considers the semantic completeness of the user's voice and historical dialogue-assisted semantic understanding, which can significantly improve the accuracy of voice signal cessation.
[0077] Based on the same inventive concept, embodiments of this specification also provide a voice signal truncation device, such as... Figure 2 As shown, the device includes: The speech recognition module 201 is used to perform real-time speech recognition on the continuously monitored speech signal during the process of monitoring the user's speech signal, and obtain the target speech recognition result, wherein the target speech recognition result includes the recognized target text segment and the target part-of-speech tag sequence corresponding to the target text segment; The first processing module 202 is used to perform stop processing on the target speech recognition result based on multiple preset speech stop strategies. If a target speech stop strategy determines that the target speech recognition result is a complete instruction, the target post-silence duration corresponding to the target speech stop strategy is used as the maximum limit duration for subsequent waiting for speech input. The post-silence duration corresponding to each preset speech stop strategy is different. The second processing module 203 is used to truncate the voice signal when no new voice signal input is detected within the target post-silence duration.
[0078] Regarding the above-mentioned apparatus, the specific implementation of each step has been described in detail in the embodiments of the speech signal truncation method provided in the specification, and will not be elaborated here.
[0079] Based on the same inventive concept, embodiments of the present invention also provide a voice signal truncation device, such as... Figure 3 The aforementioned includes a memory 304, a processor 302, and a computer program stored in the memory 304 and executable on the processor 302. When the processor 302 executes the program, it implements any of the above-described methods for truncating voice signals.
[0080] Among them, Figure 3 In this document, a bus architecture (represented by bus 300) is used. Bus 300 may include any number of interconnected buses and bridges, linking various circuits including one or more processors represented by processor 302 and memory represented by memory 304. Bus 300 may also link various other circuits such as peripheral devices, voltage regulators, and power management circuits, which are well known in the art and therefore will not be described further herein. Bus interface 305 provides an interface between bus 300 and receiver 301 and transmitter 303. Receiver 301 and transmitter 303 may be the same element, i.e., a transceiver, providing a unit for communicating with various other devices over a transmission medium. Processor 302 is responsible for managing bus 300 and general processing, while memory 304 can be used to store data used by processor 302 during operation.
[0081] It should be noted that the descriptions of each embodiment in the above embodiments have different focuses. For parts that are not described in detail in a certain embodiment, please refer to the relevant descriptions in other embodiments.
[0082] Based on the same inventive concept, embodiments of this specification provide a computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements the steps of the above-described speech signal truncation method.
[0083] Based on the same inventive concept, embodiments of this specification provide a computer program product, which includes a computer program that, when executed by a processor, loads and executes the steps of the above-described voice signal truncation method.
[0084] The functions described herein can be implemented in hardware, software executed by a processor, firmware, or any combination thereof. If implemented in software executed by a processor, the functions can be stored as one or more instructions or code on or transmitted via a computer-readable medium. Other examples and embodiments are within the scope of this invention and the appended claims. For example, due to the nature of software, the functions described above can be implemented using software executed by a processor, hardware, firmware, hardwired, or any combination thereof. Furthermore, the functional units can be integrated into a single processing unit, or each unit can exist physically separately, or two or more units can be integrated into a single unit.
[0085] In the several embodiments provided in this application, it should be understood that the disclosed technical content can be implemented in other ways. The device embodiments described above are merely illustrative; for example, the division of units can be a logical functional division, and in actual implementation, there may be other division methods. For instance, multiple units or components may be combined or integrated into another system, or some features may be ignored or not executed. Furthermore, the displayed or discussed mutual coupling, direct coupling, or communication connection may be through some interfaces; the indirect coupling or communication connection between units or modules may be electrical or other forms.
[0086] The units described as separate components may or may not be physically separate. Similarly, the components of the control device may or may not be physical units; they may be located in one place or distributed across multiple units. Some or all of the units can be selected to achieve the purpose of this embodiment, depending on actual needs.
[0087] If the integrated unit is implemented as a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present invention, in essence, or the part that contributes to the prior art, or all or part of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute all or part of the steps of the methods described in the various embodiments of the present invention. The aforementioned storage medium includes various media capable of storing program code, such as USB flash drives, read-only memory (ROM), random access memory (RAM), portable hard drives, magnetic disks, or optical disks.
[0088] The above description is merely an embodiment of the present invention and is not intended to limit the invention. Various modifications and variations can be made to the present invention by those skilled in the art. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of the present invention should be included within the scope of the claims of the present invention.
Claims
1. A method for truncating a speech signal, characterized in that, include: During the process of monitoring the user's voice signal, real-time speech recognition is performed on the continuously monitored voice signal to obtain the target speech recognition result. The target speech recognition result includes the recognized target text segment and the target part-of-speech tag sequence corresponding to the target text segment. The target speech recognition result is processed for stopping based on multiple preset speech stopping strategies. If a target speech stopping strategy determines that the target speech recognition result is a complete instruction, the target post-silence duration corresponding to the target speech stopping strategy is used as the maximum limit duration for subsequent waiting for speech input. The post-silence duration corresponding to each preset speech stopping strategy is different. If no new voice signal input is detected within the target's post-silence duration, the voice signal is truncated.
2. The method as described in claim 1, characterized in that, The multiple preset voice stop strategies include a first voice stop strategy based on instruction templates for complete instruction recognition, a second voice stop strategy based on part-of-speech tag sequences for complete instruction recognition, and a third voice stop strategy based on semantics for complete instruction recognition.
3. The method as described in claim 2, characterized in that, The post-mute durations corresponding to the first, second, and third voice stop strategies increase sequentially.
4. The method as described in claim 1 or 2, characterized in that, The step of performing stop detection processing on the target speech recognition result based on multiple preset speech stop detection strategies includes: According to the preset priority order of the multiple preset voice pausing strategies, each preset voice pausing strategy sequentially performs pausing processing on the target voice recognition result.
5. The method as described in claim 2, characterized in that, The step of performing stop detection processing on the target speech recognition result based on multiple preset speech stop detection strategies includes: During the process of pausing the target speech recognition result based on the first speech pausing strategy, the target text segment is matched with a preset instruction template library; If a target instruction template that matches the target text fragment exists in the preset instruction template library, the target field corresponding to the target slot is extracted from the target text fragment based on the target slot corresponding to the target instruction template; If the length of the target field is greater than or equal to the preset length, and the target field contains the preset keyword corresponding to the target slot, then the target speech recognition result is determined to be a complete instruction, and the first speech pausing strategy is used as the target speech pausing strategy.
6. The method as described in claim 2, characterized in that, The step of performing stop detection processing on the target speech recognition result based on multiple preset speech stop detection strategies includes: During the process of pausing the target speech recognition result based on the second speech pausing strategy, the syntactic integrity of the target part-of-speech tag sequence is determined. If the target part-of-speech tag sequence is syntactically complete, the target speech recognition result is determined to be a complete instruction, and the second speech pausing strategy is used as the target speech pausing strategy.
7. The method as described in claim 6, characterized in that, The syntactic integrity determination of the target part-of-speech tag sequence includes: If the last two consecutive part-of-speech tags of the target part-of-speech tag sequence are verbs and nouns, or verbs and pronouns, the target part-of-speech tag sequence is determined to be a complete imperative sentence syntax; If the first three parts of speech tags of the target part of speech tag sequence are noun, verb and noun in sequence, or the first three parts of speech tags are noun, verb and pronoun in sequence, the target part of speech tag sequence is determined to be a complete subject-verb-object sentence; If the last two consecutive part-of-speech tags of the target part-of-speech tag sequence are a preposition and a noun respectively, the target part-of-speech tag sequence is determined to be a complete prepositional complement syntax.
8. The method as described in claim 2, characterized in that, The step of performing stop detection processing on the target speech recognition result based on multiple preset speech stop detection strategies includes: In the process of pausing the target speech recognition result based on the third speech pausing strategy, a preset semantic recognition model is used to process the target text segment to obtain a first score that characterizes the semantic integrity of the target text segment. Based on the voice text segment from the previous round of voice interaction, the correlation between the target text segment and the voice text segment is determined to obtain a second score; Based on the first score and the second score, a target score is obtained. When the target score is greater than a preset score, the target speech recognition result is determined to be a complete instruction, and the third speech pausing strategy is used as the target speech pausing strategy.
9. The method as described in claim 1, characterized in that, If there is no target speech pausing strategy that determines the target speech recognition result as a complete instruction, the method further includes: Identify the target intent of the target speech recognition result. If the target intent is a navigation intent, extend the maximum time limit for subsequent waiting for speech input to the first time limit. If the user's historical average command duration is greater than the preset duration, the maximum time limit for waiting for subsequent voice input will be extended to the second duration. Determine the target confidence level of the target speech recognition result. If the target confidence level is less than the preset confidence level, extend the maximum time limit for subsequent waiting for speech input to a third time limit.
10. A voice signal truncation device, characterized in that, include: The speech recognition module is used to perform real-time speech recognition on continuously monitored speech signals during the process of monitoring the user's speech signals, and to obtain the target speech recognition result. The target speech recognition result includes the recognized target text segment and the target part-of-speech tag sequence corresponding to the target text segment. The first processing module is used to perform stop processing on the target speech recognition result based on multiple preset speech stop strategies. If a target speech stop strategy determines that the target speech recognition result is a complete instruction, the target post-silence duration corresponding to the target speech stop strategy is used as the maximum limit duration for subsequent waiting for speech input. The post-silence duration corresponding to each preset speech stop strategy is different. The second processing module is used to truncate the voice signal when no new voice signal input is detected within the target post-silence duration.
11. A voice signal truncation device, characterized in that, It includes a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor executes the program to implement the steps of the method as described in any one of claims 1-9.
12. A computer-readable storage medium, characterized in that, The device contains a computer program that, when executed by a processor, implements the steps of the method according to any one of claims 1-9.