Voice processing method, electronic equipment and storage medium

By combining cloud VAD and text detection models on the cloud server side, the speech recognition process is dynamically adjusted, and the recognition error problem caused by fixed periods of mute in the existing technology is solved, and accurate recognition and rapid response of speech is achieved.

CN120279920APending Publication Date: 2025-07-08AISPEECH CO LTD
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
CN202510570764.2
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-30
Publication Date
2025-07-08

AI Technical Summary

Technical Problem

In the existing VAD processing scheme, the mute period is fixed, resulting in the inability to accurately recognize the beginning and end points of the voice, especially when the pause time exceeds the preset time, the final recognition result is triggered, resulting in the inability to perform effective voice recognition.

Method used

After receiving the voice audio stream and vad timestamp information uploaded by the client on the cloud server, intermediate recognition and decoding is performed through cloud VAD, combining the text detection model to determine whether the recognition result is a complete semantic sentence, and splicing the voice audio stream if necessary, using dialogue central control to perform product-level semantic settlement and skill routing, and dynamically adjust the recognition process.

Benefits of technology

It realizes accurate recognition and splicing of voice in any scenario, especially in fast response scenarios such as on-board control, which improves recognition efficiency, ensuring the rapid response and accurate processing of voice commands.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120279920A_ABST
    Figure CN120279920A_ABST
Patent Text Reader

Abstract

The embodiment of the invention discloses a voice processing method, electronic equipment and a storage medium. The method comprises the following steps: receiving vad timestamp information, a voice start message and a voice audio stream uploaded by a client after effective audio detection; in a first preset pause duration after the audio falling edge timestamp of the current effective audio, triggering intermediate recognition decoding based on the cloud vad to obtain an intermediate recognition result, and calling a text detection model to judge whether the intermediate recognition result is a complete semantic sentence; if the intermediate recognition result is not a complete semantic sentence and a subsequent voice audio stream appears within a second preset pause duration after the audio falling edge timestamp of the current effective audio, splicing the intermediate recognition result and a subsequent text of the subsequent voice audio stream; if the sentence is a complete semantic sentence, the intermediate recognition result carries parameters and is transferred to a dialogue center controller, product-level semantic settlement is triggered based on the dialogue center controller, skill routing is completed to obtain a skill routing field, and whether callback is carried out or not is judged according to the skill routing field.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the field of speech processing, and in particular, relates to a speech processing method, an electronic device, and a storage medium. Background Art

[0002] VAD, that is, voice activity detection technology, is the abbreviation of Voice Activity Detection. The main task of this technology is to accurately locate the start and end points of speech from noisy speech. Since there is a long silence in speech, that is, to separate the silence from the actual speech. Because it is the original processing of speech data, VAD is one of the key technologies in the speech signal processing process. Its quality directly affects the success or failure. Due to the particularity of the technology itself, in the field of speech signal processing, the application of endpoint detection technology is very extensive. The first technology encountered in the recognition or acoustic model training stage of a speech recognition system is endpoint detection, which removes silence and noise as interference signals from the original data, and endpoint detection is crucial for the performance of the speech recognition system.

[0003] However, the pause time of the existing conventional vad processing scheme is fixed. The subsequent silence segment time of the input valid audio must meet a certain time to trigger the sending of the recognition end frame and then there will be a final recognition result. Once the subsequent silence segment of the input valid audio exceeds the pause time, it will be considered a complete recognition and the calculation of the final recognition result will be triggered. For example, if I want to go to a certain park... (pause time) exceeds the pause time of the conventional vad processing scheme, the calculation of the final recognition result will be triggered, resulting in the inability to recognize. Summary of the Invention

[0004] Embodiments of the present invention provide a speech processing method, an electronic device, and a storage medium, which are used to solve at least one of the above technical problems.

[0005] In a first aspect, an embodiment of the present invention provides a voice processing method for a cloud server side, including: receiving vad timestamp information, a voice start message, and a voice audio stream after valid audio detection uploaded by a client, where the vad timestamp information includes an audio rising edge timestamp and an audio falling edge timestamp; at a first preset pause duration after the audio falling edge timestamp of the current valid audio, triggering intermediate recognition and decoding based on cloud vad to obtain an intermediate recognition result, and invoking a text detection model to determine whether the intermediate recognition result is a complete semantic sentence; if the intermediate recognition result is not a complete semantic sentence, and a subsequent voice audio stream appears within a second preset pause duration after the audio falling edge timestamp of the current valid audio, splicing the intermediate recognition result and the subsequent text of the subsequent voice audio stream, where the second preset pause duration is greater than the first preset pause duration; if the intermediate recognition result is a complete semantic sentence, carrying parameters of the intermediate recognition result and transferring it to a dialogue control center, triggering product-level semantic settlement based on the dialogue control center and completing skill routing to obtain a skill routing domain, and determining whether to perform a callback according to the skill routing domain.

[0006] In a second aspect, an embodiment of the present invention further provides an electronic device, which includes: at least one processor, and a memory communicatively connected to the at least one processor, where the memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor to enable the at least one processor to execute the steps of the method in the first aspect.

[0007] In a third aspect, an embodiment of the present invention further provides a storage medium, on which a computer program is stored, characterized in that the computer program, when executed by a processor, implements the steps of the method in the first aspect.

[0008] In the method of the embodiment of the present application, after receiving the vad timestamp information, voice start message, and voice audio stream after valid audio detection uploaded by the client through the access service (ddsserver - fullduplex) on the cloud server side, when it is detected that the audio falling edge timestamp of the current valid audio reaches the first preset time, the cloud vad is used to identify and decode the current valid audio to obtain an intermediate recognition result (eof = 0 rec). After obtaining the intermediate recognition result, the text detection model is used to detect whether the intermediate recognition result is a complete semantic sentence. If it is detected that the intermediate recognition result is an incomplete semantic sentence, and there is a subsequent voice audio stream before the audio falling edge timestamp exceeds the second preset pause duration, the intermediate recognition result is concatenated with the subsequent text of the subsequent voice audio stream, where the second preset pause duration is greater than the first preset pause duration. If the recognition result is a complete semantic sentence, the intermediate recognition result is carried with parameters and transferred to the dialogue control center through the access service. After obtaining the intermediate recognition result, the dialogue control center performs product - level semantic settlement and completes skill routing to obtain the skill routing domain, and finally determines whether to perform a callback according to the skill routing domain, so that the present application can perform concatenation in any scenario. BRIEF DESCRIPTION OF THE DRAWINGS

[0009] In order to more clearly illustrate the technical solutions of the embodiments of the present invention, the following will briefly introduce the drawings required for the description of the embodiments. Obviously, the following drawings are some embodiments of the present invention. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.

[0010] Figure 1 It is a flowchart of a voice processing method for a cloud server side provided by an embodiment of the present invention; Figure 2 It is a flowchart of a voice processing method for a dialogue control center provided by an embodiment of the present invention; Figure 3 It is a flowchart of a voice processing method for a cloud vad provided by an embodiment of the present invention; Figure 4 It is a flowchart of a voice processing method for a client provided by an embodiment of the present invention; Figure 5 It is a full - link flowchart of a voice processing method provided by an embodiment of the present invention; Figure 6 It is a time - node schematic diagram of a voice processing method provided by an embodiment of the present invention; Figure 7 It is an audio schematic diagram of an alternative solution of a voice processing method provided by an embodiment of the present invention; Figure 8A framework diagram of a voice processing method provided by an embodiment of the present invention; Figure 9 An audio schematic diagram of a voice processing method provided by an embodiment of the present invention; Figure 10 A structural schematic diagram of an electronic device provided by an embodiment of the present invention. Detailed implementation manners

[0011] To make the objectives, technical solutions, and advantages of the embodiments of the present invention clearer, the technical solutions in the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings in the embodiments of the present invention. Apparently, the described embodiments are some but not all of the embodiments of the present invention. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.

[0012] It should be noted that, without conflict, the embodiments in the present application and the features in the embodiments may be combined with each other.

[0013] The present invention may be described in the general context of computer-executable instructions executed by a computer, such as program modules. Generally, program modules include routines, programs, objects, components, data structures, etc. that perform specific tasks or implement specific abstract data types. The present invention may also be practiced in a distributed computing environment where tasks are performed by remote processing devices connected through a communication network. In a distributed computing environment, program modules may be located in local and remote computer storage media including storage devices.

[0014] In the present invention, "module", "device", "system", etc. refer to related entities applied to a computer, such as hardware, a combination of hardware and software, software, or software in execution. Specifically, for example, an element may be, but is not limited to, a process running on a processor, a processor, an object, an executable element, an execution thread, a program, and / or a computer. Additionally, an application program or a script program running on a server, and the server may both be elements. One or more elements may be in a process and / or thread in execution, and the elements may be localized on one computer and / or distributed between two or more computers, and may be run by various computer-readable media. The elements may also communicate according to a signal having one or more data packets, for example, a signal from a data signal that interacts with another element in a local system, a distributed system, and / or communicates with other systems through a network on the Internet through local and / or remote processes.

[0015] Finally, it should also be noted that in this text, relational terms such as first and second are only used to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any actual relationship or order between these entities or operations. Moreover, the terms "comprising" and "including" not only include those elements, but also other elements not explicitly listed, or elements inherent to such a process, method, article, or device. Without further limitation, an element defined by the statement "comprising..." does not exclude the existence of additional identical elements in the process, method, article, or device that includes the said element.

[0016] An embodiment of the present invention provides a voice processing method, which can be applied to an electronic device. The electronic device can be a computer, a server, or other electronic products, and this application does not make any limitations in this regard.

[0017] Please refer to Figure 1 , which shows a flowchart of a voice processing method for a cloud server side provided by an embodiment of the present invention.

[0018] As Figure 1 shown, in step 101, receive the vad timestamp information, voice start message, and voice audio stream after valid audio detection uploaded by the client, where the vad timestamp information includes an audio rising edge timestamp and an audio falling edge timestamp; In step 102, at the first preset pause duration after the audio falling edge timestamp of the current valid audio, trigger intermediate recognition and decoding based on cloud vad to obtain an intermediate recognition result, and call a text detection model to determine whether the intermediate recognition result is a complete semantic sentence; In step 103, if the intermediate recognition result is not a complete semantic sentence and a subsequent voice audio stream appears within the second preset pause duration after the audio falling edge timestamp of the current valid audio, splice the subsequent text of the intermediate recognition result and the subsequent voice audio stream, where the second preset pause duration is greater than the first preset pause duration; In step 104, if the intermediate recognition result is a complete semantic sentence, transfer the intermediate recognition result with parameters to the dialogue control center, trigger product-level semantic settlement based on the dialogue control center and complete skill routing to obtain a skill routing domain, and determine whether to perform a callback according to the skill routing domain.

[0019] In this embodiment, for step 101, the cloud server receives the VAD timestamp information, voice start message, and voice audio stream after valid audio detection uploaded by the client. The VAD timestamp information includes the audio rising edge timestamp and the audio falling edge timestamp. For example, the cloud server receives the VAD timestamp information, voice start message, and voice audio stream after valid audio detection uploaded by the client through the access service (ddsserver-fullduplex). The VAD timestamp information includes the audio rising edge timestamp and the audio falling edge timestamp. The client can be an intelligent device such as a computer, a mobile phone, or a vehicle-mounted central control unit.

[0020] Then, for step 102, at the first preset pause duration after the audio falling edge timestamp of the current valid audio, the cloud VAD is triggered to perform intermediate recognition and decoding to obtain an intermediate recognition result, and a text detection model is called to determine whether the intermediate recognition result is a complete semantic sentence. For example, after detecting that the audio falling edge timestamp of the current valid audio reaches the first preset time, the cloud VAD performs recognition and decoding on the current valid audio to obtain an intermediate recognition result (eof=0 rec). After obtaining the intermediate recognition result, the text detection model is used to detect whether the intermediate recognition result is a complete semantic sentence, so as to determine whether the previous input has been finished.

[0021] Then, for step 103, if the intermediate recognition result is not a complete semantic sentence and subsequent voice audio streams appear within the second preset pause duration after the audio falling edge timestamp of the current valid audio, the subsequent text of the intermediate recognition result and the subsequent voice audio streams is spliced. The second preset pause duration is greater than the first preset pause duration. For example, if it is detected that the intermediate recognition result is an incomplete semantic sentence and subsequent voice audio streams appear before the audio falling edge timestamp exceeds the second preset pause duration, the subsequent text of the intermediate recognition result and the subsequent voice audio streams is spliced. The second preset pause duration is greater than the first preset pause duration, so that the unfinished voice can be spliced after pausing for a period of time.

[0022] Finally, for step 104, if the intermediate recognition result is a complete semantic sentence, the intermediate recognition result is carried with parameters and transferred to the dialogue central control. Based on the trigger of the dialogue central control, product-level semantic settlement is performed and skill routing is completed to obtain the skill routing domain. According to the skill routing domain, it is determined whether to perform a callback. For example, if the recognition result is a complete semantic sentence, the intermediate recognition result is carried with parameters and transferred to the dialogue central control through the access service. After obtaining the intermediate recognition result, the dialogue central control performs product-level semantic settlement and completes skill routing to obtain the skill routing domain. Finally, it is determined whether to perform a callback according to the skill routing domain.

[0023] In the method of the embodiment of the present application, after receiving the vad timestamp information, voice start message, and voice audio stream after valid audio detection uploaded from the client on the cloud server side through the access service (ddsserver-fullduplex), when it is detected that the audio falling edge timestamp of the current valid audio reaches the first preset time, the current valid audio is identified and decoded by cloud vad to obtain an intermediate recognition result (eof = 0 rec). After obtaining the intermediate recognition result, it is detected whether the intermediate recognition result is a complete semantic sentence through a text detection model. If it is detected that the intermediate recognition result is an incomplete semantic sentence and a subsequent voice audio stream appears before the audio falling edge timestamp exceeds the second preset pause duration, the intermediate recognition result is concatenated with the subsequent text of the subsequent voice audio stream, where the second preset pause duration is greater than the first preset pause duration. If the recognition result is a complete semantic sentence, the intermediate recognition result is carried with parameters and transferred to the dialogue control center through the access service. After obtaining the intermediate recognition result, the dialogue control center performs product-level semantic settlement and completes skill routing to obtain the skill routing domain, and finally determines whether to perform a callback based on the skill routing domain, so that the present application can perform concatenation in any scenario.

[0024] In some alternative embodiments, after, if the intermediate recognition result is a complete semantic sentence, carrying the intermediate recognition result with parameters and transferring it to the dialogue control center, triggering product-level semantic settlement and completing skill routing to obtain the skill routing domain by the dialogue control center, and determining whether to perform a callback based on the skill routing domain, it includes: determining whether the skill routing domain is a domain related to vehicle control. If it is a domain related to vehicle control, the callback result is sent to the cloud vad. If it is not a domain related to vehicle control, the callback result is not sent. For example, after obtaining the domain of the current conversation through skill routing, it is determined whether it is a command in the vehicle control domain. If it is a vehicle control type domain, the callback is triggered and the callback result is sent to the cloud vad. If it does not belong to the vehicle control type domain, no callback is performed, so as to ensure that commands related to vehicle control can be quickly responded to.

[0025] In some alternative embodiments, after the intermediate recognition result is a complete semantic sentence, the method includes: If the cloud VAD receives the callback result of the dialogue control within the third preset pause duration after the audio falling edge timestamp of the valid audio, it immediately sends an audio end frame and triggers the subsequent dialogue, where the third pause duration is greater than the first pause duration; if the cloud VAD does not receive the callback result of the control within the third preset pause duration after the audio falling edge timestamp of the valid audio, it sends an audio end frame at the third preset pause duration after the audio falling edge timestamp of the valid audio to obtain the final recognition result. For example, after obtaining the field of the complete semantic sentence, it detects whether the dialogue control receives the callback result at the third preset pause duration. If it receives the callback result of the dialogue control within the third preset pause duration, it immediately sends an audio end frame when receiving the callback result, and then starts to process the next dialogue. If it does not receive the callback result within the third preset pause duration, it sends an audio end frame at the third preset pause duration after the audio falling edge timestamp of the valid audio to obtain the final recognition result, so as to ensure that the valid audio that can trigger the callback result can respond quickly.

[0026] In some alternative embodiments, after obtaining the final recognition result, the method further includes: after the dialogue control receives the final recognition result, in combination with the VAD timestamp information, it determines whether the interval between the audio rising edge timestamp of the final recognition result and the previous audio falling edge timestamp meets the fourth preset pause duration, where the fourth pause duration is greater than the second preset pause duration; if it meets, it splices the previous recognition text with the final recognition result and sends it to the splicing discrimination model. If the spliced recognition result can respond, it outputs a response label and outputs the result after splicing. If the spliced recognition result cannot respond and the final recognition result can respond, it directly outputs the final recognition result without splicing; if it does not meet, it does not perform any processing and temporarily stores it in the cloud. For example, after the dialogue control obtains the final recognition result, it obtains the VAD timestamp information of the previous audio and the VAD timestamp information of this audio, and determines whether the rising edge timestamp of the final recognition result and the falling edge timestamp information of the previous audio are within the fourth preset duration, where the fourth pause duration is greater than the second preset pause duration. If it does not meet the fourth preset duration, it does not perform any processing and temporarily stores it in the cloud for use as the previous audio for subsequent audio splicing. If it meets the fourth preset pause duration, it first combines the current valid audio and the previous audio and inputs them into the model to determine whether they can respond. If they can respond, the splicing is completed, and then the spliced result is output. If only the current valid audio can respond, the final recognition result is directly output without splicing, so as to ensure that the situation of not recognizing the user's intention does not occur.

[0027] In some alternative embodiments, if the intermediate recognition result is not a complete semantic sentence and subsequent speech audio streams appear within a second preset pause duration after the audio falling-edge timestamp of the current valid audio, splicing the subsequent text of the intermediate recognition result and the subsequent speech audio streams further includes: if no subsequent audio appears within the second preset pause duration, ending the process and temporarily storing it in the cloud. For example, after determining that the intermediate recognition result is an incomplete sentence, if no subsequent audio appears within the second pause duration, ending the process and storing it in the cloud for use as the upper-segment audio for subsequent audio splicing.

[0028] Please refer to Figure 2 , which shows a flowchart of a voice processing method for dialogue central control provided by an embodiment of the present invention; As Figure 2 shown, in step 201, a complete semantic sentence transmitted by cloud vad is received; In step 202, product-level semantic settlement is triggered through the dialogue central control and a skill routing domain is obtained by completing skill routing. According to the skill routing domain, it is determined whether it is a vehicle control-related instruction; In step 203, if it is a vehicle control-related instruction, the callback result is sent to the cloud vad. If it is not in the field related to vehicle control, the callback result is not sent.

[0029] In this embodiment, for step 201, a complete semantic sentence transmitted by cloud vad is received. For example, after cloud vad determines that the intermediate recognition result belongs to a complete semantic sentence, the intermediate recognition result is input into the dialogue central control, so as to ensure that the sentence processed by the dialogue central control is a complete semantic sentence.

[0030] Then, for step 202, product-level semantic settlement is triggered through the dialogue central control and a skill routing domain is obtained by completing skill routing. According to the skill routing domain, it is determined whether it is a vehicle control-related instruction. For example, after obtaining the complete semantic sentence transmitted by cloud vad, semantic settlement is performed on the complete semantic sentence, and the skill routing domain of the complete semantic sentence is obtained through skill routing and it is determined whether to perform a callback according to the domain.

[0031] Finally, for step 203, if it is a vehicle control-related instruction, the callback result is sent to the cloud vad. If it is not in the field related to vehicle control, the callback result is not sent. For example, if it is detected that the obtained skill routing domain belongs to the vehicle control field, a callback is triggered and the callback result is sent to cloud vad. If it is detected that the obtained skill routing domain does not belong to the vehicle control field, the callback result is not sent, so as to ensure that commands belonging to the vehicle control field can be quickly transmitted to cloud vad for subsequent operations.

[0032] In the method of the embodiment of the present application, after the cloud VAD determines that the intermediate recognition result belongs to a complete semantic sentence, the intermediate recognition result is input to the dialogue control center. After obtaining the complete semantic sentence transmitted by the cloud VAD, semantic settlement is performed on the complete semantic sentence, and the skill route field of the complete semantic sentence is obtained through the skill route, and it is determined whether to perform a callback according to the field. If it is detected that the obtained skill route field belongs to the vehicle-mounted control field, the callback is triggered, and the callback result is sent to the cloud VAD. If it is detected that the obtained skill route field does not belong to the vehicle-mounted control field, the callback result is not sent, so as to ensure that commands belonging to the vehicle-mounted control field can be quickly transmitted to the cloud VAD for subsequent operations.

[0033] Please refer to Figure 3 , which shows a flowchart of a voice processing method for cloud VAD provided by an embodiment of the present invention.

[0034] As Figure 3 shown, in step 301, the VAD timestamp information, voice start message, and voice audio stream after valid audio detection uploaded by the client are received, where the VAD timestamp information includes an audio rising edge timestamp and an audio falling edge timestamp; In step 302, at the first preset pause duration after the audio falling edge timestamp of the current valid audio, intermediate recognition decoding is triggered based on the cloud VAD to obtain an intermediate recognition result, and a text detection model is called to determine whether the intermediate recognition result is a complete semantic sentence; In step 303, if the intermediate recognition result is not a complete semantic sentence, and subsequent voice audio streams appear within the second preset pause duration after the audio falling edge timestamp of the current valid audio, the subsequent text of the intermediate recognition result and the subsequent voice audio streams is spliced, where the second preset pause duration is greater than the first preset pause duration; In step 304, if the intermediate recognition result is a complete semantic sentence, the intermediate recognition result is carried with parameters and transferred to the dialogue control center.

[0035] In this embodiment, for step 301, the VAD timestamp information, voice start message, and voice audio stream after valid audio detection uploaded by the client are received, where the VAD timestamp information includes an audio rising edge timestamp and an audio falling edge timestamp. For example, the cloud server receives the VAD timestamp information, voice start message, and voice audio stream after valid audio detection uploaded by the client through the access service (ddsserver-fullduplex), where the VAD timestamp information includes an audio rising edge timestamp and an audio falling edge timestamp, and the client can be an intelligent device such as a computer, a mobile phone, or a vehicle-mounted control panel.

[0036] Then, for step 302, at the first preset pause duration after the audio falling edge timestamp of the current valid audio, trigger intermediate recognition and decoding based on cloud VAD to obtain an intermediate recognition result, and call a text detection model to determine whether the intermediate recognition result is a complete semantic sentence. For example, after detecting that the audio falling edge timestamp of the current valid audio reaches the first preset time, use cloud VAD to perform recognition and decoding on the current valid audio to obtain an intermediate recognition result (eof=0 rec). After obtaining the intermediate recognition result, use the text detection model to detect whether the intermediate recognition result is a complete semantic sentence, so as to determine whether the previous input has finished speaking.

[0037] Then, for step 303, if the intermediate recognition result is not a complete semantic sentence and a subsequent voice audio stream appears within the second preset pause duration after the audio falling edge timestamp of the current valid audio, splice the intermediate recognition result and the subsequent text of the subsequent voice audio stream, where the second preset pause duration is greater than the first preset pause duration. For example, if it is detected that the intermediate recognition result is an incomplete semantic sentence and a subsequent voice audio stream appears before the audio falling edge timestamp exceeds the second preset pause duration, splice the intermediate recognition result and the subsequent text of the subsequent voice audio stream, where the second preset pause duration is greater than the first preset pause duration, so that the unfinished speech can be spliced after pausing for a period of time.

[0038] In step 304, if the intermediate recognition result is a complete semantic sentence, transfer the intermediate recognition result with parameters to the dialogue control center. For example, if the recognition result is a complete semantic sentence, transfer the intermediate recognition result with parameters to the dialogue control center through the access service.

[0039] In the method of the embodiment of the present application, the cloud server receives the vad timestamp information, voice start message, and voice audio stream after valid audio detection uploaded by the client through the access service (ddsserver-fullduplex). Among them, the vad timestamp information includes the audio rising edge timestamp and the audio falling edge timestamp. The client can be an intelligent device such as a computer, a mobile phone, or a vehicle-mounted central control. After detecting that the audio falling edge timestamp of the current valid audio reaches the first preset time, the cloud vad is used to identify and decode the current valid audio to obtain an intermediate recognition result (eof = 0 rec). After obtaining the intermediate recognition result, the text detection model is used to detect whether the intermediate recognition result is a complete semantic sentence. If it is detected that the intermediate recognition result is an incomplete semantic sentence and a subsequent voice audio stream appears before the audio falling edge timestamp exceeds the second preset pause duration, the intermediate recognition result and the subsequent text of the subsequent voice audio stream are spliced. Among them, the second preset pause duration is greater than the first preset pause duration, so that the unfinished voice can be spliced after pausing for a period of time. If the recognition result is a complete semantic sentence, the intermediate recognition result is carried with parameters and transferred to the dialogue central control through the access service.

[0040] In some optional embodiments, after the intermediate recognition result is a complete semantic sentence and the intermediate recognition result is carried with parameters and transferred to the dialogue central control, the following steps are further included: obtaining the callback result in the dialogue and performing corresponding operations according to the callback result; if the cloud vad receives the callback result from the dialogue central control within the third preset pause duration after the audio falling edge timestamp of the valid audio, an audio end frame is sent in real time and the subsequent dialogue is triggered; if the cloud vad does not receive the callback result from the central control within the third preset pause duration after the audio falling edge timestamp of the valid audio, an audio end frame is sent at the third preset pause duration after the audio falling edge timestamp of the valid audio to obtain the final recognition result. For example, when the cloud vad obtains the callback result in the dialogue and processes according to the obtained result, if the callback result from the dialogue central control is received within the third preset pause duration, an audio end frame is sent in real time when the callback result is received, and then the next dialogue is processed. If the callback result is not received within the third preset pause duration, an audio end frame is sent at the third preset pause duration after the audio falling edge timestamp of the valid audio to obtain the final recognition result, so as to ensure that the valid audio that can trigger the callback result can respond quickly.

[0041] Please refer to Figure 4 , which shows a flowchart of a voice processing method for a client provided by an embodiment of the present invention. As Figure 4As shown, in step 401, after obtaining the audio information, voice detection is performed to determine whether it is valid audio. If it is valid audio, the vad timestamp information, voice start message, and voice audio stream of the valid audio are obtained and uploaded to the cloud server for processing; In step 402, the processing result of the cloud server is obtained and corresponding operations are performed according to the processing result.

[0042] In this embodiment, for step 401, after obtaining the audio information, voice detection is performed to determine whether it is valid audio. If it is valid audio, the vad timestamp information, voice start message, and voice audio stream of the valid audio are obtained and uploaded to the cloud server for processing. For example, after the client obtains the audio information, it detects the audio information to confirm that it is valid audio, and obtains the voice start message, voice audio stream, and vad timestamp information of the valid audio and transmits them to the cloud for processing.

[0043] Then, for step 402, the processing result of the cloud server is obtained and corresponding operations are performed according to the processing result. For example, after the cloud server finishes processing, the processing result of the cloud server is received and corresponding operations are performed according to the processing result.

[0044] In the method of this application embodiment, after the client obtains the audio information, it detects the audio information to confirm that it is valid audio, and obtains the voice start message, voice audio stream, and vad timestamp information of the valid audio and transmits them to the cloud for processing. After the cloud server finishes processing, the processing result of the cloud server is received and corresponding operations are performed according to the processing result, thereby reducing the energy consumption of the client.

[0045] Please refer to Figure 5 , which shows the full-link flowchart of a voice processing method provided by an embodiment of the present invention.

[0046] As Figure 5 shown, in step 1, the client uploads valid audio, the vad timestamp information of the valid audio, the voice start message, and the voice audio stream, where the vad timestamp information includes the audio rising edge timestamp and the audio falling edge timestamp.

[0047] In step 2, the valid audio, the vad timestamp information of the valid audio, the voice start message, and the voice audio stream uploaded by the client are received, and decoding is performed to obtain an intermediate recognition result when the audio falling edge timestamp of the valid audio reaches the T1 moment (the first preset pause duration).

[0048] Step 3: When the intermediate recognition result is obtained, use the TED model (text detection model) to identify the intermediate recognition result to determine whether the intermediate recognition result is a complete semantic sentence. If it is an incomplete semantic sentence, go to Step 4; if it is a complete semantic sentence, go to Step 5.

[0049] Step 4: If it is an incomplete semantic sentence, wait for subsequent audio for splicing within T2 (the second preset pause duration). If there is subsequent audio within T2, splice the subsequent audio and output it to the client for corresponding operations or answering questions. If there is no subsequent audio, reject the recognition and temporarily store it in the cloud.

[0050] Step 5: If it is a complete semantic sentence, input it to the dialogue control center for processing. After receiving the sentence, the dialogue control center performs product-level semantic recognition on the sentence, and inputs the recognition result to the skill route to obtain the field to which the sentence belongs, and determines whether to perform a callback based on the field. For example, a callback is performed for those belonging to the vehicle control field, while no callback is performed for those not belonging to it.

[0051] Step 6: Detect whether the cloud VAD receives a callback result within T3 (the third preset pause duration). If a callback result is received, go to Step 7; if no callback result is received, go to Step 8.

[0052] Step 7: When a callback result is received within T3, send the audio end frame and related operations to the client, and the client performs corresponding operations and triggers subsequent conversations.

[0053] Step 8: If no callback result is received within T3, send the audio end frame, obtain the final recognition result, and at the same time obtain the VAD timestamp information of the current audio and the VAD timestamp information of the previous audio.

[0054] Step 9: Detect whether the rising edge timestamp of the current audio is within T4 (the fourth preset pause duration) of the falling edge timestamp of the previous audio. If it is not within T4, reject the recognition and temporarily store the current audio in the cloud. If it is within T4, input the temporarily spliced final recognition result and the previous audio to the model to determine whether to respond. If it responds, perform splicing and output the splicing result. If it does not respond, go to Step 10.

[0055] Step 10: If the temporarily spliced audio does not respond, detect whether the current audio of the final recognition result responds. If it responds, only output the current audio. If it does not respond, reject the recognition and temporarily store the current audio in the cloud.

[0056] Please refer to Figure 6 , which shows a time node schematic diagram of a voice processing method provided by an embodiment of the present invention.

[0057] Such as Figure 6As shown, at time T1, the timestamp of the audio falling edge reaches the T1 time node. When reaching the T1 time node, it will trigger the detection to obtain the intermediate recognition result and perform text detection; At time T2, the audio falling edge reaches the T2 time node. The T2 time node is only triggered when the intermediate recognition result is detected as an incomplete sentence. Before reaching this time node, if there is subsequent audio, the intermediate recognition result will be concatenated with the subsequent audio. If no subsequent audio is found when reaching the T2 time node, it will be rejected and the intermediate recognition result will be temporarily stored in the cloud. Here, T2 time > T1 time; At time T3, the audio falling edge reaches the T3 time node. The T3 time node only appears when the intermediate recognition result is detected as a complete audio. The T3 time node is used to detect whether a callback result is received. If a callback result is detected within the T3 time node, it will send an end frame and trigger the next segment of audio. If no callback result is detected within the T3 time node, it will send an audio end frame at the T3 time to obtain the final recognition result. Here, T3 time = T2 time or greater than T2 time or less than T2 time, and this application has no limitation on this. T3 time > T1 time; T4 time is a time range between the audio rising edge timestamp of the final recognition result and the audio falling edge timestamp of the previous segment of audio. If the audio rising edge timestamp of the final recognition result and the audio falling edge timestamp of the previous segment of audio are within the T4 time, it will be temporarily concatenated and input into the model to determine whether the final recognition result and the previous segment of audio respond. If they respond, they will be concatenated and the concatenated result will be output. If the temporarily concatenated audio does not respond, it will detect whether the final recognition result responds. If it responds, only the current audio will be output. If it does not respond, it will be rejected and the current audio will be temporarily stored in the cloud.

[0058] Please refer to Figure 7 , which shows an audio schematic diagram of an alternative solution of a voice processing method provided by an embodiment of the present invention; As Figure 7 shown, generally in the alternative solution, the current vad pausetime is increased to achieve concatenation when the interval between the second sentence input and the previous input is lower than pausetime. For example Figure 7 for the two segments of audio in, when the rising edge time point Tb of the second audio is less than the vad pausetime from the falling edge time point Ta of the first audio (Tb – Ta < vad pausetime), at this time, an end frame for recognition will not be sent to the recognition, and naturally the two audios will be concatenated.

[0059] The inventor found that alternative solutions usually have the following defects: 1. In fields where quick responses are expected, such as in vehicle control (commands like opening the window), due to the overly long vad pausetime setting, the overall end-to-end response will take a relatively long time.

[0060] 2. For inputs that have been split into two sentences, there is no splicing logic. For example, in Example 1: Audio 1: Navigate to Window of the World (pause exceeding vad pausetime) Audio 2: parking lot. Example 2: Audio 1: Tomatoes (pause exceeding vad pausetime) Audio 2: How to make scrambled eggs with tomatoes? Since vad pausetime is fixed, the silent segment time following the input of valid audio must meet a certain duration before the recognition end frame is triggered and sent, resulting in the final recognition result. Once the silent segment following the input of valid audio exceeds vad pausetime, it is considered a complete recognition, and the calculation of the final recognition result is triggered.

[0061] To address these defects, a common approach is to introduce a specific module within the vad module to perform the logic of determining whether to stop. However, simply introducing this module still cannot achieve splicing in some scenarios, and generally, it is impossible to optimize the dynamic splicing characteristics from a full-link perspective.

[0062] The embodiments of this application jointly optimize the dynamic splicing characteristics from the perspective of a full-link solution: 1. The client uploads the vad timestamp information of the input audio; 2. The cloud cloudvad module calls the ted model to determine whether the current input is finished; 3. The dialogue control couples with cloudvad to perform the skill routing result callback; 4. The dialogue control combines the local vad information of the input audio to formulate the final recognition result splicing strategy to supplement the deficiencies of the cloudvad ted solution.

[0063] Please refer to Figure 8 and Figure 9 where Figure 8 shows a framework diagram of a voice processing method provided by an embodiment of the present invention, Figure 9 shows an audio schematic diagram of a voice processing method provided by an embodiment of the present invention.

[0064] As Figure 8 and Figure 9 shown; 1. Client: The client uploads the vad timestamp information after detecting the valid audio, including the rising edge timestamp and falling edge timestamp of the vad-detected audio; 2. Server: ① The access service (ddsserver-fullduplex) receives the voice start message and voice audio stream transmitted by the client; ②Cloudvad receives the voice start message and voice audio stream transmitted by the access service; ③For valid audio, when the silence audio reaches time t1 (less than vad pausetime), cloudvad triggers an identification and decoding once, calculates the intermediate recognition result eof = 0 rec (different from the previous intermediate recognition result eof = 0, var). At this time, it calls ted (text end detection module) to determine whether the current text has finished. If not, it waits for subsequent audio. If other valid audio comes within 2s, it will be concatenated. If ted determines that it has ended, at this time, a parameter will be carried in the eof = 0 rec recognition result, and the recognition result will be returned to the access service ddsserver-fullduplex and then transferred to the dialogue control center dm-dispatch-server-fullduplex; ④After receiving eof = 0 rec by the dialogue control center dm-dispatch-server-fullduplex and carrying the parameter in step 3, this recognition will trigger product-level semantic settlement. After the final semantic settlement is completed and the skill routing is completed, if it is found to be in the field related to vehicle control, the result will be called back to cloudvad at this time; for the result after the final skill routing, if the semantics are not vehicle control-related instructions, the callback is not triggered.

[0065] ⑤After receiving the callback result from the dialogue control center dm-dispatch-server-fullduplex, for the vehicle control-related field instructions above, Cloudvad immediately sends an audio end frame to trigger the subsequent dialogue.

[0066] ⑥If cloudvad does not receive the callback result from the above dm-dispatch-server-fullduplex after calling ted to determine the end, until time t2 (vad pausetime), it will send an audio end frame at this time to obtain the final recognition result eof = 1 rec.

[0067] ⑦After receiving the final recognition result eof = 1 rec by the dialogue control center, it will combine the vad timestamp information uploaded by the client to determine the vad rising edge timestamp of the current input audio and the previous audio falling edge timestamp. If it meets the time range (such as 2s), the previous text and the current text will be concatenated and sent to a model. The main function of this model is to determine whether the concatenated recognition result should be responded to and will output a response label. If the concatenated result should be responded to, it will be concatenated. If the concatenated result does not respond and the current single text can be responded to, the current text will be directly used without concatenation. This model is not limited to rule-based or model-based methods.

[0068] Implementation examples of the above solution include: 1. "I want to listen to... (pause) the song of XX"; For this type of splicing, use the cloudvad ted model to determine non-stop (the result of the ted model for "I want to listen to" is not stopped), and then the voice after pausing for a period of time (2s) can be spliced.

[0069] 2. "Open the window"; Quickly end and quickly respond.

[0070] For this type, use the cloudvad ted to determine stop (the result of the ted model for "Open the window" is stopped), and then combine it with the central control skill scheduling and skill routing result to do a quick stop. After such tests, the callback result can be received within 30ms after t1, and the end frame is sent.

[0071] 3. "How to make... (pause) scrambled eggs with tomatoes"; For this type of splicing, simply saying something like "tomatoes" generally triggers rejection (the cloud considers it an invalid input, and at this time, a rejection reply is given to the terminal side). Then, when "how to make scrambled eggs" is said within 2s, use the implementation in step 3.1 item ⑦ to achieve splicing.

[0072] In some embodiments, the embodiments of the present invention provide a non-volatile computer-readable storage medium, in which one or more programs including execution instructions are stored, and the execution instructions can be read and executed by an electronic device (including but not limited to a computer, a server, or a network device, etc.) to be used to execute any one of the above voice processing methods of the present invention.

[0073] In some embodiments, the embodiments of the present invention further provide a computer program product, the computer program product includes a computer program stored on a non-volatile computer-readable storage medium, the computer program includes program instructions, and when the program instructions are executed by a computer, the computer is made to execute any one of the above voice processing methods.

[0074] In some embodiments, the embodiments of the present invention further provide an electronic device, which includes: at least one processor, and a memory communicatively connected to the at least one processor, wherein the memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor to enable the at least one processor to execute a voice processing method.

[0075] Figure 10 It is a schematic hardware structure diagram of an electronic device for executing a voice processing method provided by another embodiment of the present application. As Figure 10 shown, the device includes: One or more processors 1010 and a memory 1020, Figure 10 Taking one processor 1010 as an example.

[0076] The device for executing a voice processing method may further include: an input device 1030 and an output device 1040.

[0077] The processor 1010, the memory 1020, the input device 1030, and the output device 1040 may be connected via a bus or other means, Figure 10 Taking connection via a bus as an example.

[0078] The memory 1020, as a non-volatile computer-readable storage medium, can be used to store non-volatile software programs, non-volatile computer-executable programs, and modules, such as the program instructions / modules corresponding to a voice processing method in an embodiment of the present application. The processor 1010 executes various functional applications and data processing of the server by running the non-volatile software programs, instructions, and modules stored in the memory 1020, that is, implements a voice processing method in the above method embodiment.

[0079] The memory 1020 may include a program storage area and a data storage area. Among them, the program storage area can store an operating system and application programs required for at least one function; the data storage area can store data created according to the use of a voice processing device, etc. In addition, the memory 1020 may include a high-speed random access memory, and may also include non-volatile memory, such as at least one magnetic disk storage device, a flash memory device, or other non-volatile solid-state storage devices. In some embodiments, the memory 1020 may optionally include a memory remotely set relative to the processor 1010, and these remote memories can be connected to a voice processing device through a network. Examples of the above network include but are not limited to the Internet, an enterprise intranet, a local area network, a mobile communication network, and combinations thereof.

[0080] The input device 1030 can receive input digital or character information, and generate signals related to the user settings and function control of a voice processing device. The output device 1040 may include a display device such as a display screen.

[0081] The one or more modules are stored in the memory 1020, and when executed by the one or more processors 1010, execute a voice processing method in any of the above method embodiments.

[0082] The above product can execute the method provided in the embodiment of the present application, and has corresponding functional modules and beneficial effects for executing the method. For technical details not described in detail in this embodiment, reference can be made to the method provided in the embodiment of the present application.

[0083] The electronic devices in the embodiments of the present application exist in various forms, including but not limited to: (1) Mobile communication devices: These devices are characterized by having mobile communication functions and mainly aim to provide voice and data communication. Such terminals include: smart phones, multimedia phones, functional phones, and low-end phones, etc.

[0084] (2) Ultra-mobile personal computer devices: These devices belong to the category of personal computers, have computing and processing functions, and generally also have the characteristic of mobile Internet access. Such terminals include: PDA, MID, and UMPC devices, etc.

[0085] (3) Portable entertainment devices: These devices can display and play multimedia content. Such devices include: audio and video players, handheld game consoles, e-books, and intelligent toys and portable vehicle navigation devices.

[0086] (4) Other airborne electronic devices with data interaction functions, such as in-vehicle device installed on a vehicle.

[0087] The device embodiments described above are merely illustrative. The units described as separate components may or may not be physically separated, and the components shown as units may or may not be physical units, that is, they may be located in one place or distributed to multiple network units. Some or all of the modules can be selected according to actual needs to achieve the purpose of the solution of this embodiment.

[0088] Through the description of the above embodiments, those skilled in the art can clearly understand that each embodiment can be implemented by means of software plus a general hardware platform, and of course, it can also be implemented by hardware. Based on this understanding, the essence of the above technical solutions, or the part that contributes to the related technologies, can be embodied in the form of a software product. This computer software product can be stored in a computer-readable storage medium, such as ROM / RAM, magnetic disk, optical disk, etc., and includes several instructions to enable a computer device (which can be a personal computer, a server, or a network device, etc.) to execute the methods described in each embodiment or some parts of the embodiments.

[0089] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present application, and are not intended to limit them; although the present application has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that they can still modify the technical solutions described in the foregoing embodiments, or perform equivalent replacements for some of the technical features; and these modifications or replacements do not make the essence of the corresponding technical solutions deviate from the spirit and scope of the technical solutions of each embodiment of the present application.

Claims

1. A voice processing method for a cloud server, comprising: Receiving the vad timestamp information, voice start message, and voice audio stream after valid audio detection uploaded by a client, wherein the vad timestamp information includes an audio rising edge timestamp and an audio falling edge timestamp; At a first preset pause duration after the audio falling edge timestamp of the current valid audio, triggering intermediate recognition and decoding based on cloud vad to obtain an intermediate recognition result, and invoking a text detection model to determine whether the intermediate recognition result is a complete semantic sentence; If the intermediate recognition result is not a complete semantic sentence and subsequent voice audio streams appear within a second preset pause duration after the audio falling edge timestamp of the current valid audio, splicing the intermediate recognition result and the subsequent text of the subsequent voice audio streams, wherein the second preset pause duration is greater than the first preset pause duration; If the intermediate recognition result is a complete semantic sentence, carrying parameters of the intermediate recognition result and transferring them to a dialogue control center, triggering product-level semantic settlement based on the dialogue control center and obtaining a skill routing domain through skill routing, and determining whether to perform a callback according to the skill routing domain.

2. The method according to claim 1, wherein, After the step of if the intermediate recognition result is a complete semantic sentence, carrying parameters of the intermediate recognition result and transferring them to a dialogue control center, triggering product-level semantic settlement based on the dialogue control center and obtaining a skill routing domain through skill routing, and determining whether to perform a callback, the method further includes: Determining whether the skill routing domain is a vehicle control-related domain. If it is a vehicle control-related domain, sending the callback result to the cloud vad; if it is not a vehicle control-related domain, not sending the callback result.

3. The method according to claim 1, wherein, After the step of if the intermediate recognition result is a complete semantic sentence, the method further includes: If the cloud vad receives the callback result from the dialogue control center within a third preset pause duration after the audio falling edge timestamp of the valid audio, sending an audio end frame in real time and triggering a subsequent dialogue, wherein the third pause duration is greater than the first pause duration; If the cloud vad does not receive the callback result from the control center within a third preset pause duration after the audio falling edge timestamp of the valid audio, sending an audio end frame at the third preset pause duration after the audio falling edge timestamp of the valid audio to obtain a final recognition result.

4. The method according to claim 3, wherein After obtaining the final recognition result, the method further includes: After receiving the final recognition result, the dialogue control center combines the vad timestamp information to determine whether the interval between the audio rising edge timestamp of the final recognition result and the previous audio falling edge timestamp meets a fourth preset pause duration, wherein the fourth pause duration is greater than the second preset pause duration; If it meets the requirement, splicing the previous recognition text and the final recognition result and sending them to a splicing discrimination model. If the spliced recognition result can respond, outputting a response label and outputting after splicing; if the spliced recognition result cannot respond and the final recognition result can respond, directly outputting the final recognition result without splicing; If it does not meet the requirement, no processing is performed and it is temporarily stored in the cloud.

5. The method according to claim 1, wherein, If the intermediate recognition result is not a complete semantic sentence, and a subsequent voice audio stream appears within a second preset pause duration after the audio falling edge timestamp of the current valid audio, splicing the subsequent text of the intermediate recognition result and the subsequent voice audio stream further includes: If no subsequent audio appears within the second preset pause duration, end the process and temporarily store it in the cloud.

6. A voice processing method for use in a dialogue control center, including: Receiving a complete semantic sentence transmitted by cloud VAD; Triggering product-level semantic settlement through the dialogue control center and completing skill routing to obtain a skill routing domain, and judging whether it is a vehicle control-related instruction according to the skill routing domain; If it is a vehicle control-related instruction, sending the callback result to the cloud VAD, and if it is not a vehicle control-related domain, not sending the callback result.

7. A voice processing method for use in cloud VAD, including: Receiving the vad timestamp information, voice start message, and voice audio stream after valid audio detection uploaded by the client, where the vad timestamp information includes an audio rising edge timestamp and an audio falling edge timestamp; At the first preset pause duration after the audio falling edge timestamp of the current valid audio, triggering intermediate recognition decoding based on cloud VAD to obtain an intermediate recognition result, and calling a text detection model to judge whether the intermediate recognition result is a complete semantic sentence; If the intermediate recognition result is not a complete semantic sentence, and a subsequent voice audio stream appears within a second preset pause duration after the audio falling edge timestamp of the current valid audio, splicing the subsequent text of the intermediate recognition result and the subsequent voice audio stream, where the second preset pause duration is greater than the first preset pause duration; If the intermediate recognition result is a complete semantic sentence, transferring the intermediate recognition result with parameters to the dialogue control center.

8. The method according to claim 7, wherein, After if the intermediate recognition result is a complete semantic sentence and transferring the intermediate recognition result with parameters to the dialogue control center, it further includes: Obtaining the callback result in the dialogue and performing corresponding operations according to the callback result; If the cloud VAD receives the callback result from the dialogue control center within a third preset pause duration after the audio falling edge timestamp of the valid audio, sending an audio end frame in real time and triggering a subsequent dialogue; If the cloud VAD does not receive the callback result from the control center within a third preset pause duration after the audio falling edge timestamp of the valid audio, sending an audio end frame at the third preset pause duration after the audio falling edge timestamp of the valid audio to obtain a final recognition result.

9. A voice processing method for use in a client, including: After obtaining audio information, performing voice detection to judge whether it is valid audio. If it is valid audio, obtaining the vad timestamp information, voice start message, and voice audio stream of the valid audio, and uploading them to the cloud server for processing; Obtaining the processing result of the cloud server and performing corresponding operations according to the processing result.

10. An electronic device, comprising: At least one processor, and a memory communicatively connected to the at least one processor, wherein the memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor to enable the at least one processor to perform the steps of the method according to any one of claims 1 to 5.

Citation Information

Cited By

  • Real-time structured extraction method and device for streaming voice

    CN121789687A