Voice Processing Method, Apparatus, Electronic Device and Medium for Audio Tail End Detection
By employing multiple VAD instances with varying silence thresholds for parallel language recognition and NLP in automotive voice interactions, the method addresses long response times, enhancing user experience by reducing latency in voice processing.
Patent Information
- Application Number
- CN202210822218.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-07-13
- Publication Date
- 2025-07-15
- Estimated Expiration
- 2042-07-13
AI Technical Summary
The voice processing response time is long in existing automotive scenarios, mainly due to the delay in speech activity detection (VAD) tail point judgment.
By performing multiple tail sound detections on the voice data received in real time, multiple language processing results are obtained, and the recognition results are compared during the detection process, and natural language processing is performed only when the recognition results are inconsistent, and the target speech recognition results are obtained in advance.
It effectively shortens the response time of voice processing, improves response efficiency, and improves user experience.
Smart Images

Figure CN115083396B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of speech processing, and particularly relates to a speech processing method, apparatus, electronic device, and medium for audio end detection. Background Art
[0002] In the prior art, in a dialogue scenario such as an automotive scenario, the voice interaction technology usually uses Voice Activity Detection (VAD for short) to determine the end point of speech. VAD detection generally has a fixed silence duration, usually in the range of 500 - 800 ms. After the user's dialogue is completed and passes through VAD detection, speech recognition, Natural Language Understanding (NLU for short), and Dialog Management (DM for short), etc. are performed on the user's dialogue. The entire process is processed serially, and finally a target dialogue corresponding to the user's dialogue is generated and output to achieve human - machine interaction.
[0003] However, when using VAD to determine the end point of speech in the existing automotive scenario, when the duration of the end - segment silence exceeds the set duration, a VAD END event is triggered. After operations such as speech model scoring and normalization processing are performed after the VAD END event is triggered, the final dialogue is returned, and then NLU and DM, etc. are processed in sequence. Since the entire processing process is serially processed, there is a problem of relatively long speech processing response time. Summary of the Invention
[0004] Embodiments of the present invention provide a speech processing method, apparatus, electronic device, and medium for audio end detection, which can effectively shorten the speech processing response time, thereby improving the response efficiency and making the user experience better.
[0005] In a first aspect of an embodiment of the present invention, a speech processing method for audio end detection is provided. The method includes:
[0006] After receiving a speech request, receive speech data in real time;
[0007] Perform N times of end - sound detection on the speech data received in real time, where N is an integer not less than 2;
[0008] Obtain M language processing results corresponding to the N times of end - sound detection, where M is a positive integer not greater than N;
[0009] Determine a target language processing result corresponding to the speech request from the M language processing results.
[0010] Optionally, obtaining the M language processing results corresponding to the N times of coda detection includes:
[0011] For the first coda detection among the N times of coda detection, if it is detected that the duration of the first end silence exceeds the first set duration, obtain the first language recognition result corresponding to the current voice data, perform natural language processing on the first language recognition result to obtain the first language processing result, and store the first language processing result in the stored recognition results in the cache;
[0012] For the i-th coda detection among the N times of coda detection, if it is detected that the duration of the i-th end silence exceeds the i-th set duration, obtain the i-th language recognition result corresponding to the current voice data. If the (i - 1)-th language recognition result is inconsistent with the i-th language recognition result, replace the stored recognition result from the (i - 1)-th language processing result with the i-th language processing result corresponding to the i-th language recognition result; if they are consistent, control the stored recognition result to remain unchanged, where i is a positive integer, and i is taken from 2 to N in sequence.
[0013] Optionally, if N = 2, obtaining the N language processing results corresponding to the N times of coda detection includes:
[0014] For the first coda detection among the N times of coda detection, if it is detected that the duration of the first end silence exceeds the first set duration, obtain the first language recognition result corresponding to the current voice data, perform natural language processing on the first language recognition result to obtain the first language processing result, store the first language processing result in the stored recognition results, and store the first language recognition result in the preprocessed language recognition results in the cache;
[0015] For the second coda detection among the N times of coda detection, if it is detected that the duration of the second end silence exceeds the second set duration, obtain the second language recognition result corresponding to the current voice data, compare the second language recognition result with the first language recognition result. If it is compared that the first language recognition result is inconsistent with the second language recognition result, replace the preprocessed language recognition result with the second language recognition result, perform natural language processing on the second language recognition result to obtain the second language processing result, and replace the stored recognition result with the second language processing result; if it is compared that the first language recognition result is consistent with the second language recognition result, control the preprocessed language recognition result to be the first language recognition result, and control the stored recognition result to be the first language processing result, where the second set duration is greater than the first set duration.
[0016] Optionally, determining the target language processing result corresponding to the voice request from the M language processing results includes:
[0017] If the first language recognition result is the same as the second language recognition result, then use the first language processing result stored in the stored recognition result as the target language processing result; if the first language recognition result is different from the second language recognition result, then use the second language processing result stored in the stored recognition result as the target language processing result.
[0018] Optionally, determining the target language processing result corresponding to the voice request from the M language processing results includes:
[0019] Compare the first language processing result with the second language processing result to obtain a first language comparison result;
[0020] Determine the target language processing result from the M language processing results according to the first language comparison result.
[0021] Optionally, if N = 3, obtaining the M language processing results corresponding to the N times of end sound detection includes:
[0022] For the first end sound detection among the N times of end sound detection, if it is detected that the first end silence duration exceeds the first set duration, obtain the first language recognition result corresponding to the current voice data, perform natural language processing on the first language recognition result to obtain a first language processing result, store the first language processing result in the stored recognition result, and store the first language recognition result in the preprocessed language recognition result in the cache;
[0023] For the second end sound detection among the N times of end sound detection, if it is detected that the second end silence duration exceeds the second set duration, obtain the second language recognition result corresponding to the current voice data, compare the second language recognition result with the first language recognition result. If it is compared that the first language recognition result is different from the second language recognition result, then replace the preprocessed language recognition result with the second language recognition result, perform natural language processing on the second language recognition result to obtain a second language processing result, and replace the stored recognition result with the second language processing result; if it is compared that the first language recognition result is the same as the second language recognition result, then control the preprocessed language recognition result to be the first language recognition result, and control the stored recognition result to be the first language processing result, where the second set duration is greater than the first set duration;
[0024] For the third vowel detection in the N - th vowel detection, if it is detected that the duration of the third trailing silence exceeds the third set duration, obtain the third language recognition result corresponding to the current voice data, compare the third language recognition result with the language recognition results stored in the pre - processed language recognition result. If it is found that the language recognition results are inconsistent, replace the pre - processed language recognition result with the third language recognition result, perform natural language processing on the third language recognition result to obtain the third language processing result, and replace the stored recognition result with the third language processing result; if it is found that the language recognition results are consistent, control the pre - processed language recognition result and the stored recognition result to remain unchanged, where the third set duration is greater than the second set duration.
[0025] Optionally, determining the target language processing result corresponding to the voice request from the M language processing results includes:
[0026] After comparing the third language recognition result with the language recognition results stored in the pre - processed language recognition result, if it is found that the language recognition results are consistent, use the currently stored recognition result in the cache as the target language processing result; if the third language recognition result is inconsistent with the second language recognition result, use the third language processing result stored in the stored recognition result as the target language processing result.
[0027] Optionally, after determining the target language processing result corresponding to the voice request from the M language processing results, the method includes:
[0028] Use the target language processing result to perform natural speech generation to generate a target dialogue corresponding to the voice request, and perform voice output on the target dialogue.
[0029] A second aspect of the embodiments of the present invention also provides a voice processing device for audio trailing - end detection. The device includes:
[0030] A voice receiving unit, configured to receive voice data in real - time after receiving a voice request;
[0031] A vowel detection unit, configured to perform N - th vowel detection on the voice data received in real - time, where N is an integer not less than 2;
[0032] A voice recognition unit, configured to obtain M language processing results corresponding to the N - th vowel detection;
[0033] A recognition result obtaining unit, configured to determine the target language processing result corresponding to the voice request from the M language processing results.
[0034] Optionally, the speech recognition unit is configured to, for the first tail - end detection among the N times of tail - end detections, if it is detected that the duration of the first tail - end silence exceeds the first set duration, obtain the first language recognition result corresponding to the current speech data, perform natural language processing on the first language recognition result to obtain a first language processing result, and store the first language processing result in the cached stored recognition result; for the i - th tail - end detection among the N times of tail - end detections, if it is detected that the duration of the i - th tail - end silence exceeds the i - th set duration, obtain the i - th language recognition result corresponding to the current speech data, and if the (i - 1) - th language recognition result is inconsistent with the i - th language recognition result, then replace the stored recognition result from the (i - 1) - th language processing result with the i - th language processing result corresponding to the i - th language recognition result; if they are consistent, then control the stored recognition result to remain unchanged, where i is a positive integer, and i is taken from 2 to N in sequence.
[0035] In a third aspect of the embodiments of the present invention, an electronic device is provided, including a memory, and one or more programs, where one or more programs are stored in the memory and are configured to be executed by one or more processors to perform operation instructions corresponding to the speech processing method for audio tail - end detection provided in the first aspect.
[0036] In a fourth aspect of the embodiments of the present invention, a computer program product is provided, characterized in that the computer program product includes computer instructions, and the computer instructions are stored in a computer - readable storage medium and are adapted to be read and executed by a processor so that a computer device having the processor executes the steps corresponding to the speech processing method for audio tail - end detection provided in the first aspect.
[0037] One or at least one of the above - mentioned technical solutions in the embodiments of the present application has at least the following technical effects:
[0038] Based on the above technical solution, after receiving a voice request, voice data is received in real time; the received voice data is subjected to N times of end - sound detection; M language processing results corresponding to the N times of end - sound detection are obtained; a target language processing result corresponding to the voice request is determined from the M language processing results; at this time, since N is an integer not less than 2, when performing the second to the Nth end - sound detection, first, if the (i - 1)th language recognition result is inconsistent with the ith language recognition result, the stored recognition result is replaced from the (i - 1)th language processing result with the ith language processing result corresponding to the ith language recognition result; if they are consistent, the stored recognition result is controlled to remain unchanged, where i is a positive integer, and i is taken from 2 to N in sequence. In this way, M language processing results can be obtained before the VAD END event, and then the target speech recognition result is determined from the M language processing results. Compared with the prior art, it is not necessary to perform speech recognition, NLU, DM, etc. after the VAD END event to obtain the target speech recognition result, so that the voice processing response time can be effectively shortened, and the response efficiency is improved, making the user experience better. BRIEF DESCRIPTION OF THE DRAWINGS
[0039] Figure 1 It is a schematic flowchart of a voice processing method for audio end - detection provided by an embodiment of the present application;
[0040] Figure 2 It is a schematic structural diagram of voice data provided by an embodiment of the present application;
[0041] Figure 3 It is a diagram showing the implementation steps of three - time end - sound detection provided by an embodiment of the present application;
[0042] Figure 4 It is a schematic flowchart of a voice processing method for user - vehicle interaction provided by an embodiment of the present application;
[0043] Figure 5 It is a block diagram of a voice processing device for audio end - detection provided by an embodiment of the present application;
[0044] Figure 6 It is a schematic structural diagram of an electronic device provided by an embodiment of the present application. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0045] The main implementation principle, specific implementation methods and corresponding beneficial effects that can be achieved of the technical solution of the embodiment of the present application are elaborated in detail below with reference to the drawings.
[0046] Embodiment
[0047] Please refer to Figure 1, an embodiment of the present application provides a voice processing method for audio end detection, and the method includes:
[0048] S101. After receiving a voice request, receive voice data in real time;
[0049] S102. Perform N times of tail - sound detection on the voice data received in real time, where N is an integer not less than 2;
[0050] S103. Obtain M language processing results corresponding to the N times of tail - sound detection, where M is a positive integer not greater than N;
[0051] S104. Determine a target language processing result corresponding to the voice request from the M language processing results.
[0052] In the embodiments of this specification, a voice processing method for audio end detection can be applied to a user - end or a cloud server. Among them, the user - end can be an automotive - end or a voice - robot - end, etc. The automotive - end can be, for example, an electric vehicle, a hybrid vehicle, and a fuel vehicle, etc. The cloud server can be, for example, a tablet computer, a notebook computer, and a desktop computer, etc. Hereinafter, the cloud server is taken as an example.
[0053] Among them, in step S101, taking the cloud server as an example, the automotive - end will automatically generate a voice request according to user operations and send the voice request to the cloud server, so that the cloud server receives the voice request. After the cloud server receives the voice request, it usually feeds back an acknowledgment signal to the automotive - end. After the automotive - end receives the acknowledgment signal, it then transmits the voice data output by the user to the cloud server in real time, so that the cloud server can receive the voice data in real time. Among them, the voice data is usually audio data.
[0054] For example, when the user is driving vehicle A and presses a certain button or utters a certain specific word such as "please turn on the radio" and other specific operations, the in - vehicle terminal of vehicle A will automatically generate a voice request and simultaneously receive the voice data output by the user. After generating the voice request, the in - vehicle terminal sends the voice request to cloud server B, and after receiving the acknowledgment signal returned by cloud server B, it sends the voice data received in real time to cloud server B, so that cloud server B can receive the voice data transmitted in real time by the in - vehicle terminal after receiving the voice request.
[0055] During the process of receiving the voice data in real time, step S102 is executed.
[0056] In step S102, perform N times of end - sound detection on the real - time received voice data, where N is an integer not less than 2. For example, N can be values such as 2, 3, 5, and 8. N can be set according to actual requirements, or can be set manually or by the device itself. For example, N can be determined according to the computing performance of the cloud server.
[0057] In the specific implementation process, when performing N times of end - sound detection on the real - time received voice data, an Automatic Speech Recognition (ASR) system can be used to perform N times of Voice Activity Detection (VAD) on the real - time received voice data.
[0058] Specifically, when performing N times of end - sound detection on the real - time received voice data, a set duration needs to be set for each end - sound detection. For the i - th end - sound detection among the N times of end - sound detection, obtain the i - th set duration corresponding to the i - th end - sound detection, obtain the i - th language recognition result corresponding to the current voice data, and perform end - sound detection on the i - th language recognition result, that is, detect whether the i - th tail - end silence duration of the i - th language recognition result exceeds the i - th set duration, where i is a positive integer and i takes values from 1 to N in sequence, so as to realize N times of end - sound detection on the real - time received voice data.
[0059] In the embodiments of this specification, when obtaining the i - th language recognition result corresponding to the current voice data, use the ASR system to perform speech recognition on the current voice data, and the recognized result is used as the i - th language recognition result.
[0060] In the actual application process, N times of audio tail-end detection can be introduced based on the ASR system, that is, an additional N set durations t1, ..., tn (t1 <... < tn < t) are added. Among them, t1 is the first set duration, t2 is the second set duration, and so on until tn is the nth set duration. Taking N = 2 as an example, two additional set durations t1 and t2 are added. When the continuous duration of the tail-end silence of the current voice data exceeds t1, a VAD detection will be triggered, which is the VAD1 event. When the continuous duration of the tail-end silence of the current voice data exceeds t2, a VAD detection will be triggered, which is the VAD2 event. After the ASR system receives the VAD1 event, it will first perform speech recognition to obtain the first language recognition result, denoted as final asr1. Then, natural language processing (Natural Language Processing, abbreviated as NLP) will be performed on the first speech recognition result to obtain the preprocessing result corresponding to VAD1, that is, the first language output result, denoted as final nlp1. Correspondingly, after receiving the VAD2 event, speech recognition can be performed first to obtain the second language recognition result, denoted as final asr2. Then, NLP will be directly performed on the second speech recognition result to obtain the preprocessing result corresponding to VAD2, that is, the second language processing result, denoted as final nlp2. Among them, NLP processing includes NLU processing and DM processing. Of course, after receiving the VAD2 event, the second language recognition result can also be compared with the first language recognition result. If they are inconsistent, NLP processing will be performed to obtain final nlp2 corresponding to VAD2. If they are consistent, NLP processing will not be performed.
[0061] For example, referring to Figure 2 , the cloud server B receives the voice data A1 sent by the vehicle terminal A, which includes the front-end voice 20 "I want to listen to Jay Chou's Qi Li Xiang", the silence s0, the back-end voice 21 "Qi Li Xiang", and the silence s1. If N = 2, the voice data will be subjected to 2 times of audio tail-end detection using t1 and t2. If t1 < s0 < t2, the first time of audio tail-end detection will be performed on the voice data using t1. At this time, the result after speech recognition of the front-end voice 20 and the silence s0 can be used as the first language recognition result, denoted as final asr1. If s1 > t2, the second time of audio tail-end detection will be performed on the voice data using t2. At this time, the result after speech recognition of the front-end voice 20, the silence s0, the back-end voice 21, and the silence s1 can be used as the second language recognition result, denoted as final asr2.
[0062] Thus, since t2 is greater than t1, when performing N times of audio tail-end detection on the real-time received voice data, multiple voice data under different tail-end mute durations can be obtained through the set N set durations. Usually, the maximum set duration among the set N set durations is greater than the pause duration of the user's output voice data (e.g., t2 > mute so), which makes the matching degree between a certain current voice data and the voice data actually output by the user higher, and the accuracy of the obtained current voice data higher.
[0063] After performing N times of tail sound detection on the real-time received voice data, step S103 is executed.
[0064] In step S103, for the first tail sound detection among the N times of tail sound detection, if it is detected that the first tail-end mute duration exceeds the first set duration, obtain the first language recognition result corresponding to the current voice data, perform natural language processing on the first language recognition result to obtain the first language processing result, and store the first language processing result in the cached stored recognition result; for the i-th tail sound detection among the N times of tail sound detection, if it is detected that the i-th tail-end mute duration exceeds the i-th set duration, obtain the i-th language recognition result corresponding to the current voice data. If the (i - 1) language recognition result is inconsistent with the i-th language recognition result, then replace the stored recognition result from the (i - 1) language processing result with the i-th language processing result corresponding to the i-th language recognition result; if they are consistent, then control the stored recognition result to remain unchanged, where i is a positive integer, and i is taken from 2 to N in sequence.
[0065] Thus, when performing N times of tail sound detection, the previous language recognition result (i.e., the (i - 1) language recognition result) can be compared with the subsequent language recognition result (i.e., the i-th language recognition result). If it is compared that the language recognition results are consistent, then update the stored recognition result in the cache, and the updated stored recognition result is the subsequent language processing result (i.e., the i-th language processing result). Thus, after completing N times of tail sound detection, the data in the stored recognition result can be used as the target language processing result. In the prior art, after tail sound detection, voice recognition, NLU, DM, etc. still need to be performed to obtain the target voice recognition result. Therefore, compared with the prior art, the technical solution of the present application can save the processing steps such as voice recognition, NLU, and DM, effectively shorten the voice processing response time, and thus improve the response efficiency, making the user experience better.
[0066] Specifically, for the i-th tail sound detection among the N times of tail sound detections, the i-th language processing result obtained from the i-th tail sound detection can be acquired, where i is a positive integer. By sequentially taking i from 1 to N, the 1st to N-th language processing results can be obtained as M language processing results. At this time, M = N.
[0067] For example, taking N = 3 as an example, three new set durations t1, t2, and t3 are set. During the 1st tail sound detection, if the 1st tail-end silence duration is greater than t1 and not greater than t2, the VAD1 event is triggered. After speech recognition, the 1st language recognition result is represented as final asr1. The 1st language recognition result is subjected to NLP processing to obtain the 1st language processing result corresponding to VAD1, represented as final nlp1. During the 2nd tail sound detection, if the 2nd tail-end silence duration is greater than t2 and not greater than t3, the VAD2 event is triggered. After speech recognition, the 2nd language recognition result is represented as final asr2. The 2nd language recognition result is subjected to NLP processing to obtain the 2nd language processing result corresponding to VAD2, represented as final nlp2. During the 3rd tail sound detection, if the 3rd tail-end silence duration is greater than t3, the VAD3 event is triggered. After speech recognition, the 3rd language recognition result is represented as final asr3. The 3rd language recognition result is subjected to NLP processing to obtain the 3rd language processing result corresponding to VAD3, represented as final nlp3.
[0068] In this way, N times of VAD detections can be performed for the same voice request, and N language processing results corresponding to the N times of VAD detections can be obtained. At this time, N = M.
[0069] In another embodiment, for the 1st tail sound detection among the N times of tail sound detections, if it is detected that the 1st tail-end silence duration exceeds the 1st set duration, the 1st language recognition result corresponding to the current voice data is acquired. The 1st language recognition result is subjected to natural language processing to obtain the 1st language processing result, and the 1st language processing result is stored in the cached recognition result. For the i-th tail sound detection among the N times of tail sound detections, if it is detected that the i-th tail-end silence duration exceeds the i-th set duration, the i-th language recognition result corresponding to the current voice data is acquired. If the (i - 1)-th language recognition result is inconsistent with the i-th language recognition result, the stored recognition result is replaced from the (i - 1)-th language processing result with the i-th language processing result corresponding to the i-th language recognition result. If they are consistent, the stored recognition result is controlled to remain unchanged, where i is a positive integer, and i is sequentially taken from 2 to N.
[0070] Thus, when the (i - 1)th language recognition result is the same as the ith language recognition result, there is no need to perform natural speech processing on the ith speech data, which can effectively reduce the computational load and thus improve the efficiency of speech processing.
[0071] Specifically, if N = 2, obtain M language processing results corresponding to N times of end - sound detection, including: for the first end - sound detection among the N times of end - sound detection, if it is detected that the duration of the first end - point silence exceeds the first set duration, obtain the first language recognition result corresponding to the current speech data, perform natural language processing on the first language recognition result to obtain the first language processing result, store the first language processing result in the stored recognition results, and store the first language recognition result in the pre - processed language recognition results cached; for the second end - sound detection among the N times of end - sound detection, if it is detected that the duration of the second end - point silence exceeds the second set duration, obtain the second language recognition result corresponding to the current speech data, compare the second language recognition result with the first language recognition result. If it is compared that the first language recognition result is inconsistent with the second language recognition result, then replace the pre - processed language recognition results with the second language recognition result, perform natural language processing on the second language recognition result to obtain the second language processing result, and replace the stored recognition results with the second language processing result; if it is compared that the first language recognition result is consistent with the second language recognition result, then control the pre - processed language recognition results to be the first language recognition result, and control the stored recognition results to be the first language processing result. Among them, the second set duration is greater than the first set duration. The second set duration can be, for example, 200 or 300 milliseconds (ms), etc., and the first set duration can be, for example, 100 or 150 ms, etc.
[0072] Specifically, when performing the first end - sound detection, it can also be detected that the duration of the first end - point silence exceeds the first set duration and does not exceed the second set duration. Then obtain the first language recognition result corresponding to the current speech data and perform natural language processing on the first language recognition result to obtain the first language processing result. Thus, the number of times of repeated end - sound detection can be reduced.
[0073] For example, see Figure 2, the cloud server B receives the voice data A1 sent by the car end A, which includes the front voice 20 of "I want to listen to Jay Chou's Seven-Li Fragrance", silence s0, and the back voice 21 of "Seven-Li Fragrance", silence s1; the first tail sound detection is performed on A1, and if it is detected that t1<s0<t2, it is determined that the current voice data is the front voice 20 and the silence s0, and the current voice data is subjected to voice recognition to obtain the first language recognition result represented by final asr1, and the final asr1 is subjected to NLP processing to obtain the first language processing result represented by final nlp1, and final nlp1 is "I want to listen to Jay Chou's"; the second tail sound detection is performed on A1, and if it is detected that s1>t2, it is determined that the current voice data is the front voice 20, silence s0, back voice 21 and silence s1, and the current voice data is subjected to voice recognition to obtain the second language recognition result represented by finalasr2, and then the finalasr2 is subjected to NLP processing to obtain the second language processing result represented by final nlp2, and final nlp2 is "I want to listen to Jay Chou's "Seven Miles of Fragrance".
[0074] In another embodiment, if N = 3, obtain M language processing results corresponding to N times of coda detection, including: for the first coda detection among the N times of coda detection, if it is detected that the first end silence duration exceeds the first set duration, obtain the first language recognition result corresponding to the current voice data, perform natural language processing on the first language recognition result to obtain the first language processing result, store the first language processing result in the stored recognition result, and store the first language recognition result in the preprocessed language recognition result in the cache; for the second coda detection among the N times of coda detection, if it is detected that the second end silence duration exceeds the second set duration, obtain the second language recognition result corresponding to the current voice data, compare the second language recognition result with the first language recognition result, if it is compared that the first language recognition result is inconsistent with the second language recognition result, then replace the preprocessed language recognition result with the second language recognition result, perform natural language processing on the second language recognition result to obtain the second language processing result, and replace the stored recognition result with the second language processing result; if it is compared that the first language recognition result is consistent with the second language recognition result, then control the preprocessed language recognition result to be the first language recognition result, and control the stored recognition result to be the first language processing result, where the second set duration is greater than the first set duration; for the third coda detection among the N times of coda detection, if it is detected that the third end silence duration exceeds the third set duration, obtain the third language recognition result corresponding to the current voice data, compare the third language recognition result with the language recognition result stored in the preprocessed language recognition result, if it is compared that the language recognition results are inconsistent, then replace the preprocessed language recognition result with the third language recognition result, perform natural language processing on the third language recognition result to obtain the third language processing result, and replace the stored recognition result with the third language processing result; if it is compared that the language recognition results are consistent, then control the preprocessed language recognition result and the stored recognition result to remain unchanged, where the third set duration is greater than the second set duration.
[0075] Specifically, when performing the first tail sound detection, it is also possible to detect that the first tail-end silence duration exceeds the first set duration and does not exceed the second set duration. Then, obtain the first language recognition result corresponding to the current voice data, perform natural language processing on the first language recognition result to obtain the first language processing result, store the first language recognition result in the preprocessed language recognition result, and store the first language processing result in the stored recognition result. And when performing the second tail sound detection, it is also possible to detect that the second tail-end silence duration exceeds the second set duration and does not exceed the third set duration. Then, obtain the second language recognition result corresponding to the current voice data, and then determine whether the second language recognition result is consistent. If it is consistent, no NLP processing is performed; if it is inconsistent, NLP is performed to obtain the second language processing result, replace the preprocessed language recognition result with the second language recognition result, and replace the stored recognition result with the second language processing result. And the specific implementation process when performing the third tail sound detection can refer to the implementation process of the second tail sound detection.
[0076] In the embodiments of this specification, when performing natural language processing on the i-th language recognition result, NLP processing can be performed on the i-th language recognition result. Among them, NLP processing includes NLU processing and DM processing.
[0077] In the embodiments of this specification, when obtaining the i-th language recognition result corresponding to the current voice data, the current voice data corresponding to the i-th language recognition result can be subjected to Automatic Speech Recognition (ASR) processing, and the obtained i-th language recognition result can be represented by final asri.
[0078] For example, taking N = 3 as an example, three new set durations t1, t2, and t3 are set. When performing the first tail sound detection, if the first tail-end silence duration is greater than t1 and not greater than t2, the VAD1 event is triggered, and the first language recognition result is obtained and represented by finalasr1. NLP processing is performed on the first language recognition result to obtain the first language processing result corresponding to VAD1 and represented by finalnlp1. Store final asr1 in the preprocessed language recognition result final asr, and store final nlp1 in the stored recognition result final nlp.
[0079] Among them, when performing the second tail sound detection, if the second tail-end silence duration is greater than t2 and not greater than t3, the VAD2 event is triggered, and the second language recognition result is obtained, denoted as final asr2. Compare final asr2 with final asr. If they are the same, control final asr and final nlp to remain unchanged, that is, the data in final asr is final asr1, and the data in final nlp is final nlp1. If they are different, perform NLP processing on the second language recognition result to obtain the second language processing result corresponding to VAD2, denoted as final nlp2. Replace final asr with final asr2 and replace final nlp with final nlp2.
[0080] In addition, when performing the third tail sound detection, if the third tail-end silence duration is greater than t3, the VAD3 event is triggered, and the third language recognition result is obtained, denoted as final asr3. Compare final asr3 with final asr. If they are the same, control final asr and final nlp to remain unchanged. If they are different, perform NLP processing on the third language recognition result to obtain the third language processing result corresponding to VAD3, denoted as final nlp3. Replace final asr with final asr3 and replace final nlp with final nlp3.
[0081] In this way, multiple VAD detections can be performed for the same voice request, and at least one language processing result corresponding to multiple VAD detections can be obtained.
[0082] After obtaining M language processing results, step S104 is executed.
[0083] In step S104, M language processing results can be compared to obtain a comparison result, and then based on the comparison result, a target language processing result is obtained from the M language processing results.
[0084] Specifically, when comparing M language processing results, the i-th language processing result can be compared with the (i + 1)-th language processing result, where i ranges from 1 to n - 1 in sequence; if it is found that the last language processing result is the same as a previous language processing result, the last language processing result is used as the target language processing result; if it is found that each language processing result is the same, the first language processing result can be used as the target language processing result.
[0085] Specifically, if N = 2, the processing result of the first language can be compared with the processing result of the second language to obtain the comparison result of the first language. If the comparison result of the first language indicates that the recognition results are consistent, the processing result of the first language is used as the target language processing result. If the comparison result of the first language indicates that the recognition results are inconsistent, the processing result of the second language is used as the target language processing result. If N = 3, the processing result of the first language can be compared with the processing result of the second language, and the processing result of the second language can be compared with the processing result of the third language. If it is compared that the processing result of the second language is inconsistent with the first speech recognition and the processing result of the second language is consistent with the processing result of the third language, the processing result of the second language or the processing result of the third language is used as the target recognition result. If the three language processing results are all consistent, any one of the three language processing results is used as the target language processing result.
[0086] Specifically, since N is an integer not less than 2, multiple end - sound detections can be performed on the real - time received voice data to obtain M language processing results. Each language processing result is obtained by performing natural language processing on the current voice data during the corresponding end - sound detection, so that M language processing results can be obtained before the VAD END event. Then, the target speech recognition result is determined from the M language processing results. Compared with the prior art, it is not necessary to perform speech recognition, NLU, DM, etc. after the VAD END event to obtain the target speech recognition result, so that the speech processing response time can be effectively shortened, thereby improving the response efficiency and making the user experience better.
[0087] In another embodiment, in order to improve the recognition efficiency of speech recognition, when determining the target language processing result corresponding to the speech request from the M language processing results, if N = 2, the following method can be adopted: after obtaining the processing result of the first language among the M language processing results, store the processing result of the first language in the stored recognition result in the cache; when obtaining the processing result of the second language among the N recognition results, if the recognition result of the first language is inconsistent with the recognition result of the second language, replace the stored recognition result from the processing result of the first language with the processing result of the second language. Of course, the recognition result of the first language can also be cached in the pre - processed language recognition result. If the recognition result of the first language is inconsistent with the recognition result of the second language, control the pre - processed language recognition result to be the recognition result of the second language; if the recognition result of the first language is consistent with the recognition result of the second language, control the pre - processed language recognition result to remain unchanged and still be the recognition result of the first language.
[0088] In another implementation, if N = 3, after obtaining the first language processing result among the M language processing results, store the first language processing result in the stored recognition result of the cache; when obtaining the second language processing result among the N recognition results, if the first language recognition result is inconsistent with the second language recognition result, replace the stored recognition result from the first language processing result with the second language processing result; when obtaining the third language processing result among the N recognition results, if the third language processing result is inconsistent with the second language processing result, replace the stored recognition result with the third language processing result. Of course, the first language recognition result can also be cached in the preprocessed language recognition result. If the first language recognition result is inconsistent with the second language recognition result, control the preprocessed language recognition result to be the second language recognition result; if the first language recognition result is consistent with the second language recognition result, control the preprocessed language recognition result to remain unchanged, still being the first language recognition result; and after obtaining the third language recognition result, compare the third language recognition result with the preprocessed language recognition result. If they are inconsistent, control the preprocessed language recognition result to be the third language recognition result; if they are consistent, control the preprocessed language recognition result to remain unchanged.
[0089] For example, refer to Figure 3 , taking N = 3 as an example, for N times of end sound detection, it includes the following steps:
[0090] Step 301: When the duration of the end silence reaches the set duration t1, obtain final asr1, where t1 is the first set duration and final asr1 is the first language recognition result;
[0091] Step S302: Update the preprocessing result to final asr1, and use the preprocessing result as the input for NLP processing to obtain final nlp1, where final nlp1 is the first speech recognition result, and the preprocessing result is the preprocessed language recognition result final asr.
[0092] Specifically, when the duration of the end silence reaches the first set duration t1, a recognition result corresponding to the currently collected audio, which is final asr1, will be obtained in advance. Send final asr1 to the NLP module for NLP processing to obtain an nlp result, that is, final nlp1, and cache it in the stored recognition result final nlp;
[0093] Step S303: When the duration of the end silence reaches the set duration t2, obtain final asr2, where t2 is the second set duration and final asr2 is the second language recognition result;
[0094] Step S304: Compare final asr2 with the preprocessing result;
[0095] If final asr2 is consistent with the preprocessing result, execute step S306; if not, execute step S305.
[0096] Specifically, since the data in the preprocessing result is final asr1 at this time, compare final asr1 with final asr2. If the comparison result is consistent, execute step S306; otherwise, execute step S305.
[0097] Step S305: Update the preprocessing result to final asr2, and use the updated preprocessing result as the input for NLP processing to update final nlp; at this time, final nlp2 is obtained after NLP processing, and the data in the updated final nlp is final nlp2.
[0098] Execute step S306 after executing step S305.
[0099] Step S306: When the duration of the tail-end silence reaches the set duration t3, obtain final asr3, where t3 is the third set duration and final asr3 is the third language recognition result;
[0100] Then execute step S307: Compare final asr3 with the preprocessing result data;
[0101] If it is determined according to step S307 that final asr2 is consistent with the preprocessing result, determine the preprocessing result as final asr1, and then compare final asr1 with final asr3; if it is determined according to step S307 that final asr2 is inconsistent with the preprocessing result, determine the updated preprocessing result as final asr2, and then compare final asr2 with final asr3.
[0102] If final asr3 is consistent with the preprocessing result, execute step S309; if not, execute step S308.
[0103] Step S308: Update the preprocessing result to final asr3, and use the updated preprocessing result as the input for NLP processing to update final nlp; at this time, final nlp3 is obtained after NLP processing, and the data in the updated final nlp is final nlp3.
[0104] After performing step S308, step S309 is performed.
[0105] Step S309: Return final nlp.
[0106] Specifically, if neither step S304 nor step S307 is executed, it is determined that the data in final nlp is finalnlp1, so the returned final nlp is determined to be final nlp1; if step S304 has been executed but step S307 has not been executed, it is determined that final nlp has been updated in step S304, so that the data in final nlp is final nlp2, and thus the returned final nlp is determined to be final nlp2; if both step S304 and step S307 have been executed, it is determined that final nlp has been updated in both step S304 and S307, so that the data in final nlp is final nlp3, and thus the returned final nlp is determined to be final nlp3.
[0107] It can be seen from this that by comparing the i-th language recognition result corresponding to the i-th end sound detection with the preprocessing result, where the value range of i is from 2 to N, if they are consistent, subsequent NLP processing is not required, and only when they are inconsistent is NLP processing performed on the i-th language recognition result. In this way, the number of NLP processing times can be effectively reduced, the NLP processing efficiency can be effectively improved, and the response time can be optimized.
[0108] In another embodiment, for the first end sound detection among N end sound detections, if it is detected that the duration of the first end silence exceeds the first set duration, the first language recognition result corresponding to the current voice data is obtained, and natural language processing is performed on the first language recognition result to obtain the first language processing result;
[0109] For the second end sound detection among N end sound detections, if it is detected that the duration of the second end silence exceeds the second set duration, the second language recognition result corresponding to the current voice data is obtained, and natural language processing is performed on the second language recognition result to obtain the second language processing result, where the current voice data at the time of the second end sound detection includes valid speech, and the second set duration is the same as the first set duration;
[0110] For the third end sound detection among N end sound detections, if it is detected that the duration of the third end silence exceeds the third set duration, the third language recognition result corresponding to the current voice data is obtained, and natural language processing is performed on the third language recognition result to obtain the third language processing result, where the third set duration is greater than the second set duration.
[0111] In the embodiments of this specification, the valid speech is human speech, etc.
[0112] For example, referring to Figure 2 , the cloud server B receives the voice data A1 sent by the vehicle terminal A, which includes the front - end voice 20 "I want to listen to Jay Chou's 'The Jasmine Flower'", the silence s0, the back - end voice 21 "The Jasmine Flower", and the silence s1. If n = 2 and the corresponding set durations are t1 and t2, then the first tail - end detection of A1 is performed using t1. If it is detected that t1 < s0 < t2, it is determined that the current voice data is the front - end voice 20 and the silence s0, and the current voice data is subjected to voice recognition to obtain final asr1, and final asr1 is subjected to NLP processing to obtain final nlp1, where final nlp1 is "I want to listen to Jay Chou's". During the second tail - end detection of A1, since the current voice data is the front - end voice 20, the silence s0, the back - end voice 21, and the silence s1, the current voice data includes the back - end voice 21 after the silence s0, which makes the current voice data include valid voice. Then, the second tail - end detection of the current voice data is continued using t1, the current voice data is subjected to voice recognition to obtain final asr2, and final asr2 is subjected to NLP processing to obtain final nlp2, where final nlp2 is "I want to listen to Jay Chou's The Jasmine Flower".
[0113] In the embodiments of this specification, the values and the number of the set durations can be set according to actual needs, or can be set by humans or devices themselves.
[0114] Specifically, when determining the values and the number of the set durations, the following test scheme can be adopted to obtain the values and the number of the set durations, which are as follows:
[0115] A1. Set 1 set duration t1, and several values [100, 200, 300, 400, 500] are taken within the range of t1 (0, t], where t = 700. The proportions of the number of times VAD1 is triggered equal to 1 and 2 in 10,000 pieces of corpus data are shown in Table 1 below:
[0116]
[0117] Table 1
[0118] As can be seen from Table 1, when the set duration is set to 300 ms, the response time is optimized by 400 ms. There is an 85% probability of not introducing additional computing power consumption and server costs, and a 9% probability of introducing additional computing power consumption and server costs once. Through the above method, the best values of the set durations and the number can be determined through comparative analysis of multiple groups of test data, so as to obtain the values of each set duration in t1 - tn and the number n.
[0119] In the embodiments of this specification, semantic understanding (Natural Language Understanding, abbreviated as NLU) refers to the process of understanding the text in the corresponding language and converting it into clear Actions and related parameters; and semantic generation (Natural Language Generation, abbreviated as NLG), which is opposite to NLU, refers to the process of generating clear Actions and related parameters and converting them into text in the corresponding language; and text-to-speech (Text To Speech, abbreviated as TTS) is used to convert text into a natural speech stream.
[0120] In another embodiment, after determining the target language processing result corresponding to the voice request from the M language processing results, natural speech can also be generated using the target language processing result, generating a target dialogue corresponding to the voice request, and performing voice output on the target dialogue.
[0121] For example, if the target language processing result is "I want to listen to Legend by Li Jian", then perform NLG processing on the target language processing result to obtain the target dialogue "Okay, playing Jasmine Flower by Jay Chou for you", and perform voice synthesis on the target dialogue and then perform voice output.
[0122] In the actual application process, as Figure 4 shown, an ASR system is set in the vehicle terminal, and the ASR system includes a VAD module 42, an ASR preprocessing module 43, an ASR module 44, an NLG module 45, a TTS module 46, and an NLP module 47. Thus, after the user 40 outputs the voice data of "I want to listen to Jasmine Flower by Jay Chou", after the vehicle terminal receives the voice data output by the user 40, the received voice data is input into the front-end signal module 41 for processing such as noise reduction and echo cancellation to obtain a front-end signal and send the front-end signal to the VAD module 42. At this time, step 1 is executed. When the VAD module 42 detects valid human voices in the front-end signal 41, it will trigger a VAD START event; and after the front-end signal reaches the ASR preprocessing module 43, a voice request is generated, and then step 2 is executed, sending the voice request to the ASR module 44, and then step 3 is executed. The ASR module 44 returns an acknowledgement signal, which can be, for example, 200 Ok.
[0123] After the vehicle terminal receives 200 Ok returned in step 3, it executes step 4, converts the front-end signal into an audio stream and sends it to the ASR module 44; then it executes step 5, and the VAD module 42 triggers a VAD1 event when it detects that the tail-end silence duration exceeds the first set duration; then it executes step 6, and sends the current voice data corresponding to VAD1 to the ASR module 44; after the ASR module 44 receives the current voice data corresponding to VAD1, it executes step 7, performs speech recognition to obtain final asr1, and returns final asr1 to the ASR preprocessing module 43; after the ASR preprocessing module 43 receives final asr1, it executes step 8, inputs final asr1 into the NLP module 47 to obtain final nlp1; after the NLP module 47 obtains final nlp1, it executes step 9, returns final nlp1 to the ASR preprocessing module 43, and caches final asr1 and final nlp1 into the corresponding final asr and final nlp.
[0124] In addition, after step 9, the VAD module 42 executes step 10, and the VAD module 42 triggers a VAD2 event when it detects that the tail-end silence duration exceeds the second set duration; then it executes step 11, and sends the current voice data corresponding to VAD2 to the ASR module 44; after the ASR module 44 receives the current voice data corresponding to VAD2, it executes step 12, performs speech recognition to obtain final asr2, and returns final asr2 to the ASR preprocessing module 43; correspondingly, the VAD module 42 executes step 13, and the VAD module 42 triggers a VAD END event when it detects that the tail-end silence duration exceeds the third set duration; then it executes step 14, and sends the current voice data corresponding to VAD END to the ASR module 44; after the ASR module 44 receives the current voice data corresponding to VAD END, it executes step 15, performs speech recognition to obtain final asr3, and returns final asr3 to the ASR preprocessing module 43. Among them, after step 12, steps 8 and 9 need to be executed. In step 8, final asr1 and final asr2 need to be compared. If they are the same, the final asr in the cache is not modified; if they are different, the final asr in the cache is replaced with final asr2, and then step 9 is executed, and the final nlp in the cache is replaced from final nlp1 with final nlp2. Correspondingly, after step 15, steps 8 and 9 also need to be executed. Among them, the specific implementation processes of steps 8 and 9 refer to steps 8 and 9 after step 12. For the sake of simplicity of the specification, they will not be elaborated here.
[0125] Further, the ASR preprocessing module 43 executes step 16, reads the data in the final nlp from the cache, represented by final nlp, and sends the final nlp to the NLG module 45; after receiving the final nlp, the NLG module 45 uses the final nlp to perform natural speech generation to generate a target dialogue. For example, the target dialogue is "Okay, play Qi Li Xiang by Jay Chou", and then executes step 17, sends the target dialogue to the TTS module 46 for speech synthesis, and outputs the target dialogue as speech.
[0126] It can be seen from this that based on the description that step 8 and step 9 need to be executed after step 12, after receiving the final asr1, it is only necessary to compare it with the final asr0. If they are the same, the cached final nlp0 is directly output as the final nlp result. Theoretically, the optimized response time is (t - t1), where t is the value range of the set duration, and t1 is the first set duration of the first VAD detection. In this way, when the latter language recognition result is the same as the previous one, NLP processing will not be performed, which can effectively reduce the computational amount, thereby improving the processing efficiency and optimizing the response time.
[0127] In another embodiment, the ASR module 44 and the NLP module 47 can be set in the cloud server. At this time, the ASR preprocessing module 43 is responsible for receiving VAD events from the VAD module 42. If it is found that the identifier of the VAD shows an increasing trend, only the first VAD event is processed. For example, when receiving VAD2, if it is found that the identifier 2 is greater than the identifier 1 in VAD1, this message event is not sent to the cloud server and is directly ignored; when the ASR module 44 receives VAD END, it directly outputs the cached final nlp1 as the final nlp to the NLG module 45. At this time, the ASR preprocessing module 43 can directly output the cached final nlp1 without waiting for the cloud server to return the final nlp after receiving VAD END. Therefore, it can save the time consumption caused by the end-cloud link interaction, and then output the returned final nlp after receiving the final nlp returned by the cloud server.
[0128] One or at least one of the above technical solutions in the embodiments of the present application has at least the following technical effects:
[0129] Based on the above technical solution, after receiving a voice request, voice data is received in real time; the voice data received in real time is subjected to N times of tail - end detection; M language processing results corresponding to the N times of tail - end detection are obtained; a target language processing result corresponding to the voice request is determined from the M language processing results; at this time, since N is an integer not less than 2, when performing the 2nd to Nth tail - end detections, first, if the (i - 1)th language recognition result is inconsistent with the ith language recognition result, the stored recognition result is replaced from the (i - 1)th language processing result with the ith language processing result corresponding to the ith language recognition result; if they are consistent, the stored recognition result is controlled to remain unchanged, where i is a positive integer, and i is taken from 2 to N in sequence. In this way, M language processing results can be obtained before the VAD END event, and then the target speech recognition result is determined from the M language processing results. Compared with the prior art, it is not necessary to perform speech recognition, NLU, DM, etc. after the VAD END event to obtain the target speech recognition result, so that the voice processing response time can be effectively shortened, and the response efficiency is improved, making the user experience better.
[0130] For the above - mentioned embodiment, a voice processing method for audio tail - end detection is provided. The embodiments of the present application also correspondingly provide a voice processing device for audio tail - end detection. Please refer to Figure 5 , the device includes:
[0131] A voice receiving unit 501, configured to receive voice data in real time after receiving a voice request;
[0132] A tail - end detection unit, configured to perform N times of tail - end detection on the voice data received in real time, where N is an integer not less than 2;
[0133] A speech recognition unit 502, configured to obtain M language processing results corresponding to the N times of tail - end detection;
[0134] A recognition result obtaining unit 503, configured to determine a target language processing result corresponding to the voice request from the M language processing results.
[0135] In an alternative embodiment, the speech recognition unit 502 is configured to, for the first vowel detection among the N vowel detections, if it is detected that the duration of the first trailing silence exceeds the first set duration, obtain the first language recognition result corresponding to the current speech data, perform natural language processing on the first language recognition result to obtain a first language processing result, and store the first language processing result in the stored recognition result in the cache; for the i-th vowel detection among the N vowel detections, if it is detected that the duration of the i-th trailing silence exceeds the i-th set duration, obtain the i-th language recognition result corresponding to the current speech data, and if the (i - 1)th language recognition result is inconsistent with the i-th language recognition result, replace the stored recognition result from the (i - 1)th language processing result with the i-th language processing result corresponding to the i-th language recognition result; if they are consistent, control the stored recognition result to remain unchanged, where i is a positive integer, and i is sequentially taken from 2 to N.
[0136] In an alternative embodiment, the speech recognition unit 502 is configured to, for the first vowel detection among the N vowel detections, if it is detected that the duration of the first trailing silence exceeds the first set duration, obtain the first language recognition result corresponding to the current speech data, perform natural language processing on the first language recognition result to obtain a first language processing result, store the first language processing result in the stored recognition result, and store the first language recognition result in the preprocessed language recognition result in the cache; for the second vowel detection among the N vowel detections, if it is detected that the duration of the second trailing silence exceeds the second set duration, obtain the second language recognition result corresponding to the current speech data, compare the second language recognition result with the first language recognition result, and if it is determined that the first language recognition result is inconsistent with the second language recognition result, replace the preprocessed language recognition result with the second language recognition result, perform natural language processing on the second language recognition result to obtain a second language processing result, and replace the stored recognition result with the second language processing result; if it is determined that the first language recognition result is consistent with the second language recognition result, control the preprocessed language recognition result to be the first language recognition result, and control the stored recognition result to be the first language processing result, where the second set duration is greater than the first set duration.
[0137] In an alternative embodiment, the recognition result acquisition unit 503 is configured to, if the first language recognition result is consistent with the second language recognition result, use the first language processing result stored in the stored recognition result as the target language processing result; if the first language recognition result is inconsistent with the second language recognition result, use the second language processing result stored in the stored recognition result as the target language processing result.
[0138] In an alternative embodiment, the recognition result acquisition unit 503 is configured to compare the first language processing result and the second language processing result to obtain a first language comparison result; and determine the target language processing result from the M language processing results according to the first language comparison result.
[0139] In an alternative embodiment, the speech recognition unit 502 is configured to, if N = 3, for the first vowel detection among the N vowel detections, if it is detected that the duration of the first end silence exceeds the first set duration, obtain the first language recognition result corresponding to the current speech data, perform natural language processing on the first language recognition result to obtain a first language processing result, store the first language processing result in the stored recognition result, and store the first language recognition result in the preprocessed language recognition result in the cache; for the second vowel detection among the N vowel detections, if it is detected that the duration of the second end silence exceeds the second set duration, obtain the second language recognition result corresponding to the current speech data, compare the second language recognition result with the first language recognition result, if it is determined that the first language recognition result is inconsistent with the second language recognition result, replace the preprocessed language recognition result with the second language recognition result, perform natural language processing on the second language recognition result to obtain a second language processing result, and replace the stored recognition result with the second language processing result; if it is determined that the first language recognition result is consistent with the second language recognition result, control the preprocessed language recognition result to be the first language recognition result, and control the stored recognition result to be the first language processing result, where the second set duration is greater than the first set duration; for the third vowel detection among the N vowel detections, if it is detected that the duration of the third end silence exceeds the third set duration, obtain the third language recognition result corresponding to the current speech data, compare the third language recognition result with the language recognition result stored in the preprocessed language recognition result, if it is determined that the language recognition results are inconsistent, replace the preprocessed language recognition result with the third language recognition result, perform natural language processing on the third language recognition result to obtain a third language processing result, and replace the stored recognition result with the third language processing result; if it is determined that the language recognition results are consistent, control the preprocessed language recognition result and the stored recognition result to remain unchanged, where the third set duration is greater than the second set duration.
[0140] In an alternative embodiment, the recognition result acquisition unit 503 is configured to, after comparing the third language recognition result with the language recognition results stored in the preprocessing language recognition result, if the recognized language recognition results are consistent, use the currently stored recognition result in the cache as the target language processing result; if the third language recognition result is inconsistent with the second language recognition result, use the third language processing result stored in the stored recognition result as the target language processing result.
[0141] In an alternative embodiment, the apparatus further includes:
[0142] A voice output unit, configured to use the target language processing result to generate natural speech, generate a target dialogue corresponding to the voice request, and perform voice output of the target dialogue.
[0143] Regarding the apparatus in the above embodiments, the specific manners in which each module performs operations have been described in detail in the embodiments related to the method, and will not be elaborated herein.
[0144] Figure 6 FIG. 800 is a block diagram of an electronic device 800 for a voice processing method for audio tail end detection according to an exemplary embodiment. For example, the electronic device 800 may be a mobile phone, a computer, a digital broadcast terminal, a messaging device, a game console, a tablet device, a medical device, a fitness device, a personal digital assistant, etc.
[0145] Referring to Figure 6 , the electronic device 800 may include one or more of the following components: a processing component 802, a memory 804, a power component 806, a multimedia component 808, an audio component 810, an input / presentation (I / O) interface 812, a sensor component 814, and a communication component 816.
[0146] The processing component 802 generally controls the overall operation of the electronic device 800, such as operations associated with display, telephone calls, data communication, camera operations, and recording operations. The processing component 802 may include one or more processors 820 to execute instructions to complete all or part of the steps of the above method. In addition, the processing component 802 may include one or more modules to facilitate the interaction between the processing component 802 and other components. For example, the processing component 802 may include a multimedia module to facilitate the interaction between the multimedia component 808 and the processing component 802.
[0147] The memory 804 is configured to store various types of data to support the operation of the device 800. Examples of such data include instructions for any application or method operating on the electronic device 800, contact data, phone book data, messages, pictures, videos, and the like. The memory 804 can be implemented by any type of volatile or non-volatile storage device or a combination thereof, such as static random access memory (SRAM), electrically erasable programmable read-only memory (EEPROM), erasable programmable read-only memory (EPROM), programmable read-only memory (PROM), read-only memory (ROM), magnetic memory, flash memory, magnetic disk, or optical disk.
[0148] The power supply component 806 provides power to various components of the electronic device 800. The power supply component 806 may include a power management system, one or more power supplies, and other components associated with generating, managing, and distributing power for the electronic device 800.
[0149] The multimedia component 808 includes a screen that provides a presentation interface between the electronic device 800 and the user. In some embodiments, the screen may include a liquid crystal display (LCD) and a touch panel (TP). If the screen includes a touch panel, the screen can be implemented as a touch screen to receive input signals from the user. The touch panel includes one or more touch sensors to sense touches, swipes, and gestures on the touch panel. The touch sensors can not only sense the boundaries of touch or swipe actions, but also detect the duration and pressure associated with the touch or swipe operation. In some embodiments, the multimedia component 808 includes a front camera and / or a rear camera. When the device 800 is in an operating mode, such as a shooting mode or a video mode, the front camera and / or the rear camera can receive external multimedia data. Each of the front camera and the rear camera can be a fixed optical lens system or have a focal length and optical zoom capabilities.
[0150] The audio component 810 is configured to present and / or input audio signals. For example, the audio component 810 includes a microphone (MIC) that is configured to receive external audio signals when the electronic device 800 is in an operating mode, such as a call mode, a recording mode, and a voice recognition mode. The received audio signals can be further stored in the memory 804 or transmitted via the communication component 816. In some embodiments, the audio component 810 further includes a speaker for presenting audio signals.
[0151] The I / O interface 812 provides an interface between the processing component 802 and a peripheral interface module, which can be a keyboard, a click wheel, buttons, etc. These buttons can include, but are not limited to: a home button, a volume button, a power-on button, and a lock button.
[0152] The sensor assembly 814 includes one or more sensors for providing an assessment of the status of various aspects of the electronic device 800. For example, the sensor assembly 814 can detect the on / off state of the device 800, the relative positioning of components, such as the display and keypad of the electronic device 800. The sensor assembly 814 can also detect a change in the position of the electronic device 800 or a component of the electronic device 800, the presence or absence of user contact with the electronic device 800, the orientation or acceleration / deceleration of the electronic device 800, and the temperature change of the electronic device 800. The sensor assembly 814 can include a proximity sensor configured to detect the presence of nearby objects without any physical contact. The sensor assembly 814 can also include a light sensor, such as a CMOS or CCD image sensor, for use in imaging applications. In some embodiments, the sensor assembly 814 can also include an acceleration sensor, a gyroscope sensor, a magnetic sensor, a pressure sensor, or a temperature sensor.
[0153] The communication component 816 is configured to facilitate communication between the electronic device 800 and other devices in a wired or wireless manner. The electronic device 800 can access a wireless network based on communication standards, such as WiFi, 2G, or 3G, or a combination thereof. In an exemplary embodiment, the communication component 816 receives a broadcast signal or broadcast-related information from an external broadcast management system via a broadcast channel. In an exemplary embodiment, the communication component 816 further includes a near field communication (NFC) module to facilitate short-range communication. For example, the NFC module can be implemented based on radio frequency identification (RFID) technology, infrared data association (IrDA) technology, ultra-wideband (UWB) technology, Bluetooth (BT) technology, and other technologies.
[0154] In an exemplary embodiment, the electronic device 800 can be implemented by one or more application specific integrated circuits (ASICs), digital signal processors (DSPs), digital signal processing devices (DSPDs), programmable logic devices (PLDs), field programmable gate arrays (FPGAs), controllers, microcontrollers, microprocessors, or other electronic components for performing the above method.
[0155] In an exemplary embodiment, a non-transitory computer-readable storage medium including instructions is also provided, such as the memory 804 including instructions that can be executed by the processor 820 of the electronic device 800 to complete the above method. For example, the non-transitory computer-readable storage medium can be a ROM, a random access memory (RAM), a CD-ROM, a magnetic tape, a floppy disk, and an optical data storage device, etc.
[0156] In addition, it should be noted that: The embodiments of the present application also provide a computer program product or a computer program. The computer program product or the computer program may include computer instructions, and the computer instructions may be stored in a computer-readable storage medium. The processor of the computer device reads the computer instructions from the computer-readable storage medium, and the processor may execute the computer instructions to enable the computer device to execute the description of the voice processing method for audio tail end detection in the corresponding embodiments described above. Figure 1 Therefore, it will not be elaborated here. In addition, the description of the beneficial effects of adopting the same method will not be elaborated either. For the technical details not disclosed in the computer program product or the computer program embodiments involved in the present application, please refer to the description of the method embodiments of the present application.
[0157] Those skilled in the art will readily conceive of other embodiments of the present invention after considering the specification and practicing the invention disclosed herein. This application is intended to cover any variations, uses, or adaptations of the present invention that follow the general principles of the present invention and include common general knowledge or conventional technical means in the technical field not disclosed in this disclosure. The specification and embodiments are to be considered exemplary only, and the true scope and spirit of the present invention are pointed out by the following claims.
[0158] It should be understood that the present invention is not limited to the exact structures already described and shown in the drawings, and various modifications and changes can be made without departing from its scope. The scope of the present invention is limited only by the appended claims.
[0159] The above are only the preferred embodiments of the present invention and are not intended to limit the present invention. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principle of the present invention shall be included within the protection scope of the present invention.
Claims
1. A voice processing method for audio end detection, characterized in that, The method includes: After receiving a voice request, receiving voice data in real time; Performing N times of end - sound detection on the voice data received in real time, where N is an integer not less than 2; Obtaining M language processing results corresponding to the N times of end - sound detection, where M is a positive integer not greater than N; Determining a target language processing result corresponding to the voice request from the M language processing results; The obtaining M language processing results corresponding to the N times of end - sound detection includes: For the first end - sound detection among the N times of end - sound detection, if it is detected that the first end - silence duration exceeds the first set duration, obtaining a first language recognition result corresponding to the current voice data, performing natural language processing on the first language recognition result to obtain a first language processing result, and storing the first language processing result in a cached storage recognition result; For the i - th end - sound detection among the N times of end - sound detection, if it is detected that the i - th end - silence duration exceeds the i - th set duration, obtaining an i - th language recognition result corresponding to the current voice data. If the (i - 1) - th language recognition result is inconsistent with the i - th language recognition result, then replacing the storage recognition result from the (i - 1) - th language processing result with the i - th language processing result corresponding to the i - th language recognition result; if they are consistent, then controlling the storage recognition result to remain unchanged, where i is a positive integer, i is taken from 2 to N in sequence. After completing the N times of end - sound detection, using the data in the storage recognition result as the target language processing result, and the first set duration and the i - th set duration increase in sequence.
2. The method according to claim 1, characterized in that, If N = 2, the obtaining N language processing results corresponding to the N times of end - sound detection includes: For the first end - sound detection among the N times of end - sound detection, if it is detected that the first end - silence duration exceeds the first set duration, obtaining a first language recognition result corresponding to the current voice data, performing natural language processing on the first language recognition result to obtain a first language processing result, storing the first language processing result in the storage recognition result, and storing the first language recognition result in a cached pre - processed language recognition result; For the second end - sound detection among the N times of end - sound detection, if it is detected that the second end - silence duration exceeds the second set duration, obtaining a second language recognition result corresponding to the current voice data, comparing the second language recognition result with the first language recognition result. If it is found that the first language recognition result is inconsistent with the second language recognition result, then replacing the pre - processed language recognition result with the second language recognition result, performing natural language processing on the second language recognition result to obtain a second language processing result, and replacing the storage recognition result with the second language processing result; if it is found that the first language recognition result is consistent with the second language recognition result, then controlling the pre - processed language recognition result to be the first language recognition result and controlling the storage recognition result to be the first language processing result, where the second set duration is greater than the first set duration.
3. The method according to claim 2, wherein Determining the target language processing result corresponding to the voice request from the M language processing results includes: If the first language recognition result is consistent with the second language recognition result, use the first language processing result stored in the stored recognition result as the target language processing result; if the first language recognition result is inconsistent with the second language recognition result, use the second language processing result stored in the stored recognition result as the target language processing result.
4. The method according to claim 3, wherein Determining the target language processing result corresponding to the voice request from the M language processing results includes: Compare the first language processing result and the second language processing result to obtain a first language comparison result; Determine the target language processing result from the M language processing results according to the first language comparison result.
5. The method according to claim 1, characterized in that, If N = 3, obtaining the M language processing results corresponding to the N times of end sound detection includes: For the first end sound detection among the N times of end sound detection, if it is detected that the first end silence duration exceeds the first set duration, obtain the first language recognition result corresponding to the current voice data, perform natural language processing on the first language recognition result to obtain a first language processing result, store the first language processing result in the stored recognition result, and store the first language recognition result in the preprocessed language recognition result in the cache; For the second end sound detection among the N times of end sound detection, if it is detected that the second end silence duration exceeds the second set duration, obtain the second language recognition result corresponding to the current voice data, compare the second language recognition result with the first language recognition result. If it is compared that the first language recognition result is inconsistent with the second language recognition result, replace the preprocessed language recognition result with the second language recognition result, perform natural language processing on the second language recognition result to obtain a second language processing result, and replace the stored recognition result with the second language processing result; if it is compared that the first language recognition result is consistent with the second language recognition result, control the preprocessed language recognition result to be the first language recognition result, and control the stored recognition result to be the first language processing result, where the second set duration is greater than the first set duration; For the third tail - end detection in the N - th tail - end detection, if it is detected that the duration of the third tail - end silence exceeds the third set duration, obtain the third language recognition result corresponding to the current voice data, compare the third language recognition result with the language recognition results stored in the pre - processed language recognition result. If it is found that the language recognition results are inconsistent, replace the pre - processed language recognition result with the third language recognition result, perform natural language processing on the third language recognition result to obtain the third language processing result, and replace the stored recognition result with the third language processing result; if it is found that the language recognition results are consistent, control the pre - processed language recognition result and the stored recognition result to remain unchanged, where the third set duration is greater than the second set duration.
6. The method according to claim 5, wherein Determining the target language processing result corresponding to the voice request from the M language processing results includes: After comparing the third language recognition result with the language recognition results stored in the pre - processed language recognition result, if it is found that the language recognition results are consistent, use the currently stored recognition result in the cache as the target language processing result; if the third language recognition result is inconsistent with the second language recognition result, use the third language processing result stored in the stored recognition result as the target language processing result.
7. The method according to any one of claims 1-6, characterized in that, After determining the target language processing result corresponding to the voice request from the M language processing results, the method includes: Use the target language processing result to perform natural speech generation, generate a target dialogue corresponding to the voice request, and perform voice output of the target dialogue.
8. A voice processing device for audio end detection, characterized in that, The device includes: A voice receiving unit, configured to receive voice data in real - time after receiving a voice request; A tail - end detection unit, configured to perform N times of tail - end detection on the voice data received in real - time, where N is an integer not less than 2; A voice recognition unit, configured to obtain M language processing results corresponding to the N times of tail - end detection; The voice recognition unit is configured to, for the first tail - end detection in the N times of tail - end detection, if it is detected that the duration of the first tail - end silence exceeds the first set duration, obtain the first language recognition result corresponding to the current voice data, perform natural language processing on the first language recognition result to obtain the first language processing result, and store the first language processing result in the stored recognition result in the cache; for the i - th tail - end detection in the N times of tail - end detection, if it is detected that the duration of the i - th tail - end silence exceeds the i - th set duration, obtain the i - th language recognition result corresponding to the current voice data. If the (i - 1) - th language recognition result is inconsistent with the i - th language recognition result, replace the stored recognition result from the (i - 1) - th language processing result with the i - th language processing result corresponding to the i - th language recognition result; if they are consistent, control the stored recognition result to remain unchanged, where i is a positive integer, and i is taken from 2 to N in sequence; the first set duration and the i - th set duration increase in sequence. An identification result acquisition unit, configured to determine a target language processing result corresponding to the voice request from the M language processing results, and after N times of tail sound detection, use the data in the stored recognition result as the target language processing result.
9. An electronic device, characterized in that, It includes a memory and one or more programs, wherein the one or more programs are stored in the memory and are configured to be executed by one or more processors for performing operation instructions included in the one or more programs corresponding to the method according to any one of claims 1 to 7.
10. A computer program product, characterized in that, The computer program product includes computer instructions, which are stored in a computer-readable storage medium and are adapted to be read and executed by a processor, so that a computer device having the processor executes the method according to any one of claims 1-7.
Citation Information
Patent Citations
Voice end-point detection device, system and method
US20180357999A1