Voice processing method and device

By performing parallel online and offline speech recognition of to-processed speech and performing arbitration processing for natural language understanding, the problem of wasting time in the wake-up sound area of ​​the speech processing system in the prior art is solved, and efficient speech processing and result availability are achieved.

CN120164461APending Publication Date: 2025-06-17BEIJING CHJ AUTOMOTIVE TECH CO LTD
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
CN202311723805.4
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2023-12-14
Publication Date
2025-06-17

AI Technical Summary

Technical Problem

In the voice processing of the wake-up sound area, the integration of offline automatic speech recognition (ASR) and online ASR requires a waste of time and cannot meet the pain points of available offline or individual online results.

Method used

By obtaining and performing online and offline speech recognition on the pending speech, the results of online and offline speech recognition are obtained, respectively, and then natural language understanding (NLU) is processed on these results. Send the NLU results to the natural language understanding result queue and output the final result through arbitration processing.

Benefits of technology

Parallel processing of offline ASR and online ASR is implemented, avoiding time waste caused by fusion, and ensuring the availability of individual offline or individual online results through arbitration processing.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120164461A_ABST
    Figure CN120164461A_ABST
Patent Text Reader

Abstract

The invention provides a voice processing method and device, and relates to the technical field of data processing. The method comprises the following steps: in response to a to-be-processed voice source from a wake-up voice area, respectively carrying out online voice recognition and offline voice recognition on the to-be-processed voice to obtain a first online voice recognition result and a first offline voice recognition result; performing online and offline natural language understanding on the first online voice recognition result and the first offline voice recognition result respectively; sending the first online natural language understanding result and the first offline natural language understanding result to a natural language understanding result queue; and performing arbitration processing on the natural language understanding result based on the natural language understanding result queue, and outputting an arbitration processing result. According to the voice processing of the wake-up voice region, the offline ASR and online ASR parallel can be realized, the situation of time waste caused by the fusion of the offline ASR and the online ASR is avoided, and the voice processing efficiency is improved. Individual offline or individual online results may be judged to be available by arbitration processing.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present disclosure relates to the field of data processing, and in particular, to a voice processing method and apparatus. Background Art

[0002] For the voice processing in the wake-up sound area of the currently available voice processing systems, the offline automatic speech recognition (Automatic Speech Recognition, hereinafter referred to as ASR) and the online ASR are first fused, and then the fused ASR results are passed through the offline natural language understanding (Natural Language Understanding, hereinafter referred to as NLU) and the online NLU. After the offline NLU results and the online NLU results are fused again, the required results are obtained. However, in the related art, the fusion of the offline ASR and the online ASR takes time, and there is a pain point that the separate offline or separate online results cannot be satisfied. Summary of the Invention

[0003] The present disclosure provides a voice processing method, apparatus, electronic device, storage medium, program product, vehicle networking module, and vehicle.

[0004] According to a first aspect of the present disclosure, there is provided a voice processing method, the method including: obtaining a voice to be processed, and in response to the voice to be processed originating from a wake-up sound area, performing online speech recognition and offline speech recognition on the voice to be processed respectively to obtain a first online speech recognition result and a first offline speech recognition result; performing online natural language understanding on the first online speech recognition result to obtain a first online natural language understanding result, and performing offline natural language understanding on the first offline speech recognition result to obtain a first offline natural language understanding result; sending the first online natural language understanding result and the first offline natural language understanding result to a natural language understanding result queue; performing arbitration processing on the natural language understanding results based on the natural language understanding result queue, and outputting the result of the arbitration processing.

[0005] In some embodiments, performing arbitration processing on the natural language understanding results based on the natural language understanding result queue and outputting the result of the arbitration processing includes: obtaining the first natural language understanding result sent to the natural language understanding result queue; if the natural language understanding result is the first online natural language understanding result, determining whether the first online natural language understanding result is valid; if the first online natural language understanding result is valid, outputting the first online natural language result.

[0006] In some embodiments, after determining whether the first online natural language understanding result is valid, the method further includes: if the first online natural language understanding result is invalid, waiting for the first offline natural language understanding result to be sent to the natural language understanding result queue; after a preset waiting time, based on the natural language understanding result queue, determining whether the first offline natural language understanding result is valid; if the first offline natural language understanding result is valid, outputting the first offline natural language understanding result.

[0007] In some embodiments, after obtaining the first natural language understanding result sent to the natural language understanding result queue, the method further includes: if the natural language understanding result is the first offline natural language understanding result, determining whether the first offline natural language understanding result requires an online information source; if the first offline natural language understanding result requires an online information source, waiting for the first online natural language understanding result to be sent to the natural language understanding result queue; after a preset waiting time, based on the natural language understanding result queue, determining whether the first online natural language understanding result is valid; if the first online natural language understanding result is valid, outputting the first online natural language understanding result.

[0008] In some embodiments, after obtaining the speech to be processed, the method further includes: in response to the speech to be processed originating from a non-awakening sound area, selecting to perform speech recognition and natural language understanding on the speech to be processed in a preset manner, where the preset manner includes online or offline.

[0009] In some embodiments, in response to the speech to be processed originating from a non-awakening sound area, selecting to perform speech recognition and natural language understanding on the speech to be processed in a preset manner includes: selecting to perform online speech recognition on the speech to be processed to obtain a second online speech recognition result; performing online natural language understanding on the second online speech recognition result to obtain a second online natural language understanding result; sending the second online natural language understanding result to the natural language understanding result queue.

[0010] In some embodiments, in response to the speech to be processed originating from a non-awakening sound area, selecting to perform speech recognition on the speech to be processed in a preset manner includes: selecting to perform offline speech recognition on the speech to be processed to obtain a second offline speech recognition result; performing offline natural language understanding on the second offline speech recognition result to obtain a second offline natural language understanding result; according to the classification of the second offline natural language understanding result, determining whether the second offline natural language understanding result requires supplementing with an online information source; in response to the second offline natural language understanding result requiring supplementing with an online information source, performing online natural language understanding on the second offline speech recognition result to obtain a second online natural language understanding result; sending the second online natural language understanding result to the natural language understanding result queue.

[0011] In some embodiments, after determining whether the second offline natural language understanding result needs to supplement online information sources, the method further includes: in response to the second offline natural language understanding result not needing to supplement online information sources, sending the second offline natural language understanding result to the natural language understanding result queue.

[0012] In some embodiments, arbitration processing of natural language understanding results is performed based on the natural language understanding result queue, and the output of the arbitration processing includes: obtaining the first natural language understanding result sent to the natural language understanding result queue; if the natural language understanding result is the second online natural language understanding result or the second offline natural language understanding result, directly outputting the natural language understanding result.

[0013] According to a second aspect of the present disclosure, there is provided a voice processing device, the device includes: a speech recognition unit, configured to obtain the speech to be processed, and in response to the speech to be processed originating from the wake-up sound area, perform online speech recognition and offline speech recognition on the speech to be processed respectively to obtain a first online speech recognition result and a first offline speech recognition result; a natural language understanding unit, configured to perform online natural language understanding on the first online speech recognition result to obtain a first online natural language understanding result, and perform offline natural language understanding on the first offline speech recognition result to obtain a first offline natural language understanding result; a transmission unit, configured to send the first online natural language understanding result and the first offline natural language understanding result to the natural language understanding result queue; an arbitration unit, configured to perform arbitration processing of natural language understanding results based on the natural language understanding result queue, and output the result of the arbitration processing.

[0014] According to a third aspect of the present disclosure, there is provided an electronic device, including: at least one processor; and a memory that is speech-processing connected to the at least one processor; wherein, the memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor so that the at least one processor can execute the method of the foregoing first aspect.

[0015] According to a fourth aspect of the present disclosure, there is provided a non-transitory computer-readable storage medium storing computer instructions, wherein the computer instructions are used to cause a computer to execute the method of the foregoing first aspect.

[0016] According to a fifth aspect of the present disclosure, there is provided a computer program product, including a computer program, and the computer program realizes the method of the foregoing first aspect when executed by a processor.

[0017] The voice processing method provided by the embodiments of the present disclosure obtains the voice to be processed. In response to the voice to be processed originating from the wake-up sound area, the voice to be processed is respectively subjected to online voice recognition and offline voice recognition to obtain a first online voice recognition result and a first offline voice recognition result; the first online voice recognition result is subjected to online natural language understanding to obtain a first online natural language understanding result, and the first offline voice recognition result is subjected to offline natural language understanding to obtain a first offline natural language understanding result; the first online natural language understanding result and the first offline natural language understanding result are sent to the natural language understanding result queue; arbitration processing of the natural language understanding results is performed based on the natural language understanding result queue, and the result of the arbitration processing is output. The method of the present disclosure can implement parallel offline ASR and online ASR for voice processing in the wake-up sound area, avoiding the situation of wasting time caused by the fusion of offline ASR and online ASR, and can determine that the separate offline or separate online results are available through arbitration processing.

[0018] It should be understood that the content described in this part is not intended to identify the key or important features of the embodiments of the present application, nor is it used to limit the scope of the present application. Other features of the present application will become easily understood through the following description. BRIEF DESCRIPTION OF THE DRAWINGS

[0019] The drawings are used to better understand the solution and do not constitute a limitation to the present disclosure. Among them:

[0020] Figure 1 is a flowchart of a voice processing method provided by an embodiment of the present disclosure;

[0021] Figure 2 is an example diagram of a voice processing method provided by an embodiment of the present disclosure;

[0022] Figure 3 is a flowchart of a voice processing method provided by an embodiment of the present disclosure;

[0023] Figure 4 is a flowchart of a voice processing method provided by an embodiment of the present disclosure;

[0024] Figure 5 is a flowchart of a voice processing method provided by an embodiment of the present disclosure;

[0025] Figure 6 is an example diagram of natural language understanding result arbitration processing provided by an embodiment of the present disclosure;

[0026] Figure 7 is a flowchart of natural language understanding result arbitration processing provided by an embodiment of the present disclosure;

[0027] Figure 8A schematic flowchart of the arbitration processing of natural language understanding results provided by an embodiment of the present disclosure;

[0028] Figure 9 A schematic structural diagram of a voice processing device provided by an embodiment of the present disclosure;

[0029] Figure 10 A schematic block diagram of an exemplary electronic device 700 provided by an embodiment of the present disclosure. Detailed implementation manners

[0030] The following describes exemplary embodiments of the present disclosure with reference to the accompanying drawings. Various details of the embodiments of the present disclosure are included to facilitate understanding, and they should be considered merely exemplary. Therefore, those of ordinary skill in the art should recognize that various changes and modifications can be made to the embodiments described herein without departing from the scope and spirit of the present disclosure. Similarly, for clarity and conciseness, descriptions of well-known functions and structures are omitted in the following description.

[0031] Currently, all in-vehicle voice systems on the market do not handle non-awakened ASR and semantic processing. For awakened users, offline ASR and online ASR are first fused, and then the fused ASR results are passed through offline NLU and online NLU. The results of offline NLU and online NLU are fused again to obtain the required results.

[0032] The following describes in detail a voice processing method, device, electronic device, storage medium, and program product proposed by the present disclosure with reference to the accompanying drawings.

[0033] Figure 1 A schematic flowchart of a voice processing method provided by an embodiment of the present disclosure. This method is applied to a voice control system, and the execution entity is a voice controller, such as an in-vehicle voice chip.

[0034] As Figure 1 shown, this method includes the following steps:

[0035] Step 101, obtain the voice to be processed. In response to the voice to be processed originating from the wake-up sound area, perform online speech recognition and offline speech recognition on the voice to be processed respectively to obtain a first online speech recognition result and a first offline speech recognition result.

[0036] In some embodiments of the present disclosure, the wake-up sound area or wake-up channel refers to the channel that clearly gives a wake-up signal. For example, in an in-vehicle scenario, when the driver says a wake-up word such as "hi, friend", the voice is received by the microphone and sent to the wake-up module. The wake-up module generates a wake-up signal, and the wake-up sound area is activated. After that, any voice spoken by the driver will be sent to the wake-up sound area.

[0037] In some embodiments of the present disclosure, parallel ASR is adopted, and the single-tone region speech recognition offline ASR and online ASR output in parallel.

[0038] In some embodiments, after obtaining the speech to be processed, it further includes: determining whether the speech to be processed originates from the wake-up sound region.

[0039] Step 102, perform online natural language understanding on the first online speech recognition result to obtain the first online natural language understanding result, and perform offline natural language understanding on the first offline speech recognition result to obtain the first offline natural language understanding result.

[0040] In some embodiments, for the processing of offline and online ASR in the wake-up sound region, the offline ASR is directly sent to the offline NLU for processing, and the online ASR is directly sent to the online NLU for processing.

[0041] Step 103, send the first online natural language understanding result and the first offline natural language understanding result to the natural language understanding result queue.

[0042] In some embodiments, the method further includes: determining whether the first online natural language understanding result and the first offline natural language understanding result in the natural language understanding result queue are complete.

[0043] In some embodiments, if the first online natural language understanding result has been sent to the natural language understanding result queue, it is necessary to wait for the first offline natural language understanding result to be sent to the natural language understanding result queue. Conversely, if the first offline natural language understanding result has been sent to the natural language understanding result queue, it is necessary to wait for the first online natural language understanding result to be sent to the natural language understanding result queue. When the first online natural language understanding result and the first offline natural language understanding result in the natural language understanding result queue are complete, the natural language understanding result to be sent to the DM for execution can be selected according to the service requirements.

[0044] In some embodiments, as Figure 2 shown, the microphone mic receives the speech to be processed, and the speech to be processed can be processed through the speech software development kit, and the processing result is sent to the ASR queue. If the speech to be processed originates from the wake-up sound region, offline ASR and online ASR are performed in parallel. The obtained offline ASR result is directly sent to the offline NLU for processing, and the obtained online ASR result is directly sent to the online NLU for processing. If the offline ASR and offline NLU are completed first, the offline result is sent to the NLU queue and waits for the online ASR and online NLU results. Conversely, if the online ASR and online NLU are completed first, the online result is sent to the NLU queue and waits for the offline ASR and offline NLU results. This process does not require the fusion of offline ASR and online ASR compared with the related technology.

[0045] Step 104: Perform arbitration processing on the natural language understanding result queue and output the result of the arbitration processing.

[0046] In some embodiments, send the result of the arbitration processing to the Dialog Manager (hereinafter referred to as DM) so that the Dialog Manager can execute.

[0047] In some embodiments, the natural language understanding result queue is a queue waiting for NLU arbitration. NLU arbitration refers to selecting an available NLU result based on the confidence level of the NLU result between offline NLU and online NLU.

[0048] In some embodiments, select an available NLU result from the offline NLU result and the online NLU result through NLU arbitration and send it to the DM for execution. Compared with the related art, it meets the requirement that a separate offline or separate online result is available. The specific NLU arbitration logic is not limited here.

[0049] In summary, according to the embodiments of the present disclosure, obtain the voice to be processed. In response to the voice to be processed originating from the wake-up sound area, perform online speech recognition and offline speech recognition on the voice to be processed respectively to obtain a first online speech recognition result and a first offline speech recognition result; perform online natural language understanding on the first online speech recognition result to obtain a first online natural language understanding result, and perform offline natural language understanding on the first offline speech recognition result to obtain a first offline natural language understanding result; send the first online natural language understanding result and the first offline natural language understanding result to the natural language understanding result queue; perform arbitration processing on the natural language understanding result queue and output the result of the arbitration processing. The method of the present disclosure can realize the parallelism of offline ASR and online ASR for the voice processing in the wake-up sound area, avoid the situation of wasting time caused by the fusion of offline ASR and online ASR, and can determine that a separate offline or separate online result is available through arbitration processing.

[0050] Based on Figure 1 the embodiments shown, Figure 3 is a schematic flowchart of a voice processing method provided by an embodiment of the present disclosure. This method is applied to a voice control system, and the execution entity is a voice controller, such as an in-vehicle voice chip.

[0051] The method includes the following steps 201-step 206.

[0052] Step 201: Obtain the voice to be processed. In response to the voice to be processed originating from the wake-up sound area, perform online speech recognition and offline speech recognition on the voice to be processed respectively to obtain a first online speech recognition result and a first offline speech recognition result.

[0053] Step 202: Perform online natural language understanding on the first online speech recognition result to obtain a first online natural language understanding result, and perform offline natural language understanding on the first offline speech recognition result to obtain a first offline natural language understanding result.

[0054] Step 203: Send the first online natural language understanding result and the first offline natural language understanding result to the natural language understanding result queue.

[0055] For the specific explanations of the above embodiments, refer to Figure 1 the content of Steps 101 - 103 in the illustrated embodiments, which will not be elaborated here.

[0056] In some embodiments of the present disclosure, after obtaining the speech to be processed, as Figure 4 shown, the method includes the following Steps 301 - 303.

[0057] In some embodiments of the present disclosure, the ASR processing in the non-wake-up sound area is single-channel, only offline ASR or only online ASR is performed. In response to the speech to be processed originating from the non-wake-up sound area, select to perform speech recognition on the speech to be processed in a preset manner.

[0058] In some embodiments of the present disclosure, the preset manner includes online. Select to perform online speech recognition and online natural language offline on the speech to be processed, and execute Steps 301 - 303.

[0059] Step 301: In response to the speech to be processed originating from the non-wake-up sound area, select to perform speech recognition on the speech to be processed in a preset manner.

[0060] Step 302: Perform online natural language understanding on the second online speech recognition result to obtain a second online natural language understanding result.

[0061] Step 303: Send the second online natural language understanding result to the natural language understanding result queue.

[0062] In some embodiments of the present disclosure, for the non-wake-up sound area, if online ASR is performed, the online NLU result obtained by sending the online ASR result to the online NLU for processing is directly sent to the natural language understanding result queue for NLU arbitration processing.

[0063] In some embodiments, the preset manner further includes offline. In response to the speech to be processed originating from the non-wake-up sound area, select to perform offline speech recognition and offline natural language offline on the speech to be processed, as Figure 5 shown, and execute Steps 401 - 405.

[0064] Step 401: In response to the speech to be processed originating from the non-wake-up sound area, select to perform offline speech recognition on the speech to be processed to obtain a second offline speech recognition result.

[0065] Step 402: Perform offline natural language understanding on the second offline speech recognition result to obtain a second offline natural language understanding result.

[0066] Step 403: Determine whether the second offline natural language understanding result needs to supplement online information sources according to the classification of the second offline natural language understanding result.

[0067] In some embodiments of the present disclosure, since some information in the NLU result is information sources that must rely on online resources, it is necessary to determine whether the second offline natural language understanding result needs to supplement online information sources.

[0068] Step 404: In response to the second offline natural language understanding result needing to supplement online information sources, perform online natural language understanding on the second offline speech recognition result to obtain a second online natural language understanding result.

[0069] Step 405: Send the second online natural language understanding result to the natural language understanding result queue.

[0070] In some embodiments, for non-awakening sound regions, if offline ASR is performed, the offline ASR result is first sent to offline NLU for processing to obtain an offline NLU result. According to the result classification in NLU, it is determined whether the offline NLU result needs to supplement online information sources. If so, the corresponding ASR is sent to online NLU for processing, and the processing result output by online NLU is sent to the natural language understanding queue for NLU result arbitration.

[0071] In some embodiments, after determining whether the second offline natural language understanding result needs to supplement online information sources, the method further includes: in response to the second offline natural language understanding result not needing to supplement online information sources, sending the second offline natural language understanding result to the natural language understanding result queue.

[0072] In some embodiments, for non-awakening sound regions, if offline ASR is performed, the offline ASR result is first sent to offline NLU for processing to obtain an offline NLU result. According to the result classification in NLU, it is determined whether the offline NLU result needs to supplement online information sources. If not, the offline ASR result is sent to the natural language understanding queue for NLU result arbitration.

[0073] Step 204: Obtain the first natural language understanding result sent to the natural language understanding result queue.

[0074] In some embodiments of the present disclosure, such as Figure 6As shown, the method further includes: determining whether the NLU result is from the wake-up sound area and determining whether the NLU result is an offline result. If the NLU result is from the wake-up sound area and is an offline result, the NLU result is the first offline NLU result; if the NLU result is from the wake-up sound area and is an online result, the NLU result is the first online NLU result; if the NLU result is from a non-wake-up sound area and is an offline result, the NLU result is the second offline NLU result; if the NLU result is from the wake-up sound area and is an offline result, the NLU result is the first offline NLU result; if the NLU result is from a non-wake-up sound area and is an offline result, the NLU result is the second offline NLU result.

[0075] In some embodiments, if the natural language understanding result sent to the natural language understanding result queue for the first time is from a non-wake-up sound area, that is, the natural language understanding result is the second online natural language understanding result or the second offline natural language understanding result, directly output the natural language understanding result, and for the non-wake-up sound area, implement separate ASR result response and semantic processing for the non-wake-up sound area.

[0076] Step 205, if the natural language understanding result is the first online natural language understanding result, determine whether the first online natural language understanding result is valid.

[0077] Step 206, if the first online natural language understanding result is valid, output the first online natural language understanding result.

[0078] In some embodiments of the present disclosure, outputting the first online natural language understanding result includes: sending the first online natural language understanding result to the DM module so that the DM module can execute.

[0079] In some embodiments of the present disclosure, as Figure 6 shown, determine whether the NLU result is from offline NLU. If the NLU result is not from offline NLU, that is, the NLU result is the first online NLU result, determine whether the NLU result is valid. If the NLU result is valid, directly send it to the DM for execution.

[0080] In some embodiments, after determining whether the first online natural language understanding result is valid, as Figure 7 shown, the method further includes: steps 501-503.

[0081] Step 501, if the first online natural language understanding result is invalid, wait for the first offline natural language understanding result to be sent to the natural language understanding result queue.

[0082] Step 502, after a preset waiting time, based on the natural language understanding result queue, determine whether the first offline natural language understanding result is valid.

[0083] Step 503, if the first offline natural language understanding result is valid, output the first offline natural language understanding result.

[0084] In some embodiments, it further includes: after a preset waiting time, if the first offline natural language understanding result has not been sent to the natural language understanding result queue, report an error.

[0085] In some embodiments of the present disclosure, as Figure 6 shown, determine whether the NLU result is available. If the NLU result is invalid, wait for the offline NLU result in the wake-up sound area. For example, if the preset waiting time is 6s, wait for the offline NLU result for 6s. If the offline NLU result is not received within 6s, report an error.

[0086] In some embodiments of the present disclosure, outputting the first offline natural language understanding result includes: sending the first offline natural language understanding result to the DM module for the DM module to execute.

[0087] In some embodiments of the present disclosure, as Figure 6 shown, determine whether the NLU result is available. If the NLU result is invalid, wait for the offline NLU result in the wake-up sound area. For example, if the preset waiting time is 6s, wait for the offline NLU result for 6s. If the offline NLU result is received within 6s, determine whether the offline NLU result is valid. If the offline NLU result is valid after arbitration, send it to the DM for execution.

[0088] In some embodiments, after obtaining the first natural language understanding result sent to the natural language understanding result queue, as Figure 8 shown, the method further includes: steps 601 - 603.

[0089] Step 601, if the natural language understanding result is the first offline natural language understanding result, determine whether the first offline natural language understanding result requires an online information source.

[0090] In some embodiments, after determining whether the first offline natural language understanding result requires an online information source, the method further includes: if the first offline natural language understanding result does not require an online information source, send the first offline natural language understanding result to the dialogue management module for execution.

[0091] In some embodiments, as Figure 6 shown, for the wake-up sound area, if the NLU result is the offline NLU processing result, i.e., the first offline NLU result, determine whether the NLU result requires an online information source. If not, directly send the NLU result to the DM for execution.

[0092] Step 602, if the first offline natural language understanding result requires an online information source, wait for the first online natural language understanding result to be sent to the natural language understanding result queue.

[0093] In some embodiments, as Figure 6 shown, for the wake-up sound area, if the NLU result is an offline NLU processing result, i.e., the first offline NLU result, determine whether the NLU result requires an online information source. If so, wait for the online NLU result of the wake-up sound area.

[0094] Step 603, after a preset waiting time, based on the natural language understanding result queue, determine whether the first online natural language understanding result is valid.

[0095] Step 604, if the first online natural language understanding result is valid, output the first online natural language understanding result.

[0096] In some embodiments, as Figure 6 shown, the preset waiting time is 6s. Wait for the online NLU result for 6s. If the online NLU result is obtained within 6s and the online NLU result is valid after arbitration in steps 501 - 503, send it to DM for execution.

[0097] In some embodiments, after waiting for the first online natural language understanding result to be sent to the natural language understanding result queue, the method further includes: after a preset waiting time, if the first online natural language understanding result has not been sent to the natural language understanding result queue, report an error.

[0098] In some embodiments, as Figure 6 shown, the preset waiting time is 6s. Wait for the online NLU result for 6s. If the online NLU result is not obtained within 6s, report an error.

[0099] In summary, the embodiments of the present disclosure further disclose the arbitration logic for NLU results in the wake-up sound area and non-wake-up sound area. For the wake-up sound area, it solves the problem of wasting time caused by the fusion of offline ASR and online ASR, satisfies the availability of separate offline or separate online results, and maximizes and fastest satisfies the semantics that can be processed fastest under the condition that NLU can only be serial, meeting the user's instructions to the greatest extent and fastest.

[0100] Corresponding to the above voice processing method, the present disclosure also proposes a voice processing device. Figure 9 It is a schematic structural diagram of a voice processing device 700 provided by an embodiment of the present disclosure. As Figure 9As shown in the figure, it includes: a speech recognition unit 710, which is used to obtain the speech to be processed, and in response to the speech to be processed originating from the wake-up sound area, perform online speech recognition and offline speech recognition on the speech to be processed respectively to obtain a first online speech recognition result and a first offline speech recognition result; a natural language understanding unit 720, which is used to perform online natural language understanding on the first online speech recognition result to obtain a first online natural language understanding result, and perform offline natural language understanding on the first offline speech recognition result to obtain a first offline natural language understanding result; a transmission unit 730, which is used to send the first online natural language understanding result and the first offline natural language understanding result to the natural language understanding result queue; an arbitration unit 740, which is used to perform arbitration processing on the natural language understanding results based on the natural language understanding result queue and output the result of the arbitration processing.

[0101] In some embodiments, the arbitration unit 740 is specifically configured to: obtain the first natural language understanding result sent to the natural language understanding result queue; if the natural language understanding result is the first online natural language understanding result, determine whether the first online natural language understanding result is valid; if the first online natural language understanding result is valid, output the first online natural language result.

[0102] In some embodiments, after determining whether the first online natural language understanding result is valid, the arbitration unit 740 is specifically configured to: if the first online natural language understanding result is invalid, wait for the first offline natural language understanding result to be sent to the natural language understanding result queue; after a preset waiting time, based on the natural language understanding result queue, determine whether the first offline natural language understanding result is valid; if the first offline natural language understanding result is valid, output the first offline natural language understanding result.

[0103] In some embodiments, after obtaining the first natural language understanding result sent to the natural language understanding result queue, the arbitration unit 740 is further configured to: if the natural language understanding result is the first offline natural language understanding result, determine whether the first offline natural language understanding result requires an online information source; if the first offline natural language understanding result requires an online information source, wait for the first online natural language understanding result to be sent to the natural language understanding result queue; after a preset waiting time, based on the natural language understanding result queue, determine whether the first online natural language understanding result is valid; if the first online natural language understanding result is valid, output the first online natural language understanding result.

[0104] In some embodiments, after obtaining the speech to be processed, the speech recognition unit 710 is configured to: in response to the speech to be processed originating from a non-wake-up sound area, select to perform speech recognition and natural language understanding on the speech to be processed in a preset manner, and the preset manner includes online or offline.

[0105] In some embodiments, the speech recognition unit 710 is specifically configured to: select to perform online speech recognition on the speech to be processed to obtain a second online speech recognition result; the natural language understanding unit 720 is specifically configured to: if the second online speech recognition result is obtained by performing online speech recognition on the speech to be processed, perform online natural language understanding on the second online speech recognition result to obtain a second online natural language understanding result; the transmission unit 730 is specifically configured to: send the second online natural language understanding result to the natural language understanding result queue.

[0106] In some embodiments, the speech recognition unit 710 is specifically configured to: select to perform offline speech recognition on the speech to be processed to obtain a second offline speech recognition result; the natural language understanding unit 720 is specifically configured to: perform offline natural language understanding on the second offline speech recognition result to obtain a second offline natural language understanding result; determine whether the second offline natural language understanding result needs to supplement an online information source according to the classification of the second offline natural language understanding result; the speech recognition unit 710 is further configured to: in response to the second offline natural language understanding result needing to supplement an online information source, perform online natural language understanding on the second offline speech recognition result to obtain a second online natural language understanding result; the transmission unit 730 is specifically configured to: send the second online natural language understanding result to the natural language understanding result queue.

[0107] In some embodiments, after determining whether the second offline natural language understanding result needs to supplement an online information source, the transmission unit 730 is further configured to: in response to the second offline natural language understanding result not needing to supplement an online information source, directly output the second offline natural language understanding result.

[0108] In some embodiments, the arbitration unit 740 is specifically configured to: obtain the first natural language understanding result sent to the natural language understanding result queue; if the natural language understanding result is the second online natural language understanding result or the second offline natural language understanding result, directly output the natural language understanding result.

[0109] In summary, according to the embodiments of the present disclosure, the device realizes parallel offline ASR and online ASR for the wake-up sound area throughout the vehicle and at all times, responds to the separate ASR results when not waking up, and the fastest processed semantics under the condition that NLU can only be in series, maximizing and fastest satisfying the user's voice commands through the speech recognition unit, natural language understanding unit, transmission unit, and arbitration unit.

[0110] It should be noted that since the device embodiments of the present disclosure correspond to the above method embodiments, the foregoing explanations of the method embodiments also apply to the devices of this embodiment. The principles are the same. For the details not disclosed in the device embodiments, reference may be made to the above method embodiments, and no further elaboration will be provided in the present disclosure.

[0111] According to an embodiment of the present disclosure, the present disclosure also provides an electronic device, a readable storage medium, and a computer program product.

[0112] Figure 10 FIG. shows a schematic block diagram of an exemplary electronic device 800 that can be used to implement embodiments of the present disclosure. The electronic device is intended to represent various forms of digital computers, such as, for example, a laptop computer, a desktop computer, a workbench, a personal digital assistant, a server, a blade server, a mainframe computer, and other suitable computers. The electronic device can also represent various forms of mobile devices, such as, for example, a personal digital processor, a cellular phone, a smart phone, a wearable device, and other similar computing devices. The components shown herein, their connections and relationships, and their functions are merely exemplary and are not intended to limit the implementation of the present disclosure described and / or claimed herein.

[0113] As Figure 10 shown, the device 800 includes a computing unit 801 that can perform various appropriate actions and processes according to a computer program stored in a ROM (Read-Only Memory) 802 or a computer program loaded from a storage unit 808 into a RAM (Random Access Memory) 803. In the RAM 803, various programs and data required for the operation of the device 800 can also be stored. The computing unit 801, the ROM 802, and the RAM 803 are connected to each other via a bus 804. An I / O (Input / Output) interface 805 is also connected to the bus 804.

[0114] Multiple components in the device 800 are connected to the I / O interface 805, including: an input unit 806, such as a keyboard, a mouse, etc.; an output unit 807, such as various types of displays, speakers, etc.; a storage unit 808, such as a magnetic disk, an optical disk, etc.; and a voice processing unit 809, such as a network card, a modem, a wireless voice processing transceiver, etc. The voice processing unit 809 allows the device 800 to exchange information / data with other devices via a computer network such as the Internet and / or various telecommunication networks.

[0115] The computing unit 801 can be various general-purpose and / or special-purpose processing components with processing and computing capabilities. Some examples of the computing unit 801 include, but are not limited to, a CPU (Central Processing Unit), a GPU (Graphic Processing Units), various dedicated AI (Artificial Intelligence) computing chips, various computing units running machine learning model algorithms, a DSP (Digital Signal Processor), and any suitable processor, controller, microcontroller, etc. The computing unit 801 executes the various methods and processes described above, such as the voice processing method. For example, in some embodiments, the voice processing method can be implemented as a computer software program, which is tangibly included in a machine-readable medium, such as the storage unit 808. In some embodiments, part or all of the computer program can be loaded and / or installed onto the device 800 via the ROM 802 and / or the voice processing unit 809. When the computer program is loaded into the RAM 803 and executed by the computing unit 801, one or more steps of the methods described above can be executed. Alternatively, in other embodiments, the computing unit 801 can be configured to execute the aforementioned voice processing method in any other suitable manner (e.g., by means of firmware).

[0116] The various embodiments of the systems and techniques described above in this document can be implemented in digital electronic circuit systems, integrated circuit systems, FPGAs (Field Programmable Gate Arrays), ASICs (Application-Specific Integrated Circuits), ASSPs (Application Specific Standard Products), SOCs (System On Chip), CPLDs (Complex Programmable Logic Devices), computer hardware, firmware, software, and / or combinations thereof. These various embodiments can include: being implemented in one or more computer programs that can be executed and / or interpreted on a programmable system including at least one programmable processor, which can be a dedicated or general-purpose programmable processor, receiving data and instructions from a storage system, at least one input device, and at least one output device, and transmitting the data and instructions to the storage system, the at least one input device, and the at least one output device.

[0117] The program code for implementing the methods of the present disclosure may be written in any combination of one or more programming languages. These program codes may be provided to a processor or controller of a general-purpose computer, a special-purpose computer, or other programmable data processing device, such that when the program codes are executed by the processor or controller, the functions / operations specified in the flowchart and / or block diagram are implemented. The program codes may be executed entirely on the machine, partially on the machine, executed partially on the machine and partially on a remote machine as an independent software package, or executed entirely on a remote machine or server.

[0118] In the context of the present disclosure, a machine-readable medium may be a tangible medium that can contain or store a program for use by or in connection with an instruction execution system, apparatus, or device. A machine-readable medium may be a machine-readable signal medium or a machine-readable storage medium. A machine-readable medium may include, but is not limited to, electronic, magnetic, optical, electromagnetic, infrared, or semiconductor systems, apparatus, or devices, or any suitable combination of the foregoing. More specific examples of a machine-readable storage medium would include an electrical connection based on one or more wires, a portable computer disk, a hard disk, a RAM, a ROM, an EPROM (Electrically Programmable Read-Only-Memory), or a flash memory, an optical fiber, a CD-ROM (Compact Disc Read-Only Memory), an optical storage device, a magnetic storage device, or any suitable combination of the foregoing.

[0119] In order to provide interaction with a user, the systems and techniques described herein may be implemented on a computer having: a display device for displaying information to the user (e.g., a CRT (Cathode-Ray Tube) or an LCD (Liquid Crystal Display) monitor); and a keyboard and a pointing device (e.g., a mouse or a trackball) through which the user can provide input to the computer. Other kinds of devices may also be used to provide interaction with the user; for example, the feedback provided to the user may be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and the input from the user may be received in any form (including acoustic input, voice input, or tactile input).

[0120] The systems and techniques described herein can be implemented in a computing system including backend components (e.g., as a data server), or a computing system including middleware components (e.g., an application server), or a computing system including frontend components (e.g., a user computer having a graphical user interface or a web browser through which a user can interact with an implementation of the systems and techniques described herein), or a computing system including any combination of such backend components, middleware components, or frontend components. The components of the system can be interconnected by digital data voice processing in any form or medium (e.g., a voice processing network). Examples of voice processing networks include: LAN (Local Area Network), WAN (Wide Area Network), the Internet, and blockchain networks.

[0121] A computer system can include a client and a server. The client and the server are generally remote from each other and typically interact through a voice processing network. The client-server relationship is created by computer programs running on respective computers and having a client-server relationship with each other. The server can be a cloud server, also known as a cloud computing server or a cloud host, which is a host product in the cloud computing service system, addressing the deficiencies of traditional physical hosts and VPS services ("Virtual Private Server", or simply "VPS") in terms of high management difficulty and weak business scalability. The server can also be a server of a distributed system, or a server combined with blockchain.

[0122] It should be noted that artificial intelligence is a discipline that studies how to make a computer simulate certain human thought processes and intelligent behaviors (such as learning, reasoning, thinking, planning, etc.), and it has both hardware-level and software-level technologies. Artificial intelligence hardware technologies generally include technologies such as sensors, dedicated artificial intelligence chips, cloud computing, distributed storage, and big data processing; artificial intelligence software technologies mainly include several major directions such as computer vision technology, speech recognition technology, natural language processing technology, and machine learning / deep learning, big data processing technology, and knowledge graph technology.

[0123] It should be understood that the various forms of processes shown above can be used, with steps reordered, added or deleted. For example, the steps described in this disclosure can be executed in parallel, sequentially or in different orders, as long as the desired results of the technical solutions disclosed in this disclosure can be achieved, and this is not limited herein. The above specific embodiments do not constitute a limitation on the protection scope of this disclosure. Those skilled in the art should understand that various modifications, combinations, sub-combinations and substitutions can be made according to design requirements and other factors. Any modifications, equivalent substitutions and improvements made within the spirit and principles of this disclosure shall be included within the protection scope of this disclosure.

Claims

1. A voice processing method, characterized in that, The method includes: Obtaining the speech to be processed, and in response to the speech to be processed originating from the wake-up sound area, performing online speech recognition and offline speech recognition on the speech to be processed respectively to obtain a first online speech recognition result and a first offline speech recognition result; Performing online natural language understanding on the first online speech recognition result to obtain a first online natural language understanding result, and performing offline natural language understanding on the first offline speech recognition result to obtain a first offline natural language understanding result; Sending the first online natural language understanding result and the first offline natural language understanding result to the natural language understanding result queue; Performing arbitration processing on the natural language understanding results based on the natural language understanding result queue, and outputting the result of the arbitration processing.

2. The method according to claim 1, characterized in that, The performing arbitration processing on the natural language understanding results based on the natural language understanding result queue and outputting the result of the arbitration processing includes: Obtaining the first natural language understanding result sent to the natural language understanding result queue; If the natural language understanding result is the first online natural language understanding result, determining whether the first online natural language understanding result is valid; If the first online natural language understanding result is valid, outputting the first online natural language result.

3. The method according to claim 2, characterized in that, After determining whether the first online natural language understanding result is valid, the method further includes: If the first online natural language understanding result is invalid, waiting for the first offline natural language understanding result to be sent to the natural language understanding result queue; After a preset waiting time, based on the natural language understanding result queue, determining whether the first offline natural language understanding result is valid; If the first offline natural language understanding result is valid, outputting the first offline natural language understanding result.

4. The method according to claim 2, characterized in that, After obtaining the first natural language understanding result sent to the natural language understanding result queue, the method further includes: If the natural language understanding result is the first offline natural language understanding result, determining whether the first offline natural language understanding result requires an online information source; If the first offline natural language understanding result requires an online information source, waiting for the first online natural language understanding result to be sent to the natural language understanding result queue; After a preset waiting time, based on the natural language understanding result queue, determining whether the first online natural language understanding result is valid; If the first online natural language understanding result is valid, outputting the first online natural language understanding result.

5. The method according to claim 1, characterized in that, After obtaining the speech to be processed, the method further includes: In response to the speech to be processed originating from a non-wake-up sound area, selecting to perform speech recognition and natural language understanding on the speech to be processed in a preset manner, where the preset manner includes online or offline.

6. The method according to claim 5, characterized in that, The selecting to perform speech recognition and natural language understanding on the speech to be processed in a preset manner in response to the speech to be processed originating from a non-wake-up sound area includes: Selecting to perform online speech recognition on the speech to be processed to obtain a second online speech recognition result; Performing online natural language understanding on the second online speech recognition result to obtain a second online natural language understanding result; Send the second online natural language understanding result to the natural language understanding result queue.

7. The method according to claim 5, characterized in that, The selecting to perform speech recognition and natural language understanding on the to-be-processed speech in a preset manner in response to the to-be-processed speech originating from a non-awakening sound area includes: Select to perform offline speech recognition on the to-be-processed speech to obtain a second offline speech recognition result; Perform offline natural language understanding on the second offline speech recognition result to obtain a second offline natural language understanding result; Judge whether the second offline natural language understanding result needs to supplement an online information source according to the classification of the second offline natural language understanding result; In response to the second offline natural language understanding result needing to supplement an online information source, perform online natural language understanding on the second offline speech recognition result to obtain a second online natural language understanding result; Send the second online natural language understanding result to the natural language understanding result queue.

8. The method according to claim 7, characterized in that, After judging whether the second offline natural language understanding result needs to supplement an online information source, the method further includes: In response to the second offline natural language understanding result not needing to supplement an online information source, send the second offline natural language understanding result to the natural language understanding result queue.

9. The method according to any one of claims 6 to 8, characterized in that The outputting the result of the arbitration process based on the natural language understanding result queue for the natural language understanding result includes: Obtain the first natural language understanding result sent to the natural language understanding result queue; If the natural language understanding result is the second online natural language understanding result or the second offline natural language understanding result, directly output the natural language understanding result.

10. A voice processing device, characterized in that The apparatus includes: A speech recognition unit, configured to obtain the to-be-processed speech, and in response to the to-be-processed speech originating from an awakening sound area, perform online speech recognition and offline speech recognition on the to-be-processed speech respectively to obtain a first online speech recognition result and a first offline speech recognition result; A natural language understanding unit, configured to perform online natural language understanding on the first online speech recognition result to obtain a first online natural language understanding result, and perform offline natural language understanding on the first offline speech recognition result to obtain a first offline natural language understanding result; A transmission unit, configured to send the first online natural language understanding result and the first offline natural language understanding result to the natural language understanding result queue; An arbitration unit, configured to perform an arbitration process on the natural language understanding result based on the natural language understanding result queue and output the result of the arbitration process.

11. An electronic device, comprising: At least one processor; And A memory communicatively connected to the at least one processor; wherein, The memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor so that the at least one processor can execute the method according to any one of claims 1-9.

12. A non-transitory computer-readable storage medium storing computer instructions, wherein The computer instructions are used to cause the computer to execute the method according to any one of claims 1-9.

13. A computer program product, comprising a computer program, wherein the computer program, when executed by a processor, implements the method according to any one of claims 1-9.

Citation Information

Patent Citations

  • Voice recognition result display method and device

    CN103634321A

  • Voice processing method, device and system

    CN114724564A

  • Voice interaction method, device and equipment and computer storage medium

    CN116416986A