Voice processing method and device

By receiving and sorting speech recognition results in different voice regions in the scenario of multi-channel ASR engine operation, and directly performing parallel processing of multiple-channel ASR engines, the problems of underutilization of resources and audio processing delays are solved, and efficient voice processing and excellent user experience are achieved.

CN120164467APending Publication Date: 2025-06-17BEIJING CO WHEELS TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202311719564.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2023-12-14
Publication Date
2025-06-17

AI Technical Summary

Technical Problem

In the scenario where multiple ASR engines work and can only be used by single NLU engines, the existing technology causes the audio queue to need serial processing, resources are not fully utilized, memory consumption is high, and audio streaming processes occupy the entire engine, resulting in other sound areas being unable to respond in time.

Method used

By receiving voice information from N different voice regions, voice recognition is performed separately, stored in a subqueue in the voice recognition queue, and sorting and processing from small to large in the first time of the subqueue, the audio is directly sent to the multi-channel ASR engine for parallel processing.

Benefits of technology

It realizes the ASR engine that directly parallel processing without cached audio is achieved, making full use of resources, shortening processing time, reducing memory consumption, and prioritizing first-arrival voice information to improve user experience.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120164467A_ABST
    Figure CN120164467A_ABST
Patent Text Reader

Abstract

The invention provides a voice processing method and device, and relates to the technical field of data processing. The method comprises the following steps: receiving voice information of N different voice areas, wherein N is a positive integer greater than or equal to 2; the voice information is recognized through voice recognition engines corresponding to the N different voice areas, and a plurality of voice recognition results corresponding to the voice information of the N different voice areas are obtained; storing a plurality of voice recognition results corresponding to the voice information of the N different voice areas into sub-queues corresponding to the N different voice areas in a voice recognition queue; and sorting the sub-queues according to the sequence of the first time of the sub-queues from small to large, and taking out the voice recognition result stored in the head of the sub-queue with the first sorting for execution. The method is applied to a scene in which multiple paths of ASR engines work and only a single path of NLU engine can be used, and ASR results from different voice areas are stored and sorted to be taken out, so that the voice recognition efficiency is improved, and the voice recognition efficiency is improved. And the verbal skill of the user can be responded accurately as soon as possible.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present disclosure relates to the field of data processing, and in particular, to a voice processing method and apparatus. Background Art

[0002] Multi-channel parallel ASR (Automatic Speech Recognition) means that multiple microphones simultaneously collect voices of different people in different sound zones, and corresponding ASR results are output by respective ASR engines at the same time, and the multi-channel ASR engines do not affect each other. Single-channel NLU (Natural Language Understanding) means that since the NLU consumes too much CPU resources, only a single NLU engine can be processing at the same time. Even if there are multiple ASR results, they can only be processed serially one by one.

[0003] In the scenario where multi-channel ASR engines are working and only a single-channel NLU engine is used, related technologies make the audio queue serial, sort and cache the audio, and only send one-way audio data to the ASR engine each time, so as to ensure that only one-way audio is sent to the NLU engine each time. Then, after the NLU result is executed on the processing side, the audio queue is opened for the next ASR recognition. In related technologies, the audio processing takes time to become ASR, and caching the audio consumes more time than directly executing the processing. The resources consumed by the ASR engine are controllable and not fully utilized. The audio data is too large, and caching too much audio data causes excessive memory consumption. Moreover, in related technologies, the audio streaming processing occupies the entire engine, resulting in that even if other sound zones complete the speech first, they still cannot respond in time. Summary of the Invention

[0004] The present disclosure provides a voice processing method, apparatus, electronic device, storage medium, program product, vehicle networking module, and vehicle.

[0005] According to a first aspect of the present disclosure, there is provided a voice processing method, the method including: receiving voice information of N different sound zones, where N is a positive integer greater than or equal to 2; respectively identifying the voice information through voice recognition engines corresponding to the N different sound zones to obtain a plurality of voice recognition results corresponding to the voice information of the N different sound zones; storing the plurality of voice recognition results corresponding to the voice information of the N different sound zones into sub-queues corresponding to the N different sound zones in a voice recognition queue, where the N different sound zones respectively correspond to different sub-queues, and each sub-queue includes a head and a tail, and the head and the tail of the sub-queue are respectively used to store a voice recognition result; sorting the sub-queues in ascending order of a first time of the sub-queues, and taking out the voice recognition result stored in the head of the sub-queue sorted first for execution, where the first time is the start time of the voice information corresponding to the voice recognition result stored in the head of the sub-queue.

[0006] In some embodiments, before sorting the sub - queues in ascending order of the first time of the sub - queues, the method further includes: performing voice activity detection on the voice information of N different pitch regions to obtain detection results corresponding to the voice information of the N different pitch regions, where the detection results at least include the start time of the voice information; storing multiple speech recognition results corresponding to the voice information of the N different pitch regions into the sub - queues corresponding to the N different pitch regions in the speech recognition queue, including: after binding the multiple speech recognition results corresponding to the voice information of the N different pitch regions and the corresponding detection results, storing them into the sub - queues corresponding to the N different pitch regions in the speech recognition queue.

[0007] In some embodiments, storing multiple speech recognition results corresponding to the voice information of N different pitch regions into the sub - queues corresponding to the speech recognition queue includes: for any one of the N different pitch regions, in response to receiving the first speech recognition result of the pitch region, creating a first sub - queue in the speech recognition queue and storing the first speech recognition result at the head of the first sub - queue, where the first speech recognition result is the first speech recognition result corresponding to the voice information of the pitch region, the first sub - queue includes a head and a tail, and the speech recognition queue is initially an empty queue.

[0008] In some embodiments, after creating a first sub - queue in the speech recognition queue in response to receiving the first speech recognition result of the pitch region and storing the first speech recognition result at the head of the first sub - queue, the method further includes: in response to receiving the second speech recognition result of the pitch region, storing the second speech recognition result at the tail of the first sub - queue, where the second speech recognition result is the second speech recognition result corresponding to the voice information of the pitch region; in response to receiving the third speech recognition result of the pitch region, deleting the second speech recognition result from the first sub - queue and storing the third speech recognition result at the tail of the first sub - queue, where the third speech recognition result is the third speech recognition result corresponding to the voice information of the pitch region.

[0009] In some embodiments, sorting the sub - queues in ascending order of the first time of the sub - queues and taking out and executing the speech recognition result stored at the head of the sub - queue with the first sorting includes: in response to the first time of the first sub - queue being the smallest, taking out and executing the first speech recognition result of the first sub - queue; moving the third speech recognition result of the first pitch region from the tail of the first sub - queue to the head of the first sub - queue.

[0010] In some embodiments, after retrieving and executing the first speech recognition result of the first sub-queue, the method further includes: in response to the completion of the execution of the first speech recognition result of the first sub-queue, obtaining the start time of the speech information corresponding to the speech recognition result stored at the tail of the first sub-queue; comparing the start time with the first time of the sub-queues corresponding to multiple sound zones, obtaining a comparison result, where the multiple sound zones include the other sound zones except the sound zone corresponding to the first sub-queue among the N sound zones; and re-sorting the first sub-queue and the sub-queues corresponding to the multiple sound zones according to the comparison result.

[0011] In some embodiments, after retrieving and executing the third speech recognition result of the first sub-queue, the method further includes: in response to the first speech recognition result and the third speech recognition result of the first sub-queue being retrieved and no other speech recognition results from the sound zone being received, deleting the first sub-queue from the speech recognition queue.

[0012] According to a second aspect of the present disclosure, there is provided a speech processing apparatus, the apparatus includes: a receiving unit, configured to receive speech information of N different sound zones, where N is a positive integer greater than or equal to 2; a speech recognition unit, configured to respectively recognize the speech information through speech recognition engines corresponding to the N different sound zones, obtaining multiple speech recognition results corresponding to the speech information of the N different sound zones; a storage unit, configured to store the multiple speech recognition results corresponding to the speech information of the N different sound zones into sub-queues corresponding to the N different sound zones in a speech recognition queue, where the N different sound zones respectively correspond to different sub-queues, a sub-queue includes a head and a tail, and the head and the tail of the sub-queue are respectively configured to store a speech recognition result; a sorting and retrieving unit, configured to sort the sub-queues in ascending order of the first time of the sub-queues, and retrieve and execute the speech recognition result stored at the head of the sub-queue with the first sorting, where the first time is the start time of the speech information corresponding to the speech recognition result stored at the head of the sub-queue.

[0013] According to a third aspect of the present disclosure, there is provided an electronic device, including: at least one processor; and a memory speech-processing connected to the at least one processor; where the memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor to enable the at least one processor to execute the method of the foregoing first aspect.

[0014] According to a fourth aspect of the present disclosure, there is provided a non-transitory computer-readable storage medium storing computer instructions, where the computer instructions are used to cause a computer to execute the method of the foregoing first aspect.

[0015] According to a fifth aspect of the present disclosure, there is provided a computer program product, including a computer program, where the computer program implements the method of the foregoing first aspect when executed by a processor.

[0016] The voice processing method provided by the embodiments of the present disclosure includes receiving voice information of N different voice regions, where N is a positive integer greater than or equal to 2; respectively identifying the voice information through voice recognition engines corresponding to the N different voice regions to obtain a plurality of voice recognition results corresponding to the voice information of the N different voice regions; storing the plurality of voice recognition results corresponding to the voice information of the N different voice regions into sub-queues corresponding to the N different voice regions in a voice recognition queue, where the N different voice regions respectively correspond to different sub-queues, and each sub-queue includes a head and a tail, and the head and the tail of the sub-queue are respectively used to store a voice recognition result; sorting the sub-queues in ascending order of the first time of the sub-queues, and taking out the voice recognition result stored in the head of the sub-queue with the first sorting to execute, where the first time is the start time of the voice information corresponding to the voice recognition result stored in the head of the sub-queue. The method of the present disclosure is applied to the scenario where multiple ASR engines work and only a single NLU engine can be used. Without caching the audio, the audio is directly sent to multiple ASR engines for parallel processing, making full use of the controllable resources consumed by the ASR engines, shortening the processing time, reducing the memory consumption, storing the ASR results from different voice regions by voice region, and sorting and taking out for processing according to the start time of the user's voice information, so that when different users in multiple voice regions issue instructions simultaneously, the user who finishes speaking the instruction first is preferentially processed, improving the user experience.

[0017] It should be understood that the content described in this part is not intended to identify the key or important features of the embodiments of the present application, nor is it used to limit the scope of the present application. Other features of the present application will become easily understandable through the following description. BRIEF DESCRIPTION OF THE DRAWINGS

[0018] The drawings are used to better understand the solution and do not constitute a limitation to the present disclosure. Among them:

[0019] Figure 1 is a flowchart of a voice processing method provided by an embodiment of the present disclosure;

[0020] Figure 2 is a flowchart of a voice processing method provided by an embodiment of the present disclosure;

[0021] Figure 3 is an example diagram of a voice processing method provided by an embodiment of the present disclosure;

[0022] Figure 4 is an example diagram of a voice processing method provided by an embodiment of the present disclosure;

[0023] Figure 5 is an example diagram of a voice processing method provided by an embodiment of the present disclosure;

[0024] Figure 6 Schematic diagram of a voice processing method provided by an embodiment of the present disclosure;

[0025] Figure 7 Schematic structural diagram of a voice processing device provided by an embodiment of the present disclosure;

[0026] Figure 8 Schematic block diagram of an exemplary electronic device 800 provided by an embodiment of the present disclosure. Detailed implementation manners

[0027] The following describes exemplary embodiments of the present disclosure with reference to the accompanying drawings. Various details of the embodiments of the present disclosure are included to facilitate understanding, and they should be considered merely exemplary. Therefore, those of ordinary skill in the art should recognize that various changes and modifications can be made to the embodiments described herein without departing from the scope and spirit of the present disclosure. Similarly, for clarity and conciseness, descriptions of well-known functions and structures are omitted below.

[0028] The following describes in detail a voice processing method, device, electronic device, storage medium, and program product proposed by the present disclosure with reference to the accompanying drawings.

[0029] Figure 1 Flowchart of a voice processing method provided by an embodiment of the present disclosure. The method is applied to a voice control system, and the execution entity is a voice controller, such as an in-vehicle voice chip.

[0030] As Figure 1 shown, the method includes the following steps:

[0031] Step 101: Receive voice information from N different sound zones, where N is a positive integer greater than or equal to 2.

[0032] In some embodiments of the present disclosure, it is assumed that there are multiple continuous or discontinuous sound zones in a physical space, and each sound zone is provided with a microphone for picking up voices in the corresponding sound zone. For example, in an in-vehicle scenario, the driver's area, passenger area, and rear row area in the vehicle space are set as different sound zones, and corresponding microphones are set in each sound zone for voice reception.

[0033] Step 102: Recognize the voice information through voice recognition engines corresponding to N different sound zones respectively to obtain multiple voice recognition results corresponding to the voice information of the N different sound zones.

[0034] In some embodiments of the present disclosure, the input voice is recognized by a voice recognition engine and converted into an executable instruction, that is, the voice recognition result is executable text information.

[0035] In some embodiments of the present disclosure, multi-channel parallel ASR is adopted to support receiving voices from N different voice regions simultaneously. The ASR recognition engines corresponding to the N different voice regions perform voice recognition simultaneously and output their respective ASR results, and the N different ASR recognition engines do not affect each other.

[0036] Step 103: Store the multiple voice recognition results corresponding to the voice information from N different voice regions into the sub-queues corresponding to the N different voice regions in the voice recognition queue.

[0037] Among them, the N different voice regions respectively correspond to different sub-queues. The sub-queue includes a head and a tail, and the head and the tail of the sub-queue are respectively used to store a voice recognition result.

[0038] In some embodiments, since multi-channel ASR is adopted while NLU can only process in a single channel, the present disclosure creates an ASR queue. The ASR queue is initially an empty queue. When the first ASR result is received, the ASR queue is triggered to create a sub-queue.

[0039] In some embodiments, the ASR results from different voice regions are stored in different sub-queues. For example, the ASR queue is initially an empty queue. When the first ASR result triggers, item A is created as sub-queue A, and queue A can store the first and last two ASRs. When ASR results from other voice regions enter the ASR queue, queue item B is created as sub-queue B.

[0040] In some embodiments, the head and the tail of the sub-queue can respectively store a voice recognition result, and at most two can be stored. The first voice recognition result corresponding to the voice region is put into the head, and the second voice recognition result is put into the tail. After that, every time a new voice recognition result is received, it replaces the voice recognition result at the end of the queue.

[0041] In some embodiments, after the voice recognition result at the head of the sub-queue is taken out for processing, the voice recognition result at the tail moves forward and is put into the head.

[0042] In some embodiments, when all the voice recognition results in the sub-queue are taken out for processing, the sub-queue is deleted from the voice recognition queue.

[0043] Step 104: Sort the sub-queues in ascending order of the first time of the sub-queue, and execute the voice recognition result stored in the head of the sub-queue with the first sorting.

[0044] Among them, the first time is the start time of the voice information corresponding to the voice recognition result stored in the head of the sub-queue.

[0045] In some embodiments, before arranging the sub - queues according to the first time of the sub - queues, the method further includes: performing voice activity detection on the voice information of N different pitch regions to obtain detection results corresponding to the voice information of N different pitch regions, where the voice activity detection results at least include the start time of the voice information.

[0046] In some embodiments, voice activity detection (VAD) is performed on the voices of different pitch regions to obtain VAD information of the voices. The VAD information includes the start point (vad begin) of the voice, the end point (vad end) of the voice, and the record number (recordId) of the current conversation.

[0047] In some embodiments, it further includes: binding a plurality of speech recognition results corresponding to the voice information of N different pitch regions and the corresponding detection results. Thus, the detection results corresponding to the voice information will be stored in the sub - queue of the speech recognition queue together with the corresponding speech recognition results.

[0048] In some embodiments, the sub - queues are arranged in ascending order of the first time of the sub - queues to form a new queue. The new queue includes the speech recognition results and the start time of the voice information corresponding to the speech recognition results. The speech recognition result stored at the head of the sub - queue ranked first is taken out from the queue and executed.

[0049] It can be understood that the start time of the voice information corresponding to the speech recognition result stored at the head of the sub - queue ranked first is the smallest. Thus, the voice information received earliest is preferentially processed.

[0050] In some embodiments, the ASR results of different pitch regions and the VAD information of the corresponding voices are stored together in the new queue. The stored content includes but is not limited to the ASR text, the start point (vad begin) of the voice, the end point (vad end) of the voice, and the record number (recordId) of the current conversation. The record number is used to identify that the voice corresponding to the current ASR text is a complete sentence to prevent errors.

[0051] In one embodiment, a new data queue is dynamically created according to the ASR data generated by multiple pitch regions, that is, the data of multiple sub - queues in the ASR queue. The storage method of the new data queue is the same as that of the sub - queues. Only the first and last two ASR data of each pitch region are stored. Each pitch region is sorted according to the vad begin of the first ASR data in the pitch region. The ASR data stored at the head of the pitch region ranked the most forward is taken out and executed. After the execution is completed, the ASR data stored at the head of the next pitch region ranked the most forward is taken out and executed.

[0052] In summary, according to the embodiments of the present disclosure, voice information in N different voice regions is received, where N is a positive integer greater than or equal to 2; the voice information is respectively recognized by voice recognition engines corresponding to the N different voice regions to obtain multiple voice recognition results corresponding to the voice information in the N different voice regions; the multiple voice recognition results corresponding to the voice information in the N different voice regions are stored in sub-queues corresponding to the N different voice regions in a voice recognition queue, where the N different voice regions respectively correspond to different sub-queues, and each sub-queue includes a head and a tail, and the head and the tail of the sub-queue are respectively used to store a voice recognition result; the sub-queues are sorted in descending order of the first time of the sub-queues, and the voice recognition result stored in the head of the sub-queue with the first sorting is taken for execution, where the first time is the start time of the voice information corresponding to the voice recognition result stored in the head of the sub-queue. The method of the present disclosure is applied to the scenario where multiple ASR engines work and only a single NLU engine can be used. Without caching the audio, the audio is directly sent to multiple ASR engines for parallel processing, making full use of the fact that the resources consumed by the ASR engines are controllable, shortening the processing time, reducing the memory consumption, storing the ASR results from different voice regions by voice region, and sorting and taking out the processing according to the start time of the user's voice information, so that when different users in multiple voice regions issue instructions at the same time, the user who finishes speaking the instruction first is preferentially processed, improving the user experience.

[0053] Based on Figure 1 the embodiments shown, Figure 2 FIG. is a schematic flowchart of a voice processing method provided by an embodiment of the present disclosure. This method is applied to a voice control system, and the execution entity is a voice controller, such as an in-vehicle voice chip.

[0054] The method includes the following steps 201-2010.

[0055] Step 201, receive voice information in N different voice regions, where N is a positive integer greater than or equal to 2.

[0056] In some embodiments of the present disclosure, it is set that there are multiple continuous or discontinuous voice regions in a physical space, and each voice region is provided with a microphone for picking up the voice of the corresponding voice region. For example, in an in-vehicle scenario, the driver's area, the co-driver's area, and the rear row area in the vehicle space are set as different voice regions, and each voice region is provided with a corresponding microphone for receiving voice.

[0057] Step 202, respectively recognize the voice information through voice recognition engines corresponding to the N different voice regions to obtain multiple voice recognition results corresponding to the voice information in the N different voice regions.

[0058] In some embodiments of the present disclosure, the input speech is subjected to speech recognition by a speech recognition engine to convert it into executable instructions, that is, the speech recognition result is executable text information.

[0059] In some embodiments of the present disclosure, multi-channel parallel ASR is adopted to support receiving voices in N different voice regions simultaneously. The ASR recognition engines corresponding to the N different voice regions perform speech recognition simultaneously and output their respective ASR results, and the N different ASR recognition engines do not affect each other.

[0060] In the embodiments of the present disclosure, storing multiple speech recognition results corresponding to the speech information in N different voice regions into the sub-queues corresponding to the N different voice regions in the speech recognition queue includes steps 203-step 20.

[0061] It should be noted that in the embodiments of the present disclosure, taking two different voice regions including the first voice region and the second voice region in the N different voice regions as an example to reflect the ASR queue creation and storage method and logic of the present disclosure solution. The N different voice regions also include three or more different voice regions. The first voice region and the second voice region are any two different voice regions in the N different voice regions, which does not limit the application scope of the present disclosure solution.

[0062] Step 203, for any one of the N different voice regions, in response to receiving the first speech recognition result of the voice region, create a first sub-queue in the speech recognition queue, and store the first speech recognition result at the head of the first sub-queue.

[0063] Wherein, the first speech recognition result is the first speech recognition result corresponding to the speech information of the voice region. The first sub-queue includes a head and a tail, and the speech recognition queue is initially an empty queue.

[0064] In one implementation manner of the present disclosure, as Figure 3 shown, the N different voice regions at least include the first voice region and the second voice region. Multi-channel parallel ASR is adopted. The speech received through the microphone mic1 enters the first voice region, and the speech received through the microphone mic2 enters the second voice region. The first voice region and the second voice region perform ASR simultaneously to obtain the ASR result.

[0065] In one implementation manner of the present disclosure, as Figure 3 shown, the first speech recognition result received by the ASR queue is mic1_asr1 from the first voice region. A sub-queue 1 is created in the ASR queue. The sub-queue 1 includes a head and a tail, and mic1_asr1 is stored at the head of the sub-queue 1.

[0066] Step 204, in response to receiving the second speech recognition result of the voice region, store the second speech recognition result at the tail of the first sub-queue.

[0067] Among them, the second speech recognition result is the second speech recognition result corresponding to the speech information of the voice zone.

[0068] Step 205, in response to receiving the third speech recognition result of the voice zone, delete the second speech recognition result from the first sub-queue, and store the third speech recognition result at the tail of the first sub-queue.

[0069] Among them, the third speech recognition result is the third speech recognition result corresponding to the speech information of the voice zone.

[0070] In an implementation manner of the present disclosure, as Figure 3 shown, in response to receiving mic1_asr2 from the first voice zone, store it at the tail of sub-queue 1. In response to receiving mic1_asr3 from the first voice zone, Figure 3 the mic1_asr2 in the dashed box in the figure indicates deleting mic1_asr2 from sub-queue 1 and storing mic1_asr3 at the tail of sub-queue 1.

[0071] For example, in a vehicle-mounted scenario, the driver continuously utters three voice commands: "Open the window", "Check the weather", and "Turn on the air conditioner". Cache the two commands "Open the window" and "Turn on the air conditioner", which can prevent excessive command blocking of the queue by retaining the start and end commands when consecutive command transmissions are blocked.

[0072] It should be noted that the ASRs of different voice zones are processed in parallel. During this period, there may also be two different voice zones, three different voice zones... N different voice zones with parallel ASR processing, which is not limited in the present disclosure.

[0073] In another implementation manner of the present disclosure, taking the scenario of three different voice zones as an example to illustrate the method and logic for creating and storing the ASR queue of the present solution. As Figure 5 shown, the three different voice zones include the first voice zone, the second voice zone, and the third voice zone, corresponding to Figure 5 the mic1, mic2, and mic3 shown respectively. Using a multi-channel parallel ASR engine, the three different voice zones simultaneously execute ASR to obtain ASR results.

[0074] For mic1, the first ASR result that reaches the ASR queue and originates from mic1 is mic1_asr1, the second is mic1_asr2, and the third is mic1_asr3. When mic1_asr1 arrives, create sub-queue 1 and store mic1_asr1 at the head of sub-queue 1; when mic1_asr2 arrives, store mic1_asr2 at the tail of sub-queue 1; when mic1_asr3 arrives, delete mic1_asr2 from sub-queue 1 and put mic1_asr3 at the tail of sub-queue 1.

[0075] For mic2, the first ASR result that reaches the ASR queue and originates from mic2 is mic2_asr1, and the second is mic2_asr2. When mic2_asr1 arrives, create sub-queue 2 and store mic2_asr1 at the head of sub-queue 2; when mic2_asr2 arrives, store mic2_asr2 at the tail of sub-queue 2.

[0076] For mic3, the first ASR result that reaches the ASR queue and originates from mic3 is mic3_asr1. When mic3_asr1 arrives, create sub-queue 3 and store mic3_asr1 at the head of sub-queue 3.

[0077] Step 206, in response to the first time of the first sub-queue being the smallest, take out and execute the first speech recognition result of the first sub-queue.

[0078] In an embodiment of the present disclosure, according to the first time of the first sub-queue and the first times of the sub-queues corresponding to multiple sound zones, sort the first sub-queue and other multiple sound zone sub-queues, where the multiple sound zones include other sound zones except the sound zone corresponding to the first sub-queue among the N sound zones. In response to the first time of the first sub-queue being less than or equal to the first time of the second sub-queue, arrange the first sub-queue before the second sub-queue, and take out and execute the first speech recognition result of the first sub-queue.

[0079] Step 207, move the third speech recognition result from the tail of the first sub-queue to the head of the first sub-queue.

[0080] In an implementation manner, taking the example that there are two sub-queues in the speech recognition queue, including the first sub-queue and the second sub-queue, the first time is the start time of the speech information corresponding to the speech recognition result stored at the head of the sub-queue. Thus, the first time of the first sub-queue is the start time of the speech information corresponding to the first speech recognition result of the first sound zone, and the first time of the second sub-queue is the start time of the speech information corresponding to the first speech recognition result of the second sound zone.

[0081] In some embodiments, the method further includes: performing voice activity detection on the speech information corresponding to the first speech recognition result of the first voice region, the third speech recognition result, and the first speech recognition result of the second voice region, to obtain corresponding voice activity detection results, where the voice activity detection results at least include the start time of the speech information.

[0082] In one implementation manner of the present disclosure, as Figure 4 shown, if the start (vadbegin) time of the speech of mic1_asr1 is the smallest, the start time of the speech information of mic1_asr3 is the second smallest, and the start time of the speech information of mic2_asr1 is the largest, then sub-queue 1 is arranged in front of sub-queue 2. First, take out mic1_asr1 in sub-queue 1 and send mic1_asr1 to the NLU engine for processing, and then hand it over to the corresponding other application for response. For example, if the instruction of mic1_asr1 is "open the window", it is handed over to the application responsible for controlling the window for response. After mic1_asr1 is taken out and executed, put mic1_asr3 in sub-queue 1 into the head.

[0083] In another embodiment, taking the example that the speech recognition queue has three sub-queues including the first sub-queue, the second sub-queue, and the third sub-queue, as Figure 6 shown, there are sub-queue 1, sub-queue 2, and sub-queue 3. Assume that the start time of the speech information corresponding to mic1_asr1 is the smallest, followed by mic3_asr1, and then mic2_asr1, mic2_asr2, and mic1_asr3 in sequence. Sort sub-queue 1, sub-queue 2, and sub-queue 3 according to the start time of the speech information for the ASR result at their heads. Since mic1_asr1 < mic3_asr1 < mic2_asr1, the sub-queue sorted at the front is sub-queue 1, followed by sub-queue 2, and the last one is sub-queue 3. First, take out mic1_asr1 in sub-queue 1 and execute it. After mic1_asr1 is executed, the ASR result stored at the head of sub-queue 1 is mic1_asr3.

[0084] Step 208, in response to the completion of the execution of the first speech recognition result of the first sub-queue, obtain the start time of the speech information corresponding to the speech recognition result stored at the tail of the first sub-queue.

[0085] Step 209, compare the start time with the first time of the sub-queues corresponding to multiple voice regions to obtain a comparison result, where the multiple voice regions include other voice regions except the voice region corresponding to the first sub-queue among the N voice regions.

[0086] Step 210, re-sort the first sub-queue and the sub-queues corresponding to the multiple voice regions according to the comparison result.

[0087] In one embodiment, taking the example that there are a first sub - queue and a second sub - queue in the speech recognition queue, as Figure 4 shown, there are sub - queue 1 and sub - queue 2. In response to mic1_asr1 in queue 1 having been taken out for execution, and mic1_asr3 being at the head of sub - queue 1, compare the start time of the speech information of mic1_asr3 with the start time of the speech information of mic1_asr3, and re - sort sub - queue 1 and sub - queue 2. Suppose the start (vad begin) time of the speech of mic1_asr1 is the smallest, the start time of the speech information of mic1_asr3 is the second, and the start time of the speech information of mic2_asr1 is the largest. Since the start time of the speech information of mic1_asr3 is earlier than the start time of the speech information of mic2_asr1, after re - arrangement, the sub - queue ranked at the front is sub - queue 1. Take out mic1_asr3 and send mic1_asr3 to the NLU engine for processing, and then hand it over to the corresponding other application for response.

[0088] In another embodiment, taking the example that there are a first sub - queue, a second sub - queue and a third sub - queue in the speech recognition queue, as Figure 6 shown, there are sub - queue 1, sub - queue 2 and sub - queue 3. In response to mic1_asr1 in queue 1 having been taken out for execution, and mic1_asr3 being at the head of sub - queue 1, compare the start time of the speech information of mic1_asr3 with the start time of the speech information of mic2_asr1 and the start time of the speech information of mic3_asr1. The comparison result is mic3_asr1 < mic2_asr1 < mic1_asr3. After re - arranging according to the comparison result, the sub - queue ranked at the front is sub - queue 3, followed by sub - queue 2, and the last - ranked is sub - queue 1. Take out mic3_asr1 in sub - queue 3 for execution.

[0089] Step 2011: In response to the third speech recognition result and the third speech recognition result of the first sub - queue being taken out and no other speech recognition results from the sound area corresponding to the first sub - queue being received, delete the first sub - queue from the speech recognition queue.

[0090] In one implementation manner of the present disclosure, as Figure 4 shown, after taking out and executing mic1_asr3, there is no other ASR data in sub - queue 1, and delete sub - queue 1 from the ASR queue.

[0091] In some embodiments of the present disclosure, it further includes: taking out and executing the first speech recognition result of the second sub - queue. In response to the first speech recognition result of the second sub - queue being taken out and no other speech recognition results from the second sound area being received, delete the second sub - queue from the speech recognition queue.

[0092] In one implementation of the present disclosure, as Figure 4 shown, it further includes taking out and executing mic2_asr1, deleting sub-queue 2 from the ASR queue, and waiting for a new ASR result to trigger the ASR queue.

[0093] Similarly, in another implementation of the present disclosure, as Figure 6 shown, after mic3_asr1 is executed and there are no other ASR results in sub-queue 3, sub-queue 3 is deleted from the ASR queue, and the start time of the voice information of mic2_asr1 is compared with the start time of the voice information of mic1_asr3, and re-sorted. Since mic2_asr1 < mic1_asr3, the sorting of sub-queue 2 is the most forward. Take out mic2_asr1 in sub-queue 2 and execute it. After mic2_asr1 is executed, the ASR result stored at the head of sub-queue 2 is mic2_asr2. Compare the start time of the voice information of mic2_asr2 with the start time of the voice information of mic1_asr3, and re-sort. Since mic2_asr2 < mic1_asr3, the sorting of sub-queue 2 is the most forward. Take out mic2_asr2 and execute it. There are no other ASR results in sub-queue 2, so sub-queue 2 is deleted from the ASR queue. Finally, take out mic1_asr3 in sub-queue 1 and execute it. There are no other ASR results in sub-queue 1, so sub-queue 1 is deleted from the ASR queue, and wait for a new ASR result to trigger the ASR queue.

[0094] In summary, according to the embodiments of the present disclosure, there is no need to cache audio, and the audio is directly sent to multiple ASR engines for parallel processing. The resources consumed by the ASR engines are fully utilized and controllable, the processing time is shortened, the memory consumption is reduced, the ASR results from different voice regions are stored in different sub-queues according to the voice regions, and sorted and taken out for processing according to the start time of the user's voice information. When an ASR result is executed, the sub-queues are re-sorted to realize that when instructions are sent to different users in multiple voice regions at the same time, the user who finishes speaking the instruction first is given priority for processing, improving the user experience.

[0095] Corresponding to the above voice processing method, the present disclosure also proposes a voice processing device. Figure 7 It is a schematic structural diagram of a voice processing device 700 provided by an embodiment of the present disclosure. As Figure 7 shown, it includes:

[0096] A receiving unit 710, configured to receive voice information from N different voice regions, where N is a positive integer greater than or equal to 2.

[0097] A voice recognition unit 720, configured to respectively recognize voice information through voice recognition engines corresponding to N different voice regions, and obtain a plurality of voice recognition results corresponding to the voice information of the N different voice regions;

[0098] A storage unit 730, configured to store a plurality of voice recognition results corresponding to the voice information of the N different voice regions into sub-queues corresponding to the N different voice regions in a voice recognition queue, where the N different voice regions respectively correspond to different sub-queues, and each sub-queue includes a head and a tail, and the head and the tail of the sub-queue are respectively used to store a voice recognition result;

[0099] A sorting and fetching unit 740, configured to sort the sub-queues in ascending order of the first time of the sub-queues, and execute the voice recognition result stored in the head of the sub-queue with the first sorting result, where the first time is the start time of the voice information corresponding to the voice recognition result stored in the head of the sub-queue.

[0100] In some embodiments, it further includes a detection module, configured to perform voice activity detection on the voice information of the N different voice regions before sorting the sub-queues in ascending order of the first time of the sub-queues, and obtain detection results corresponding to the voice information of the N different voice regions, where the detection results at least include the start time of the voice information; the storage unit 730 is specifically configured to: after binding a plurality of voice recognition results corresponding to the voice information of the N different voice regions and the corresponding detection results, store them into sub-queues corresponding to the N different voice regions in the voice recognition queue.

[0101] In some embodiments, the storage unit 730 is specifically configured to: in response to receiving a first voice recognition result of a voice region among the N different voice regions, create a first sub-queue in the voice recognition queue, and store the first voice recognition result into the head of the first sub-queue, where the first voice recognition result is the first voice recognition result corresponding to the voice information of the voice region, the first sub-queue includes a head and a tail, and the voice recognition queue is initially an empty queue.

[0102] In some embodiments, the storage unit 730 is specifically configured to: in response to receiving a second voice recognition result of the voice region, store the second voice recognition result into the tail of the first sub-queue, where the second voice recognition result is the second voice recognition result corresponding to the voice information of the voice region; in response to receiving a third voice recognition result of the voice region, delete the second voice recognition result from the first sub-queue, and store the third voice recognition result into the tail of the first sub-queue, where the third voice recognition result is the third voice recognition result corresponding to the voice information of the voice region.

[0103] In some embodiments, the sorting and fetching unit 740 is specifically configured to: in response to the first time in the first sub-queue being the smallest, fetch and execute the first speech recognition result of the first sub-queue; move the third speech recognition result from the tail of the first sub-queue to the head of the first sub-queue.

[0104] In some embodiments, the sorting and fetching unit 740 is further configured to: after the execution of the first speech recognition result of the first sub-queue is completed, obtain the start time of the speech information corresponding to the speech recognition result stored at the tail of the first sub-queue; compare the start time with the first time of the sub-queues corresponding to multiple sound zones, where the multiple sound zones include the other sound zones except the sound zone corresponding to the first word among the N sound zones, to obtain a comparison result; and re-sort the first sub-queue and the sub-queues corresponding to the multiple sound zones according to the comparison result.

[0105] In some embodiments, the storage unit 730 is further configured to: after fetching and executing the third speech recognition result of the first sub-queue, in response to the first speech recognition result and the third speech recognition result of the first sub-queue being fetched and no other speech recognition results from the sound zones being received, delete the first sub-queue from the speech recognition queue.

[0106] In summary, according to the embodiments of the present disclosure, the device stores and sorts and fetches the ASR results from different sound zones through the speech recognition unit, the storage unit, and the sorting and fetching unit, so as to respond to the user's speech as quickly and accurately as possible.

[0107] It should be noted that since the device embodiments of the present disclosure correspond to the above method embodiments, the foregoing explanations of the method embodiments also apply to the devices in this embodiment. The principles are the same. For the details not disclosed in the device embodiments, reference may be made to the above method embodiments, and no further elaboration will be provided in the present disclosure.

[0108] According to the embodiments of the present disclosure, the present disclosure also provides an electronic device, a readable storage medium, and a computer program product.

[0109] Figure 8 FIG. shows a schematic block diagram of an exemplary electronic device 800 that can be used to implement the embodiments of the present disclosure. The electronic device is intended to represent various forms of digital computers, such as, for example, a laptop computer, a desktop computer, a workbench, a personal digital assistant, a server, a blade server, a mainframe computer, and other suitable computers. The electronic device can also represent various forms of mobile devices, such as, for example, a personal digital processor, a cellular phone, a smart phone, a wearable device, and other similar computing devices. The components shown herein, their connections and relationships, and their functions are merely illustrative and are not intended to limit the implementation of the present disclosure described and / or claimed herein.

[0110] AsFigure 8 As shown, the device 800 includes a computing unit 801, which can perform various appropriate actions and processes according to a computer program stored in a ROM (Read-Only Memory) 802 or a computer program loaded from a storage unit 808 to a RAM (Random Access Memory) 803. In the RAM 803, various programs and data required for the operation of the device 800 can also be stored. The computing unit 801, the ROM 802, and the RAM 803 are connected to each other via a bus 804. An I / O (Input / Output) interface 805 is also connected to the bus 804.

[0111] A number of components in the device 800 are connected to the I / O interface 805, including: an input unit 806, such as a keyboard, a mouse, etc.; an output unit 807, such as various types of displays, speakers, etc.; a storage unit 808, such as a disk, an optical disk, etc.; and a voice processing unit 809, such as a network card, a modem, a wireless voice processing transceiver, etc. The voice processing unit 809 allows the device 800 to exchange information / data with other devices through a computer network such as the Internet and / or various telecommunication networks.

[0112] The computing unit 801 may be a variety of general and / or special processing components with processing and computing capabilities. Some examples of the computing unit 801 include, but are not limited to, CPU (Central Processing Unit), GPU (Graphic Processing Units), various dedicated AI (Artificial Intelligence) computing chips, various computing units running machine learning model algorithms, DSP (Digital Signal Processor), and any appropriate processor, controller, microcontroller, etc. The computing unit 801 performs the various methods and processes described above, such as a speech processing method. For example, in some embodiments, the speech processing method may be implemented as a computer software program, which is tangibly contained in a machine-readable medium, such as a storage unit 808. In some embodiments, part or all of the computer program may be loaded and / or installed on the device 800 via ROM 802 and / or speech processing unit 809. When the computer program is loaded into RAM 803 and executed by the computing unit 801, one or more steps of the method described above may be performed. Alternatively, in other embodiments, the computing unit 801 may be configured to execute the aforementioned speech processing method in any other appropriate manner (for example, by means of firmware).

[0113] The various embodiments of the systems and techniques described above in this specification can be implemented in digital electronic circuitry, integrated circuit systems, FPGAs (Field Programmable Gate Arrays), ASICs (Application-Specific Integrated Circuits), ASSPs (Application Specific Standard Products), SOCs (System On Chip), CPLDs (Complex Programmable Logic Devices), computer hardware, firmware, software, and / or combinations thereof. These various embodiments can include: being implemented in one or more computer programs that are executable and / or interpretable on a programmable system including at least one programmable processor, which can be a special-purpose or general-purpose programmable processor that receives data and instructions from, and transmits data and instructions to, a storage system, at least one input device, and at least one output device.

[0114] The program code for implementing the methods of the present disclosure can be written in any combination of one or more programming languages. These program codes can be provided to a processor or controller of a general-purpose computer, a special-purpose computer, or other programmable data processing device, such that the program codes, when executed by the processor or controller, cause the functions / operations specified in the flowchart and / or block diagram to be implemented. The program code can be executed entirely on the machine, partially on the machine, as a stand-alone software package partially on the machine and partially on a remote machine, or entirely on a remote machine or server.

[0115] In the context of this disclosure, a machine-readable medium can be a tangible medium that can contain or store a program for use by or in connection with an instruction execution system, apparatus, or device. A machine-readable medium can be a machine-readable signal medium or a machine-readable storage medium. A machine-readable medium can include, but is not limited to, electronic, magnetic, optical, electromagnetic, infrared, or semiconductor systems, apparatus, or devices, or any suitable combination of the foregoing. More specific examples of a machine-readable storage medium would include an electrical connection based on one or more wires, a portable computer diskette, a hard disk, a RAM, a ROM, an EPROM (Electrically Programmable Read-Only Memory), or a flash memory, an optical fiber, a CD-ROM (Compact Disc Read-Only Memory), an optical storage device, a magnetic storage device, or any suitable combination of the foregoing.

[0116] To provide for interaction with a user, the systems and techniques described herein can be implemented on a computer having: a display device (e.g., a CRT (Cathode-Ray Tube) or LCD (Liquid Crystal Display) monitor) for displaying information to the user; and a keyboard and a pointing device (e.g., a mouse or a trackball) by which the user can provide input to the computer. Other kinds of devices can also be used to provide for interaction with the user; for example, feedback provided to the user can be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and input from the user can be received in any form (including acoustic input, voice input, or tactile input).

[0117] The systems and techniques described herein can be implemented in a computing system that includes back-end components (e.g., as a data server), or a computing system that includes middleware components (e.g., an application server), or a computing system that includes front-end components (e.g., a user computer having a graphical user interface or a web browser through which the user can interact with an implementation of the systems and techniques described herein), or a computing system that includes any combination of such back-end components, middleware components, or front-end components. The components of the system can be interconnected by any form or medium of digital data voice processing (e.g., a voice processing network). Examples of a voice processing network include: a LAN (Local Area Network), a WAN (Wide Area Network), the Internet, and a blockchain network.

[0118] A computer system may include a client and a server. The client and the server are generally far from each other and usually interact via a voice processing network. The relationship between the client and the server is generated by computer programs running on the respective computers and having a client-server relationship with each other. The server may be a cloud server, also known as a cloud computing server or a cloud host, which is a host product in the cloud computing service system, solving the defects of difficult management and weak business scalability existing in traditional physical hosts and VPS services ("Virtual Private Server", or simply "VPS"). The server may also be a server of a distributed system or a server combined with a blockchain.

[0119] Among them, it should be noted that artificial intelligence is a discipline that studies how to make a computer simulate certain thinking processes and intelligent behaviors of humans (such as learning, reasoning, thinking, planning, etc.), and there are both hardware-level technologies and software-level technologies. Artificial intelligence hardware technologies generally include technologies such as sensors, dedicated artificial intelligence chips, cloud computing, distributed storage, and big data processing; artificial intelligence software technologies mainly include several major directions such as computer vision technology, speech recognition technology, natural language processing technology, and machine learning / deep learning, big data processing technology, and knowledge graph technology.

[0120] It should be understood that various forms of processes shown above can be used, reordering, adding, or deleting steps. For example, the steps recorded in the present disclosure can be executed in parallel, sequentially, or in a different order, as long as the desired results of the technical solutions disclosed in the present disclosure can be achieved, and no limitations are imposed herein. The above specific implementation manners do not constitute a limitation on the protection scope of the present disclosure. Those skilled in the art should understand that various modifications, combinations, sub-combinations, and substitutions can be made according to design requirements and other factors. Any modifications, equivalent substitutions, and improvements made within the spirit and principles of the present disclosure shall be included within the protection scope of the present disclosure.

Claims

1. A voice processing method, characterized in that, The method includes: Receiving voice information in N different pitch ranges, where N is a positive integer greater than or equal to 2; Respectively identifying the voice information through voice recognition engines corresponding to the N different pitch ranges to obtain a plurality of voice recognition results corresponding to the voice information in the N different pitch ranges; Storing the plurality of voice recognition results corresponding to the voice information in the N different pitch ranges into sub-queues corresponding to the N different pitch ranges in a voice recognition queue, where the N different pitch ranges respectively correspond to different sub-queues, and each sub-queue includes a head and a tail, and the head and the tail of the sub-queue are respectively used to store a voice recognition result; Sorting the sub-queues in ascending order of the first time of the sub-queues, and taking and executing the voice recognition result stored in the head of the sub-queue with the first sorting order, where the first time is the start time of the voice information corresponding to the voice recognition result stored in the head of the sub-queue.

2. The method according to claim 1, characterized in that, Before sorting the sub-queues in ascending order of the first time of the sub-queues, the method further includes: Performing voice activity detection on the voice information in the N different pitch ranges to obtain detection results corresponding to the voice information in the N different pitch ranges, where the detection results at least include the start time of the voice information; The storing the plurality of voice recognition results corresponding to the voice information in the N different pitch ranges into sub-queues corresponding to the N different pitch ranges in a voice recognition queue includes: After binding the plurality of voice recognition results corresponding to the voice information in the N different pitch ranges and the corresponding detection results, storing them into sub-queues corresponding to the N different pitch ranges in a voice recognition queue.

3. The method according to claim 2, characterized in that, The storing the plurality of voice recognition results corresponding to the voice information in the N different pitch ranges into sub-queues corresponding to the voice recognition queue includes: For any one of the N different pitch ranges, in response to receiving a first voice recognition result of the pitch range, creating a first sub-queue in the voice recognition queue and storing the first voice recognition result into the head of the first sub-queue, where the first voice recognition result is the first voice recognition result corresponding to the voice information of the pitch range, the first sub-queue includes a head and a tail, and the voice recognition queue is initially an empty queue.

4. The method according to claim 3, characterized in that, After creating a first sub-queue in the voice recognition queue and storing the first voice recognition result into the head of the first sub-queue in response to receiving the first voice recognition result of the pitch range, the method further includes: In response to receiving a second voice recognition result of the pitch range, storing the second voice recognition result into the tail of the first sub-queue, where the second voice recognition result is the second voice recognition result corresponding to the voice information of the pitch range; In response to receiving a third voice recognition result of the pitch range, deleting the second voice recognition result from the first sub-queue and storing the third voice recognition result into the tail of the first sub-queue, where the third voice recognition result is the third voice recognition result corresponding to the voice information of the pitch range.

5. The method according to claim 4, characterized in that, Sorting the sub - queues in ascending order of the first time of the sub - queues, and taking out and executing the speech recognition result stored at the head of the sub - queue with the first sorting includes: In response to the first time of the first sub - queue being the smallest, taking out and executing the first speech recognition result of the first sub - queue; Moving the third speech recognition result from the tail of the first sub - queue to the head of the first sub - queue.

6. The method according to claim 5, characterized in that, After taking out and executing the first speech recognition result of the first sub - queue, the method further includes: In response to the completion of the execution of the first speech recognition result of the first sub - queue, obtaining the start time of the speech information corresponding to the speech recognition result stored at the tail of the first sub - queue; Comparing the start time with the first time of the sub - queues corresponding to multiple sound regions, obtaining a comparison result, where the multiple sound regions include other sound regions in the N sound regions except the sound region corresponding to the first sub - queue; Re - sorting the first sub - queue and the sub - queues corresponding to the multiple sound regions according to the comparison result.

7. The method according to any one of claim 5, characterized in that, The method further includes: In response to both the first speech recognition result and the third speech recognition result of the first sub - queue being taken out and no speech recognition result corresponding to other speech information from the sound region being received, deleting the first sub - queue from the speech recognition queue.

8. A voice processing device, characterized in that, The device includes: A receiving unit, configured to receive speech information of N different sound regions, where N is a positive integer greater than or equal to 2; A speech recognition unit, configured to respectively recognize the speech information through speech recognition engines corresponding to the N different sound regions, and obtain multiple speech recognition results corresponding to the speech information of the N different sound regions; A storage unit, configured to store the multiple speech recognition results corresponding to the speech information of the N different sound regions into the sub - queues corresponding to the N different sound regions in the speech recognition queue, where the N different sound regions respectively correspond to different sub - queues, the sub - queue includes a head and a tail, and the head and the tail of the sub - queue are respectively used to store a speech recognition result; A sorting and taking - out unit, configured to sort the sub - queues in ascending order of the first time of the sub - queues, take out and execute the speech recognition result stored at the head of the sub - queue with the first sorting, where the first time is the start time of the speech information corresponding to the speech recognition result stored at the head of the sub - queue.

9. An electronic device, comprising: At least one processor; And A memory communicatively connected to the at least one processor; wherein, The memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor so that the at least one processor can execute the method according to any one of claims 1 - 7.

10. A non-transitory computer-readable storage medium storing computer instructions, wherein, The computer instructions are used to cause the computer to execute the method according to any one of claims 1 - 7.

11. A computer program product comprising a computer program which, when executed by a processor, implements the method according to any one of claims 1-7.