Speech processing method and apparatus, device, and medium

By obtaining multiple audio data in the voice processing method and comparing the energy values ​​of the first and other audio data, dynamically adjusting the wake-up word triggering time, the problem of balancing the voice interaction response speed and sound source positioning accuracy is solved, and the user experience is improved.

WO2025130669A1PCT designated stage expired Publication Date: 2025-06-26NIO TECH ANHUI CO LTD
View PDF 5 Cites 0 Cited by

Patent Information

Application Number
PCT/CN2024/137667
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2023-12-22
Filing Date
2024-12-09
Publication Date
2025-06-26

AI Technical Summary

Technical Problem

The existing voice processing methods cannot balance the voice interaction response speed and sound source positioning accuracy, resulting in a decrease in user wake-up rate and an increase in interaction delay in vehicle-mounted scenarios.

Method used

By acquiring the multi-channel audio data in the driving device, when the first audio data triggers the wake-up word, the first energy value and the second energy value are obtained, and based on the comparison results, decide whether the other audio data is needed for wake-up word triggering, or dynamically allocate the wake-up word triggering time of the other audio data.

Benefits of technology

On the premise of ensuring the accuracy of sound source positioning, reduce the interaction waiting time and improve the user experience.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN2024137667_26062025_PF_FP_ABST
    Figure CN2024137667_26062025_PF_FP_ABST
Patent Text Reader

Abstract

The present application relates to the technical field of speech processing, and specifically provides a speech processing method and apparatus, a device, and a medium, aiming at solving the problem that existing speech processing methods cannot achieve the balance between the speech interaction response speed and the sound source localization accuracy. To achieve this objective, the method of the present application comprises: acquiring multiple paths of audio data in a driving device; when the first path of audio data first triggers a wake-up word, acquiring a first energy value and second energy values and comparing the first energy value with the second energy values, wherein the first energy value is an energy value corresponding to the first path of audio data, and the second energy values are energy values corresponding to the remaining paths of audio data among the multiple paths of audio data; and if the first energy value is greater than or equal to the second energy values, the remaining paths of audio data not needing to perform wake-up word triggering, or otherwise, dynamically allocating wake-up word triggering durations of the remaining paths of audio data. Such a configuration can reduce an interaction waiting duration while guaranteeing the sound source localization accuracy as much as possible.
Need to check novelty before this filing date? Find Prior Art

Description

Voice processing method, device, equipment and medium

[0001] This application claims priority to Chinese patent application No. 202311797456.0, filed on December 22, 2023, with the invention name “Speech processing method, device, equipment and medium”. The entire contents of the above Chinese patent application are incorporated into this application by reference. Technical Field

[0002] The present application relates to the field of speech processing technology, and specifically provides a speech processing method, apparatus, device and medium. Background Art

[0003] Voice interaction systems have become essential for smart cockpits. However, in-vehicle voice interaction is severely impacted by in-vehicle noise, including the motor, multimedia, air conditioning, and human voices, significantly increasing the difficulty of voice interaction systems recognizing user voices. To ensure user wake-up rates in in-vehicle scenarios, existing technologies typically incorporate a front-end voice processing module (ECNR: Echo Cancellation and Noise Reduction) into voice interaction systems. This module suppresses ambient noise and enhances the effective voice signal, then sends the processed audio signal to the wake-up engine for wake-up processing.

[0004] Due to the limitations of the front-end speech processing module (ECNR) algorithm, the audio collected by the microphone in the voice interaction system's non-human voice location still retains some wake-up word after processing. Therefore, to accurately locate the human voice, it is usually necessary to trigger the wake-up word on all channels before making the next judgment. However, if the wake-up word is triggered on all channels, it is easy to cause a long waiting time, introducing a large delay and seriously affecting the user experience. If the channel that first triggers the wake-up word is directly used as the location of the human voice, the positioning accuracy will drop sharply.

[0005] Accordingly, the art needs a new speech processing method to solve the above problems. Summary of the Invention

[0006] This application aims to solve the above technical problems, namely, to solve the problem that existing speech processing methods cannot balance the speech interaction response speed and the sound source localization accuracy.

[0007] To achieve the above objectives, in a first aspect, the present application provides a speech processing method, applied to a driving device, comprising the following steps:

[0008] Acquiring multi-channel audio data in the driving device;

[0009] When the first channel of audio data triggers the wake-up word for the first time, obtaining a first energy value and a second energy value and comparing the first energy value with the second energy value, wherein the first energy value is an energy value corresponding to the first channel of audio data, and the second energy value is an energy value corresponding to the remaining channels of audio data in the multiple channels of audio data;

[0010] If the first energy value is greater than or equal to the second energy value, the remaining audio data is not required to trigger the wake-up word;

[0011] If the first energy value is less than the second energy value, the wake-up word triggering duration of the remaining audio data is dynamically allocated.

[0012] In an optional technical solution of the above-mentioned speech processing method, before obtaining the first energy value and the second energy value, the method further includes:

[0013] The first channel of audio data and the remaining channels of audio data are filtered respectively based on preset rules.

[0014] In an optional technical solution of the above-mentioned voice processing method, the first channel of audio data and the remaining channels of audio data each include multiple sampling points, and the step of “filtering the first channel of audio data and the remaining channels of audio data based on preset rules” includes:

[0015] Performing preset energy value screening on multiple sampling points of the first channel of audio data and the remaining channels of audio data respectively;

[0016] After completing the preset energy value screening, the first channel of audio data and the remaining channels of audio data are respectively screened for preset durations.

[0017] In an optional technical solution of the above-mentioned speech processing method, the step of “respectively screening the plurality of sampling points of the first channel of audio data and the remaining channels of audio data for preset energy values” includes:

[0018] Performing a first preset energy value screening on each sampling point in the first channel of audio data and the remaining channels of audio data respectively;

[0019] Alternatively, second preset energy value screening is performed on preset sampling points in the first channel of audio data and the remaining channels of audio data respectively.

[0020] In an optional technical solution of the above-mentioned voice processing method, before filtering the first channel of audio data and the remaining channels of audio data based on preset rules, the method further includes:

[0021] Filtering the first channel of audio data based on the wake-up word to obtain first channel of audio data containing only the wake-up word;

[0022] Obtain a time zone corresponding to the first channel of audio data containing only the wake-up word, and filter the remaining channels of audio data based on the time zone.

[0023] In the optional technical solution of the above-mentioned voice processing method, the step of "dynamically allocating the wake-up word triggering duration of the remaining audio data" includes:

[0024] obtaining a difference between the first energy value and the second energy value;

[0025] Based on the difference, the wake-up word triggering duration of the remaining audio data is dynamically allocated.

[0026] In an optional technical solution of the above-mentioned voice processing method, the remaining audio data includes X channels of audio data, where X is greater than 1 and is a positive integer. After dynamically allocating the wake-up word trigger duration of the remaining audio data, the method further includes:

[0027] When Y channels of audio data among the remaining channels of audio data trigger the wake-up word, comparing energy values ​​corresponding to the Y channels of audio data, where Y is greater than 1 and less than or equal to X, and Y is a positive integer;

[0028] Based on the comparison result of the energy values ​​corresponding to the Y-channel audio data, the user's location is determined.

[0029] In a second aspect, the present application further provides a speech processing device, applied to a driving device, the device comprising:

[0030] an audio data acquisition module, configured to acquire multi-channel audio data in the driving device;

[0031] an energy value acquisition and comparison module, configured to, when the first channel of audio data first triggers the wake-up word, acquire a first energy value and a second energy value and compare the first energy value with the second energy value, wherein the first energy value is an energy value corresponding to the first channel of audio data, and the second energy value is an energy value corresponding to the remaining channels of audio data in the multiple channels of audio data;

[0032] A first processing module is configured to, if the first energy value is greater than or equal to the second energy value, eliminate the need for the remaining audio data to trigger the wake-up word;

[0033] The second processing module is configured to dynamically allocate the wake-up word triggering duration of the remaining audio data if the first energy value is less than the second energy value.

[0034] In a third aspect, the present application further provides a computer device comprising a processor and a storage device, wherein the storage device is suitable for storing a plurality of program codes, and the program codes are suitable for being loaded and run by the processor to execute any of the above-mentioned speech processing methods.

[0035] In a fourth aspect, the present application further provides a computer-readable storage medium, wherein a plurality of program codes are stored in the computer-readable storage medium, and the program codes are suitable for being loaded and run by a processor to execute any one of the above-mentioned speech processing methods.

[0036] Those skilled in the art will understand that, in the technical solution of the present application, by obtaining multiple channels of audio data in the driving device; when the first channel of audio data triggers the wake-up word for the first time, the first energy value and the second energy value are obtained and the first energy value is compared with the second energy value, wherein the first energy value is the energy value corresponding to the first channel of audio data, and the second energy value is the energy value corresponding to the remaining channels of audio data in the multiple channels of audio data; if the first energy value is greater than or equal to the second energy value, the remaining channels of audio data are not required to trigger the wake-up word; if the first energy value is less than the second energy value, the wake-up word triggering duration of the remaining channels of audio data is dynamically allocated. Such a setting can reduce the interaction waiting time while ensuring the accuracy of sound source positioning as much as possible, thereby improving the user experience.

[0037] Furthermore, the first channel of audio data and the remaining channels of audio data each contain multiple sampling points, and filtering the first channel of audio data and the remaining channels of audio data based on preset rules includes: filtering the multiple sampling points of the first channel of audio data and the remaining channels of audio data by preset energy values; after completing the preset energy value filtering, filtering the first channel of audio data and the remaining channels of audio data by preset durations. This setting can avoid the user's pronunciation affecting the calculation of the audio energy value, thereby interfering with the allocation of the wake-up word trigger duration based on the energy value, further improving the user experience. BRIEF DESCRIPTION OF THE DRAWINGS

[0038] The disclosure of this application will be more easily understood with reference to the accompanying drawings. Those skilled in the art will readily appreciate that these drawings are for illustrative purposes only and are not intended to limit the scope of protection of this application. Furthermore, similar numbers in the figures represent similar components, where:

[0039] FIG1 is a flow chart showing the main steps of a speech processing method according to an embodiment of the present application;

[0040] FIG2 is a schematic flow chart of main steps for filtering the first channel of audio data and the remaining channels of audio data based on preset rules according to an embodiment of the present application;

[0041] 3 is a flowchart illustrating the main steps before filtering the first channel of audio data and the remaining channels of audio data based on preset rules according to an embodiment of the present application;

[0042] FIG4 is a flowchart illustrating the main steps of dynamically allocating the wake-up word trigger duration of the remaining audio data according to an embodiment of the present application;

[0043] FIG5 is a schematic diagram of a detailed flow chart of a speech processing method according to an embodiment of the present application;

[0044] FIG6 is a flow chart of a speech processing method according to the present application in an actual application scenario;

[0045] FIG7 is a schematic diagram of a main structural block diagram of a voice device according to an embodiment of the present application;

[0046] FIG8 is a schematic diagram of the main structure of a computer device according to an embodiment of the present application.

[0047] List of reference numerals:

[0048] 11: audio data acquisition module; 12: energy value acquisition and comparison module; 13: first processing module; 14: second processing module. DETAILED DESCRIPTION

[0049] Some embodiments of the present application are described below with reference to the accompanying drawings. Those skilled in the art should understand that these embodiments are only used to explain the technical principles of the present application and are not intended to limit the scope of protection of the present application.

[0050] In the description of this application, "module" and "processor" may include hardware, software, or a combination of both. A module may include hardware circuitry, various suitable sensors, communication ports, and memory. It may also include software components, such as program code, or a combination of software and hardware. A processor may be a central processing unit, a microprocessor, an image processor, a digital signal processor, or any other suitable processor. A processor has data and / or signal processing capabilities. A processor may be implemented in software, hardware, or a combination of both. Non-transitory computer-readable storage media include any suitable medium capable of storing program code, such as magnetic disks, hard disks, optical disks, flash memory, read-only memory, random access memory, etc. The term "A and / or B" refers to all possible combinations of A and B, such as only A, only B, or both A and B. The terms "at least one of A or B" or "at least one of A and B" have similar meanings to "A and / or B" and may include only A, only B, or both A and B. The singular forms "a" and "the" may also include the plural forms.

[0051] As described in the background technology section, to address the problem that existing speech processing methods cannot balance the speech interaction response speed and the sound source localization accuracy, the present application provides a speech processing method.

[0052] Referring to FIG1 , FIG1 is a flow chart illustrating the main steps of a speech processing method according to an embodiment of the present application. The speech processing method can be executed by a driving device, a server, or both a server and a driving device. As shown in FIG1 , the speech processing method of the present application is applied to a driving device, and the method comprises the following steps:

[0053] Step S101: Acquire multi-channel audio data in the driving device.

[0054] Specifically, the driving device is provided with various sound receiving devices, such as microphones. When the user speaks in the driving device, the various sound receiving devices can receive the sound to obtain the multi-channel audio data in the driving device.

[0055] Step S102: When the first audio data triggers the wake-up word for the first time, obtain the first energy value and the second energy value and compare the first energy value with the second energy value, where the first energy value is the energy value corresponding to the first audio data, and the second energy value is the energy value corresponding to the remaining audio data in the multiple audio data.

[0056] Specifically, a wake-up word is a technology based on intelligent speech recognition. It can wake up the voice interaction system on the device by recognizing specific spoken wake-up words, allowing the voice interaction system to receive user commands and complete corresponding actions. When using the voice interaction system, users can wake up the voice interaction system by speaking the wake-up word and perform corresponding operations or queries based on the instructions issued by the wake-up word. Energy values ​​can represent the differences between different audio data and can be expressed in various ways. For example, energy values ​​can be expressed as at least one of amplitude, the mean of the sum of squared amplitudes, and decibels.

[0057] Step S103: If the first energy value is greater than or equal to the second energy value, the remaining audio data is not required to trigger the wake-up word.

[0058] Specifically, if the energy value of the first audio data that triggers the wake-up word is greater than or equal to the energy value of the remaining audio data that does not trigger the wake-up word, the wake-up word triggering time will no longer be allocated to the remaining audio data, and the location of the first audio data will be used as the location of the user. This avoids the situation where a longer wake-up word triggering time is spent due to the low energy value of the audio data, thereby improving the voice triggering efficiency.

[0059] Step S104: If the first energy value is less than the second energy value, the wake-up word triggering duration of the remaining audio data is dynamically allocated.

[0060] Specifically, if the energy value of the first audio data that triggers the wake-up word is less than the energy value of the remaining audio data that does not trigger the wake-up word, the wake-up word triggering duration is dynamically allocated to the remaining audio data based on the first energy value and the second energy value. This makes it easier to locate the user's sound source based on the wake-up word triggering results of the remaining audio data, thereby ensuring the accuracy of locating the user's sound source as much as possible.

[0061] In some embodiments, the remaining audio data includes multiple audio data, wherein the remaining audio data includes audio whose energy value is less than or equal to the first audio data, and also includes audio whose energy value is greater than the first audio data. Then the method is: there is no need to trigger the wake-up word for the audio in the remaining audio data whose energy value is less than or equal to the first audio data; dynamically allocate the wake-up word triggering duration of the audio in the remaining audio data whose energy value is greater than the first audio data.

[0062] Based on the above steps S101 to S104, the present application obtains multiple channels of audio data in the driving device; when the first channel of audio data triggers the wake-up word for the first time, obtains the first energy value and the second energy value and compares the first energy value with the second energy value, wherein the first energy value is the energy value corresponding to the first channel of audio data, and the second energy value is the energy value corresponding to the remaining channels of audio data in the multiple channels of audio data; if the first energy value is greater than or equal to the second energy value, the remaining channels of audio data are not required to trigger the wake-up word; if the first energy value is less than the second energy value, the wake-up word triggering duration of the remaining channels of audio data is dynamically allocated. Such a setting can reduce the interaction waiting time while ensuring the accuracy of sound source positioning as much as possible, thereby improving the user experience.

[0063] Next, step S102 and step S104 are further explained.

[0064] In some embodiments, before obtaining the first energy value and the second energy value, the method further includes: filtering the first channel of audio data and the remaining channels of audio data based on preset rules.

[0065] Referring to FIG. 2 , FIG. 2 is a flow chart illustrating the main steps of filtering the first channel of audio data and the remaining channels of audio data based on preset rules according to one embodiment of the present application. As shown in FIG. 2 , in some embodiments, the first channel of audio data and the remaining channels of audio data each include multiple sampling points, and filtering the first channel of audio data and the remaining channels of audio data based on the preset rules includes the following steps:

[0066] Step S201: performing preset energy value screening on a plurality of sampling points of the first channel of audio data and the remaining channels of audio data respectively.

[0067] Step S202: After completing the preset energy value screening, the first channel of audio data and the remaining channels of audio data are screened for preset durations respectively.

[0068] Specifically, after the audio data is received by the sound receiving device, it is converted into a discrete digital signal, i.e., a waveform file, through sampling, quantization, and encoding. Different sampling rates will cause the received audio data to contain different sampling points. For example, when the sampling rate is 16K, the audio data with a duration of 1 second contains 16,000 sampling points; when the sampling rate is 8K, the audio data with a duration of 1 second contains 8,000 sampling points. Preset energy value screening of multiple sampling points of the first audio data and the remaining audio data can avoid the influence of different speaking methods on the calculation of the energy value. For example, if the audio data contains some segments where the user pauses or the user prolongs the tail sound, it will affect the overall energy value of the audio data. Therefore, it is necessary to perform preset energy value screening on the audio data. After completing the preset energy value screening, the audio data needs to be further screened for a preset duration. The setting of the preset duration is related to the length of the wake-up word. That is, the longer the wake-up word, the longer the preset duration, and the shorter the wake-up word, the smaller the preset duration, to ensure that all wake-up words are included in the audio data. After completing the preset energy value screening and the preset time length screening for the first channel of audio data and the remaining channels of audio data respectively, the energy value of the first channel of audio data and the energy value of the remaining channels of audio data are respectively obtained.

[0069] In some embodiments, performing preset energy value screening on multiple sampling points of the first channel of audio data and the remaining channels of audio data respectively includes the following steps:

[0070] Step S301: performing a first preset energy value screening on each sampling point in the first channel of audio data and the remaining channels of audio data.

[0071] Step S302: Alternatively, a second preset energy value screening is performed on the preset sampling points in the first channel of audio data and the remaining channels of audio data respectively.

[0072] Specifically, when performing a first preset energy value screening on each sampling point in the first channel of audio data and the remaining channels of audio data, the energy value of each sampling point in the first channel of audio data and the remaining channels of audio data is obtained, and each sampling point in the first channel of audio data and the remaining channels of audio data is sorted in descending order based on the energy value, so as to facilitate the subsequent first preset energy value screening. When performing a second preset energy value screening on the preset sampling points in the first channel of audio data and the remaining channels of audio data, the energy values ​​of the preset sampling points in the first channel of audio data and the remaining channels of audio data are obtained, and the preset sampling points in the first channel of audio data and the remaining channels of audio data are sorted in descending order based on the energy value, so as to facilitate the subsequent second preset energy value screening.

[0073] Exemplarily, the energy value of each sampling point can be represented by the amplitude of each sampling point. Then, the first preset energy value screening is performed on each sampling point in the first channel of audio data and the remaining channels of audio data, respectively, as follows: determining whether the amplitude of each sampling point in the first channel of audio data is greater than the first preset amplitude; if so, retaining the sampling point in the first channel of audio data; otherwise, deleting the sampling point from the first channel of audio data; the first preset energy value screening of the remaining channels of audio data is similar to the first energy value screening of the first channel of audio data, and will not be repeated here.

[0074] The energy values ​​of the preset sampling points can be represented by the mean of the sum of the squares of the amplitudes of the preset sampling points. Then, the second preset energy value screening is performed on the preset sampling points in the first channel of audio data and the remaining channels of audio data, respectively, as follows: determining whether the mean of the sum of the squares of the amplitudes of each preset sampling point in the first channel of audio data is greater than the second preset mean of the sum of the squares of the amplitudes; if so, retaining the preset sampling points in the first channel of audio data; otherwise, deleting the preset sampling points from the first channel of audio data; the second preset energy value screening of the remaining channels of audio data is similar to the second energy value screening of the first channel of audio data, and is not further described here.

[0075] The above setting method of the first preset energy value screening and the setting method of the second preset energy value screening are only for illustrative purposes, and can be selected according to actual needs in actual applications.

[0076] Referring to FIG. 3 , FIG. 3 is a flow chart illustrating the main steps before filtering the first channel of audio data and the remaining channels of audio data based on preset rules according to an embodiment of the present application. As shown in FIG. 3 , in some embodiments, before filtering the first channel of audio data and the remaining channels of audio data based on preset rules, the method further includes the following steps:

[0077] Step S401: filtering the first channel of audio data based on the wake-up word to obtain the first channel of audio data containing only the wake-up word.

[0078] Step S402: Obtain the time zone corresponding to the first channel of audio data containing only the wake-up word, and filter the remaining channels of audio data based on the time zone.

[0079] Specifically, in addition to the wake-up word, the first channel of audio data may also contain other audios, so it is necessary to filter the first channel of audio data to obtain the first channel of audio data that only contains the wake-up word. Obtain the corresponding time area of ​​the wake-up word in the first channel of audio data, that is, the generation time of the wake-up word in the first channel of audio data, and then filter the remaining channels of audio data to obtain the remaining channels of audio data that only contain the corresponding time area, so as to facilitate the subsequent triggering of the wake-up word on the remaining channels of audio data that only contain the corresponding time area. After obtaining the first channel of audio data that only contains the wake-up word, filter the first channel of audio data that only contains the wake-up word based on the preset rules; similarly, after obtaining the remaining channels of audio data that only contain the corresponding time area, filter the remaining channels of audio data that only contain the corresponding time area based on the preset rules.

[0080] Refer to Figure 4, which is a flowchart of the main steps of dynamically allocating the wake-up word trigger duration of the remaining audio data according to an embodiment of the present application. As shown in Figure 4, in some embodiments, dynamically allocating the wake-up word trigger duration of the remaining audio data includes the following steps:

[0081] Step 501: Obtain a difference between a first energy value and a second energy value.

[0082] Step 502: Based on the difference, dynamically allocate the wake-up word triggering duration of the remaining audio data.

[0083] Specifically, the greater the difference between the first energy value and the second energy value, the longer the wake-up word trigger time allocated to the remaining audio data; conversely, the smaller the difference between the first energy value and the second energy value, the shorter the wake-up word trigger time allocated to the remaining audio data. This setting can allocate sufficient wake-up word trigger time to audio with large energy values ​​in the remaining audio data to ensure the accuracy of subsequent positioning based on the user's sound source; at the same time, it can also allocate appropriate wake-up word trigger time to audio with large energy values ​​in the remaining audio data to avoid increasing the delay in voice interaction.

[0084] In some embodiments, the remaining audio data includes X channels of audio data, where X is greater than 1 and is a positive integer. After dynamically allocating the wake-up word trigger duration for the remaining audio data, the method further includes the following steps:

[0085] Step S601: When Y channels of audio data among the remaining channels of audio data trigger the wake-up word, the energy values ​​corresponding to the Y channels of audio data are compared, where Y is greater than 1 and less than or equal to X, and Y is a positive integer.

[0086] Step S602: Determine the user's location based on the comparison result of the energy values ​​corresponding to the Y-channel audio data.

[0087] Specifically, by comparing the energy values ​​corresponding to the Y-channel audio data, the audio data with the largest energy value in the Y-channel audio data is obtained, and the direction of the audio data with the largest energy value in the Y-channel audio data is used as the direction of the user.

[0088] Referring to FIG5 , FIG5 is a schematic flow chart of detailed steps of a speech processing method according to an embodiment of the present application. The speech processing method can be executed by a driving device, a server, or both a server and a driving device. As shown in FIG5 , the speech processing method of the present application is applied to a driving device, and the method includes the following steps:

[0089] Step S701: Acquire multi-channel audio data in the driving device.

[0090] Step S702: When the first channel of audio data triggers the wake-up word for the first time, the first channel of audio data is filtered based on the wake-up word to obtain the first channel of audio data containing only the wake-up word; the time area corresponding to the first channel of audio data containing only the wake-up word is obtained, and the remaining channels of audio data are filtered based on the time area to obtain the remaining channels of audio data containing only the corresponding time area.

[0091] Step S703: Based on preset rules, the first channel of audio data containing only the wake-up word and the remaining channels of audio data containing only the corresponding time zone are filtered to obtain the first channel of valid audio and the remaining channels of valid audio.

[0092] Step S704: Obtain the energy value of the first channel of valid audio and the energy values ​​of the remaining channels of valid audio.

[0093] Step S705: Compare the energy value of the first channel of valid audio with the energy values ​​of the remaining channels of valid speech.

[0094] Step S706: If the energy value of the first channel of valid audio is greater than or equal to the energy values ​​of the remaining channels of valid audio, there is no need to trigger the wake-up word using only the remaining channels of audio data in the corresponding time zone.

[0095] Step S707: If the energy value of the first channel of valid audio is less than the energy values ​​of the remaining channels of valid audio, the wake-up word triggering duration of the remaining channels of audio data containing only the corresponding time zone is dynamically allocated.

[0096] Step S708: The remaining valid audio channels include X valid audio channels, where X is greater than 1 and is a positive integer. When Y valid audio channels among the remaining valid audio channels trigger the wake-up word, the energy values ​​corresponding to the Y valid audio channels are compared, where Y is greater than 1 and less than or equal to X, and Y is a positive integer.

[0097] Step S709: Determine the user's location based on the comparison result of the energy values ​​corresponding to the Y-channel effective audio.

[0098] Specifically, the flow chart of the voice processing method involved in the actual application scenario of the present application can be shown in Figure 6. Microphones are set in N directions in the driving device for sound reception, where N is greater than 1 and N is a positive integer. The audio data of the N directions obtained are sent to the front-end voice processing module (ECNR) for processing, and audio data of M channels are obtained based on the audio data of the N directions that have been processed, where M is greater than 1 and M is a positive integer. M wake-up engines are started to synchronously process the audio data of M channels. When the first channel audio data triggers the wake-up word for the first time, the first channel audio data containing only the wake-up word and the remaining (M-1) channel audio data containing only the corresponding time zone are obtained; based on the preset rules, the first channel audio data containing only the wake-up word and the remaining (M-1) channel audio data containing only the corresponding time zone are filtered to obtain the first channel valid audio and the remaining (M-1) channel valid audio. The energy values ​​of the first channel valid audio and the remaining (M-1) channel valid audio are obtained respectively to dynamically allocate the wake-up word triggering duration, thereby ensuring the accuracy of positioning based on the user sound source.

[0099] Furthermore, the speech processing method of the present application can be applied to V2X scenarios. Specifically, the Internet of Vehicles (IoV) refers to the network communication technology applied to vehicles. This technology is based on the intra-vehicle network, inter-vehicle network and in-vehicle mobile Internet. According to the agreed communication protocols and data exchange standards, it is a large-scale system network for wireless communication and information exchange between vehicles and X (vehicles, roads, people and the cloud, etc.) (Vehicle to X, V2X), that is, it can realize real-time online communication between vehicles, vehicles and facilities, vehicles and the cloud, etc.

[0100] It should be noted that the multi-channel audio data involved in the embodiments of the present disclosure are all audio data authorized by the user or fully authorized by all parties. The actions such as obtaining audio data involved in the embodiments of the present disclosure are all performed after authorization by the user, the object, or after full authorization by all parties.

[0101] It should be noted that although the various steps are described in a specific order in the above embodiments, those skilled in the art will appreciate that, in order to achieve the effects of the present application, the different steps do not necessarily have to be performed in this order; they can be performed simultaneously (in parallel) or in other orders, and such variations are within the scope of protection of the present application. Furthermore, all of the above embodiments can be combined in any manner to form optional embodiments of the present application, and will not be described in detail here.

[0102] Furthermore, the present application also provides a speech processing device.

[0103] Referring to Figure 7 , Figure 7 is a schematic block diagram of the main structure of a speech processing device according to an embodiment of the present application. As shown in Figure 7 , the speech processing device in this embodiment of the present application primarily includes an audio data acquisition module 11, an energy value acquisition and comparison module 12, a first processing module 13, and a second processing module 14. In some embodiments, one or more of the audio data acquisition module 11, the energy value acquisition and comparison module 12, the first processing module 13, and the second processing module 14 can be combined into a single module. In some embodiments, the audio data acquisition module 11 can be configured to acquire multiple channels of audio data in the driving device; the energy value acquisition and comparison module 12 can be configured to acquire a first energy value and a second energy value and compare the first energy value with the second energy value when the first channel of audio data triggers the wake-up word for the first time, wherein the first energy value is the energy value corresponding to the first channel of audio data, and the second energy value is the energy value corresponding to the remaining channels of audio data in the multiple channels of audio data; the first processing module 13 can be configured to, if the first energy value is greater than or equal to the second energy value, then there is no need for the remaining channels of audio data to trigger the wake-up word; the second processing module 14 can be configured to, if the first energy value is less than the second energy value, dynamically allocate the wake-up word triggering duration of the remaining channels of audio data.

[0104] In some embodiments, the energy value acquisition and comparison module 12 screens the first channel of audio data and the remaining channels of audio data based on preset rules.

[0105] In some embodiments, the energy value acquisition and comparison module 12 performs preset energy value screening on multiple sampling points of the first audio data and the remaining audio data respectively; after completing the preset energy value screening, the first audio data and the remaining audio data are respectively screened for preset durations.

[0106] In some embodiments, the energy value acquisition and comparison module 12 performs a first preset energy value screening on each sampling point in the first channel of audio data and the remaining channels of audio data, respectively; or, performs a second preset energy value screening on a preset number of sampling points in the first channel of audio data and the remaining channels of audio data, respectively.

[0107] In some embodiments, the energy value acquisition and comparison module 12 filters the first channel of audio data based on the wake-up word to obtain the first channel of audio data containing only the wake-up word; obtains the time area corresponding to the first channel of audio data containing only the wake-up word, and filters the remaining channels of audio data based on the time area.

[0108] In some embodiments, the second processing module 14 obtains the difference between the first energy value and the second energy value; based on the difference, dynamically allocates the wake-up word triggering duration of the remaining audio data.

[0109] In some embodiments, the device also includes a position determination module. When Y-channel audio data among the remaining audio data triggers the wake-up word, the position determination module compares the energy values ​​corresponding to the Y-channel audio data, where Y is greater than 1 and less than or equal to X, and Y is a positive integer; based on the comparison result of the energy values ​​corresponding to the Y-channel audio data, the user's position is determined.

[0110] It will be understood by those skilled in the art that all or part of the processes in the method for implementing the above embodiment of the present application can also be completed by instructing the relevant hardware through a computer program. The computer program can be stored in a computer-readable storage medium. When the computer program is executed by a processor, it can implement the steps of each of the above method embodiments. The computer program includes computer program code, which can be in source code form, object code form, executable file or some intermediate form. The computer-readable storage medium can include: any entity or device, medium, USB flash drive, mobile hard disk, magnetic disk, optical disk, computer memory, read-only memory, random access memory, electric carrier signal, telecommunication signal and software distribution medium that can carry the computer program code. It should be noted that the content contained in the computer-readable storage medium can be appropriately increased or decreased according to the requirements of legislation and patent practice in the jurisdiction. For example, in some jurisdictions, according to legislation and patent practice, computer-readable storage media do not include electric carrier signals and telecommunication signals.

[0111] Furthermore, the present application also provides a computer device.

[0112] Refer to Figure 8, which is a schematic diagram of the main structure of a computer device according to an embodiment of the present application. As shown in Figure 8, the computer device in the embodiment of the present application mainly includes a storage device 21 and a processor 22. The storage device 21 can be configured to store a program for executing the speech processing method of the above-mentioned method embodiment, and the processor 22 can be configured to execute the program in the storage device, which includes but is not limited to a program for executing the speech processing method of the above-mentioned method embodiment. For ease of explanation, only the parts related to the embodiment of the present application are shown. For specific technical details not disclosed, please refer to the method section of the embodiment of the present application.

[0113] In the embodiment of the present application, the computer device may be a control device device formed by various electronic devices. In some possible implementations, the computer device may include multiple storage devices 21 and multiple processors 22. The program for executing the speech processing method of the above-mentioned method embodiment can be divided into multiple subroutines, and each subroutine can be loaded and run by the processor to execute different steps of the speech processing method of the above-mentioned method embodiment. Specifically, each subroutine can be stored in different storage devices 21 respectively, and each processor 22 can be configured to execute the program in one or more storage devices 21 to jointly implement the speech processing method of the above-mentioned method embodiment, that is, each processor 22 executes different steps of the speech processing method of the above-mentioned method embodiment respectively to jointly implement the speech processing method of the above-mentioned method embodiment.

[0114] The multiple processors 22 may be processors deployed on the same device. For example, the computer device may be a high-performance device composed of multiple processors, and the multiple processors 22 may be processors configured on the high-performance device. Furthermore, the multiple processors 22 may be processors deployed on different devices. For example, the computer device may be a server cluster, and the multiple processors 22 may be processors on different servers in the server cluster.

[0115] Furthermore, the present application also provides a computer-readable storage medium. In a computer-readable storage medium embodiment according to the present application, the computer-readable storage medium can be configured to store a program for executing the speech processing method of the above-mentioned method embodiment, and the program can be loaded and run by the processor to implement the above-mentioned speech processing method. For ease of explanation, only the parts related to the embodiment of the present application are shown. For specific technical details not disclosed, please refer to the method part of the embodiment of the present application. The computer-readable storage medium can be a storage device formed by various electronic devices. Optionally, the computer-readable storage medium in the embodiment of the present application is a non-transitory computer-readable storage medium.

[0116] Furthermore, it should be understood that since the configuration of each module is merely for the purpose of illustrating the functional units of the apparatus of the present application, the physical devices corresponding to these modules may be the processor itself, or a portion of the software in the processor, a portion of the hardware, or a combination of software and hardware. Therefore, the number of modules in the figure is merely illustrative.

[0117] Those skilled in the art will appreciate that the various modules in the device can be adaptively split or merged. Such splitting or merging of specific modules will not cause the technical solution to deviate from the principles of this application. Therefore, the technical solutions after splitting or merging will fall within the scope of protection of this application.

[0118] Thus far, the technical solutions of the present application have been described in conjunction with the preferred embodiments shown in the accompanying drawings. However, it is readily understood by those skilled in the art that the scope of protection of the present application is obviously not limited to these specific embodiments. Without departing from the principles of the present application, those skilled in the art may make equivalent changes or substitutions to the relevant technical features, and the technical solutions after such changes or substitutions will fall within the scope of protection of the present application.

Claims

1. A speech processing method, applied to a driving device, characterized in that: The method comprises the following steps: Acquiring multi-channel audio data in the driving device; When the first channel of audio data triggers the wake-up word for the first time, obtaining a first energy value and a second energy value and comparing the first energy value with the second energy value, wherein the first energy value is an energy value corresponding to the first channel of audio data, and the second energy value is an energy value corresponding to the remaining channels of audio data in the multiple channels of audio data; If the first energy value is greater than or equal to the second energy value, the remaining audio data is not required to trigger the wake-up word; If the first energy value is less than the second energy value, the wake-up word trigger duration of the remaining audio data is dynamically allocated.

2. The speech processing method according to claim 1, characterized in that: Before obtaining the first energy value and the second energy value, the method further includes: The first channel of audio data and the remaining channels of audio data are respectively screened based on preset rules.

3. The speech processing method according to claim 2, characterized in that: The first channel of audio data and the remaining channels of audio data respectively include a plurality of sampling points, and the step of "respectively filtering the first channel of audio data and the remaining channels of audio data based on a preset rule" includes: Performing preset energy value screening on multiple sampling points of the first channel of audio data and the remaining channels of audio data respectively; After completing the preset energy value screening, the first channel of audio data and the remaining channels of audio data are screened for preset durations respectively.

4. The speech processing method according to claim 3, characterized in that: The step of "respectively screening the plurality of sampling points of the first audio data and the remaining audio data by preset energy values" includes: Performing first preset energy value screening on each sampling point in the first channel of audio data and the remaining channels of audio data respectively; Alternatively, a second preset energy value screening is performed on preset sampling points in the first channel of audio data and the remaining channels of audio data respectively.

5. The speech processing method according to claim 2, characterized in that: Before respectively screening the first channel of audio data and the remaining channels of audio data based on a preset rule, the method further includes: Filtering the first channel of audio data based on the wake-up word to obtain the first channel of audio data containing only the wake-up word; The time zone corresponding to the first channel of audio data containing only the wake-up word is obtained, and the remaining channels of audio data are filtered based on the time zone.

6. The speech processing method according to claim 1, characterized in that: The step of "dynamically allocating the wake-up word trigger duration of the remaining audio data" includes: Obtaining a difference between the first energy value and the second energy value; Based on the difference, the wake-up word triggering duration of the remaining audio data is dynamically allocated.

7. The speech processing method according to claim 1, characterized in that: The remaining audio data includes X audio data, where X is greater than 1 and is a positive integer. After the wake-up word trigger duration of the remaining audio data is dynamically allocated, the method further includes: When Y channels of audio data among the remaining channels of audio data trigger the wake-up word, compare the energy values ​​corresponding to the Y channels of audio data, where Y is greater than 1 and less than or equal to X, and Y is a positive integer; Based on the comparison result of the energy values ​​corresponding to the Y-channel audio data, the user's location is determined.

8. A speech processing device, applied to driving equipment, characterized in that: The device comprises: An audio data acquisition module, configured to acquire multi-channel audio data in the driving device; An energy value acquisition and comparison module is configured to acquire a first energy value and a second energy value and compare the first energy value with the second energy value when the first channel of audio data triggers the wake-up word for the first time, wherein the first energy value is an energy value corresponding to the first channel of audio data, and the second energy value is an energy value corresponding to the remaining channels of audio data in the multiple channels of audio data; A first processing module is configured to trigger the wake-up word without using the remaining audio data if the first energy value is greater than or equal to the second energy value; The second processing module is configured to dynamically allocate the wake-up word triggering duration of the remaining audio data if the first energy value is less than the second energy value.

9. A computer device comprising a processor and a storage device, wherein the storage device is suitable for storing a plurality of program codes, characterized in that: The program code is suitable for being loaded and run by the processor to execute the speech processing method according to any one of claims 1 to 7.

10. A computer-readable storage medium storing a plurality of program codes, characterized in that: The program code is suitable for being loaded and run by a processor to execute the speech processing method according to any one of claims 1 to 7.

Citation Information

Patent Citations

  • Driver sound localization system and method for automobile

    CN102819009A

  • Always-on Low-power Keyword Spotting

    CN104049707A

  • Voice control method, device and system, vehicle and storage medium

    CN115346527A

  • Voice processing method and device, equipment and medium

    CN117542361A

  • Wake-word processing in an electronic device

    WO2023196695A1