Voice recognition method and electronic equipment

CN120937076APending Publication Date: 2025-11-11HONOR DEVICE CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202380096236.7
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2023-07-10
Publication Date
2025-11-11

AI Technical Summary

Technical Problem

In the audio scenario, the simultaneous call of multiple sound recognition modules results in an increase in power consumption of electronic devices and cannot effectively reduce power consumption.

Method used

By running the AAD module in ADSP, the module associates multiple sound recognition modules, extracts the audio features of the audio signal, and calls the target sound recognition module based on the features instead of calling all modules at the same time.

Benefits of technology

It achieves the effect of reducing the power consumption of electronic equipment in the audio normally on scenario, while improving the accuracy and efficiency of sound recognition.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120937076A_ABST
    Figure CN120937076A_ABST
Patent Text Reader

Abstract

The invention discloses a sound recognition method and electronic equipment, relates to the technical field of electronics, and is used for reducing the power consumption of the electronic equipment with a plurality of sound recognition modules in an audio normally-open scene. The method is applied to an audio digital signal processor, a sound activity detection module runs in the audio digital signal processor, the sound activity detection module is associated with a plurality of sound recognition modules, and the method comprises the following steps: acquiring an audio signal corresponding to sound acquired by a microphone; extracting audio features of the audio signals; calling a target voice recognition module based on the audio features, and performing voice recognition on the audio signals; the target voice recognition module is a voice recognition module corresponding to an audio feature range where the audio features are located in the plurality of voice recognition modules, and the plurality of audio feature ranges correspond to the plurality of voice recognition modules.
Need to check novelty before this filing date? Find Prior Art

Description

Voice recognition method and electronic device Technical Field

[0001] The embodiments of the present application relate to the field of electronic technology, and in particular to a sound recognition method and electronic device. Background Art

[0002] In the scenario of always-on audio (AON Audio), the power consumption (or power consumption) requirements of electronic devices are very strict. In the scenario of always-on audio, the voice recognition module of the electronic device is in the always-on state. When the voice recognition module runs for a long time, it will occupy more memory and processor computing power, resulting in higher power consumption of the electronic device. In order to reduce the power consumption of the electronic device, the voice recognition module of the electronic device is not in the always-on state. The electronic device usually runs a low-power acoustic active detection (AAD) module. When the AAD module detects that the audio signal meets the conditions for voice recognition, it calls the voice recognition module, and the voice recognition module performs voice recognition on the audio signal, thereby reducing the power consumption of the electronic device.

[0003] However, some electronic devices may include multiple sound recognition modules for performing different sound recognition functions, such as keyword spotting (KWS), acoustic context detection (ACD), and acoustic scene classification (ASC). When the AAD module detects that an audio signal meets the conditions for sound recognition, it calls all sound recognition modules to perform sound recognition on the audio signal, including sound recognition modules that do not need to be called, resulting in a poor effect of reducing power consumption.

[0004] Summary of the Invention

[0005] Embodiments of the present application provide a voice recognition method and an electronic device for reducing power consumption of an electronic device having multiple voice recognition modules in a scenario where the audio is always on.

[0006] To achieve the above objectives, the embodiments of the present application adopt the following technical solutions:

[0007] In a first aspect, a sound recognition method is provided, which is applied to an audio digital signal processor (ADSP), wherein an AAD module is run in the ADSP, the AAD module is associated with multiple sound recognition modules, and the multiple sound recognition modules correspond to multiple audio feature ranges. The method includes: obtaining a first audio signal corresponding to a first sound collected by a microphone; extracting audio features of the first audio signal; calling a target sound recognition module based on the audio features of the first audio signal to perform sound recognition on the first audio signal; the target sound recognition module is a sound recognition module corresponding to an audio feature range where the audio features of the first audio signal are located among multiple sound recognition modules, and the multiple audio feature ranges correspond to the multiple sound recognition modules; obtaining a second audio signal corresponding to a second sound collected by the microphone; extracting audio features of the second audio signal; and not running any sound recognition module among the multiple sound recognition modules based on the audio features of the second audio signal, wherein the audio feature range where the audio features of the second audio signal are located has no corresponding sound recognition module.

[0008] The sound recognition method provided by the present application is executed by ADSP. First, the AAD module in the ADSP obtains the audio signal corresponding to the sound collected by the microphone, and extracts the audio features of the audio signal. Then, by comparing the audio features of the audio signal with multiple audio feature ranges, the sound recognition module corresponding to the audio feature range where the audio features of the audio signal are located is determined from multiple sound recognition modules, that is, the target sound recognition module, and the target sound recognition module is called. Since the audio features of the audio signal belong to different audio feature ranges under different sound sources and different audio signal collection scenarios, the suitable sound recognition modules are also different. Therefore, the target sound recognition module that is suitable for the audio features of the audio signal can be screened out from multiple sound recognition modules, and the target sound recognition module is called to perform sound recognition on the audio signal to obtain a recognition result, instead of calling all sound recognition modules to perform sound recognition on the audio signal. Alternatively, if the audio feature range to which the audio features of the audio signal belong is not found, no sound recognition module is run. This reduces the power consumption of electronic devices with multiple sound recognition modules in audio-on scenarios.

[0009] That is, in this application, in the audio always-on scenario, the low-power sound recognition module running in the ADSP is in the always-on state, and the multiple sound recognition modules associated with the AAD module are not called. When the sound recognition module detects that the audio features of the audio signal match a certain sound recognition module (i.e., the target sound recognition module), the target sound recognition module is called, and other sound recognition modules other than the target sound recognition module are still not called, or, if there is no audio feature of the audio signal that falls within the audio feature range, no sound recognition module will be run. Therefore, reducing the computing power of the ADSP when performing sound recognition can reduce the power consumption of the ADSP, and thereby reduce the power consumption of electronic devices with multiple sound recognition modules in the audio always-on scenario.

[0010] In one possible embodiment, the electronic device further includes an application processor (AP); after calling a target sound recognition module based on the audio features of the first audio signal and performing sound recognition on the audio signal, the method further includes: if the recognition result of the target sound recognition module meets the AP wake-up condition, waking up the AP in a dormant state; otherwise, not waking up the AP in a dormant state.

[0011] That is to say, if the target sound recognition module performs sound recognition on the audio signal and obtains a recognition result that meets the AP's wake-up condition, the AP in the dormant state will be awakened; otherwise, the AP in the dormant state will not be awakened, thereby reducing the power consumption of the AP, and further reducing the power consumption of electronic devices with multiple sound recognition modules in the audio-always-on scenario.

[0012] In one possible embodiment, the audio features of the first audio signal include the center frequency and sound pressure level of a preset frequency band of the first audio signal. Each audio feature range includes a frequency band and a sound pressure level range. The above-mentioned calling of the target sound recognition module based on the first audio feature includes: based on the center frequency and sound pressure level of the preset frequency band of the first audio signal, searching for the first audio feature range from multiple audio feature ranges, wherein the frequency band in the first audio feature range includes the center frequency of the preset frequency band of the first audio signal, and the sound pressure level range in the first audio feature range includes the sound pressure level of the preset frequency band of the first audio signal; after finding the first audio feature range, calling the target sound recognition module corresponding to the first audio feature range.

[0013] The sound pressure level of an audio signal can, to a certain extent, distinguish the scenes it represents, and can be used as an audio feature of the audio signal. Furthermore, even if the sound pressure level of an audio signal is the same, the sound sources can vary significantly. However, the center frequencies of audio signals from different sound sources can be different. Therefore, in addition to using the sound pressure level of an audio signal as an audio feature, the center frequency of an audio signal can also be used as an audio feature of the audio signal.

[0014] The frequency band in the audio feature range can be represented by a frequency interval. If the center frequency of the audio signal is within the frequency interval, then the center frequency of the audio signal is determined to be within the frequency band in the audio feature range. The sound pressure level range in the audio feature range can be represented by a sound pressure level interval or a sound pressure level threshold. When the sound pressure level range in the audio feature range is represented by a sound pressure level interval, if the sound pressure level of the audio signal is within the sound pressure level interval, then the sound pressure level of the audio signal is determined to be within the sound pressure level range in the audio feature range. When the sound pressure level range in the audio feature range is represented by a sound pressure level threshold, the meaning of the sound pressure level range is the sound pressure level greater than the sound pressure level threshold. That is, if the sound pressure level of the audio signal is greater than the sound pressure level threshold, then the sound pressure level of the audio signal is determined to be within the sound pressure level range in the audio feature range.

[0015] In one possible implementation, extracting audio features of a first audio signal includes: performing bandpass filtering on the first audio signal within a preset frequency band to obtain components of the first audio signal within the preset frequency band; calculating the geometric mean of the upper and lower frequency limits of the components of the first audio signal within the preset frequency band, where the center frequency of the preset frequency band of the first audio signal is the geometric mean; obtaining the amplitude h corresponding to the center frequency of the components of the preset frequency band of the first audio signal, and calculating the sound pressure p of the first audio signal within the preset frequency band based on the amplitude h. Specifically, sound pressure p = amplitude h * standard atmospheric pressure. Then, based on the sound pressure p of the first audio signal within the preset frequency band, calculating the sound pressure level L of the first audio signal within the preset frequency band. Specifically, sound pressure level L = 20log10(p / p0), where p0 is a reference sound pressure.

[0016] In one possible implementation, extracting audio features of the second audio signal includes: performing bandpass filtering on the second audio signal within a preset frequency band to obtain components of the second audio signal within the preset frequency band; calculating the geometric mean of the upper and lower frequency limits of the components of the second audio signal within the preset frequency band, where the center frequency of the preset frequency band of the second audio signal is the geometric mean; obtaining the amplitude h corresponding to the center frequency of the components of the second audio signal within the preset frequency band, and calculating the sound pressure p of the second audio signal within the preset frequency band based on the amplitude h. Specifically, sound pressure p = amplitude h * standard atmospheric pressure. Then, based on the sound pressure p of the second audio signal within the preset frequency band, calculating the sound pressure level L of the second audio signal within the preset frequency band. Specifically, sound pressure level L = 20log10(p / p0), where p0 is a reference sound pressure.

[0017] In one possible embodiment, the audio features of the second audio signal include the center frequency and sound pressure level of a preset frequency band of the second audio signal, and each audio feature range includes a frequency band and a sound pressure level range; and not running any of the multiple sound recognition modules based on the audio features of the second audio signal includes: searching for the second audio feature range from the multiple audio feature ranges based on the center frequency and sound pressure level of the second audio signal, wherein the frequency band in the second audio feature range includes the center frequency of the second audio signal, and the sound pressure level range in the second audio feature range includes the sound pressure level of the second audio signal; and not running any of the multiple sound recognition modules if the second audio feature range is not found. This saves computing power and reduces power consumption.

[0018] In one possible embodiment, the audio features of the first audio signal are filter bank (FBank) features of the first audio signal obtained by extracting filter bank (FBank) features from the first audio signal, and the audio features of the first audio signal are FBank features. Each audio feature range is an FBank feature range; calling a target sound recognition module based on the audio features of the first audio signal includes: searching for a third audio feature range that includes the FBank features of the first audio signal from multiple audio feature ranges; and calling a target sound recognition module corresponding to the third audio feature range.

[0019] FBank feature extraction extracts audio features in a manner similar to that of the human ear (essentially reflecting the frequency characteristics of the audio), which is equivalent to simulating the sound recognition method of the human ear and can improve the performance of audio feature extraction.

[0020] Among them, the FBank feature may include a vector composed of multiple FBank coefficients, and the FBank feature range may be represented by a FBank feature threshold or a FBank feature interval. When the FBank feature range is represented by a FBank feature threshold, the FBank feature range includes multiple FBank feature thresholds, and the multiple FBank feature thresholds correspond to multiple FBank coefficients. The multiple FBank feature thresholds may include valid FBank feature thresholds and invalid FBank feature thresholds, wherein the valid FBank feature threshold may refer to a FBank feature threshold with a non-null value, and the invalid FBank feature threshold may refer to a null value. For each FBank coefficient corresponding to a valid FBank feature threshold, if the FBank coefficient is greater than its corresponding valid FBank feature threshold, it is determined that the FBank feature of the audio signal is within the FBank feature range. It should be understood that in the present application, for an invalid FBank feature threshold, it can be assumed that the FBank coefficient corresponding to the invalid FBank feature threshold meets the requirements.

[0021] When the FBank feature range is represented by an FBank feature interval, the FBank feature range includes multiple FBank feature intervals, and the multiple FBank feature intervals correspond to multiple FBank coefficients. The multiple FBank feature intervals may include valid FBank feature intervals, and may also include invalid FBank feature intervals, wherein a valid FBank feature interval may refer to a FBank feature interval with a non-empty value, and an invalid FBank feature interval may refer to a null value. For each FBank coefficient corresponding to a valid FBank feature interval, if the FBank coefficient is within its corresponding valid FBank feature interval, it is determined that the FBank feature of the audio signal is within the FBank feature range. It should be understood that in the present application, for an invalid FBank feature interval, it can be assumed that the MFCC coefficient corresponding to the invalid FBank feature interval meets the requirements.

[0022] In one possible implementation, the audio feature of the second audio signal is an FBank feature of the second audio signal obtained by performing FBank feature extraction on the second audio signal, where each audio feature range is an FBank feature range; not running any of the multiple sound recognition modules based on the audio feature of the second audio signal includes: searching for a fourth audio feature range that includes the FBank feature of the second audio signal from the multiple audio feature ranges; and not running any of the multiple sound recognition modules if the fourth audio feature range is not found. This saves computing power and reduces power consumption.

[0023] In one possible implementation, the audio features of the first audio signal are MFCC features of the first audio signal obtained by extracting Mel frequency cepstral coefficients (MFCC) features from the audio signal. Each audio feature range is an MFCC feature range. Calling a target sound recognition module based on the audio features of the first audio signal includes: searching for a fifth audio feature range that includes the MFCC features from multiple audio feature ranges; calling a target sound recognition module corresponding to the fifth audio feature range, and having the target sound recognition module perform sound recognition on the audio signal.

[0024] MFCC feature extraction takes the logarithm of FBank features and then performs a discrete cosine transform (DCT) to obtain MFCC features in vector form. MFCC feature extraction filters out the correlation between adjacent FBank coefficients, resulting in audio features with better recognition.

[0025] Among them, the MFCC feature may include a vector composed of multiple MFCC coefficients, and the MFCC feature range may be represented by an MFCC feature threshold or an MFCC feature interval. When the MFCC feature range is represented by an MFCC feature threshold, the MFCC feature range includes multiple MFCC feature thresholds, and the multiple MFCC feature thresholds correspond to multiple MFCC coefficients. The multiple MFCC feature thresholds may include valid MFCC feature thresholds and invalid MFCC feature thresholds, wherein the valid MFCC feature threshold may refer to an MFCC feature threshold with a non-empty value, and the invalid MFCC feature threshold may refer to an empty value. For each MFCC coefficient corresponding to the valid MFCC feature threshold, if the MFCC coefficient is greater than its corresponding valid MFCC feature threshold, it is determined that the MFCC feature of the audio signal is within the MFCC feature range. It should be understood that in this application, for invalid MFCC feature thresholds, it can be assumed that the MFCC coefficient corresponding to the invalid MFCC feature threshold meets the requirements.

[0026] When the MFCC feature range is represented by an MFCC feature interval, the MFCC feature range includes multiple MFCC feature intervals, and the multiple MFCC feature intervals correspond to multiple MFCC coefficients. The multiple MFCC feature intervals may include valid MFCC feature intervals and may also include invalid MFCC feature intervals, wherein a valid MFCC feature interval may refer to an MFCC feature interval with a non-empty value, and an invalid MFCC feature interval may refer to a null value. For each MFCC coefficient corresponding to a valid MFCC feature interval, if the MFCC coefficient is within its corresponding valid MFCC feature interval, it is determined that the MFCC features of the audio signal are within the MFCC feature range. It should be understood that in the present application, for an invalid MFCC feature interval, it can be assumed that the MFCC coefficient corresponding to the invalid MFCC feature interval meets the requirements, and the MFCC coefficient is not compared with the MFCC feature interval in the MFCC feature range.

[0027] In one possible implementation, the audio features of the second audio signal are MFCC features of the second audio signal obtained by performing MFCC feature extraction on the second audio signal, where each audio feature range is an MFCC feature range; and not running any of the multiple sound recognition modules based on the audio features of the second audio signal includes: searching, from the multiple audio feature ranges, for a sixth audio feature range that includes the MFCC features of the second audio signal; and not running any of the multiple sound recognition modules if the sixth audio feature range is not found. This saves computing power and reduces power consumption.

[0028] In one possible embodiment, the target sound recognition module includes a first sound recognition module and a second sound recognition module. The method further includes: obtaining a first recognition result of the first sound recognition module recognizing an audio signal; and transmitting the first recognition result to the second sound recognition module, whereby the second sound recognition module performs sound recognition on the audio signal based on the first recognition result. The AAD module can control the two sound recognition modules, forwarding the recognition results of the sound recognition modules via the AAD module, thereby assisting the sound recognition module receiving the recognition results to more accurately recognize the audio signal.

[0029] In one possible embodiment, the target sound recognition module includes a first sound recognition module and a second sound recognition module, and the method further includes: sending a first indication message to the first sound recognition module, the first indication message being used to instruct the first sound recognition module to send a first recognition result of recognizing an audio signal to the second sound recognition module. The AAD module can control the two sound recognition modules, and under the instruction of the AAD module, one of the sound recognition modules obtains the recognition result of the other sound recognition module, thereby assisting the sound recognition module receiving the recognition result to more accurately perform sound recognition on the audio signal. In this embodiment, the second sound recognition module passively receives the first recognition result from the first sound recognition module.

[0030] In one possible embodiment, the target sound recognition module includes a first sound recognition module and a second sound recognition module, and the method further includes: sending a second indication message to the second sound recognition module, the second indication message being used to instruct the second sound recognition module to obtain a first recognition result of recognizing an audio signal from the first sound recognition module. The AAD module can control the two sound recognition modules, and under the instruction of the AAD module, one of the sound recognition modules obtains the recognition result of the other sound recognition module, thereby assisting the sound recognition module receiving the recognition result to more accurately recognize the audio signal. In this embodiment, the second sound recognition module actively obtains the first recognition result from the first sound recognition module.

[0031] In a possible implementation, the method further includes: sending audio features to a target sound recognition module; wherein the audio features are used by the target sound recognition module to perform sound recognition on the audio signal, and can assist the target sound recognition module in performing sound recognition on the audio signal.

[0032] In a second aspect, a chip system is provided, including: an ADSP and an interface circuit; the interface circuit is used to receive an audio signal; the ADSP is used to execute the method as described in the first aspect and any embodiment thereof, an audio AAD module is run in the ADSP, and the AAD module is associated with multiple sound recognition modules; the AAD module is used to obtain a first audio signal corresponding to a first sound collected by a microphone; extract the audio features of the first audio signal; call a target sound recognition module based on the audio features of the first audio signal to perform sound recognition on the first audio signal; the target sound recognition module is a sound recognition module corresponding to an audio feature range where the audio features of the first audio signal are located among multiple sound recognition modules, and multiple audio feature ranges correspond to multiple sound recognition modules; obtain a second audio signal corresponding to a second sound collected by the microphone; extract the audio features of the second audio signal; based on the audio features of the second audio signal, no sound recognition module among the multiple sound recognition modules is run, wherein the audio feature range where the audio features of the second audio signal are located has no corresponding sound recognition module.

[0033] In one possible embodiment, the chip system also includes an AP; the target sound recognition module is used to wake up the AP in a dormant state if the recognition result of the sound recognition of the first audio signal meets the wake-up condition of the application processor; if the recognition result of the sound recognition of the first audio signal does not meet the wake-up condition of the application processor, then the AP in a dormant state is not woken up.

[0034] In a third aspect, an electronic device is provided, comprising: a microphone, an ADSP, and a memory; the ADSP is coupled to the microphone and the memory; computer code is stored in the memory, and the code includes instructions. When the ADSP executes the instructions, the electronic device executes the method described in the first aspect and any embodiment thereof.

[0035] In a possible implementation, an AP is further included, and the ADSP is coupled to the AP. The ADSP is used to wake up the AP in a dormant state.

[0036] In a fourth aspect, a computer-readable storage medium is provided, comprising instructions, which, when executed on an electronic device, enable the electronic device to execute the method as described in the first aspect and any embodiment thereof.

[0037] In a fifth aspect, a computer program product comprising instructions is provided. When the instructions are executed on the electronic device, the electronic device executes the method as described in the first aspect and any embodiment thereof.

[0038] The technical effects of the second to fifth aspects refer to the technical effects of the first aspect and any of its embodiments and will not be repeated here. BRIEF DESCRIPTION OF THE DRAWINGS

[0039] FIG1 is a schematic diagram of a Mel filter provided in an embodiment of the present application;

[0040] FIG2 is a schematic diagram of an application scenario of an electronic device provided in an embodiment of the present application;

[0041] FIG3 is a schematic diagram of an embodiment of the present application in which each voice recognition module is provided with an AAD module;

[0042] FIG4 is a schematic diagram of a method of multiple voice recognition modules multiplexing an AAD module according to an embodiment of the present application;

[0043] FIG5 is a schematic diagram of a core concept of the present application provided in an embodiment of the present application;

[0044] FIG6 is a schematic structural diagram of an electronic device provided in an embodiment of the present application;

[0045] FIG7 is a schematic diagram of a software structure in an electronic device provided in an embodiment of the present application;

[0046] FIG8 is a flow chart of a sound recognition method according to an embodiment of the present application;

[0047] FIG9 is a schematic diagram of an audio feature including the center frequency and sound pressure level of an audio signal provided by an embodiment of the present application;

[0048] FIG10 is a schematic diagram of an audio feature including an FBank feature of an audio signal provided by an embodiment of the present application;

[0049] FIG11 is a schematic diagram showing an audio feature including Mel frequency cepstral coefficients (MFCC) features of an audio signal provided by an embodiment of the present application;

[0050] FIG12 is a schematic diagram showing an FBank feature range represented by an FBank feature threshold according to an embodiment of the present application;

[0051] FIG13 is a schematic diagram showing an FBank feature range represented by an FBank feature interval according to an embodiment of the present application;

[0052] FIG14 is a schematic diagram showing an MFCC feature range represented by an MFCC feature threshold according to an embodiment of the present application;

[0053] FIG15 is a schematic diagram showing an MFCC feature range represented by an MFCC feature interval according to an embodiment of the present application;

[0054] FIG16 is a schematic diagram of a first embodiment of the present application, in which a voice recognition module obtains a recognition result of another voice recognition module;

[0055] FIG17 is a schematic diagram of a second embodiment of the present application, in which a voice recognition module obtains a recognition result of another voice recognition module;

[0056] FIG18 is a schematic diagram of a third embodiment of the present application, in which a voice recognition module obtains a recognition result of another voice recognition module;

[0057] FIG19 is a schematic structural diagram of a chip system provided in an embodiment of the present application. DETAILED DESCRIPTION

[0058] The terms "first", "second", etc. involved in the embodiments of the present application are only used to distinguish features of the same type and cannot be understood as indicating relative importance, quantity, order, etc.

[0059] The terms "exemplary" or "for example" in the embodiments of this application are used to indicate examples, illustrations, or descriptions. Any embodiment or design described in this application as "exemplary" or "for example" should not be construed as being preferred or advantageous over other embodiments or designs. Rather, the use of words such as "exemplary" or "for example" is intended to present the relevant concepts in a concrete manner.

[0060] The terms "coupling" and "connection" involved in the embodiments of this application should be understood in a broad sense. For example, they may refer to a physical direct connection, or an indirect connection achieved through electronic devices, such as a connection achieved through resistors, inductors, capacitors or other electronic devices.

[0061] The term "module" as used in the embodiments of the present application may refer to a functional module implemented in software form, such as program code, algorithm implementation, etc. For example, an AAD module and a voice recognition module may refer to a functional module implemented in software form, such as program code, algorithm implementation, etc.

[0062] First, some concepts involved in this application are described.

[0063] A bandpass filter (BPF) allows audio signals within a preset frequency band to pass through while blocking audio signals in other frequency bands. After the audio signal passes through the BPF within the preset frequency band, the audio signal within the preset frequency band is obtained. It should be understood that the preset frequency band can be understood as the filter frequency band of the filter and can be pre-configured.

[0064] Filter bank (FBank) feature extraction: can refer to extracting FBank features from audio signals. Among them, the process of extracting FBank features from audio signals generally includes: pre-emphasis, framing, windowing, short-time Fourier transform (STFT), Mel filtering, de-averaging, etc. of the audio signal. The method of the embodiment of the present application involves Mel filtering. The following introduces the relevant content of Mel filtering to facilitate understanding of the method provided by the embodiment of the present application.

[0065] Mel filtering can be implemented using multiple filters. The filters used in Mel filtering are called Mel filters. Each Mel filter has a corresponding filtering frequency band. Different Mel filters have different frequency ranges and widths. The closer a Mel filter's corresponding filtering frequency band is to the low-frequency region, the narrower the corresponding filtering frequency band is. The closer a Mel filter's corresponding filtering frequency band is to the high-frequency region, the wider the corresponding filtering frequency band is.

[0066] For example, a Mel filter can include a set of M triangular filters. Assume that the center frequency of the mth triangular filter is f(m), and the frequency response of the mth triangular filter is Hm(k), where k represents discrete sampling and m = 1, 2, ..., M. Please refer to Figure 1, which shows a schematic diagram of the filter frequency bands and frequency responses corresponding to each triangular filter in the set of triangular filters.

[0067] As shown in Figure 1, the filtering frequency band of triangular filter 1 is f(0)-f(2), and the filtering frequency band of triangular filter 2 is f(1)-f(3). Compared with triangular filter 2, the filtering frequency band of triangular filter 1 is closer to the low-frequency region; compared with triangular filter 2, the filtering frequency band of triangular filter 1 is narrower.

[0068] As shown in Figure 1, the filtering frequency band of triangular filter 3 is f(2)-f(4), and the filtering frequency band of triangular filter 4 is f(2)-f(4). Compared with triangular filter 3, the filtering frequency band of triangular filter 4 is closer to the low-frequency region; compared with triangular filter 4, the filtering frequency band of triangular filter 3 is narrower.

[0069] The human ear's sound filtering principle is similar to that of a set of filters (i.e., multiple filters). Specifically, the human ear can only hear sounds within a few preset frequency bands (equivalent to the filtering bands); each filter in a set of filters also only allows audio signals within a single preset frequency band (equivalent to the filtering band) to pass through. A set of filters, when combined, can allow audio signals from several preset frequency bands to pass through.

[0070] As mentioned above, the filtering principle of the set of filters (i.e., multiple filters) used in FBank feature extraction is similar to the filtering principle of the human ear. Therefore, FBank feature extraction can extract features from audio signals in a manner similar to that of the human ear, which is equivalent to simulating the sound recognition method of the human ear and can improve the performance of audio feature extraction.

[0071] It can be seen that the FBank feature extraction based on Mel filtering can simulate the principle of human ear filtering sound, thereby extracting audio features of audio signals suitable for human ear perception.

[0072] The FBank feature extracted from the audio signal may include M FBank coefficients. Each FBank coefficient corresponds to the result of filtering the audio signal with a triangular filter, so the FBank feature can be represented by an M-dimensional vector including M FBank coefficients, for example<FBank1,FBank2...FBankM> Since the filtering frequency bands of two adjacent triangular filters in FBank overlap, the correlation between two adjacent FBank coefficients (such as FBank1 and FBank2) in the obtained FBank features is high.

[0073] Mel frequency cepstral coefficients (MFCC) feature extraction: By taking the logarithm of the M FBank coefficients in the FBank feature and then performing discrete cosine transform (DCT), the MFCC feature consisting of M MFCC coefficients can be obtained. That is, the MFCC feature can be represented by an M-dimensional vector consisting of M MFCC coefficients, for example<MFCC1,MFCC2...MFCCM> Compared with FBank feature extraction, MFCC feature extraction filters out the correlation between two adjacent FBank coefficients in the FBank feature, and has better recognition when used for sound feature extraction.

[0074] Keyword spotting (KWS): used to detect whether a preset sentence (i.e., wake-up word) is included in a piece of audio. KWS is often used for voice wake-up of electronic devices that are in sleep mode. For example, in order to reduce power consumption, when users are not using electronic devices such as mobile phones and smart speakers, these electronic devices are usually in sleep mode, but the audio is still kept on to facilitate voice wake-up by the user. When the user says a preset sentence (for example, "Hello, YOYO"), the electronic device detects through KWS that the user has said the preset sentence, and then resumes from sleep mode to working mode.

[0075] Acoustic context detection (ACD): This function analyzes audio signals to identify the event (i.e., the sound source category) corresponding to the audio signal. For example, as shown in FIG2 , electronic device 100 can analyze the audio signal and identify the event category as a knock on the door, a moving vehicle, a male or female voice, a dog barking, or a cat meowing.

[0076] Acoustic scene classification (ASC): This function analyzes audio signals to identify the scene in which the electronic device was located when the sound corresponding to the audio signal was collected. For example, as shown in FIG2 , the electronic device 100 can analyze the audio signal to detect whether the sound corresponding to the audio signal is a sound in a traffic scene (e.g., subway or car sounds), an indoor scene, or an outdoor scene (e.g., a stadium).

[0077] Acoustic activity detection (AAD): It is used to distinguish noise and valid audio (such as speech, object sounds, animal sounds, etc.) in an audio signal.

[0078] In the embodiments of this application, KWS, ASC, ACD, etc. are collectively referred to as voice recognition. These KWS, ASC, and ACD are often implemented as software modules in electronic devices such as mobile phones and smart speakers. For example, these KWS, ASC, and ACD can be implemented as software code and run in the ADSP of an electronic device.

[0079] In the scenario where audio is always on, the power consumption requirements of electronic devices are very strict. When these sound recognition modules are running on ADSP, they will occupy a lot of memory and processor computing power. If they are kept on all the time, the power consumption of electronic devices will be higher. Therefore, when the electronic device is in standby mode, not only is the AP in sleep mode to reduce the power consumption of the electronic device, but also the AAD module that occupies very little memory and processor computing power can be run in ADSP. When the AAD module detects that the audio signal meets the conditions for sound recognition (which may contain valid information, such as voice, object sound, animal sound, etc.), these sound recognition modules will be called for sound recognition, thereby reducing the power consumption of the electronic device.

[0080] In one possible implementation, an AAD module may be provided for each voice recognition module to call the corresponding voice recognition module. For example, as shown in FIG3 , an AAD module 21 is provided for the KWS module 24 to call the KWS module 24, an AAD module 22 is provided for the ASC module 25 to call the ASC module 25, and an AAD module 23 is provided for the ACD module 26 to call the ACD module 26. However, running multiple AAD modules will also increase the power consumption of the electronic device, making it difficult to achieve the goal of reducing power consumption.

[0081] Based on this, another possible implementation method is to reduce power consumption by multiple voice recognition modules reusing one AAD module. For example, as shown in Figure 4, the AAD module 31 can be associated with the KWS module 32, the ASC module 33, and the ACD module 34. However, in the AAD module reuse scheme, based on the same calling conditions, the AAD module 31 will call all associated voice recognition modules, including voice recognition modules that do not need to be called, which is also difficult to achieve the purpose of reducing power consumption. Moreover, the same calling conditions cannot be adjusted for each voice recognition module, so it will affect the accuracy of voice recognition.

[0082] In order to solve the above problems, an embodiment of the present application provides a sound recognition method, which can be applied to the ADSP of an electronic device. The electronic device includes an ADSP and an AP, and an AAD module is running in the ADSP. The AAD module is associated with multiple sound recognition modules (such as a KWS module, an ACD module, an ASC module, etc.), and multiple sound recognition modules correspond to multiple audio feature ranges. After the AAD module obtains the audio signal, it extracts the audio features of the audio signal; based on the audio features of the audio signal, it calls the sound recognition module (i.e., the target sound recognition module) corresponding to the audio feature range where the audio features of the audio signal are located from multiple sound recognition modules. The target sound recognition module performs sound recognition on the audio signal to obtain a recognition result. If the recognition result meets the wake-up condition of the AP, the target sound recognition module wakes up the AP in the dormant state, and reduces the power consumption of the electronic device with multiple sound recognition modules in the audio-on scenario by reducing the power consumption of the ADSP and the power consumption of the AP.

[0083] It should be noted that this application takes ADSP as an example, but is not intended to be limited to this. ADSP can also be replaced by a processor with audio processing capabilities such as a system control processor (SCP).

[0084] As shown in Figure 5, multiple sound recognition modules in ADSP 51 (for example, sound recognition module 1-sound recognition module N) can reuse the same AAD module. In addition, in the method provided in the embodiment of the present application, corresponding audio feature ranges (for example, audio feature range 1-audio feature range N) can also be configured for multiple sound recognition modules respectively. After obtaining the audio signal corresponding to the sound collected by the microphone, the AAD module can extract the audio features of the audio signal, compare the audio features of the audio signal with each audio feature range, and call the corresponding sound recognition module (that is, the target sound recognition module) if the audio feature is within a certain audio feature range. In addition, if the target sound recognition module performs sound recognition on the audio signal and the recognition result obtains that meets the AP wake-up condition, the AP 52 in the dormant state is awakened; otherwise, the AP 52 in the dormant state is not awakened.

[0085] By adopting this solution, first, only one AAD module needs to be run instead of multiple AAD modules; second, corresponding audio feature ranges can be configured for different sound recognition modules, so that the sound recognition module can be called more accurately without calling unnecessary sound recognition modules. Moreover, when the called sound recognition module performs sound recognition and the recognition result meets the AP wake-up condition, the AP in dormant state is woken up, thereby reducing the power consumption of electronic devices in the audio-on scenario.

[0086] For example, the AAD module described in the embodiments of the present application may also be referred to as a voice activity detection (VAD) module, a sound wake-up module, etc., and the specific name is not limited.

[0087] When the electronic device is powered on, the ADSP can run the AAD module, and even if the AP enters sleep mode, the AAD module remains in a constant-on state. The voice recognition module is selectively called by the AAD module based on the audio signal.

[0088] The electronic device provided in the embodiments of the present application may be a handheld device or a vehicle-mounted device with a microphone, such as a mobile phone, a tablet, a smart speaker, a laptop computer, a PDA, a mobile internet device (MID), a virtual reality (VR) device, an augmented reality (AR) device, a wireless device in industrial control, a wireless device in self-driving, a wireless device in remote medical surgery, a wireless device in a smart grid, a wireless device in transportation safety, a wireless device in a smart city, a wireless device in a smart home, a cellular phone, a handheld device with wireless communication capabilities, a personal computer (PC), a computing device or other processing device connected to a wireless modem, a wearable device, a terminal device in a 5G network, or a terminal device in a future-evolved public land mobile network (PLMN), etc. The embodiments of the present application are not limited thereto.

[0089] As shown in FIG6 , the electronic device 100 provided in an embodiment of the present application includes an AP 1011, an ADSP 1012, a memory 102, an AD decoder 103, and a MIC 104. The AP 1011 and the ADSP 1012 may be integrated into a system on chip (SOC) 101. In other embodiments, the AD decoder 103 may also be integrated into the SOC 101, which is not limited in this embodiment of the present application.

[0090] The microphone (MIC) 104 is used to collect sound to obtain a corresponding analog audio signal and output it to the AD decoder 103. The AD decoder 103 is used to perform AD conversion on the analog audio signal to obtain a digital audio signal and output it to the ADSP 1012.

[0091] Memory 102 may be volatile memory or non-volatile memory, or may include both volatile and non-volatile memory. Among them, non-volatile memory may be read-only memory (ROM), programmable read-only memory (PROM), erasable programmable read-only memory (EPROM), electrically erasable programmable read-only memory (EEPROM), or flash memory. Volatile memory may be random access memory (RAM), which is used as an external cache. Many forms of RAM are available, such as static random access memory (SRAM), dynamic random access memory (DRAM), synchronous dynamic random access memory (SDRAM), double data rate synchronous dynamic random access memory (DDR SDRAM), enhanced synchronous dynamic random access memory (ESDRAM), synchronous link dynamic random access memory (SLDRAM), and direct rambus RAM (DR RAM). The memory 102 may store computer program instructions for execution by the AP 1011 and the ADSP 1012 .

[0092] AP 1011 can be a field programmable gate array (FPGA), an application specific integrated circuit (ASIC), a central processing unit (CPU), a network processor (NP), a microcontroller unit (MCU), a programmable logic device (PLD), or other integrated circuit. AP 1011 can send configuration parameters to ADSP 1012 to configure corresponding audio feature ranges for multiple voice recognition modules. When AP 1011 is in a dormant state, AP 1011 can also receive a wake-up signal from ADSP 1012. In a scenario where audio is always on, AP 1011 is typically in a dormant state. At this time, AP 1011 can reduce its operating frequency and shut down threads that consume a lot of power (such as threads for graphics rendering and voice recognition), while retaining necessary threads (such as the main thread and the thread for receiving wake-up signals from ADSP 1012). If AP 1011 includes multiple cores, AP 1011 can also retain one core and shut down the remaining cores. These methods can reduce the power consumption of the AP 1011 in the sleep state.

[0093] The ADSP 1012 can be used to process the audio signal from the AD decoder 103 in real time by executing the computer program instructions stored in the memory 102, thereby executing the sound recognition method involved in this application. When the recognition result of the sound recognition meets the AP wake-up condition, the AP 1011 is woken up and the audio signal is sent to the AP 1011 for further processing.

[0094] As shown in Figure 7, AP 1011 runs on the Android operating system. Taking the software architecture of AP 1011 as an example, the software architecture of AP 1011 may include a driver layer, a kernel layer, a hardware abstraction layer (HAL), a framework layer, and an application layer.

[0095] The driver layer includes hardware drivers for driving hardware resources. For example, the driver layer includes an audio driver, which is used to implement communication of audio data between the AP 1011 and the ADSP 1012.

[0096] The kernel layer includes the operating system (OS) kernel, which manages the system's processes, memory, file system, and network system.

[0097] HAL is used to provide a virtual hardware platform to abstract the hardware. HAL hides the hardware interface details, making the code hardware-independent and portable across multiple platforms. For example, HAL includes an audio HAL, which abstracts the audio path.

[0098] The framework layer provides application programming interfaces (APIs) and system resource services to applications in the application layer. For example, the framework layer includes the ADSP API, where the speech recognition module can implement speech recognition functions.

[0099] The application layer can include various user-oriented applications. For example, the application layer can include a voice assistant that can call the voice recognition module in the framework layer to implement voice control functions. For example, the voice assistant can start or close applications by recognizing voice, assist in autonomous driving by recognizing traffic sounds, enable camera recording by recognizing unusual sounds, and adjust audio playback volume by recognizing ambient noise.

[0100] As shown in Figure 7, the software architecture of ADSP 1012 includes an AD decoder driver, an AAD module, and multiple sound recognition modules associated with the AAD module. The multiple sound recognition modules correspond to multiple audio feature ranges and may include the KWS module, ACD module, ASC module, etc. In the audio-always-on scenario, ADSP 1012 can continuously run the AAD module so that the AAD module can continuously detect audio signals and selectively call the sound recognition module based on the audio features of the audio signal.

[0101] The AD decoder 103 performs AD conversion on the analog audio signal collected by the MIC 104, obtains a digital audio signal and sends it to the ADSP 1012. The AD decoder driver obtains the digital audio signal from the AD decoder 103 and sends it to the AAD module. The AAD module can call the target sound recognition module based on the audio features of the audio signal, wherein the target sound recognition module is a sound recognition module among multiple sound recognition modules that corresponds to the audio feature range where the audio features of the audio signal are located. The target sound recognition module performs sound recognition on the audio signal to obtain a recognition result. If the sound recognition result meets the AP wake-up condition, the AP 1011 in the dormant state is awakened, and the audio signal is sent to the AP 1011 for further processing.

[0102] Exemplarily, AP 1011 is in a dormant state, MIC 104 collects the user's voice, the AD decoder drives the audio signal of the voice to the ADD module, the AAD module calls the KWS module based on the audio features of the audio signal being within the audio feature range of the voice, and the KWS module performs sound recognition on the audio signal. If the KWS module detects that the audio signal includes a wake-up word (for example, "Hello, YOYO"), the KWS module determines that the recognition result of the sound recognition meets the AP's wake-up condition, wakes up the dormant AP 1011, and sends subsequent audio signals to AP 1011. The speech recognition module of the framework layer in AP 1011 performs speech recognition on the subsequent audio signals. If the KWS module detects that the audio signal does not include a wake-up word, the dormant AP 1011 is not woken up.

[0103] As shown in FIG8 , the voice recognition method performed by the ADSP provided in the embodiment of the present application includes S101 to S104:

[0104] S101. The AAD module obtains an audio signal (including a first audio signal and a second audio signal) corresponding to the sound collected by a microphone.

[0105] The microphone (MIC) 104 collects sound to obtain a corresponding analog audio signal and outputs it to the AD decoder 103. The AD decoder 103 performs AD conversion on the analog audio signal to obtain a digital audio signal and outputs it to the AAD module in the ADSP 1012. The AAD module in the ADSP 1012 can store the audio signal in an internal buffer.

[0106] The AAD module can obtain a first audio signal corresponding to a first sound collected by a microphone, and obtain a second audio signal corresponding to the first sound collected by a microphone, wherein the first sound is different from the second sound, and the first audio signal and the second audio signal are different. In the following text, the first audio signal is an audio signal for which a corresponding sound recognition module can be found, and the second audio signal is an audio signal for which a corresponding sound recognition module cannot be found.

[0107] It should be noted that, to avoid repetition, in this application, where there is no need to distinguish between the first audio signal and the second audio signal, that is, where the first audio signal and the second audio signal have commonalities, they are uniformly replaced by audio signal.

[0108] S102. The AAD module extracts audio features of the audio signal (including audio features of the first audio signal and audio features of the second audio signal).

[0109] The AAD module can extract the audio features of the first audio signal and the audio features of the second audio signal. Since the method of extracting audio features is the same for any audio signal, this application does not distinguish between how to extract the audio features of the first audio signal and the audio features of the second audio signal. The application will be described using the method of extracting the audio features of the audio signal as an example.

[0110] This application exemplarily lists three possible ways to extract audio features of an audio signal, but is not intended to be limited thereto. Other forms may also be used to extract audio features of an audio signal.

[0111] Implementation method one: The audio features may include the center frequency and sound pressure level of the preset frequency band of the audio signal. As shown in Figure 9, the AAD module can obtain the components of the preset frequency band of the audio signal by performing band-pass filtering on the audio signal in the preset frequency band (i.e., passing it through a band-pass filter). The geometric mean of the upper frequency limit and the lower frequency limit of the components of the preset frequency band of the audio signal is calculated as the center frequency of the preset frequency band of the audio signal. Afterwards, the amplitude h corresponding to the center frequency in the components of the preset frequency band of the audio signal can be obtained, and the sound pressure p of the preset frequency band of the audio signal can be calculated based on the amplitude h. Specifically, sound pressure p = amplitude h * standard atmospheric pressure. Then, based on the sound pressure of the components of the preset frequency band of the audio signal, the sound pressure level L of the preset frequency band of the audio signal is calculated. Specifically, sound pressure level L = 20log10(p / p0), where p0 is the reference sound pressure.

[0112] For example, Table 1 lists the sound pressure levels (in decibels) of different audio signals in various typical scenarios. As can be seen from Table 1, the sound pressure levels of audio signals can be used to distinguish the scenarios represented by the audio signals to a certain extent, and the sound pressure levels can be used as audio features of the audio signals.

[0113] Table 1

[0114] Table 2 lists the center frequencies and sound pressure levels of audio signals from different sound sources. As can be seen from Table 2, even if the sound pressure level of an audio signal is the same (e.g., 50dB), its source may be different. For example, an audio signal with the same sound pressure level could be a human voice, a piano, or even the sound of an escalator. However, the center frequencies of audio signals from different sound sources may be different. Therefore, in addition to using the sound pressure level of an audio signal as an audio feature, the center frequency of the audio signal can also be used as an audio feature.

[0115] Table 2

[0116] Table 3 lists the center frequency and sound pressure level of the audio signal of the same sound source. It can be seen from Table 3 that the center frequency and sound pressure level of the audio signal of the same sound source may be different. Therefore, for implementation method one, when comparing the audio features of the audio signal with the audio feature range, the center frequency of the audio signal can be compared with the frequency band corresponding to the sound recognition module, and the sound pressure level of the audio signal can be compared with the sound pressure level range corresponding to the sound recognition module. That is, the audio feature range of implementation method one below will include two ranges: frequency band and sound pressure level range. For details, see the description of the audio feature range below.

[0117] Table 3

[0118] Implementation Method 2: The audio feature can be an FBank feature. As shown in Figure 10, the FBank feature of the audio signal can be obtained by performing FBank feature extraction on the audio signal. The FBank feature can be in vector form. For details on how to perform FBank feature extraction, please refer to the previous description and will not be repeated here.

[0119] Implementation Method 3: The audio features can be MFCC features. As shown in Figure 11, MFCC features can be obtained by extracting them from the audio signal. These MFCC features can be in vector form. For details on how to extract MFCC features, refer to the previous description and will not be repeated here.

[0120] S103. The AAD module calls a target sound recognition module based on the first audio feature, and the target sound recognition module performs sound recognition on the audio signal. Based on the audio feature of the second audio signal, none of the multiple sound recognition modules is run.

[0121] The "call" involved in the embodiments of the present application can also be expressed as "run", "load", etc.

[0122] The target sound recognition module is a sound recognition module corresponding to the audio feature range where the audio feature of the first audio signal is located among multiple sound recognition modules, and multiple audio feature ranges correspond to multiple sound recognition modules. That is to say, the audio feature of the first audio signal is compared with multiple audio feature ranges respectively. If the audio feature of the first audio signal is within a certain audio feature range, the corresponding sound recognition module is the target sound recognition module. For example, the audio feature of the audio signal shown in Figure 5 is within audio feature range 1, then the corresponding sound recognition module 1 is the target sound recognition module. In addition, the audio feature of the second audio signal is compared with multiple audio feature ranges respectively. If the audio feature of the second audio signal is not within any audio feature range, that is, there is no corresponding sound recognition module in the audio feature range where the audio feature of the second audio signal is located, then no sound recognition module is run, thereby saving computing power and reducing power consumption.

[0123] If the first audio feature is within multiple audio feature ranges, the sound recognition modules corresponding to these audio feature ranges are all target sound recognition modules. In other words, the target sound recognition module may include one or more sound recognition modules. For example, if the audio features of the audio signal shown in FIG5 are within audio feature range 1 and audio feature range 2, then the corresponding sound recognition modules 1 and 2 are the target sound recognition modules.

[0124] It should be noted that in order to further reduce the waste of computing power of electronic devices and lower power consumption, in the embodiment of the present application, two sound recognition modules of the same type can be assigned to two different audio feature ranges. This avoids calling the same type of sound recognition modules with the same audio feature range, thereby obtaining the same recognition results, wasting computing power and increasing power consumption.

[0125] For example, suppose that both voice recognition module 1 and voice recognition module 2 are KWS modules, but voice recognition module 1 is used for keyword retrieval in a noisy environment (such as a subway station), while voice recognition module 2 is used for keyword retrieval in a quiet environment. These two voice recognition modules process audio signals differently for voice recognition. For example, voice recognition module 1 will filter out noise with a higher gain. Therefore, the audio feature ranges corresponding to these two KWS modules are different.

[0126] For two sound recognition modules of different types, the two corresponding audio feature ranges may be different, partially overlapping, or the same. The embodiment of the present application does not limit this. For example, the sound recognition module 1 shown in Figure 5 is an ACD module, and the sound recognition module 2 is a KWS module, which correspond to the audio feature range 1 and the audio feature range 2, respectively. The audio feature range 1 and the audio feature range 2 may be different, partially overlapping, or the same. By calling two different types of sound recognition modules to perform sound recognition on the audio signal, two recognition results for different purposes can be obtained, and there will be no waste of computing power. Therefore, the embodiment of the present application does not limit the value relationship of the audio feature ranges corresponding to different types of sound recognition modules.

[0127] The following describes step S103 in conjunction with implementation methods 1 to 3 in step S102.

[0128] Implementation method 1 corresponding to step S102: The audio features of the audio signal include the center frequency and sound pressure level of the audio signal after bandpass filtering. In this implementation method, each audio feature range includes a frequency band and a sound pressure level range. For example, each audio feature range can be represented by a key-value pair including a frequency band and a sound pressure level range, and multiple audio feature ranges can be represented by multiple key-value pairs including frequency bands and sound pressure level ranges. For example, audio feature range 1 can be represented as <frequency band 1, sound pressure level range 1>, audio feature range 2 can be represented as <frequency band 2, sound pressure level range 2>, ..., audio feature range N can be represented as <frequency band N, sound pressure level range N>.

[0129] Accordingly, step S103 can be implemented as follows: based on the center frequency and sound pressure level of the first audio signal after bandpass filtering, searching for a first audio feature range from multiple audio feature ranges, wherein the frequency band in the first audio feature range includes the center frequency of the first audio signal after bandpass filtering, and the sound pressure level range in the first audio feature range includes the sound pressure level of the first audio signal after bandpass filtering. If the first audio feature range is found, calling the target sound recognition module corresponding to the first audio feature range, and having the target sound recognition module perform sound recognition on the audio signal. Based on the center frequency and sound pressure level of the second audio signal after bandpass filtering, searching for a second audio feature range from multiple audio feature ranges, wherein the frequency band in the second audio feature range includes the center frequency of the second audio signal after bandpass filtering, and the sound pressure level range in the second audio feature range includes the sound pressure level of the second audio signal after bandpass filtering. If the second audio feature range is not found, not running any of the multiple sound recognition modules, thereby saving computing power and reducing power consumption.

[0130] Exemplarily, as shown in FIG9 , if the center frequency of the first audio signal is within frequency band 1 and the sound pressure level of the first audio signal is within sound pressure level range 1, the AAD module calls the sound recognition module 1 corresponding to <frequency band 1, sound pressure level range 1>.

[0131] It should be noted that the frequency band in the audio feature range can be represented by a frequency interval (e.g., 85Hz-300Hz). If the center frequency of the audio signal is within this frequency interval (e.g., 150Hz), the center frequency of the audio signal is determined to be within the frequency band in the audio feature range. The sound pressure level range in the audio feature range can be represented by a sound pressure level interval or a sound pressure level threshold. When the sound pressure level range in the audio feature range is represented by a sound pressure level interval (e.g., 45dB-55dB), if the sound pressure level of the audio signal (e.g., 50dB) is within this sound pressure level interval, the sound pressure level of the audio signal is determined to be within the sound pressure level range in the audio feature range. When the sound pressure level range in the audio feature range is represented by a sound pressure level threshold, the sound pressure level range means the sound pressure level greater than the sound pressure level threshold. For example, if the sound pressure level of the audio signal (e.g., 50dB) is greater than the sound pressure level threshold (e.g., 45dB), the sound pressure level of the audio signal (50dB) is determined to be within the sound pressure level range (e.g., greater than 45dB) in the audio feature range.

[0132] A typical application scenario of implementation method 1 is as follows:

[0133] Assume that the sound recognition module 1 is a KWS module, the frequency band 1 of the first key-value pair is the frequency band of human voice 85Hz-300Hz (including the center frequency 85Hz-180Hz of the male audio signal, the center frequency 165Hz-255Hz of the female audio signal, and the center frequency 250Hz-300Hz of the child's audio signal in Table 2), and the sound pressure level range 1 of the first key-value pair is the sound pressure level range of human voice 45db-55db (the sound pressure level of the male, female and child audio signals in Table 2 is 50db). Then when the user says a preset sentence (for example, "Hello, YOYO"), the AAD module recognizes that the center frequency of the audio signal is 150Hz, which is within the frequency band of human voice, and recognizes that the sound pressure level of the audio signal is 50db, which is within the sound pressure level range of human voice. Then, the AAD module determines that the target sound recognition module is a KWS module and calls the KWS module so that the KWS module performs keyword retrieval on the audio signal.

[0134] Assume that the sound recognition module 2 is an ASC module, the frequency band 2 of the second key-value pair is 400Hz-500Hz, and the sound pressure level range 2 of the second key-value pair is 45db-55db. Then when playing the piano in a concert hall, the AAD module recognizes that the center frequency of the audio signal is 440Hz, which is within the frequency band 2 (400Hz-500Hz), and recognizes that the sound pressure level of the audio signal is 53db, which is within the sound pressure level range 2 (45db-55db). Then, the AAD module determines that the target sound recognition module is the ASC module and calls the ASC module so that the ASC module can recognize the scene in which the audio is collected.

[0135] Implementation method 2 corresponding to step S102: the audio feature is the FBank feature of the audio signal, and each audio feature range is an FBank feature range. At this time, step S103 can be implemented in the following way: searching for a third audio feature range including the FBank feature of the first audio signal from multiple audio feature ranges; if the third audio feature range is found, calling the target sound recognition module corresponding to the third audio feature range, and having the target sound recognition module perform sound recognition on the audio signal. Searching for a fourth audio feature range including the FBank feature of the second audio signal from multiple audio feature ranges; if the fourth audio feature range is not found, not running any of the multiple sound recognition modules, thereby saving computing power and reducing power consumption.

[0136] For example, as shown in Figure 10, from multiple audio feature ranges (FBank feature range 1, FBank feature range 2...FBank feature range N), it is found that the FBank feature of the audio signal is within FBank feature range 1, that is, FBank feature range 1 includes the FBank feature of the audio signal, then the AAD module determines that the sound recognition module 1 corresponding to FBank feature range 1 is the target sound recognition module, and calls sound recognition module 1.

[0137] As mentioned above, the FBank feature is an M-dimensional vector (e.g.<FBank1,FBank2...FBankM> ), so each FBank feature range can pass M FBank feature thresholds (e.g. ) or M FBank feature intervals (e.g.<h11-h12,h21-h22...hM1-hM2> ) to indicate.<h1>

[0138] Each FBank feature threshold or FBank feature interval in each FBank feature range corresponds to an FBank coefficient in the FBank feature and is used for comparison with the FBank coefficient. For example, the first FBank feature threshold H1 or the first FBank feature interval h11-h12 in the FBank feature range corresponds to the first FBank coefficient FBank1 in the FBank feature and is used for comparison with the first FBank coefficient FBank1; the second FBank feature threshold H2 or the second FBank feature interval h21-h22 corresponds to the second FBank coefficient FBank2 in the FBank feature and is used for comparison with the second FBank coefficient FBank2, and so on.

[0139] When the FBank feature range is represented by an FBank feature threshold, the FBank feature range includes multiple FBank feature thresholds, and the multiple FBank feature thresholds correspond to multiple FBank coefficients. The multiple FBank feature thresholds may include valid FBank feature thresholds and may also include invalid FBank feature thresholds, wherein the valid FBank feature threshold may refer to a FBank feature threshold with a non-null value, and the invalid FBank feature threshold may refer to a null value. For each FBank coefficient corresponding to the valid FBank feature threshold, if the FBank coefficient is greater than its corresponding valid FBank feature threshold, it is determined that the FBank feature of the audio signal is within the FBank feature range. It should be understood that in the present application, for an invalid FBank feature threshold, it can be assumed that the FBank coefficient corresponding to the invalid FBank feature threshold meets the requirements. That is, as long as there is a valid FBank feature threshold in the FBank feature range, and the corresponding FBank coefficient in the FBank feature of the audio signal is not greater than the valid FBank feature threshold, it is determined that the FBank feature of the audio signal is not within the FBank feature range.

[0140] For example, as shown in Figure 12, assuming that the M FBank feature thresholds of the FBank feature range They are all valid FBank feature thresholds. If the first FBank coefficient FBank1 of the FBank feature of the audio signal is greater than the first FBank feature threshold H1 of the FBank feature range, the second FBank coefficient FBank2 of the FBank feature is greater than the second FBank feature threshold H2 of the FBank feature range, and so on, the Mth FBank coefficient FBankM of the FBank feature is greater than the Mth FBank feature threshold HM of the FBank feature range, then it is determined that the FBank feature of the audio signal is within the FBank feature range. For another example, taking the FBank feature including four FBank coefficients (for example, <4, 5, 0, 1>), and the FBank feature range including two valid FBank feature thresholds (for example, <3, 2, NULL, NULL>, NULL represents an invalid FBank feature threshold) as an example, the first FBank coefficient 4 of the FBank feature is greater than the first FBank feature threshold 3 of the FBank feature range, and the second FBank coefficient 5 of the FBank feature is greater than the second FBank feature threshold 2 of the FBank feature range, then it is determined that the FBank feature <4, 5, 0, 1> of the audio signal is within the FBank feature range <3, 2, NULL, NULL>. For another example, assuming that the FBank feature includes four FBank coefficients (for example, <4, 0, 0, 1>), since the second FBank coefficient 0 of the FBank feature is not greater than the second FBank feature threshold 2 of the FBank feature range, it is determined that the FBank feature <4, 0, 0, 1> of the audio signal is not within the FBank feature range <3, 2, NULL, NULL>.<h1>

[0141] When the FBank feature range is represented by an FBank feature interval, the FBank feature range includes multiple FBank feature intervals, and the multiple FBank feature intervals correspond to multiple FBank coefficients. The multiple FBank feature intervals may include valid FBank feature intervals and may also include invalid FBank feature intervals, wherein a valid FBank feature interval may refer to a FBank feature interval with a non-empty value, and an invalid FBank feature interval may refer to a null value. For each FBank coefficient corresponding to a valid FBank feature interval, if the FBank coefficient is within its corresponding valid FBank feature interval, it is determined that the FBank feature of the audio signal is within the FBank feature range. It should be understood that in the present application, for an invalid FBank feature interval, it can be assumed that the MFCC coefficient corresponding to the invalid FBank feature interval meets the requirements. That is, as long as there is a valid FBank feature interval in the FBank feature range, and the corresponding FBank coefficient in the FBank feature of the audio signal is not within the valid FBank feature interval, it is determined that the FBank feature of the audio signal is not within the FBank feature range.

[0142] For example, as shown in Figure 13, assuming that the FBank feature range has M FBank feature intervals<h11-h12,h21-h22...hM1-hM2> They are all valid FBank feature intervals. If the first FBank coefficient FBank1 of the FBank feature of the audio signal is within the first FBank feature interval h11-h12, the second FBank coefficient FBank2 of the FBank feature is within the second FBank feature interval h21-h22, and so on, the Mth FBank coefficient FBankM of the FBank feature is within the Mth FBank feature interval hM1-hM2, then it is determined that the FBank feature of the audio signal is within the FBank feature range. For another example, taking the FBank feature including four FBank coefficients (for example, <4, 5, 0, 1>), and the FBank feature range including two valid FBank feature intervals (for example, <3-6, 2-8, NULL, NULL>, NULL represents an invalid FBank feature interval) as an example, the first FBank coefficient 4 of the FBank feature is within the first FBank feature interval 3-6 of the FBank feature range, and the second FBank coefficient 5 of the FBank feature is within the second FBank feature interval 2-8 of the FBank feature range, then it is determined that the FBank feature <4, 5, 0, 1> of the audio signal is within the FBank feature range <3-6, 2-8, NULL, NULL>. For another example, assuming that the FBank feature includes four FBank coefficients (for example, <4, 0, 0, 1>), since the second FBank coefficient 0 of the FBank feature is not within the second FBank feature interval 2-8 of the FBank feature range, it is determined that the FBank feature <4, 0, 0, 1> of the audio signal is not within the FBank feature range <3-6, 2-8, NULL, NULL>.

[0143] It should be noted that the M FBank feature thresholds in each FBank feature range are uniformly used. To express, or, the M FBank feature intervals in each FBank feature range are uniformly expressed as<h11-h12,h21-h22...hM1-hM2> To represent, but the values ​​of the same position in different FBank feature ranges may be different, for example, the value of the first FBank feature threshold H1 in FBank feature range 1 may be different from the value of the first FBank feature threshold H1 in FBank feature range 2.<h1>

[0144] A typical application scenario of the second implementation method is as follows: For example, as shown in Figure 10, it is assumed that both the sound recognition module 1 and the sound recognition module 2 are KWS modules, but the sound recognition module 1 is used for keyword retrieval in a noisy environment (such as a subway station), and the sound recognition module 2 is used for keyword retrieval in a quiet environment (such as a library). FBank feature range 1 is different from FBank feature range 2. When the FBank feature of the audio signal is in FBank feature range 1, the AAD module determines that the target sound recognition module is a KWS module and calls the KWS module so that the KWS module performs keyword retrieval on the audio signal in a noisy environment (such as a subway station), for example, filtering out noise in the audio signal with a higher gain, thereby improving the accuracy of keyword retrieval.

[0145] For another example, assume that voice recognition module 1 is a KWS module and voice recognition module 2 is an ACD module. When the FBank feature of the audio signal is within FBank feature range 1, the AAD module determines that the target voice recognition module is a KWS module and calls the KWS module so that the KWS module can perform keyword search on the audio signal.

[0146] Implementation method three corresponding to step S102: the audio feature is the MFCC feature of the audio signal, and each audio feature range is an MFCC feature range. At this time, step S103 can be implemented in the following manner: searching for a fifth audio feature range including the MFCC feature of the first audio signal from multiple audio feature ranges; if the fifth audio feature range is found, calling the target sound recognition module corresponding to the fifth audio feature range. Searching for a sixth audio feature range including the MFCC feature of the second audio signal from multiple audio feature ranges; if the sixth audio feature range is not found, not running any of the multiple sound recognition modules, thereby saving computing power and reducing power consumption.

[0147] Exemplarily, as shown in Figure 11, from multiple audio feature ranges (MFCC feature range 1, MFCC feature range 2...MFCC feature range N), it is found that the MFCC features of the audio signal are within MFCC feature range 1, that is, MFCC feature range 1 includes the MFCC features of the audio signal, then the AAD module determines that the sound recognition module 1 corresponding to the MFCC feature range 1 is the target sound recognition module, and calls the sound recognition module 1.

[0148] Similar to FBank features, MFCC features as mentioned above are M-dimensional vectors (e.g.<MFCC1,MFCC2...MFCCM> ), so each MFCC feature range can be passed through M MFCC feature thresholds (e.g.<C1,C2...CM> ) or M MFCC feature intervals (e.g.<c11-c12,c21-c22...cM1-cM2> ) to indicate.

[0149] Each MFCC feature threshold or MFCC feature interval in each MFCC feature range corresponds to an MFCC coefficient in the MFCC feature and is used to compare with the MFCC coefficient. For example, the first MFCC feature threshold C1 or the first MFCC feature interval c11-c12 in the MFCC feature range corresponds to the first MFCC coefficient MFCC1 in the MFCC feature and is used to compare with the first MFCC coefficient MFCC1; the second MFCC feature threshold C2 or the second MFCC feature interval c21-c22 corresponds to the second MFCC coefficient MFCC2 in the MFCC feature and is used to compare with the second MFCC coefficient MFCC2, and so on.

[0150] When the MFCC feature range is represented by an MFCC feature threshold, the MFCC feature range includes multiple MFCC feature thresholds, and the multiple MFCC feature thresholds correspond to multiple MFCC coefficients. The multiple MFCC feature thresholds may include valid MFCC feature thresholds and may also include invalid MFCC feature thresholds, wherein the valid MFCC feature threshold may refer to an MFCC feature threshold with a non-null value, and the invalid MFCC feature threshold may refer to a null value. For each MFCC coefficient corresponding to the valid MFCC feature threshold, if the MFCC coefficient is greater than its corresponding valid MFCC feature threshold, it is determined that the MFCC features of the audio signal are within the MFCC feature range. It should be understood that in this application, for an invalid MFCC feature threshold, it can be assumed that the MFCC coefficient corresponding to the invalid MFCC feature threshold meets the requirements. That is, as long as there is a valid MFCC feature threshold in the MFCC feature range, and the corresponding MFCC coefficient in the MFCC features of the audio signal is not greater than the valid MFCC feature threshold, it is determined that the MFCC features of the audio signal are not within the MFCC feature range.

[0151] For example, as shown in FIG14, assuming that the M MFCC feature thresholds of the MFCC feature range<C1,C2...CM> All are valid MFCC feature thresholds. If the first MFCC feature coefficient MFCC1 of the MFCC feature of the audio signal is greater than the first MFCC feature threshold C1 of the MFCC feature range, the second MFCC feature coefficient MFCC2 of the MFCC feature is greater than the first MFCC feature threshold C2 of the MFCC feature range, and so on, the Mth MFCC feature coefficient MFCCM of the MFCC feature is greater than the Mth MFCC feature threshold CM of the MFCC feature range, then it is determined that the MFCC feature of the audio signal is within the MFCC feature range. For another example, taking the MFCC feature including four MFCC coefficients (for example, <0.4, 0.5, 0, 0.1>), and the MFCC feature range including two valid MFCC feature thresholds (for example, <0.3, 0.2, NULL, NULL>, NULL represents an invalid MFCC feature threshold) as an example, the first MFCC coefficient 0.4 of the MFCC feature is greater than the first MFCC feature threshold 0.3 of the MFCC feature range, and the second MFCC coefficient 0.5 of the MFCC feature is greater than the second MFCC feature threshold 0.2 of the MFCC feature range, then it is determined that the MFCC feature <0.4, 0.5, 0, 0.1> of the audio signal is within the MFCC feature range <0.3, 0.2, NULL, NULL>. For another example, assuming that the MFCC feature includes four MFCC coefficients (for example, <0.4, 0, 0, 0.1>), since the second MFCC coefficient 0 of the MFCC feature is not greater than the second MFCC feature threshold 2 of the MFCC feature range, the MFCC feature <0.4, 0, 0, 0.1> of the audio signal is not within the MFCC feature range <0.3, 0.2, NULL, NULL>.

[0152] When the MFCC feature range is represented by an MFCC feature interval, the MFCC feature range includes multiple MFCC feature intervals, and the multiple MFCC feature intervals correspond to multiple MFCC coefficients. The multiple MFCC feature intervals may include valid MFCC feature intervals and may also include invalid MFCC feature intervals, wherein a valid MFCC feature interval may refer to an MFCC feature interval with a non-empty value, and an invalid MFCC feature interval may refer to a null value. For each MFCC coefficient corresponding to a valid MFCC feature interval, if the MFCC coefficient is within its corresponding valid MFCC feature interval, it is determined that the MFCC features of the audio signal are within the MFCC feature range. It should be understood that in this application, for an invalid MFCC feature interval, it can be assumed that the MFCC coefficient corresponding to the invalid MFCC feature interval meets the requirements, and the MFCC coefficient is not compared with the MFCC feature interval in the MFCC feature range. That is, as long as there is a valid MFCC feature interval in the MFCC feature range, and the corresponding coefficient in the MFCC features of the audio signal is not within the valid MFCC feature interval, it is determined that the MFCC features of the audio signal are not within the MFCC feature range.

[0153] For example, as shown in FIG15 , assuming that the MFCC feature range is M MFCC feature intervals<c11-c12,c21-c22...cM1-cM2> All of them are valid MFCC feature intervals. If the first MFCC coefficient MFCC1 of the MFCC feature of the audio signal is within the first MFCC feature interval c11-c12, the second MFCC coefficient MFCC2 of the MFCC feature is within the second MFCC feature interval c21-c22, and so on, the Mth MFCC coefficient MFCCM of the MFCC feature is within the Mth MFCC feature interval cM1-cM2, then it is determined that the MFCC feature of the audio signal is within the MFCC feature range. For another example, taking the MFCC feature including four MFCC coefficients (for example, <0.4, 0.5, 0, 0.1>), and the MFCC feature range including two valid MFCC feature intervals (for example, <0.3-0.6, 0.2-0.8, NULL, NULL>, NULL represents an invalid MFCC feature interval) as an example, the first MFCC coefficient 0.4 of the MFCC feature is within the first MFCC feature interval 0.3-0.6 of the MFCC feature range, and the second MFCC coefficient 0.5 of the MFCC feature is within the second MFCC feature interval 0.2-0.8 of the MFCC feature range, then it is determined that the MFCC feature <0.4, 0.5, 0, 0.1> of the audio signal is within the MFCC feature range <0.3-0.6, 0.2-0.8, NULL, NULL>. For another example, assuming that the MFCC feature includes four MFCC coefficients (for example, <0.4, 0, 0, 0.1>), since the second MFCC coefficient 0 of the MFCC feature is not in the second MFCC feature interval 0.2-0.8 of the MFCC feature range, the MFCC feature <0.4, 0, 0, 0.1> of the audio signal is not in the MFCC feature range <0.3-0.6, 0.2-0.8, NULL, NULL>.

[0154] It should be noted that the M MFCC feature thresholds in each MFCC feature range are uniformly used.<C1,C2...CM> To express, or, although the M MFCC feature intervals in each MFCC feature range are uniformly expressed<c11-c12,c21-c22...cM1-cM2> However, the values ​​of the same position in different MFCC feature ranges may be different. For example, the value of the first MFCC feature threshold C1 in MFCC feature range 1 may be different from the value of the first MFCC feature threshold C1 in MFCC feature range 2.

[0155] The typical application scenario of implementation method three is similar to the typical application scenario of implementation method two. For example, as shown in Figure 11, it is assumed that the sound recognition module 1 and the sound recognition module 2 are both KWS modules, but the sound recognition module 1 is used for keyword retrieval in a noisy environment (such as a subway station), and the sound recognition module 2 is used for keyword retrieval in a quiet environment (such as a library). The MFCC feature range 1 is different from the MFCC feature range 2. When the MFCC feature of the audio signal is in the MFCC feature range 1, the AAD module determines that the target sound recognition module is a KWS module and calls the KWS module so that the KWS module performs keyword retrieval on the audio signal in a noisy environment (such as a subway station), for example, filtering out noise in the audio signal with a higher gain, thereby improving the accuracy of keyword retrieval.

[0156] For another example, assume that voice recognition module 1 is a KWS module and voice recognition module 2 is an ACD module. When the MFCC features of the audio signal are within MFCC feature range 1, the AAD module determines that the target voice recognition module is the KWS module and calls the KWS module so that the KWS module can perform keyword search on the audio signal.

[0157] The difference between Implementation Method 2 and Implementation Method 3 relative to Implementation Method 1 is that in Implementation Method 1, when matching the center frequency and sound pressure level of the audio signal, as long as the center frequency of the audio signal is within a single frequency band and the sound pressure level is within a single sound pressure level range, the corresponding target sound recognition module can be determined, and the conditions are relatively loose. However, for Implementation Methods 2 and 3, since both FBank characteristics and MFCC are audio features obtained by filtering multiple triangular filters in different frequency bands, the corresponding target sound recognition module can only be determined if the multiple frequency bands of the audio signal are within the filtering frequency band of the triangular filter. The conditions are stricter, but the corresponding target sound recognition module can be determined more accurately.

[0158] In addition, the AAD module provided in the embodiment of the present application can also control two sound recognition modules, so that one of the sound recognition modules obtains the recognition result of the other sound recognition module, thereby assisting the sound recognition module that receives the recognition result to more accurately perform sound recognition on the audio signal. A typical application scenario is as follows: For example, as shown in Figures 10 and 11, assuming that the sound recognition module 1 is an ASC module and the sound recognition module 2 is a KWS module, the audio features of the audio signal include not only the features of the human voice frequency band, but also the features of the noisy environment frequency band. The AAD module determines that the target sound recognition module includes the sound recognition module 2 based on the features of the human voice frequency band, and determines that the target sound recognition module also includes the sound recognition module 1 based on the features of the noisy environment frequency band. The AAD module then calls the sound recognition module 1 and the sound recognition module 2, and the sound recognition module 1 and the sound recognition module 2 perform sound recognition. The sound recognition module 1 performs sound recognition on the audio signal and obtains that the sound collection scene is a noisy environment (such as a subway station). The sound recognition module 2 filters out the noise in the audio signal with a higher gain based on the scene of the noisy environment, and then performs sound recognition on the audio signal after the noise is filtered out, thereby improving the accuracy of keyword retrieval.

[0159] The following describes how the AAD module controls multiple sound recognition modules and two sound recognition modules, so that one sound recognition module obtains the recognition result of another sound recognition module.

[0160] In a possible implementation, taking the example of a target sound recognition module including a first sound recognition module and a second sound recognition module, the second sound recognition module can obtain the recognition result of the first sound recognition module through the transfer of the AAD module. Specifically, as shown in Figure 16, the first sound recognition module (for example, sound recognition module 1) obtains a first recognition result by performing sound recognition on an audio signal. After the AAD module obtains the first recognition result from the first sound recognition module, it sends the first recognition result to the second sound recognition module (for example, sound recognition module 2). The first recognition result can assist the second sound recognition module in performing sound recognition on the audio signal.

[0161] In another possible implementation, still taking the example of the target sound recognition module including the first sound recognition module and the second sound recognition module, the first sound recognition module performs sound recognition on the audio signal to obtain a first recognition result. Controlled by the AAD module, the second sound recognition module can directly obtain the first recognition result from the first sound recognition module. Specifically, the second sound recognition module can directly obtain the first recognition result from the first sound recognition module in at least the following two ways:

[0162] Method 1: The second voice recognition module passively receives the first recognition result from the first voice recognition module. For example, as shown in Figure 17, the AAD module can send a first indication message to the first voice recognition module (e.g., voice recognition module 1), and the first indication message is used to instruct the first voice recognition module to send the first recognition result to the second voice recognition module (e.g., voice recognition module 2). The first recognition result can assist the second voice recognition module in voice recognition of the audio signal.

[0163] Method 2: The second sound recognition module actively obtains the first recognition result from the first sound recognition module. For example, as shown in Figure 18, the AAD module can send a second indication message to the second sound recognition module (e.g., sound recognition module 2). The second indication message is used to instruct the second sound recognition module to obtain the first recognition result from the first sound recognition module (e.g., sound recognition module 1) (e.g., the second sound recognition module requests the first recognition result from the first sound recognition module). The first recognition result can assist the second sound recognition module in performing sound recognition on the audio signal.

[0164] In addition, the AAD module provided in the embodiment of the present application can also send the extracted audio features to the target sound recognition module to assist the target sound recognition module in performing sound recognition on the audio signal. A typical application scenario is as follows: For example, as shown in Figure 10, assuming that the sound recognition module 1 is a KWS module, the audio features of the sound signal include not only the audio features of the human voice frequency band, but also the audio features of the noisy environment frequency band. The AAD module determines that the target sound recognition module is the sound recognition module 1 based on the fact that the audio features of the human voice frequency band are within the audio feature range 1, and then calls the sound recognition module 1 for sound recognition. In addition, the AAD module sends the audio features of the noisy environment frequency band to the sound recognition module 1. The sound recognition module 2 filters out the noise in the audio signal with a higher gain based on the audio features of the noisy environment, and then performs keyword retrieval on the audio signal that has been filtered out of the noise, thereby improving the accuracy of the keyword retrieval.

[0165] S104: If the target voice recognition module performs voice recognition on the audio signal and obtains a recognition result that meets the AP wake-up condition, the target voice recognition module wakes up the AP in the dormant state; otherwise, the AP in the dormant state is not woken up.

[0166] The description of the AP being in the dormant state refers to the description of AP 1011 in FIG. 7 , which will not be repeated here.

[0167] For the KWS scenario, the target voice recognition module performs voice recognition on the audio signal to obtain a recognition result that meets the AP's wake-up condition, which may mean that the KWS module detects that the audio signal includes a wake-up word.

[0168] For example, assuming that the AAD module calls the KWS module (i.e., the target sound recognition module) based on the audio features of the audio signal being within the audio feature range of the speech, the KWS module performs sound recognition on the audio signal. If the KWS module detects that the audio signal includes a wake-up word (for example, "Hello, YOYO"), the KWS module determines that the recognition result obtained by the sound recognition meets the AP's wake-up condition, and then wakes up the AP in a dormant state and sends subsequent audio signals to the AP. The speech recognition module in the AP can perform speech recognition on subsequent audio signals. If the KWS module detects that the audio signal does not include a wake-up word, the dormant AP will not be woken up to reduce the power consumption of the electronic device.

[0169] For the ASC scenario, the target sound recognition module performs sound recognition on the audio signal and obtains a recognition result that meets the AP wake-up condition, which may mean that the ASC module detects that the collection scene of the audio signal belongs to a preset scene (such as a concert hall, subway station, etc.).

[0170] For example, assuming that the AAD module calls the ASC module (i.e., the target sound recognition module) based on the audio features of the audio signal being within the audio feature range of a noisy room, the ASC module performs sound recognition on the audio signal. If the ASC module detects that the audio signal acquisition scene is a concert hall, the ASC module determines that the recognition result obtained by the sound recognition meets the AP wake-up condition, and then wakes up the AP in a dormant state and sends the subsequent audio signal to the AP. The music recognition module in the AP can perform music recognition on subsequent audio signals, such as recognizing the name of music. If the ASC module detects that the audio signal acquisition scene is a subway station, the dormant AP will not be woken up to reduce the power consumption of the electronic device.

[0171] For the ACD scenario, the target sound recognition module performs sound recognition on the audio signal to obtain a recognition result that meets the AP wake-up condition, which may mean that the ACD module detects that the audio signal is generated from a preset event (such as a knock sound generated by knocking on the door, an animal call, etc.).

[0172] For example, assuming that the AAD module calls the ACD module (i.e., the target sound recognition module) based on the fact that the audio features of the audio signal are within the audio feature range of living noise, the ACD module performs sound recognition on the audio signal. If the ACD module detects that the audio signal includes a knock on the door (i.e., the event is a knock on the door), the ASC module determines that the recognition result obtained by the sound recognition meets the AP's wake-up condition, and then wakes up the AP in a dormant state and sends subsequent audio signals to the AP. The hearing-impaired assistance application in the AP can display a dialog box "Someone is knocking on the door" to assist users with hearing impairments. If the ASC module detects that the audio signal includes animal calls, the dormant AP will not be woken up to reduce the power consumption of the electronic device.

[0173] The sound recognition method and electronic device provided in the embodiment of the present application are executed by ADSP. First, the AAD module in the ADSP obtains the audio signal corresponding to the sound collected by the microphone, and extracts the audio features of the audio signal. Then, by comparing the audio features of the audio signal with multiple audio feature ranges, the sound recognition module corresponding to the audio feature range where the audio features of the audio signal are located is determined from multiple sound recognition modules, that is, the target sound recognition module, and the target sound recognition module is called. Due to different sound sources and different audio signal collection scenarios, the audio features of the audio signal belong to different audio feature ranges, and the suitable sound recognition modules are also different. Therefore, the target sound recognition module suitable for the audio features of the audio signal can be screened out from multiple sound recognition modules, and the target sound recognition module is called to perform sound recognition on the audio signal to obtain a recognition result, rather than calling all sound recognition modules to perform sound recognition on the audio signal. Thereby reducing the power consumption of electronic devices with multiple sound recognition modules in audio-on scenarios.

[0174] As shown in FIG19 , an embodiment of the present application further provides a chip system. The chip system 190 includes an ADSP 1901 and an interface circuit 1902. Optionally, it may also include an AP 1903. ADSP 1901, interface circuit 1902, and AP 1903 may be interconnected via a line. Interface circuit 1902 may be used to communicate with other devices (e.g., the AD decoder described above). ADSP 1901 is used to perform the various functions or steps in the above-described method embodiments.

[0175] Exemplarily, the interface circuit 1902 is used to receive audio signals. An AAD module runs in the ADSP. The AAD module is associated with multiple sound recognition modules. The sound activity detection module is always on in the ADSP. After the electronic device where the ADSP is located is turned on, the AAD module can run in the ADSP. The above-mentioned multiple sound recognition modules are selectively triggered and called by the AAD module based on the audio signal. Specifically, the sound activity detection module selectively triggers and calls the target sound recognition module based on the audio signal to perform sound recognition on the audio signal. If the recognition result of the sound recognition of the audio signal by the target sound recognition module meets the wake-up condition of AP 1902, the dormant AP 1902 is woken up; otherwise, the dormant AP 1902 is not woken up.

[0176] An embodiment of the present application further provides an electronic device comprising: a microphone, an ADSP, and a memory. Optionally, an AP is further included. The ADSP is coupled to the microphone, the memory, and the AP. The memory stores computer code, which includes instructions. When the ADSP executes the instructions, the electronic device can perform the various functions or steps in the above-described method embodiment. The structure of the electronic device can refer to the structure shown in Figure 6.

[0177] An embodiment of the present application also provides a computer-readable storage medium, which includes instructions. When the instructions are executed on the above-mentioned electronic device, the electronic device executes each step in the above-mentioned method embodiment, for example, executes the method shown in Figure 8.

[0178] An embodiment of the present application further provides a computer program product comprising instructions, which, when executed on the electronic device, enables the electronic device to execute the various steps in the above method embodiment, such as executing the method shown in FIG8 .

[0179] Regarding the technical effects of the chip system, computer-readable storage medium, and computer program product, refer to the technical effects of the previous method embodiments.

[0180] It should be understood that in the various embodiments of the present application, the size of the serial numbers of the above-mentioned processes does not mean the order of execution. The execution order of each process should be determined by its function and internal logic, and should not constitute any limitation on the implementation process of the embodiments of the present application.

[0181] Those skilled in the art will appreciate that the modules and algorithm steps of each example described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, or a combination of computer software and electronic hardware. Whether these functions are performed in hardware or software depends on the specific application and design constraints of the technical solution. Professional and technical personnel can use different methods to implement the described functions for each specific application, but such implementation should not be considered beyond the scope of this application.

[0182] Those skilled in the art will clearly understand that, for the convenience and brevity of description, the specific working processes of the systems, devices and modules described above can refer to the corresponding processes in the aforementioned method embodiments and will not be repeated here.

[0183] In the several embodiments provided in this application, it should be understood that the disclosed systems, devices, and methods can be implemented in other ways. For example, the device embodiments described above are merely illustrative. For example, the division of the modules is merely a logical function division. In actual implementation, there may be other division methods, such as multiple modules or components can be combined or integrated into another device, or some features can be ignored or not executed. Another point is that the mutual coupling or direct coupling or communication connection shown or discussed can be through some interfaces, indirect coupling or communication connection of devices or modules, which can be electrical, mechanical or other forms.

[0184] The modules described as separate components may or may not be physically separate, and the components shown as modules may or may not be physical modules, that is, they may be on a single device or distributed across multiple devices. Some or all of the modules may be selected based on actual needs to achieve the purpose of this embodiment.

[0185] In addition, the functional modules in the various embodiments of the present application may be integrated into one device, or each module may exist physically separately, or two or more modules may be integrated into one device.

[0186] In the above embodiments, all or part of the embodiments may be implemented by software, hardware, firmware, or any combination thereof. When implemented using a software program, all or part of the embodiments may be implemented in the form of a computer program product. The computer program product includes one or more computer instructions. When the computer program instructions are loaded and executed on a computer, all or part of the processes or functions described in the embodiments of the present application are generated. The computer may be a general-purpose computer, a special-purpose computer, a computer network, or other programmable device. The computer instructions may be stored in a computer-readable storage medium or transmitted from one computer-readable storage medium to another computer-readable storage medium. For example, the computer instructions may be transmitted from one website, computer, server, or data center to another website, computer, server, or data center via a wired (e.g., coaxial cable, optical fiber, digital subscriber line (DSL)) or wireless (e.g., infrared, wireless, microwave, etc.) method. The computer-readable storage medium may be any available medium that a computer can access or a data storage device such as a server or data center that includes one or more media integrated therein. The available medium may be a magnetic medium (eg, a floppy disk, a hard disk, a magnetic tape), an optical medium (eg, a DVD), or a semiconductor medium (eg, a solid state disk (SSD)).

[0187] The above description is merely a specific embodiment of the present application, but the scope of protection of the present application is not limited thereto. Any changes or substitutions that can be easily conceived by a person skilled in the art within the technical scope disclosed in this application should be included in the scope of protection of this application. Therefore, the scope of protection of this application should be based on the scope of protection of the claims.

Claims

1. A sound recognition method, characterized in that: The method is applied to an audio digital signal processor, in which a sound activity detection module is run, and the sound activity detection module is associated with a plurality of sound recognition modules. The method comprises: Acquire a first audio signal corresponding to a first sound collected by a microphone; extracting audio features of the first audio signal; Based on the audio feature of the first audio signal, a target sound recognition module is called to perform sound recognition on the first audio signal; the target sound recognition module is a sound recognition module corresponding to an audio feature range where the audio feature of the first audio signal is located among the multiple sound recognition modules, and the multiple audio feature ranges correspond to the multiple sound recognition modules; Acquire a second audio signal corresponding to a second sound collected by the microphone; extracting audio features of the second audio signal; Based on the audio feature of the second audio signal, any sound recognition module among the multiple sound recognition modules is not run, wherein the audio feature range where the audio feature of the second audio signal is located has no corresponding sound recognition module.

2. The method according to claim 1, characterized in that: The audio digital signal processor is coupled to an application processor; after calling a target sound recognition module based on the audio feature of the first audio signal and performing sound recognition on the audio signal, the method further includes: If the recognition result of the target sound recognition module meets the wake-up condition of the application processor, the application processor in the dormant state is woken up; otherwise, the application processor in the dormant state is not woken up.

3. The method according to claim 1 or 2, characterized in that: The audio feature of the first audio signal includes a center frequency and a sound pressure level of a preset frequency band of the first audio signal, and each audio feature range includes a frequency band and a sound pressure level range; calling a target sound recognition module based on the audio feature of the first audio signal includes: Based on the center frequency and the sound pressure level of the preset frequency band of the first audio signal, searching for a first audio feature range from the multiple audio feature ranges, wherein the frequency band in the first audio feature range includes the center frequency of the preset frequency band of the first audio signal, and the sound pressure level range in the first audio feature range includes the sound pressure level of the preset frequency band of the first audio signal; The first audio feature range is found, and a target sound recognition module corresponding to the first audio feature range is called.

4. The method according to claim 3, characterized in that: The extracting the audio feature of the first audio signal includes: Performing a bandpass filter in a preset frequency band on the first audio signal to obtain a component in the preset frequency band of the first audio signal; Calculating a geometric mean of an upper frequency limit and a lower frequency limit of a component of a preset frequency band of the first audio signal, wherein a center frequency of the preset frequency band of the first audio signal is the geometric mean; obtaining a sound pressure of the preset frequency band of the first audio signal according to an amplitude of a component of the preset frequency band of the first audio signal corresponding to a center frequency of the first audio signal; A sound pressure level of the preset frequency band of the first audio signal is obtained according to the sound pressure of the preset frequency band of the first audio signal.

5. The method according to any one of claims 1 to 4, characterized in that The audio feature of the second audio signal includes a center frequency and a sound pressure level of a preset frequency band of the second audio signal, and each audio feature range includes a frequency band and a sound pressure level range; and the method of not running any of the multiple sound recognition modules based on the audio feature of the second audio signal includes: Based on the center frequency and the sound pressure level of the second audio signal, searching for a second audio feature range from the multiple audio feature ranges, wherein the frequency band in the second audio feature range includes the center frequency of the second audio signal, and the sound pressure level range in the second audio feature range includes the sound pressure level of the second audio signal; If the second audio feature range is not found, any sound recognition module among the multiple sound recognition modules will not be run.

6. The method according to claim 5, characterized in that The extracting the audio feature of the second audio signal comprises: Performing a bandpass filter in a preset frequency band on the second audio signal to obtain a component in the preset frequency band of the second audio signal; Calculating a geometric mean of an upper frequency limit and a lower frequency limit of a component of a preset frequency band of the second audio signal, wherein a center frequency of the preset frequency band of the second audio signal is the geometric mean; obtaining a sound pressure of the preset frequency band of the second audio signal according to an amplitude of a component of the preset frequency band of the second audio signal corresponding to the center frequency of the first audio signal; The sound pressure level of the preset frequency band of the second audio signal is obtained according to the sound pressure of the preset frequency band of the second audio signal.

7. The method according to claim 1 or 2, characterized in that: The audio feature of the first audio signal is an FBank feature of the first audio signal obtained by performing filter bank FBank feature extraction on the first audio signal, and each audio feature range is an FBank feature range; calling the target sound recognition module based on the audio feature of the first audio signal includes: Searching for a third audio feature range including the FBank feature of the first audio signal from the plurality of audio feature ranges; The third audio feature range is found, and a target sound recognition module corresponding to the third audio feature range is called.

8. The method according to any one of claims 1, 2 or 7, characterized in that The audio feature of the second audio signal is an FBank feature of the second audio signal obtained by performing FBank feature extraction on the second audio signal, and each audio feature range is an FBank feature range; the audio feature based on the second audio signal does not run any of the multiple sound recognition modules, including: Searching for a fourth audio feature range including the FBank feature of the second audio signal from the plurality of audio feature ranges; If the fourth audio feature range is not found, any sound recognition module among the multiple sound recognition modules will not be run.

9. The method according to claim 1 or 2, characterized in that: The audio feature of the first audio signal is an MFCC feature of the first audio signal obtained by extracting Mel-frequency cepstral coefficient MFCC features of the first audio signal, and each audio feature range is an MFCC feature range; calling the target sound recognition module based on the audio feature of the first audio signal includes: Searching for a fifth audio feature range including the MFCC feature from the plurality of audio feature ranges; The fifth audio feature range is found, and a target sound recognition module corresponding to the fifth audio feature range is called.

10. The method according to any one of claims 1, 2 or 9, characterized in that: The audio feature of the second audio signal is an MFCC feature of the second audio signal obtained by performing MFCC feature extraction on the second audio signal, and each audio feature range is an MFCC feature range; and the method of not running any of the multiple sound recognition modules based on the audio feature of the second audio signal includes: Searching for a sixth audio feature range including the MFCC features of the second audio signal from the plurality of audio feature ranges; If the sixth audio feature range is not found, any of the multiple sound recognition modules will not be run.

11. The method according to any one of claims 1 to 10, characterized in that: The target sound recognition module includes a first sound recognition module and a second sound recognition module, and the method further includes: Obtaining a first recognition result of the audio signal recognized by the first sound recognition module; The first recognition result is sent to the second sound recognition module, and the second sound recognition module performs sound recognition on the audio signal according to the first recognition result.

12. The method according to any one of claims 1 to 10, characterized in that: The target sound recognition module includes a first sound recognition module and a second sound recognition module, and the method further includes: First indication information is sent to the first sound recognition module, where the first indication information is used to instruct the first sound recognition module to send a first recognition result of recognizing the audio signal to the second sound recognition module.

13. The method according to any one of claims 1 to 10, characterized in that: The target sound recognition module includes a first sound recognition module and a second sound recognition module, and the method further includes: Second indication information is sent to the second sound recognition module, where the second indication information is used to instruct the second sound recognition module to obtain a first recognition result of recognizing the audio signal from the first sound recognition module.

14. The method according to any one of claims 1 to 13, characterized in that: Also includes: The audio feature is sent to the target sound recognition module; wherein the audio feature is used by the target sound recognition module to perform sound recognition on the audio signal.

15. A chip system, characterized in that: include: Audio digital signal processors and interface circuits; The interface circuit is used to receive an audio signal; the audio digital signal processor is used to execute the method according to any one of claims 1 to 14, a sound activity detection module is running in the audio digital signal processor, and the sound activity detection module is associated with multiple sound recognition modules; The sound activity detection module is used to: obtain a first audio signal corresponding to a first sound collected by a microphone; extract audio features of the first audio signal; call a target sound recognition module based on the audio features of the first audio signal to perform sound recognition on the first audio signal; the target sound recognition module is a sound recognition module corresponding to an audio feature range where the audio features of the first audio signal are located among the multiple sound recognition modules, and the multiple audio feature ranges correspond to the multiple sound recognition modules; obtain a second audio signal corresponding to a second sound collected by the microphone; and extract audio features of the second audio signal; Based on the second audio signal The audio feature of the second audio signal does not run any of the multiple sound recognition modules, wherein the audio feature range where the audio feature of the second audio signal is located has no corresponding sound recognition module.

16. The chip system according to claim 15, characterized in that: The chip system also includes an application processor; The target sound recognition module is used to: if the recognition result of sound recognition on the first audio signal meets the wake-up condition of the application processor, then wake up the application processor in a dormant state; if the recognition result of sound recognition on the first audio signal does not meet the wake-up condition of the application processor, then do not wake up the application processor in a dormant state.

17. An electronic device, characterized in that: include: A microphone, an audio digital signal processor, and a memory; the audio digital signal processor is coupled to the microphone and the memory; computer code is stored in the memory, and the computer code includes instructions. When the audio digital signal processor executes the instructions, the electronic device executes the method as described in any one of claims 1-14.

18. The electronic device according to claim 17, characterized in that: The system also includes an application processor, the audio digital signal processor is coupled to the application processor, and the audio digital signal processor is used to wake up the application processor in a dormant state.

19. A computer-readable storage medium, characterized in that: The method comprises instructions, which, when executed on an electronic device, cause the electronic device to execute the method according to any one of claims 1 to 14.