Audio processing method and device

By employing a frame skipping strategy and scene-adaptive feature extraction in electronic devices, the high power consumption problem of voice wake-up function is solved, achieving low-power and high-accuracy audio recognition.

CN121415768APending Publication Date: 2026-01-27HONOR DEVICE CO LTD
View PDF 4 Cites 0 Cited by

Patent Information

Application Number
CN202410955910.9
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2024-07-16
Publication Date
2026-01-27

AI Technical Summary

Technical Problem

Electronic devices consume significant power during voice interaction functions such as voice wake-up, which affects the device's energy efficiency.

Method used

By employing a frame skipping strategy in electronic devices, features are extracted from audio signals at regular intervals. By combining keywords from different scenarios, the feature extraction strategy is adjusted to reduce power consumption and improve recognition accuracy.

Benefits of technology

This technology reduces the power consumption of electronic devices in voice interaction functions, while improving the accuracy of audio recognition and the user experience.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121415768A_ABST
    Figure CN121415768A_ABST
Patent Text Reader

Abstract

The embodiment of the invention provides an audio processing method and device, and relates to the technical field of terminals. The method comprises: in a first scene, performing feature extraction on a first audio signal collected by an electronic device every N frames of audio signals, and performing audio recognition based on the extracted first audio features, N being related to a first keyword pre-collected in the first scene, and N being an integer greater than 0; in the second scene, feature extraction is carried out on a second audio signal collected by the electronic equipment every M frames of audio signals, audio recognition is carried out based on extracted second audio features, M is related to a second keyword collected in advance in the second scene, and M is an integer larger than 0. Therefore, in different application scenes, feature extraction can be performed on part of frames of voice signals based on the frame skipping strategies corresponding to the application scenes, and the power consumption of the electronic equipment can be reduced.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of terminal technology, and in particular to audio processing methods and apparatus. Background Technology

[0002] To enhance the convenience and intelligence of human-computer interaction, voice interaction functions are widely used in electronic devices. Examples include voice wake-up and voice recognition. For instance, an electronic device can collect audio signals emitted by a user. By extracting features and recognizing keywords from the audio signals, the device can identify keywords within the audio signal and then execute corresponding processes based on those keywords. For example, if the electronic device recognizes a wake-up word in the audio signal, it is activated.

[0003] However, electronic devices may consume a significant amount of power while implementing voice interaction functions such as voice wake-up. Summary of the Invention

[0004] This application provides an audio processing method and apparatus, applicable to the field of terminal technology. By enabling electronic devices to extract features from partial frames of audio signals, the power consumption of the electronic devices in implementing voice interaction functions is reduced.

[0005] In a first aspect, embodiments of this application propose an audio processing method applied to an electronic device. In a first scenario, every N frames of audio signal, features are extracted from the first audio signal collected by the electronic device, and audio recognition is performed based on the extracted first audio features. N is an integer greater than 0, and N is related to a first keyword pre-collected in the first scenario. In a second scenario, every M frames of audio signal, features are extracted from the second audio signal collected by the electronic device, and audio recognition is performed based on the extracted second audio features. M is related to a second keyword pre-collected in the second scenario, and M is an integer greater than 0.

[0006] The first scenario can be a voice wake-up application scenario. N can be the number of frames skipped in frame skipping strategy 1 in method 800, i.e., XY. When the wake-up word is "Hello, Youyou," N can be 2 as described below. The first audio feature can be, for example, speech feature 1 in method 800. Audio recognition can also be understood as keyword recognition in S803. The pre-collected first keyword can be, for example, "you" in "Hello, Youyou" as described below. The second scenario can be, for example, a navigation application scenario, a music application scenario, or an office application scenario. The second keyword can be a character from a high-frequency speech word in the second scenario. The second keywords M and N are preset. M and N can be different. Alternatively, M and N can also be the same if the first keyword and the second keyword are the same.

[0007] The audio processing method of this application allows the electronic device to preset frame skipping strategies corresponding to different scenarios. The frame skipping strategy is essentially the strategy used by the electronic device to extract features from the audio signal. For example, the frame skipping strategy for the first scenario is N frames, and the frame skipping strategy for the second scenario is M frames. This allows the electronic device to extract features from the audio signal it collects at least every one frame. This reduces the power consumption of the feature extraction process, thereby reducing the power consumption for voice interaction functionality.

[0008] Furthermore, different application scenarios can have different frame skipping strategies. For example, M and N can be different. Each application scenario's frame skipping strategy is determined based on the keywords in that application scenario. This allows electronic devices to perform feature extraction based on the frame skipping strategy corresponding to that application scenario. This reduces the power consumption of feature extraction while minimizing the omission of audio features from some words, thus improving the accuracy of feature extraction and consequently increasing the accuracy of audio recognition.

[0009] In one possible implementation, the first keyword is the wake word of the electronic device; audio recognition based on the extracted first audio features includes: determining whether the first audio signal contains a wake word based on the first audio features; the method further includes: if the first audio signal contains a wake word, the electronic device displays a first interface, which is the interface of the electronic device after being woken up by voice.

[0010] The wake word could be, for example, "Hello, Youyou" as described below, and the first keyword could be "you" from "Hello, Youyou". The first interface could be, for example,... Figure 1 The interface (b) in the text.

[0011] In this way, the power consumption of electronic devices in detecting whether the audio signal contains a wake-up word is relatively small, thus reducing the power consumption of electronic devices in implementing voice wake-up functions.

[0012] In one possible implementation, the first keyword is the word with the shortest first pronunciation duration in the wake word, and the first pronunciation duration is the shorter of the initial consonant pronunciation duration and the final vowel pronunciation duration of a word.

[0013] It can be understood that when a character includes an initial and a final, the first pronunciation duration of the character is the shorter one between the pronunciation duration of the initial and the pronunciation duration of the final of the character; when a character includes an initial but does not include a final, the first pronunciation duration of the character is the pronunciation duration of the initial of the character; when a character includes a final but does not include an initial, the first pronunciation duration of the character is the pronunciation duration of the final of the character. The first pronunciation duration of "你" can be the pronunciation duration of the initial of "你", that is, 30 - 70 ms; the first pronunciation duration of "好" can be the pronunciation duration of the initial of "好", that is, 70 - 130 ms; the first pronunciation duration of "悠" can be the pronunciation duration of the final of "悠", that is, 405 ms.

[0014] In this way, when N and the first keyword are determined, during the process of extracting features from the first audio signal for every N frames of audio signals by the electronic device, the situation of missing the audio features of any character in the wake-up word can be reduced, making the accuracy of audio recognition by the electronic device relatively high.

[0015] In a possible implementation manner, N is determined based on a first ratio, and the first ratio is the ratio of the pronunciation duration of the initial of the first keyword to the duration of one frame of speech signal.

[0016] Among them, the pronunciation duration of the initial of the first keyword can be the pronunciation duration of the initial of "你" in the following text, that is, 30 - 70 ms. The duration of one frame of speech signal can be, for example, 10 ms in the following text. The first ratio can be 30 / t in the following text, and when the duration of one frame of speech signal is 10 ms, the first ratio can be the ratio of 30 ms to 10 ms. N is, for example, 2.

[0017] In this way, the duration of N frames of speech signals can be less than the pronunciation duration of the initial and the pronunciation duration of the final of any character in the wake-up word, so that during the process of extracting features from the first audio signal for every N frames of audio signals by the electronic device, the situation of missing the audio features of any character in the wake-up word can be reduced, making the accuracy of audio recognition by the electronic device relatively high.

[0018] In a possible implementation manner, in the second scenario, for every M frames of audio signals, features are extracted from the second audio signal collected by the electronic device, including: when the electronic device is in a preset mode and the first application is running in the foreground of the electronic device, obtaining a first frame skipping strategy corresponding to the first application, where the first frame skipping strategy includes M; collecting the second audio signal; and based on the first frame skipping strategy, extracting features from the second audio signal to obtain second audio features.

[0019] The second scenario is, for example, a navigation application scenario. The first application is, for example, a navigation application. The preset mode can be the preset state described below. The first frame skipping strategy can be frame skipping strategy a. The M frames are, for example, the q1 frames included in frame skipping strategy a. The second audio feature can be speech feature 3 in method 900. Alternatively, if the second scenario is, for example, a music application scenario, then the first frame skipping strategy is frame skipping strategy b; or, if the second scenario is, for example, an office application scenario, then the first frame skipping strategy is frame skipping strategy c. Based on the first frame skipping strategy, feature extraction of the second audio signal can be, for example, performing feature extraction on 1 frame of audio signal every M frames of audio signal.

[0020] In this way, in the second scenario, the electronic device can extract features from part of the speech signal based on the frame skipping strategy corresponding to the second scenario, so that the power consumption of the electronic device in recognizing the keywords corresponding to the second scenario is relatively small.

[0021] In one possible implementation, the second keyword is the word with the shortest first pronunciation duration among at least one keyword, where the first pronunciation duration is the shorter of the initial consonant pronunciation duration and the final vowel pronunciation duration of each word, and the at least one keyword is a pre-collected high-frequency speech word corresponding to the first application.

[0022] The term "high-frequency speech words" does not imply that the frequency of use of speech words must exceed a certain value. High-frequency speech words can be speech words commonly used in an application scenario, as determined by experience or experimentation. For example, in the second scenario, which is a navigation application scenario, at least one keyword could be "navigate home," "navigate to the company," or "navigate to the park," etc.

[0023] In this way, in the second scenario, during the feature extraction process of the second audio signal by the electronic device every M frames of audio signal, the possibility of missing the audio features of at least one word in a keyword can be reduced, resulting in a higher accuracy of audio recognition by the electronic device.

[0024] In one possible implementation, audio recognition based on the extracted second audio features includes: determining whether the second audio signal contains a target keyword corresponding to the first application, based on the second audio features, wherein the target keyword is part or all of the keywords in at least one keyword.

[0025] The target keywords corresponding to the first application can be, for example, keywords related to the first application that are pre-installed in the electronic device. For instance, if the second application is a navigation application, at least one keyword can include, but is not limited to, one or more of the following: "Navigate home," "Navigate to the company," "Navigate to the park," "End navigation," or "Exit navigation." The target keywords can be, for example, any one or more of the following: "Navigate home," "Navigate to the company," or "End navigation."

[0026] It should be understood that the implementation of this solution is similar to that of S907, as can be seen in the following description.

[0027] In this way, electronic devices can identify target keywords with low power consumption and execute the corresponding processes, resulting in a better user experience.

[0028] In one possible implementation, the preset mode is any of the following: in-vehicle mode, sports mode, or conference mode.

[0029] Among them, the preset mode can be the preset state described below; the vehicle mode can be the vehicle state described below; the sports mode can be the sports state described below; and the meeting mode can be the meeting state described below.

[0030] In this way, when it is inconvenient for users to operate electronic devices with their hands, they can use voice commands to instruct the electronic devices to perform the corresponding processes, resulting in a better user experience.

[0031] In one possible implementation, the method further includes: in the third scene, every S frames of audio signal, performing feature extraction on the third audio signal collected by the electronic device, and performing audio recognition based on the extracted third audio features, where S is related to the third keyword collected in advance in the third scene, and S is an integer greater than 0.

[0032] The third scenario can differ from the second scenario. For example, the third scenario could be a music application scenario, where S represents the number of skipped frames included in the frame skipping strategy b. The number of skipped frames refers to the number of frames between two consecutive feature extraction operations. The third keyword can be a character from a high-frequency speech word in the third scenario. S and M can be different.

[0033] In this way, in different application scenarios, electronic devices can extract features from the audio signals they collect at least once per frame. This results in lower power consumption for feature extraction and higher accuracy in audio recognition.

[0034] In one possible implementation, the third keyword is the word with the shortest first pronunciation duration among the high-frequency speech words corresponding to the second application. The first pronunciation duration is the shorter of the initial consonant pronunciation duration and the final vowel pronunciation duration of a word. In the third scenario, every S frames of audio signal, feature extraction is performed on the third audio signal collected by the electronic device, including: when the electronic device is in a preset mode and the second application is running in the foreground of the electronic device, obtaining the second frame skipping strategy corresponding to the second application, the second frame skipping strategy including S; collecting the third audio signal; and extracting features from the third audio signal based on the second frame skipping strategy to obtain the third audio features.

[0035] The second application differs from the first application. The second application could be, for example, a music application or an office application. The high-frequency speech words corresponding to the second application could be pre-collected speech words related to the second application. These high-frequency speech words could be determined based on experience; alternatively, they could be pre-collected, for example, speech words used by the first user more than or equal to a first threshold within a preset time period. Speech words used more than or equal to the first threshold are all related to the second application, meaning the electronic device needs to execute the processes corresponding to these speech words through the second application. The third scenario could be, for example, an office application scenario, where the high-frequency speech words corresponding to the second application could be "play the previous page" and "play the next page," etc.

[0036] The third scenario could be, for example, a music application scenario, where the second frame skipping strategy could be frame skipping strategy b; or, the third scenario could be, for example, an office application scenario, where the second frame skipping strategy could be frame skipping strategy c. Based on the second frame skipping strategy, feature extraction of the third audio signal could be performed, for example, by extracting features from one frame of audio signal every S frames of audio signal.

[0037] In this way, in the third scenario, the electronic device can also extract features from some frame audio signals, and the power consumption of the electronic device to realize voice interaction function in the third scenario is relatively small.

[0038] In one possible implementation, the electronic device includes a digital signal processor (DSP) and a microphone. In a first scenario, every N frames of audio signal, feature extraction is performed on the first audio signal acquired by the electronic device, and audio recognition is performed based on the extracted first audio features. This includes: the microphone acquiring the first audio signal; the DSP acquiring the first audio signal from the microphone; the DSP extracting features from one frame of audio signal every N frames of audio signal to obtain the first audio features; and the DSP determining whether the first audio signal contains a wake-up word of the electronic device based on the first audio features.

[0039] In this way, the power consumption of the electronic device in acquiring the first audio feature is relatively low. Furthermore, because DSPs are highly efficient, low-power, and capable of real-time processing audio signals, acquiring the first audio feature via a DSP results in lower power consumption and higher efficiency for the electronic device.

[0040] Optionally, the electronic device may also include an AP, and the method further includes: if the DSP determines that the first audio signal includes a wake-up word of the electronic device, the DSP instructs the AP to display a first interface.

[0041] In one possible implementation, N is 2.

[0042] It should be understood that the first scenario can be an application scenario of voice wake-up, and the wake-up word of the electronic device can be "Hello, Youyou". Then N can be 2. See the description below for details.

[0043] Secondly, embodiments of this application provide an audio processing apparatus, which may be an electronic device, a chip, or a chip system within an electronic device. The audio processing apparatus may include an audio acquisition unit and a processing unit. When the audio processing apparatus is an electronic device, the audio acquisition unit may be a microphone. The audio acquisition unit is used to perform the step of acquiring audio signals, so that the electronic device implements an audio processing method described in the first aspect or any possible implementation of the first aspect. When the audio processing apparatus is an electronic device, the processing unit may be a processor, such as a DSP. The audio processing apparatus may further include a storage unit, which may be a memory. The storage unit is used to store instructions, and the processing unit executes the instructions stored in the storage unit to make the electronic device implement an audio processing method described in the first aspect or any possible implementation of the first aspect. When the audio processing apparatus is a chip or a chip system within an electronic device, the processing unit may be a processor. The processing unit executes the instructions stored in the storage unit to make the electronic device implement an image processing method described in the first aspect or any possible implementation of the first aspect. The storage unit can be a storage unit within the chip (e.g., a register, cache, etc.) or a storage unit located outside the chip within the electronic device (e.g., a read-only memory, random access memory, etc.).

[0044] For example, an audio acquisition unit is used to acquire a first audio signal.

[0045] The processing unit is used to extract features from the acquired first audio signal every N frames of audio signal, and to perform audio recognition based on the extracted first audio features.

[0046] Thirdly, embodiments of this application provide an electronic device including a processor and a memory, the memory for storing code instructions, and the processor for running the code instructions to perform the methods described in the first aspect or any possible implementation of the first aspect.

[0047] Fourthly, embodiments of this application provide a computer-readable storage medium storing a computer program or instructions that, when executed on a computer, cause the computer to perform the methods described in the first aspect or any possible implementation thereof.

[0048] Fifthly, embodiments of this application provide a computer program product including a computer program, which, when run on a computer, causes the computer to perform the methods described in the first aspect or any possible implementation thereof.

[0049] Sixthly, this application provides a chip or chip system including at least one processor and a communication interface. The communication interface and the at least one processor are interconnected via a circuit. The at least one processor is used to run computer programs or instructions to perform the methods described in the first aspect or any possible implementation of the first aspect. The communication interface in the chip can be an input / output interface, pins, or circuits, etc.

[0050] In one possible implementation, the chip or chip system described above in this application further includes at least one memory storing instructions. The memory can be an internal storage unit of the chip, such as a register or cache, or it can be a storage unit of the chip itself (e.g., read-only memory, random access memory, etc.).

[0051] It should be understood that the second to sixth aspects of this application correspond to the technical solutions of the first aspect of this application, and the beneficial effects achieved by each aspect and the corresponding feasible implementation are similar, and will not be repeated here. Attached Figure Description

[0052] Figure 1 A schematic diagram illustrating an application scenario provided in an embodiment of this application;

[0053] Figure 2 A schematic diagram illustrating another application scenario provided by an embodiment of this application;

[0054] Figure 3 This is a schematic diagram of a keyword recognition process;

[0055] Figure 4 This is a schematic diagram of a feature extraction process;

[0056] Figure 5 This is a schematic diagram illustrating a feature extraction process provided in an embodiment of this application;

[0057] Figure 6 A schematic block diagram of the hardware structure of the electronic device provided in the embodiments of this application;

[0058] Figure 7 A schematic diagram illustrating the interaction between software and hardware in an electronic device provided in an embodiment of this application;

[0059] Figure 8 A flowchart illustrating an audio processing method provided in an embodiment of this application;

[0060] Figure 9 A flowchart illustrating another audio processing method provided in an embodiment of this application;

[0061] Figure 10 A schematic diagram of a decoding process provided in an embodiment of this application;

[0062] Figure 11 This is a schematic diagram of an audio processing method provided in an embodiment of this application;

[0063] Figure 12 This is a schematic block diagram of an audio processing apparatus provided in an embodiment of this application. Detailed Implementation

[0064] To facilitate a clear description of the technical solutions in the embodiments of this application, some terms and technologies involved in the embodiments of this application will be briefly introduced below:

[0065] 1. Voice wake-up is an operation or technology that activates an electronic device by recognizing specific patterns in a user's voice. Voice wake-up technology relies on sophisticated speech recognition algorithms and machine learning to recognize and respond to audio signals.

[0066] 2. Front-end signal processing: This refers to a processing procedure in audio signal processing by electronic devices. Front-end signal processing may include, but is not limited to, one or more of the following: preprocessing, such as noise reduction and / or gain control, which can improve the quality of the audio signal; filtering, such as using high-pass filters, low-pass filters, or band-pass filters to remove noise; pre-emphasis processing, such as compensating for the attenuation of high-frequency components by enhancing the high frequencies of the audio signal; or framing and windowing processing, such as dividing the continuous audio signal into short frames and windowing each frame to reduce edge effects.

[0067] Frame splitting: This can be achieved by sliding a fixed-length window across the audio signal. The window length (i.e., frame length) is typically chosen to be 10–30 milliseconds, while the overlap between windows (i.e., frame shift) is usually about half the frame length. This overlapping method helps avoid abrupt changes between adjacent frames, making the transition between frames smoother.

[0068] Windowing refers to applying a window function to each frame after framing to reduce boundary effects at the ends of the frame, thereby improving the accuracy of spectral analysis. The function of the window function is to smooth the signal at the ends of the frame, gradually transitioning it to zero, thus reducing spectral leakage caused by abrupt frame truncation.

[0069] 3. Frame: In the field of audio signal processing, a frame refers to a short segment of a continuous audio signal. Each frame typically contains a certain number of sampling points, which represent the amplitude changes of the audio signal within that time period. A frame can be an audio signal of a certain duration, such as 10ms, 20ms, or 30ms. This duration is called the frame length, or the time length of each frame.

[0070] 4. Endpoint detection: Also known as voice activity detection (VAD), it is used to determine the start and end positions of each frame of audio signal so that the audio signal from the start to the end position can be processed subsequently. It can remove silent segments and noise segments in the audio signal, improving the efficiency and accuracy of audio processing.

[0071] 5. Feature extraction: This is a crucial step in audio signal processing and speech recognition. The purpose of feature extraction is to convert the raw audio signal into a set of feature parameters that can better represent speech information.

[0072] The speech features of an audio signal can be one or more of the following: mel-frequency cepstral coefficients (MFCC), perceptual linear prediction (PLP), filter bank energies (FBANK), linear predictive coding (LPC), linear predictive cepstral coefficients (LPCC), power-normalized cepstral coefficients (PNCC), formant frequencies, or pitch, etc.

[0073] It should be understood that speech features can also be called acoustic features, etc., and this application does not make a specific limitation on them.

[0074] Initials, finals, and syllables are fundamental concepts in the phonological system of Chinese (and some other languages), collectively constituting the pronunciation structure of Chinese. The following is a detailed explanation of these three concepts.

[0075] 6. Initials are the beginning part of a syllable and are usually a consonant. In Chinese pinyin, there are 21 initials (sometimes including the zero initial), which are: b, p, m, f, d, t, n, l, g, k, h, j, q, x, zh, ch, sh, r, z, c, and s. Initials play a role in guiding the pronunciation in a syllable. For example, in the syllable "妈 (mā)", the initial is "m".

[0076] 7. Finals are the main part of a syllable and usually include one or more vowels, and sometimes also include a consonant. In Chinese pinyin, there are 39 finals (including single finals, compound finals, and nasal finals), which are: Single finals: a, o, e, i, u, and ü; Compound finals: ai, ei, ui, ao, ou, iu, ie, üe, and er; and Nasal finals: an, en, in, un, ün, ang, eng, ing, and ong. Finals play a role in the main pronunciation in a syllable. For example, in the syllable "妈 (mā)", the final is "a".

[0077] 8. Syllables are the basic units of speech and are composed of initials and finals. In Chinese, one syllable usually corresponds to one Chinese character. Syllables can be divided into the following types: Simple syllables, that is, only finals without initials; Compound syllables, that is, composed of initials and finals; and Zero-initial syllables, that is, without initials and only finals.

[0078] 9. Phonemes are the smallest speech units in a language. They cannot be further decomposed but can be used to distinguish different words. For example, in English, the words "bat" and "pat" have different meanings only because the first phoneme is different ( / b / and / p / ). Phonemes can be divided into vowels and consonants.

[0079] 10. Phoneme sequence refers to the arrangement of a series of phonemes. Phonemes are the smallest speech units in speech and are the basic sound units for distinguishing different words. In a speech recognition system, the phoneme sequence is the output of the acoustic model, which represents the representation of the input audio signal at the phoneme level.

[0080] 11. Word sequence, in a speech recognition system, refers to a series of words finally output by the decoder. These words are the text representation of the input audio signal. The word sequence is the ultimate goal of speech recognition because it represents the system's understanding of the speech content.

[0081] In a speech recognition system, the concepts of initials, finals, and syllables are used to construct pronunciation dictionaries and acoustic models. The following are some specific applications.

[0082] Applications in pronunciation dictionaries: Pronunciation dictionaries contain combinations of initials and finals for each Chinese character; Applications in acoustic models: Training models to recognize different initials and finals, thereby recognizing complete syllables; or, Applications in decoders: Using combinations of initials and finals to generate possible word sequences, and using language models to select the most likely word sequences.

[0083] 12. An acoustic model is used to map the acoustic features of an audio signal (such as MFCC, PLP, etc.) to corresponding phonemes or syllables. Acoustic models are usually trained using machine learning algorithms. Commonly used models include Gaussian Mixture Model-Hidden Markov Model (GMM-HMM), Deep Neural Network (DNN), Convolutional Neural Network (CNN), and Long Short-Term Memory Network (LSTM).

[0084] 13. A language model is used to predict the probability of word sequences based on contextual information, thereby improving the accuracy of speech recognition. Language models are usually built based on statistical or neural network methods. Common language models include n-gram models and neural network-based language models, such as recurrent neural networks (RNNs), Transformer models, and long short-term memory (LSTM) models.

[0085] Among them, the n-gram model calculates the probability of a word sequence based on n-gram statistics of words or characters. For example, the trigram model considers the joint probability of the current word and the two preceding words.

[0086] Neural network language models: These models use neural networks (such as RNN, LSTM, and Transformer) to model the probability of word sequences, enabling them to capture longer contextual information.

[0087] 14. A pronunciation dictionary, also known as a dictionary or vocabulary list, is a mapping table. It maps each word to its corresponding phoneme sequence. In a speech recognition system, the pronunciation dictionary acts as a bridge, connecting the output (phoneme sequence) of the acoustic model with the input (word sequence) of the language model.

[0088] A pronunciation dictionary is typically a text file, with each line containing a word and its corresponding phoneme sequence. Examples: hello HH AH L OW, etc.

[0089] 15. A decoder is a key component whose main task is to convert the extracted speech features into text. The decoder combines an acoustic model, a language model, and a pronunciation dictionary, and finds the most likely word sequence through a search algorithm.

[0090] 16. Voiceless stops refer to the sounds where the glottis is closed and then quickly opened, and the air passes through the oral cavity without an obvious air burst.

[0091] Exemplarily, [b] in "爸(bà)" and "本(běn)"; [d] in "的(de)" and "大(dà)"; [g] in "个(gè)" and "高(gāo)".

[0092] 17. Aspirated stops refer to the sounds where the glottis is closed and then quickly opened, and the air passes through the oral cavity accompanied by an obvious air burst.

[0093] Exemplarily, [p] in "怕(pà)" and "朋友(péng yǒu)"; [t] in "他(tā)" and "天(tiān)"; [k] in "看(kàn)" and "空(kōng)".

[0094] 18. Voiceless affricates refer to the sounds where the glottis is closed and then quickly opened, and the air passes through the oral cavity and generates friction in the oral cavity without an obvious air burst.

[0095] Exemplarily, [z] in "资(zī)" and "字(zì)"; [zh] in "知(zhī)" and "中(zhōng)"; [j] in "鸡(jī)" and "家(jiā)".

[0096] 19. Aspirated affricates refer to the sounds where the glottis is closed and then quickly opened, and the air passes through the oral cavity and generates friction in the oral cavity, accompanied by an obvious air burst.

[0097] Exemplarily, [c] in "次(cì)" and "菜(cài)"; [ch] in "吃(chī)" and "车(chē)"; [q] in "七(qī)" and "(qián)前".

[0098] 20. Nasals refer to the sounds produced by the air passing through the nasal cavity. In Chinese, common nasals are [m], [n], and [ŋ].

[0099] Exemplarily, [m] in "妈妈(māma)"; [n] in "你(nǐ)".

[0100] 21. Fricative: It refers to the sound produced by friction when air flows through a narrow part of the oral cavity. In Chinese, common fricatives include [f], [s], [∫] (sh), and [x], etc.

[0101] 22. Tone sandhi: It refers to the phenomenon that the tone changes in a specific phonetic environment. Chinese is a tonal language, and the change of tone will affect the meaning of words.

[0102] Exemplarily, "不(bù)" changes to the second tone when it meets the fourth tone. For example, 不对(búduì). "一(yī)" changes to the second tone when it meets the fourth tone and changes to the fourth tone when it meets other tones. For example, 一样(yíyàng), 一天(yìtiān).

[0103] 23. Low-power mode or sleep state: It refers to the working state in which the device saves energy by reducing power consumption when it does not need to run at full power. These modes are usually used to extend the battery life or reduce energy consumption.

[0104] Among them, low-power mode (low power mode): The low-power mode is an energy-saving mode widely used in various devices, usually achieved by reducing the processor frequency, decreasing the screen brightness, turning off unnecessary wireless communication modules, etc.

[0105] The sleep state can also be understood as the sleep mode or hibernate mode.

[0106] Sleep mode (sleep mode): The sleep mode can reduce the power consumption of electronic devices. Usually, most activities are paused, and only the basic memory state is retained to quickly resume the working state.

[0107] Hibernate mode (hibernate mode): An electronic device in hibernate mode can save the current working state to the hard disk or other non-volatile memories, and then completely turn off the power. When resuming, the device continues to run from the saved state. Taking a smartwatch as an example, when the smartwatch is in low-power mode or sleep state, the smartwatch can but is not limited to having one or several of the following characteristics: the screen is turned off (screen off) or dimmed; the processor frequency is reduced; the sensors are turned off or the activities are reduced. For example, GPS, heart rate monitoring, etc. will be turned off or the sampling frequency is reduced; the wireless communication module is turned off or enters the low-power state. For example, Wi-Fi, Bluetooth, and cellular network modules may be turned off or enter the low-power mode; notifications and synchronizations are reduced; or, enter deep sleep. That is, in the case of long-term inactivity, the smartwatch may enter the deep sleep state, almost completely stopping all activities, and only retaining the basic clock function.

[0108] W24. Floor function. When s is a decimal number, Find the integer that is smaller than s and closest to s.

[0109] 25. Other terms

[0110] In the embodiments of this application, terms such as "first" and "second" are used to distinguish identical or similar items with substantially the same function and purpose. For example, "first chip" and "second chip" are used only to distinguish different chips and do not limit their order of execution. Those skilled in the art will understand that terms such as "first" and "second" do not limit the quantity or execution order, and that "first" and "second" do not necessarily imply that they are different.

[0111] It should be noted that, in the embodiments of this application, the terms "exemplary" or "for example" are used to indicate examples, illustrations, or descriptions. Any embodiment or design scheme described as "exemplary" or "for example" in this application should not be construed as being more preferred or advantageous than other embodiments or design schemes. Specifically, the use of terms such as "exemplary" or "for example" is intended to present the relevant concepts in a specific manner.

[0112] In this application embodiment, "at least one" refers to one or more, and "more than one" refers to two or more. "And / or" describes the relationship between related objects, indicating that three relationships can exist. For example, A and / or B can represent: A alone, A and B simultaneously, or B alone, where A and B can be singular or plural. The character " / " generally indicates that the preceding and following related objects are in an "or" relationship. "At least one of the following" or similar expressions refer to any combination of these items, including any combination of single or plural items. For example, at least one of a, b, or c can represent: a, b, c, ab, a--c, bc, or abc, where a, b, and c can be single or multiple.

[0113] 26. Electronic equipment

[0114] The electronic devices in this application embodiment may include handheld devices with facial recognition function, vehicle-mounted devices, etc. For example, some electronic devices include: mobile phones, tablets, PDAs, laptops, mobile internet devices (MIDs), wearable devices, virtual reality (VR) devices, augmented reality (AR) devices, wireless terminals in industrial control, wireless terminals in self-driving vehicles, wireless terminals in remote medical surgery, wireless terminals in smart grids, wireless terminals in transportation safety, wireless terminals in smart cities, wireless terminals in smart homes, cellular phones, cordless phones, session initiation protocol (SIP) phones, wireless local loop (WLL) stations, personal digital assistants (PDAs), handheld devices with wireless communication capabilities, computing devices or other processing devices connected to a wireless modem, in-vehicle devices, wearable devices, terminal devices in 5G networks, or future evolution of public land mobile communication networks. Terminal devices in a network (PLMN), etc., are not limited to this in the embodiments of this application.

[0115] By way of example and not limitation, in this embodiment, the electronic device can also be a wearable device. Wearable devices, also known as wearable smart devices, are a general term for devices that utilize wearable technology to intelligently design and develop everyday wearables, such as glasses, gloves, watches, clothing, and shoes. Wearable devices are portable devices that are worn directly on the body or integrated into the user's clothing or accessories. Wearable devices are not merely hardware devices, but also achieve powerful functions through software support, data interaction, and cloud interaction. Broadly speaking, wearable smart devices include those that are feature-rich, large in size, and can achieve complete or partial functions without relying on a smartphone, such as smartwatches or smart glasses, as well as those that focus on a specific type of application function and require the use of other devices such as smartphones, such as various smart bracelets and smart jewelry for vital sign monitoring.

[0116] Furthermore, in this embodiment of the application, the electronic device can also be a terminal device in the Internet of Things (IoT) system. IoT is an important part of the future development of information technology. Its main technical feature is to connect objects to the network through communication technology, thereby realizing an intelligent network of human-machine interconnection and object-to-object interconnection.

[0117] The electronic devices in the embodiments of this application may also be referred to as: terminal equipment, user equipment (UE), mobile station (MS), mobile terminal (MT), access terminal, user unit, user station, mobile station, mobile station, remote station, remote terminal, mobile device, user terminal, terminal, wireless communication equipment, user agent, or user device, etc.

[0118] In this embodiment, the electronic device or various network devices include a hardware layer, an operating system layer running on top of the hardware layer, and an application layer running on top of the operating system layer. The hardware layer includes hardware such as a central processing unit (CPU), a memory management unit (MMU), and memory (also called main memory). The operating system can be any one or more computer operating systems that implement business processing through processes, such as Linux, Unix, Android, iOS, or Windows. The application layer includes applications such as browsers, address books, word processing software, and instant messaging software.

[0119] To provide a convenient way for users to interact with computers, they can now wake up electronic devices with their voice and give voice commands to complete various tasks. This voice-based human-computer interaction function can also be called a voice interaction function or a smart voice function.

[0120] The following are combined with Figure 1 and Figure 2 Examples of application scenarios for voice interaction functions are provided.

[0121] Figure 1 This is a schematic diagram of application scenario 100 provided in an embodiment of this application. Application scenario 100 is a scenario of voice wake-up of an electronic device. The electronic device that implements the voice wake-up function can be, for example, Figure 1 The smartwatch shown.

[0122] like Figure 1As shown in interface (a), with the smartwatch constantly displaying the time, the user can speak a voice message containing a wake-up word, assuming the wake-up word is "Hello, Youyou". Correspondingly, the smartwatch can pick up the audio signal through its microphone, and upon recognizing the wake-up word "Hello, Youyou" in the audio signal, the smartwatch is awakened. The awakened smartwatch can then display... Figure 1 The interface (b) in the text.

[0123] It should be understood that Figure 1 For example only. Figure 1 The interface (a) can also be replaced with the interface of the smartwatch when it is off. That is, the smartwatch can also be woken up when the screen is off. This application embodiment does not specifically limit the state of the smartwatch before it is woken up by voice.

[0124] It should be noted that "Hello, Youyou" is for illustrative purposes only, and the wake word for electronic devices can also be "Hello, Youyou," "Hello, YOYO," "Little A, Little A," etc. This application does not impose specific limitations on this.

[0125] Figure 2 This is a schematic diagram of application scenario 200 provided in an embodiment of this application. Application scenario 200 is a scenario where a user navigates an electronic device using voice commands. The electronic device implementing this voice interaction function can be, for example, a... Figure 2 The phone shown.

[0126] like Figure 2 As shown in interface (a), the phone displays the interface of a navigation application. When the phone displays the navigation application interface, the user can give voice instructions to the electronic device to navigate; for example, the user's voice command might be "Navigate home." Correspondingly, the phone can pick up the audio signal through its microphone, and if the phone recognizes that the audio signal includes "Navigate home," it can begin navigation, with "home" as the destination. The geographical location corresponding to "home" can be pre-set by the user in the navigation application. For example, when navigating home, the phone might display... Figure 2 The interface (b) in the text.

[0127] It should be understood that Figure 2 The interface (b) in the example is only. During the navigation process of the electronic device, the interface of the navigation application can be updated in real time. This application does not specifically limit this.

[0128] Whether it's voice wake-up or other voice interaction functions, electronic devices need to perform keyword recognition on the acquired audio signal, that is, determine whether the audio signal contains predefined keywords. Each of these predefined keywords can be used to trigger different processes executed by the electronic device. Keyword spotting (KWS), also known as keyword detection, is a technology that detects predefined keywords from continuous audio signals and is one of the key technologies in the field of speech recognition. After detecting predefined keywords, the electronic device can perform tasks based on these keywords.

[0129] For example, in Figure 1 In the illustrated application scenario, the predefined keyword serves as the wake-up word for the electronic device. This wake-up word can be one set by the user beforehand. The electronic device is activated when it detects that the wake-up word is present in the audio signal. Figure 2 In the illustrated application scenario, the predefined keyword could be "navigate home." This keyword includes the user's instruction to the electronic device to navigate to the geographical location corresponding to "home." Therefore, the electronic device can navigate to the geographical location corresponding to home based on the recognized keyword.

[0130] Currently, electronic devices can perform keyword recognition in the following ways.

[0131] Figure 3 This is a schematic diagram illustrating the process of keyword recognition for an electronic device. For example... Figure 3 As shown, firstly, the electronic device can acquire audio signals through an audio acquisition module, and can convert the audio signals from sound signals into electrical signals. The audio acquisition module can be, for example, a microphone or a microphone.

[0132] Subsequently, the electronic device can perform front-end signal processing on the electrical signals. Front-end signal processing may include, for example, preprocessing, pre-emphasis processing, and framing and windowing processing. Preprocessing may include, but is not limited to, filtering, noise reduction, or silence removal. Through front-end signal processing, the electronic device can obtain multiple frames of audio signals.

[0133] Subsequently, the electronic device can perform endpoint detection on multiple frames of audio signals. This allows the electronic device to determine the start and end positions of each frame of audio signal within the multi-frame audio signal. For ease of description, the start position of each frame of audio signal can be called the start point of that frame, and the end position of each frame can be called the end point. Within each frame, the audio signal corresponding to the start point to the end point can be called the speech portion, and the remaining portion can be the non-speech portion, such as a silence segment or a noise segment.

[0134] Next, the electronic device extracts features from the speech portion of each frame of the multi-frame audio signal, thus obtaining speech features. Speech features can also be called audio features or acoustic features, which are the characteristics of the speech portion of each frame of the audio signal. Speech features may include, but are not limited to, one or more of the following features: MFCC, FBANK, PLP, LPCC, or PNCC, etc.

[0135] The electronic device can then input the speech features into the decoder, which can then output keywords.

[0136] The process of keyword recognition based on speech features by the decoder can be as follows: The decoder maps speech features to corresponding phonemes or syllables based on an acoustic model, thus obtaining a phoneme sequence. The decoder then converts the phoneme sequence into a word sequence based on a pronunciation dictionary. Next, the decoder predicts the reasonableness of the word sequence based on a language model and outputs the most reasonable (highest probability) word sequence. The decoder can then compare this most reasonable word sequence with at least one preset keyword in the electronic device. If the most reasonable word sequence includes a preset keyword 1 in the electronic device, the electronic device determines that keyword 1 has been detected.

[0137] Afterwards, the electronic device can execute the process corresponding to keyword 1. For example, if keyword 1 is the wake word of the electronic device, "Hello, Youyou," then the electronic device can be woken up.

[0138] exist Figure 3 During the process, the electronic device needs to extract features from each frame of the multi-frame audio signal. For example, as shown... Figure 4 As shown, assuming the frame length of the audio signal is 10ms, taking the first frame audio signal in a multi-frame audio signal as an example, the electronic device extracts features from each of the nine frames of audio signals.

[0139] It should be noted that although electronic devices extract features from audio signals in units of frames, the actual duration of each frame of audio signal from which features are extracted may be longer than the frame length. This is mainly due to one or more of the following reasons.

[0140] 1. Smoothing effect of window functions: When applying windowing, the window function (such as the Hamming or Hanning window) is not limited to the frame length but extends beyond the frame boundaries. This smoothing helps reduce abrupt changes at frame boundaries, thereby reducing spectral leakage. Therefore, when electronic devices perform feature extraction, they are actually extracting features from the entire audio signal segment covered by the window function. The length of this segment may be slightly longer than the frame length, especially when the window function has a long tail.

[0141] 2. Feature Extraction Algorithm Requirements: Some feature extraction algorithms may need to consider signals within a certain range before and after a frame to more accurately describe the characteristics of that frame. For example, when extracting features such as Mel-frequency cepstral coefficients (MFCCs) of an audio signal, signals from several frames before and after the current frame may be used for preprocessing (such as pre-emphasis, frame segmentation and windowing) and subsequent calculations. Thus, although the final speech features correspond to a specific frame, the length of the audio signals involved in the calculation process may be greater than the original length of that frame.

[0142] 3. Signal Processing Delay: In real-time or near-real-time audio processing systems, the complexity of signal processing algorithms and the limitations of hardware processing speed may introduce a certain processing delay. This delay may manifest as the fact that when processing the current frame, a portion of the audio signal from the future (i.e., later in time) has already been considered. However, this delay is usually not directly caused by the feature extraction algorithm itself, but rather determined by the processing flow of the entire system.

[0143] 4. Overlapping parts of frame shift: In practical applications, due to the existence of frame shift, there will be some overlap between adjacent frames, which helps to smoothly transition and process changes in audio signals.

[0144] It should be understood that Figure 4 The 11ms length of each audio frame used for feature extraction shown is merely an example. The length of each audio frame used for feature extraction by the electronic device can be longer or shorter; for example, it could be 12ms. Furthermore, the lengths of any two audio frames used for feature extraction can be the same or different. For example, the duration for feature extraction of the first audio frame could be 11ms, and the duration for feature extraction of the second audio frame could be 12ms. This does not constitute a limitation on the embodiments of this application.

[0145] Therefore, electronic devices need to extract features from each frame of a multi-frame audio signal to obtain the speech features of each frame; and then perform keyword recognition based on the speech features of each frame. This results in high power consumption for the electronic devices, which may in turn affect their performance.

[0146] For example, taking a smartwatch as an electronic device, a smartwatch may include two different types of microcontroller units (MCUs) or system-on-chips (SoCs). Keyword recognition can be performed in one of the MCUs or SoCs. Since this MCU or SoC is already responsible for operational health, trajectory algorithms, and announcements, it provides less memory for keyword recognition.

[0147] Therefore, there is an urgent need to provide a method to reduce the power consumption of electronic devices in extracting features from audio signals, thereby reducing the power consumption of electronic devices in implementing voice interaction functions such as voice wake-up.

[0148] In view of this, this application provides an audio processing method in which a frame skipping strategy can be preset in the electronic device. In different application scenarios, the electronic device can extract features from Y frames of audio signal every X frames of audio signal based on the frame skipping strategy corresponding to the application scenario, where X and Y are positive integers. This eliminates the need for the electronic device to extract features from all frames of audio signal, resulting in lower power consumption for feature extraction and thus lower power consumption when implementing voice interaction functions such as voice wake-up.

[0149] It should be understood that a frame skipping strategy may include, for example, X frames, meaning that the electronic device extracts features from the acquired audio signal every X frames of audio signal. Alternatively, the frame skipping strategy can also be understood as: in every X+Y frames of audio signal, no feature extraction is performed on X frames of audio signal, and feature extraction is performed on Y frames of audio signal. If Y is 1, then the frame skipping strategy is to perform feature extraction on 1 frame of audio signal every X frames of audio signal (i.e., no feature extraction is performed on X frames of audio signal or every X frames of audio signal).

[0150] For example, let's continue with the scenario of voice wake-up of an electronic device. Suppose the wake-up phrase of the electronic device is "Hello Youyou". And the frame skipping strategy corresponding to the preset voice wake-up scenario in the electronic device is to extract features from one frame of audio signal every two frames (or skip two frames). Alternatively, this frame skipping strategy can also be represented as two frames.

[0151] and Figure 4 The feature extraction process shown is similar. Taking the audio signals from frame 1 to frame 9 in an audio signal as an example, if the electronic device extracts features from one frame of audio signal every two frames, the feature extraction process of the electronic device is as follows: Figure 5 As shown. The electronic device does not extract features from the first and second frames of audio signals, but extracts features from the third frame; it does not extract features from the fourth and fifth frames, but extracts features from the sixth frame; it does not extract features from the seventh and eighth frames, but extracts features from the ninth frame, and so on. In the multi-frame audio signals acquired by the electronic device, it can skip the 3×a-2 and 3×a-1 frames, and only extract features from the 3×a frame, where a is a positive integer.

[0152] It should be understood that Figure 5For illustrative purposes only, an electronic device may also extract features from the audio signal of frame 3×a-2 but not from the audio signals of frames 3×a-1 and 3×a; or, the electronic device may also extract features from the audio signal of frame 3×a-1 but not from the audio signals of frames 3×a-2 and 3×a, etc. For the sake of brevity, these examples will not be shown here.

[0153] It should also be understood that, in the embodiments of this application, skipping the c-frame audio signal means not performing feature extraction on the c-frame audio signal. For the sake of brevity, this will not be elaborated further below.

[0154] In this way, electronic devices can extract features from some frames of audio signals, resulting in lower power consumption for keyword recognition and thus lower power consumption for voice interaction functions such as voice wake-up.

[0155] The following is combined with Figures 6 to 10 The technical solutions of this application and how they solve the aforementioned technical problems are described in detail with specific embodiments. The following specific embodiments can be implemented independently or in combination with each other. Identical or similar concepts or processes may not be described again in some embodiments.

[0156] The embodiments shown in this application can be executed by an electronic device, which may be a terminal device. The specific form and number of the devices shown are merely examples and should not constitute any limitation on the implementation of the methods provided in this application. The audio processing method of the embodiments of this application will be described in detail below from the perspective of interaction between internal modules of an electronic device.

[0157] It should be understood that an electronic device can be the electronic device itself, or a chip, chip system, or processor that supports the electronic device in implementing audio processing methods, or a logic module or software that can implement all or part of the functions of the electronic device.

[0158] To facilitate understanding this solution, we will first combine... Figure 6 Describe the hardware structure of the electronic device.

[0159] Figure 6 This is a schematic diagram of the structure of the electronic device 100 provided in an embodiment of this application. Figure 6As shown, the electronic device 100 may include a processor 110, an external memory interface 120, an internal memory 121, a universal serial bus (USB) interface 130, a charging management module 140, a power management module 141, a battery 142, an antenna 1, an antenna 2, a mobile communication module 150, a wireless communication module 160, an audio module 170, a speaker 170A, a receiver 170B, a microphone 170C, a headphone jack 170D, a sensor module 180, buttons 190, a motor 191, an indicator 192, a camera 193, a display screen 194, and a subscriber identification module (SIM) card interface 195, etc. The sensor module 180 may include a pressure sensor 180A, a gyroscope sensor 180B, a barometric pressure sensor 180C, a magnetic sensor 180D, an accelerometer sensor 180E, a distance sensor 180F, a proximity sensor 180G, a fingerprint sensor 180H, a temperature sensor 180J, a touch sensor 180K, an ambient light sensor 180L, a bone conduction sensor 180M, etc.

[0160] It is understood that the structures illustrated in the embodiments of this application do not constitute a specific limitation on the electronic device 100. In other embodiments of this application, the electronic device 100 may include more or fewer components than illustrated, or combine some components, or split some components, or have different component arrangements. The illustrated components may be implemented in hardware, software, or a combination of software and hardware.

[0161] Processor 110 may include one or more processing units, such as application processors (APs) and digital signal processors (DSPs). The application processor includes the application layer, application framework layer, algorithm layer, and kernel layer of the software architecture adopted by electronic device 500. Different processing units may be independent devices or integrated into one or more processors.

[0162] The audio module 170 is used to convert digital audio signals into analog audio signals for output, and also to convert analog audio inputs into digital audio signals. The audio module 170 can also be used for encoding and decoding audio signals. In some embodiments, the audio module 170 may be located in the processor 110, or some functional modules of the audio module 170 may be located in the processor 110.

[0163] Microphone 170C, also known as a "microphone" or "voice transducer," is used to convert sound signals into electrical signals. When making a phone call or sending a voice message, the user can speak by bringing their mouth close to microphone 170C, inputting the sound signal into microphone 170C. Terminal device 100 may be equipped with at least one microphone 170C. In some embodiments, terminal device 100 may be equipped with two microphones 170C, which, in addition to collecting sound signals, can also perform noise reduction. In other embodiments, terminal device 100 may be equipped with three, four, or more microphones 170C, which can collect sound signals, reduce noise, identify the sound source, and perform directional recording, etc.

[0164] In one possible implementation, the DSP can acquire audio signals from the microphone 170C; the AP can determine a frame skipping strategy based on the application scenario and instruct the DSP on the frame skipping strategy; then, the DSP can extract keywords from the acquired audio signals based on the frame skipping strategy from the AP; and / or, the DSP can acquire audio signals from the microphone 170C; extract features from the audio signals based on the preset frame skipping strategy 1 to obtain voice features 1; and determine whether the audio signals contain the wake-up word of the electronic device 100 based on the voice features 1.

[0165] The AP and DSP can exchange information through methods such as calling interfaces.

[0166] The controller can generate operation control signals based on the instruction opcode and timing signals to control the fetching and execution of instructions. The processor 110 can also include a memory for storing instructions and data.

[0167] In some embodiments, the processor 110 may include one or more interfaces. It is understood that the interface connection relationships between the modules illustrated in the embodiments of this application are merely illustrative and do not constitute a structural limitation of the electronic device 500.

[0168] The power management module 141 is used to connect the battery 142, the charging management module 140, and the processor 110.

[0169] The software system of an electronic device can adopt a layered architecture, event-driven architecture, microkernel architecture, microservice architecture, or cloud architecture. This embodiment of the invention uses a layered architecture system as an example to illustrate the software structure of an electronic device.

[0170] For example, taking electronic devices as smart wearable devices (such as smartwatches) as an example, Figure 7 This is a schematic diagram illustrating the interaction between the software and hardware of a smart wearable device, as provided in an embodiment of this application.

[0171] like Figure 7As shown, the layered architecture divides the software into several layers, each with a clear role and function. Layers communicate with each other through software interfaces. In some embodiments, the software architecture of a smart wearable device can be divided into four layers, from top to bottom: the application (APP) layer, the application framework layer, the algorithm layer, and the kernel layer. These four layers can be configured within the application layer (AP).

[0172] In addition to the four layers mentioned above, smart wearable devices can also include a hardware layer, which may include, for example, an audio acquisition module. The audio acquisition module could be a microphone 170C. The microphone 170C can be understood as hardware within the smart wearable device, used to acquire audio signals. The smart wearable device can also include a DSP, which can interact with the AP (Access Controller).

[0173] The specific software architecture in AP is as follows.

[0174] The application layer, also known as the application layer, can include a series of applications. For example, an application may include: a watch face application, a health application, a navigation application, a music application, an office application, a calendar application, a user interface (UI) module, and a software development kit (SDK), etc.

[0175] The first SDK is used to identify the state of the smart wearable device, such as determining whether the device is in a preset state. If the device is in the preset state, the decision module, based on instructions from other applications, instructs the DSP to send a frame skipping strategy. These other applications could be navigation apps, music apps, or office applications, etc.

[0176] Preset states can be, for example, in-vehicle status, in a meeting, or in motion. In-vehicle status can be understood as the smart wearable device recognizing that the user is in a vehicle, such as when the user is wearing the smart wearable device while driving a vehicle.

[0177] Optionally, the first SDK can determine whether the smart wearable device is in a vehicle-mounted state through any of the following methods: the first SDK can acquire sensor data and determine whether the smart wearable device is in a vehicle-mounted state based on the sensor data; the first SDK can detect the movement status and location of the smart wearable device through a GPS module and determine whether the smart wearable device is in a vehicle-mounted state based on the movement status and location of the smart wearable device; the first SDK can acquire the Bluetooth connection status of the smart wearable device and determine whether the smart wearable device is in a vehicle-mounted state based on the Bluetooth connection status, for example, if it is determined that the smart wearable device is connected to the vehicle's Bluetooth system, it is determined that the smart wearable device is in a vehicle-mounted state; the first SDK can acquire the Wi-Fi connection status of the smart wearable device and determine whether the smart wearable device is in a vehicle-mounted state based on the Wi-Fi connection status, for example, if it is determined that the smart wearable device is connected to the vehicle's Wi-Fi, it is determined that the smart wearable device is in a vehicle-mounted state; or, the first SDK can determine whether the smart wearable device is in a vehicle-mounted state by determining whether the smart wearable device is plugged into the vehicle's universal serial bus (USB), for example, if the smart wearable device is plugged into the vehicle's USB port, it is determined that the smart wearable device is in a vehicle-mounted state.

[0178] It should be understood that the aforementioned sensor data may be, for example, data acquired by an accelerometer and / or data acquired by a gyroscope. Based on the sensor data, the first SDK can determine whether the smart wearable device conforms to the state during vehicle operation.

[0179] Meeting status can be understood as the smart wearable device recognizing that the user is in a meeting, such as when the user presents a document through the smart wearable device during a meeting.

[0180] Optionally, the first SDK can determine whether the smart wearable device is in a meeting state through any of the following methods: The first SDK can obtain time and calendar information, which may include the date and calendar events. Calendar events can be events pre-set by the user, such as a meeting marked in the calendar application for 2023-04-15 from 14:00 to 16:00. In this way, the first SDK can determine the current time based on the date in the time and calendar information, and determine whether the smart wearable device is in a meeting state by judging whether the current time falls within the time period corresponding to the calendar event (which may be a meeting event marked in the calendar). For example, if the current time is determined to be within the time period corresponding to the calendar event, the smart wearable device is determined to be in a meeting state.

[0181] The state of motion can be understood as the smart wearable device recognizing that the user is engaged in some form of exercise, such as running, cycling, or fitness.

[0182] Optionally, the first SDK can determine whether the smart wearable device is in motion by any of the following methods: the first SDK can acquire sensor data and determine whether the smart wearable device is in motion based on the sensor data; the first SDK can detect the user's heart rate by a heart rate sensor and determine whether the smart wearable device is in motion based on the user's heart rate; or, the first SDK can detect the movement status and location of the smart wearable device by a GPS module and determine whether the smart wearable device is in operation based on the movement status and location of the smart wearable device.

[0183] In addition to the methods described above, users can optionally manually set the state of the smart wearable device, such as manually setting it to be in a vehicle state, meeting state, or sports state, so that the first SDK can determine the state of the smart wearable device. Alternatively, the first SDK can also determine whether the smart wearable device is in a preset state through other methods, which are not specifically limited in this embodiment.

[0184] In the default state, it may be inconvenient for users to manually operate the smart wearable device. Therefore, in the default state, the first SDK can send a frame skipping strategy to the DSP through the decision module, so that the DSP can continue to perform keyword recognition, and then the electronic device can execute the corresponding process based on the recognized keywords.

[0185] For example, in vehicle mode, the navigation application instructs the first SDK navigation application to start; in response to the instruction of the navigation application, the decision module can instruct the DSP on the frame skipping strategy corresponding to the navigation application scenario, so that the DSP can identify keywords in the navigation application scenario based on the frame skipping strategy corresponding to the navigation application scenario.

[0186] Navigation apps can be used for navigation. For example, a navigation app can display a navigation interface based on the user's voice instructions.

[0187] In some possible implementations, after the navigation application is launched, the navigation application may instruct the first SDK navigation application to launch.

[0188] Music apps can be used to play music. For example, based on a user's voice commands, a music app can start playing music, pause music, switch to the next song, or switch to the previous song.

[0189] In some possible implementations, after the music application is launched, the music application can instruct the first SDK music application to launch.

[0190] Office applications can be used to present documents. For example, based on a user's voice command, an office application can switch the presented document to the next page; or based on a user's voice command, an office application can switch the presented document to the previous page, etc.

[0191] In some possible implementations, after the office application is launched, the office application can instruct the first SDK office application to launch.

[0192] The UI module can be used to manage the interface displayed on a smart wearable device. In this embodiment, the UI module can be used to display the interface of the smart wearable device after it is woken up, based on instructions from the power management service, such as displaying... Figure 1 The interface (b) in the text.

[0193] The application framework layer provides application programming interfaces (APIs) and a programming framework for applications in the application layer. The application framework layer includes some predefined functions. For example... Figure 7 As shown, the application framework layer may include a decision-making module and a speech recognition module, etc.

[0194] The decision module can obtain the frame skipping strategy corresponding to each application scenario and instruct the DSP on the frame skipping strategy corresponding to the current application scenario.

[0195] For example, in a navigation application scenario, the navigation application instructs the first SDK to start the navigation application; the first SDK may instruct the decision module to transmit the frame skipping strategy corresponding to the navigation application scenario to the DSP.

[0196] Optionally, the electronic device may be equipped with a frame skipping strategy configuration library, which includes at least one application scenario and a frame skipping strategy corresponding to each application scenario.

[0197] The decision-making module can determine the frame skipping strategy corresponding to the current application scenario based on the frame skipping strategy configuration library. For example, if the current scenario is navigation, the decision-making module can determine the frame skipping strategy corresponding to the navigation scenario based on the frame skipping strategy configuration library.

[0198] For example, the frame skipping strategy configuration library includes at least one application scenario and the frame skipping strategy corresponding to each application scenario in the at least one application scenario, as shown in Table 1.

[0199] Table 1

[0200] Application scenarios Frame skipping strategy Navigation application scenarios Frame skipping strategy a Music application scenarios Frame skipping strategy b Office application scenarios Frame skipping strategy c

[0201] The navigation application scenario can also be understood as a scenario where the navigation application runs in the foreground. For example, the navigation application instructs the first SDK to launch itself, and thus the navigation application runs in the foreground. Similarly, the music application scenario can be understood as a scenario where the music application runs in the foreground. For example, the music application instructs the first SDK to launch itself, and thus the music application runs in the foreground. The office application scenario can also be understood as a scenario where office-related applications run in the foreground. For example, the office-related application instructs the first SDK to launch itself, and thus the office-related application runs in the foreground.

[0202] A frame skipping strategy can include q frames, meaning that feature extraction is performed on one frame of audio signal every q frames, where q is a positive integer; for example, frame skipping strategy a can include q1 frames, q2 frames, and q3 frames. Alternatively, a frame skipping strategy can also include q frames and p frames, meaning that feature extraction is performed on p frames of audio signal signal every q frames, where q and p are positive integers; for example, frame skipping strategy a can include q1 and p1 frames, q2 and p2 frames, and q3 and p3 frames, etc. q, q1, q2, q3, p, p1, p2, and p3 are all positive integers.

[0203] The speech recognition module, also known as the speech recognition engine, is used to perform speech recognition on audio signals from the audio input management service. Through speech recognition, the module can determine the user's instructions and, based on those instructions, instruct the appropriate modules to execute corresponding processes. For example, if the speech recognition module determines that the user intends to set an alarm, it can then instruct the alarm clock application to do so.

[0204] The system service layer provides the core services and functions of the operating system, supporting application execution. It can be used to allocate hardware resources for smart wearable devices. It also provides interfaces for the application layer and application framework layer to access the underlying hardware. The system service layer may include power management services and audio input management services, among others.

[0205] The power management service is responsible for coordinating and managing the power state of smart wearable devices. It can be used to manage the sleep and wake-up states of smart wearable devices, enabling them to enter a low-power mode (sleep state) when not in use and to quickly wake up when needed (e.g., to handle wake-up events).

[0206] For example, upon detecting a wake-up event, the power management service can reinitialize disabled or suspended peripherals, such as displays, sensors, and communication modules; restart or resume suspended system services, such as audio input management services; and instruct the UI module to switch interfaces, such as displaying the interface of the electronic device after it has been woken up. Figure 1 The interface (b), etc.

[0207] In some possible implementations, the power management service can also be replaced by a low-power management module or a low-power management service, etc. In other words, the low-power management service can also exist in the form of a power management service.

[0208] The audio input management service is responsible for managing the audio input function of smart wearable devices, including managing microphones and processing audio data. For example, the functions of the audio input management service may include: managing the enabled and disabled states of microphone 170C, turning microphone 170C on or off as needed; acquiring audio signals from the audio input management module and transmitting them to applications or services that require audio input, etc.

[0209] The kernel layer can be considered the layer between hardware and software. The kernel refers to system software that provides functions such as a hardware abstraction layer, disk and file system control, and multitasking. The kernel is the core of an operating system and its most fundamental part. It manages system processes, memory, device drivers, file and network systems, and determines system performance and stability. It is part of the software that provides secure access to computer hardware for numerous applications; this access is limited, and the kernel determines when a program can operate on a particular part of the hardware and for how long.

[0210] This kernel layer can be the operating system kernel (OS kernel).

[0211] The kernel layer may include a power management unit and an audio input management module.

[0212] The power management unit can indicate a wake-up event to the power management service based on the wake-up signal from the DSP, so that the power management service can process the wake-up event.

[0213] The audio input management module can acquire audio signals from the DSP and transmit these audio signals to the audio input management service.

[0214] The above describes the layered architecture included in the AP. For the DSP that interacts with the AP, a keyword recognition module can be configured within the DSP.

[0215] The keyword recognition module can acquire audio signals and frame skipping strategies from the AP (decision module), and perform keyword recognition on the audio signals based on the frame skipping strategies; and / or, in the application scenario of voice wake-up, the keyword recognition module can acquire audio signals and perform keyword recognition on the audio signals based on the preset frame skipping strategy 1.

[0216] Optionally, the keyword recognition module may include multiple modules, such as keyword recognition module A, keyword recognition module B, and keyword recognition module C. Different modules among these can be used to detect different keywords. For example, keyword recognition module A detects wake words, that is, whether the audio signal contains a wake word; keyword recognition module B detects "play music," that is, whether the audio signal contains "play music," etc.

[0217] It should be understood that Figure 7 For illustrative purposes only, the application processing unit (AP) of a smart wearable device may include more or fewer layers, such as an algorithm layer and a hardware abstraction layer. In this case, the frame skipping strategy configuration library can be set in the algorithm layer. Each layer may also include more or fewer modules; for example, the application framework layer may also include rendering components (UIKit). This application does not impose any specific limitations on this.

[0218] It should be noted that when the electronic device is not a smart wearable device, the layered architecture of the electronic device can also be referenced. Figure 7 For the sake of brevity, they will not be shown one by one here.

[0219] Below, in Figure 7 Based on the software and hardware modules shown, and taking the wake-up word "Hello, Youyou" as an example, this paper provides a detailed explanation of the audio processing flow of the electronic device during the voice wake-up process.

[0220] Figure 8 This is a flowchart illustrating an audio processing method 800 provided in an embodiment of this application. Figure 8 As shown, method 800 includes the following steps:

[0221] S801, DSP acquires audio signal from microphone 1.

[0222] Among them, the audio signal 1 obtained by the DSP from the microphone can be understood as an electrical signal.

[0223] S802 and DSP extract features from audio signal 1 based on the preset frame skipping strategy 1 to obtain speech features 1.

[0224] For example, S802 can be implemented in the following way: the DSP performs front-end signal processing on the audio signal 1 to obtain multiple frames of audio signal 1; the DSP performs feature extraction on the multiple frames of audio signal 1 based on a preset frame skipping strategy 1 to obtain speech features 1.

[0225] It should be understood that the implementation method of the DSP performing front-end signal processing on audio signal 1 to obtain multiple frames of audio signal 1 is similar to... Figure 3The implementation method of obtaining multi-frame audio signals by front-end signal processing of electrical signals in electronic devices is similar, as described above, and will not be repeated here.

[0226] Frame skipping strategy 1 is preset in the electronic device. Frame skipping strategy 1 can perform feature extraction on Y frames of speech signal every X frames of audio signal. X and Y are positive integers. For example, X can be 2, and Y can be 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19 ...

[0227] For example, frame skipping strategy 1 can be to extract features from one frame of audio signal every two frames (skipping two frames). The method of feature extraction for multiple frames of audio signal 1 is similar to... Figure 5 The method shown is similar, and you can refer to the description above.

[0228] S803, the DSP performs keyword recognition based on speech feature 1. That is, the DSP determines the word sequence based on speech feature 1 and determines whether the word sequence includes the wake-up word "Hello, Youyou"; if it does, then execute S804 to S815; if it does not, this process ends, or it can continue to perform keyword recognition on other audio signals from the microphone.

[0229] It should be understood that the implementation method of S803 is different from... Figure 3 The electronic device in China uses a decoder to recognize keyword 1 based on speech features in a similar way, as described above, and will not be repeated here.

[0230] The S804 and DSP send a wake-up signal to the power management unit.

[0231] S805. Based on the wake-up signal, the power management unit indicates a wake-up event to the power management service.

[0232] S806, Power Management Service instructs the UI module to update the interface. For example, the UI module can switch the interface of the electronic device to... Figure 1 The interface shown in (b) allows users to determine that the electronic device has been woken up based on the switched interface.

[0233] S807, the power management service instructs the audio input management service to receive and process subsequent audio signals.

[0234] The subsequent audio signal can be understood as the audio signal collected by the microphone after collecting audio signal 1.

[0235] It should be understood that S807 can be executed before or after S806, and S807 can also be executed in parallel with S806. This application does not make any specific limitations on this.

[0236] S808, Audio Input Management Service instructs the audio input management module to collect subsequent audio signals.

[0237] In this way, the audio input management module can obtain subsequent audio signals from the DSP, enabling the electronic device to perform speech recognition based on the subsequent audio signals.

[0238] It should be understood that after waking up the electronic device, the user typically continues to instruct the device to perform the next operation via voice. Therefore, after waking up the electronic device via S801 to 808, the electronic device can continue to collect subsequent audio signals and perform voice recognition on these signals to execute the next operation according to the user's instructions. For example, after being woken up, the electronic device can display... Figure 1 In the interface (b), the electronic device can continue to collect audio signals; for example, if the audio signal subsequently collected by the electronic device is "navigate home", the electronic device can navigate to the geographical location corresponding to "home" through the navigation application.

[0239] The process of electronic devices acquiring subsequent audio signals and performing speech recognition can be shown in S809 to S815.

[0240] S809, DSP acquires audio signals from the microphone 2.

[0241] The S810 and DSP process the audio signal 2 to obtain the processed audio signal 2.

[0242] Process 1 may include, but is not limited to, one or more of the following: noise reduction, such as removing background noise to improve the clarity of audio signal 2; echo cancellation, such as eliminating echoes caused by speaker sounds picked up by the microphone; gain control, such as dynamically adjusting the gain of audio signal 2 to keep it within a suitable range; signal enhancement, such as enhancing a portion of the audio signal in audio signal 2 to improve the clarity of audio signal 2; frequency analysis and filtering, such as analyzing the frequency components of audio signal 2 and performing necessary filtering; or feature extraction, etc.

[0243] Thus, because the DSP is efficient, low-power, and capable of real-time processing audio signals, the processing of audio signal 2 by the DSP can reduce the subsequent processing of audio signal 2 by the AP, thereby reducing the power consumption of the AP.

[0244] S811, the audio input management module obtains the processed audio signal 2 from the DSP.

[0245] S812, the audio input management module instructs the audio signal 2 after processing 1 to the audio input management service.

[0246] S813, The audio input management service processes the audio signal 2 after processing 1 to obtain the audio signal 2 after processing 2.

[0247] Processing 2 may include, but is not limited to, one or more of the following: further noise reduction and echo cancellation; advanced gain control; advanced signal enhancement; format conversion and data processing, i.e., converting the audio signal 2 after processing 1 into a format suitable for subsequent processing or transmission, such as sampling rate conversion, quantization bit adjustment, etc.; audio buffer management, i.e., managing the buffer of the audio signal to ensure the continuity and stability of the data stream; or feature extraction and processing, etc.

[0248] It should be understood that there are the following differences between Process 1 and Process 2: Further noise reduction and echo cancellation in Process 2 are optional; for example, if noise reduction and echo cancellation in Process 1 meet the requirements, further noise reduction and echo cancellation may not be included in Process 2. If noise reduction and echo cancellation in Process 1 do not meet the requirements, the audio input management service may perform further noise reduction and echo cancellation. Compared to gain control in Process 1, advanced gain control in Process 2 involves more complex gain adjustments to adapt to different application scenarios. Compared to signal enhancement in Process 1, advanced signal enhancement in Process 2 uses more complex algorithms to enhance the signal and improve the accuracy of speech recognition. Compared to feature extraction in Process 1, feature extraction and processing in Process 2 involves more complex feature extraction and processing to provide high-quality input data for the speech recognition engine.

[0249] Therefore, it can be seen that the audio signal processing 1 performed by the DSP and the audio input management service processing 2 are similar in some aspects, but their specific responsibilities and processing depths differ. Compared to the DSP's processing 1, the audio input management service's processing 2 is usually designed to meet the needs of specific applications, such as speech recognition.

[0250] It should be noted that the feature extraction involved in Process 1 and Process 2 may not be based on a preset frame skipping strategy, that is, the electronic device performs feature extraction on each frame of audio signal.

[0251] It should be understood that in some possible implementations, S810 may also be optional, meaning that the DSP may not process the audio signal 2.

[0252] S813, The audio input management service instructs the speech recognition module to process the audio signal 2 after processing 2.

[0253] S814, the speech recognition module performs speech recognition on the processed audio signal 2.

[0254] In this way, the voice recognition module can determine the command corresponding to audio signal 2. For example, if audio signal 2 could be "Please play music," then the electronic device can determine that the command corresponding to audio signal 2 is "Play music." The electronic device can then play music through a music application.

[0255] It is understood that electronic devices can acquire audio signals through the DSP in low-power mode (or sleep mode) and detect whether the audio signal contains a wake-up word. Alternatively, electronic devices can continuously acquire audio signals through the DSP and detect whether the audio signal contains a wake-up word. Alternatively, electronic devices can acquire audio signals through the DSP and detect whether the audio signal contains a wake-up word in other situations. This application does not specifically limit the scenarios for electronic devices to implement voice wake-up functions.

[0256] When an electronic device is in a low-power mode (or sleep state), and the DSP acquires an audio signal and detects whether the audio signal contains a wake-up word, the DSP can determine that the electronic device is in a low-power mode (or sleep state) using any of the following methods. That is, it can be triggered by any of the following methods 800.

[0257] Method 1: The power management unit instructs the DSP electronic device to be in low-power mode (or sleep mode).

[0258] The power management unit (PMU) is responsible for managing the power status of the device. The PMU can instruct the DSP electronic device to enter a low-power mode (or sleep mode) via the power management module, thus enabling the DSP to begin keyword recognition.

[0259] Method 2: Determine whether the electronic device is in low-power mode (or sleep mode) by using a low-power timer.

[0260] Electronic devices can be equipped with low-power timers. DSPs can use low-power timers to periodically check the system's power status.

[0261] Method 3: Before entering low-power mode (or sleep mode), the operating system can send a specific command to the DSP, instructing the DSP to start performing keyword recognition.

[0262] Method 4: Instruct the DSP electronic device to enter a low-power mode (or sleep state) via hardware signals.

[0263] Exemplarily, when the electronic device enters the low-power mode (or sleep state), the state of the hardware signal changes, and the DSP can start performing keyword recognition by detecting this change.

[0264] In method 800, when the wake-up word of the electronic device is "Hello, Youyou", assuming that the preset frame-skipping strategy 1 in the electronic device is to extract features from 1 frame of audio signal every 2 frames of audio signals. The implementation of extracting features from 1 frame of audio signal every 2 frames of audio signals is similar to Figure 5 the manner shown. Although the electronic device does not extract features from all frames of audio signals, the accuracy of the electronic device in recognizing whether the audio signal includes "Hello, Youyou" still meets the requirements. Below, the determination method of the preset frame-skipping strategy 1 will be described in detail.

[0265] Table 2 shows the pronunciation durations corresponding to the initials of common pronunciation methods in Chinese. The durations shown in Table 2 can be data collected through experiments or empirical data.

[0266] Table 2

[0267] initial consonant in monosyllabic initials in words initial consonants in singing Duration / ms Change ratio Change ratio unaspirated plosives 2-10 0.9 1.1 Aspirated plosive 40-70 0.8 1.5 unaspirated affricates 20-70 0.7 1.5 Aspiration fricative 75-140 0.8 1.5 nasal 30-70 0.7 1.6 fricative 70-130 0.8 1.6 Voice change 40-100 0.6 1.6

[0268] Table 3 shows the pronunciation durations of the finals in the high-level tone and monosyllables. Among them, the high-level tone can be one of the tones in Mandarin Chinese, belonging to one of the four tones, and the high-level tone is the first tone among the four tones.

[0269] Table 3

[0270]

[0271] Combining Table 2 and Table 3, it can be seen that in Chinese pronunciation, the pronunciation duration of the finals is usually longer than that of the initials.

[0272] Combining Table 2 or Table 3, the pronunciation durations of the finals and initials of some Chinese words can be determined. In "Hello, Youyou", the initial of "你 (nǐ)" is n, which belongs to a nasal sound. Combining Table 2, the pronunciation duration of the initial in "你 (nǐ)" is 30 - 70 ms. That is, in the process of the electronic device recognizing "你 (nǐ)", the time interval between feature extractions of two adjacent frames needs to be less than 30 ms, otherwise, the electronic device may not detect the initial of "你 (nǐ)" and thus cannot recognize "你 (nǐ)".

[0273] Since the pronunciation duration of the finals is usually longer than that of the initials, when the time interval between feature extractions of two adjacent frames is less than but exceeds 30 ms, the electronic device can detect the finals of "你 (nǐ)".

[0274] The initial consonant of "好" is h, which belongs to fricatives. Referring to Table 2, the pronunciation duration of the initial consonant in "好" is 70 - 130 ms. That is, during the process of the electronic device recognizing "好", the time interval for feature extraction between two adjacent frames needs to be less than 70 ms. Otherwise, the electronic device may not detect the initial consonant of "好" and thus cannot recognize "好".

[0275] Since the pronunciation duration of the final is usually longer than that of the initial consonant, when the time interval for feature extraction between two adjacent frames is less than 70 ms, the electronic device can detect the final of "好".

[0276] "悠" belongs to the zero initial syllable y, and the final of "悠" is ou. Referring to Table 3, the pronunciation duration of the final of "悠" is 405 ms. That is, during the process of the electronic device recognizing "悠", the time interval for feature extraction between two adjacent frames needs to be less than 405 ms. Otherwise, the electronic device may not detect the final of "悠" and thus cannot recognize "悠".

[0277] In order for the electronic device to detect each syllable in "你好,悠悠", the time interval for feature extraction between two adjacent frames needs to be less than 30 ms. In this way, in the time interval for feature extraction between two adjacent frames, the initial and final of "你" will not be missed, nor will the initial and final of "好" be missed, and the final of "悠" will not be missed either. Therefore, when the time interval for feature extraction between two adjacent frames is less than 30 ms, the electronic device can recognize the acoustic features of each character in the wake-up word, enabling the electronic device to recognize "你好,悠悠".

[0278] Assume that after the electronic device performs frame division on the audio signal, the duration of one frame of audio signal is t. Then when 30 / t is a decimal, the electronic device can perform feature extraction on 1 frame of audio signal every frames of audio signal.

[0279] Exemplarily, when the duration of one frame of audio signal is 10 ms after the electronic device performs frame division, the electronic device can perform feature extraction on 1 frame of audio signal every 2 frames of audio signal.

[0280] Or, assume that the electronic device performs feature extraction on Y frames of audio signal every X frames of audio signal. X can also be less than or 30 / t - 1, and / or, Y can be 1 or greater than 1. As long as the time interval for feature extraction between two adjacent frames is less than 30 ms. Exemplarily, the electronic device can also perform feature extraction every The electronic device can extract features from multiple consecutive audio frames, either in frames of audio signal or every 30 / t-1 frames. For example, if the duration of one audio frame is 10ms, the electronic device can extract features from two consecutive audio frames every two frames. Alternatively, the electronic device can also extract features from every... The frame audio signal performs feature extraction on a single frame of audio signal, etc. For the sake of brevity, these steps will not be shown in detail here.

[0281] Similarly, the frame skipping strategy corresponding to each application scenario in at least one application scenario included in the frame skipping strategy configuration library can also be determined in the above manner.

[0282] For example, in navigation applications, the user's voice is related to navigation. To enable the DSP to recognize each character in the keywords, in the frame skipping strategy 'a' corresponding to the navigation application scenario, the time interval between feature extraction between two adjacent frames needs to be less than the pronunciation duration 1 of any character in the high-frequency speech words used in the navigation application scenario. Specifically, when a character includes both an initial consonant and a final vowel, the pronunciation duration 1 is the shorter of the initial consonant's pronunciation duration and the final vowel's pronunciation duration; when a character includes an initial consonant but not a final vowel, the pronunciation duration 1 is the initial consonant's pronunciation duration; and when a character includes a final vowel but not an initial consonant, the pronunciation duration 1 is the final vowel's pronunciation duration.

[0283] It should be understood that high-frequency speech words in an application scenario can be pre-collected speech words that users frequently utter in that application scenario. High frequency does not limit the frequency of speech words to exceeding a certain value. High-frequency speech words can be words set based on experience. For example, in a navigation application scenario, high-frequency speech words are predefined navigation-related speech words, such as "navigate home," "navigate to the company," "navigate to the park," etc. High-frequency speech words can also be pre-collected speech words. For example, for a certain number of users, within a certain time period, speech words uttered by these users more than a threshold of 1 in an application scenario can be collected. This application embodiment does not specifically limit the method for determining high-frequency speech words.

[0284] For example, in navigation applications, high-frequency speech terms could be: "navigate home," "navigate to the company," "navigate to the park," etc. For the initial consonant and final vowel pronunciation durations of each word in the audio signals of "navigate home," "navigate to the company," and "navigate to the park," let's assume the shortest duration is T1. Then, in navigation applications, the time interval for feature extraction between two adjacent frames needs to be less than T1. With a frame duration of t, the frame skipping strategy a can be determined as follows: when T1 / t is a decimal, the electronic device can skip frames every... Frame audio signal, feature extraction is performed on 1 frame audio signal. It is the integer obtained by rounding down the quotient of T1 divided by t. When T1 / t is a positive integer, the electronic device can extract features from one frame of audio signal every T1 / t-1 frames of audio signal.

[0285] Alternatively, in navigation applications, suppose the electronic device extracts features from Y frames of audio signal every X frames. X can also be less than... Or T1 / t-1, and / or, Y can be 1 or greater than 1. This is acceptable as long as the time interval between feature extractions of two adjacent frames is less than T1 ms. For example, the electronic device can also perform feature extraction every... Feature extraction is performed on multiple consecutive frames of audio signals, either frame-by-frame or every T1 / t-1 frames; alternatively, the electronic device may also perform feature extraction on every... Feature extraction is performed on a single frame of audio signal. For the sake of brevity, these features will not be shown individually here.

[0286] Similarly, frame skipping strategies for other application scenarios, such as frame skipping strategy b and frame skipping strategy c, can also be determined based on the above method and then preset in the electronic device. For the sake of simplicity, they will not be shown one by one here.

[0287] It should be noted that the above are merely examples. In various application scenarios, the high-frequency speech words in the audio signal can also include wake-up words. For instance, in a navigation application scenario, the high-frequency speech words could be: "Hello Youyou, navigate home," "Hello Youyou, navigate to the company," "Hello Youyou, navigate to the park," etc. Therefore, the frame skipping strategy corresponding to each application scenario can also be determined based on other high-frequency speech words. For example, frame skipping strategy a corresponding to the navigation application scenario can be determined based on the initial consonant pronunciation duration and / or final vowel pronunciation duration of each character in the audio signals of "Hello Youyou, navigate home," "Hello Youyou, navigate to the company," "Hello Youyou, navigate to the park," etc. In this case, the method for determining the frame skipping strategy corresponding to each application scenario can also be referred to the description above, and will not be repeated here.

[0288] When the electronic device is in a preset state, it can perform audio processing based on a preset frame skipping strategy in the following manner.

[0289] Figure 9 This is a flowchart illustrating another audio processing method 900 provided in an embodiment of this application. Figure 9 As shown, method 900 includes the following steps:

[0290] S901, The navigation application transmits the first instruction to the first SDK.

[0291] The first instruction can be used, for example, to instruct the first SDK navigation application to run in the foreground, to instruct the first SDK navigation application to start, or to instruct the first SDK to issue a frame skipping strategy corresponding to the navigation application scenario.

[0292] Based on the first instruction, the first SDK can determine that it is currently in a navigation application scenario, or that the navigation application is running in the foreground, etc.

[0293] S902, The first SDK determines that the electronic device is in a preset state.

[0294] The preset state can be in-vehicle state, running state, or meeting state, etc. The method by which the first SDK determines that the electronic device is in the preset state can be found in the description above, and will not be repeated here.

[0295] It should be understood that S902 can be executed before S901, or after S902, or S901 and S902 can be executed in parallel. This application does not make any specific limitations on this.

[0296] S903. When the first SDK determines that the electronic device is in a preset state and is currently in a navigation application scenario, it instructs the decision module to issue the frame skipping strategy corresponding to the navigation application scenario.

[0297] S904. Based on the instructions of the first SDK, the decision module can determine the frame skipping strategy a corresponding to the navigation application scenario based on the frame skipping strategy configuration library, and instruct the DSP on the frame skipping strategy a.

[0298] Here, frame skipping strategy a can include q1 frames, meaning that feature extraction is performed on one frame of audio signal every q1 frames of audio signal. q1 is a positive integer. Alternatively, frame skipping strategy a can include q1 frames and p1 frames, meaning that feature extraction is performed on p1 frames of audio signal every q1 frames of audio signal, etc.

[0299] S905, Based on the frame skipping strategy a indicated by the decision module, the DSP can determine that keyword recognition needs to be performed based on frame skipping strategy a. The DSP acquires the audio signal 3 from the microphone.

[0300] S906 and DSP extract speech features 3 from audio signal 3 based on frame skipping strategy a.

[0301] S907 and DSP perform keyword recognition based on speech feature 3.

[0302] Keyword recognition can be achieved, for example, by determining whether the audio signal 3 contains at least one keyword corresponding to the navigation application scenario based on speech features 3. For instance, the DSP determines the word sequence corresponding to the audio signal 3 based on speech features 3 and determines whether the word sequence corresponds to one of the at least one keyword. This at least one keyword can be preset in the electronic device. This at least one keyword may include, but is not limited to, one or more of the following: "Navigate home," "Navigate to the company," "Navigate to the park," "End navigation," or "Exit navigation," etc. The DSP can determine whether the audio signal 3 corresponds to one of the at least one keyword; if so, the DSP can instruct the AP to perform the operation corresponding to that keyword. For example, if the keyword is "Navigate home," the DSP can instruct the navigation application to navigate to the geographical location corresponding to "home."

[0303] It should be understood that the implementation of S906 and S907 is similar to that of S802 and S803, and can be referred to the description above, which will not be repeated here.

[0304] It should be noted that, in the embodiments of this application, the DSP's determination that an audio signal corresponds to a keyword can be as follows: the DSP determines that the word sequence determined based on the audio signal includes all the characters in the keyword; or, the DSP determines that the word sequence determined based on the audio signal has the same meaning as the keyword; or, the DSP determines that the word sequence determined based on the audio signal has a similarity greater than a threshold of 2 with the keyword, etc. The embodiments of this application do not limit the specific method by which the DSP determines that an audio signal corresponds to a keyword.

[0305] The DSP can convert audio signal 3 into text that can be recognized by electronic devices by performing keyword recognition on speech feature 3. This text can be understood as instruction 1 that can be recognized by electronic devices.

[0306] S908, DSP Instruction Navigation Application Processing Instruction 1.

[0307] It is understood that the execution of S908 by an electronic device may involve the interaction of multiple modules. The embodiments of this application do not limit the implementation of S908.

[0308] Furthermore, the navigation application can execute corresponding processes based on instruction 1. For example, if instruction 1 is "Navigate home", the navigation application can generate new voice prompts, continuously update the current location, and update voice prompts in real time.

[0309] It should be understood that Method 900 is an audio processing method illustrated using a navigation application scenario as an example. In other application scenarios, the implementation of the audio processing method is similar to Method 900, the difference being that the application sending the instruction to the first SDK is different, and the frame skipping strategy indicated by the decision module to the DSP may be different. For example, in a music application scenario, the music application can indicate that the first SDK is in a music application scenario; the frame skipping strategy indicated by the decision module to the DSP can be frame skipping strategy b. In an office application scenario, an office application can indicate that the first SDK is in an office application scenario; the frame skipping strategy indicated by the decision module to the DSP can be frame skipping strategy c, etc. For the sake of simplicity, they will not be shown one by one here.

[0310] It is understandable that users may continue to issue voice commands after S908; therefore, electronic devices can continue to perform voice recognition in the manner described in S809 to S815. Please refer to the description above; it will not be repeated here.

[0311] It should be noted that, in this embodiment of the application, after the electronic device extracts features from the audio signal to obtain speech features, it can then decode the accumulated speech features of multiple frames of audio signals after accumulating them, that is, determine the words corresponding to the multiple frames of audio signals based on their speech features. For example, as shown... Figure 10 As shown, the electronic device can decode the speech features of the accumulated 10 frames of audio signals after accumulating the speech features of each 10 frames, for example, to identify the words corresponding to those 10 frames of audio signals.

[0312] Because the amount of speech feature information in a single frame of audio signal is limited, accumulating speech features from multiple frames helps capture the complete information of keywords. Furthermore, the limited speech features of a single frame may not be sufficient to accurately identify keywords. By accumulating speech features from multiple frames, more contextual information can be provided, thereby improving the accuracy of keyword recognition.

[0313] It should be understood that electronic devices can also be decoded in other ways. This application does not specifically limit this method.

[0314] As can be seen from Method 900, electronic devices can achieve voice interaction functions other than voice wake-up through the following process.

[0315] Figure 11 This is a schematic diagram illustrating an audio processing procedure provided in an embodiment of this application. Figure 11 As shown, firstly, before the electronic device leaves the factory, a frame skipping strategy is preset in the electronic device. For example, a frame skipping strategy configuration library is set in the electronic device. The frame skipping strategy configuration library includes at least one application scenario and a frame skipping strategy corresponding to each application scenario in the at least one application scenario.

[0316] It should be understood that the frame skipping strategy corresponding to each application scenario can be determined by the R&D personnel and pre-configured in the electronic device, or it can be determined by other electronic devices and transmitted to the electronic device before it leaves the factory. This application does not specifically limit the method of pre-configuring the frame skipping strategy.

[0317] Subsequently, the electronic device can issue a frame skipping strategy corresponding to the current application scenario. For example, S901 to S904 in method 900. The electronic device can then perform keyword recognition based on this frame skipping strategy. For example, S905 to S908 in method 900.

[0318] Afterwards, the electronic device can continue to perform voice recognition. For example, after S908, the electronic device can continue to perform voice recognition through S809 to S815.

[0319] As can be seen from methods 800 and 900, the audio processing method provided in this application embodiment can extract features from some frame audio signals during keyword recognition in the implementation of various voice interaction functions such as voice wake-up. This reduces the power consumption of the electronic device during keyword recognition, thereby reducing the power consumption of the electronic device in implementing various voice interaction functions such as voice wake-up.

[0320] It should be understood that in the above embodiments, the relative size of each step number does not indicate the order in which the steps are executed, but is determined by the internal logic between each step.

[0321] It should be noted that the module names involved in the embodiments of this application can all be defined as other names, as long as they can achieve the function of each module, and no specific restrictions are placed on the module names.

[0322] It should also be noted that methods 800 and 900 can be implemented individually or in combination. For example, an electronic device can execute both method 800 and method 900; or an electronic device can execute either method 800 or method 900. This application does not impose specific limitations in this regard.

[0323] Furthermore, the user information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data used for analysis, stored data, displayed data, etc.) involved in the embodiments of this application are all information and data authorized by the user or fully authorized by all parties. The collection, use and processing of related data must comply with the relevant laws, regulations and standards of relevant countries and regions, and corresponding operation entry points are provided for users to choose to authorize or refuse.

[0324] The image processing method of the present application embodiments has been described above. The apparatus for performing the above method provided in the present application embodiments is described below. Those skilled in the art will understand that the methods and apparatus can be combined with and referenced by each other, and the related apparatus provided in the present application embodiments can perform the steps in the above list sorting method.

[0325] Figure 12 This is a schematic block diagram of an audio processing device 1200 provided in an embodiment of this application. The device 1200 includes a processor 1201, a communication interface 1202, and a memory 1203. The processor 1201, communication interface 1202, and memory 1203 communicate with each other via internal connection paths. The memory 1203 stores instructions, and the processor 1201 executes the instructions stored in the memory 1203. The communication interface 1202 can be used to send signals to other devices (e.g., the processor 1201 or a touchscreen of an electronic device) and to receive signals from other devices (e.g., the memory 1203). Exemplarily, the communication interface 1202 reads instructions stored in the memory 1203 and sends the instructions to the processor 1201.

[0326] It should be understood that the device 1200 may specifically be an electronic device as described in the above embodiments, and may be used to execute the various steps and / or processes corresponding to the electronic device in the above method embodiments. Optionally, the memory 1203 may include read-only memory and random access memory, and provide instructions and data to the processor. A portion of the memory may also include non-volatile random access memory. For example, the memory may also store device type information. The processor 1201 may be used to execute instructions stored in the memory, and when the processor 1201 executes instructions stored in the memory, the processor 1201 is used to execute the various steps and / or processes of the above method embodiments.

[0327] It should be understood that, in the embodiments of this application, the processor may be a central processing unit (CPU), or it may be other general-purpose processors, digital signal processors (DSPs), application-specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. The general-purpose processor may be a microprocessor or any conventional processor.

[0328] In implementation, each step of the above method can be completed by integrated logic circuits in the processor's hardware or by instructions in software. The steps of the method disclosed in the embodiments of this application can be directly manifested as execution by a hardware processor, or as a combination of hardware and software modules within the processor. The software modules can reside in random access memory, flash memory, read-only memory, programmable read-only memory, electrically erasable programmable memory, registers, or other mature storage media in the art. This storage medium is located in memory, and the processor executes the instructions in the memory, combining them with its hardware to complete the steps of the above method. To avoid repetition, detailed descriptions are omitted here.

[0329] The image processing method provided in this application can be applied to electronic devices with communication functions. The electronic devices include terminal devices, and the specific device form of the terminal devices can be referred to the above-described related descriptions, which will not be repeated here.

[0330] This application provides a terminal device, which includes a processor and a memory; the memory stores computer execution instructions; the processor executes the computer execution instructions stored in the memory, causing the terminal device to perform the above-described method.

[0331] This application provides a chip. The chip includes a processor, which is used to call a computer program in memory to execute the technical solutions in the above embodiments. Its implementation principle and technical effects are similar to those in the related embodiments described above, and will not be repeated here.

[0332] This application also provides a computer-readable storage medium. The computer-readable storage medium stores a computer program. When the computer program is executed by a processor, it implements the methods described above. The methods described in the above embodiments can be implemented wholly or partially by software, hardware, firmware, or any combination thereof. If implemented in software, the functionality can be stored as one or more instructions or code on or transmitted over the computer-readable medium. The computer-readable medium can include computer storage media and communication media, and can also include any medium that can transfer a computer program from one place to another. The storage medium can be any target medium accessible by a computer.

[0333] In one possible implementation, a computer-readable medium may include RAM, ROM, compact disc read-only memory (CD-ROM) or other optical disc storage, disk storage or other magnetic storage devices, or any other medium targeted to carry or to store the required program code in the form of instructions or data structures, and accessible by a computer. Furthermore, any connection is appropriately referred to as a computer-readable medium. For example, if software is transmitted from a website, server, or other remote source using coaxial cable, fiber optic cable, twisted pair, DSL, or wireless technologies such as infrared, radio, and microwave, then coaxial cable, fiber optic cable, twisted pair, DSL, or wireless technologies such as infrared, radio, and microwave are included in the definition of medium. As used herein, disks and optical discs include optical discs, laser discs, optical discs, Digital Versatile Discs (DVDs), floppy disks, and Blu-ray discs, where disks typically reproduce data magnetically, while optical discs optically reproduce data using lasers. Combinations of the above should also be included within the scope of computer-readable media.

[0334] This application provides a computer program product, which includes a computer program that, when run, causes a computer to perform the above-described method.

[0335] This application describes embodiments of methods, apparatus (systems), and computer program products according to embodiments of this application with reference to flowchart illustrations and / or block diagrams. It should be understood that each block of the flowchart illustrations and / or block diagrams, and combinations of blocks in the flowchart illustrations and / or block diagrams, can be implemented by computer program instructions. These computer program instructions can be provided to a processing unit of a general-purpose computer, special-purpose computer, embedded processor, or other programmable device to produce a machine, such that the instructions, which execute via the processing unit of the computer or other programmable data processing device, generate instructions for implementing the flowchart illustrations. Figure 1 One or more processes and / or boxes Figure 1 A device that provides the functions specified in one or more boxes.

[0336] The above specific embodiments further illustrate the purpose, technical solution, and beneficial effects of the present invention. It should be understood that the above are merely specific embodiments of the present invention and are not intended to limit the scope of protection of the present invention. Any modifications, equivalent substitutions, improvements, etc., made on the basis of the technical solution of the present invention should be included within the scope of protection of the present invention.

Claims

1. An audio processing method, characterized in that, Applied to electronic devices, the method includes: In the first scenario, every N frames of audio signal, feature extraction is performed on the first audio signal collected by the electronic device, and audio recognition is performed based on the extracted first audio features. The N is related to the first keyword collected in advance in the first scenario, and the N is an integer greater than 0. In the second scenario, every M frames of audio signal, features are extracted from the second audio signal collected by the electronic device, and audio recognition is performed based on the extracted second audio features. M is related to the second keyword collected in advance in the second scenario, and M is an integer greater than 0.

2. The method according to claim 1, characterized in that, The first keyword is the wake word of the electronic device; The audio recognition based on the extracted first audio features includes: Based on the first audio feature, determine whether the wake-up word is included in the first audio signal; The method further includes: When the wake-up word is included in the first audio signal, the electronic device displays a first interface, which is the interface of the electronic device after it is woken up by voice.

3. The method according to claim 2, characterized in that, The first keyword is the word with the shortest first pronunciation duration in the wake word, and the first pronunciation duration is the shorter of the initial consonant pronunciation duration and the final vowel pronunciation duration of a word.

4. The method according to claim 3, characterized in that, The N is determined based on a first ratio, which is the ratio of the duration of the initial consonant pronunciation of the first keyword to the duration of a frame of speech signal.

5. The method according to any one of claims 1 to 4, characterized in that, In the second scenario, every M frames of audio signal, feature extraction is performed on the second audio signal collected by the electronic device, including: When the electronic device is in a preset mode and a first application is running in the foreground of the electronic device, a first frame skipping strategy corresponding to the first application is obtained, and the first frame skipping strategy includes M; Acquire the second audio signal; Based on the first frame skipping strategy, feature extraction is performed on the second audio signal to obtain the second audio features.

6. The method according to claim 5, characterized in that, The second keyword is the word with the shortest first pronunciation duration among at least one keyword. The first pronunciation duration is the shorter of the initial consonant pronunciation duration and the final vowel pronunciation duration of a word. The at least one keyword is a high-frequency speech word corresponding to the first application that has been collected in advance.

7. The method according to claim 6, characterized in that, The audio recognition based on the extracted second audio features includes: Based on the second audio feature, it is determined whether the second audio signal includes the target keyword corresponding to the first application, wherein the target keyword is part or all of the keywords in the at least one keyword.

8. The method according to any one of claims 5 to 7, characterized in that, The preset mode is any one of the following: vehicle mode, sports mode, or conference mode.

9. The method according to any one of claims 1 to 8, characterized in that, The method further includes: In the third scenario, every S frames of audio signal, features are extracted from the third audio signal collected by the electronic device, and audio recognition is performed based on the extracted third audio features. The S is related to the third keyword collected in advance in the third scenario, and the S is an integer greater than 0.

10. The method according to claim 9, characterized in that, The third keyword is the word with the shortest first pronunciation duration among the high-frequency speech words corresponding to the second application. The first pronunciation duration is the shorter of the initial consonant pronunciation duration and the final vowel pronunciation duration of a word. In the third scenario, every S frames of audio signal, feature extraction is performed on the third audio signal collected by the electronic device, including: When the electronic device is in a preset mode and the second application is running in the foreground of the electronic device, a second frame skipping strategy corresponding to the second application is obtained, and the second frame skipping strategy includes the S; Acquire the third audio signal; Based on the second frame skipping strategy, feature extraction is performed on the third audio signal to obtain the third audio features.

11. The method according to any one of claims 1 to 10, characterized in that, The electronic device includes a digital signal processor (DSP) and a microphone; In the first scenario, every N frames of audio signal, feature extraction is performed on the first audio signal collected by the electronic device, and audio recognition is performed based on the extracted first audio features, including: The microphone collects the first audio signal; The DSP acquires the first audio signal from the microphone; The DSP extracts features from one frame of audio signal every N frames of audio signal to obtain the first audio feature; The DSP determines whether the first audio signal contains the wake word of the electronic device based on the first audio feature.

12. The method according to any one of claims 1 to 11, characterized in that, The value of N is 2.

13. An electronic device, characterized in that, The electronic device includes: one or more processors, memory, and a microphone; The memory is coupled to the one or more processors, the memory being used to store computer program code, the computer program code including computer instructions, the one or more processors calling the computer instructions to cause the electronic device to perform the method as described in any one of claims 1 to 12, and the microphone being used to perform the steps related to acquiring audio signals as described in any one of claims 1 to 12.

14. A chip system, characterized in that, The chip system is applied to an electronic device, the chip system including one or more processors, the one or more processors being used to invoke computer instructions to cause the electronic device to perform the method as described in any one of claims 1 to 12.

15. A computer-readable storage medium, characterized in that, The computer-readable storage medium includes computer instructions that, when executed on an electronic device, cause the electronic device to perform the method as described in any one of claims 1 to 12.

16. A computer program product, characterized in that, The computer program product includes computer program code that, when run on an electronic device, causes the electronic device to perform the method as described in any one of claims 1 to 12.

Citation Information

Patent Citations

  • Voice wakeup method and voice wakeup device

    CN105741838A

  • Voice data processing method and device, electronic equipment and medium

    CN111768764A

  • Voice wake-up method and device and storage medium

    CN116129878A

  • Behavior recognition method and apparatus, terminal device, and readable storage medium

    WO2022134983A1