Deep learning keyword recognition method based on gas-bone conduction dual-mode
Through the deep learning method based on the dual-mode of air-bone conduction, the Audiomer-L neural network model is constructed, which solves the problem of low signal-to-noise ratio of air-to-conductance signals in harsh acoustic environments, and realizes the simultaneous processing of air-to-conductance signals and bone-to-conductance signals, improving the accuracy and robustness of keyword recognition.
Patent Information
- Application Number
- CN202510222640.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-26
- Publication Date
- 2025-05-30
AI Technical Summary
The current technology has low signal-to-noise ratio of gas conduction signals in harsh acoustic environments, resulting in a decrease in the accuracy of keyword recognition structure or even failure. It has only studied gas conduction signals and failed to effectively process bone conduction signals.
The deep learning keyword recognition method based on the dual-mode of the air-bone conduction is adopted. By constructing the Audiomer-L neural network model, combining the learning vector module, convolutional attention module, multi-layer perceptron module, etc., the simultaneous processing of the air-bone conduction signal and the bone conduction signal is achieved, and the robustness of the recognition is improved.
It significantly improves the accuracy of keyword recognition in noisy environments, enhances detection performance, and has the advantages of low parameter quantity and low operation complexity, proving the effectiveness of gas conduction signals and bone conduction signals in keyword recognition.
Smart Images

Figure CN120071907A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of speech recognition, and particularly relates to a deep learning keyword recognition method based on a dual-mode of air conduction and bone conduction. Background Art
[0002] With the development of device computing power, the accumulation of speech keyword recognition technology, and the abundance of speech data, applications related to speech keyword recognition have begun to be truly popularized in our daily lives. Keyword recognition technology is a speech recognition technology that focuses on recognizing specific keywords or phrases in speech signals. Different from a comprehensive speech recognition system, the goal of keyword recognition is to monitor and detect specific words or phrases in a speech stream and trigger corresponding actions when these keywords are detected.
[0003] It is particularly necessary to study keyword recognition technology that can maintain high performance in a low signal-to-noise ratio environment. This technology can enable people to accurately receive and execute key information in a noisy environment and reduce communication latency. The air-conduction microphone collects vibration signals transmitted through the air, while the bone-conduction microphone collects vibration signals passing through the jawbone, human tissues, etc. In scenarios with a relatively harsh acoustic environment, the signal-to-noise ratio of the signals received by the microphone is low, while the bone-conduction signal is less affected by environmental noise.
[0004] The patent with the application number 202110101761.6 discloses "A Lightweight Neural Network Speech Keyword Recognition Method Based on Hierarchical Quantization", which makes full use of the advantages of a large reduction in the number of parameters and computational complexity brought by depthwise separable convolution and the attention mechanism to label the importance of features on different channels during the convolution process, thereby improving the accuracy and speed of model recognition. The patent with the application number 202310612970.6 discloses "A Speech Keyword Recognition Method, Device and Related Equipment", which improves the performance of a small-sample model by extracting the Mel Frequency Cepstral Coefficient (MFCC) features of the speech to be recognized and inputting the MFCC features into a keyword model to obtain the recognition result of the speech to be recognized.
[0005] The above research methods solve the problem of improving the keyword recognition performance of a terminal under a low signal-to-noise ratio. However, they only study the keyword recognition of air-conduction signals and only have a mechanism for targeted processing of air-conduction signals. In a scenario with a relatively harsh acoustic environment, the signal-to-noise ratio of the signals received by the air-conduction microphone is low, which will lead to a decrease in the accuracy of the recognition result or even failure. Summary of the Invention
[0006] The purpose of the present invention is to provide a network structure with low computational complexity to achieve keyword recognition for both air-conduction signals and bone-conduction signals simultaneously, realize robust recognition in a noisy environment, and verify the effectiveness of air-conduction signals and bone-conduction signals in keyword recognition.
[0007] To achieve the object of the present invention, the present invention provides a deep learning keyword recognition method based on air-bone conduction dual-mode, including the following steps:
[0008] Step 1: Synchronously record pure air-conducted speech s a and pure bone-conducted speech s b in a noise-free environment. Add environmental noise d to the air-conducted speech s a to obtain noisy air-conducted speech s = s a + d. Construct a data set [(s, s b ), s a , and then divide the data set into a training set, a validation set, and a test set;
[0009] Step 2: Cut the speech data of the training set into multiple small segments of speech according to a fixed length;
[0010] Step 3: Connect one learning vector module, n convolutional attention modules, n one-dimensional convolutional network layers, and one multi-layer perceptron MLP module in a residual connection manner to construct an Audiomer-L neural network model;
[0011] Step 31: Input the original audio waveforms of the noisy air-conducted speech and the pure bone-conducted speech into the learning vector module, and output aggregated information feature Z oi :
[0012] The learning vector module connects a learnable vector of 128 frames to the beginning of the original audio waveforms of the noisy air-conducted speech and the pure bone-conducted speech, so as to ensure that the first layer of the encoding network can aggregate classification-specific information at the beginning of the input sequence; at the same time, the learning vector module performs mapping processing on the spliced learnable vector and the original audio waveform to form an aggregated information feature Z carrying high-dimensional classification feature information oi
[0013] Step 32: Input the aggregated information feature Z oi into a processing unit composed of n convolutional attention modules and n one-dimensional convolutional network layers connected in a residual manner to obtain a feature map Z o ;
[0014] Step 33: Input the first frame z o of the feature map Z o into a multi-layer perception module composed of two linear fully connected layers, and output the final keyword recognition classification result to complete the construction of the Audiomer-L neural network model;
[0015] The convolutional attention module includes two one-dimensional convolutional modules with a squeeze-and-excitation mechanism and a Performer attention module; among them, the one-dimensional convolutional module with a squeeze-and-excitation mechanism is called a convolutional module with a squeeze-and-excitation mechanism, and the Performer attention module is used to learn the tensor features counted by the convolutional module with a squeeze-and-excitation mechanism; given the output of the previous module as the input of the current convolutional attention module, the input of the k-th convolutional attention module is denoted as Z ik ; In the two one-dimensional convolutional modules, the strides of the convolutional operations are set to S / 2 and S respectively, and according to the input feature Z of the module ik a pre-query tensor and a context tensor are generated respectively. The context tensor is used as the key K and value V input to the Performer attention module, and the pre-query tensor is used for query Q; the final output Z of the convolutional attention module is obtained through a residual connection between the query Q and the output of the Performer attention module ok ;
[0016] Step 4: Use the data of the training set after cutting in Step 2 to train the Audiomer-L neural network model to obtain the trained Audiomer-L neural network model;
[0017] Step 5: Feed the data of the test set into the trained Audiomer-L neural network model to obtain the final Audiomer-L neural network model, and output the keyword recognition result.
[0018] Compared with the prior art, the significant progress of the present invention lies in: (1) The present invention adopts the low-weight deep neural network Audiomer-L to conduct keyword recognition research on air-conducted signals and bone-conducted signals. Through simulation experiments, it shows that using Audiomer-L to conduct keyword recognition on the input air-conducted signals and bone-conducted signals significantly improves the accuracy, has higher detection performance, and has the advantages of low parameter quantity and low computational complexity; (2) The present invention ensures that the spatial resolution of K and V is half of the spatial resolution of Q by setting different strides in the one-dimensional convolutional module, which can save computing resources while also bringing performance improvement; (3) The residual connection can improve the gradient flow and can significantly improve the performance of the model with only a small increase in parameters.
[0019] To more clearly illustrate the functional characteristics and structural parameters of the present invention, the following further explains in conjunction with the drawings and specific embodiments. Description of the Drawings
[0020] The accompanying drawings described herein are used to provide a further understanding of the present invention and form a part of this application. The illustrative embodiments of the present invention and their descriptions are used to explain the present invention and do not constitute an improper limitation of the present invention. In the drawings:
[0021] Figure 1 is a flowchart of the steps of the present invention;
[0022] Figure 2 is a schematic diagram of the framework structure of the Audiomer-L neural network model of the present invention;
[0023] Figure 3 is a schematic diagram of the structure of the convolutional attention module of the present invention. Detailed implementation manners
[0024] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments; based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of the present invention.
[0025] A deep learning keyword recognition method based on air-bone conduction dual-mode of the present invention, in combination with Figure 1 , includes the following steps:
[0026] Step 1. Synchronously record pure air-conducted speech s a and pure bone-conducted speech s b in a noise-free environment, add environmental noise d to the air-conducted speech s a to obtain noisy air-conducted speech s = s a + d, construct a data set [(s, s b ), s a , and then divide the data set into a training set, a validation set, and a test set;
[0027] Step 1-1. In this example, a total of 500 air-conducted and bone-conducted common command word voices from 15 speakers were collected. The total duration of these data is 4.15 hours, and each speaker contributed approximately 17 minutes of speech. The noise data set selected the noise data set in the DNS competition as the air-conducted signal noise;
[0028] Step 1-2: To enhance the robustness of the model during the recognition process, data augmentation was performed on the data. For air-conducted signals, various augmentation methods were adopted, including adding noise, adjusting the speech rate and pitch. In this embodiment, the signal-to-noise ratio (SNR) of the added noise was set to [5-10] dB, the adjustment range of the speech rate was [0.75, 0.85, 0.95, 1.05, 1.15, 1.25, 1.35, 1.45, 1.55], indicating that the speech rate changed from 75% to 155% of the original rate, and the adjustment range of the pitch was [-3, -2, -1, 0, 1, 2, 3]. For bone-conducted signals, the adjustment of the speech rate was completed through the publicly available librosa library;
[0029] Step 1-3: The method of dividing the dataset into a training set, a validation set, and a test set was such that 80% of the dataset was the training set, 10% was the validation set, and 10% was the test set.
[0030] Step 2: The speech data in the training set was cut into multiple small segments of speech with a fixed length. In this embodiment, the fixed length was 2 seconds;
[0031] Step 3: Combine Figure 2 , connect one learning vector module, n convolutional attention modules, n one-dimensional convolutional network layers, and one multi-layer perceptron (MLP) module in a residual link manner to construct the Audiomer-L neural network model. The Audiomer-L neural network model is a high-efficiency neural network designed specifically for keyword recognition tasks;
[0032] Step 31: Input the original audio waveforms of the noisy air-conducted speech and the clean bone-conducted speech into the learning vector module, and output the aggregated information feature Z oi ; The specific processing process is that the learning vector module connects a learnable vector c with a length of 128 frames to the beginning of the original audio waveforms s of the noisy air-conducted speech and the clean bone-conducted speech to ensure that the first layer of the encoding network can aggregate classification-specific information at the beginning of the input sequence; at the same time, the learning vector module forms the aggregated information feature Z carrying high-dimensional classification feature information through mapping processing of the concatenated learnable vector and the original audio waveform oi ;
[0033] Step 32: Input the aggregated information feature into a processing unit composed of n convolutional attention modules and n one-dimensional convolutional network layers connected in a residual manner to obtain the feature map Z o ;
[0034] Step 33: The first frame z of the feature map Z o oIt is input into a multi-layer perceptron (MLP) module composed of two linear fully-connected layers, and the final keyword recognition classification result is output to complete the construction of the Audiomer-L neural network model.
[0035] Combined with Figure 3 , the convolutional attention module includes two one-dimensional convolutional modules with a squeeze-and-excitation mechanism and a Performer attention module; among them, the one-dimensional convolutional module with a squeeze-and-excitation mechanism is called a convolutional module with a squeeze-and-excitation mechanism, and the Performer attention module is used to learn the tensor features counted by the convolutional module with a squeeze-and-excitation mechanism; given the output of the previous module as the input of the current convolutional attention module, the input of the k-th convolutional attention module is denoted as Z ik ; in the two one-dimensional convolutional modules, the strides of the convolutional operations are set to S / 2 and S respectively, and according to the input feature Z of the module ik a pre-query tensor and a context tensor are generated respectively. The context tensor is used as the key K and value V input to the Performer attention module, and the pre-query tensor is used for query Q; the final output Z of the convolutional attention module is obtained through a residual connection between the query Q and the output of the Performer attention module ok .
[0036] In this embodiment, the encoding part of the Audiomer-L neural network model is composed of 11 convolutional attention modules and 11 one-dimensional convolutional neural network layers in a residual connection manner.
[0037] Step 4: Use the data of the training set segmented in Step 2 to train the Audiomer-L neural network model to obtain a trained Audiomer-L neural network model;
[0038] Step 5: Send the data of the test set into the final Audiomer-L neural network model obtained from the trained Audiomer-L neural network model, and output the keyword recognition result.
[0039] Compared with the classical method keyword Transformer, the method of this embodiment can have higher detection performance and has the advantages of low parameter quantity and low operation complexity.
[0040] Table 1 shows the parameter quantity and operation volume of this network. It can be seen that the Audiomer-L neural network model has fewer network parameters compared with most other models. In addition, the Audiomer-L neural network model is only 0.8M and the operation complexity is 0.088 GFlops.
[0041] Table 1 Network parameter quantity and operation complexity
[0042]
[0043] Table 2 Recognition accuracy of the Audiomer-L neural network model for air-conducted signals and bone-conducted signals respectively under the scenario of DNS random noise
[0044]
[0045] Table 2 shows the keyword recognition accuracy of the Audiomer-L neural network for air-conducted signals and bone-conducted signals with different signal-to-noise ratios as inputs; the Audiomer-L neural network has a relatively high recognition accuracy for all signal types; and the recognition accuracy of bone-conducted signals is higher than that of air-conducted signals in the 5dB - 15dB noise environment;
[0046] In summary, although there is energy loss in the high-frequency band of bone-conducted signals, they still contain speech information that can be used for keyword recognition.
[0047] It should be noted that in this article, relational terms such as first and second are only used to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any actual relationship or order between these entities or operations. Moreover, the term "comprising", "including" or any other variant thereof is intended to cover non-exclusive inclusion, so that a process, method, article or device comprising a series of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such process, method, article or device.
[0048] Although the embodiments of the present invention have been shown and described, it will be understood by those of ordinary skill in the art that various changes, modifications, substitutions and variations can be made in these embodiments without departing from the principles and spirit of the present invention, and the scope of the present invention is defined by the appended claims and their equivalents.
Claims
1. A deep learning keyword recognition method based on air-bone dual-mode, characterized in that: The following steps are involved: Step 1: Record pure air-conducted speech synchronously in a noise-free environment a and pure bone conduction speech b , the air-conducted speech a Add the environmental noise d to get the noisy air-conducted speech s=s a +d, construct the data set [(s, s b ), s a ], and then divide the data set into training set, validation set, and test set; Step 2: cutting the speech data of the training set into multiple short speech segments according to fixed length; Step 3: Connect one learning vector module, n convolutional attention modules, n one-dimensional convolutional network layers, and one multi-layer perceptron MLP module in a residual link manner to construct an Audiomer-L neural network model; Step 4: train the Audiomer-L neural network model using the data of the training set cut in step 2 to obtain a trained Audiomer-L neural network model; Step 5: Send the data of the test set to the trained Audiomer-L neural network model to obtain the final Audiomer-L neural network model, and output the keyword recognition result.
2. The deep learning keyword recognition method based on air-bone dual-mode according to claim 1 is characterized in that: The step 3 comprises the following steps: Step 31: Input the original audio waveforms of the noisy air-conducted speech and the pure bone-conducted speech into a learning vector module, and output the aggregated information feature Z oi : Step 32: The aggregate information feature Z oi Input is a processing unit consisting of n convolutional attention modules and n one-dimensional convolutional network layers connected in a residual manner, and the feature map Z is obtained. o ; Step 33: transform the feature map Z o The first frame of z o The data is input into the multi-layer perceptron (MLP) module consisting of two linear fully connected layers, and the final keyword recognition and classification results are output to complete the construction of the Audiomer-L neural network model.
3. The deep learning keyword recognition method based on air-bone dual-mode according to claim 2 is characterized in that: The step 31 is specifically as follows: The learning vector module connects the learnable vector with a length of 128 frames to the beginning of the original audio waveform of the noisy air-conducted speech and the clean bone-conducted speech, so as to ensure that the first layer of the encoding network can aggregate the classification-specific information at the beginning of the input sequence; at the same time, the learning vector module processes the spliced learnable vector and the original audio waveform through mapping to form an aggregated information feature Z carrying high-dimensional classification feature information. oi .
4. The deep learning keyword recognition method based on air-bone dual-mode according to claim 3 is characterized in that: The convolutional attention module includes two one-dimensional convolutional modules with compression excitation mechanism and a Performer attention module; wherein the one-dimensional convolutional module with compression excitation mechanism is called a convolutional module with compression excitation mechanism, and the Performer attention module is used to learn the tensor features counted by the convolutional module with compression excitation mechanism; given the output of the previous module as the input of the current convolutional attention module, the input of the kth convolutional attention module is represented as Z ik ; In the two one-dimensional convolution modules, the strides of the convolution operations are set to S / 2 and S respectively, and according to the input feature Z of the module ik Generate a pre-query tensor and a context tensor respectively. The context tensor is used to input the key K and value V of the Performer attention module, and the pre-query tensor is used to query Q. The query Q is connected to the output of the Performer attention module through a residual connection to obtain the final output Z of the convolutional attention module. ok .
5. The deep learning keyword recognition method based on air-bone dual-mode according to claim 1 is characterized in that: The fixed length in step 2 is 2 seconds.
6. The deep learning keyword recognition method based on air-bone dual-mode according to claim 1 is characterized in that: 80% of the data set in step 1 is a training set, 10% is a validation set, and 10% is a test set.
7. The deep learning keyword recognition method based on air-bone dual-mode according to claim 1 is characterized in that: There are 11 convolutional attention modules and 11 one-dimensional convolutional neural network layers.
Citation Information
Patent Citations
A lightweight neural network speech keyword recognition method based on hierarchical quantization
CN112786021B
Voice keyword recognition method and device and related equipment
CN116564279A