Voice wake-up method, electronic device, storage medium and computer program product
By performing feature extraction and depth-wise separable convolution processing on the sound signal on the end device, the computing resource limitations of the voice wake-up system on low-power devices are overcome, and efficient voice wake-up performance is improved in complex scenarios.
Patent Information
- Application Number
- CN202511189911.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-08-25
- Publication Date
- 2025-09-19
- Estimated Expiration
- 2045-08-25
AI Technical Summary
On low-power embedded devices and mobile terminals, the deployment of existing voice wake-up systems is limited by computing resources and storage space. Deep neural networks are difficult to apply under low latency and low power requirements, making the design of an efficient end-side voice wake-up system a challenge.
By extracting features from the sound signal, two-dimensional signal features are obtained. Depthwise separable convolution processing is used, including two-dimensional depth-by-depth convolution and two-dimensional point-by-point convolution, to reduce the number of parameters and calculations and improve voice wake-up performance.
While ensuring detailed information about time-frequency domain features, the algorithm's parameter count and computational complexity are greatly reduced, improving the performance of end-side voice wake-up in complex scenarios and enabling precise voice command triggering.
Smart Images

Figure CN120673760A_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of signal processing technology, and in particular to a voice wake-up method, electronic device, storage medium, and computer program product. Background Art
[0002] In recent years, with the widespread adoption of smart devices (such as smartwatches, AR / VR glasses, and wireless headphones), on-device voice wake-up (Keyword Spotting (KWS)) technology has become a key entry point for human-computer interaction. Traditional voice wake-up systems typically rely on cloud-based processing, but due to network latency, privacy concerns, and power consumption constraints, on-device voice wake-up has gradually become a research hotspot. Current voice wake-up algorithms can be categorized into traditional machine learning algorithms and deep learning algorithms. Traditional machine learning algorithms typically have a core process consisting of feature extraction and keyword matching. While these algorithms have low computational complexity, they are limited in handling complex scenarios. Deep learning-based approaches utilize models such as feedforward neural networks, convolutional neural networks, and recurrent neural networks to enhance voice wake-up, achieving excellent performance. However, these models have high parameter counts and computational complexity, making them difficult to implement on on-device devices.
[0003] In real-world scenarios, especially on low-power embedded devices and mobile terminals, the deployment of voice wake-up systems is often limited by computing resources and storage space. Deep neural networks, while offering excellent performance, struggle to meet low latency and power requirements. Therefore, designing efficient KWS systems within resource-constrained environments remains a pressing challenge.
[0004] The above content is only used to assist in understanding the technical solution of this application and does not constitute an admission that the above content is prior art. Summary of the Invention
[0005] The main purpose of this application is to provide a voice wake-up method, electronic device, storage medium and computer program product, aiming to improve the voice wake-up performance of the voice wake-up (KWS) task in complex scenarios on a terminal device with low parameter count and low computing resources.
[0006] To achieve the above objectives, the present application proposes a voice wake-up method, which includes: Acquiring a sound signal, and performing feature extraction on the sound signal to obtain two-dimensional signal features of at least two channels, wherein the two dimensions of the two-dimensional signal features are a time dimension and a frequency dimension; Performing convolution processing on the two-dimensional signal features of each channel to obtain a feature map of at least one channel, wherein the convolution processing includes at least one depthwise separable convolution processing, and the depthwise separable convolution processing includes two-dimensional depthwise convolution processing and two-dimensional pointwise convolution processing; The voice wake-up word recognition result is obtained based on the feature map recognition of each channel.
[0007] Optionally, the step of performing convolution processing on the two-dimensional signal features of each channel to obtain a feature map of at least one channel includes: The two-dimensional signal features of each channel are convolved using at least two preset depth-wise separable convolution structures connected in sequence to obtain a feature map of at least one channel, wherein the input data of the first depth-wise separable convolution structure is the two-dimensional signal features of each channel, and the input data of each depth-wise separable convolution structure after the first depth-wise separable convolution structure is the output data of the previous depth-wise separable convolution structure.
[0008] Optionally, for each of the depthwise separable convolution structures, the step of performing convolution processing on input data using the depthwise separable convolution structure includes: Inputting the input data into the depthwise separable convolution structure to perform convolution processing to obtain a processing result; The input data is added to the processing result to obtain output data of the depthwise separable convolution structure.
[0009] Optionally, the depthwise separable convolution structure includes a depthwise convolution module, a first normalized activation module, a pointwise convolution module and a second normalized activation module connected in sequence; the depthwise convolution module includes a first type of convolution kernel corresponding to the two-dimensional features of each channel in the input data, and the first type of convolution kernel is used to perform convolution processing on the two-dimensional features of the corresponding channel; the point-by-point convolution module includes at least one second type of convolution kernel, and the size of the second type of convolution kernel is 1×1×M, where M is the number of output channels of the first normalized activation module.
[0010] Optionally, the step of performing convolution processing on the two-dimensional signal features of each channel using at least two preset sequentially connected depthwise separable convolution structures to obtain a feature map of at least one channel includes: Performing convolution processing on the two-dimensional signal features of each channel using at least two preset depthwise separable convolution structures connected in sequence to obtain output data corresponding to each of the depthwise separable convolution structures; The output data are added together to obtain a feature map of at least one channel.
[0011] Optionally, the step of extracting features from the sound signal to obtain two-dimensional signal features of at least two channels includes: Extracting features from the sound signal to obtain FBANK features; The FBANK features are subjected to channel expansion and dimension transformation to obtain two-dimensional signal features of at least two channels.
[0012] Optionally, the step of obtaining a voice wake-up word recognition result based on the feature map recognition of each channel includes: Input the feature map of each channel into a preset classification layer for classification to obtain a probability value; If the probability value is greater than a preset threshold, a recognition result is obtained that the sound signal contains the voice wake-up word; If the probability value is less than or equal to the preset threshold, a recognition result is obtained that the sound signal does not contain the voice wake-up word.
[0013] In addition, to achieve the above-mentioned purpose, the present application also proposes an electronic device, which includes: a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the computer program is configured to implement the steps of the voice wake-up method as described above.
[0014] In addition, to achieve the above-mentioned purpose, the present application also proposes a storage medium, which is a computer-readable storage medium. A computer program is stored on the storage medium, and when the computer program is executed by the processor, the steps of the voice wake-up method described above are implemented.
[0015] In addition, to achieve the above-mentioned purpose, the present application also provides a computer program product, which includes a computer program. When the computer program is executed by a processor, the steps of the voice wake-up method as described above are implemented.
[0016] One or more technical solutions proposed in this application have at least the following technical effects: By extracting features from the sound signal, two-dimensional signal features of at least two channels are obtained, and the two dimensions are the time dimension and the frequency dimension, providing a data basis for fully capturing the feature details in the time-frequency domain; convolution processing is performed on the two-dimensional signal features of each channel, and the convolution processing includes at least one depth-wise separable convolution processing, and the depth-wise separable convolution processing includes two-dimensional depth-by-depth convolution processing and two-dimensional point-by-point convolution processing. The feature map of at least one channel obtained by the convolution processing is recognized to obtain the voice wake-up word recognition result. Compared with ordinary two-dimensional convolution, while ensuring the full capture of the feature details in the time-frequency domain, the algorithm's parameter amount and computational complexity are greatly reduced, and the performance of the end-side voice wake-up in complex scenarios is improved, ultimately achieving accurate voice command triggering when the end-side device is in standby state. BRIEF DESCRIPTION OF THE DRAWINGS
[0017] The accompanying drawings, which are incorporated in and constitute a part of this specification, illustrate embodiments consistent with the present application and, together with the description, serve to explain the principles of the present application.
[0018] In order to more clearly illustrate the embodiments of the present application or the technical solutions in the prior art, the following briefly introduces the drawings required for use in the embodiments or the description of the prior art. Obviously, for ordinary technicians in this field, other drawings can be obtained based on these drawings without any creative work.
[0019] Figure 1 A flowchart of the first embodiment of the voice wake-up method of this application is provided; Figure 2 A schematic diagram of two-dimensional depth-wise convolution provided in one embodiment of the present application; Figure 3 A schematic diagram of two-dimensional point-by-point convolution provided in one embodiment of the present application; Figure 4 A schematic diagram of a neural network structure provided in one embodiment of the present application; Figure 5 This is a schematic diagram of the device structure of the hardware operating environment involved in the voice wake-up method in the embodiment of the present application.
[0020] The purpose, features and advantages of this application will be further explained in conjunction with the embodiments and with reference to the accompanying drawings. DETAILED DESCRIPTION
[0021] It should be understood that the specific embodiments described herein are merely used to explain the technical solutions of the present application and are not intended to limit the present application.
[0022] It should be noted that, in the description of this application specification and the appended claims, the terms "first", "second", etc. are only used to distinguish the description and cannot be understood as indicating or implying relative importance.
[0023] In order to better understand the technical solution of the present application, a detailed description will be given below in conjunction with the accompanying drawings and specific implementation methods.
[0024] In real-world scenarios, especially on low-power embedded devices and mobile terminals, the deployment of voice wake-up systems is often limited by computing resources and storage space. Deep neural networks, while offering excellent performance, struggle to meet low latency and power requirements. Therefore, designing efficient KWS systems within resource-constrained environments remains a pressing challenge.
[0025] The embodiments of the present application provide a solution for improving the performance of voice wake-up (KWS) tasks in complex scenarios on end-side devices with low parameter counts and low computing resources. By extracting features from the sound signal, two-dimensional signal features are obtained for at least two channels, with the two dimensions being time and frequency, providing a data foundation for fully capturing detailed information about time-frequency domain features. Convolution processing is performed on the two-dimensional signal features of each channel, and the convolution processing includes at least one depthwise separable convolution process, which includes two-dimensional depth-wise convolution processing and two-dimensional point-wise convolution processing. The feature map of at least one channel obtained by the convolution processing is recognized to obtain a voice wake-up word recognition result. Compared to ordinary two-dimensional convolution, this greatly reduces the algorithm's parameter count and computational complexity while ensuring that detailed information about time-frequency domain features is fully captured. This improves the performance of end-side voice wake-up in complex scenarios, ultimately achieving accurate voice command triggering when the end-side device is in standby mode.
[0026] It should be noted that the execution subject of this embodiment can be an electronic device with data processing, network communication and program running functions, such as AR (Augmented Reality) / VR (Virtual Reality) glasses, helmets and other head-mounted devices, or tablet computers, personal computers, mobile phones and other devices. Figure 1 , Figure 1 This is a flowchart of the first embodiment of the voice wake-up method of this application. In this embodiment, the voice wake-up method includes steps S10 to S30: Step S10: Acquire a sound signal, perform feature extraction on the sound signal to obtain two-dimensional signal features of at least two channels, wherein the two dimensions of the two-dimensional signal features are a time dimension and a frequency dimension.
[0027] The sound signal can be acquired through a built-in or external microphone. The method for acquiring the sound signal is not limited in this embodiment. In one feasible embodiment, to ensure the timeliness of voice wake-up, the sound signal acquired by the microphone in real time can be acquired and processed in real time. In one feasible embodiment, to improve the accuracy of voice wake-up, voice activity detection (VAD) can be performed on the sound signal acquired by the microphone. Voice wake-up recognition is performed only when speech is detected in the sound signal. In other words, the acquired sound signal can be a sound signal containing speech.
[0028] The purpose of feature extraction of sound signals is to extract key parameters that can effectively represent the speech content from the original waveform, so as to accurately recognize the voice wake-up word in the subsequent steps. In this embodiment, feature extraction is performed on the sound signal to obtain two-dimensional signal features of at least two channels. C represents the number of channels, T represents the time dimension, and F represents the frequency dimension. The dimensions of the two-dimensional signal features of at least two channels can be expressed as C×T×F, where C≥2.
[0029] In a specific embodiment, to obtain two-dimensional signal features in two dimensions, namely time and frequency, the time-domain sound signal can be framed, and each frame can be Fourier transformed to obtain a time-varying spectrogram, which serves as the basis for subsequent feature extraction. This embodiment does not limit the feature extraction method for the sound signal; for example, Mel-Frequency Cepstral Coefficients (MFCCs) can be extracted as signal features.
[0030] In this embodiment, there is no limitation on the method for obtaining two-dimensional signal features of at least two channels. For example, the two-dimensional signal features of a single channel obtained by feature extraction of a single-channel sound signal can be channel-expanded to expand them into two-dimensional signal features of at least two channels. The method of channel expansion is not limited here. For another example, a microphone array including at least two microphones can be used to acquire sound signals of at least two channels, and feature extraction can be performed on the sound signals of each channel to obtain two-dimensional signal features of a single channel corresponding to each channel, and then the two-dimensional signal features of at least two channels can be obtained by combining them. For another example, at least two feature extraction methods can be used to extract features of a single-channel sound signal to obtain at least two two-dimensional signal features, with each two-dimensional signal feature being used as a channel to obtain two-dimensional signal features of at least two channels.
[0031] In one feasible implementation, to improve the accuracy of voice wake-up, the FBANK (Filter Bank Energies) of the sound signal can be extracted as a signal feature. The step S10 includes S101-S102: Step S101: extract features from the sound signal to obtain FBANK features.
[0032] In a specific implementation, the sound signal can be segmented to obtain a series of short-time frames (for example, 20-40ms per frame, with a frame shift of 10-20ms); an STFT (Short-Time Fourier Transform) is performed on each frame signal to obtain a spectrum, which is then passed through a set of Mel filter banks, each filter corresponding to a specific frequency band (the center frequency is nonlinearly distributed according to the Mel scale, with low frequencies dense and high frequencies sparse), each Mel filter outputs an energy value, and the logarithm of the energy value is taken; the resulting FBANK feature is a two-dimensional matrix of T'×F', where T' is the number of frames, F' is the number of Mel filters, and the element in the i-th row and j-th column is the result of taking the logarithm of the energy value output by the j-th Mel filter in the i-th frame.
[0033] It should be noted that the FBANK-based feature extraction method compresses the spectral energy into a frequency band (Mel band) that simulates the auditory characteristics of the human ear. Each filter output represents the "intensity" of the frequency band. Compared with the original STFT spectrum, the Mel filter bank greatly reduces the dimensionality of the frequency dimension (for example, from several hundred to 40 or 80), while retaining the most important auditory-related information and conforming to the nonlinear characteristics of frequency perception. While reducing computational complexity, it can improve the accuracy of voice wake-up.
[0034] Step S102 : performing channel expansion and dimension transformation on the FBANK features to obtain two-dimensional signal features of at least two channels.
[0035] The extracted FBANK features can be expressed as 1×T'×F', that is, a single-channel T'×F' two-dimensional matrix. The FBANK features can be channel-expanded and dimensionally transformed to obtain C×T×F signal features, where C ≥ 2. The relationship between T and T', as well as the relationship between F and F', can be set as needed and are not limited here.
[0036] The order of channel expansion and dimensionality conversion is not restricted in this embodiment. That is, channel expansion can be performed first and then dimensionality conversion, or dimensionality conversion can be performed first and then channel expansion, depending on the needs. Through channel expansion and dimensionality conversion, signal features are transformed to a number of channels and dimensions more suitable for subsequent convolution processing, while reducing redundant information between frequencies, thereby improving the expressiveness of features and facilitating improved voice wake-up accuracy.
[0037] There are many ways to achieve channel expansion, which are not limited in this embodiment. For example, taking the FBANK feature as an example, a convolution kernel with a shape of C×1×1×1 can be used to convolve the FBANK feature, expand the FBANK feature into a signal feature of C×T'×F', and then perform a dimensional transformation to obtain a signal feature of C×T×F. There are many ways to achieve dimensional transformation, which are not limited in this embodiment. For example, a linear layer can be set to perform a linear transformation on the FBANK feature to obtain a two-dimensional matrix of T×F, and then the two-dimensional matrix of T×F can be channel expanded to obtain a signal feature of C×T×F.
[0038] Step S20, performing convolution processing on the two-dimensional signal features of each channel to obtain a feature map of at least one channel, wherein the convolution processing includes at least one depth-wise separable convolution processing, and the depth-wise separable convolution processing includes two-dimensional depth-wise convolution processing and two-dimensional point-wise convolution processing.
[0039] The two-dimensional depth-wise convolution process is to perform a two-dimensional convolution on each input channel separately, with the same number of channels as the input. Figure 2 , a schematic diagram of two-dimensional depth-wise convolution processing is given. Taking the input of three-channel two-dimensional features as an example, the figure uses three 3×3 convolution kernels to perform separate two-dimensional convolutions, outputting three-channel two-dimensional features. Two-dimensional convolution refers to the convolution kernel moving in the time and frequency dimensions to perform convolution processing. Through two-dimensional depth-wise convolution processing, the number of model parameters and the amount of computation can be effectively reduced. At the same time, compared with one-dimensional convolution, two-dimensional convolution can jointly capture local patterns in the two-dimensional time-frequency structure of speech (such as formant trajectories and phoneme boundaries), significantly improving feature discrimination and noise robustness compared to one-dimensional convolution. Two-dimensional point-by-point convolution processing is a two-dimensional convolution that does not change the input feature dimension. The convolution kernel size is 1×1×M, where M is the number of output channels of the previous layer.
[0040] Since the two-dimensional depth-wise convolution process performs convolution operations on each channel of the input layer independently, it does not effectively utilize the feature information of different channels at the same spatial position. In order to fully utilize the feature information between different channels, the feature map output by the two-dimensional depth-wise convolution process is weighted and summed in the depth direction through two-dimensional point-by-point convolution, thereby effectively utilizing the feature information of different channels at the same spatial position. Figure 3 , a schematic diagram of two-dimensional point-by-point convolution processing is given. In the figure, taking the two-dimensional features of three channels as input as an example, three convolution kernels of size 1×1×3 are used for convolution processing, and the two-dimensional features of three channels are output.
[0041] Compared with ordinary two-dimensional convolution, depthwise separable convolution processing ensures the capture of detailed information of time-frequency domain features. It can not only greatly reduce the number of algorithm parameters and computational complexity, but also improve the performance of on-device voice wake-up in complex scenarios.
[0042] In one feasible implementation, to enhance the modeling capability and the ability to capture information at different levels, the two-dimensional signal features of each channel are convolved by stacking at least two depth-wise separable convolution structures. Specifically, step S20 includes: Step S201, using at least two preset depth-separable convolution structures connected in sequence to perform convolution processing on the two-dimensional signal features of each channel to obtain a feature map of at least one channel, wherein the input data of the first depth-separable convolution structure is the two-dimensional signal features of each channel, and the input data of each depth-separable convolution structure after the first depth-separable convolution structure is the output data of the previous depth-separable convolution structure.
[0043] The depth-wise separable convolution structure includes a depth-wise convolution module and a point-wise convolution module, which are used to perform two-dimensional depth-wise convolution processing and two-dimensional point-by-point convolution processing on the input data. The number of depth-wise separable convolution structures can be set in advance as needed and is not limited in this embodiment. The depth-wise separable convolution structures are connected in sequence, that is, the two-dimensional signal features of each channel are convolved by the first depth-wise separable convolution structure to obtain output data (also called intermediate feature map), and then the output data is input into the second depth-wise separable convolution structure for processing to obtain output data, and so on, to obtain the output data obtained by the last depth-wise separable convolution structure; at least the feature map of at least one channel of the final output is obtained based on the output data of the last depth-wise separable convolution structure.
[0044] In this embodiment, the stacking of multiple layers of depth-wise separable convolutional structures increases the receptive field of the model, thereby enhancing the model's ability to model contextual information, thereby further improving the accuracy of voice wake-up.
[0045] Step S30: obtaining a voice wake-up word recognition result based on the feature map recognition of each channel.
[0046] After obtaining the feature graph of at least one channel, a voice wake-up word recognition result can be obtained based on the feature graph of each channel. The voice wake-up word recognition result can be a recognition result indicating whether the sound signal contains the voice wake-up word. In this embodiment, the method of identifying the voice wake-up word recognition result based on the feature graph is not limited.
[0047] In one feasible implementation, the step S30 includes S301 to S303: Step S301: Input the feature map of each channel into a preset classification layer for classification to obtain a probability value.
[0048] Step S302: If the probability value is greater than a preset threshold, a recognition result is obtained that the sound signal contains a voice wake-up word.
[0049] Step S303: If the probability value is less than or equal to the preset threshold, a recognition result is obtained that the sound signal does not contain the voice wake-up word.
[0050] The classification layer can be implemented using a linear layer, and the implementation method of the classification layer is not limited in this embodiment. The preset threshold can be set as needed, for example, set to 0.8. The feature map of each channel is input into the classification layer for classification, and a probability value in the range of 0-1 is output. The probability value is compared with the preset threshold. If it is greater than the preset threshold, the recognition result of the voice wake-up word included in the sound signal is output. Otherwise, the recognition result of the voice wake-up word not included in the sound signal is output.
[0051] It should be noted that this embodiment does not limit the specific training method for the neural network model containing linear layers, depthwise separable convolutional structures, and classification layers. Training samples can use sound signals containing the target wake-up word as positive samples, and sound signals containing background noise, similar words, or irrelevant speech as negative samples.
[0052] In this embodiment, two-dimensional signal features of at least two channels are obtained by feature extraction of the sound signal, and the two dimensions are the time dimension and the frequency dimension, providing a data basis for fully capturing the feature details of the time-frequency domain; convolution processing is performed on the two-dimensional signal features of each channel, and the convolution processing includes at least one depth-wise separable convolution processing, and the depth-wise separable convolution processing includes two-dimensional depth-by-depth convolution processing and two-dimensional point-by-point convolution processing. The feature map of at least one channel obtained by the convolution processing is recognized to obtain the voice wake-up word recognition result. Compared with ordinary two-dimensional convolution, while ensuring the full capture of the feature details of the time-frequency domain, the number of parameters and the amount of calculation of the algorithm are greatly reduced, and the performance of the end-side voice wake-up in complex scenarios is improved, and finally accurate voice command triggering is achieved when the end-side device is in standby state.
[0053] Based on the above-mentioned first embodiment, a second embodiment of the voice wake-up method of the present application is proposed. In this embodiment, the same or similar contents as those of the above-mentioned first embodiment can be referred to the above introduction and will not be repeated hereafter. In this embodiment, for each of the depthwise separable convolution structures, the step of convolution processing the input data using the depthwise separable convolution structure in step S201 includes S2011~S2012: Step S2011: Input the input data into the depthwise separable convolution structure to perform convolution processing to obtain a processing result.
[0054] It should be noted that, for the first depth-wise separable convolution structure, the two-dimensional signal features of each channel are input into the depth-wise separable convolution structure for convolution processing to obtain the output processing result; for the second and subsequent depth-wise separable convolution structures, the output data of the previous depth-wise separable convolution structure is used as data input and input into the depth-wise separable convolution structure for convolution processing to obtain the processing result.
[0055] Step S2012: Add the input data to the processing result to obtain output data of the depthwise separable convolution structure.
[0056] In order to better retain the input information and improve the stability of training, the input data of each depth-wise separable convolutional structure is added to the processing result obtained by convolution processing on the input data, and the result of the addition is used as the output data of the depth-wise separable convolutional structure.
[0057] In a specific embodiment, to facilitate the addition of input data and processed results, the size of the input data can be the same as the size of the processed result, for example, if both are C×T×F, then the input data and the processed result can be added element-by-element, and the size of the resulting output data is also C×T×F. The size of the convolutional layer in the depthwise separable convolutional structure can be designed so that the size of both the input data and the output data is C×T×F.
[0058] In a feasible embodiment, the depthwise separable convolution structure includes a depthwise convolution module, a first normalized activation module, a pointwise convolution module and a second normalized activation module connected in sequence; the depthwise convolution module includes a first type of convolution kernel corresponding to the two-dimensional features of each channel in the input data, and the first type of convolution kernel is used to perform convolution processing on the two-dimensional features of the corresponding channel; the point-by-point convolution module includes at least one second type of convolution kernel, and the size of the second type of convolution kernel is 1×1×M, where M is the number of output channels of the first normalized activation module.
[0059] It should be noted that the convolution kernel in the depth-by-depth convolution module is called the first type of convolution kernel, and the convolution kernel in the point-by-point convolution module is called the second type of convolution kernel, to distinguish them. The size of the first type of convolution kernel can be set as needed, and is not limited in the present embodiment. For example, a 3×3 convolution kernel is used. The number of first-type convolution kernels is the same as the number of channels of the input data, and each convolution kernel is responsible for performing convolution processing on the two-dimensional features of a channel in the input data. The size of the second-type convolution kernel is 1×1×M, which is responsible for performing convolution processing on the two-dimensional features of each channel of the input data, and outputting the processing result of one channel. Multiple second-type convolution kernels are used to obtain the processing results of multiple channels. The number of second-type convolution kernels can be set as needed, and is not limited in the present embodiment. For example, in order to make the sizes of the input data and output data C×T×F, the number of second-type convolution kernels can be C.
[0060] The first normalized activation module and the second normalized activation module both include a normalization module and an activation module. The normalization module is used to perform batch normalization processing, and the activation module is used to perform activation processing, such as activation processing through the Relu activation function (Rectified Linear Unit). Through batch normalization processing and activation processing, the data distribution can be stabilized, nonlinearity can be introduced, and the generalization ability of the model can be improved.
[0061] Based on the above depth-wise separable convolution structure, the process of processing input data may include: first inputting the input data into the depth-wise convolution module, and performing convolution processing on each channel of the input data through each first type of convolution kernel in the depth-wise convolution module, and the result obtained is called the first intermediate output data for distinction; inputting the first intermediate output data into the first normalization activation module, and performing batch normalization processing and activation processing, and the result obtained is called the second intermediate output data for distinction; inputting the second intermediate output data into the point-by-point convolution module, and performing convolution processing through each second type of convolution kernel in the point-by-point convolution module, and the result obtained is called the third intermediate output data for distinction; inputting the third intermediate output data into the second normalization activation module, and performing batch normalization processing and activation processing, and the result obtained is the processing result of the depth-wise separable convolution structure; the processing result can be added to the input data to obtain the output data of the depth-wise separable convolution structure.
[0062] In one feasible implementation, step S201 includes S2013-S2014: Step S2013: Perform convolution processing on the two-dimensional signal features of each channel using at least two preset depthwise separable convolution structures connected in sequence to obtain output data corresponding to each of the depthwise separable convolution structures.
[0063] Step S2014: Add the output data to obtain a feature map of at least one channel.
[0064] To further enhance the ability to capture hidden features in sound signals, in this embodiment, the two-dimensional signal features of each channel can be convolved based on at least two sequentially connected depthwise separable convolutional structures. After obtaining the output data corresponding to each depthwise separable convolutional structure, the output data of each depthwise separable convolutional structure is added together to obtain a feature map for at least one channel. Specifically, if the output data of each depthwise separable convolutional structure is of size C×T×F, the output data can be element-wise added together to obtain a feature map of size C×T×F.
[0065] By adding the output features of each depth-wise separable convolutional structure, making full use of the hidden features at different levels, and then performing recognition to obtain the voice wake-up word recognition results, the accuracy of voice wake-up can be further improved.
[0066] In one possible implementation, the following Figure 4 The neural network structure shown processes the two-dimensional signal features extracted from the sound signal. Figure 4 In the process, the two-dimensional signal features of at least two channels are transformed by the linear layer and then input into N stacked depthwise separable convolution structures for convolution processing to obtain a feature map of at least one channel. The feature map is input into the classification layer for classification to obtain a probability value. The speech wake-up word recognition result is obtained by comparing the probability value with the threshold. Each depthwise separable convolution structure includes a depthwise convolution module, a first normalized activation module, a pointwise convolution module, and a second normalized activation module. The input data of each depthwise separable convolution structure is added to the processing result to serve as the output data of the depthwise separable convolution structure; the output data of each depthwise separable convolution structure are added to obtain a feature map of at least one channel for inputting into the classification layer. Through this neural network structure, a speech wake-up architecture with time-frequency depthwise separable convolution is realized. This architecture significantly reduces the number of parameters and computational complexity of the model, and ultimately achieves accurate voice command triggering in the standby state of the end-side device.
[0067] An embodiment of the present application provides an electronic device, which includes: at least one processor; and a memory communicatively connected to the at least one processor; wherein the memory stores instructions that can be executed by the at least one processor, and the instructions are executed by the at least one processor to enable the at least one processor to execute the voice wake-up method in the above embodiment.
[0068] Reference below Figure 5, which shows a schematic diagram of the structure of an electronic device suitable for implementing the embodiments of the present application. The electronic devices in the embodiments of the present application may include, but are not limited to, mobile terminals such as mobile phones, laptop computers, digital broadcast receivers, PDAs (Personal Digital Assistants), PADs (Portable Application Descriptions), PMPs (Portable Media Players), in-vehicle terminals (such as in-vehicle navigation terminals), and fixed terminals such as digital TVs and desktop computers. Figure 5 The electronic device shown is merely an example and should not limit the functions and scope of use of the embodiments of the present application.
[0069] like Figure 5 As shown, the electronic device may include a processing device 1001 (e.g., a central processing unit, a graphics processing unit, etc.), which can perform various appropriate actions and processes based on programs stored in a read-only memory 1002 or programs loaded from a storage device 1003 into a random access memory 1004. The random access memory 1004 also stores various programs and data required for the operation of the electronic device. The processing device 1001, the read-only memory 1002, and the random access memory 1004 are interconnected via a bus 1005. An input / output interface 1006 is also connected to the bus. Typically, the following systems may be connected to the input / output interface 1006: an input device 1007 including, for example, a touchscreen, touchpad, keyboard, mouse, image sensor, microphone, accelerometer, gyroscope, etc.; an output device 1008 including, for example, a liquid crystal display (LCD), speaker, vibrator, etc.; a storage device 1003 including, for example, a magnetic tape or hard disk; and a communication device 1009. The communication device 1009 may allow the electronic device to communicate with other devices wirelessly or wired to exchange data. Although the figures show electronic devices with various systems, it should be understood that it is not required to implement or have all the systems shown. More or fewer systems may be implemented or have instead.
[0070] In particular, according to the embodiments disclosed in the present application, the processes described above with reference to the flowcharts can be implemented as computer software programs. For example, the embodiments disclosed in the present application include a computer program product comprising a computer program carried on a computer-readable medium, the computer program comprising program code for executing the method shown in the flowchart. In such an embodiment, the computer program can be downloaded and installed from a network via a communication device, or installed from a storage device 1003, or installed from a read-only memory 1002. When the computer program is executed by the processing device 1001, the above-mentioned functions defined in the method of the embodiment disclosed in the present application are performed.
[0071] The electronic device provided in the embodiment of the present application adopts the voice wake-up method in the above embodiment. Compared with the prior art, the beneficial effects of the electronic device provided in the present application are the same as the beneficial effects of the voice wake-up method provided in the above embodiment, and the other technical features in the electronic device are the same as the features disclosed in the method of the previous embodiment, which will not be repeated here.
[0072] It should be understood that the various parts disclosed in this application can be implemented using hardware, software, firmware, or a combination thereof. In the description of the above embodiments, specific features, structures, materials, or characteristics can be combined in any one or more embodiments or examples in a suitable manner.
[0073] The above description is merely a specific embodiment of the present application, but the scope of protection of the present application is not limited thereto. Any changes or substitutions that can be easily conceived by a person skilled in the art within the technical scope disclosed in this application should be included in the scope of protection of this application. Therefore, the scope of protection of this application should be based on the scope of protection of the claims.
[0074] An embodiment of the present application provides a computer-readable storage medium having computer-readable program instructions (ie, a computer program) stored thereon, and the computer-readable program instructions are used to execute the voice wake-up method in the above embodiment.
[0075] The computer-readable storage medium provided in the embodiments of the present application may be, for example, a USB flash drive, but is not limited to electrical, magnetic, optical, electromagnetic, infrared, or semiconductor systems or devices, or any combination thereof. More specific examples of computer-readable storage media may include, but are not limited to: an electrical connection having one or more wires, a portable computer disk, a hard disk, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination thereof. In this embodiment, the computer-readable storage medium may be any tangible medium that contains or stores a program that can be used by or in conjunction with an instruction execution system or device. The program code contained on the computer-readable storage medium may be transmitted using any suitable medium, including but not limited to: wires, optical cables, RF (Radio Frequency), etc., or any suitable combination thereof.
[0076] The computer-readable storage medium may be included in the electronic device, or may exist independently without being incorporated into the electronic device.
[0077] The computer-readable storage medium carries one or more programs. When the one or more programs are executed by an electronic device, the electronic device performs the functions defined in the method of the embodiment disclosed in this application.
[0078] Computer program code for performing the operations of the present application may be written in one or more programming languages, or a combination thereof, including object-oriented programming languages such as Java, Smalltalk, C++, and conventional procedural programming languages such as "C" or similar programming languages. The program code may be executed entirely on the user's computer, partially on the user's computer, as a stand-alone software package, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In the case of a remote computer, the remote computer may be connected to the user's computer via any type of network, including a local area network (LAN) or a wide area network (WAN), or may be connected to an external computer (e.g., via the Internet using an Internet service provider).
[0079] The flow charts and block diagrams in the accompanying drawings illustrate the possible architecture, functions and operations of the systems, methods and computer program products according to various embodiments of the present application. In this regard, each box in the flow chart or block diagram can represent a module, program segment or a part of code, and the module, program segment or a part of code contains one or more executable instructions for realizing the specified logical function. It should also be noted that in some alternative implementations, the functions marked in the box can also occur in a different order than that marked in the accompanying drawings. For example, two boxes represented in succession can actually be executed substantially in parallel, and they can sometimes be executed in the opposite order, depending on the functions involved. It should also be noted that each box in the block diagram and / or flow chart, and the combination of the boxes in the block diagram and / or flow chart can be implemented by a dedicated hardware-based system that performs the specified function or operation, or can be implemented by a combination of dedicated hardware and computer instructions.
[0080] The modules described in the embodiments of the present application may be implemented in software or hardware, wherein the name of a module does not necessarily limit the unit itself.
[0081] The readable storage medium provided in the embodiment of the present application is a computer-readable storage medium, which stores computer-readable program instructions (i.e., computer program) for executing the above-mentioned voice wake-up method. Compared with the prior art, the beneficial effects of the computer-readable storage medium provided in the embodiment of the present application are the same as the beneficial effects of the voice wake-up method provided in the above-mentioned embodiment, and will not be repeated here.
[0082] An embodiment of the present application further provides a computer program product, including a computer program, which implements the steps of the above-mentioned voice wake-up method when executed by a processor.
[0083] Compared with the prior art, the beneficial effects of the computer program product provided in the embodiment of the present application are the same as the beneficial effects of the voice wake-up method provided in the above embodiment, and will not be repeated here.
[0084] The above description is only part of the embodiments of the present application and does not limit the patent scope of the present application. All equivalent structural transformations made by using the contents of the present application specification and drawings under the technical concept of the present application, or direct / indirect application in other related technical fields are included in the patent protection scope of the present application.
Claims
1. A voice wake-up method, characterized in that: The voice wake-up method includes: Acquiring a sound signal, and performing feature extraction on the sound signal to obtain two-dimensional signal features of at least two channels, wherein the two dimensions of the two-dimensional signal features are a time dimension and a frequency dimension; Performing convolution processing on the two-dimensional signal features of each channel to obtain a feature map of at least one channel, wherein the convolution processing includes at least one depthwise separable convolution processing, and the depthwise separable convolution processing includes two-dimensional depthwise convolution processing and two-dimensional pointwise convolution processing; The voice wake-up word recognition result is obtained based on the feature map recognition of each channel.
2. The voice wake-up method according to claim 1, wherein: The step of performing convolution processing on the two-dimensional signal features of each channel to obtain a feature map of at least one channel includes: The two-dimensional signal features of each channel are convolved using at least two preset convolution structures connected in sequence to obtain a feature map of at least one channel, wherein the input data of the first depth-separable convolution structure is the two-dimensional signal features of each channel, and the input data of each depth-separable convolution structure after the first depth-separable convolution structure is the output data of the previous depth-separable convolution structure.
3. The voice wake-up method according to claim 2, wherein: For each of the depthwise separable convolution structures, the step of performing convolution processing on input data using the depthwise separable convolution structure includes: Inputting the input data into the depthwise separable convolution structure to perform convolution processing to obtain a processing result; The input data is added to the processing result to obtain output data of the depthwise separable convolution structure.
4. The voice wake-up method according to claim 3, wherein: The depthwise separable convolution structure includes a depthwise convolution module, a first normalized activation module, a pointwise convolution module and a second normalized activation module connected in sequence; the depthwise convolution module includes a first type of convolution kernel corresponding to the two-dimensional features of each channel in the input data, and the first type of convolution kernel is used to perform convolution processing on the two-dimensional features of the corresponding channel; the pointwise convolution module includes at least one second type of convolution kernel, and the size of the second type of convolution kernel is 1×1×M, where M is the number of output channels of the first normalized activation module.
5. The voice wake-up method according to claim 2, wherein: The step of performing convolution processing on the two-dimensional signal features of each channel using at least two preset sequentially connected depthwise separable convolution structures to obtain a feature map of at least one channel includes: Performing convolution processing on the two-dimensional signal features of each channel using at least two preset depthwise separable convolution structures connected in sequence to obtain output data corresponding to each of the depthwise separable convolution structures; The output data are added together to obtain a feature map of at least one channel.
6. The voice wake-up method according to claim 1, wherein: The step of extracting features from the sound signal to obtain two-dimensional signal features of at least two channels includes: Extracting features from the sound signal to obtain FBANK features; The FBANK features are subjected to channel expansion and dimension transformation to obtain two-dimensional signal features of at least two channels.
7. The voice wake-up method according to any one of claims 1 to 6, characterized in that: The step of obtaining a voice wake-up word recognition result based on the feature graph recognition of each channel includes: Input the feature map of each channel into a preset classification layer for classification to obtain a probability value; If the probability value is greater than a preset threshold, a recognition result is obtained that the sound signal contains the voice wake-up word; If the probability value is less than or equal to the preset threshold, a recognition result is obtained that the sound signal does not contain the voice wake-up word.
8. An electronic device, characterized in that: The electronic device includes: a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the computer program is configured to implement the steps of the voice wake-up method according to any one of claims 1 to 7.
9. A storage medium, characterized in that: The storage medium is a computer-readable storage medium, and a computer program is stored on the storage medium. When the computer program is executed by a processor, the steps of the voice wake-up method according to any one of claims 1 to 7 are implemented.
10. A computer program product, characterized in that The computer program product includes a computer program, and when the computer program is executed by a processor, the steps of the voice wake-up method according to any one of claims 1 to 7 are implemented.
Citation Information
Patent Citations
Neural network model establishing method and voice waking method, device thereof, medium and equipment
CN109448719A
Voice awakening method and device, electronic equipment and storage medium
CN111933111A
Lightweight neural network voice keyword recognition method based on hierarchical quantification
CN112786021A
Voice recognition method and device, electronic equipment and computer readable storage medium
CN113628612A
Audio identification method and system for juveniles
CN113793602A