Voice activity detection method and device, computer equipment and storage medium
By using pre-trained neural networks to supplement missing features in speech activity detection, the accuracy problem of traditional methods when dealing with noise and incomplete data is solved, and more efficient speech activity detection is achieved.
Patent Information
- Application Number
- CN202311702242.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2023-12-12
- Publication Date
- 2025-06-13
AI Technical Summary
When traditional speech activity detection methods process voice signals with noise, it is difficult to accurately detect the start and end points of the speech signal, especially when there is incomplete or noise contaminated in the input data.
By obtaining audio sampled data, the acoustic features of the audio frame are extracted and the feature map is generated. The feature map is then inputted into the pre-trained neural network for processing, generating a predictive feature map after the missing features is supplemented, and finally performing speech activity detection based on the predicted feature map.
This method can more accurately restore the broken feature information by supplementing the missing features of the pre-trained neural network, thereby improving the accuracy of speech activity detection.
Smart Images

Figure CN120148563A_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the technical field of speech processing, and in particular, to a method, apparatus, computer device, and storage medium for speech activity detection. Background Art
[0002] With the development of speech processing technology, speech activity detection technology has emerged. Speech activity detection not only involves issues of digital signal processing, but also involves auditory perception characteristics and human speech features. At the same time, the diversity of noise increases the difficulty of speech activity detection. It is very difficult to determine the starting point and ending point of a speech signal from a noisy speech signal.
[0003] In traditional technologies, a deep learning-based method is used, for example, a convolutional neural network structure is adopted to complete the extraction of robust feature expressions. It has achieved a good detection effect on non-stationary noise signals, but the model cannot notice the incomplete parts or the parts contaminated by noise in the entire input data. Therefore, when there are incomplete parts or parts contaminated by noise in the input speech data, it is easy to cause inaccurate speech activity detection. Summary of the Invention
[0004] Based on this, it is necessary to provide a method, apparatus, computer device, and storage medium for speech activity detection that can improve the accuracy of speech activity detection for the above technical problems.
[0005] In a first aspect, a method for speech activity detection is provided. The method includes:
[0006] Obtaining audio sampling data;
[0007] Extracting acoustic features of each audio frame in the audio sampling data, and generating a feature map according to the acoustic features of each audio frame;
[0008] Inputting the feature map into a pre-trained neural network to obtain a predicted feature map corresponding to the feature map output by the pre-trained neural network; wherein, the predicted feature map is a feature map after supplementing the missing features of the feature map; and
[0009] Performing speech activity detection according to the predicted feature map.
[0010] In some embodiments, extracting acoustic features of each audio frame in the audio sampling data, and generating a feature map according to the acoustic features of each audio frame includes: dividing the audio sampling data into frames to obtain each audio frame; extracting the corresponding filter bank features of each audio frame as acoustic features; and concatenating the filter bank features to obtain a feature map.
[0011] In some embodiments, the pre-trained neural network includes a chunk embedding layer, a masking layer, an encoder layer, a decoding embedding layer, and a decoder layer. Inputting a feature map into the pre-trained neural network to obtain a predicted feature map output by the pre-trained neural network, including: dividing the feature map into multiple chunks through the chunk embedding layer, and performing dimensionality reduction processing on each chunk to obtain a chunk vector corresponding to each chunk; performing a random masking operation on the multiple chunk vectors through the masking layer, and combining the unmasked chunk vectors that are not masked to generate a chunk sequence and inputting it into the encoding layer; encoding the chunk sequence through the encoding layer to obtain a first hidden feature; concatenating the masked chunk vectors with the first hidden feature through the decoding embedding layer to obtain a concatenated feature; and decoding the concatenated feature through the decoder layer to obtain a second hidden feature, and performing regression processing on the second hidden feature to obtain a predicted feature map.
[0012] In some embodiments, the encoder layer includes a first normalization layer, a multi-head attention mechanism layer, a second normalization layer, and a multi-layer perceptron layer. Encoding the chunk sequence through the encoding layer to obtain a first hidden feature, including: processing the chunk sequence through the first normalization layer to obtain a first output; performing a calculation based on an attention matrix on the first output through the multi-head attention mechanism layer to obtain a second output; concatenating the second output with the chunk sequence to obtain a third output; processing the third output through the second normalization layer to obtain a fourth output; mapping the fourth output through the multi-layer perceptron layer to obtain a fifth output; and concatenating the fifth output with the third output to obtain a first hidden feature.
[0013] In some embodiments, the method further includes: training the pre-trained neural network according to the mean squared error loss function and based on the stochastic gradient descent optimization algorithm.
[0014] In some embodiments, performing voice activity detection according to the predicted feature map, including: classifying the predicted feature map based on a fully connected layer to obtain a classification output; and normalizing the classification output to obtain the probability distribution of each speech frame falling into a voice activity frame and a non-voice activity frame, and performing voice activity detection according to the probability distribution.
[0015] In some embodiments, a method for training a pre-trained neural network includes: collecting vehicle noise and vehicle audio to generate training samples; preprocessing the training samples to obtain sample feature maps corresponding to the training samples; constructing a pre-trained neural network; wherein the pre-trained neural network includes a block embedding layer, a masking layer, an encoder layer, a decoding embedding layer, and a decoder layer; sequentially processing the sample feature maps through the block embedding layer, the masking layer, the encoder layer, the decoding embedding layer, and the decoder layer to obtain predicted sample feature maps; and optimizing the model parameters of the block embedding layer, the masking layer, the encoder layer, the decoding embedding layer, and the decoder layer according to the training samples, the predicted sample feature maps, and the mean squared error loss function.
[0016] In a second aspect, a voice activity detection device is provided, which includes:
[0017] An audio data acquisition module for acquiring audio sampling data;
[0018] A feature map generation module for extracting acoustic features of each audio frame in the audio sampling data and generating a feature map according to the acoustic features of each audio frame;
[0019] A pre-trained neural network module for inputting the feature map into a pre-trained neural network to obtain a predicted feature map corresponding to the feature map output by the pre-trained neural network; wherein the predicted feature map is a feature map after supplementing missing features of the feature map; and
[0020] A fully connected classification processing module for performing voice activity detection according to the predicted feature map.
[0021] In a third aspect, a computer device is provided, including a memory, a processor, and a computer program stored on the memory and executable on the processor. When the processor executes the computer program, the steps of the voice activity detection method according to any embodiment of the first aspect are implemented.
[0022] In a fourth aspect, a computer-readable storage medium is provided, on which a computer program is stored. When the computer program is executed by a processor, the steps of the voice activity detection method according to any embodiment of the first aspect are implemented.
[0023] The above voice activity detection method, device, computer device and storage medium extract features from audio sampling data to obtain corresponding feature maps, and perform prediction calculations through a pre-trained neural network to predict missing partial data in the feature maps, and then restore incomplete feature information to generate predicted feature maps. Voice activity detection is performed based on the predicted feature maps supplemented with missing features. By using this method, the missing partial features in the data can be fully predicted by using neural network machine learning technology, and then the incomplete feature information can be restored. Therefore, voice activity detection can be performed based on more complete predicted feature maps, thereby improving the accuracy of voice activity detection. BRIEF DESCRIPTION OF THE DRAWINGS
[0024] Figure 1 FIG. is an application environment diagram of the voice activity detection method in some embodiments;
[0025] Figure 2 FIG. is a schematic flowchart of the voice activity detection method in some embodiments;
[0026] Figure 3 FIG. is a schematic structural diagram of the pre-trained neural network in some embodiments;
[0027] Figure 4 FIG. is a schematic flowchart of the steps of inputting a feature map into the pre-trained neural network to obtain the predicted feature map output by the pre-trained neural network in some embodiments;
[0028] Figure 5 FIG. is a schematic structural diagram of the encoder layer of the pre-trained neural network involved in the present application in some embodiments;
[0029] Figure 6 FIG. is a schematic structural diagram of the voice activity detection device in some embodiments;
[0030] Figure 7 FIG. is an internal structure diagram of a computer device in some embodiments. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0031] In order to make the objectives, technical solutions and advantages of the present application clearer, the present application will be further described in detail below with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present application and are not used to limit the present application.
[0032] The voice activity detection method provided by the present application can be applied to an application environment as shown in Figure 1 . Among them, Figure 1The vehicle 100 shown may include an in-vehicle terminal 110. The in-vehicle terminal 110 may include at least one memory and at least one processor. A computer program is stored in the at least one memory. When the computer program is executed by the at least one processor, a vehicle pose calculation method according to an exemplary embodiment of the present application is executed. Here, the in-vehicle terminal 110 does not have to be a single electronic device, and may also be an aggregate of any devices or circuits that can execute the above computer program alone or jointly.
[0033] In the in-vehicle terminal 110, the processor may include a central processing unit (CPU), a graphics processing unit (GPU), a programmable logic device, a dedicated processor system, a microcontroller, or a microprocessor. By way of example and not limitation, the processor may also include an analog processor, a digital processor, a microprocessor, a multi-core processor, a processor array, a network processor, etc.
[0034] In the in-vehicle terminal 110, the processor may run the computer program stored in the memory. The computer program may be divided into one or more modules / units (such as computer program 1, computer program 2,...). One or more modules / units are stored in the memory and executed by the processor to complete the method of the embodiment of the present application. One or more modules / units may be a series of computer program instructions capable of completing specific functions, and the instructions may describe the execution process of the computer program in the in-vehicle terminal device. For example, the detail compensation model in the embodiment of the present application may be one of the modules / units.
[0035] The memory may be integrated with the processor. For example, RAM or flash memory may be arranged within an integrated circuit microprocessor, etc. In addition, the memory may include independent devices, such as external disk drives, storage arrays, or other storage devices that can be used by any database system. The memory and the processor may be operatively coupled, or may communicate with each other, for example, through an I / O port, a network connection, etc., so that the processor can read the files stored in the memory.
[0036] In addition, the in-vehicle terminal 110 may further include a display device (such as a liquid crystal display) and a user interaction interface (such as a keyboard, a mouse, a touch input device, etc.). All components of the in-vehicle terminal 110 may be connected to each other via a bus and / or a network.
[0037] Specifically, the in-vehicle terminal 110 acquires audio sampling data; extracts the acoustic features of each audio frame in the audio sampling data, and generates a feature map according to the acoustic features of each audio frame; inputs the feature map into a pre-trained neural network to obtain a predicted feature map corresponding to the feature map output by the pre-trained neural network; wherein, the predicted feature map is a feature map after missing features of the feature map are supplemented; and performs voice activity detection according to the predicted feature map.
[0038] In some other embodiments, the voice activity detection method provided in this application can also be used in other application scenarios. For example, it can be applied to a computer device. It should be noted that the execution subject can be a configuration device for virtual network card resources, and this device can be implemented as part or all of the computer device through software, hardware, or a combination of software and hardware. Among them, the computer device can be a terminal, a client, or a server. The server can be a single server or a server cluster composed of multiple servers. The terminal can be an in-vehicle terminal, a smart audio device, a smart phone, a personal computer, a tablet computer, a wearable device, a smart robot, and other smart hardware devices, etc.
[0039] In some embodiments, as Figure 2 shown, a voice activity detection method is provided. Taking the case where this method is applied to the Figure 1 in-vehicle terminal as an example, it includes the following steps:
[0040] Step S202: Obtain audio sampling data.
[0041] Specifically, voice data can be collected through an audio sampling device and audio sampling can be performed to obtain audio sampling data. In the application scenario of an in-vehicle terminal, voice commands issued by the vehicle user can be collected, or the interactive voice during the conversation between the vehicle user and the in-vehicle terminal can be collected, and the collected voice commands or interactive voice can be used as audio sampling data.
[0042] Step S204: Extract the acoustic features of each audio frame in the audio sampling data, and generate a feature map according to the acoustic features of each audio frame.
[0043] Specifically, the audio sampling data can be framed, and the corresponding acoustic features can be extracted from each of the divided voice frames respectively, and the corresponding acoustic features extracted from each voice frame are concatenated to generate a feature map.
[0044] Step S206: Input the feature map into a pre-trained neural network to obtain a predicted feature map corresponding to the feature map output by the pre-trained neural network; where the predicted feature map is a feature map after supplementing the missing features of the feature map.
[0045] Among them, the pre-trained neural network can be a neural network for data prediction based on supervised learning technology, which is used to detect the missing part and predict the missing features according to the input feature map, and output a predicted feature map after supplementing the missing features.
[0046] Specifically, the feature map can be input into the pre-trained neural network, and the pre-trained neural network predicts the missing features such as the incomplete features or the parts contaminated by noise in the feature map to generate a predicted feature map after supplementing the missing features.
[0047] Step S208: Perform voice activity detection based on the predicted feature graph.
[0048] Specifically, deep learning-based voice activity detection is performed based on the predicted feature map after missing feature supplementation, and the voice segments and non-voice segments in the voice signal are distinguished by the predicted feature map after missing feature supplementation corresponding to the audio sampling data and the deep learning algorithm, thereby obtaining the result of voice activity detection.
[0049] In the above-mentioned voice activity detection method, features are extracted from audio sampling data to obtain a corresponding feature map, and a pre-trained neural network is used to perform prediction calculations to predict the missing data in the feature map, and then the incomplete feature information is restored to generate a predicted feature map, and voice activity detection is performed based on the predicted feature map after the missing features are supplemented. By adopting this method, it is possible to fully utilize the neural network machine learning technology to predict the missing features in the data, and then restore the incomplete feature information. Therefore, voice activity detection can be performed based on a more complete predicted feature map, thereby improving the accuracy of voice activity detection.
[0050] In some embodiments, extracting acoustic features of each audio frame in the audio sampling data and generating a feature map according to each acoustic feature of each audio frame includes: framing the audio sampling data to obtain each audio frame; extracting filter group features corresponding to each audio frame as acoustic features; and connecting the filter group features in series to obtain the feature map.
[0051] In this embodiment, the input audio sampling data is framed and the FBANK (Filter Bank) features of each frame are extracted respectively. By first framing and then extracting acoustic features, the efficiency and quality of feature extraction can be improved. Moreover, by connecting the acoustic features of each frame in series, a feature map that more comprehensively reflects the feature information can be generated.
[0052] In some embodiments, reference Figure 3 and Figure 4 As shown, Figure 3 shows a schematic diagram of the structure of a pre-trained neural network in some embodiments, Figure 4 A flow chart of the steps of inputting a feature map into a pre-trained neural network to obtain a predicted feature map output by the pre-trained neural network in some embodiments is shown. The pre-trained neural network may include a block embedding layer 302, a mask layer 304, an encoder layer 306, a decoding embedding layer 308, and a decoder layer 310. The step of inputting a feature map into the pre-trained neural network to obtain a predicted feature map output by the pre-trained neural network may specifically include the following steps:
[0053] Step S402: Divide the feature map into multiple patches through a patch embedding layer, and perform dimensionality reduction processing on each patch to obtain the patch vector corresponding to each patch.
[0054] Step S404: Perform a random masking operation on multiple patch vectors through a masking layer, and combine the unmasked patch vectors that are not masked to generate a patch sequence and input it into the encoding layer.
[0055] Step S406: Perform encoding processing on the patch sequence through the encoding layer to obtain the first hidden feature;
[0056] Step S408: Concatenate the masked patch vectors with the first hidden feature through a decoding embedding layer to obtain a concatenated feature; and
[0057] Step S410: Perform decoding processing on the concatenated feature through the decoder layer to obtain the second hidden feature, and perform regression processing on the second hidden feature to obtain a predicted feature map.
[0058] In the above embodiments, through block embedding processing of the feature map, combined with random masking operations, unmasked patch vectors that are not masked are screened and combined, and after changing from a two-dimensional image to a one-dimensional patch sequence, they are input into the encoding layer for autoencoding to obtain the first hidden feature. Then, the masked feature is combined with the first hidden feature, so as to perform missing feature prediction and reconstruction, obtain a new hidden feature as the second hidden feature, and generate a predicted feature map according to the reconstructed second hidden feature, thereby improving the accuracy of restoring incomplete feature information.
[0059] In some embodiments, as shown in Figure 5 shown, Figure 5 shows a schematic structural diagram of the encoder layer of the pre-trained neural network involved in the present application in some embodiments. Among them, the encoder layer may include a first Layer Normalization (LN) layer, a multi-head self-attention (MSA) layer, a second normalization layer, and a multi-layer perceptron (MLP) layer. The connection relationship between each layer can be referred to Figure 5 shown.
[0060] Exemplarily, performing encoding processing on the patch sequence through the encoding layer shown in Figure 5 to obtain the first hidden feature includes:
[0061] The segmented sequence is normalized by a first normalization layer to obtain a first output; the first output is calculated based on an attention matrix through a multi-head attention mechanism layer to obtain a second output; the second output is concatenated with the segmented sequence to obtain a third output; the third output is normalized by a second normalization layer to obtain a fourth output; the fourth output is mapped through a multi-layer perceptron layer to obtain a fifth output; and the fifth output is concatenated with the third output to obtain a first hidden feature.
[0062] Exemplarily, the calculation formulas for processing based on the first and second normalization (Layer Norm, LN) layers, the multi-head self-attention (MSA) layer, and the multi-layer perceptron (MLP) layer of the encoder layer can be referred to as follows:
[0063] z′ l = MSA(LN(z l-1 )) + z l-1
[0064] z l = MLP(LN(z′ l )) + z′ l
[0065] where z l-1 represents the segmented sequence input at the (i - 1)-th iteration, LN(z l-1 ) represents the first output obtained after normalization processing, MSA(LN(z l-1 )) represents the second output obtained after calculation based on the attention matrix, z′ l represents the third output obtained after concatenating the second output with the segmented sequence, LN(z′ l ) represents the fourth output obtained after normalizing the third output, MLP(LN(z′ l )) represents the fifth output obtained after the mapping calculation of the fourth output through the multi-layer perceptron layer, and z represents the first hidden feature obtained by concatenating the fifth output with the third output.
[0066] In some embodiments, the structure of the decoder and the calculation formulas of each layer can be configured in a consistent manner corresponding to the structure and calculation formulas of the encoder, which will not be elaborated here.
[0067] In some embodiments, the method further includes: training the pre-trained neural network according to the mean squared error loss function and based on the stochastic gradient descent optimization algorithm.
[0068] In this embodiment, during the training process of the pre-trained neural network, the mean squared error (MSE) loss function can be used to optimize its parameters, and the model can be trained based on the stochastic gradient descent optimization algorithm so that the trained pre-trained neural network can perform a recovery operation on the concatenated feature maps to obtain predicted feature maps.
[0069] In some embodiments, performing voice activity detection based on the predicted feature maps includes: classifying the predicted feature maps based on a fully connected layer to obtain a classification output; and performing a normalization process on the classification output to obtain the probability distributions of each speech frame falling into voice activity frames and non-voice activity frames, and performing voice activity detection based on the probability distributions.
[0070] In this embodiment, the predicted feature maps obtained based on the pre-trained neural network can be used as intermediate features and input into a fully connected layer for classification processing, and the classification output obtained after the classification processing can be further normalized (Softmax processing), so as to obtain the probability distributions for determining whether each speech frame belongs to a voice activity frame or a non-voice activity frame, and the voice detection result can be obtained by comparing the probabilities of each speech frame with a preset threshold.
[0071] In some embodiments, for example, in the application scenario of applying this voice activity detection method to the voice activity detection of in-vehicle terminals, the pre-trained neural network can be trained as follows:
[0072] 1. Collect in-vehicle noise and in-vehicle audio to generate training samples.
[0073] Specifically, the collected in-vehicle noise and in-vehicle speech can be used to construct an audio data set for model training, and the data in the audio data set can be used as training samples. Among them, the in-vehicle noise can include five situations: the noise with the window open on the road, the noise A with the window closed on the road, the noise B with the window closed on the road, the noise with the window open in the parking lot, and the noise with the window closed in the parking lot. The noise A with the window closed on the road and the noise B with the window closed on the road are noises from different recording sections.
[0074] 2. Preprocess the training samples to obtain the sample feature maps corresponding to the training samples.
[0075] Specifically, preprocessing the training samples can include the following steps: dividing the training samples into frames, extracting the acoustic features of each speech and concatenating them to obtain sample feature maps.
[0076] 3. Construct a pre-trained neural network; where the pre-trained neural network includes a block embedding layer, a masking layer, an encoder layer, a decoding embedding layer, and a decoder layer.
[0077] 4. The sample feature map is successively processed through a block embedding layer, a masking layer, an encoder layer, a decoded embedding layer, and a decoder layer to obtain a predicted sample feature map.
[0078] Specifically, the process of training the pre-trained neural network using the training samples is similar to the process of generating the predicted feature map using the trained pre-trained neural network, and may specifically include: dividing the sample feature map into multiple blocks through the block embedding layer, and performing dimensionality reduction processing on each block to obtain block vectors corresponding to each block; performing a random masking operation on the multiple block vectors through the masking layer, and combining the unmasked block vectors that are not masked to generate a block sequence and input it into the encoding layer; encoding the block sequence through the encoding layer to obtain a first sample hidden feature; splicing the masked block vectors with the first sample hidden feature through the decoded embedding layer to obtain a sample splicing feature; and decoding the sample splicing feature through the decoder layer to obtain a second sample hidden feature, and performing regression processing on the second sample hidden feature to obtain a predicted sample feature map.
[0079] 5. Optimize each model parameter of the block embedding layer, the masking layer, the encoder layer, the decoded embedding layer, and the decoder layer according to the training sample, the predicted sample feature map, and the least mean square error loss function.
[0080] Specifically, through the training samples and the predicted sample feature maps corresponding to each training sample, and based on the least mean square error loss function, supervised machine learning training is performed on the pre-trained neural network to optimize each learnable model parameter of each layer, so as to obtain the trained pre-trained neural network.
[0081] In this embodiment, through the above training method of the pre-trained neural network, the supervised learning technology is fully utilized to predict the missing part in the data, and then the incomplete feature information is restored. Moreover, this algorithm can also reduce the dependence on external information, capture the internal relationship of the data or features, and optimize the result of model training.
[0082] It should be understood that although Figure 2 and Figure 4 the steps in the flowchart are shown in sequence according to the arrows, these steps are not necessarily executed in the order indicated by the arrows. Unless otherwise clearly stated in this article, the execution of these steps has no strict order limit, and these steps can be executed in other orders. Moreover, Figure 2 and Figure 4At least a part of the steps therein may include multiple sub-steps or multiple stages, and these sub-steps or stages are not necessarily executed and completed at the same moment, but can be executed at different moments, and the execution order of these sub-steps or stages is not necessarily sequential, but can be executed alternately or in turn with at least a part of other steps or sub-steps or stages of other steps.
[0083] In some embodiments, as Figure 6 shown, a voice activity detection device is provided, including: an audio data acquisition module 602, a feature map generation module 604, a pre-trained neural network module 606, and a fully connected classification processing module 608, wherein:
[0084] The audio data acquisition module 602 is configured to obtain audio sampling data;
[0085] The feature map generation module 604 is configured to extract acoustic features of each audio frame in the audio sampling data, and generate a feature map according to the acoustic features of each audio frame;
[0086] The pre-trained neural network module 606 is configured to input the feature map into a pre-trained neural network to obtain a predicted feature map corresponding to the feature map output by the pre-trained neural network; wherein, the predicted feature map is a feature map after missing features of the feature map are supplemented; and
[0087] The fully connected classification processing module 608 is configured to perform voice activity detection according to the predicted feature map.
[0088] In some embodiments, the feature map generation module 604 frames the audio sampling data to obtain each audio frame; extracts filter bank features corresponding to each audio frame as acoustic features; and concatenates the filter bank features to obtain a feature map.
[0089] In some embodiments, the pre-trained neural network module 606 divides the feature map into multiple blocks through a block embedding layer, and performs dimensionality reduction processing on each block to obtain block vectors corresponding to each block; performs a random masking operation on the multiple block vectors through a masking layer, and combines the unmasked block vectors that are not masked to generate a block sequence and inputs it into an encoding layer; encodes the block sequence through the encoding layer to obtain a first hidden feature; splices the masked block vectors that are masked with the first hidden feature through a decoding embedding layer to obtain a spliced feature; and decodes the spliced feature through a decoder layer to obtain a second hidden feature, and performs regression processing on the second hidden feature to obtain a predicted feature map.
[0090] In some embodiments, the pre-trained neural network module 606 processes the segmented sequence through the first normalization layer to obtain a first output; performs a calculation based on the attention matrix on the first output through the multi-head attention mechanism layer to obtain a second output; concatenates the second output with the segmented sequence to obtain a third output; processes the third output through the second normalization layer to obtain a fourth output; maps the fourth output through a multi-layer perceptron layer to obtain a fifth output; and concatenates the fifth output with the third output to obtain a first hidden feature.
[0091] In some embodiments, the pre-trained neural network module 606 is further configured to train the pre-trained neural network according to the mean squared error loss function and based on the stochastic gradient descent optimization algorithm.
[0092] In some embodiments, the fully connected classification processing module 608 performs classification processing on the predicted feature map based on the fully connected layer to obtain a classification output; and performs normalization processing on the classification output to obtain the probability distribution of each speech frame falling into the speech activity frame and the non-speech activity frame, and performs speech activity detection according to the probability distribution.
[0093] In some embodiments, the pre-trained neural network module 606 is further configured to collect in-vehicle noise and in-vehicle audio to generate training samples; preprocess the training samples to obtain sample feature maps corresponding to the training samples; construct a pre-trained neural network; wherein the pre-trained neural network includes a segmented embedding layer, a masking layer, an encoder layer, a decoding embedding layer, and a decoder layer; sequentially process the sample feature maps through the segmented embedding layer, the masking layer, the encoder layer, the decoding embedding layer, and the decoder layer to obtain predicted sample feature maps; and optimize the model parameters of the segmented embedding layer, the masking layer, the encoder layer, the decoding embedding layer, and the decoder layer according to the training samples, the predicted sample feature maps, and the mean squared error loss function.
[0094] For the specific limitations of the speech activity detection device, reference may be made to the limitations on the speech activity detection method in the foregoing text, which will not be elaborated herein. Each module in the above speech activity detection device can be implemented in whole or in part by software, hardware, and their combination. The above-mentioned modules can be embedded in the processor of the computer device in hardware form or be independent of it, or can be stored in the memory of the computer device in software form, so that the processor can call and execute the operations corresponding to the above-mentioned modules.
[0095] In some embodiments, a computer device is provided. The computer device can be a terminal, and its internal structure diagram can be as Figure 7As shown in the figure. The computer device includes a processor, a memory, a network interface, a display screen, and an input device connected through a system bus. Among them, the processor of the computer device is used to provide computing and control capabilities. The memory of the computer device includes a non-volatile storage medium and an internal memory. The non-volatile storage medium stores an operating system and a computer program. The internal memory provides an environment for the operation of the operating system and the computer program in the non-volatile storage medium. The network interface of the computer device is used to communicate with an external terminal through a network connection. When the computer program is executed by the processor, it implements a voice activity detection method. The display screen of the computer device can be a liquid crystal display screen or an electronic ink display screen. The input device of the computer device can be a touch layer covered on the display screen, or a button, a trackball, or a touchpad set on the shell of the computer device, or an external keyboard, a touchpad, or a mouse, etc.
[0096] Those skilled in the art can understand that Figure 7 the structure shown in the figure is only a block diagram of some structures related to the solution of this application, and does not constitute a limitation on the computer device to which the solution of this application is applied. The specific computer device may include more or fewer components than those shown in the figure, or combine some components, or have different component arrangements.
[0097] In some embodiments, a computer device is provided, including a memory, a processor, and a computer program stored on the memory and executable on the processor. When the processor executes the computer program, the following steps are implemented: obtaining audio sampling data; extracting the acoustic features of each audio frame in the audio sampling data, and generating a feature map according to the acoustic features of each audio frame; inputting the feature map into a pre-trained neural network to obtain a predicted feature map corresponding to the feature map output by the pre-trained neural network; wherein, the predicted feature map is the feature map after supplementing the missing features of the feature map; and performing voice activity detection according to the predicted feature map.
[0098] In some embodiments, when the processor executes the computer program, the following steps are further implemented: framing the audio sampling data to obtain each audio frame; extracting the filter bank features corresponding to each audio frame as acoustic features; and concatenating the filter bank features to obtain a feature map.
[0099] In some embodiments, when the processor executes the computer program, the following steps are further implemented: dividing the feature map into multiple blocks through a block embedding layer, and performing dimensionality reduction processing on each block to obtain a block vector corresponding to each block; performing a random masking operation on the multiple block vectors through a masking layer, and combining the unmasked block vectors that are not masked to generate a block sequence and inputting it into an encoding layer; performing encoding processing on the block sequence through the encoding layer to obtain a first hidden feature; splicing the masked block vectors that are masked with the first hidden feature through a decoding embedding layer to obtain a spliced feature; and performing decoding processing on the spliced feature through a decoder layer to obtain a second hidden feature, and performing regression processing on the second hidden feature to obtain a predicted feature map.
[0100] In some embodiments, when the processor executes the computer program, the following steps are further implemented: processing the block sequence through a first normalization layer to obtain a first output; performing a calculation based on an attention matrix on the first output through a multi-head attention mechanism layer to obtain a second output; splicing the second output with the block sequence to obtain a third output; processing the third output through a second normalization layer to obtain a fourth output; mapping the fourth output through a multi-layer perceptron layer to obtain a fifth output; and splicing the fifth output with the third output to obtain a first hidden feature.
[0101] In some embodiments, when the processor executes the computer program, the following steps are further implemented: training the pre-trained neural network according to the mean squared error loss function and based on the stochastic gradient descent optimization algorithm.
[0102] In some embodiments, when the processor executes the computer program, the following steps are further implemented: performing classification processing on the predicted feature map based on a fully connected layer to obtain a classification output; and performing normalization processing on the classification output to obtain the probability distribution of each speech frame falling into a speech active frame and a non-speech active frame, and performing speech activity detection according to the probability distribution.
[0103] In some embodiments, when the processor executes the computer program, the following steps are further implemented: collecting vehicle noise and vehicle audio to generate training samples; preprocessing the training samples to obtain a sample feature map corresponding to the training samples; constructing a pre-trained neural network; wherein, the pre-trained neural network includes a block embedding layer, a masking layer, an encoder layer, a decoding embedding layer, and a decoder layer; sequentially processing the sample feature map through the block embedding layer, the masking layer, the encoder layer, the decoding embedding layer, and the decoder layer to obtain a predicted sample feature map; and optimizing the model parameters of the block embedding layer, the masking layer, the encoder layer, the decoding embedding layer, and the decoder layer according to the training samples, the predicted sample feature map, and the mean squared error loss function.
[0104] In some embodiments, a computer-readable storage medium is provided, on which a computer program is stored. When the computer program is executed by a processor, the following steps are implemented: obtaining audio sampling data; extracting acoustic features of each audio frame in the audio sampling data, and generating a feature map according to the acoustic features of each audio frame; inputting the feature map into a pre-trained neural network to obtain a predicted feature map corresponding to the feature map output by the pre-trained neural network; wherein the predicted feature map is a feature map after supplementing missing features of the feature map; and performing voice activity detection according to the predicted feature map.
[0105] In some embodiments, when the computer program is executed by a processor, the following steps are further implemented: framing the audio sampling data to obtain each audio frame; extracting filter bank features corresponding to each audio frame as acoustic features; and concatenating the filter bank features to obtain a feature map.
[0106] In some embodiments, when the computer program is executed by a processor, the following steps are further implemented: dividing the feature map into multiple blocks through a block embedding layer, and performing dimensionality reduction processing on each block to obtain a block vector corresponding to each block; performing a random masking operation on the multiple block vectors through a masking layer, and combining the unmasked block vectors that are not masked to generate a block sequence and inputting it into an encoding layer; encoding the block sequence through the encoding layer to obtain a first hidden feature; splicing the masked block vectors that are masked with the first hidden feature through a decoding embedding layer to obtain a spliced feature; and decoding the spliced feature through a decoder layer to obtain a second hidden feature, and performing regression processing on the second hidden feature to obtain a predicted feature map.
[0107] In some embodiments, when the computer program is executed by a processor, the following steps are further implemented: processing the block sequence through a first normalization layer to obtain a first output; performing a calculation based on an attention matrix on the first output through a multi-head attention mechanism layer to obtain a second output; splicing the second output with the block sequence to obtain a third output; processing the third output through a second normalization layer to obtain a fourth output; mapping the fourth output through a multi-layer perceptron layer to obtain a fifth output; and splicing the fifth output with the third output to obtain a first hidden feature.
[0108] In some embodiments, when the computer program is executed by a processor, the following steps are further implemented: training the pre-trained neural network according to the mean squared error loss function and based on the stochastic gradient descent optimization algorithm.
[0109] In some embodiments, when the computer program is executed by a processor, the following steps are further implemented: classifying the predicted feature map based on the fully connected layer to obtain a classification output; and normalizing the classification output to obtain the probability distribution of each speech frame falling into a speech active frame and a non-speech active frame, and performing speech activity detection according to the probability distribution.
[0110] In some embodiments, when the computer program is executed by a processor, the following steps are further implemented: collecting in-vehicle noise and in-vehicle audio to generate training samples; preprocessing the training samples to obtain sample feature maps corresponding to the training samples; constructing a pre-trained neural network; wherein the pre-trained neural network includes a block embedding layer, a masking layer, an encoder layer, a decoding embedding layer, and a decoder layer; sequentially processing the sample feature maps through the block embedding layer, the masking layer, the encoder layer, the decoding embedding layer, and the decoder layer to obtain predicted sample feature maps; and optimizing the model parameters of the block embedding layer, the masking layer, the encoder layer, the decoding embedding layer, and the decoder layer according to the training samples, the predicted sample feature maps, and the least mean square error loss function.
[0111] Those of ordinary skill in the art can understand that all or part of the processes of implementing the methods in the above embodiments can be completed by instructing relevant hardware through a computer program. The computer program can be stored in a non-volatile computer-readable storage medium. When the computer program is executed, it can include the processes of the embodiments of the above methods. Among them, any reference to a memory, storage, database, or other medium used in the various embodiments provided in the present application can include non-volatile and / or volatile memories. Non-volatile memories can include read-only memory (ROM), programmable ROM (PROM), electrically programmable ROM (EPROM), electrically erasable programmable ROM (EEPROM), or flash memory. Volatile memories can include random access memory (RAM) or external cache memory. By way of illustration and not limitation, RAM is available in various forms, such as static RAM (SRAM), dynamic RAM (DRAM), synchronous DRAM (SDRAM), double data rate SDRAM (DDR SDRAM), enhanced SDRAM (ESDRAM), synchronous link (Synchlink) DRAM (SLDRAM), memory bus (Rambus) direct RAM (RDRAM), direct memory bus dynamic RAM (DRDRAM), and memory bus dynamic RAM (RDRAM), etc.
[0112] The technical features of the above embodiments can be combined arbitrarily. For the sake of concise description, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, it should be considered as the scope described in this specification.
[0113] The embodiments described above merely represent several implementation manners of the present application. The description is relatively specific and detailed, but it should not be construed as a limitation on the scope of the invention patent. It should be noted that for those of ordinary skill in the art, without departing from the concept of the present application, several modifications and improvements can still be made, and these all fall within the protection scope of the present application. Therefore, the protection scope of the patent of the present application shall be subject to the appended claims.
Claims
1. A voice activity detection method, the method comprises: obtaining audio sampling data; extracting acoustic features of each audio frame in the audio sampling data, and generating a feature map according to the acoustic features of each audio frame; inputting the feature map into a pre-trained neural network to obtain a predicted feature map corresponding to the feature map output by the pre-trained neural network; wherein, the predicted feature map is a feature map after supplementing missing features of the feature map; and performing voice activity detection according to the predicted feature map.
2. The method according to claim 1, wherein, the extracting acoustic features of each audio frame in the audio sampling data and generating a feature map according to the acoustic features of each audio frame comprises: framing the audio sampling data to obtain each audio frame; extracting corresponding filter bank features of each audio frame as the acoustic features; and concatenating the filter bank features to obtain the feature map.
3. The method according to claim 1, wherein, the pre-trained neural network comprises a block embedding layer, a masking layer, an encoder layer, a decoding embedding layer and a decoder layer, and the inputting the feature map into the pre-trained neural network to obtain the predicted feature map output by the pre-trained neural network comprises: dividing the feature map into multiple blocks through the block embedding layer, and performing dimensionality reduction processing on each block to obtain block vectors corresponding to each block; performing a random masking operation on multiple block vectors through the masking layer, and combining unmasked block vectors that are not masked to generate a block sequence and inputting the block sequence into the encoding layer; performing encoding processing on the block sequence through the encoding layer to obtain a first hidden feature; concatenating the masked block vectors that are masked with the first hidden feature through the decoding embedding layer to obtain a concatenated feature; and performing decoding processing on the concatenated feature through the decoder layer to obtain a second hidden feature, and performing regression processing on the second hidden feature to obtain the predicted feature map.
4. The method according to claim 3, wherein, the encoder layer comprises a first normalization layer, a multi-head attention mechanism layer, a second normalization layer and a multi-layer perceptron layer, and the performing encoding processing on the block sequence through the encoding layer to obtain a first hidden feature comprises: processing the block sequence through the first normalization layer to obtain a first output; performing a calculation based on an attention matrix on the first output through the multi-head attention mechanism layer to obtain a second output; concatenating the second output with the block sequence to obtain a third output; processing the third output through the second normalization layer to obtain a fourth output; mapping the fourth output through the multi-layer perceptron layer to obtain a fifth output; and concatenating the fifth output with the third output to obtain the first hidden feature.
5. The method according to claim 1, wherein, the method further comprises: training the pre-trained neural network according to the mean squared error loss function and based on the stochastic gradient descent optimization algorithm.
6. The method according to claim 1, wherein, the voice activity detection based on the predicted feature map includes: classifying the predicted feature map based on a fully connected layer to obtain a classification output; and performing normalization processing on the classification output to obtain the probability distributions of each of the voice frames falling into voice activity frames and non-voice activity frames, and performing voice activity detection according to the probability distributions.
7. The method according to claim 1, wherein, the training method of the pre-trained neural network includes: collecting vehicle noise and vehicle audio to generate training samples; performing preprocessing on the training samples to obtain sample feature maps corresponding to the training samples; constructing the pre-trained neural network; wherein, the pre-trained neural network includes a block embedding layer, a masking layer, an encoder layer, a decoding embedding layer, and a decoder layer; successively processing the sample feature maps through the block embedding layer, the masking layer, the encoder layer, the decoding embedding layer, and the decoder layer to obtain predicted sample feature maps; and optimizing the model parameters of the block embedding layer, the masking layer, the encoder layer, the decoding embedding layer, and the decoder layer according to the training samples, the predicted sample feature maps, and the least mean square error loss function.
8. A voice activity detection device, wherein, the device includes: an audio data acquisition module for acquiring audio sampling data; a feature map generation module for extracting acoustic features of each audio frame in the audio sampling data and generating a feature map according to the acoustic features of each audio frame; a pre-trained neural network module for inputting the feature map into a pre-trained neural network to obtain a predicted feature map corresponding to the feature map output by the pre-trained neural network; wherein, the predicted feature map is a feature map with missing features supplemented; and a fully connected classification processing module for performing voice activity detection according to the predicted feature map.
9. A computer device, including a memory, a processor, and a computer program stored on the memory and executable on the processor, wherein, when the processor executes the computer program, the steps of the method according to any one of claims 1 to 7 are implemented.
10. A computer-readable storage medium, on which a computer program is stored, wherein, when the computer program is executed by a processor, the steps of the method according to any one of claims 1 to 7 are implemented.