Voice activity detection method and device, computer equipment and storage medium

Through the combination of convolutional neural network, graph convolutional neural network and multi-head attention mechanism network, the problem of poor feature expression in traditional methods is solved, and the accuracy of speech activity detection is improved, especially in complex noise environments.

CN120220736APending Publication Date: 2025-06-27MOBILITY ASIA SMART TECH CO LTD
View PDF 0 Cites 2 Cited by

Patent Information

Application Number
CN202311702240.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2023-12-12
Publication Date
2025-06-27

AI Technical Summary

Technical Problem

When traditional speech activity detection methods deal with non-stationary noise signals, they cannot effectively utilize the correlation between different parts of the input signal, resulting in poor feature expression effects, which in turn affects the detection accuracy.

Method used

The method of using convolutional neural network combined with graph convolutional neural network and multi-head attention mechanism network is used to enhance the correlation and dependency of feature expression through spatial relationship construction and rotational position encoding.

Benefits of technology

It improves the accuracy of speech activity detection, enhances the feature expression effect, and can better deal with speech activity detection in complex noise environments.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120220736A_ABST
    Figure CN120220736A_ABST
Patent Text Reader

Abstract

The invention relates to a voice activity detection method and device, computer equipment and a storage medium. The method comprises the following steps: acquiring audio sampling data, and extracting acoustic features from the audio sampling data; inputting the acoustic features into a convolutional neural network to obtain a first intermediate feature vector; inputting the first intermediate feature vector into a graph convolutional neural network for spatial relationship construction to obtain a second intermediate feature vector containing spatial relationship information among a plurality of feature channels; inputting the second intermediate feature vector into a multi-head attention mechanism network, and performing multi-head attention calculation based on rotation position coding on the second intermediate feature vector through the multi-head attention mechanism network to obtain a third intermediate feature vector; and performing voice activity detection according to the third intermediate feature vector. By adopting the method, the accuracy of voice activity detection can be improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the technical field of speech processing, and particularly to a method and device for speech activity detection, a computer device, and a storage medium. Background Art

[0002] With the development of speech processing technology, speech activity detection technology has emerged. Speech activity detection not only involves problems of digital signal processing, but also involves auditory perception characteristics and human speech features. At the same time, the diversity of noise increases the difficulty of speech activity detection. It is very difficult to determine the starting point and ending point of a speech signal from a noisy speech signal.

[0003] In traditional technologies, a method based on deep learning is used, such as adopting the structure of a convolutional neural network to complete the extraction of robust feature expressions. It has a certain improvement in the detection effect of non-stationary noise signals compared with the statistical method. However, due to the inability to notice the correlation between different parts of the entire input, the feature expression effect is poor, resulting in inaccurate speech activity detection. Summary of the Invention

[0004] Based on this, it is necessary to provide a method and device for speech activity detection, a computer device, and a storage medium that can improve speech activity detection for the above technical problems.

[0005] In a first aspect, a method for speech activity detection is provided, and the method includes:

[0006] Obtain audio sampling data, and extract acoustic features from the audio sampling data;

[0007] Input the acoustic features into a convolutional neural network to obtain a first intermediate feature vector;

[0008] Input the first intermediate feature vector into a graph convolutional neural network for spatial relationship construction to obtain a second intermediate feature vector containing spatial relationship information between multiple feature channels;

[0009] Input the second intermediate feature vector into a multi-head attention mechanism network, and perform multi-head attention calculation based on rotational position encoding on the second intermediate feature vector through the multi-head attention mechanism network to obtain a third intermediate feature vector; and

[0010] Perform speech activity detection according to the third intermediate feature vector.

[0011] In some embodiments, the convolutional neural network includes at least two convolutional units and a long short-term memory network layer. Each convolutional unit includes a convolutional layer, a batch normalization layer, an activation function layer, and a max pooling layer. Among them, the sizes of the convolutional kernels in the convolutional layers of each convolutional unit are different. Inputting the acoustic features into the convolutional neural network to obtain a first intermediate feature vector includes:

[0012] Inputting the acoustic features into the convolutional layers, batch normalization layers, activation function layers, and max pooling layers of each convolutional unit in sequence to obtain a convolutional feature vector; and

[0013] Inputting the convolutional feature vector into the long short-term memory network layer for processing to obtain a first intermediate feature vector with context temporal dependence representation.

[0014] In some embodiments, the graph convolutional neural network includes multiple graph convolutional structure units. Inputting the first intermediate feature vector into the graph convolutional neural network to obtain a second intermediate feature vector includes:

[0015] Inputting the first intermediate feature vector into multiple graph convolutional structure units connected in sequence; and generating a second intermediate feature vector according to the output of the last graph convolutional structure unit; where the calculation formula of each graph convolutional structure unit is:

[0016]

[0017] Among them, represents the output of the l-th layer of graph convolutional structure unit, g represents the graph convolution function, represents the diagonal node degree matrix, is the adjacency matrix, is the weight matrix of the (l - 1)-th layer of graph convolution, represents the input first intermediate feature vector.

[0018] In some embodiments, performing multi-head attention calculation based on rotational position encoding on the second intermediate feature vector through a multi-head attention mechanism network to obtain a third intermediate feature vector, including:

[0019] Adding corresponding position information to each element in the second intermediate feature vector;

[0020] Constructing Q matrix, K matrix, and V matrix for multi-head attention mechanism calculation with different position information expressions respectively according to the second intermediate feature vector with added position information;

[0021] Generating multiple head attention calculation units according to each Q matrix, K matrix, and V matrix respectively; and

[0022] Generate a multi - head attention mechanism network according to each head attention calculation unit and the weight matrix of the learned parameters, and perform multi - head attention calculation according to the multi - head attention mechanism network to obtain a third intermediate feature vector.

[0023] In some embodiments, acquiring audio sampling data and extracting acoustic features from the audio sampling data includes:

[0024] Framing the audio sampling data to obtain each audio frame; and

[0025] Extracting filter bank features corresponding to each audio frame respectively as acoustic features.

[0026] In some embodiments, performing voice activity detection according to the third intermediate feature vector includes:

[0027] Performing classification mapping processing on the third intermediate feature vector based on a fully - connected layer to obtain a classification output; and

[0028] Performing normalization processing on the classification output to obtain the probabilities that each speech frame falls into a voice activity frame or a non - voice activity frame, and performing voice activity detection according to the probabilities.

[0029] In some embodiments, the method further includes:

[0030] Collecting vehicle - borne noise and vehicle - borne audio to generate training samples;

[0031] Extracting sample acoustic features from the training samples;

[0032] Inputting the sample acoustic features into a convolutional neural network to obtain a first sample intermediate feature vector;

[0033] Inputting the first sample intermediate feature vector into a graph convolutional neural network for spatial relationship construction to obtain a second sample intermediate feature vector containing spatial relationship information between multiple feature channels;

[0034] Inputting the second sample intermediate feature vector into a multi - head attention mechanism network, and performing multi - head attention calculation based on rotational position encoding on the second sample intermediate feature vector through the multi - head attention mechanism network to obtain a third sample intermediate feature vector; and

[0035] Optimizing and training the parameters of the convolutional neural network, the graph convolutional neural network, and the multi - head attention mechanism network according to the sample third intermediate feature vector, the training samples, based on the cross - entropy loss function and the stochastic gradient descent optimization algorithm.

[0036] In a second aspect, a voice activity detection device is provided, and the device includes:

[0037] An acoustic feature extraction module, configured to obtain audio sampling data and extract acoustic features from the audio sampling data;

[0038] A convolutional neural network module, configured to input the acoustic features into a convolutional neural network to obtain a first intermediate feature vector;

[0039] A graph convolutional neural network module, configured to input the first intermediate feature vector into a graph convolutional neural network for spatial relationship construction to obtain a second intermediate feature vector containing spatial relationship information between multiple feature channels;

[0040] A multi-head attention mechanism network module, configured to input the second intermediate feature vector into a multi-head attention mechanism network, and perform multi-head attention calculation based on rotational position encoding on the second intermediate feature vector through the multi-head attention mechanism network to obtain a third intermediate feature vector; and

[0041] A voice activity detection module, configured to perform voice activity detection according to the third intermediate feature vector.

[0042] A computer device, including a memory, a processor, and a computer program stored on the memory and executable on the processor, wherein when the processor executes the computer program, the steps of the voice activity detection method according to any embodiment of the first aspect are implemented.

[0043] A computer-readable storage medium, having stored thereon a computer program, wherein when the computer program is executed by a processor, the steps of the voice activity detection method according to any embodiment of the first aspect are implemented.

[0044] In the above voice activity detection method, device, computer device, and storage medium, by combining the first intermediate feature vector output by the convolutional neural network into the graph convolutional neural network for spatial relationship construction, and combining the second intermediate feature vector output by the graph convolutional neural network into the multi-head attention mechanism network for multi-head attention calculation based on rotational position encoding, the correlation between information in different parts and different subspaces of the input can be enhanced, and the dependency relationships between various heads and the dependency relationships between the position information of features can be effectively fused, so as to be able to generate more meaningful feature expressions, enhance the effect of feature expressions, and further improve the accuracy of voice activity detection. Description of the Drawings

[0045] Figure 1 It is an application environment diagram of the voice activity detection method in some embodiments;

[0046] Figure 2 It is a flowchart of the voice activity detection method in some embodiments;

[0047] Figure 3Schematic structural diagram of a voice activity detection model applied in a voice activity detection method in some embodiments;

[0048] Figure 4 Schematic structural diagram of a convolutional neural network in some embodiments;

[0049] Figure 5 Schematic flowchart of steps for performing multi - head attention calculation based on rotational position encoding on a second intermediate feature vector through a multi - head attention mechanism network to obtain a third intermediate feature vector in some embodiments;

[0050] Figure 6 Schematic block diagram of a voice activity detection device in some embodiments;

[0051] Figure 7 Internal structure diagram of a computer device in some embodiments. Detailed implementation manners

[0052] In order to make the objectives, technical solutions and advantages of the present application clearer, the present application will be further described in detail below with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present application and are not used to limit the present application.

[0053] The voice activity detection method provided by the present application can be applied to an application environment as Figure 1 shown. Among them, Figure 1 the vehicle 100 shown can include an in - vehicle terminal 110. The in - vehicle terminal 110 can include at least one memory and at least one processor. A computer program is stored in at least one memory. When the computer program is executed by at least one processor, a vehicle pose calculation method according to an exemplary embodiment of the present application is executed. Here, the in - vehicle terminal 110 does not have to be a single electronic device, but can also be any aggregate of devices or circuits that can execute the above computer program alone or jointly.

[0054] In the in - vehicle terminal 110, the processor can include a central processing unit (CPU), a graphics processing unit (GPU), a programmable logic device, a dedicated processor system, a microcontroller or a microprocessor. By way of example and not limitation, the processor can also include an analog processor, a digital processor, a microprocessor, a multi - core processor, a processor array, a network processor, etc.

[0055] In the vehicle-mounted terminal 110, a processor can run a computer program stored in a memory. The computer program can be divided into one or more modules / units (such as computer program 1, computer program 2, ……). One or more modules / units are stored in the memory and executed by the processor to complete the method of the embodiments of the present application. One or more modules / units can be a series of computer program instructions capable of completing specific functions, and the instructions can describe the execution process of the computer program in the vehicle-mounted terminal device. For example, the detail compensation model in the embodiments of the present application can be one of the modules / units.

[0056] The memory can be integrated with the processor. For example, RAM or flash memory can be arranged inside an integrated circuit microprocessor, etc. In addition, the memory can include independent devices, such as an external disk drive, a storage array, or other storage devices that can be used by any database system. The memory and the processor can be operatively coupled, or can communicate with each other, for example, through an I / O port, a network connection, etc., so that the processor can read files stored in the memory.

[0057] In addition, the vehicle-mounted terminal 110 can further include a display device (such as a liquid crystal display) and a user interaction interface (such as a keyboard, a mouse, a touch input device, etc.). All components of the vehicle-mounted terminal 110 can be connected to each other via a bus and / or a network.

[0058] Specifically, the vehicle-mounted terminal 110 acquires audio sampling data, extracts acoustic features from the audio sampling data; inputs the acoustic features into a convolutional neural network to obtain a first intermediate feature vector; inputs the first intermediate feature vector into a graph convolutional neural network for spatial relationship construction to obtain a second intermediate feature vector including spatial relationship information between multiple feature channels; inputs the second intermediate feature vector into a multi-head attention mechanism network, and performs multi-head attention calculation based on rotational position encoding on the second intermediate feature vector through the multi-head attention mechanism network to obtain a third intermediate feature vector; and performs voice activity detection according to the third intermediate feature vector.

[0059] In some other embodiments, the voice activity detection method provided by the present application can also be used in other application scenarios. For example, it can be applied to a computer device. It should be noted that the execution subject can be a configuration device for virtual network card resources, and this device can be implemented as part or all of the computer device through software, hardware, or a combination of software and hardware. Among them, the computer device can be a terminal, a client, or a server. The server can be a single server or a server cluster composed of multiple servers. The terminal can be other intelligent hardware devices such as a vehicle-mounted terminal, a smart audio device, a smart phone, a personal computer, a tablet computer, a wearable device, and a smart robot, etc.

[0060] In some embodiments, such as Figure 2 shown, a voice activity detection method is provided. Taking the vehicle terminal in Figure 1 as an example, illustratively, the voice activity detection method can be implemented based on the Figure 3 shown voice activity detection model. Among them, the voice detection model can at least include a convolutional neural network 10, a graph convolutional neural network 20, and a multi-head attention mechanism network 30. Specifically, it can include the following steps:

[0061] Step S202: Obtain audio sampling data and extract acoustic features from the audio sampling data.

[0062] Specifically, voice data can be collected by an audio sampling device and audio sampling can be performed to obtain audio sampling data. In the application scenario of a vehicle terminal, voice commands issued by the vehicle user can be collected, or the interactive voice during the conversation between the vehicle user and the vehicle terminal can be collected, and the collected voice commands or interactive voice can be used as audio sampling data. Then, feature extraction can be performed on the audio sampling data. Among them, the acoustic feature is a feature representing the acoustic information of the audio sampling, such as it can include energy features, time-domain features, or frequency-domain features, etc. Further, for example, the time-domain feature can include FBANK (Filter Bank) features.

[0063] Step S204: Input the acoustic features into the convolutional neural network to obtain a first intermediate feature vector.

[0064] Specifically, the acoustic features are input into the convolutional neural network for convolution operations. Among them, the convolutional neural network is used to perform convolution operations on the input acoustic features and perform dimension integration processing on the features after convolution operations, etc., to output an intermediate feature vector that has undergone convolution operations and whose vector dimension is adapted to the next unit as the first intermediate feature vector.

[0065] Step S206: Input the first intermediate feature vector into the graph convolutional neural network to construct spatial relationships and obtain a second intermediate feature vector containing spatial relationship information between multiple feature channels.

[0066] Specifically, the first intermediate feature vector output by the convolutional neural network can be input into the graph convolutional neural network, and the graph convolutional neural network performs spatial relationship modeling processing on the first intermediate feature vector to obtain an intermediate feature vector containing spatial relationship information as the second intermediate feature vector for output.

[0067] Step S208: Input the second intermediate feature vector into the multi-head attention mechanism network, and perform multi-head attention calculation based on rotational position encoding on the second intermediate feature vector through the multi-head attention mechanism network to obtain a third intermediate feature vector.

[0068] Specifically, on the basis of spatial correlation construction, the position information of the features is further combined through rotational position encoding, and multi-head attention is used to calculate the weight distribution for enhancing feature representation. Through the multi-head attention calculation based on rotational position encoding, important information can be screened and emphasized from the feature representations of different parts. The multi-head attention mechanism models the information in different subspaces, and combines with the spatial correlation between different feature channels modeled by the graph convolutional neural network, which can improve and enhance the feature representation of the spatial relationship between channels, so as to obtain the third intermediate feature vector after enhanced feature representation.

[0069] Step S210: Perform voice activity detection according to the third intermediate feature vector.

[0070] Specifically, the third intermediate feature vector can be used as the input vector to perform deep learning-based voice activity detection. For example, the third intermediate feature vector can be input into a traditional voice activity detection model to judge the voice segments and / or non-voice segments in the audio sampling data; or a fully connected layer can be added after the multi-head attention mechanism network, and the third intermediate feature vector is subjected to classification mapping processing through the fully connected layer, so as to judge the voice segments and / or non-voice segments in the audio sampling data, and then voice activity detection is realized.

[0071] The above voice activity detection method can enhance the correlation between the information of different parts and different subspaces of the input by combining the first intermediate feature vector output by the convolutional neural network into the graph convolutional neural network for spatial relationship construction, and combining the second intermediate feature vector output by the graph convolutional neural network into the multi-head attention mechanism network for multi-head attention calculation based on rotational position encoding, and effectively fuse the dependencies between each head and the dependencies between the position information of the features, so as to be able to generate more meaningful feature representations, enhance the effect of feature representation, and further improve the accuracy of voice activity detection.

[0072] In some embodiments, refer to Figure 4 , Figure 4 shows a schematic structural diagram of a convolutional neural network in some embodiments. The convolutional neural network includes at least two convolutional units and a long short-term memory network layer. Each convolutional unit respectively includes a convolutional layer, a batch normalization layer, an activation function layer, and a max pooling layer. Among them, the sizes of the convolutional kernels of the convolutional layers in each convolutional unit are different. Specifically, inputting the acoustic features into the convolutional neural network to obtain the first intermediate feature vector may include the following steps:

[0073] Input the acoustic features into each convolutional layer, batch normalization layer, activation function layer, and max pooling layer of each convolutional unit in sequence to obtain a convolutional feature vector; and

[0074] The convolutional feature vector is input into a long short-term memory network layer for processing to obtain a first intermediate feature vector with a context temporal dependence representation.

[0075] In this embodiment, by adding at least one long short-term memory (LSTM) layer at the end of multiple convolutional units, the temporal dependence relationship between features can be enhanced, so as to obtain a first intermediate feature vector with a context temporal dependence representation, providing a basis for subsequent processing of the model.

[0076] In some embodiments, the graph convolutional neural network includes multiple graph convolutional layers, and the graph convolutional neural network includes multiple graph convolutional structure units. Inputting the first intermediate feature vector into the graph convolutional neural network to obtain a second intermediate feature vector includes: inputting the first intermediate feature vector into multiple sequentially connected graph convolutional structure units; and generating a second intermediate feature vector according to the output of the last graph convolutional structure unit; wherein, the calculation formula of each graph convolutional structure unit is:

[0077]

[0078] Wherein, represents the output of the l-th layer of graph convolutional structure unit, g represents the graph convolution function, represents the diagonal node degree matrix, is the adjacency matrix, is the weight matrix of the l-1-th layer of graph convolution, represents the input first intermediate feature vector.

[0079] Specifically, the diagonal node degree matrix adjacency matrix weight matrix can be constructed based on the graph convolution operation principle on the graph convolutional structure unit. Among them, the parameters of the adjacency matrix and the weight matrix are learnable parameters. The graph convolutional neural network can be trained in advance through sample data to obtain the adjacency matrix and the weight matrix after parameter learning. Using the adjacency matrix

[0080] and the weight matrix after parameter learning to construct the spatial relationship of the input first intermediate feature vector, and taking the output of the last graph convolutional layer as the second intermediate feature vector. Exemplarily, each graph convolutional structure unit may specifically include a graph convolutional layer, a dropout layer, and a ReLU layer.

[0080] In the above embodiments, by constructing a plurality of graph convolution structure units and performing multi-layer graph convolution operations on the first intermediate feature vector through the above formula to construct spatial relationships, the second intermediate feature vector containing spatial relationship information between multiple feature channels can be obtained more accurately.

[0081] In some embodiments, referring to Figure 5 as shown, Figure 5 FIG. shows a schematic flowchart of steps for obtaining a third intermediate feature vector by performing multi-head attention calculation based on rotational position encoding on the second intermediate feature vector through a multi-head attention mechanism network in some embodiments, which may specifically include:

[0082] Step S502: Add corresponding position information to each element in the second intermediate feature vector.

[0083] Specifically, the second intermediate feature vector output by the graph convolutional neural network is The multi-head attention mechanism based on rotational position encoding can first incorporate the position information into x t .

[0084] Step S504: Construct Q matrix, K matrix, and V matrix for multi-head attention mechanism calculation with different position information expressions respectively according to the second intermediate feature vector with added position information.

[0085] Specifically, after incorporating the position information, x t is converted into queries, keys, and values, and their expressions can be referred to as follows:

[0086]

[0087]

[0088]

[0089] where m and n are position information, and W q , W k , W v are learnable weight matrices, which can be expressed as:

[0090]

[0091] where Θ = {θ i = 10000 -2(i-1) / d , i ∈ [1, 2,..., d / 2]}; d = 160 is the dimension of the embedding vector.

[0092] Step S506: Generate multiple head attention calculation units according to each Q matrix, K matrix, and V matrix.

[0093] Specifically, the head attention calculation units can be generated according to the following formula:

[0094]

[0095] where i represents the index of the head attention calculation unit.

[0096] Step S508: Generate a multi-head attention mechanism network according to each head attention calculation unit and the weight matrix of the learned parameters, and perform multi-head attention calculation according to the multi-head attention mechanism network to obtain a third intermediate feature vector.

[0097] The output of a single-layer multi-head attention mechanism network can be expressed as:

[0098] MultiHead = Concat(head0, … head i …, head H-1 )W O

[0099] where head i represents the i-th head attention calculation unit, H represents the number of head attention calculation units, and W O is a parameter-learnable weight matrix in the multi-head attention mechanism network, which can be pre-trained and learned according to the sample data.

[0100] In some embodiments, obtaining audio sampling data and extracting acoustic features from the audio sampling data includes: framing the audio sampling data to obtain each audio frame; and extracting filter bank features corresponding to each audio frame as acoustic features.

[0101] In this embodiment, by framing the input audio sampling data and separately extracting the FBANK (Filter Bank) features of each frame, the efficiency and quality of feature extraction can be improved by first framing and then extracting acoustic features.

[0102] In some embodiments, performing voice activity detection according to the third intermediate feature vector includes: performing classification mapping processing on the third intermediate feature vector based on a fully connected layer to obtain a classification output; and performing normalization processing on the classification output to obtain the probability that each speech frame falls into a speech activity frame or a non-speech activity frame, and performing voice activity detection according to the probability.

[0103] In this embodiment, the third intermediate feature vector processed by the multi-head attention mechanism based on rotational position encoding can be input into a layer normalization layer for processing, and then input into a fully connected layer for fully connected layer mapping to obtain a classification output, and the classification output is normalized (Softmax processing), so as to obtain the probabilities of each speech frame falling into a speech activity frame or a non-speech activity frame. By comparing the probabilities of each speech frame falling into a speech activity frame or a non-speech activity frame with a preset probability threshold, the result of speech activity detection is obtained.

[0104] In some embodiments, the convolutional neural network, graph convolutional neural network, and multi-head attention mechanism network involved in the speech activity detection method of the embodiments of the present application can be pre-trained, so that the learnable parameters in each network layer are trained and optimized.

[0105] Exemplarily, in the application scenario of an in-vehicle terminal, the speech activity detection method may further include a process of training the model, which may specifically include the following steps:

[0106] 1. Collect in-vehicle noise and in-vehicle audio to generate training samples. Specifically, the collected in-vehicle noise and in-vehicle speech can be used to construct an audio data set for model training, and the data in the audio data set is used as training samples. Among them, the in-vehicle noise can include five situations: the noise with the window open on the road, the noise A with the window closed on the road, the noise B with the window closed on the road, the noise with the window open in the parking lot, and the noise with the window closed in the parking lot. The noise A with the window closed on the road and the noise B with the window closed on the road are noises from different recording sections.

[0107] 2. Extract sample acoustic features from the training samples.

[0108] 3. Input the sample acoustic features into a convolutional neural network to obtain a first sample intermediate feature vector.

[0109] 4. Input the first sample intermediate feature vector into a graph convolutional neural network for spatial relationship construction to obtain a second sample intermediate feature vector containing spatial relationship information between multiple feature channels.

[0110] 5. Input the second sample intermediate feature vector into a multi-head attention mechanism network, and perform multi-head attention calculation based on rotational position encoding on the second sample intermediate feature vector through the multi-head attention mechanism network to obtain a third sample intermediate feature vector.

[0111] 6. Optimize and train the parameters of the convolutional neural network, graph convolutional neural network, and multi-head attention mechanism network according to the sample third intermediate feature vector, training samples, and based on the cross-entropy loss function and the stochastic gradient descent optimization algorithm.

[0112] It should be understood that although Figure 2 and Figure 5 each step in the flowcharts is shown in sequence according to the indication of the arrows, these steps are not necessarily executed in sequence according to the order indicated by the arrows. Unless otherwise clearly stated in this document, there is no strict order restriction for the execution of these steps, and these steps can be executed in other orders. Moreover, Figure 2 and Figure 5 at least a part of the steps in

[0113] In some embodiments, as Figure 6 shown, a voice activity detection device is provided, including: an acoustic feature extraction module 602, a convolutional neural network module 604, a graph convolutional neural network module 606, a multi-head attention mechanism network module 608, and a voice activity detection module 610, where:

[0114] The acoustic feature extraction module 602 is configured to obtain audio sampling data and extract acoustic features from the audio sampling data;

[0115] The convolutional neural network module 604 is configured to input the acoustic features into a convolutional neural network to obtain a first intermediate feature vector;

[0116] The graph convolutional neural network module 606 is configured to input the first intermediate feature vector into a graph convolutional neural network for spatial relationship construction to obtain a second intermediate feature vector including spatial relationship information between multiple feature channels;

[0117] The multi-head attention mechanism network module 608 is configured to input the second intermediate feature vector into a multi-head attention mechanism network, and perform multi-head attention calculation based on rotational position encoding on the second intermediate feature vector through the multi-head attention mechanism network to obtain a third intermediate feature vector; and

[0118] The voice activity detection module 610 is configured to perform voice activity detection according to the third intermediate feature vector.

[0119] In some embodiments, the convolutional neural network module 604 sequentially inputs the acoustic features into each convolutional layer, batch normalization layer, activation function layer, and max pooling layer of each convolutional unit to obtain a convolutional feature vector; and inputs the convolutional feature vector into a long short-term memory network layer for processing to obtain a first intermediate feature vector with context temporal dependency representation.

[0120] In some embodiments, the graph convolutional neural network module 606 inputs the first intermediate feature vector into a plurality of graph convolutional structure units connected in sequence; and generates a second intermediate feature vector according to the output of the last graph convolutional structure unit; wherein, the calculation formula of each graph convolutional structure unit is:

[0121]

[0122] Wherein, represents the output of the l-th layer of graph convolutional structure unit, g represents the graph convolutional function, represents the diagonal node degree matrix, is the adjacency matrix, is the weight matrix of the l-1-th layer of graph convolution, represents the input first intermediate feature vector.

[0123] In some embodiments, the multi-head attention mechanism network module 608 adds corresponding position information to each element in the second intermediate feature vector; constructs a Q matrix, a K matrix, and a V matrix for multi-head attention mechanism calculation with different position information expressions according to the second intermediate feature vector with added position information; generates a plurality of head attention calculation units according to each Q matrix, K matrix, and V matrix; and generates a multi-head attention mechanism network according to each head attention calculation unit and the weight matrix of the learned parameters, and performs multi-head attention calculation according to the multi-head attention mechanism network to obtain a third intermediate feature vector.

[0124] In some embodiments, the acoustic feature extraction module 602 frames the audio sampling data to obtain each audio frame; and extracts the filter bank features corresponding to each audio frame as acoustic features.

[0125] In some embodiments, the voice activity detection module 610 performs classification mapping processing on the third intermediate feature vector based on a fully connected layer to obtain a classification output; and performs normalization processing on the classification output to obtain the probability that each speech frame falls into a voice activity frame or a non-voice activity frame, and performs voice activity detection according to the probability.

[0126] In some embodiments, it further includes a model training module, which is used to collect in-vehicle noise and in-vehicle audio to generate training samples; extract sample acoustic features from the training samples; input the sample acoustic features into a convolutional neural network to obtain a first sample intermediate feature vector; input the first sample intermediate feature vector into a graph convolutional neural network for spatial relationship construction to obtain a second sample intermediate feature vector containing spatial relationship information between multiple feature channels; input the second sample intermediate feature vector into a multi-head attention mechanism network, and perform multi-head attention calculation based on rotational position encoding on the second sample intermediate feature vector through the multi-head attention mechanism network to obtain a third sample intermediate feature vector; and optimize and train the parameters of the convolutional neural network, the graph convolutional neural network, and the multi-head attention mechanism network according to the sample third intermediate feature vector, the training samples, and based on the cross-entropy loss function and the stochastic gradient descent optimization algorithm.

[0127] For the specific limitations of the voice activity detection device, reference can be made to the limitations on the voice activity detection method in the above text, which will not be elaborated here. Each module in the above voice activity detection device can be implemented in whole or in part by software, hardware, and their combination. The above modules can be embedded in the processor in the computer device in hardware form or be independent of it, or stored in the memory in the computer device in software form, so as to facilitate the processor to call and execute the operations corresponding to the above modules.

[0128] In some embodiments, a computer device is provided. The computer device can be a terminal, and its internal structure diagram can be as Figure 7 shown. The computer device includes a processor, a memory, a network interface, a display screen, and an input device connected through a system bus. Among them, the processor of the computer device is used to provide computing and control capabilities. The memory of the computer device includes a non-volatile storage medium and an internal memory. The non-volatile storage medium stores an operating system and a computer program. The internal memory provides an environment for the operation of the operating system and the computer program in the non-volatile storage medium. The network interface of the computer device is used to communicate with an external terminal through a network connection. When the computer program is executed by the processor, it implements a voice activity detection method. The display screen of the computer device can be a liquid crystal display screen or an electronic ink display screen. The input device of the computer device can be a touch layer covering the display screen, or a button, a trackball, or a touchpad provided on the shell of the computer device, or an external keyboard, touchpad, or mouse, etc.

[0129] Those skilled in the art can understand that Figure 7The structure shown is only a block diagram of some structures related to the solution of this application, and does not constitute a limitation on the computer device to which the solution of this application is applied. The specific computer device may include more or fewer components than those shown in the figure, or combine some components, or have a different component layout.

[0130] In some embodiments, a computer device is provided, including a memory, a processor, and a computer program stored on the memory and executable on the processor. When the processor executes the computer program, the following steps are implemented: obtaining audio sampling data, and extracting acoustic features from the audio sampling data; inputting the acoustic features into a convolutional neural network to obtain a first intermediate feature vector; inputting the first intermediate feature vector into a graph convolutional neural network for spatial relationship construction to obtain a second intermediate feature vector including spatial relationship information between multiple feature channels; inputting the second intermediate feature vector into a multi-head attention mechanism network, and performing multi-head attention calculation based on rotational position encoding on the second intermediate feature vector through the multi-head attention mechanism network to obtain a third intermediate feature vector; and performing voice activity detection according to the third intermediate feature vector.

[0131] In some embodiments, when the processor executes the computer program, the following steps are further implemented: sequentially inputting the acoustic features into each convolutional layer, batch normalization layer, activation function layer, and max pooling layer of each convolutional unit to obtain a convolutional feature vector; and inputting the convolutional feature vector into a long short-term memory network layer for processing to obtain a first intermediate feature vector with a context temporal dependence representation.

[0132] In some embodiments, when the processor executes the computer program, the following steps are further implemented: inputting the first intermediate feature vector into a plurality of graph convolutional structure units connected in sequence; and generating a second intermediate feature vector according to the output of the last graph convolutional structure unit; wherein, the calculation formula of each graph convolutional structure unit is:

[0133]

[0134] Wherein, represents the output of the l-th layer of the graph convolutional structure unit, g represents the graph convolutional function, represents the diagonal node degree matrix, is the adjacency matrix, is the weight matrix of the graph convolution of the (l - 1)-th layer, represents the input first intermediate feature vector.

[0135] In some embodiments, when the processor executes the computer program, the following steps are further implemented: adding corresponding position information to each element in the second intermediate feature vector; respectively constructing a Q matrix, a K matrix, and a V matrix with different position information expressions for multi-head attention mechanism calculation based on the second intermediate feature vector with added position information; generating a plurality of head attention calculation units according to each Q matrix, K matrix, and V matrix; and generating a multi-head attention mechanism network according to each head attention calculation unit and the weight matrix of the learned parameters, and performing multi-head attention calculation according to the multi-head attention mechanism network to obtain a third intermediate feature vector.

[0136] In some embodiments, when the processor executes the computer program, the following steps are further implemented: framing the audio sampling data to obtain each audio frame; and extracting filter bank features corresponding to each audio frame as acoustic features.

[0137] In some embodiments, when the processor executes the computer program, the following steps are further implemented: performing classification mapping processing on the third intermediate feature vector based on a fully connected layer to obtain a classification output; and performing normalization processing on the classification output to obtain the probability that each speech frame falls into a speech active frame or a non-speech active frame, and performing speech activity detection according to the probability.

[0138] In some embodiments, when the processor executes the computer program, the following steps are further implemented: collecting in-vehicle noise and in-vehicle audio to generate training samples; extracting sample acoustic features from the training samples; inputting the sample acoustic features into a convolutional neural network to obtain a first sample intermediate feature vector; inputting the first sample intermediate feature vector into a graph convolutional neural network for spatial relationship construction to obtain a second sample intermediate feature vector containing spatial relationship information between multiple feature channels; inputting the second sample intermediate feature vector into a multi-head attention mechanism network, and performing multi-head attention calculation based on rotational position encoding on the second sample intermediate feature vector through the multi-head attention mechanism network to obtain a third sample intermediate feature vector; and optimizing and training the parameters of the convolutional neural network, the graph convolutional neural network, and the multi-head attention mechanism network according to the sample third intermediate feature vector, the training samples, based on the cross-entropy loss function and the stochastic gradient descent optimization algorithm.

[0139] In some embodiments, a computer-readable storage medium is provided, on which a computer program is stored. When the computer program is executed by a processor, the following steps are implemented: obtaining audio sampling data, and extracting acoustic features from the audio sampling data; inputting the acoustic features into a convolutional neural network to obtain a first intermediate feature vector; inputting the first intermediate feature vector into a graph convolutional neural network for spatial relationship construction to obtain a second intermediate feature vector including spatial relationship information between multiple feature channels; inputting the second intermediate feature vector into a multi-head attention mechanism network, and performing multi-head attention calculation based on rotational position encoding on the second intermediate feature vector through the multi-head attention mechanism network to obtain a third intermediate feature vector; and performing voice activity detection according to the third intermediate feature vector.

[0140] In some embodiments, when the computer program is executed by a processor, the following steps are further implemented: sequentially inputting the acoustic features into each convolutional layer, batch normalization layer, activation function layer, and max pooling layer of each convolutional unit to obtain a convolutional feature vector; and inputting the convolutional feature vector into a long short-term memory network layer for processing to obtain a first intermediate feature vector with a context time series dependence representation.

[0141] In some embodiments, when the computer program is executed by a processor, the following steps are further implemented: inputting the first intermediate feature vector into a plurality of graph convolutional structure units connected in sequence; and generating a second intermediate feature vector according to the output of the last graph convolutional structure unit; wherein, the calculation formula of each graph convolutional structure unit is:

[0142]

[0143] Wherein, represents the output of the l-th layer of graph convolutional structure unit, g represents the graph convolution function, represents the diagonal node degree matrix, is the adjacency matrix, is the weight matrix of the l-1-th layer of graph convolution, represents the input first intermediate feature vector.

[0144] In some embodiments, when the computer program is executed by a processor, the following steps are further implemented: adding corresponding position information to each element in the second intermediate feature vector; respectively constructing a Q matrix, a K matrix, and a V matrix for multi-head attention mechanism calculation with different position information expressions according to the second intermediate feature vector with added position information; respectively generating a plurality of head attention calculation units according to each Q matrix, K matrix, and V matrix; and generating a multi-head attention mechanism network according to each head attention calculation unit and the weight matrix of the learned parameters, and performing multi-head attention calculation according to the multi-head attention mechanism network to obtain a third intermediate feature vector.

[0145] In some embodiments, when the computer program is executed by a processor, the following steps are further implemented: frame the audio sampling data to obtain each audio frame; and extract the filter bank features respectively corresponding to each audio frame as acoustic features.

[0146] In some embodiments, when the computer program is executed by a processor, the following steps are further implemented: perform classification mapping processing on the third intermediate feature vector based on a fully connected layer to obtain a classification output; and perform normalization processing on the classification output to obtain the probabilities that each speech frame falls into a speech active frame or a non-speech active frame, and perform speech activity detection according to the probabilities.

[0147] In some embodiments, when the computer program is executed by a processor, the following steps are further implemented: collect in-vehicle noise and in-vehicle audio to generate training samples; extract sample acoustic features from the training samples; input the sample acoustic features into a convolutional neural network to obtain a first sample intermediate feature vector; input the first sample intermediate feature vector into a graph convolutional neural network for spatial relationship construction to obtain a second sample intermediate feature vector containing spatial relationship information between multiple feature channels; input the second sample intermediate feature vector into a multi-head attention mechanism network, and perform multi-head attention calculation based on rotational position encoding on the second sample intermediate feature vector through the multi-head attention mechanism network to obtain a third sample intermediate feature vector; and optimize and train the parameters of the convolutional neural network, the graph convolutional neural network, and the multi-head attention mechanism network according to the sample third intermediate feature vector, the training samples, and based on a cross-entropy loss function and a stochastic gradient descent optimization algorithm.

[0148] Those of ordinary skill in the art can understand that all or part of the processes in the methods of the above embodiments can be completed by instructing relevant hardware through a computer program. The computer program can be stored in a non-volatile computer-readable storage medium. When the computer program is executed, it can include the processes of the embodiments of the above methods. Among them, any reference to a memory, storage, database, or other medium used in the embodiments provided in the present application can include non-volatile and / or volatile memories. Non-volatile memory can include read-only memory (ROM), programmable ROM (PROM), electrically programmable ROM (EPROM), electrically erasable programmable ROM (EEPROM), or flash memory. Volatile memory can include random access memory (RAM) or external cache memory. By way of illustration and not limitation, RAM is available in many forms, such as static RAM (SRAM), dynamic RAM (DRAM), synchronous DRAM (SDRAM), double data rate SDRAM (DDR SDRAM), enhanced SDRAM (ESDRAM), synchronous link (Synchlink), DRAM (SLDRAM), Rambus direct RAM (RDRAM), direct memory bus dynamic RAM (DRDRAM), and Rambus dynamic RAM (RDRAM), etc.

[0149] The technical features of the above embodiments can be combined arbitrarily. For the sake of brevity of description, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, it should be considered as the scope described in this specification.

[0150] The above-described embodiments merely represent several implementation manners of the present application. Their descriptions are relatively specific and detailed, but they should not be construed as limiting the scope of the invention patent. It should be noted that for those of ordinary skill in the art, without departing from the concept of the present application, several modifications and improvements can still be made, and these all belong to the protection scope of the present application. Therefore, the protection scope of the patent of the present application should be subject to the appended claims.

Claims

1. A voice activity detection method, the method comprising: Obtaining audio sampling data, and extracting acoustic features from the audio sampling data; Inputting the acoustic features into a convolutional neural network to obtain a first intermediate feature vector; Inputting the first intermediate feature vector into a graph convolutional neural network for spatial relationship construction to obtain a second intermediate feature vector containing spatial relationship information between multiple feature channels; Inputting the second intermediate feature vector into a multi-head attention mechanism network, and performing multi-head attention calculation based on rotational position encoding on the second intermediate feature vector through the multi-head attention mechanism network to obtain a third intermediate feature vector; and Performing voice activity detection according to the third intermediate feature vector.

2. The method according to claim 1, wherein The convolutional neural network includes at least two convolutional units and a long short-term memory network layer. Each of the convolutional units includes a convolutional layer, a batch normalization layer, an activation function layer, and a max pooling layer. Among them, the sizes of the convolutional kernels of the convolutional layers in each of the convolutional units are different. The inputting the acoustic features into the convolutional neural network to obtain a first intermediate feature vector includes: Sequentially inputting the acoustic features into the convolutional layers, the batch normalization layers, the activation function layers, and the max pooling layers of each of the convolutional units to obtain a convolutional feature vector; and Inputting the convolutional feature vector into the long short-term memory network layer for processing to obtain the first intermediate feature vector with a context time series dependence representation.

3. The method according to claim 1, wherein The graph convolutional neural network includes a plurality of graph convolutional structure units. The inputting the first intermediate feature vector into the graph convolutional neural network to obtain a second intermediate feature vector includes: Inputting the first intermediate feature vector into a plurality of the graph convolutional structure units connected in sequence; and generating the second intermediate feature vector according to the output of the last graph convolutional structure unit; wherein, the calculation formula of each of the graph convolutional structure units is: Among them, represents the output of the graph convolutional structure unit of the l-th layer, g represents the graph convolution function, represents the diagonal node degree matrix, is the adjacency matrix, is the weight matrix of the graph convolution of the (l-1)-th layer, represents the input first intermediate feature vector.

4. The method according to claim 1, wherein The performing multi-head attention calculation based on rotational position encoding on the second intermediate feature vector through the multi-head attention mechanism network to obtain a third intermediate feature vector includes: Adding corresponding position information to each element in the second intermediate feature vector; Respectively constructing a Q matrix, a K matrix, and a V matrix for multi-head attention mechanism calculation with different position information expressions according to the second intermediate feature vector with added position information; Generating a plurality of head attention calculation units according to each of the Q matrix, the K matrix, and the V matrix; and Generating a multi-head attention mechanism network according to each of the head attention calculation units and a weight matrix of learned parameters, and performing multi-head attention calculation according to the multi-head attention mechanism network to obtain the third intermediate feature vector.

5. The method according to claim 1, characterized in that, The obtaining audio sampling data and extracting acoustic features from the audio sampling data includes: Framing the audio sampling data to obtain each audio frame; and Extracting filter bank features corresponding to each of the audio frames as the acoustic features.

6. The method according to claim 5, wherein The performing voice activity detection according to the third intermediate feature vector includes: Performing classification mapping processing on the third intermediate feature vector based on the fully connected layer to obtain a classification output; and Performing normalization processing on the classification output to obtain the probabilities that each of the speech frames fall into speech active frames or non-speech active frames, and performing voice activity detection according to the probabilities.

7. The method according to any one of claims 1 to 6, characterized in that The method further includes: Collecting in-vehicle noise and in-vehicle audio to generate training samples; Extracting sample acoustic features from the training samples; Inputting the sample acoustic features into the convolutional neural network to obtain a first sample intermediate feature vector; Inputting the first sample intermediate feature vector into a graph convolutional neural network for spatial relationship construction to obtain a second sample intermediate feature vector including spatial relationship information between multiple feature channels; Inputting the second sample intermediate feature vector into a multi-head attention mechanism network, and performing multi-head attention calculation based on rotational position encoding on the second sample intermediate feature vector through the multi-head attention mechanism network to obtain a third sample intermediate feature vector; and Optimally training the parameters of the convolutional neural network, the graph convolutional neural network, and the multi-head attention mechanism network according to the sample third intermediate feature vector, the training samples, and based on the cross-entropy loss function and the stochastic gradient descent optimization algorithm.

8. A voice activity detection device, characterized in that, The device includes: An acoustic feature extraction module, configured to obtain audio sampling data and extract acoustic features from the audio sampling data; A convolutional neural network module, configured to input the acoustic features into a convolutional neural network to obtain a first intermediate feature vector; A graph convolutional neural network module, configured to input the first intermediate feature vector into a graph convolutional neural network for spatial relationship construction to obtain a second intermediate feature vector including spatial relationship information between multiple feature channels; A multi-head attention mechanism network module, configured to input the second intermediate feature vector into a multi-head attention mechanism network, and performing multi-head attention calculation based on rotational position encoding on the second intermediate feature vector through the multi-head attention mechanism network to obtain a third intermediate feature vector; and A voice activity detection module, configured to perform voice activity detection according to the third intermediate feature vector.

9. A computer device, comprising a memory, a processor, and a computer program stored on the memory and executable on the processor, characterized in that, When the processor executes the computer program, the steps of the method according to any one of claims 1 to 7 are implemented.

10. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by the processor, the steps of the method according to any one of claims 1 to 7 are implemented.

Citation Information

Cited By

  • Voice activity signal detection method and voice activity signal detection system

    CN120452428A

  • Live broadcast book-speaking noise processing method, system and device, and medium

    CN120636436A