Voice activity detection method, electronic device, and storage medium
Patent Information
- Application Number
- CN202510632713.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-15
- Publication Date
- 2026-09-11
- Estimated Expiration
- 2045-05-15
AI Technical Summary
然而,基于能量或谱特征的语音活动检测方法在复杂声学环境中的鲁棒性较差,易误检或漏检,语音活动检测的准确性和稳定性较低
[0012]本发明实施例提供一种语音活动检测方法、电子设备及存储介质,本发明实施例通过第一卷积神经模块对利用多通道音频数据构建的初始邻接矩阵进行优化处理,可以得到用于描述声音拓扑关系的目标邻接矩阵,并且通过第二卷积神经模块对多通道音频数据进行语音特征提取处理,可以得到目标语音特征图,然后通过图卷积神经网络对用于描述声音拓扑关系的目标邻接矩阵和目标语音特征图进行图卷积处理,从而融合得到更加全面且准确的目标特征向量,这样通过分类网络对目标特征向量进行分类处理,能够准确地得到多通道音频数据的语音活动检测信息,使得无论是在安静环境、静止场景等简单声学环境下,还是在嘈杂环境、运动场景等复杂声学环境下,均可以稳定准确地进行语音活动检测,提高了语音活动检测的准确性和稳定性。
Smart Images

Figure CN120564767B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of speech processing technology, and in particular to a speech activity detection method, electronic device, and storage medium. Background Technology
[0002] Voice Activity Detection (VAD) is a fundamental and critical task in the field of speech signal processing, aiming to accurately identify speech segments from a continuous audio stream. VAD plays a pre-processor role in Automatic Speech Recognition (ASR), Keyword-spotting (KWS), speech enhancement, and call noise reduction. The detection accuracy and response speed of VAD directly affect the performance of subsequent core tasks such as ASR, KWS, speech enhancement, and call noise reduction. Current VAD methods mainly include those based on energy or spectral features. However, energy or spectral feature-based VAD methods have poor robustness in complex acoustic environments, are prone to false positives or false negatives, and exhibit low accuracy and stability. Summary of the Invention
[0003] This invention provides a method, electronic device, and storage medium for detecting voice activity, aiming to improve the accuracy and stability of voice activity detection.
[0004] In a first aspect, embodiments of the present invention provide a speech activity detection method applied to an electronic device, the electronic device including a microphone array and a speech activity detection model stored matching the microphone array, the speech activity detection model including a first convolutional neural module, a second convolutional neural module, a graph convolutional neural network, and a classification network, the method including:
[0005] Acquire multi-channel audio data collected by the microphone array;
[0006] An initial adjacency matrix is constructed based on the multi-channel audio data, and the initial adjacency matrix is optimized by the first convolutional neural module to obtain the target adjacency matrix.
[0007] The second convolutional neural module performs speech feature extraction processing on the multi-channel audio data to obtain a target speech feature map.
[0008] The target feature vector is obtained by performing graph convolution processing on the target adjacency matrix and the target speech feature map using the graph convolutional neural network.
[0009] The target feature vector is classified using the classification network to obtain speech activity detection information from the multi-channel audio data.
[0010] In a second aspect, embodiments of the present invention also provide an electronic device, the electronic device including a processor, a memory, a computer program stored in the memory and executable by the processor, and a data bus for implementing communication between the processor and the memory, wherein when the computer program is executed by the processor, it implements the voice activity detection method as described in the first aspect.
[0011] Thirdly, embodiments of the present invention also provide a storage medium for computer-readable storage, wherein the storage medium stores one or more programs that can be executed by one or more processors to implement the voice activity detection method as described in the first aspect.
[0012] This invention provides a speech activity detection method, electronic device, and storage medium. The invention optimizes an initial adjacency matrix constructed from multi-channel audio data using a first convolutional neural module to obtain a target adjacency matrix describing sound topology. A second convolutional neural module extracts speech features from the multi-channel audio data to obtain a target speech feature map. A graph convolutional neural network then performs graph convolution on the target adjacency matrix and the target speech feature map, fusing them to obtain a more comprehensive and accurate target feature vector. By classifying the target feature vector using a classification network, speech activity detection information from multi-channel audio data can be accurately obtained. This allows for stable and accurate speech activity detection in both simple acoustic environments (quiet, static, etc.) and complex acoustic environments (noisy, moving, etc.), improving the accuracy and stability of speech activity detection. Attached Figure Description
[0013] To more clearly illustrate the technical solutions of the embodiments of this application, the drawings used in the description of the embodiments will be briefly introduced below. Obviously, the drawings described below are some embodiments of this application. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0014] Figure 1 This is a flowchart illustrating a voice activity detection method provided in an embodiment of the present invention;
[0015] Figure 2 This is a schematic diagram of the structure of the speech activity detection model in an embodiment of the present invention;
[0016] Figure 3This is another structural schematic diagram of the speech activity detection model in this embodiment of the invention;
[0017] Figure 4 This is another structural schematic diagram of the speech activity detection model in this embodiment of the invention;
[0018] Figure 5 This is another structural schematic diagram of the speech activity detection model in this embodiment of the invention;
[0019] Figure 6 This is a schematic block diagram of the structure of an electronic device provided in an embodiment of the present invention. Detailed Implementation
[0020] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some, not all, of the embodiments of the present invention. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.
[0021] The flowchart shown in the attached diagram is for illustrative purposes only and does not necessarily include all content and operations / steps, nor does it necessarily have to be performed in the order described. For example, some operations / steps can be broken down, combined, or partially merged, so the actual execution order may change depending on the actual situation.
[0022] It should be understood that the terminology used in this specification is for the purpose of describing particular embodiments only and is not intended to limit the invention. As used in this specification and the appended claims, the singular forms “a,” “an,” and “the” are intended to include the plural forms unless the context clearly indicates otherwise.
[0023] Voice Activity Detection (VAD) is a fundamental and critical task in the field of speech signal processing, aiming to accurately identify speech segments from a continuous audio stream. VAD plays a pre-processor role in Automatic Speech Recognition (ASR), Keyword-spotting (KWS), speech enhancement, and call noise reduction. The detection accuracy and response speed of VAD directly affect the performance of subsequent core tasks such as ASR, KWS, speech enhancement, and call noise reduction. Current VAD methods mainly include those based on energy or spectral features. However, energy or spectral feature-based VAD methods have poor robustness in noisy environments, are prone to false positives or false negatives, and exhibit low accuracy and stability.
[0024] To address the aforementioned issues, this invention provides a speech activity detection method, electronic device, and storage medium. This invention optimizes an initial adjacency matrix constructed from multi-channel audio data using a first convolutional neural module to obtain a target adjacency matrix describing sound topology. A second convolutional neural module extracts speech features from the multi-channel audio data to obtain a target speech feature map. A graph convolutional neural network then performs graph convolution on the target adjacency matrix and the target speech feature map, fusing them to obtain a more comprehensive and accurate target feature vector. By classifying the target feature vector using a classification network, speech activity detection information from multi-channel audio data can be accurately obtained. This enables stable and accurate speech activity detection in both simple acoustic environments (quiet, static, etc.) and complex acoustic environments (noisy, moving, etc.), improving the accuracy and stability of speech activity detection.
[0025] In this embodiment of the invention, the voice activity detection method can be applied to electronic devices, including mobile phones, tablets, laptops, desktop computers, personal digital assistants, and wearable devices. Wearable devices can include head-mounted displays, smart bracelets, and ring devices. Head-mounted displays can include augmented reality (AR) glasses, AR helmets, mixed reality (MR) glasses, and MR helmets.
[0026] The following detailed description of some embodiments of the present invention is provided in conjunction with the accompanying drawings. Unless otherwise specified, the following embodiments and features can be combined with each other.
[0027] Please refer to Figure 1 , Figure 1 This is a flowchart illustrating a voice activity detection method provided in an embodiment of the present invention.
[0028] like Figure 1 As shown, the voice activity detection method includes steps S101 to S105.
[0029] Step S101: Acquire multi-channel audio data collected by the microphone array.
[0030] In this embodiment, the microphone array includes at least two microphones. A microphone array is an array formed by arranging microphones at different locations in space according to a certain shape rule. It is a device for spatially sampling sound input propagating in space, and the collected multi-channel audio data contains the spatial location information of the audio. Based on the topology of the microphone array, it can be classified as a linear array, planar array, volume array, etc. The multi-channel audio data includes dual-channel audio data, four-channel audio data, or eight-channel audio data, etc.
[0031] In some embodiments, acquiring multi-channel audio data collected by the microphone array may include: in response to a detected voice activity detection command, acquiring multi-channel audio data collected by the microphone array. This voice activity detection command can be manually triggered by the user or automatically triggered by the electronic device. For example, in some scenarios, a user needs to convert their speech into text. In this case, in response to the user's triggering of the voice input button displayed on the electronic device, a voice activity detection command is generated, and in response to the voice activity detection command, the microphone array is controlled to collect multi-channel audio data. Subsequently, voice activity detection is performed on the multi-channel audio data before speech recognition. As another example, when the electronic device is in a call state, in order to achieve call noise reduction, the electronic device automatically triggers a voice activity detection command and, in response to the voice activity detection command, controls the microphone array to collect multi-channel audio data. Subsequently, voice activity detection is performed on the multi-channel audio data before call noise reduction.
[0032] In some embodiments, the electronic device stores a voice activity detection model that matches the microphone array, such as Figure 2 As shown, the speech activity detection model includes a first convolutional neural module 110, a second convolutional neural module 120, a graph convolutional neural network 130, and a classification network 140. The first convolutional neural module 110 and the second convolutional neural module 120 are connected in parallel. The first convolutional neural module 110 and the second convolutional neural module 120 are respectively connected to the graph convolutional neural network 130, and the graph convolutional neural network 130 is connected to the classification network 140. The speech activity detection model is pre-trained on a preset neural network model using multiple training samples. These training samples include audio samples and labeled speech activity detection information. The audio samples include multi-channel audio data collected by a microphone array. It is understood that the number of microphones in the microphone array may vary, resulting in the same network structure for the speech activity detection model matched to the microphone array, but with different model parameters.
[0033] For example, if a microphone array in an electronic device includes two microphones, the speech activity detection model matching the two-microphone array is pre-trained on a preset neural network model using multiple first training samples. These first training samples include first audio samples and labeled speech activity detection information. The first audio samples include dual-channel audio data collected by the microphone array consisting of two microphones. As another example, if a microphone array in an electronic device includes four microphones, the speech activity detection model matching the four-microphone array is pre-trained on a preset neural network model using multiple second training samples. The second training samples include second audio samples and labeled speech activity detection information. The second audio samples include four-channel audio data collected by the microphone array consisting of four microphones.
[0034] Step S102: Construct an initial adjacency matrix based on multi-channel audio data, and optimize the initial adjacency matrix through the first convolutional neural module to obtain the target adjacency matrix.
[0035] In this embodiment, the initial adjacency matrix is a three-dimensional structure with dimensions n×n×N. The first and second dimensions (n×n) reflect the connectivity between the n channels, and the third dimension N represents the number of sampling points in a frame of multi-channel audio data, where n is the number of microphones in the microphone array. For example, if the microphone array includes two microphones and a frame of multi-channel audio data contains 800 sampling points, then the initial adjacency matrix has dimensions 2×2×800.
[0036] In some embodiments, constructing an initial adjacency matrix based on multi-channel audio data includes: determining the number of sampling points in the multi-channel audio data; determining the target dimension of the adjacency matrix based on the number of sampling points and the number of microphones in the microphone array; constructing a blank adjacency matrix of the target dimension, and setting each element in the blank adjacency matrix to a preset value to obtain the initial adjacency matrix. The preset value can be set based on actual conditions, and this embodiment of the invention does not specifically limit it. For example, if the preset value is 1, the number of microphones in the microphone array is 2, and the number of sampling points in one frame of multi-channel audio data is N, then the initial adjacency matrix can be represented as: A (0) ∈R 2×2×N , A (0) Let R be the initial adjacency matrix. 2×2×N Denotes the initial adjacency matrix A (0) The dimension is 2×2×N. Let A be the initial adjacency moment. (0)The element with index ijk is defined in the matrix. In this embodiment, each element in the initial adjacency matrix has a preset value, enabling the initial adjacency matrix to express the relationships between multiple microphone channels.
[0037] In some embodiments, the first convolutional neural module includes multiple cascaded convolutional neural layers. Each convolutional neural layer is used to perform convolution, normalization, and nonlinear activation processing on the adjacency matrix output by the previous convolutional neural layer. That is, each convolutional neural layer includes a convolutional layer, a normalization layer, and a nonlinear activation function layer. After convolving the output of the previous convolutional neural layer by the convolutional layer, the convolution result is normalized by the normalization layer, and finally, the normalized result is nonlinearly activated by the nonlinear activation function layer. For example, the convolution, normalization, and nonlinear activation processing of the adjacency matrix output by the previous convolutional neural layer by each convolutional neural layer can be represented as: A (l) =σ(Norm(W) (l) *A (l-1) +b (l) A (l) A is the adjacency matrix output by the l-th convolutional neural layer. (l-1) It is the adjacency matrix output by the (l-1)th convolutional neural layer, where σ represents the activation function, Norm represents the normalization operation, and W... (l) and b (l) These represent the convolution kernel and bias parameters of the l-th convolutional neural layer, respectively.
[0038] In some embodiments, the first convolutional neural module includes a first convolutional neural layer and a second convolutional neural layer, wherein the first convolutional neural layer is connected to the second convolutional neural layer. Optimizing the initial adjacency matrix using the first convolutional neural module to obtain the target adjacency matrix may include: performing convolution, normalization, and nonlinear activation on the initial adjacency matrix using the first convolutional neural layer to obtain the first adjacency matrix; and performing convolution, normalization, and nonlinear activation on the first adjacency matrix using the second convolutional neural layer to obtain the target adjacency matrix.
[0039] In some embodiments, such as Figure 3As shown, the first convolutional neural module 110 includes a first convolutional neural layer 111, a second convolutional neural layer 112, and a third convolutional neural layer 113. The first convolutional neural layer 111 is connected to the second convolutional neural layer 112, and the second convolutional neural layer 112 is connected to the third convolutional neural layer 113. Optimizing the initial adjacency matrix using the first convolutional neural module to obtain the target adjacency matrix can include: performing convolution, normalization, and nonlinear activation on the initial adjacency matrix using the first convolutional neural layer to obtain the first adjacency matrix; performing convolution, normalization, and nonlinear activation on the first adjacency matrix using the second convolutional neural layer to obtain the second adjacency matrix; and performing convolution, normalization, and nonlinear activation on the second adjacency matrix using the third convolutional neural layer to obtain the target adjacency matrix.
[0040] For example, the initial adjacency matrix A (0) After processing by the first convolutional neural module, the first convolutional neural module outputs the first adjacency matrix A. (1) =σ(Norm(W) (1) *A (0) +b (1) After the first adjacency matrix is processed by the second convolutional neural layer, the second convolutional neural module outputs the second adjacency matrix A. (2) =σ(Norm(W) (2) *A (1) +b (2) After the second adjacency matrix is processed by the third convolutional neural layer, the third convolutional neural layer outputs the third adjacency matrix A. (3) =σ(Norm(W) (3) *A (2) +b (2) At this point, the third adjacency matrix A can be used. (3) =σ(Norm(W) (3) *A (2) +b (2) The target adjacency matrix is determined.
[0041] Step S103: The second convolutional neural module is used to extract speech features from the multi-channel audio data to obtain the target speech feature map.
[0042] In this embodiment, the target speech feature map is used to describe the time-frequency features and semantic features of multi-channel audio data. The time-frequency features may include time-frequency energy, time-frequency edges and textures, short-time harmonic structure, etc., and the semantic features may include phonemes or phoneme combinations, intonation, etc.
[0043] In some embodiments, the second convolutional neural module includes multiple convolutional pooling modules connected in series. Each convolutional pooling module includes a convolutional module and a pooling layer. The number of convolutional layers in the convolutional pooling modules at different levels is different, and the parameters of the pooling layers are also different. The convolutional module performs convolution processing on the speech feature map output by the pooling layer of the previous level, and the pooling layer performs adaptive average pooling processing on the speech feature map output by the convolutional module at the same level. Starting from the second convolutional pooling module, the process by which each convolutional pooling module processes the speech feature map output by the convolutional pooling module of the previous level can be represented as follows: F (m) W is the speech feature map output by the convolutional module in the m-th layer's convolutional pooling module. (m) and b (m) Let represent the convolution kernel and bias parameters of the convolutional pooling module in layer m, respectively. Let σ represent the activation function, and Conv() represent the convolution operation. It is the speech feature map output by the pooling layer in the m-th convolutional pooling module. It is the speech feature map output by the pooling layer in the (m-1)th convolutional pooling module, and AdaptiveAvgPool() represents adaptive average pooling.
[0044] In some embodiments, such as Figure 4 As shown, the second convolutional neural module 120 includes a first convolutional module 121, a first pooling layer 122, a second convolutional module 123, and a second pooling layer 124. The second convolutional neural module performs speech feature extraction processing on multi-channel audio data to obtain a target speech feature map. This process includes: convolving the multi-channel audio data using the first convolutional module to obtain a first speech feature map; performing adaptive average pooling on the first speech feature map using the first pooling layer to obtain a second speech feature map; convolving the second speech feature map using the second convolutional module to obtain a third speech feature map; and performing adaptive average pooling on the third speech feature map using the second pooling layer to obtain the target speech feature map. The first convolutional module includes two convolutional layers, and the second convolutional module includes three convolutional layers.
[0045] For example, the first speech feature map is represented as: F (1) =σ(Conv(Q),W (1) )+b (1) The second speech feature map is represented as follows: The third speech feature map is represented as follows: The target speech feature map is represented as follows: Q represents multi-channel audio data, W (1) and b (1)W represents the convolution kernel and bias parameter of the first convolutional module, respectively. (2) and b (2) These represent the convolution kernel and bias parameters of the second convolution module, respectively.
[0046] Step S104: Perform graph convolution processing on the target adjacency matrix and the target speech feature map using a graph convolutional neural network to obtain the target feature vector.
[0047] This embodiment uses a graph convolutional neural network to perform graph convolution processing on the target adjacency matrix and the target speech feature map, thereby achieving full fusion of sound topological relationships and high-dimensional speech feature information, which can improve the robustness and accuracy of speech activity detection in complex acoustic environments.
[0048] Step S105: Classify the target feature vector through a classification network to obtain speech activity detection information of multi-channel audio data.
[0049] In this embodiment, the voice activity detection information of the multi-channel audio data includes a first tag or a second tag. The first tag indicates that the multi-channel audio data is voice data, and the second tag indicates that the multi-channel audio data is not voice data. The specific forms of the first and second tags can be set based on actual circumstances, and this embodiment of the invention does not impose specific limitations on them. It is understood that after obtaining the voice activity detection information of the multi-channel audio data, the voice activity detection information can be used to perform functions such as automatic speech recognition, keyword wake-up, voice enhancement, and call noise reduction according to actual business needs. For example, if the voice activity detection information determines that the multi-channel audio data is voice data, speech recognition is performed on the multi-channel audio data to obtain the corresponding speech recognition result, and the speech recognition result is displayed. However, if the voice activity detection information determines that the multi-channel audio data is not voice data, then no speech recognition operation is performed; instead, the multi-channel audio data is deleted.
[0050] In some embodiments, the classification network may include one or more fully connected layers and a classification layer. For example, as... Figure 5 As shown, the classification network 140 includes a first fully connected layer 141, a second fully connected layer 142, and a classification layer 143. The classification network performs classification processing on the target feature vector to obtain speech activity detection information of the multi-channel audio data. This includes: processing the target feature vector through the first fully connected layer to obtain a first feature vector; processing the first feature vector through the second fully connected layer to obtain a second feature vector; and classifying the second feature vector through the classification layer to obtain speech activity detection information of the multi-channel audio data. The classification layer may include a softmax layer.
[0051] Current speech activity detection methods, besides those based on energy or spectral features, also include those based on deep neural networks (DNN, CNN, LSTM, etc.). Deep neural network-based methods require the deployment of a deep neural network-based speech activity detection model, but these models are large, computationally complex, and demand significant computing and storage resources. This makes them difficult to adapt to electronic devices with limited computing and storage resources, compromising the efficiency of speech activity detection on such devices and potentially leading to increased inference latency and battery consumption.
[0052] The embodiments of the present invention provide, as follows Figure 5 The speech activity detection model shown is smaller and requires less computation compared to speech activity detection models based on deep neural networks (DNN, CNN, LSTM). It features real-time response and low power consumption, effectively adapting to electronic devices and IoT nodes with limited computing and storage resources. Furthermore, the first convolutional neural module in the speech activity detection model can automatically construct a speech topology graph (target adjacency matrix) based on the actual spatial layout of the microphone array in the electronic device, enabling dynamic transmission of spatial perception information between different channels. This design abandons the traditional manual graph construction method of GCN, possesses strong generalization ability, and can adapt to microphone arrays in different models or types of electronic devices. Experiments show that in typical AR scenarios such as noisy streets, crowded environments, and rapid head movements, this design significantly improves the accuracy and stability of speech activity detection.
[0053] Furthermore, the second convolutional neural module in the speech activity detection model can extract frame-level speech features from the audio spectrogram and fully integrate sound topological relationships with high-dimensional speech feature information through graph convolutional neural networks. This significantly improves the robustness and accuracy of speech activity detection in complex acoustic environments (such as wearer movement, multi-source interference, and strong background noise). This design overcomes the bottlenecks of traditional CNNs, such as difficulty in capturing inter-channel dependencies and unstable RNN training, enabling the speech activity detection model to better understand low-volume, intermittent speech and non-standard speech behaviors in wearable interactions, thus enhancing the user's voice interaction experience. Moreover, the speech activity detection model supports end-to-end training, has good platform compatibility, can adapt to different microphone arrangements and different speech input formats, and can be seamlessly integrated with front-end speech function modules such as voice wake-up, ASR, and speech enhancement. It is widely used in key scenarios such as voice interaction, far-field pickup, and private voice input in electronic devices, significantly enhancing the intelligent voice capabilities and user interaction experience of electronic devices.
[0054] In some embodiments, the voice activity detection method provided by this invention further includes: in response to at least one microphone in the microphone array malfunctioning, determining the number of microphones in the microphone array that are not malfunctioning; in response to the number of microphones in the microphone array that are not malfunctioning being greater than or equal to 2, downloading a voice activity detection model corresponding to the number of microphones in the microphone array that are not malfunctioning from a server, and updating the voice activity detection model currently stored in the electronic device to the downloaded voice activity detection model. The server stores voice activity detection models corresponding to different numbers of microphones. In this embodiment, when at least one microphone in the microphone array malfunctions, a voice activity detection model corresponding to the number of microphones in the microphone array that are not malfunctioning is downloaded from the server, thereby updating the stored voice activity detection model. This allows the updated voice activity detection model to match the microphones in the microphone array that are not malfunctioning, enabling the updated voice activity detection model to accurately detect voice activity in the multi-channel audio data collected by the microphones in the microphone array that are not malfunctioning.
[0055] In some embodiments, downloading a speech activity detection model corresponding to the number of microphones in the microphone array that are not malfunctioning from a server and updating the speech activity detection model currently stored in the electronic device to the downloaded speech activity detection model includes: downloading model parameters corresponding to the number of microphones in the microphone array that are not malfunctioning from a server, and updating the model parameters of the speech activity detection model currently stored in the electronic device to the downloaded model parameters, thereby updating the speech activity detection model. The server stores model parameters for speech activity detection models corresponding to different numbers of microphones. This embodiment does not require updating the network structure of the speech activity detection model; it only updates the model parameters, reducing the time required to update the speech activity detection model and improving the update efficiency.
[0056] In some embodiments, in response to the number of microphones without faults in the microphone array being greater than or equal to 2, multi-channel audio data collected by the microphone array is acquired; an initial adjacency matrix is constructed based on the multi-channel audio data, and the initial adjacency matrix is optimized by a first convolutional neural module to obtain a target adjacency matrix; speech feature extraction processing is performed on the multi-channel audio data by a second convolutional neural module to obtain a target speech feature map; graph convolution processing is performed on the target adjacency matrix and the target speech feature map by a graph convolutional neural network to obtain a target feature vector; the target feature vector is classified by a classification network to obtain speech activity detection information of the multi-channel audio data; in response to the number of microphones without faults in the microphone array being 1, single-channel audio data collected by the microphones without faults in the microphone array is acquired; the short-time energy and short-time zero-crossing rate of the single-channel audio data are determined, and the speech activity detection information of the single-channel audio data is determined based on the short-time energy and short-time zero-crossing rate of the single-channel audio data. In this embodiment, when the microphone array includes at least two non-faulty microphones, speech activity detection is performed on multi-channel audio data using a speech activity detection model. When the microphone array includes only one non-faulty microphone, speech activity detection is performed on single-channel audio data using short-time energy and short-time zero-crossing rate. This ensures that the electronic device can still perform speech activity detection even when some microphones in the microphone array fail, thus improving the stability of speech activity detection.
[0057] Please see Figure 6 , Figure 6 This is a schematic block diagram of the structure of an electronic device provided in an embodiment of the present invention.
[0058] like Figure 6 As shown, the electronic device 100 includes a processor 101 and a memory 102, which are connected via a bus 103, such as an I2C (Inter-integrated Circuit) bus.
[0059] Specifically, processor 101 provides computing and control capabilities to support the operation of the entire electronic device. Processor 101 can be a Central Processing Unit (CPU), but it can also be other general-purpose processors, digital signal processors (DSPs), application-specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. The general-purpose processor can be a microprocessor or any conventional processor.
[0060] Specifically, the memory 102 can be a Flash chip, a read-only memory (ROM) disk, an optical disk, a USB flash drive, or a portable hard drive, etc.
[0061] Those skilled in the art will understand that Figure 6 The structures shown are merely block diagrams of some structures related to the embodiments of the present invention, and do not constitute a limitation on the electronic devices to which the embodiments of the present invention are applied. Specific electronic devices may include more or fewer components than those shown in the figures, or combine certain components, or have different component arrangements.
[0062] The processor 101 is used to run a computer program stored in the memory 102, and implements any of the voice activity detection methods provided in the embodiments of the present invention when executing the computer program.
[0063] In one embodiment, the electronic device includes a microphone array and a speech activity detection model matched with the microphone array. The speech activity detection model includes a first convolutional neural module, a second convolutional neural module, a graph convolutional neural network, and a classification network. The processor 101 is configured to run a computer program stored in a memory and, when executing the computer program, perform the following steps:
[0064] Acquire multi-channel audio data collected by the microphone array;
[0065] An initial adjacency matrix is constructed based on the multi-channel audio data, and the initial adjacency matrix is optimized by the first convolutional neural module to obtain the target adjacency matrix.
[0066] The second convolutional neural module performs speech feature extraction processing on the multi-channel audio data to obtain a target speech feature map.
[0067] The target feature vector is obtained by performing graph convolution processing on the target adjacency matrix and the target speech feature map using the graph convolutional neural network.
[0068] The target feature vector is classified using the classification network to obtain speech activity detection information from the multi-channel audio data.
[0069] In some embodiments, the first convolutional neural module includes a first convolutional neural layer, a second convolutional neural layer, and a third convolutional neural layer. When the processor 101 optimizes the initial adjacency matrix using the first convolutional neural module to obtain the target adjacency matrix, it is configured to:
[0070] The initial adjacency matrix is obtained by performing convolution, normalization, and nonlinear activation on the first convolutional neural layer;
[0071] The second convolutional neural layer performs convolution, normalization, and nonlinear activation on the first adjacency matrix to obtain the second adjacency matrix.
[0072] The target adjacency matrix is obtained by performing convolution, normalization, and nonlinear activation on the second adjacency matrix through the third convolutional neural layer.
[0073] In some embodiments, the second convolutional neural module includes a first convolutional module, a first pooling layer, a second convolutional module, and a second pooling layer. When the processor 101 performs speech feature extraction processing on the multi-channel audio data through the second convolutional neural module to obtain a target speech feature map, it is used to:
[0074] The first convolution module performs convolution processing on the multi-channel audio data to obtain a first speech feature map.
[0075] The first speech feature map is obtained by adaptive average pooling through the first pooling layer.
[0076] The second speech feature map is processed by the second convolution module to obtain the third speech feature map;
[0077] The target speech feature map is obtained by adaptive average pooling of the third speech feature map through the second pooling layer.
[0078] In some embodiments, the classification network includes a first fully connected layer, a second fully connected layer, and a classification layer. When the processor 101 performs classification processing on the target feature vector through the classification network to obtain speech activity detection information of the multi-channel audio data, it is configured to:
[0079] The target feature vector is processed by the first fully connected layer to obtain a first feature vector;
[0080] The second feature vector is obtained by processing the first feature vector through the second fully connected layer;
[0081] The second feature vector is classified by the classification layer to obtain the speech activity detection information of the multi-channel audio data.
[0082] In some embodiments, when the processor 101 constructs an initial adjacency matrix based on the multi-channel audio data, it is configured to:
[0083] The number of sampling points for the multi-channel audio data is determined, and the target dimension of the adjacency matrix is determined based on the number of sampling points and the number of microphones contained in the microphone array.
[0084] Construct a blank adjacency matrix for the target dimension, and set each element in the blank adjacency matrix to a preset value to obtain the initial adjacency matrix.
[0085] In some embodiments, the processor 101 is further configured to implement:
[0086] In response to a failure of at least one microphone in the microphone array, determine the number of microphones in the microphone array that are not faulty;
[0087] In response to the quantity being greater than or equal to 2, a voice activity detection model corresponding to the quantity is downloaded from the server, and the voice activity detection model currently stored in the electronic device is updated to the downloaded voice activity detection model.
[0088] In some embodiments, when the processor 101 acquires the multi-channel audio data collected by the microphone array, it is configured to:
[0089] In response to a number of microphones in the microphone array that are not faulty being greater than or equal to 2, the multi-channel audio data collected by the microphone array is acquired.
[0090] In some embodiments, the processor 101 is further configured to implement:
[0091] In response to the fact that the number of microphones in the microphone array that are not faulty is 1, the single-channel audio data collected by the microphones in the microphone array that are not faulty is acquired.
[0092] The short-time energy and short-time zero-crossing rate of the single-channel audio data are determined, and the speech activity detection information of the single-channel audio data is determined based on the short-time energy and short-time zero-crossing rate of the single-channel audio data.
[0093] It should be noted that those skilled in the art will understand that, for the sake of convenience and brevity, the specific working process of the electronic device described above can be referred to the corresponding process in the aforementioned speech activity detection method embodiments, and will not be repeated here.
[0094] This invention also provides a storage medium for computer-readable storage, wherein the storage medium stores one or more programs that can be executed by one or more processors to implement any of the voice activity detection methods provided in the specification of this invention.
[0095] The storage medium can be an internal storage unit of the electronic device described in the foregoing embodiments, such as a hard drive or memory of the electronic device. Alternatively, the storage medium can be an external storage device of the electronic device, such as a plug-in hard drive, Smart Media Card (SMC), Secure Digital (SD) card, or Flash Card.
[0096] It will be understood by those skilled in the art that all or some of the steps, systems, or apparatuses disclosed above, and their functional modules / units, can be implemented as software, firmware, hardware, or suitable combinations thereof. In hardware embodiments, the division between functional modules / units mentioned in the above description does not necessarily correspond to the division of physical components; for example, a physical component may have multiple functions, or a function or step may be performed collaboratively by several physical components. Some or all physical components may be implemented as software executed by a processor, such as a central processing unit, digital signal processor, or microprocessor, or as hardware, or as an integrated circuit, such as an application-specific integrated circuit. Such software may be distributed on a computer-readable medium, which may include computer storage media (or non-transitory media) and communication media (or transient media). As is known to those skilled in the art, the term computer storage media includes volatile and non-volatile, removable and non-removable media implemented in any method or technology for storing information (such as computer-readable instructions, data structures, program modules, or other data). Computer storage media include, but are not limited to, RAM, ROM, EEPROM, flash memory or other memory technologies, CD-ROM, digital versatile disc (DVD) or other optical disc storage, magnetic cartridges, magnetic tape, disk storage or other magnetic storage devices, or any other medium that can be used to store desired information and can be accessed by a computer. Furthermore, it is well known to those skilled in the art that communication media typically contain computer-readable instructions, data structures, program modules, or other data in modulated data signals such as carrier waves or other transmission mechanisms, and may include any information delivery medium.
[0097] It should be understood that the term "and / or" as used in this specification and the appended claims refers to any combination and all possible combinations of one or more of the associated listed items, and includes such combinations. It should be noted that, herein, the terms "comprising," "including," or any other variations thereof are intended to cover a non-exclusive inclusion, such that a process, method, article, or system that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such a process, method, article, or system. Without further limitation, an element defined by the phrase "comprising one..." does not exclude the presence of other identical elements in the process, method, article, or system that includes that element.
[0098] The sequence numbers of the above embodiments of the present invention are merely for descriptive purposes and do not represent the superiority or inferiority of the embodiments. The above descriptions are only specific embodiments of the present invention, but the scope of protection of the present invention is not limited thereto. Any person skilled in the art can easily conceive of various equivalent modifications or substitutions within the technical scope disclosed in the present invention, and these modifications or substitutions should all be covered within the scope of protection of the present invention. Therefore, the scope of protection of the present invention should be determined by the scope of the claims.
Claims
1. A method for detecting speech activity, characterized in that, Applied to an electronic device, the electronic device including a microphone array and storing a speech activity detection model matched with the microphone array, the speech activity detection model including a first convolutional neural module, a second convolutional neural module, a graph convolutional neural network, and a classification network, the method includes: Acquire multi-channel audio data collected by the microphone array; An initial adjacency matrix is constructed based on the multi-channel audio data, and the initial adjacency matrix is optimized by the first convolutional neural module to obtain the target adjacency matrix. The second convolutional neural module performs speech feature extraction processing on the multi-channel audio data to obtain a target speech feature map. The target feature vector is obtained by performing graph convolution processing on the target adjacency matrix and the target speech feature map using the graph convolutional neural network. The target feature vector is classified by the classification network to obtain the speech activity detection information of the multi-channel audio data; The first convolutional neural module includes a first convolutional neural layer, a second convolutional neural layer, and a third convolutional neural layer. The step of optimizing the initial adjacency matrix using the first convolutional neural module to obtain the target adjacency matrix includes: performing convolution, normalization, and nonlinear activation on the initial adjacency matrix using the first convolutional neural layer to obtain a first adjacency matrix; performing convolution, normalization, and nonlinear activation on the first adjacency matrix using the second convolutional neural layer to obtain a second adjacency matrix; and performing convolution, normalization, and nonlinear activation on the second adjacency matrix using the third convolutional neural layer to obtain the target adjacency matrix. The second convolutional neural module includes a first convolutional module, a first pooling layer, a second convolutional module, and a second pooling layer. The step of extracting speech features from the multi-channel audio data using the second convolutional neural module to obtain a target speech feature map includes: performing convolution processing on the multi-channel audio data using the first convolutional module to obtain a first speech feature map; performing adaptive average pooling processing on the first speech feature map using the first pooling layer to obtain a second speech feature map; and performing convolution processing on the second speech feature map using the second convolutional module to obtain a third speech feature map. The target speech feature map is obtained by adaptive average pooling of the third speech feature map through the second pooling layer.
2. The speech activity detection method according to claim 1, characterized in that, The classification network includes a first fully connected layer, a second fully connected layer, and a classification layer. The classification process, which classifies the target feature vector using the classification network to obtain speech activity detection information from the multi-channel audio data, includes: The target feature vector is processed by the first fully connected layer to obtain a first feature vector; The second feature vector is obtained by processing the first feature vector through the second fully connected layer; The second feature vector is classified by the classification layer to obtain the speech activity detection information of the multi-channel audio data.
3. The speech activity detection method according to claim 1, characterized in that, The step of constructing an initial adjacency matrix based on the multi-channel audio data includes: The number of sampling points for the multi-channel audio data is determined, and the target dimension of the adjacency matrix is determined based on the number of sampling points and the number of microphones contained in the microphone array. Construct a blank adjacency matrix for the target dimension, and set each element in the blank adjacency matrix to a preset value to obtain the initial adjacency matrix.
4. The speech activity detection method according to claim 1, characterized in that, The method further includes: In response to a failure of at least one microphone in the microphone array, determine the number of microphones in the microphone array that are not faulty; In response to the quantity being greater than or equal to 2, a voice activity detection model corresponding to the quantity is downloaded from the server, and the voice activity detection model currently stored in the electronic device is updated to the downloaded voice activity detection model.
5. The speech activity detection method according to any one of claims 1-4, characterized in that, The acquisition of multi-channel audio data collected by the microphone array includes: In response to a number of microphones in the microphone array that are not faulty being greater than or equal to 2, the multi-channel audio data collected by the microphone array is acquired.
6. The voice activity detection method according to claim 5, characterized in that, The method further includes: In response to the fact that the number of microphones without faults in the microphone array is 1, single-channel audio data collected by the microphones without faults in the microphone array is acquired. The short-time energy and short-time zero-crossing rate of the single-channel audio data are determined, and the speech activity detection information of the single-channel audio data is determined based on the short-time energy and short-time zero-crossing rate of the single-channel audio data.
7. An electronic device, characterized in that, The electronic device includes a processor, a memory, a computer program stored in the memory and executable by the processor, and a data bus for establishing communication between the processor and the memory, wherein when the computer program is executed by the processor, it implements the steps of the voice activity detection method as described in any one of claims 1 to 6.
8. A storage medium for computer-readable storage, characterized in that, The storage medium stores one or more programs, which can be executed by one or more processors to implement the steps of the voice activity detection method according to any one of claims 1 to 6.
Citation Information
Patent Citations
Data processing method, electronic device, storage medium and computer program product
CN118942491A
Adaptive targeting for proactive voice notifications
US12020690B1