Streaming multichannel end-to-end speech recognition method based on spherical harmonic function transformation
By using spherical harmonic function transformation and deep learning modules for dereverberation and beamforming in streaming multi-channel speech recognition, the problem of high computational complexity in the prior art is solved, and the computing speed and accuracy of streaming speech recognition are improved.
Patent Information
- Application Number
- CN202411987165.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2024-12-31
- Publication Date
- 2025-05-13
- Estimated Expiration
- 2044-12-31
AI Technical Summary
In the prior art, the high computational complexity and slow computational speed are caused by the CUSIDE-Array method being unfavorable for voice signal processing in streaming situations.
The streaming multi-channel end-to-end speech recognition method based on spherical harmonic function transformation is adopted to extract spatial information of spectrum data through spherical harmonic transformation, and combine the convolution block attention module and self-attention channel combiner to perform dereverberation and beam formation, reducing operation complexity.
It reduces the computational complexity and computational volume of streaming speech recognition, improves the computational speed, and is suitable for multi-channel voice signal processing in streaming situations.
Smart Images

Figure CN119993122A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of speech recognition technology, and in particular to a streaming multi-channel end-to-end speech recognition method based on spherical harmonic function transformation. Background Art
[0002] Multi-channel Automatic Speech Recognition (ASR) systems have been continuously studied and improved because they can improve the robustness and accuracy of speech recognition through multi-channel input, especially in far-field acoustic environments. A beamforming front end is usually introduced before the ASR back end to utilize the spatial information of multi-channel speech signals for speech enhancement.
[0003] In multi-channel speech signal recognition, the key is to make good use of the spatial information in the multi-channel signal and to effectively integrate and process the spatial clues. For streaming recognition in multi-channel scenarios, the existing technology uses the CUSIDE-Array (Chunking, Simulating Future Context, and Decoding- Array) method. Based on the CUSIDE framework, the CUSIDE-Array method uses a neural beamforming filter front end and performs end-to-end training to achieve streaming multi-channel speech recognition. The beamforming front end used by CUSIDE-Array is a mask-based minimum variance distortionless response (MVDR) beamforming filter, which estimates the mask for each channel separately, then calculates the spatial covariance matrix of the signal and noise, and brings it into the beamforming filter parameter estimation formula to obtain the parameters of the beamforming filter.
[0004] However, the mask-based MVDR beamforming (Beamforming module) in CUSIDE-Array involves the calculation of multiple spatial covariance matrices and multiple matrix inversion operations when calculating the parameters of the beamforming filter. The operation complexity is high and the operation speed is slow, which makes the CUSIDE-Array method unfavorable for speech signal processing in streaming scenarios. Summary of the invention
[0005] The present invention provides a streaming multi-channel end-to-end speech recognition method based on spherical harmonic function transformation, which is used to solve the defects of the prior art that the CUSIDE-Array method is not conducive to speech signal processing in a streaming situation due to high computational complexity and slow computational speed, thereby reducing computational complexity and computational amount and improving computational speed.
[0006] In a first aspect, the present invention provides a streaming multi-channel end-to-end speech recognition method based on spherical harmonics transform, the method comprising the following steps: Acquire a multi-channel signal waveform to be identified; the multi-channel signal waveform to be identified includes at least two frames of multi-channel signals; Performing spherical harmonic transformation on the spectrum data corresponding to each frame of the multi-channel signal to obtain the spectrum data corresponding to each frame of the multi-channel signal after the number of channels is transformed, and converting the spectrum data corresponding to each frame of the multi-channel signal after the number of channels is transformed into the first amplitude spectrum data of each frame; the spectrum data is a time-frequency domain signal; Based on the first amplitude spectrum data of each frame, a convolution block attention module is used to determine the first amplitude spectrum data of each frame after dereverberation, and according to the first amplitude spectrum data of each frame after dereverberation, a self-attention channel combiner is used to determine the beam signal after beamforming; According to the beam signal after beamforming, the target category corresponding to the multi-channel signal waveform to be identified is obtained by using the target ASR network recognition.
[0007] According to a streaming multi-channel end-to-end speech recognition method based on spherical harmonic function transformation provided by the present invention, the spherical harmonic transformation is performed on the spectrum data corresponding to each frame of the multi-channel signal to obtain the spectrum data corresponding to each frame of the multi-channel signal after the number of channels is transformed, including: Based on the order of the spherical harmonic transformation function, the number of channels after transformation is determined; Based on the number of channels after the transformation, determining the spatial information corresponding to the spectrum data corresponding to each of the frame multi-channel signals; Based on the spatial information corresponding to the frequency spectrum data corresponding to each of the frame multi-channel signals, the frequency spectrum data corresponding to each of the frame multi-channel signals after the number of channels is converted is determined.
[0008] According to a streaming multi-channel end-to-end speech recognition method based on spherical harmonic function transformation provided by the present invention, the convolution block attention module includes a channel attention module and a spatial attention module; based on the first amplitude spectrum data of each frame, the convolution block attention module is used to determine the first amplitude spectrum data of each frame after dereverberation, including: Input the first amplitude spectrum data of each frame into the channel attention module to obtain the channel attention matrix corresponding to the first amplitude spectrum data of each frame; Determine the second amplitude spectrum data of each frame based on the channel attention matrix and the first amplitude spectrum data of each frame; Input the second amplitude spectrum data of each frame into the spatial attention module to obtain a spatial attention matrix corresponding to the second amplitude spectrum data of each frame; Based on the spatial attention matrix and the second amplitude spectrum data of each of the frames, the first amplitude spectrum data of each of the frames after dereverberation is determined.
[0009] According to a streaming multi-channel end-to-end speech recognition method based on spherical harmonic function transformation provided by the present invention, the beam signal after beamforming is determined by using a self-attention channel combiner according to the first amplitude spectrum data of each frame after dereverberation, including: Utilizing the fully connected function of the activation layer in the self-attention channel combiner, the first amplitude spectrum data of each frame after dereverberation is activated to obtain the key vector, query vector and value vector corresponding to the first amplitude spectrum data of each frame after dereverberation; the dimension of the first amplitude spectrum data of each frame after dereverberation is determined based on the number of frames, channels and frequencies included in the multi-channel signal waveform to be identified; Determining a self-attention matrix for a key based on the query vector, the key vector, and the transformed number of channels; Determining a self-attention matrix for a value based on the self-attention matrix and the value vector; Based on the self-attention matrix of the values and the first amplitude spectrum data of each frame after dereverberation, the beam signal after beamforming is determined.
[0010] According to a streaming multi-channel end-to-end speech recognition method based on spherical harmonic function transformation provided by the present invention, the target category corresponding to the multi-channel signal waveform to be identified is obtained by using a target ASR network to identify the beam signal after beamforming, including: Post-filtering the beam signal after the beamforming using a multi-head self-attention module to obtain a filtered beam signal; The ASR network is used to identify the target category corresponding to the multi-channel signal waveform to be identified based on the filtered beam signal.
[0011] According to a streaming multi-channel end-to-end speech recognition method based on spherical harmonic function transformation provided by the present invention, before performing spherical harmonic transformation on the spectrum data corresponding to each frame multi-channel signal to obtain the spectrum data corresponding to each frame multi-channel signal after the number of channels is transformed, the method further includes: The multi-channel signal waveform to be identified is subjected to short-time Fourier transform frame by frame to obtain frequency spectrum data corresponding to each frame of the multi-channel signal.
[0012] In a second aspect, the present invention further provides a streaming multi-channel end-to-end speech recognition device based on spherical harmonic function transformation, the device comprising the following modules: An acquisition module, used for acquiring a multi-channel signal waveform to be identified; the multi-channel signal waveform to be identified includes at least two frames of multi-channel signals; A beamforming module is used to perform spherical harmonic transformation on the spectrum data corresponding to each frame of the multi-channel signal to obtain the spectrum data corresponding to each frame of the multi-channel signal after the number of channels is transformed, and convert the spectrum data corresponding to each frame of the multi-channel signal after the number of channels is transformed into the first amplitude spectrum data of each frame; the spectrum data is a time-frequency domain signal; Based on the first amplitude spectrum data of each frame, a convolution block attention module is used to determine the first amplitude spectrum data of each frame after dereverberation, and according to the first amplitude spectrum data of each frame after dereverberation, a self-attention channel combiner is used to determine the beam signal after beamforming; The identification module is used to obtain the target category corresponding to the multi-channel signal waveform to be identified by using the target ASR network according to the beam signal after the beamforming.
[0013] In a third aspect, the present invention also provides an electronic device, comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein when the processor executes the computer program, the processor implements a streaming multi-channel end-to-end speech recognition method based on spherical harmonic transform as described in any one of the above.
[0014] In a fourth aspect, the present invention further provides a non-transitory computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements a streaming multi-channel end-to-end speech recognition method based on spherical harmonic function transform as described in any one of the above.
[0015] In a fifth aspect, the present invention further provides a computer program product, comprising a computer program, which, when executed by a processor, implements any of the above-described streaming multi-channel end-to-end speech recognition methods based on spherical harmonic transform.
[0016] The present invention provides a streaming multi-channel end-to-end speech recognition method based on spherical harmonic function transformation. First, a multi-channel signal waveform to be recognized is obtained, wherein the multi-channel signal waveform to be recognized includes at least two frames of multi-channel signals; then, spherical harmonic transformation is performed on the spectrum data corresponding to each frame of the multi-channel signal to obtain the spectrum data corresponding to each frame of the multi-channel signal after the number of channels is changed, and the spectrum data corresponding to each frame of the multi-channel signal after the number of channels is changed is converted into the first amplitude spectrum data of each frame, and the spectrum data is a time-frequency domain signal; based on the first amplitude spectrum data of each frame, a convolution block attention module is used to determine the first amplitude spectrum data of each frame after dereverberation, and according to the first amplitude spectrum data of each frame after dereverberation, a self-attention channel combiner is used to determine the beam signal after beamforming; further, according to the beam signal after beamforming, a target ASR network is used to identify the target category corresponding to the multi-channel signal waveform to be recognized.
[0017] The present invention uses spherical harmonic transform SHT to extract spatial information from spectrum data corresponding to each frame of multi-channel signal, obtains the transformed signal, that is, the spectrum data corresponding to each frame of multi-channel signal after the number of channels is changed, and then sent to the convolution block attention module for dereverberation, and then sent to the self-attention channel combiner SACC. The self-attention channel combiner SACC uses the self-attention mechanism to beamform the signal, and then sends it to the ASR network for recognition. The amount of beamforming calculation is small. The present invention reduces the calculation complexity and amount of streaming speech recognition, improves the calculation speed, and is suitable for multi-channel speech signal processing in streaming scenarios. BRIEF DESCRIPTION OF THE DRAWINGS
[0018] In order to more clearly illustrate the technical solutions in the present invention or the prior art, the following briefly introduces the drawings required for use in the embodiments or the description of the prior art. Obviously, the drawings described below are some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying creative work.
[0019] Figure 1 It is a flow chart of a streaming multi-channel end-to-end speech recognition method based on spherical harmonic transformation provided by the present invention.
[0020] Figure 2 It is a schematic diagram of the framework of the streaming multi-channel end-to-end speech recognition method based on spherical harmonic function transformation provided by the present invention.
[0021] Figure 3 It is a structural schematic diagram of the self-attention channel combiner SACC provided by the present invention.
[0022] Figure 4 It is one of the structural schematic diagrams of the convolutional block attention module CBAM provided by the present invention.
[0023] Figure 5 This is the second structural diagram of the convolutional block attention module CBAM provided by the present invention.
[0024] Figure 6 It is a structural schematic diagram of a streaming multi-channel end-to-end speech recognition device based on spherical harmonic function transformation provided by the present invention.
[0025] Figure 7 It is a structural schematic diagram of the electronic device provided by the present invention. DETAILED DESCRIPTION
[0026] In order to make the purpose, technical solution and advantages of the present invention clearer, the technical solution of the present invention will be clearly and completely described below in conjunction with the drawings of the present invention. Obviously, the described embodiments are part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of the present invention.
[0027] In order to more clearly understand the various embodiments provided by the present invention, the technical contents involved in the present invention are first introduced as follows: Compared with single-channel speech signals, multi-channel speech signals have more advantages in signal processing because they have more spatial information. However, in multi-channel speech signal recognition, the key is to make good use of the spatial information in multi-channel signals, and effectively integrating and processing spatial clues is still a challenge.
[0028] Traditional spatial filtering methods include delay-and-beamformers, minimum variance distortion-free response (MVDR) beamformers, super-directional beamformers, etc. These exploit phase and timing differences between microphones to preferentially extract signals from certain directions. While these methods work well, their performance depends on reliable estimation of spatial information, which can be challenging to accurately estimate in noisy conditions.
[0029] In recent years, deep learning has made great progress in multi-channel speech enhancement. For example, MMUB has taken advantage of the complementary advantages of deep learning and array processing to achieve state-of-the-art performance in speech enhancement tasks. However, the current streaming multi-channel end-to-end speech recognition system CUSIDE-Array has the following problems: 1. CUSIDE-Array directly connects the short-time Fourier transform (STFT) from each microphone as the model input. They rely on the powerful modeling ability of neural networks to mine the spatial information of the sound source. However, the traditional STFT representation is difficult to express the spatial information of the sound source.
[0030] 2. Mask-based MVDR beamforming (Beamforming module) in CUSIDE-Array This model estimates the mask for each channel separately. The effect of beamforming is extremely dependent on the accurate estimation of the mask. Independent per-channel processing cannot capture the inter-channel dependencies and spatial relationships that provide valuable context, affecting the subsequent recognition structure.
[0031] 3. The MVDR method in CUSIDE-Array involves the calculation of multiple spatial covariance matrices and multiple matrix inversion operations when calculating beamforming coefficients. The operation complexity is high and the operation speed is slow, which is not conducive to speech signal processing in streaming scenarios.
[0032] In view of the above shortcomings, the present invention provides a streaming multi-channel end-to-end speech recognition method based on spherical harmonic function transformation, which reduces the calculation complexity, improves the calculation speed, and improves the accuracy of model recognition.
[0033] Combine the following Figure 1-Figure 7 The present invention describes a streaming multi-channel end-to-end speech recognition method based on spherical harmonic function transformation.
[0034] Figure 1 is one of the flow charts of the streaming multi-channel end-to-end speech recognition method based on spherical harmonic function transformation provided by the present invention, such as Figure 1 As shown, the method comprises the following steps: Step 101, obtaining a multi-channel signal waveform to be identified; the multi-channel signal waveform to be identified includes at least two frames of multi-channel signals; Specifically, it should be noted that the executor of the present invention is an electronic device, which is used to implement streaming multi-channel end-to-end speech recognition, reduce the amount of calculation, and improve the calculation speed.
[0035] In the embodiment of the present invention, in view of the shortcomings of the CUSIDE-Array method of the streaming multi-channel end-to-end speech recognition system, we provide a streaming multi-channel end-to-end speech recognition method based on spherical harmonic function transformation. Figure 2 is a schematic diagram of the framework of the streaming multi-channel end-to-end speech recognition method based on spherical harmonic function transformation provided by the present invention, such as Figure 2 As shown, the present invention uses Spherical Harmonic Transformation (SHT) plus a series of deep learning modules such as Self-Attention Channel Combinator (SACC) to replace the mask-based MVDR beamforming module, and achieves good results. Figure 3 : is a schematic diagram of the structure of the self-attention channel combiner SACC provided by the present invention, such as Figure 3 As shown, a series of deep learning modules include SHT transform, amplitude spectrum conversion, CBAM, SACC, and MESA.
[0036] The method of this embodiment is implemented by the following steps: First, a multi-channel signal waveform to be identified is obtained, wherein the multi-channel signal waveform to be identified includes at least two frames of multi-channel signals.
[0037] Before beamforming, the multi-channel signal waveform to be identified is transformed by short-time Fourier transform (STFT) to transform the signal into a time-frequency domain signal, that is, spectrum data. B represents the batch size, C represents the number of channels, N represents the number of sampling points, T represents the number of frames, and F represents the frequency.
[0038] Step 102, performing spherical harmonic transformation on the spectrum data corresponding to each frame of the multi-channel signal to obtain the spectrum data corresponding to each frame of the multi-channel signal after the number of channels is transformed, and converting the spectrum data corresponding to each frame of the multi-channel signal after the number of channels is transformed into the first amplitude spectrum data of each frame; the spectrum data is a time-frequency domain signal; Specifically, after obtaining the spectrum data corresponding to each frame of the multi-channel signal (the spectrum data is a time-frequency domain signal), the present invention introduces a method of using the order of spherical harmonic transform (SHT) as an auxiliary model input.
[0039] In this embodiment, based on the order of spherical harmonic transform SHT, the spectrum data corresponding to each frame of multi-channel signal is transformed to obtain the spectrum data corresponding to each frame of multi-channel signal after the number of channels is transformed. The order of spherical harmonic transform SHT concisely represents the spatial distribution. It has two advantages: 1. Effectively capture spatial information: Unlike the short-time Fourier transform (STFT), the spherical harmonic transform (SHT) mainly captures the spatial distribution characteristics of the sound field. Based on the spherical harmonic theory, SHT identifies the spatial properties of signals from different directions and their mutual relationship on the microphone channel. This spatial capture is crucial for multi-microphone speech enhancement. In a multi-microphone array, this transform can skillfully capture the spatial direction of the sound, thereby more accurately distinguishing the target speech from the background noise.
[0040] 2. Improve spatial resolution: Spherical harmonics form the complete basis of functions defined on a sphere. Therefore, any spherical function can be accurately represented as a linear combination of spherical harmonics. SHT helps to accurately describe the spatial distribution of the sound field. The order of spherical harmonics determines the granularity of spatial feature capture. While low orders describe broad spatial patterns, high orders describe more subtle spatial differences. By choosing the right order, the desired spatial resolution can be achieved. If the spatial information in these SHCs can be fully utilized, this may help to make up for the shortcomings of descriptive spatial modeling in current mainstream methods. Taking advantage of this can greatly improve the performance and robustness of multi-channel speech enhancement.
[0041] Furthermore, the spectrum data corresponding to the multi-channel signal of each frame after the number of channels is changed can be converted into the first amplitude spectrum data Magnitude of each frame, so as to facilitate the subsequent deep learning module to process and obtain the beam signal. The deep learning module includes, for example, a convolutional block attention module (CBAM) and a self-attention channel combiner (SACC).
[0042] Step 103: Based on the first amplitude spectrum data of each frame, a convolution block attention module is used to determine the first amplitude spectrum data of each frame after dereverberation, and a self-attention channel combiner is used to determine the beam signal after beamforming according to the first amplitude spectrum data of each frame after dereverberation; Specifically, after the first amplitude spectrum data of each frame is obtained by conversion, the first amplitude spectrum data of each frame after dereverberation is determined by using a convolutional block attention module CBAM. The convolutional block attention module CBAM module can well capture the spatial dependency between channels to enhance feature representation. The present invention uses four layers of convolutional block attention modules CBAM in cascade to simulate the dereverberation process.
[0043] Furthermore, the beam signal after beamforming can be determined by using the self-attention channel combiner SACC according to the first amplitude spectrum data of each frame after dereverberation. The calculation process of the self-attention channel combiner SACC can be approximated with the MVDR beamforming process. The self-attention matrix w att The dimensions of the value matrix are similar to the matrix dimensions required for beamforming filter coefficient calculation, so the effects achieved are similar.
[0044] Step 104: According to the beam signal after beamforming, the target ASR network is used to identify the target category corresponding to the multi-channel signal waveform to be identified.
[0045] Specifically, after the above processing, a beam signal S after beamforming can be obtained. Further, the beam signal after beamforming is sent to a target multi-channel automatic speech recognition (Automatic Speech Recognition, ASR) network, and the target ASR network is used to identify the target category corresponding to the multi-channel signal waveform to be identified.
[0046] For example, the target ASR network includes a conformer based on convolution and self-attention mechanisms. The conformer combines the ability of convolution to model local information and the ability of transformer to model global information, achieving excellent results in speech recognition tasks. A conformer block mainly includes a feed forward module (FFN), a convolution module (Conv), and a multi-headed self-attention module (MHSA). The recognition result output by the last conformer is determined as the target category corresponding to the multi-channel signal waveform to be recognized.
[0047] The method provided in this embodiment first obtains a multi-channel signal waveform to be identified, wherein the multi-channel signal waveform to be identified includes at least two frames of multi-channel signals; then, spherical harmonic transform is performed on the spectrum data corresponding to each frame of the multi-channel signal to obtain the spectrum data corresponding to each frame of the multi-channel signal after the number of channels is changed, and the spectrum data corresponding to each frame of the multi-channel signal after the number of channels is changed is converted into the first amplitude spectrum data of each frame, and the spectrum data is a time-frequency domain signal; based on the first amplitude spectrum data of each frame, a convolution block attention module is used to determine the first amplitude spectrum data of each frame after dereverberation, and according to the first amplitude spectrum data of each frame after dereverberation, a self-attention channel combiner is used to determine the beam signal after beamforming; and then, according to the beam signal after beamforming, a target ASR network is used to identify the target category corresponding to the multi-channel signal waveform to be identified.
[0048] The present invention uses spherical harmonic transform SHT to extract spatial information from spectrum data corresponding to each frame of multi-channel signal, obtains the transformed signal, that is, the spectrum data corresponding to each frame of multi-channel signal after the number of channels is changed, and then sent to the convolution block attention module for dereverberation, and then sent to the self-attention channel combiner SACC. The self-attention channel combiner SACC uses the self-attention mechanism to beamform the signal, and then sends it to the ASR network for recognition. The amount of beamforming calculation is small. The present invention reduces the calculation complexity and amount of streaming speech recognition, improves the calculation speed, and is suitable for multi-channel speech signal processing in streaming scenarios.
[0049] According to a streaming multi-channel end-to-end speech recognition method based on spherical harmonic function transformation provided by the present invention, spherical harmonic transformation is performed on the spectrum data corresponding to each frame of multi-channel signal to obtain the spectrum data corresponding to each frame of multi-channel signal after the number of channels is transformed, including: Based on the order of the spherical harmonic transformation function, the number of channels after transformation is determined; Based on the transformed number of channels, determining the spatial information corresponding to the spectrum data corresponding to each frame of the multi-channel signal; Based on the spatial information corresponding to the frequency spectrum data corresponding to the multi-channel signal of each frame, the frequency spectrum data corresponding to the multi-channel signal of each frame after the number of channels is converted is determined.
[0050] Specifically, in some embodiments, the specific implementation process of performing spherical harmonic transformation on the spectrum data corresponding to each frame of the multi-channel signal to obtain the spectrum data corresponding to each frame of the multi-channel signal after the number of transformed channels in step 102 includes the following steps: First, based on the order of the spherical harmonic transformation function, the number of channels after transformation is determined. Among them, the order of the spherical harmonic transformation function determines the number and complexity of the spherical harmonic function. In practical applications, the order of spherical harmonics is limited because complex spherical functions need to be approximated by truncation. In practical applications, the appropriate order is usually selected according to the required accuracy and computing resource limitations. A common practice is to take the first few orders (such as 2nd or 3rd order) to approximate complex spherical functions. For example, the number of channels D after transformation is calculated by the following formula: D = (SHT order + 1)^2 Furthermore, based on the transformed number of channels, spatial information corresponding to the spectral data corresponding to each frame of the multi-channel signal can be determined, and based on the spatial information corresponding to the spectral data corresponding to each frame of the multi-channel signal, spectral data corresponding to each frame of the multi-channel signal after the number of channels is transformed can be determined.
[0051] Spherical harmonic transform (SHT) mainly captures the spatial distribution characteristics of the sound field. Based on the spherical harmonic theory, SHT identifies the spatial properties of signals from different directions and their mutual relationship on the microphone channel, which enables the subsequent neural network to make full use of the spatial information.
[0052] The method provided in this embodiment determines the number of channels after transformation based on the order of the spherical harmonic transformation function, determines the spatial information corresponding to the spectrum data corresponding to each frame of the multi-channel signal based on the number of channels after transformation, and then determines the spectrum data corresponding to each frame of the multi-channel signal after the number of channels is transformed based on the spatial information corresponding to the spectrum data corresponding to each frame of the multi-channel signal. This transformation can skillfully capture the spatial direction of the sound, thereby more accurately distinguishing the target speech and background noise. It can enable the subsequent neural network to make full use of the spatial information and improve the accuracy of recognition.
[0053] According to a streaming multi-channel end-to-end speech recognition method based on spherical harmonic function transformation provided by the present invention, the convolution block attention module includes a channel attention module and a spatial attention module; based on the first amplitude spectrum data of each frame, the convolution block attention module is used to determine the first amplitude spectrum data of each frame after reverberation, including: Input the first amplitude spectrum data of each frame into the channel attention module to obtain the channel attention matrix corresponding to the first amplitude spectrum data of each frame; Determine the second amplitude spectrum data of each frame based on the channel attention matrix and the first amplitude spectrum data of each frame; Input the second amplitude spectrum data of each frame into the spatial attention module to obtain the spatial attention matrix corresponding to the second amplitude spectrum data of each frame; Based on the spatial attention matrix and the second amplitude spectrum data of each frame, the first amplitude spectrum data of each frame after dereverberation is determined.
[0054] Specifically, in some embodiments, the convolutional block attention module CBAM includes a channel attention module (CAM) and a spatial attention module (SAM). Figure 4 It is one of the structural diagrams of the convolutional block attention module CBAM provided by the present invention, such as Figure 4 As shown, the first amplitude spectrum data of each frame of the input data first passes through the channel attention module CAM, and then passes through the spatial attention module SAM, and the obtained attention matrix is superimposed on the first amplitude spectrum data of each frame to obtain the first amplitude spectrum data of each frame after output dedeverberation.
[0055] Correspondingly, in step 103, based on the first amplitude spectrum data of each frame, a specific implementation process of using the convolution block attention module to determine the first amplitude spectrum data of each frame after dereverberation includes the following steps: First, the first amplitude spectrum data of each frame is input into the channel attention module CAM to obtain the channel attention matrix corresponding to the first amplitude spectrum data of each frame. For example, Figure 5 This is the second structural diagram of the convolutional block attention module CBAM provided by the present invention, such as Figure 5 As shown, the input feature map is input into the channel attention module CAM, and after the maximum pooling, average pooling, and sharing layer, the channel attention matrix M is finally obtained. C Then, based on the channel attention matrix and the first amplitude spectrum data of each frame, the second amplitude spectrum data of each frame is determined.
[0056] Furthermore, the second amplitude spectrum data of each frame is input into the spatial attention module SAM to obtain the spatial attention matrix corresponding to the second amplitude spectrum data of each frame, such as Figure 5 As shown in the figure, the second amplitude spectrum data of each frame superimposed with the channel attention matrix is input into the spatial attention module SAM, and after maximum pooling, average pooling, convolution, and linearization, the spatial attention matrix is finally obtained.
[0057] Then, based on the spatial attention matrix and the second amplitude spectrum data of each frame, the first amplitude spectrum data of each frame after dereverberation is determined. For example, the final attention matrix is superimposed on the first amplitude spectrum data of each frame to obtain the first amplitude spectrum data of each frame after dereverberation.
[0058] The method provided in this embodiment, the convolution block attention module includes a channel attention module and a spatial attention module, the first amplitude spectrum data of each frame is input into the channel attention module, the channel attention matrix corresponding to the first amplitude spectrum data of each frame is obtained, and the second amplitude spectrum data of each frame is determined based on the channel attention matrix and the first amplitude spectrum data of each frame; then, the second amplitude spectrum data of each frame is input into the spatial attention module, the spatial attention matrix corresponding to the second amplitude spectrum data of each frame is obtained, and finally the first amplitude spectrum data of each frame after dereverberation is determined based on the spatial attention matrix and the second amplitude spectrum data of each frame. These two modules capture the inter-channel dependency well to enhance the feature representation, and the subsequent speech recognition based on the beam signal after beamforming has a high accuracy.
[0059] According to a streaming multi-channel end-to-end speech recognition method based on spherical harmonic function transformation provided by the present invention, a beam signal after beamforming is determined by a self-attention channel combiner according to the first amplitude spectrum data of each frame after dereverberation, including: The first amplitude spectrum data of each frame after dereverberation is activated by using the fully connected function of the activation layer in the self-attention channel combiner to obtain the key vector, query vector and value vector corresponding to the first amplitude spectrum data of each frame after dereverberation; the dimension of the first amplitude spectrum data of each frame after dereverberation is determined based on the number of frames, channels and frequencies included in the multi-channel signal waveform to be identified; Determine the key self-attention matrix based on the query vector, the key vector, and the transformed number of channels; Based on the self-attention matrix and the value vector, determine the self-attention matrix of the value; Based on the self-attention matrix of the value and the first amplitude spectrum data of each frame after dereverberation, the beam signal after beamforming is determined.
[0060] Specifically, in some embodiments, the specific implementation process of step 103 determining the beam signal after beamforming by using the self-attention channel combiner according to the first amplitude spectrum data of each frame after dereverberation includes the following steps: First, the fully connected function of the activation layer in the self-attention channel combiner SACC is used to calculate the first amplitude spectrum data X of each frame after dereverberation. magActivate to obtain the key vector key, query vector query and value vector value corresponding to the first amplitude spectrum data of each frame after dereverberation, wherein the dimension of the first amplitude spectrum data of each frame after dereverberation is determined based on the number of frames T, the number of channels C and the number of frequencies F included in the multi-channel signal waveform to be identified.
[0061] Furthermore, the final beam signal after beamforming is obtained through the following processing: (1) Based on the query vector, key vector, and the transformed number of channels, determine the key self-attention matrix: Among them, X mag ∈R T×C×F ,query∈R T×C×D’ , key∈R T×C×D’ in, represents the self-attention matrix of the key, represents the query vector, represents the key vector, ' represents the embedding vector dimension obtained after the first amplitude spectrum data of each frame passes through the fully connected layer.
[0062] (2) Based on the self-attention matrix and the value vector, determine the self-attention matrix of the value: in, ∈R T×C×C , ∈R T×C×1 in, The self-attention matrix representing the value, represents the self-attention matrix of the key, Represents a vector of values.
[0063] (3) Based on the self-attention matrix of the value and the first amplitude spectrum data of each frame after dereverberation, the beam signal after beamforming is determined.
[0064] in, ∈R T×C×1 , X mag ∈R T×C×F ∈R T×F Represents the beam signal after beamforming.
[0065] The method provided in this embodiment, first, uses the fully connected function of the activation layer in the self-attention channel combiner to activate the first amplitude spectrum data of each frame after dereverberation, and obtains the key vector, query vector and value vector corresponding to the first amplitude spectrum data of each frame after dereverberation, wherein the dimension of the first amplitude spectrum data of each frame after dereverberation is determined based on the number of frames, channels and frequencies included in the multi-channel signal waveform to be identified; then, based on the query vector, the key vector and the number of channels after the transformation, the self-attention matrix of the key is determined, based on the self-attention matrix and the value vector, the self-attention matrix of the value is determined, based on the self-attention matrix of the value and the first amplitude spectrum data of each frame after dereverberation, the beam signal after beamforming is determined. The present invention uses SHT to express the spatial information of the sound source, so that the subsequent network can conveniently use the spatial information to process the signal. In addition, a series of deep learning modules are used to replace the complex mask-based MVDR beamforming module, effectively capturing the channel dependencies and spatial relationships that provide valuable context, and reducing the complexity of calculation.
[0066] According to a streaming multi-channel end-to-end speech recognition method based on spherical harmonic function transformation provided by the present invention, according to the beam signal after beamforming, the target category corresponding to the multi-channel signal waveform to be identified is obtained by using the target ASR network recognition, including: The beam signal after beamforming is post-filtered by using a multi-head self-attention module to obtain a filtered beam signal; The target ASR network is used to identify the target category corresponding to the multi-channel signal waveform to be identified based on the filtered beam signal.
[0067] Specifically, in some embodiments, the backend identification process in step 103 can be implemented by the following steps: First, the beam signal after beamforming is post-filtered using the Multi-Headed Self-Attention Module (MHSA) to obtain the filtered beam signal. It pays attention to the time information and acts as a filter in the frequency domain to filter out the residual noise.
[0068] Furthermore, the target ASR network is used to identify the target category corresponding to the multi-channel signal waveform to be identified based on the filtered beam signal.
[0069] The target ASR network is based on a conformer with convolution and self-attention mechanisms. The conformer combines the ability of convolution to model local information with the transformer. [1]The ability to model global information has achieved excellent results in speech recognition tasks. A conformer block mainly consists of a feed forward module (FFN), a convolution module (Conv), and a multi-headed self-attention module (MHSA), which will be introduced in detail below.
[0070] The feed-forward module sequentially connects a layer normalization layer, a linear layer, and a swish activation layer ( ), a dropout layer, a linear layer, and a dropout layer. The input and output use residual connections.
[0071] The convolution module includes a layer normalization layer, a convolution layer, a Glu (gated linear unit) activation layer, a one-dimensional depth convolution layer, a batch normalization layer, a swish activation layer, a convolution layer and a dropout layer. The input and output also use residual connections.
[0072] The multi-head self-attention module consists of a layer normalization layer, a multi-headed self-attention layer (MHSA) using relative positional embedding, and a dropout layer. The input and output also use residual connections. The definition of the multi-headed self-attention layer is as follows: in represents the mth head, 𝐐𝑚, 𝐊𝑚, 𝐕𝑚 represent the query vector Query, key vector Key and value vector Value corresponding to the mth head respectively, where are all parameters to be trained. 𝐐, 𝐊, 𝐕 are the inputs of MHSA. When MHSA is used as an encoder, 𝐐, 𝐊, 𝐕 have the same values, which are all input time-frequency features or output features of the previous conformer block. The attention in MHSA is defined as the dot product attention: A Conformer module is composed of the above three blocks. For the input 𝑥𝑖 of the i-th conformer block, its output 𝑦𝑖 is defined as follows: Based on the output of the last conformer block, the streaming recognition result, that is, the target category corresponding to the multi-channel signal waveform, is further determined to achieve streaming multi-channel end-to-end speech recognition.
[0073] The target ASR network is obtained by training two initial ASR networks with shared weights, such as Figure 2 As shown, the enhanced streaming signal and the enhanced whole sentence signal of the sample data are respectively input into the two initial ASR networks with shared weights, and the whole sentence recognition results and the streaming recognition results can be obtained respectively. Based on the whole sentence recognition results, the streaming recognition results and the true results (the labels corresponding to the sample data), the non-streaming loss and the streaming loss are calculated respectively. When the loss is within a certain range, it is determined that the model training is completed.
[0074] The method provided in this embodiment uses a multi-head self-attention module to perform post-filtering processing on the beam signal after beamforming to obtain the filtered beam signal, and uses the target ASR network to identify the target category corresponding to the multi-channel signal waveform to be identified based on the filtered beam signal. The accuracy of streaming speech recognition in the present invention is relatively high.
[0075] According to a streaming multi-channel end-to-end speech recognition method based on spherical harmonic function transformation provided by the present invention, before performing spherical harmonic transformation on the spectrum data corresponding to each frame of the multi-channel signal to obtain the spectrum data corresponding to each frame of the multi-channel signal after the number of channels is transformed, the method further includes: The multi-channel signal waveform to be identified is subjected to short-time Fourier transform frame by frame to obtain the spectrum data corresponding to each frame of the multi-channel signal.
[0076] Specifically, in some embodiments, before performing spherical harmonic transformation on the spectrum data corresponding to each frame of the multi-channel signal to obtain the spectrum data corresponding to each frame of the multi-channel signal after the number of transformed channels is obtained, the following steps are also included: The multi-channel signal waveform to be identified is subjected to short-time Fourier transform frame by frame to obtain the spectrum data corresponding to each frame of the multi-channel signal. For each data block, a certain number of frames on the left and right of the data block are spliced into context frames. These spliced frames are collectively referred to as context-sensitive blocks, which are fed into the front-end beamformer. The enhanced single-channel frame is then fed to the back-end ASR encoder. Note that the output of the ASR encoder of these context frames is not included in the calculation of the final ASR loss.
[0077] Among them, to perform short-time Fourier transform (STFT) on the recognized multi-channel signal waveform frame by frame to obtain the spectrum data corresponding to each frame of the multi-channel signal, the following specific implementation steps can be followed: (1) Signal preprocessing: Integrate multi-channel signals into a matrix form, where each row represents the signal data of one channel.
[0078] (2) Window function selection: Select an appropriate window function for the signal of each channel. Commonly used window functions include rectangular window, Hamming window and Hanning window. The choice of window function will affect the trade-off between time resolution and frequency resolution.
[0079] (3) Frame segmentation: Determine the window length (frame_length) and frame shift (frame_step). The window length determines the number of samples contained in each frame, while the frame shift determines the degree of overlap between frames. Usually, the window length is between 20ms and 40ms, and the frame step is between 10ms and 20ms.
[0080] (4) Windowing: Window processing is performed on the signal of each channel, that is, the signal is multiplied by the window function.
[0081] (5) Frame overlap: In order to reduce the discontinuity of the spectrum, there is usually a certain overlap between adjacent frames.
[0082] (6) Fast Fourier Transform (FFT): Perform FFT on each windowed frame to obtain the spectrum data of each frame.
[0083] (7) Spectrum calculation: Calculate the amplitude and phase of the spectrum of each frame. Usually we focus on the amplitude information.
[0084] (8) Construct the STFT matrix: Organize the spectral data of all frames into a two-dimensional matrix, where each column represents the spectral data of all channels at a time point.
[0085] The method provided in this embodiment performs short-time Fourier transform on the multi-channel signal waveform to be identified frame by frame to obtain spectrum data corresponding to each frame of the multi-channel signal, thereby converting the multi-channel input signal into complex spectrum features and then dividing the speech into non-overlapping language blocks.
[0086] The architecture provided by the present invention uses a waveform as input, and first transforms the signal into a time-frequency domain signal through STFT transformation, where B represents the batch size, C represents the number of channels, N represents the number of sampling points, T represents the number of frames, and F represents the frequency. Then use SHT to extract spatial information to obtain the transformed signal, where D represents the number of channels after the transformation, which is related to the order of SHT, D=(SHT order + 1)^2. Then it is sent to the CBAMs module, which uses multiple CBAM modules for cascading, and CBAM uses residual links to approximate the process of de-reverberation. Then it is sent to SACC, which uses a self-attention mechanism to approximate the beamforming process of the signal. Then it is sent to the MHSA module for post-filtering processing to filter out residual noise. Then it is converted to Fbank features, where M represents the order of the Mel filter. Then it is sent to the ASR network for recognition, and Num_class represents the possible class.
[0087] The following is a description of the streaming multi-channel end-to-end speech recognition device based on spherical harmonic function transformation provided by the present invention. The streaming multi-channel end-to-end speech recognition device based on spherical harmonic function transformation described below and the streaming multi-channel end-to-end speech recognition method based on spherical harmonic function transformation described above can be referenced to each other.
[0088] Figure 6 is a structural diagram of a streaming multi-channel end-to-end speech recognition device based on spherical harmonic function transformation provided by the present invention, such as Figure 6 As shown, the streaming multi-channel end-to-end speech recognition device 600 based on spherical harmonics transformation includes the following modules: An acquisition module 610 is used to acquire a multi-channel signal waveform to be identified; the multi-channel signal waveform to be identified includes at least two frames of multi-channel signals; The beamforming module 620 is used to perform spherical harmonic transformation on the spectrum data corresponding to each frame of the multi-channel signal to obtain the spectrum data corresponding to each frame of the multi-channel signal after the number of channels is transformed, and convert the spectrum data corresponding to each frame of the multi-channel signal after the number of channels is transformed into the first amplitude spectrum data of each frame; the spectrum data is a time-frequency domain signal; Based on the first amplitude spectrum data of each frame, a convolution block attention module is used to determine the first amplitude spectrum data of each frame after dereverberation, and according to the first amplitude spectrum data of each frame after dereverberation, a self-attention channel combiner is used to determine the beam signal after beamforming; The identification module 630 is used to obtain the target category corresponding to the multi-channel signal waveform to be identified by using the target ASR network according to the beam signal after the beamforming.
[0089] The device provided in this embodiment includes an acquisition module 610, which is used to acquire a multi-channel signal waveform to be identified, wherein the multi-channel signal waveform to be identified includes at least two frames of multi-channel signals; then, a beamforming module 620 is used to perform a spherical harmonic transform on the spectrum data corresponding to each frame of the multi-channel signal to obtain the spectrum data corresponding to each frame of the multi-channel signal after the number of channels is changed, and convert the spectrum data corresponding to each frame of the multi-channel signal after the number of channels is changed into the first amplitude spectrum data of each frame, and the spectrum data is a time-frequency domain signal; based on the first amplitude spectrum data of each frame, a convolution block attention module is used to determine the first amplitude spectrum data of each frame after dereverberation, and based on the first amplitude spectrum data of each frame after dereverberation, a self-attention channel combiner is used to determine the beam signal after beamforming; further, an identification module 630 is used to obtain the target category corresponding to the multi-channel signal waveform to be identified by using a target ASR network according to the beam signal after beamforming.
[0090] The present invention uses spherical harmonic transform SHT to extract spatial information from spectrum data corresponding to each frame of multi-channel signal, obtains the transformed signal, that is, the spectrum data corresponding to each frame of multi-channel signal after the number of channels is changed, and then sent to the convolution block attention module for dereverberation, and then sent to the self-attention channel combiner SACC. The self-attention channel combiner SACC uses the self-attention mechanism to beamform the signal, and then sends it to the ASR network for recognition. The amount of beamforming calculation is small. The present invention reduces the calculation complexity and amount of streaming speech recognition, improves the calculation speed, and is suitable for multi-channel speech signal processing in streaming scenarios.
[0091] According to a streaming multi-channel end-to-end speech recognition device 600 based on spherical harmonic function transformation provided by the present invention, the beamforming module 620 is specifically used for: Based on the order of the spherical harmonic transformation function, the number of channels after transformation is determined; Based on the number of channels after the transformation, determining the spatial information corresponding to the spectrum data corresponding to each of the frame multi-channel signals; Based on the spatial information corresponding to the frequency spectrum data corresponding to each of the frame multi-channel signals, the frequency spectrum data corresponding to each of the frame multi-channel signals after the number of channels is converted is determined.
[0092] According to a streaming multi-channel end-to-end speech recognition device 600 based on spherical harmonic function transformation provided by the present invention, the convolution block attention module includes a channel attention module and a spatial attention module; the beamforming module 620 is further used for: Input the first amplitude spectrum data of each frame into the channel attention module to obtain the channel attention matrix corresponding to the first amplitude spectrum data of each frame; Determine the second amplitude spectrum data of each frame based on the channel attention matrix and the first amplitude spectrum data of each frame; Input the second amplitude spectrum data of each frame into the spatial attention module to obtain a spatial attention matrix corresponding to the second amplitude spectrum data of each frame; Based on the spatial attention matrix and the second amplitude spectrum data of each of the frames, the first amplitude spectrum data of each of the frames after dereverberation is determined.
[0093] According to a streaming multi-channel end-to-end speech recognition device 600 based on spherical harmonic transformation provided by the present invention, the beamforming module 620 is further used for: Utilizing the fully connected function of the activation layer in the self-attention channel combiner, the first amplitude spectrum data of each frame after dereverberation is activated to obtain the key vector, query vector and value vector corresponding to the first amplitude spectrum data of each frame after dereverberation; the dimension of the first amplitude spectrum data of each frame after dereverberation is determined based on the number of frames, channels and frequencies included in the multi-channel signal waveform to be identified; Determining a self-attention matrix for a key based on the query vector, the key vector, and the transformed number of channels; Determining a self-attention matrix for a value based on the self-attention matrix and the value vector; Based on the self-attention matrix of the values and the first amplitude spectrum data of each frame after dereverberation, the beam signal after beamforming is determined.
[0094] According to a streaming multi-channel end-to-end speech recognition device 600 based on spherical harmonic function transformation provided by the present invention, the recognition module 630 is specifically used for: Post-filtering the beam signal after the beamforming using a multi-head self-attention module to obtain a filtered beam signal; The ASR network is used to identify the target category corresponding to the multi-channel signal waveform to be identified based on the filtered beam signal.
[0095] According to a streaming multi-channel end-to-end speech recognition device 600 based on spherical harmonic transformation provided by the present invention, the beamforming module 620 is also used to The multi-channel signal waveform to be identified is subjected to short-time Fourier transform frame by frame to obtain frequency spectrum data corresponding to each frame of the multi-channel signal.
[0096] Figure 7 An example of a physical structure diagram of an electronic device is shown in FIG. Figure 7As shown, the electronic device may include: a processor 710, a communication interface 720, a memory 730 and a communication bus 740, wherein the processor 710, the communication interface 720 and the memory 730 communicate with each other through the communication bus 740. The processor 710 may call the logic instructions in the memory 730 to execute a streaming multi-channel end-to-end speech recognition method based on spherical harmonic function transformation, the method comprising: Acquire a multi-channel signal waveform to be identified; the multi-channel signal waveform to be identified includes at least two frames of multi-channel signals; Performing spherical harmonic transformation on the spectrum data corresponding to each frame of the multi-channel signal to obtain the spectrum data corresponding to each frame of the multi-channel signal after the number of channels is transformed, and converting the spectrum data corresponding to each frame of the multi-channel signal after the number of channels is transformed into the first amplitude spectrum data of each frame; the spectrum data is a time-frequency domain signal; Based on the first amplitude spectrum data of each frame, a convolution block attention module is used to determine the first amplitude spectrum data of each frame after dereverberation, and according to the first amplitude spectrum data of each frame after dereverberation, a self-attention channel combiner is used to determine the beam signal after beamforming; According to the beam signal after beamforming, the target category corresponding to the multi-channel signal waveform to be identified is obtained by using the target ASR network recognition.
[0097] In addition, the logic instructions in the above-mentioned memory 730 can be implemented in the form of a software functional unit and can be stored in a computer-readable storage medium when it is sold or used as an independent product. Based on this understanding, the technical solution of the present invention can be essentially or partly embodied in the form of a software product that contributes to the prior art. The computer software product is stored in a storage medium, including several instructions to enable a computer device (which can be a personal computer, a server, or a network device, etc.) to perform all or part of the steps of the method described in each embodiment of the present invention. The aforementioned storage medium includes: U disk, mobile hard disk, read-only memory (ROM, Read-Only Memory), random access memory (RAM, Random Access Memory), disk or optical disk, etc. Various media that can store program codes.
[0098] On the other hand, the present invention further provides a computer program product, the computer program product comprising a computer program, the computer program can be stored on a non-transitory computer-readable storage medium, and when the computer program is executed by a processor, the computer can execute the streaming multi-channel end-to-end speech recognition method based on spherical harmonic function transformation provided by the above methods, the method comprising: Acquire a multi-channel signal waveform to be identified; the multi-channel signal waveform to be identified includes at least two frames of multi-channel signals; Performing spherical harmonic transformation on the spectrum data corresponding to each frame of the multi-channel signal to obtain the spectrum data corresponding to each frame of the multi-channel signal after the number of channels is transformed, and converting the spectrum data corresponding to each frame of the multi-channel signal after the number of channels is transformed into the first amplitude spectrum data of each frame; the spectrum data is a time-frequency domain signal; Based on the first amplitude spectrum data of each frame, a convolution block attention module is used to determine the first amplitude spectrum data of each frame after dereverberation, and according to the first amplitude spectrum data of each frame after dereverberation, a self-attention channel combiner is used to determine the beam signal after beamforming; According to the beam signal after beamforming, the target category corresponding to the multi-channel signal waveform to be identified is obtained by using the target ASR network recognition.
[0099] In another aspect, the present invention further provides a non-transitory computer-readable storage medium having a computer program stored thereon, which is implemented when the computer program is executed by a processor to perform the streaming multi-channel end-to-end speech recognition method based on spherical harmonic function transformation provided by the above methods, the method comprising: Acquire a multi-channel signal waveform to be identified; the multi-channel signal waveform to be identified includes at least two frames of multi-channel signals; Performing spherical harmonic transformation on the spectrum data corresponding to each frame of the multi-channel signal to obtain the spectrum data corresponding to each frame of the multi-channel signal after the number of channels is transformed, and converting the spectrum data corresponding to each frame of the multi-channel signal after the number of channels is transformed into the first amplitude spectrum data of each frame; the spectrum data is a time-frequency domain signal; Based on the first amplitude spectrum data of each frame, a convolution block attention module is used to determine the first amplitude spectrum data of each frame after dereverberation, and according to the first amplitude spectrum data of each frame after dereverberation, a self-attention channel combiner is used to determine the beam signal after beamforming; According to the beam signal after beamforming, the target category corresponding to the multi-channel signal waveform to be identified is obtained by using the target ASR network recognition.
[0100] The device embodiments described above are merely illustrative, wherein the units described as separate components may or may not be physically separated, and the components displayed as units may or may not be physical units, that is, they may be located in one place, or they may be distributed on multiple network units. Some or all of the modules may be selected according to actual needs to achieve the purpose of the scheme of this embodiment. Ordinary technicians in this field can understand and implement it without paying creative labor.
[0101] Through the description of the above implementation methods, those skilled in the art can clearly understand that each implementation method can be implemented by means of software plus a necessary general hardware platform, and of course, can also be implemented by hardware. Based on this understanding, the above technical solution is essentially or the part that contributes to the prior art can be embodied in the form of a software product, and the computer software product can be stored in a computer-readable storage medium, such as ROM / RAM, a disk, an optical disk, etc., including a number of instructions for a computer device (which can be a personal computer, a server, or a network device, etc.) to execute the methods described in each embodiment or some parts of the embodiments.
[0102] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit it. Although the present invention has been described in detail with reference to the aforementioned embodiments, those skilled in the art should understand that they can still modify the technical solutions described in the aforementioned embodiments, or make equivalent replacements for some of the technical features therein. However, these modifications or replacements do not deviate the essence of the corresponding technical solutions from the spirit and scope of the technical solutions of the embodiments of the present invention.
[0103] [1] Vaswani et al., "Attention Is All You Need".
Claims
1. A streaming multi-channel end-to-end speech recognition method based on spherical harmonic function transformation, characterized in that: include: Acquire a multi-channel signal waveform to be identified; the multi-channel signal waveform to be identified includes at least two frames of multi-channel signals; Performing spherical harmonic transformation on the spectrum data corresponding to each frame of the multi-channel signal to obtain the spectrum data corresponding to each frame of the multi-channel signal after the number of channels is transformed, and converting the spectrum data corresponding to each frame of the multi-channel signal after the number of channels is transformed into the first amplitude spectrum data of each frame; the spectrum data is a time-frequency domain signal; Based on the first amplitude spectrum data of each frame, a convolution block attention module is used to determine the first amplitude spectrum data of each frame after dereverberation, and according to the first amplitude spectrum data of each frame after dereverberation, a self-attention channel combiner is used to determine the beam signal after beamforming; According to the beam signal after beamforming, the target category corresponding to the multi-channel signal waveform to be identified is obtained by using the target ASR network recognition.
2. The streaming multi-channel end-to-end speech recognition method based on spherical harmonic function transformation according to claim 1 is characterized in that: The spherical harmonic transformation is performed on the spectrum data corresponding to each frame of the multi-channel signal to obtain the spectrum data corresponding to each frame of the multi-channel signal after the number of channels is transformed, including: Based on the order of the spherical harmonic transformation function, the number of channels after transformation is determined; Based on the number of channels after the transformation, determining the spatial information corresponding to the spectrum data corresponding to each of the frame multi-channel signals; Based on the spatial information corresponding to the frequency spectrum data corresponding to each of the frame multi-channel signals, the frequency spectrum data corresponding to each of the frame multi-channel signals after the number of channels is converted is determined.
3. The streaming multi-channel end-to-end speech recognition method based on spherical harmonic function transformation according to claim 1 is characterized in that: The convolutional block attention module includes a channel attention module and a spatial attention module; The method of determining the first amplitude spectrum data of each frame after dereverberation by using a convolution block attention module based on the first amplitude spectrum data of each frame comprises: Input the first amplitude spectrum data of each frame into the channel attention module to obtain the channel attention matrix corresponding to the first amplitude spectrum data of each frame; Determine the second amplitude spectrum data of each frame based on the channel attention matrix and the first amplitude spectrum data of each frame; Input the second amplitude spectrum data of each frame into the spatial attention module to obtain a spatial attention matrix corresponding to the second amplitude spectrum data of each frame; Based on the spatial attention matrix and the second amplitude spectrum data of each of the frames, the first amplitude spectrum data of each of the frames after dereverberation is determined.
4. The streaming multi-channel end-to-end speech recognition method based on spherical harmonics transform according to claim 2 is characterized in that: The step of determining the beam signal after beamforming by using a self-attention channel combiner according to the first amplitude spectrum data of each frame after dereverberation comprises: Utilizing the fully connected function of the activation layer in the self-attention channel combiner, the first amplitude spectrum data of each frame after dereverberation is activated to obtain the key vector, query vector and value vector corresponding to the first amplitude spectrum data of each frame after dereverberation; the dimension of the first amplitude spectrum data of each frame after dereverberation is determined based on the number of frames, channels and frequencies included in the multi-channel signal waveform to be identified; Determining a self-attention matrix for a key based on the query vector, the key vector, and the transformed number of channels; Determining a self-attention matrix for a value based on the self-attention matrix and the value vector; Based on the self-attention matrix of the values and the first amplitude spectrum data of each frame after dereverberation, the beam signal after beamforming is determined.
5. The streaming multi-channel end-to-end speech recognition method based on spherical harmonic function transformation according to claim 1 is characterized in that: The step of obtaining the target category corresponding to the multi-channel signal waveform to be identified by using the target ASR network according to the beam signal after the beamforming, comprises: Post-filtering the beam signal after the beamforming using a multi-head self-attention module to obtain a filtered beam signal; The ASR network is used to identify the target category corresponding to the multi-channel signal waveform to be identified based on the filtered beam signal.
6. The streaming multi-channel end-to-end speech recognition method based on spherical harmonic function transformation according to claim 1 is characterized in that: Before performing spherical harmonic transformation on the spectrum data corresponding to each of the frame multi-channel signals to obtain the spectrum data corresponding to each of the frame multi-channel signals after the number of transformed channels is obtained, the method further includes: The multi-channel signal waveform to be identified is subjected to short-time Fourier transform frame by frame to obtain frequency spectrum data corresponding to each frame of the multi-channel signal.
7. A streaming multi-channel end-to-end speech recognition device based on spherical harmonic function transformation, characterized in that: include: An acquisition module, used for acquiring a multi-channel signal waveform to be identified; the multi-channel signal waveform to be identified includes at least two frames of multi-channel signals; A beamforming module is used to perform spherical harmonic transformation on the spectrum data corresponding to each frame of the multi-channel signal to obtain the spectrum data corresponding to each frame of the multi-channel signal after the number of channels is transformed, and convert the spectrum data corresponding to each frame of the multi-channel signal after the number of channels is transformed into the first amplitude spectrum data of each frame; the spectrum data is a time-frequency domain signal; Based on the first amplitude spectrum data of each frame, a convolution block attention module is used to determine the first amplitude spectrum data of each frame after dereverberation, and according to the first amplitude spectrum data of each frame after dereverberation, a self-attention channel combiner is used to determine the beam signal after beamforming; The identification module is used to obtain the target category corresponding to the multi-channel signal waveform to be identified by using the target ASR network according to the beam signal after the beamforming.
8. An electronic device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein: When the processor executes the computer program, the streaming multi-channel end-to-end speech recognition method based on spherical harmonic transform as described in any one of claims 1 to 6 is implemented.
9. A non-transitory computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the streaming multi-channel end-to-end speech recognition method based on spherical harmonic transform as described in any one of claims 1 to 6 is implemented.
10. A computer program product, comprising a computer program, characterized in that When the computer program is executed by a processor, the streaming multi-channel end-to-end speech recognition method based on spherical harmonic transform as described in any one of claims 1 to 6 is implemented.
Citation Information
Patent Citations
End-to-end far-field speech recognition method and system
CN111179920A
Voice change determination method and device, storage medium and electronic device
CN116013279A
Complex de-reverberation speech enhancement method based on multi-head attention mechanism and Bi-LSTM
CN119107963A
System and method of pattern recognition in very high-dimensional space
US20020077817A1
Systems and vehicles that provide speech recognition system notifications
US20150006166A1