A streaming multi-channel end-to-end speech recognition method based on spherical harmonic function transform
By replacing the beamforming module in the CUSIDE-Array with the spherical harmonic transformation (SHT) and deep learning modules SACC and CBAM, the high computational complexity of multi-channel speech recognition systems in streaming scenarios is solved, thereby improving computational speed and recognition accuracy.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- CHINA NAT POSTAL & TELECOMM APPLIANCES CORP
- Filing Date
- 2024-12-31
- Publication Date
- 2026-07-31
AI Technical Summary
In existing technologies, the multi-channel speech recognition system CUSIDE-Array suffers from high computational complexity and slow processing speed in streaming scenarios, resulting in low efficiency in speech signal processing.
The Spherical Harmonic Transform (SHT) is used to replace the mask-based MVDR beamforming module in the CUSIDE-Array. Combined with the Self-Attention Channel Combiner (SACC) and the Convolutional Block Attention (CBAM) module, beamforming and recognition of multi-channel signals are performed.
It reduces the computational complexity and workload of streaming speech recognition, improves the processing speed, and enhances the robustness and accuracy of multi-channel speech signal processing.
Smart Images

Figure CN119993122B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of speech recognition technology, and in particular to a streaming multichannel end-to-end speech recognition method based on spherical harmonic function transformation. Background Technology
[0002] Multi-channel automatic speech recognition (ASR) systems have been continuously researched and improved due to their ability to enhance the robustness and accuracy of speech recognition through multi-channel input, especially in far-field acoustic environments. Typically, a beamforming front-end is introduced before the ASR back-end to utilize the spatial information of the multi-channel speech signals for speech enhancement.
[0003] In multi-channel speech signal recognition, the key is to make good use of the spatial information in the multi-channel signals and to effectively integrate and process spatial cues. For streaming recognition in multi-channel scenarios, existing technologies use the CUSIDE-Array (Chunking, Simulating Future Context, and Decoding-Array) method. The CUSIDE-Array method, based on the CUSIDE framework, uses a neural beamforming filter front-end and performs end-to-end training to achieve streaming multi-channel speech recognition. The beamforming front-end used by CUSIDE-Array is a mask-based minimum variance distortionless response (MVDR) beamforming filter, which estimates the mask for each channel separately, then calculates the spatial covariance matrix of the signal and noise, and substitutes it into the beamforming filter parameter estimation formula to obtain the parameters of the beamforming filter.
[0004] However, the mask-based MVDR beamforming module in CUSIDE-Array involves the calculation of multiple spatial covariance matrices and multiple matrix inversion operations when calculating the parameters of the beamforming filter. This results in high computational complexity and slow processing speed, making the CUSIDE-Array method unsuitable for speech signal processing in streaming scenarios. Summary of the Invention
[0005] This invention provides a streaming multi-channel end-to-end speech recognition method based on spherical harmonic function transformation, which solves the defects of the existing technology, such as high computational complexity and slow computation speed, which makes the CUSIDE-Array method unsuitable for speech signal processing in streaming scenarios. This invention reduces computational complexity and computational load, and improves computational speed.
[0006] In a first aspect, the present invention provides a streaming multi-channel end-to-end speech recognition method based on spherical harmonic function transform, the method comprising the following steps:
[0007] Acquire the multi-channel signal waveform to be identified; the multi-channel signal waveform to be identified includes at least two frames of multi-channel signal;
[0008] A spherical harmonic transform is performed on the spectral data corresponding to the multi-channel signal of each frame to obtain the spectral data corresponding to the multi-channel signal of each frame after the number of transformed channels, and the spectral data corresponding to the multi-channel signal of each frame after the number of transformed channels is converted into the first amplitude spectrum data of each frame; the spectral data is a time-frequency domain signal.
[0009] Based on the first amplitude spectrum data of each frame, the first amplitude spectrum data of each frame after de-reverberation is determined by the convolutional block attention module, and the beam signal after beamforming is determined by the self-attention channel combiner according to the first amplitude spectrum data of each frame after de-reverberation.
[0010] Based on the beam signal after beamforming, the target category corresponding to the multi-channel signal waveform to be identified is obtained using the target ASR network.
[0011] According to the present invention, a streaming multi-channel end-to-end speech recognition method based on spherical harmonic transformation is provided, wherein performing spherical harmonic transformation on the spectral data corresponding to each frame of multi-channel signals to obtain the spectral data corresponding to each frame of multi-channel signals after the number of transformed channels includes:
[0012] The number of channels after transformation is determined based on the order of the spherical harmonic transform function;
[0013] Based on the transformed number of channels, determine the spatial information corresponding to the spectral data of each frame multi-channel signal;
[0014] Based on the spatial information corresponding to the spectral data of each frame multi-channel signal, the spectral data of each frame multi-channel signal after determining the number of transformation channels is obtained.
[0015] According to the present invention, a streaming multi-channel end-to-end speech recognition method based on spherical harmonic function transform is provided, wherein the convolutional block attention module includes a channel attention module and a spatial attention module; the step of determining the de-reverberated first amplitude spectrum data of each frame based on the first amplitude spectrum data of each frame using the convolutional block attention module includes:
[0016] The first amplitude spectrum data of each frame is input into the channel attention module to obtain the channel attention matrix corresponding to the first amplitude spectrum data of each frame.
[0017] Based on the channel attention matrix and the first amplitude spectrum data of each frame, the second amplitude spectrum data of each frame is determined;
[0018] The second amplitude spectrum data of each frame is input into the spatial attention module to obtain the spatial attention matrix corresponding to the second amplitude spectrum data of each frame.
[0019] Based on the spatial attention matrix and the second amplitude spectrum data of each frame, the first amplitude spectrum data of each frame after de-reverberation is determined.
[0020] According to the present invention, a streaming multichannel end-to-end speech recognition method based on spherical harmonic function transform is provided, wherein determining the beam signal after beamforming using a self-attention channel combiner based on the first amplitude spectrum data of each frame after reverberation is removed includes:
[0021] The fully connected function of the activation layer in the self-attention channel combiner is used to activate the first amplitude spectrum data of each frame after dérification, thereby obtaining the key vector, query vector, and value vector corresponding to the first amplitude spectrum data of each frame after dérification; the dimension of the first amplitude spectrum data of each frame after dérification is determined based on the number of frames, the number of channels, and the number of frequencies included in the multi-channel signal waveform to be identified;
[0022] Based on the query vector, the key vector, and the transformed number of channels, determine the self-attention matrix of the key;
[0023] Based on the self-attention matrix and the value vector, determine the self-attention matrix of the value;
[0024] Based on the self-attention matrix of the value and the first amplitude spectrum data of each frame after déreverberation, the beam signal after beamforming is determined.
[0025] According to the present invention, a streaming multi-channel end-to-end speech recognition method based on spherical harmonic function transform is provided, wherein the step of identifying the target category corresponding to the multi-channel signal waveform to be recognized by using a target ASR network based on the beamformed beam signal includes:
[0026] The beam signal after beamforming is post-filtered using a multi-head self-attention module to obtain a filtered beam signal.
[0027] Using an ASR network, the target category corresponding to the multi-channel signal waveform to be identified is obtained based on the filtered beam signal.
[0028] According to the present invention, a streaming multi-channel end-to-end speech recognition method based on spherical harmonic transformation is provided, which, before performing spherical harmonic transformation on the spectral data corresponding to each frame of multi-channel signals to obtain the spectral data corresponding to each frame of multi-channel signals after the number of transformed channels, further includes:
[0029] The multi-channel signal waveform to be identified is subjected to frame-by-frame short-time Fourier transform to obtain the spectral data corresponding to each frame of the multi-channel signal.
[0030] Secondly, the present invention also provides a streaming multi-channel end-to-end speech recognition device based on spherical harmonic function transform, the device comprising the following modules:
[0031] An acquisition module is used to acquire the waveform of a multi-channel signal to be identified; the waveform of the multi-channel signal to be identified includes at least two frames of multi-channel signal.
[0032] The beamforming module is used to perform spherical harmonic transformation on the spectral data corresponding to the multi-channel signals of each frame to obtain the spectral data corresponding to the multi-channel signals of each frame after the number of transformed channels, and to convert the spectral data corresponding to the multi-channel signals of each frame after the number of transformed channels into the first amplitude spectrum data of each frame; the spectral data is a time-frequency domain signal.
[0033] Based on the first amplitude spectrum data of each frame, the first amplitude spectrum data of each frame after de-reverberation is determined by the convolutional block attention module, and the beam signal after beamforming is determined by the self-attention channel combiner according to the first amplitude spectrum data of each frame after de-reverberation.
[0034] The identification module is used to identify the target category corresponding to the multi-channel signal waveform to be identified by using the target ASR network based on the beam signal after beamforming.
[0035] Thirdly, the present invention also provides an electronic device, including a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor executes the computer program to implement the streaming multichannel end-to-end speech recognition method based on spherical harmonic function transformation as described above.
[0036] Fourthly, the present invention also provides a non-transitory computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements the streaming multichannel end-to-end speech recognition method based on spherical harmonic function transformation as described above.
[0037] Fifthly, the present invention also provides a computer program product, including a computer program that, when executed by a processor, implements the streaming multichannel end-to-end speech recognition method based on spherical harmonic function transformation as described above.
[0038] The present invention provides a streaming multi-channel end-to-end speech recognition method based on spherical harmonic transformation. First, it acquires the waveform of the multi-channel signal to be recognized, wherein the waveform includes at least two frames of multi-channel signal. Then, it performs spherical harmonic transformation on the spectral data corresponding to each frame of multi-channel signal to obtain the spectral data corresponding to each frame of multi-channel signal after the channel number transformation. The spectral data corresponding to each frame of multi-channel signal after the channel number transformation is then converted into first amplitude spectrum data for each frame, where the spectral data is a time-frequency domain signal. Based on the first amplitude spectrum data of each frame, a convolutional block attention module is used to determine the dedevered first amplitude spectrum data of each frame. Based on the dedevered first amplitude spectrum data of each frame, a self-attention channel combiner is used to determine the beam signal after beamforming. Finally, based on the beamformed beam signal, a target ASR network is used to identify the target category corresponding to the multi-channel signal waveform to be recognized.
[0039] This invention uses Spherical Harmonic Transform (SHT) to extract spatial information from the spectral data of each frame of multi-channel signals, obtaining the transformed signal, i.e., the spectral data of each frame of multi-channel signals after changing the number of channels. This data is then fed into a convolutional block attention module for de-reverberation, and then into a Self-Attention Channel Combiner (SACC). The SACC uses a self-attention mechanism to perform beamforming on the signal, which is then fed into an ASR network for recognition. Beamforming has a relatively low computational load. This invention reduces the computational complexity and workload of streaming speech recognition, improves the computational speed, and is suitable for multi-channel speech signal processing in streaming scenarios. Attached Figure Description
[0040] To more clearly illustrate the technical solutions in this invention or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are some embodiments of this invention. For those skilled in the art, other drawings can be obtained from these drawings without creative effort.
[0041] Figure 1 This is a flowchart illustrating the streaming multichannel end-to-end speech recognition method based on spherical harmonic function transformation provided by the present invention.
[0042] Figure 2 This is a schematic diagram of the framework of the streaming multichannel end-to-end speech recognition method based on spherical harmonic function transformation provided by the present invention.
[0043] Figure 3 This is a schematic diagram of the self-attention channel combiner SACC provided by the present invention.
[0044] Figure 4This is one of the structural schematic diagrams of the Convolutional Block Attention (CBAM) module provided by the present invention.
[0045] Figure 5 This is the second schematic diagram of the structure of the Convolutional Block Attention (CBAM) module provided by the present invention.
[0046] Figure 6 This is a schematic diagram of the structure of the streaming multichannel end-to-end speech recognition device based on spherical harmonic function transformation provided by the present invention.
[0047] Figure 7 This is a schematic diagram of the structure of the electronic device provided by the present invention. Detailed Implementation
[0048] To make the objectives, technical solutions, and advantages of this invention clearer, the technical solutions of this invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some, not all, of the embodiments of this invention. All other embodiments obtained by those skilled in the art based on the embodiments of this invention without creative effort are within the scope of protection of this invention.
[0049] To more clearly understand the various embodiments provided by the present invention, the technical content involved in the present invention will first be described as follows:
[0050] Compared to single-channel speech signals, multi-channel speech signals have advantages in signal processing due to their greater spatial information. However, in multi-channel speech signal recognition, effectively utilizing the spatial information from multiple channels and efficiently integrating and processing spatial cues remains a challenge.
[0051] Traditional spatial filtering methods include delay and beamforming, minimum variance distortionless response (MVDR) beamforming, and superdirectional beamforming. These methods utilize phase and timing differences between microphones to preferentially extract signals from certain directions. While these methods can work well, their performance depends on a reliable estimation of spatial information, which can be challenging under noisy conditions.
[0052] In recent years, deep learning has made great progress in multi-channel speech enhancement. Systems like MMUB leverage the complementary advantages of deep learning and array processing to achieve state-of-the-art performance in speech enhancement tasks. However, current streaming multi-channel end-to-end speech recognition systems like CUSIDE-Array suffer from the following problems:
[0053] 1. CUSIDE-Array directly connects the Short Time Fourier Transform (STFT) from each microphone as model input. They rely on the powerful modeling capabilities of neural networks to extract spatial information about the sound source. However, traditional STFT representations struggle to express the spatial information of the sound source.
[0054] 2. Mask-based MVDR beamforming (Beamforming module) in CUSIDE-Array: This model performs mask estimation for each channel individually. The beamforming effect is highly dependent on the accurate estimation of the mask. However, the independent per-channel processing cannot capture the inter-channel dependencies and spatial relationships that provide valuable context, which affects the subsequent identification of the structure.
[0055] 3. The MVDR method in CUSIDE-Array involves calculating multiple spatial covariance matrices and performing multiple matrix inversion operations when calculating beamforming coefficients. This results in high computational complexity and slow processing speed, making it unsuitable for speech signal processing in streaming scenarios.
[0056] To address the aforementioned shortcomings, this invention provides a streaming multi-channel end-to-end speech recognition method based on spherical harmonic function transformation, which reduces computational complexity, increases computational speed, and improves the accuracy of model recognition.
[0057] The following is combined Figures 1-7 This invention describes a streaming multichannel end-to-end speech recognition method based on spherical harmonic function transform.
[0058] Figure 1 This is one of the flowcharts illustrating the streaming multichannel end-to-end speech recognition method based on spherical harmonic function transform provided by the present invention, such as... Figure 1 As shown, the method includes the following steps:
[0059] Step 101: Obtain the waveform of the multi-channel signal to be identified; the waveform of the multi-channel signal to be identified includes at least two frames of multi-channel signal;
[0060] Specifically, it should be noted that the execution subject of this invention is an electronic device, used to realize streaming multi-channel end-to-end speech recognition, reduce the amount of computation, and improve the computation speed.
[0061] In this embodiment of the invention, in order to address the shortcomings of the CUSIDE-Array method in streaming multichannel end-to-end speech recognition systems, we provide a streaming multichannel end-to-end speech recognition method based on spherical harmonic function transformation. Figure 2 This is a schematic diagram of the framework of the streaming multichannel end-to-end speech recognition method based on spherical harmonic function transform provided by the present invention, as shown below. Figure 2As shown, this invention uses Spherical Harmonic Transformation (SHT) plus a series of deep learning modules such as Self-Attention Channel Combinator (SACC) to replace the mask-based MVDR beamforming module, achieving good results. Figure 3 This is a schematic diagram of the self-attention channel combiner SACC provided by the present invention, as shown below. Figure 3 As shown, a series of deep learning modules include SHT transform, amplitude spectrum transform, CBAM, SACC, and MESA.
[0062] The method in this embodiment is implemented through the following steps:
[0063] First, acquire the waveform of the multi-channel signal to be identified, which includes at least two frames of multi-channel signal.
[0064] Before beamforming, the waveform of the multi-channel signal to be identified is first subjected to a Short-Time Fourier Transform (STFT) to transform the signal into a time-frequency domain signal, i.e., spectral data. Here, B represents the batch size, C represents the number of channels, N represents the number of sampling points, T represents the number of frames, and F represents the frequency.
[0065] Step 102: Perform spherical harmonic transformation on the spectral data corresponding to the multi-channel signals of each frame to obtain the spectral data corresponding to the multi-channel signals of each frame after the number of channels has been transformed, and convert the spectral data corresponding to the multi-channel signals of each frame after the number of channels has been transformed into the first amplitude spectrum data of each frame; the spectral data is a time-frequency domain signal;
[0066] Specifically, after obtaining the spectral data (which is a time-frequency domain signal) corresponding to the multi-channel signal of each frame, this invention introduces a method that uses the order of Spherical Harmonic Transformation (SHT) as an auxiliary model input.
[0067] In this embodiment, the spectral data corresponding to the multi-channel signal of each frame are subjected to spherical harmonic transformation (SHT) based on the order of the Spherical Harmonic Transform (SHT), resulting in the spectral data corresponding to the multi-channel signal of each frame after the number of transformed channels. The order of the SHT clearly represents the spatial distribution. It has two advantages:
[0068] 1. Effective Capture of Spatial Information: Unlike the Short Time Fourier Transform (STFT), the Spherical Harmonic Transform (SHT) primarily captures the spatial distribution characteristics of the sound field. Based on spherical harmonic theory, SHT identifies the spatial properties of signals from different directions and their interrelationships in the microphone channels. This spatial capture is crucial for multi-microphone speech enhancement. In multi-microphone arrays, this transform can skillfully capture the spatial direction of sound, thereby more accurately distinguishing target speech from background noise.
[0069] 2. Improved Spatial Resolution: Spherical harmonics (SHCs) form the complete basis for functions defined on a sphere. Therefore, any spherical function can be precisely represented as a linear combination of spherical harmonics. SHTs help to accurately describe the spatial distribution of the sound field. The order of the spherical harmonics determines the granularity of spatial feature capture. While lower orders describe a wide range of spatial patterns, higher orders describe more subtle spatial differences. By choosing an appropriate order, the desired spatial resolution can be achieved. If the spatial information in these SHCs can be fully utilized, it may help compensate for the shortcomings of descriptive spatial modeling in current mainstream methods. Utilizing this can significantly improve the performance and robustness of multichannel speech enhancement.
[0070] Furthermore, the spectral data corresponding to the multi-channel signals of each frame after changing the number of channels can be converted into the first amplitude spectrum data (Magnitude) of each frame, which is convenient for subsequent deep learning modules to process and obtain beam signals. Deep learning modules include, for example, a Convolutional Block Attention Module (CBAM) and a Self-Attention Channel Combinator (SACC).
[0071] Step 103: Based on the first amplitude spectrum data of each frame, use the convolutional block attention module to determine the first amplitude spectrum data of each frame after dération, and use the self-attention channel combiner to determine the beam signal after beamforming based on the first amplitude spectrum data of each frame after dération.
[0072] Specifically, after obtaining the first amplitude spectrum data of each frame through conversion, the first amplitude spectrum data of each frame after dreverberation is first determined using a Convolutional Block Attention (CBAM) module. The CBAM module can effectively capture the spatial dependencies between channels to enhance feature representation. This invention uses four cascaded CBAM layers to simulate the dreverberation process.
[0073] Furthermore, the beam signal after beamforming can be determined using the self-attention channel combiner (SACC) based on the first amplitude spectrum data of each frame after reverberation removal. The calculation process of the SACC can be approximated by the MVDR beamforming process, where the self-attention matrix w...att Since the dimensions of the value matrix are similar to those of the matrix required for calculating the coefficients of the beamforming filter, the achieved results are similar.
[0074] Step 104: Based on the beam signal after beamforming, use the target ASR network to identify the target category corresponding to the multi-channel signal waveform to be identified.
[0075] Specifically, after the above processing, the beam signal S after beamforming can be obtained. Further, the beam signal after beamforming is sent to the target multi-channel automatic speech recognition (ASR) network, and the target category corresponding to the multi-channel signal waveform to be recognized is obtained by using the target ASR network.
[0076] For example, the target ASR network includes a conformer based on convolution and self-attention mechanisms. The conformer combines the ability of convolution to model local information with the ability of transformer to model global information, achieving excellent results in speech recognition tasks. A conformer block mainly consists of a feed-forward module (FFN), a convolution module (Conv), and a multi-head self-attention module (MHSA). The recognition result output by the last conformer is used to determine the target category corresponding to the multi-channel signal waveform to be recognized.
[0077] The method provided in this embodiment first acquires the multi-channel signal waveform to be identified, wherein the multi-channel signal waveform to be identified includes at least two frames of multi-channel signals; then, spherical harmonic transformation is performed on the spectral data corresponding to each frame of multi-channel signals to obtain the spectral data corresponding to each frame of multi-channel signals after the number of channels is transformed, and the spectral data corresponding to each frame of multi-channel signals after the number of channels is transformed is converted into the first amplitude spectrum data of each frame, the spectral data being a time-frequency domain signal; based on the first amplitude spectrum data of each frame, the first amplitude spectrum data of each frame after dération is determined using a convolutional block attention module, and based on the first amplitude spectrum data of each frame after dération, the beam signal after beamforming is determined using a self-attention channel combiner; furthermore, based on the beam signal after beamforming, the target category corresponding to the multi-channel signal waveform to be identified is obtained using a target ASR network.
[0078] This invention uses Spherical Harmonic Transform (SHT) to extract spatial information from the spectral data of each frame of multi-channel signals, obtaining the transformed signal, i.e., the spectral data of each frame of multi-channel signals after changing the number of channels. This data is then fed into a convolutional block attention module for de-reverberation, and then into a Self-Attention Channel Combiner (SACC). The SACC uses a self-attention mechanism to perform beamforming on the signal, which is then fed into an ASR network for recognition. Beamforming has a relatively low computational load. This invention reduces the computational complexity and workload of streaming speech recognition, improves the computational speed, and is suitable for multi-channel speech signal processing in streaming scenarios.
[0079] According to the present invention, a streaming multi-channel end-to-end speech recognition method based on spherical harmonic transformation is provided, which performs spherical harmonic transformation on the spectral data corresponding to the multi-channel signals of each frame to obtain the spectral data corresponding to the multi-channel signals of each frame after the number of transformed channels, including:
[0080] The number of channels after transformation is determined based on the order of the spherical harmonic transform function;
[0081] Based on the transformed number of channels, determine the spatial information corresponding to the spectral data of the multi-channel signal in each frame;
[0082] Based on the spatial information corresponding to the spectral data of the multi-channel signal in each frame, the spectral data of the multi-channel signal in each frame after the number of transformed channels is determined.
[0083] Specifically, in some embodiments, the specific implementation process of performing spherical harmonic transformation on the spectral data corresponding to each frame of multi-channel signals to obtain the number of transformed channels in step 102 includes the following steps:
[0084] First, the number of channels after transformation is determined based on the order of the spherical harmonic transform function. The order of the spherical harmonic transform function determines the number and complexity of the spherical harmonic function. In practical applications, the order of the spherical harmonics is finite because truncation is needed to approximate complex spherical functions. In practice, an appropriate order is usually chosen based on the required accuracy and computational resource constraints. A common practice is to use the first few orders (such as 2nd or 3rd order) to approximate complex spherical functions. For example, the number of channels D after transformation is calculated using the following formula:
[0085] D = (SHT order + 1)^2
[0086] Furthermore, based on the transformed number of channels, the spatial information corresponding to the spectral data of each frame of multi-channel signals can be determined, and based on the spatial information corresponding to the spectral data of each frame of multi-channel signals after the transformation of the number of channels, the spectral data corresponding to each frame of multi-channel signals can be determined.
[0087] Spherical Harmonic Transform (SHT) primarily captures the spatial distribution characteristics of the sound field. Based on spherical harmonic theory, SHT identifies the spatial properties of signals from different directions and their interrelationships in the microphone channel, allowing subsequent neural networks to fully utilize spatial information.
[0088] The method provided in this embodiment determines the number of channels after transformation based on the order of the spherical harmonic transform function. Based on the number of channels after transformation, it determines the spatial information corresponding to the spectral data of each frame of multi-channel signals. Then, based on the spatial information corresponding to the spectral data of each frame of multi-channel signals, it determines the spectral data corresponding to each frame of multi-channel signals after the channel number transformation. This transformation can skillfully capture the spatial direction of sound, thereby more accurately distinguishing target speech from background noise. This allows subsequent neural networks to fully utilize spatial information, improving recognition accuracy.
[0089] According to the present invention, a streaming multi-channel end-to-end speech recognition method based on spherical harmonic function transform is provided. The convolutional block attention module includes a channel attention module and a spatial attention module. Based on the first amplitude spectrum data of each frame, the convolutional block attention module determines the de-reverberated first amplitude spectrum data of each frame, including:
[0090] The first amplitude spectrum data of each frame is input into the channel attention module to obtain the channel attention matrix corresponding to the first amplitude spectrum data of each frame;
[0091] Based on the channel attention matrix and the first amplitude spectrum data of each frame, determine the second amplitude spectrum data of each frame;
[0092] Input the second amplitude spectrum data of each frame into the spatial attention module to obtain the spatial attention matrix corresponding to the second amplitude spectrum data of each frame;
[0093] Based on the spatial attention matrix and the second amplitude spectrum data of each frame, the first amplitude spectrum data of each frame after dreverberation is determined.
[0094] Specifically, in some embodiments, the convolutional block attention module (CBAM) includes a channel attention module (CAM) and a spatial attention module (SAM). Figure 4 This is one of the structural schematic diagrams of the Convolutional Block Attention (CBAM) module provided by the present invention, as shown below. Figure 4 As shown, the first amplitude spectrum data of each frame of input data first passes through the channel attention module CAM, and then through the spatial attention module SAM. The resulting attention matrix is superimposed on the first amplitude spectrum data of each frame to obtain the first amplitude spectrum data of each frame after de-reverberation.
[0095] Correspondingly, the specific implementation process of determining the de-reverberated first amplitude spectrum data of each frame based on the first amplitude spectrum data of each frame using the convolutional block attention module in step 103 includes the following steps:
[0096] First, the first amplitude spectrum data of each frame is input into the channel attention module (CAM) to obtain the channel attention matrix corresponding to the first amplitude spectrum data of each frame. For example, Figure 5 This is the second structural schematic diagram of the Convolutional Block Attention (CBAM) module provided by the present invention, as shown below. Figure 5 As shown, the input feature map is input into the channel attention module CAM, and after max pooling, average pooling, and a shared layer, the channel attention matrix M is finally obtained. C Then, based on the channel attention matrix and the first amplitude spectrum data of each frame, the second amplitude spectrum data of each frame is determined.
[0097] Furthermore, the second amplitude spectrum data of each frame is input into the spatial attention module (SAM) to obtain the spatial attention matrix corresponding to the second amplitude spectrum data of each frame, such as... Figure 5 As shown, the second amplitude spectrum data of each frame, which is superimposed with the channel attention matrix, is input into the spatial attention module SAM. After max pooling, average pooling, convolution, and linearization, the spatial attention matrix is finally obtained.
[0098] Furthermore, based on the spatial attention matrix and the second amplitude spectrum data of each frame, the first amplitude spectrum data of each frame after dredging is determined. For example, the final attention matrix is superimposed on the first amplitude spectrum data of each frame to obtain the output first amplitude spectrum data of each frame after dredging.
[0099] The method provided in this embodiment includes a convolutional block attention module comprising a channel attention module and a spatial attention module. First amplitude spectrum data of each frame is input to the channel attention module to obtain the channel attention matrix corresponding to the first amplitude spectrum data of each frame. Based on the channel attention matrix and the first amplitude spectrum data of each frame, second amplitude spectrum data of each frame is determined. Then, the second amplitude spectrum data of each frame is input to the spatial attention module to obtain the spatial attention matrix corresponding to the second amplitude spectrum data of each frame. Finally, based on the spatial attention matrix and the second amplitude spectrum data of each frame, the déreverberated first amplitude spectrum data of each frame is determined. These two modules effectively capture the inter-channel dependencies to enhance feature representation, resulting in high accuracy for subsequent speech recognition based on the beamformed signal.
[0100] According to the present invention, a streaming multichannel end-to-end speech recognition method based on spherical harmonic function transform is provided, which determines the beam signal after beamforming based on the first amplitude spectrum data of each frame after reverberation removal using a self-attention channel combiner, including:
[0101] The fully connected function of the activation layer in the self-attention channel combiner is used to activate the first amplitude spectrum data of each frame after dérification, so as to obtain the key vector, query vector and value vector corresponding to the first amplitude spectrum data of each frame after dérification; the dimension of the first amplitude spectrum data of each frame after dérification is determined based on the number of frames, the number of channels and the number of frequencies of the multi-channel signal waveform to be identified.
[0102] The self-attention matrix of the key is determined based on the query vector, the key vector, and the number of channels after transformation.
[0103] Determine the self-attention matrix of the value based on the self-attention matrix and the value vector;
[0104] Based on the self-attention matrix of the value and the first amplitude spectrum data of each frame after dreverberation, the beam signal after beamforming is determined.
[0105] Specifically, in some embodiments, step 103, which uses a self-attention channel combiner to determine the beam signal after beamforming based on the first amplitude spectrum data of each frame after reverberation, includes the following steps:
[0106] First, using the fully connected function of the activation layer in the self-attention channel combiner (SACC), the first amplitude spectrum data X of each frame after reverberation is processed. mag Activation is performed to obtain the key vector, query vector, and value vector corresponding to the first amplitude spectrum data of each frame after dération. The dimension of the first amplitude spectrum data of each frame after dération is determined based on the number of frames T, the number of channels C, and the number of frequencies F of the multi-channel signal waveform to be identified.
[0107] Furthermore, the final beam signal after beamforming is obtained through the following processing:
[0108] (1) Determine the self-attention matrix of the key based on the query vector, the key vector, and the number of transformed channels:
[0109] Among them, X mag ∈R T×C×F ,query∈R T×C×D’ , key∈R T×C×D’
[0110]
[0111] in, The self-attention matrix represents the key. Represents the query vector. Represents the key vector. ' represents the dimension of the embedding vector obtained after the first amplitude spectrum data of each frame passes through the fully connected layer.
[0112] (2) Determine the self-attention matrix of the value based on the self-attention matrix and the value vector:
[0113] in, ∈R T×C×C , ∈R T×C×1
[0114]
[0115] in, The self-attention matrix represents the values. The self-attention matrix represents the key. Represents a value vector.
[0116] (3) Based on the self-attention matrix of the value and the first amplitude spectrum data of each frame after de-reverberation, determine the beam signal after beamforming.
[0117] in, ∈R T×C×1 X mag ∈R T×C×F
[0118]
[0119] ∈R T×F This indicates the beam signal after beamforming.
[0120] The method provided in this embodiment firstly activates the first amplitude spectrum data of each frame after déresonance using the fully connected function of the activation layer in the self-attention channel combiner, obtaining the key vector, query vector, and value vector corresponding to the first amplitude spectrum data of each frame after déresonance. The dimension of the first amplitude spectrum data of each frame after déresonance is determined based on the number of frames, channels, and frequencies included in the multi-channel signal waveform to be identified. Then, based on the query vector, key vector, and the number of transformed channels, the self-attention matrix of the key is determined; based on the self-attention matrix and value vector, the self-attention matrix of the value is determined; and based on the self-attention matrix of the value and the first amplitude spectrum data of each frame after déresonance, the beam signal after beamforming is determined. This invention uses SHT to represent the spatial information of the sound source, enabling subsequent networks to conveniently utilize spatial information to process the signal. Furthermore, a series of deep learning modules are used to replace the complex mask-based MVDR beamforming module, effectively capturing inter-channel dependencies and spatial relationships that provide valuable context, and reducing computational complexity.
[0121] According to the present invention, a streaming multi-channel end-to-end speech recognition method based on spherical harmonic function transform is provided. Based on the beamformed signal, a target ASR network is used to identify the target category corresponding to the multi-channel signal waveform to be recognized, including:
[0122] The beam signal after beamforming is post-filtered using a multi-head self-attention module to obtain the filtered beam signal.
[0123] Using the target ASR network, the target category corresponding to the multi-channel signal waveform to be identified is obtained based on the filtered beam signal.
[0124] Specifically, in some embodiments, the backend identification process in step 103 can be implemented through the following steps:
[0125] First, a multi-headed self-attention module (MHSA) is used to perform post-filtering on the beamformed signal, resulting in a filtered beamform. This process incorporates time information and acts as a filter in the frequency domain to remove residual noise.
[0126] Furthermore, using the target ASR network, the target category corresponding to the multi-channel signal waveform to be identified is obtained based on the filtered beam signal.
[0127] The target ASR network is based on a convolutional and self-attention mechanism called a Conformer. The Conformer combines the ability of convolution to model local information with that of a transformer. [1] Its ability to model global information has yielded excellent results in speech recognition tasks. A conformer block mainly consists of a feed-forward module (FFN), a convolution module (Conv), and a multi-head self-attention module (MHSA), which will be described in detail below.
[0128] The feedforward module successively connects a layer normalization layer, a linear layer, and a swish activation layer. The system consists of one dropout layer, one linear layer, and one dropout layer. Residual connections are used for the input and output.
[0129] The convolutional module consists of a layer normalization layer, a convolutional layer, a Glu (gated linear unit) activation layer, a one-dimensional depthwise convolutional layer, a batch normalization layer, a swish activation layer, another convolutional layer, and a dropout layer. Residual connections are also used for both input and output.
[0130] The multi-head self-attention module consists of a layer normalization layer, a multi-head self-attention layer (MHSA) using relative positional embedding, and a dropout layer. Residual connections are also used between the input and output. The multi-head self-attention layer is defined as follows:
[0131]
[0132]
[0133]
[0134] in Let represent the m-th header. 𝐐𝑚, 𝐊𝑚, and 𝐕𝑚 represent the query vector Query, key vector Key, and value vector Value corresponding to the m-th header, respectively. These are all parameters to be trained. 𝐐, 𝐊, and 𝐕 are the inputs to MHSA. When MHSA is used as the encoder, 𝐐, 𝐊, and 𝐕 have the same value, representing either the spectral features at the input or the output features of the previous conformer block. The attention mechanism in MHSA is defined as dot product attention.
[0135]
[0136] A Conformer module is obtained by combining the above three blocks. For the input 𝑥𝑖 of the i-th Conformer block, its output 𝑦𝑖 is defined as follows:
[0137]
[0138]
[0139]
[0140]
[0141] Based on the output of the last conformer block, the streaming recognition result is further determined, that is, the target category corresponding to the multi-channel signal waveform, thus realizing streaming multi-channel end-to-end speech recognition.
[0142] The target ASR network is obtained by training two initial ASR networks with shared weights, such as... Figure 2As shown, the enhanced streaming signal and the enhanced whole sentence signal of the sample data are input into two initial ASR networks with shared weights, respectively, to obtain whole sentence recognition results and streaming recognition results. Based on the whole sentence recognition results, streaming recognition results and the real results (the labels corresponding to the sample data), the non-streaming loss and streaming loss are calculated respectively until the loss is within a certain range, then the model training is considered complete.
[0143] The method provided in this embodiment uses a multi-head self-attention module to perform post-filtering processing on the beam signal after beamforming to obtain a filtered beam signal. Using a target ASR network, the target category corresponding to the multi-channel signal waveform to be identified is obtained based on the filtered beam signal. The accuracy of streaming speech recognition in this invention is relatively high.
[0144] According to the streaming multi-channel end-to-end speech recognition method based on spherical harmonic transformation provided by the present invention, before performing spherical harmonic transformation on the spectral data corresponding to the multi-channel signals of each frame to obtain the spectral data corresponding to the multi-channel signals of each frame after the number of transformed channels, the method further includes:
[0145] The multi-channel signal waveform to be identified is subjected to frame-by-frame short-time Fourier transform to obtain the spectral data corresponding to the multi-channel signal in each frame.
[0146] Specifically, in some embodiments, before obtaining the number of transformed channels by performing spherical harmonic transformation on the spectral data corresponding to the multi-channel signals of each frame, the following steps are also included:
[0147] The multi-channel signal waveform to be identified is subjected to frame-by-frame short-time Fourier transform to obtain the spectral data corresponding to each frame of the multi-channel signal. For each data block, a certain number of frames to the left and right of the data block are concatenated into context frames. These concatenated frames are collectively referred to as context-sensitive blocks, which are fed into the front-end beamformer. The enhanced single-channel frames are then fed into the back-end ASR encoder. Note that the output of the ASR encoder for these context frames is not included in the final ASR loss calculation.
[0148] To obtain the spectral data corresponding to each frame of the identified multi-channel signal waveforms by performing frame-by-frame Short Time Fourier Transform (STFT), the following specific implementation steps can be followed:
[0149] (1) Signal preprocessing: integrate multi-channel signals into a matrix form, with each row representing the signal data of one channel.
[0150] (2) Window function selection: Select an appropriate window function for each channel signal. Commonly used window functions include rectangular window, Hamming window, and Hanning window. The choice of window function will affect the trade-off between time resolution and frequency resolution.
[0151] (3) Frame segmentation: Determine the window length (frame_length) and frame shift (frame_step). The window length determines the number of samples contained in each frame, while the frame shift determines the degree of overlap between frames. Typically, the window length is between 20ms and 40ms, and the frame shift is between 10ms and 20ms.
[0152] (4) Windowing: Window the signal of each channel, that is, multiply the signal by a window function.
[0153] (5) Frame overlap: In order to reduce the discontinuity of the spectrum, there is usually some overlap between adjacent frames.
[0154] (6) Fast Fourier Transform (FFT): Perform FFT on each windowed frame to obtain the spectral data of each frame.
[0155] (7) Spectrum calculation: Calculate the amplitude and phase of the spectrum of each frame. Usually, we are concerned with the amplitude information.
[0156] (8) Construct the STFT matrix: Organize the spectrum data of all frames into a two-dimensional matrix, where each column represents the spectrum data of all channels at a time point.
[0157] The method provided in this embodiment performs frame-by-frame short-time Fourier transform on the multi-channel signal waveform to be identified, and obtains the spectral data corresponding to each frame of multi-channel signal. This realizes the conversion of multi-channel input signals into complex spectral features, and then divides speech into non-overlapping speech blocks.
[0158] The architecture provided by this invention uses waveforms as input. First, the signal undergoes STFT transformation to convert it into a time-frequency domain signal, where B represents the batch size, C represents the number of channels, N represents the number of sampling points, T represents the number of frames, and F represents the frequency. Then, SHT is used to extract spatial information, obtaining the transformed signal, where D represents the number of transformed channels, which is related to the order of SHT: D = (SHT order + 1)^2. This signal is then fed into a CBAMs module, which uses multiple cascaded CBAMs, employing residual linking to approximate the dereverberation process. Next, it is fed into SACC, which uses a self-attention mechanism to approximate the beamforming process. Then, it is fed into an MHSA module for post-filtering to remove residual noise. Finally, it is converted into F-bank features, where M represents the order of the Mel filter. Finally, it is fed into an ASR network for identification, where Num_class represents the possible classes.
[0159] The following describes the streaming multi-channel end-to-end speech recognition device based on spherical harmonic function transformation provided by the present invention. The streaming multi-channel end-to-end speech recognition device based on spherical harmonic function transformation described below can be referred to in correspondence with the streaming multi-channel end-to-end speech recognition method based on spherical harmonic function transformation described above.
[0160] Figure 6 This is a schematic diagram of the structure of the streaming multi-channel end-to-end speech recognition device based on spherical harmonic function transform provided by the present invention, as shown below. Figure 6 As shown, the streaming multichannel end-to-end speech recognition device 600 based on spherical harmonic function transform includes the following modules:
[0161] The acquisition module 610 is used to acquire the multi-channel signal waveform to be identified; the multi-channel signal waveform to be identified includes at least two frames of multi-channel signal.
[0162] The beamforming module 620 is used to perform spherical harmonic transformation on the spectral data corresponding to the multi-channel signals of each frame to obtain the spectral data corresponding to the multi-channel signals of each frame after the number of transformed channels, and to convert the spectral data corresponding to the multi-channel signals of each frame after the number of transformed channels into the first amplitude spectrum data of each frame; the spectral data is a time-frequency domain signal.
[0163] Based on the first amplitude spectrum data of each frame, the first amplitude spectrum data of each frame after de-reverberation is determined by the convolutional block attention module, and the beam signal after beamforming is determined by the self-attention channel combiner according to the first amplitude spectrum data of each frame after de-reverberation.
[0164] The identification module 630 is used to identify the target category corresponding to the multi-channel signal waveform to be identified by using the target ASR network based on the beam signal after beamforming.
[0165] The apparatus provided in this embodiment includes an acquisition module 610 for acquiring a multi-channel signal waveform to be identified, wherein the multi-channel signal waveform to be identified includes at least two frames of multi-channel signals; then, a beamforming module 620 for performing spherical harmonic transformation on the spectral data corresponding to each frame of multi-channel signals to obtain spectral data corresponding to each frame of multi-channel signals after the number of channels has been transformed, and converting the spectral data corresponding to each frame of multi-channel signals after the number of channels has been transformed into first amplitude spectrum data for each frame, wherein the spectral data is a time-frequency domain signal; based on the first amplitude spectrum data for each frame, a convolutional block attention module is used to determine the first amplitude spectrum data for each frame after dération, and based on the first amplitude spectrum data for each frame after dération, a self-attention channel combiner is used to determine the beam signal after beamforming; furthermore, an identification module 630 is used to identify the target category corresponding to the multi-channel signal waveform to be identified using a target ASR network based on the beam signal after beamforming.
[0166] This invention uses Spherical Harmonic Transform (SHT) to extract spatial information from the spectral data of each frame of multi-channel signals, obtaining the transformed signal, i.e., the spectral data of each frame of multi-channel signals after changing the number of channels. This data is then fed into a convolutional block attention module for de-reverberation, and then into a Self-Attention Channel Combiner (SACC). The SACC uses a self-attention mechanism to perform beamforming on the signal, which is then fed into an ASR network for recognition. Beamforming has a relatively low computational load. This invention reduces the computational complexity and workload of streaming speech recognition, improves the computational speed, and is suitable for multi-channel speech signal processing in streaming scenarios.
[0167] According to the present invention, a streaming multichannel end-to-end speech recognition device 600 based on spherical harmonic function transform is provided, wherein the beamforming module 620 is specifically used for:
[0168] The number of channels after transformation is determined based on the order of the spherical harmonic transform function;
[0169] Based on the transformed number of channels, determine the spatial information corresponding to the spectral data of each frame multi-channel signal;
[0170] Based on the spatial information corresponding to the spectral data of each frame multi-channel signal, the spectral data of each frame multi-channel signal after determining the number of transformation channels is obtained.
[0171] According to the present invention, a streaming multichannel end-to-end speech recognition device 600 based on spherical harmonic function transform is provided, wherein the convolutional block attention module includes a channel attention module and a spatial attention module; the beamforming module 620 is further used for:
[0172] The first amplitude spectrum data of each frame is input into the channel attention module to obtain the channel attention matrix corresponding to the first amplitude spectrum data of each frame.
[0173] Based on the channel attention matrix and the first amplitude spectrum data of each frame, the second amplitude spectrum data of each frame is determined;
[0174] The second amplitude spectrum data of each frame is input into the spatial attention module to obtain the spatial attention matrix corresponding to the second amplitude spectrum data of each frame.
[0175] Based on the spatial attention matrix and the second amplitude spectrum data of each frame, the first amplitude spectrum data of each frame after de-reverberation is determined.
[0176] According to the present invention, a streaming multichannel end-to-end speech recognition device 600 based on spherical harmonic function transform is provided, wherein the beamforming module 620 is further used for:
[0177] The fully connected function of the activation layer in the self-attention channel combiner is used to activate the first amplitude spectrum data of each frame after dérification, thereby obtaining the key vector, query vector, and value vector corresponding to the first amplitude spectrum data of each frame after dérification; the dimension of the first amplitude spectrum data of each frame after dérification is determined based on the number of frames, the number of channels, and the number of frequencies included in the multi-channel signal waveform to be identified;
[0178] Based on the query vector, the key vector, and the transformed number of channels, determine the self-attention matrix of the key;
[0179] Based on the self-attention matrix and the value vector, determine the self-attention matrix of the value;
[0180] Based on the self-attention matrix of the value and the first amplitude spectrum data of each frame after déreverberation, the beam signal after beamforming is determined.
[0181] According to the present invention, a streaming multichannel end-to-end speech recognition device 600 based on spherical harmonic function transform is provided, wherein the recognition module 630 is specifically used for:
[0182] The beam signal after beamforming is post-filtered using a multi-head self-attention module to obtain a filtered beam signal.
[0183] Using an ASR network, the target category corresponding to the multi-channel signal waveform to be identified is obtained based on the filtered beam signal.
[0184] According to the present invention, a streaming multichannel end-to-end speech recognition device 600 based on spherical harmonic function transform is provided, wherein the beamforming module 620 is further used for
[0185] The multi-channel signal waveform to be identified is subjected to frame-by-frame short-time Fourier transform to obtain the spectral data corresponding to each frame of the multi-channel signal.
[0186] Figure 7 An example is a schematic diagram of the physical structure of an electronic device, such as... Figure 7 As shown, the electronic device may include: a processor 710, a communications interface 720, a memory 730, and a communication bus 740, wherein the processor 710, the communications interface 720, and the memory 730 communicate with each other via the communication bus 740. The processor 710 can call logic instructions in the memory 730 to execute a streaming multichannel end-to-end speech recognition method based on spherical harmonic function transform, the method including:
[0187] Acquire the multi-channel signal waveform to be identified; the multi-channel signal waveform to be identified includes at least two frames of multi-channel signal;
[0188] A spherical harmonic transform is performed on the spectral data corresponding to the multi-channel signal of each frame to obtain the spectral data corresponding to the multi-channel signal of each frame after the number of transformed channels, and the spectral data corresponding to the multi-channel signal of each frame after the number of transformed channels is converted into the first amplitude spectrum data of each frame; the spectral data is a time-frequency domain signal.
[0189] Based on the first amplitude spectrum data of each frame, the first amplitude spectrum data of each frame after de-reverberation is determined by the convolutional block attention module, and the beam signal after beamforming is determined by the self-attention channel combiner according to the first amplitude spectrum data of each frame after de-reverberation.
[0190] Based on the beam signal after beamforming, the target category corresponding to the multi-channel signal waveform to be identified is obtained using the target ASR network.
[0191] Furthermore, the logical instructions in the aforementioned memory 730 can be implemented as software functional units and, when sold or used as independent products, can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present invention, essentially, or the part that contributes to the prior art, or a part of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute all or part of the steps of the methods described in the various embodiments of the present invention. The aforementioned storage medium includes various media capable of storing program code, such as USB flash drives, portable hard drives, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical disks.
[0192] On the other hand, the present invention also provides a computer program product, the computer program product comprising a computer program that can be stored on a non-transitory computer-readable storage medium, wherein when the computer program is executed by a processor, the computer is able to execute the streaming multi-channel end-to-end speech recognition method based on spherical harmonic function transform provided by the above methods, the method comprising:
[0193] Acquire the multi-channel signal waveform to be identified; the multi-channel signal waveform to be identified includes at least two frames of multi-channel signal;
[0194] A spherical harmonic transform is performed on the spectral data corresponding to the multi-channel signal of each frame to obtain the spectral data corresponding to the multi-channel signal of each frame after the number of transformed channels, and the spectral data corresponding to the multi-channel signal of each frame after the number of transformed channels is converted into the first amplitude spectrum data of each frame; the spectral data is a time-frequency domain signal.
[0195] Based on the first amplitude spectrum data of each frame, the first amplitude spectrum data of each frame after de-reverberation is determined by the convolutional block attention module, and the beam signal after beamforming is determined by the self-attention channel combiner according to the first amplitude spectrum data of each frame after de-reverberation.
[0196] Based on the beam signal after beamforming, the target category corresponding to the multi-channel signal waveform to be identified is obtained using the target ASR network.
[0197] In another aspect, the present invention also provides a non-transitory computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements the streaming multichannel end-to-end speech recognition method based on spherical harmonic function transform provided by the methods described above, the method comprising:
[0198] Acquire the multi-channel signal waveform to be identified; the multi-channel signal waveform to be identified includes at least two frames of multi-channel signal;
[0199] A spherical harmonic transform is performed on the spectral data corresponding to the multi-channel signal of each frame to obtain the spectral data corresponding to the multi-channel signal of each frame after the number of transformed channels, and the spectral data corresponding to the multi-channel signal of each frame after the number of transformed channels is converted into the first amplitude spectrum data of each frame; the spectral data is a time-frequency domain signal.
[0200] Based on the first amplitude spectrum data of each frame, the first amplitude spectrum data of each frame after de-reverberation is determined by the convolutional block attention module, and the beam signal after beamforming is determined by the self-attention channel combiner according to the first amplitude spectrum data of each frame after de-reverberation.
[0201] Based on the beam signal after beamforming, the target category corresponding to the multi-channel signal waveform to be identified is obtained using the target ASR network.
[0202] The device embodiments described above are merely illustrative. The units described as separate components may or may not be physically separate. The components shown as units may or may not be physical units; that is, they may be located in one place or distributed across multiple network units. Some or all of the modules can be selected to achieve the purpose of this embodiment according to actual needs. Those skilled in the art can understand and implement this without any creative effort.
[0203] Through the above description of the embodiments, those skilled in the art can clearly understand that each embodiment can be implemented by means of software plus necessary general-purpose hardware platforms, and of course, it can also be implemented by hardware. Based on this understanding, the above technical solutions, in essence or the part that contributes to the prior art, can be embodied in the form of a software product. This computer software product can be stored in a computer-readable storage medium, such as ROM / RAM, magnetic disk, optical disk, etc., and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute the methods described in the various embodiments or some parts of the embodiments.
[0204] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, and not to limit them; although the present invention has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some of the technical features; and these modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of the present invention.
[0205] [1]Vaswani et al., "Attention Is All You Need".
Claims
1. A streaming multi-channel end-to-end speech recognition method based on spherical harmonic function transform, characterized in that, include: Acquire the multi-channel signal waveform to be identified; the multi-channel signal waveform to be identified includes at least two frames of multi-channel signal; A spherical harmonic transform is performed on the spectral data corresponding to the multi-channel signal of each frame to obtain the spectral data corresponding to the multi-channel signal of each frame after the number of transformed channels, and the spectral data corresponding to the multi-channel signal of each frame after the number of transformed channels is converted into the first amplitude spectrum data of each frame; the spectral data is a time-frequency domain signal. Based on the first amplitude spectrum data of each frame, a convolutional block attention module is used to determine the dedevered first amplitude spectrum data of each frame, and a self-attention channel combiner is used to determine the beamforming signal based on the dedevered first amplitude spectrum data of each frame. The convolutional block attention module includes a four-layer cascaded CBAM module, which uses residual linking to approximate the dedevering process. The self-attention channel combiner uses a self-attention mechanism to approximate the beamforming process of the multi-channel signal. The step of determining the beamforming signal based on the dedevered first amplitude spectrum data of each frame using the self-attention channel combiner includes: using the self-attention channel combiner... The fully connected function of the activation layer in the device activates the first amplitude spectrum data of each frame after dérebraization to obtain the key vector, query vector, and value vector corresponding to the first amplitude spectrum data of each frame after dérebraization. The dimension of the first amplitude spectrum data of each frame after dérebraization is determined based on the number of frames, channels, and frequencies of the multi-channel signal waveform to be identified. Based on the query vector, the key vector, and the transformed number of channels, the self-attention matrix of the key is determined. Based on the self-attention matrix and the value vector, the self-attention matrix of the value is determined. Based on the self-attention matrix of the value and the first amplitude spectrum data of each frame after dérebraization, the beam signal after beamforming is determined. Based on the beamformed signal, a target ASR network is used to identify the target category corresponding to the multi-channel signal waveform to be identified. This identification includes: using a multi-head self-attention module to perform post-filtering on the beamformed signal to obtain a filtered beamformed signal. The multi-head self-attention module focuses on time information and acts as a filter in the frequency domain to remove residual noise. The target ASR network is then used to identify the target category based on the filtered beamformed signal. The target category corresponds to the multi-channel signal waveform; the target ASR network is a Conformer network based on convolution and self-attention mechanisms, the Conformer network includes a feedforward module, a convolution module and a multi-head self-attention module; the target ASR network is obtained by training two initial ASR networks with shared weights. During training, the data-enhanced streaming signal and the whole sentence signal are respectively input into the two initial ASR networks with shared weights to obtain streaming recognition results and whole sentence recognition results. Based on the whole sentence recognition results, the streaming recognition results and the real results, the non-streaming loss and streaming loss are calculated respectively until the loss is within a certain range.
2. The streaming multi-channel end-to-end speech recognition method based on spherical harmonic function transform according to claim 1, characterized in that, The step of performing spherical harmonic transformation on the spectral data corresponding to each frame of multi-channel signals to obtain the spectral data corresponding to each frame of multi-channel signals after the number of transformed channels includes: The number of channels after transformation is determined based on the order of the spherical harmonic transform function; Based on the transformed number of channels, determine the spatial information corresponding to the spectral data of each frame multi-channel signal; Based on the spatial information corresponding to the spectral data of each frame multi-channel signal, the spectral data of each frame multi-channel signal after determining the number of transformation channels is obtained.
3. The streaming multi-channel end-to-end speech recognition method based on spherical harmonic function transform according to claim 1, characterized in that, The convolutional block attention module includes a channel attention module and a spatial attention module; The step of determining the de-reverberated first amplitude spectrum data of each frame using a convolutional block attention module based on the first amplitude spectrum data of each frame includes: The first amplitude spectrum data of each frame is input into the channel attention module to obtain the channel attention matrix corresponding to the first amplitude spectrum data of each frame. Based on the channel attention matrix and the first amplitude spectrum data of each frame, the second amplitude spectrum data of each frame is determined; The second amplitude spectrum data of each frame is input into the spatial attention module to obtain the spatial attention matrix corresponding to the second amplitude spectrum data of each frame. Based on the spatial attention matrix and the second amplitude spectrum data of each frame, the first amplitude spectrum data of each frame after de-reverberation is determined.
4. The streaming multi-channel end-to-end speech recognition method based on spherical harmonic function transform according to claim 1, characterized in that, Before performing spherical harmonic transformation on the spectral data corresponding to each of the multi-channel signals in each frame to obtain the spectral data corresponding to each of the multi-channel signals after the number of transformed channels, the method further includes: The multi-channel signal waveform to be identified is subjected to frame-by-frame short-time Fourier transform to obtain the spectral data corresponding to each frame of the multi-channel signal.
5. A streaming multi-channel end-to-end speech recognition device based on spherical harmonic function transform, characterized in that, include: An acquisition module is used to acquire the waveform of a multi-channel signal to be identified; the waveform of the multi-channel signal to be identified includes at least two frames of multi-channel signal. The beamforming module is used to perform spherical harmonic transformation on the spectral data corresponding to the multi-channel signals of each frame to obtain the spectral data corresponding to the multi-channel signals of each frame after the number of transformed channels, and to convert the spectral data corresponding to the multi-channel signals of each frame after the number of transformed channels into the first amplitude spectrum data of each frame; the spectral data is a time-frequency domain signal. Based on the first amplitude spectrum data of each frame, the de-reverberated first amplitude spectrum data of each frame is determined using a convolutional block attention module, and the beam signal after beamforming is determined using a self-attention channel combiner based on the de-reverberated first amplitude spectrum data of each frame; the convolutional block attention module includes a four-layer cascaded CBAM module, which uses residual linking to approximate the de-reverberation process; the self-attention channel combiner uses a self-attention mechanism to approximate the beamforming process of the multi-channel signal; the beamforming module is specifically used to: utilize the self-attention... The fully connected function of the activation layer in the intention channel combiner activates the first amplitude spectrum data of each frame after dération, obtaining the key vector, query vector, and value vector corresponding to the first amplitude spectrum data of each frame after dération; the dimension of the first amplitude spectrum data of each frame after dération is determined based on the number of frames, channels, and frequencies included in the multi-channel signal waveform to be identified; the self-attention matrix of the key is determined based on the query vector, the key vector, and the number of transformed channels; the self-attention matrix of the value is determined based on the self-attention matrix and the value vector. Based on the self-attention matrix of the value and the first amplitude spectrum data of each frame after déreverberation, the beam signal after beamforming is determined; The identification module is used to identify the target category corresponding to the multi-channel signal waveform to be identified based on the beamformed beam signal using a target ASR network. Specifically, the identification module uses a multi-head self-attention module to perform post-filtering on the beamformed beam signal to obtain a filtered beam signal. The multi-head self-attention module focuses on time information and acts as a filter in the frequency domain to remove residual noise. The target ASR network is used to identify the target category corresponding to the multi-channel signal waveform to be identified based on the filtered beam signal. The target ASR network is a Conformer network based on convolution and self-attention mechanisms. The Conformer network includes a feedforward module, a convolution module, and a multi-head self-attention module. The target ASR network is obtained by training two initial ASR networks with shared weights. During training, the data-enhanced streaming signal and the complete sentence signal are input into the two initial ASR networks with shared weights to obtain streaming recognition results and complete sentence recognition results. Based on the complete sentence recognition results, the streaming recognition results, and the true results, non-streaming loss and streaming loss are calculated respectively until the loss is within a certain range.
6. An electronic device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that, When the processor executes the computer program, it implements the streaming multichannel end-to-end speech recognition method based on spherical harmonic function transformation as described in any one of claims 1 to 4.
7. A non-transitory computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by the processor, it implements the streaming multichannel end-to-end speech recognition method based on spherical harmonic function transformation as described in any one of claims 1 to 4.
8. A computer program product, comprising a computer program, characterized in that, When the computer program is executed by the processor, it implements the streaming multichannel end-to-end speech recognition method based on spherical harmonic function transformation as described in any one of claims 1 to 4.