A speech recognition method, apparatus and device

CN115831122BActive Publication Date: 2026-09-18CHINA MOBILE COMM LTD RES INST +1
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202111098391.1
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2021-09-18
Publication Date
2026-09-18
Estimated Expiration
2041-09-18

AI Technical Summary

Technical Problem

[0006]本发明的目的在于提供一种语音识别方法、装置及设备,以解决现有技术中针对多信道传输的语音识别方案识别性能差的问题

Benefits of technology

[0058]In the above scheme, the speech recognition method encodes the speech data to be recognized using autoencoders corresponding to at least two existing channels to obtain encoded data corresponding to each existing channel; it then uses a self-attention mechanism to obtain at least two abstract spatial features of the speech data based on the encoded data; it further uses the self-attention mechanism to obtain the intrinsic relationship information between each of the abstract spatial features; and finally, it decodes the speech data based on the intrinsic relationship information to obtain the recognition result of the speech data. This scheme can realize an integrated robust speech recognition scheme that integrates multiple transmission channels. It can use a dual attention mechanism to learn the intrinsic relationship between different channels and automatically determine the proportion of information learned under each transmission channel, thereby achieving good recognition of multiple channels transmitted to the ASR port. Specifically, the embedding generated by the multi-transmission channel autoencoder (i.e., pre-trained autoencoder structure) can replace the Mel features used in the prior art, which can eliminate the influence of the existing filter bank feature acquisition method on the sensitivity of channel differences to a certain extent. In addition, the dual attention mechanism can be used to realize the autonomous learning of the intrinsic relationship between different channels without relying on prior channel information (i.e., prior channel feature labels), thereby realizing an integrated robust speech recognition scheme. Furthermore, besides improving the speech recognition accuracy for N (N greater than 1) transmission channels known during model training, for newly added transmission channels, the use of N pre-trained auto-encoders (which can be understood as channel encoders) can effectively simulate the embedding features of the new channels (specifically, for the speech data of the new transmission channels, the different contributions of the existing channels can be generated through the channel self-attention mechanism to obtain the fused channel information expression), eliminating the need for multiple retraining and system updates, thus reducing memory and computational resource consumption (specifically, reducing the data accumulation required when the transmission channels change, the computational resource consumption and time consumption caused by model retraining, reducing the impact of parameter differences caused by the channel, and improving recognition accuracy; that is, it does not require a large amount of corresponding data for training, avoiding the need for model retraining to consume a large amount of computational resources and a long time for specific update deployment); it effectively solves the problem of poor recognition performance of existing speech recognition schemes for multi-channel transmission.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115831122B_ABST
    Figure CN115831122B_ABST
Patent Text Reader

Abstract

The application provides a speech recognition method, device and equipment, wherein the speech recognition method comprises: using at least two existing channel corresponding automatic encoders to encode the speech data to be recognized to obtain the encoding data corresponding to each existing channel; using a self-attention mechanism to obtain at least two abstract space features of the speech data according to the encoding data; using the self-attention mechanism to obtain the internal connection information between each abstract space feature; and decoding according to the internal connection information to obtain the recognition result of the speech data. The scheme can realize an integrated robust speech recognition scheme fusing multiple transmission channels, can use a double attention mechanism to learn the internal relationship of different channels, automatically determine the proportion of the learned information under each transmission channel, and thus realize good recognition of the multiple channel transmission to the ASR port; and the scheme can well solve the poor recognition performance of the speech recognition scheme for the multiple channel transmission in the prior art.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of communication technology, and in particular to a speech recognition method, apparatus, and device. Background Technology

[0002] With the rapid development of the information age, numerous artificial intelligence (AI) technologies have spurred the practical application of smart terminal devices, such as smart speakers, in-vehicle navigation systems, mobile phone assistants, and smart homes. Among the various AI technologies, voice, as the most convenient human-computer interaction method, has made automatic speech recognition (ASR) an indispensable component of smart devices. In daily life and work, due to changing usage scenarios, multi-device communication (such as mobile phones, tablets, and laptops) is becoming increasingly common. Consequently, countless devices experience a decline in speech recognition performance across multiple transmission channels due to differences in audio acquisition and transmission between different terminals.

[0003] Under the influence of multiple transmission channels, the audio signals generated during communication can experience a sharp performance degradation when transmitted to the invoked speech recognition model due to channel differences. This is because different channels have unique response functions, or because of audio compression processing implemented by devices to limit signal transmission. To address this, existing technologies provide end-to-end speech recognition models.

[0004] However, in existing technologies, end-to-end speech recognition models use Mel filter banks to extract features, which relies on the statistical properties of acoustic signals and filter design, thus exacerbating the sensitivity to differences in channel response. They cannot handle the differences in speech from multiple transmission channels, leading to a decrease in recognition performance when encountering audio signals transmitted from multiple channels.

[0005] As can be seen from the above, existing speech recognition solutions for multi-channel transmission suffer from problems such as poor recognition performance. Summary of the Invention

[0006] The purpose of this invention is to provide a speech recognition method, apparatus, and device to solve the problem of poor recognition performance of existing speech recognition schemes for multi-channel transmission.

[0007] To address the aforementioned technical problems, embodiments of the present invention provide a speech recognition method, comprising:

[0008] The speech data to be recognized is encoded using at least two existing channel-corresponding autoencoders to obtain the encoded data corresponding to each existing channel.

[0009] Using a self-attention mechanism, at least two abstract spatial features of the speech data are obtained based on the encoded data;

[0010] By utilizing the self-attention mechanism, the intrinsic relationship information between the various abstract space features is obtained;

[0011] The recognition result of the speech data is obtained by decoding based on the inherent connection information.

[0012] Optional, also includes:

[0013] Using a self-attention mechanism, the inherent correlation information between the existing channels is obtained based on the encoded data;

[0014] The method of utilizing a self-attention mechanism to obtain at least two abstract spatial features of the speech data based on the encoded data includes:

[0015] When the source channel of the speech data does not belong to the existing channel, a self-attention mechanism is used to obtain at least two abstract spatial features of the speech data based on the encoded data and the inherent correlation information.

[0016] Optionally, the step of utilizing a self-attention mechanism to obtain the intrinsic correlation information between the existing channels based on the encoded data includes:

[0017] Using a self-attention mechanism, spatial response function correlation information between the existing channels is obtained based on the encoded data.

[0018] Optionally, the step of utilizing a self-attention mechanism to obtain at least two abstract spatial features of the speech data based on the encoded data includes:

[0019] Using Formula 1 and a self-attention mechanism, at least two abstract spatial features of the speech data are obtained based on the encoded data.

[0020] Formula one is as follows:

[0021] C out =∑G(en) j )×F(en j ), j = 1···N;

[0022] C out Representing the at least two abstract space features, G(en) j F(en) represents the weighted weight matrix for self-attention. j ) represents the channel response abstract feature output matrix of the autoencoder, and N represents the number of existing channels.

[0023] This invention also provides a voice recognition device, comprising:

[0024] The first encoding module is used to encode the speech data to be recognized using at least two existing channel-corresponding autoencoders to obtain the encoded data corresponding to each existing channel.

[0025] The first acquisition module is used to acquire at least two abstract spatial features of the speech data based on the encoded data using a self-attention mechanism.

[0026] The second acquisition module is used to acquire the intrinsic relationship information between the various abstract space features by utilizing a self-attention mechanism;

[0027] The first decoding module is used to decode the speech data based on the inherent connection information to obtain the recognition result of the speech data.

[0028] Optional, also includes:

[0029] The third acquisition module is used to acquire the intrinsic correlation information between the existing channels based on the encoded data using a self-attention mechanism.

[0030] The method of utilizing a self-attention mechanism to obtain at least two abstract spatial features of the speech data based on the encoded data includes:

[0031] When the source channel of the speech data does not belong to the existing channel, a self-attention mechanism is used to obtain at least two abstract spatial features of the speech data based on the encoded data and the inherent correlation information.

[0032] Optionally, the step of utilizing a self-attention mechanism to obtain the intrinsic correlation information between the existing channels based on the encoded data includes:

[0033] Using a self-attention mechanism, spatial response function correlation information between the existing channels is obtained based on the encoded data.

[0034] Optionally, the step of utilizing a self-attention mechanism to obtain at least two abstract spatial features of the speech data based on the encoded data includes:

[0035] Using Formula 1 and a self-attention mechanism, at least two abstract spatial features of the speech data are obtained based on the encoded data.

[0036] Formula one is as follows:

[0037] C out =∑G(en) j )×F(en j ), j = 1···N;

[0038] C out Representing the at least two abstract space features, G(en) j F(en) represents the weighted weight matrix for self-attention. j ) represents the channel response abstract feature output matrix of the autoencoder, and N represents the number of existing channels.

[0039] This invention also provides a voice recognition device, including: a processor and a transceiver;

[0040] The processor is used to encode the speech data to be recognized using at least two existing channel-corresponding autoencoders to obtain encoded data corresponding to each existing channel.

[0041] Using a self-attention mechanism, at least two abstract spatial features of the speech data are obtained based on the encoded data;

[0042] By utilizing the self-attention mechanism, the intrinsic relationship information between the various abstract space features is obtained;

[0043] The recognition result of the speech data is obtained by decoding based on the inherent connection information.

[0044] Optionally, the processor is further configured to:

[0045] Using a self-attention mechanism, the inherent correlation information between the existing channels is obtained based on the encoded data;

[0046] The method of utilizing a self-attention mechanism to obtain at least two abstract spatial features of the speech data based on the encoded data includes:

[0047] When the source channel of the speech data does not belong to the existing channel, a self-attention mechanism is used to obtain at least two abstract spatial features of the speech data based on the encoded data and the inherent correlation information.

[0048] Optionally, the step of utilizing a self-attention mechanism to obtain the intrinsic correlation information between the existing channels based on the encoded data includes:

[0049] Using a self-attention mechanism, spatial response function correlation information between the existing channels is obtained based on the encoded data.

[0050] Optionally, the step of utilizing a self-attention mechanism to obtain at least two abstract spatial features of the speech data based on the encoded data includes:

[0051] Using Formula 1 and a self-attention mechanism, at least two abstract spatial features of the speech data are obtained based on the encoded data.

[0052] Formula one is as follows:

[0053] C out =∑G(en) j )×F(en j ), j = 1···N;

[0054] C out Representing the at least two abstract space features, G(en) j F(en) represents the weighted weight matrix for self-attention. j ) represents the channel response abstract feature output matrix of the autoencoder, and N represents the number of existing channels.

[0055] This invention also provides a voice recognition device, including a memory, a processor, and a program stored in the memory and executable on the processor; when the processor executes the program, it implements the above-described voice recognition method.

[0056] This invention also provides a readable storage medium storing a program that, when executed by a processor, implements the steps in the above-described speech recognition method.

[0057] The beneficial effects of the above-described technical solution of the present invention are as follows:

[0058] In the above scheme, the speech recognition method encodes the speech data to be recognized using autoencoders corresponding to at least two existing channels to obtain encoded data corresponding to each existing channel; it then uses a self-attention mechanism to obtain at least two abstract spatial features of the speech data based on the encoded data; it further uses the self-attention mechanism to obtain the intrinsic relationship information between each of the abstract spatial features; and finally, it decodes the speech data based on the intrinsic relationship information to obtain the recognition result of the speech data. This scheme can realize an integrated robust speech recognition scheme that integrates multiple transmission channels. It can use a dual attention mechanism to learn the intrinsic relationship between different channels and automatically determine the proportion of information learned under each transmission channel, thereby achieving good recognition of multiple channels transmitted to the ASR port. Specifically, the embedding generated by the multi-transmission channel autoencoder (i.e., pre-trained autoencoder structure) can replace the Mel features used in the prior art, which can eliminate the influence of the existing filter bank feature acquisition method on the sensitivity of channel differences to a certain extent. In addition, the dual attention mechanism can be used to realize the autonomous learning of the intrinsic relationship between different channels without relying on prior channel information (i.e., prior channel feature labels), thereby realizing an integrated robust speech recognition scheme. Furthermore, besides improving the speech recognition accuracy for N (N greater than 1) transmission channels known during model training, for newly added transmission channels, the use of N pre-trained auto-encoders (which can be understood as channel encoders) can effectively simulate the embedding features of the new channels (specifically, for the speech data of the new transmission channels, the different contributions of the existing channels can be generated through the channel self-attention mechanism to obtain the fused channel information expression), eliminating the need for multiple retraining and system updates, thus reducing memory and computational resource consumption (specifically, reducing the data accumulation required when the transmission channels change, the computational resource consumption and time consumption caused by model retraining, reducing the impact of parameter differences caused by the channel, and improving recognition accuracy; that is, it does not require a large amount of corresponding data for training, avoiding the need for model retraining to consume a large amount of computational resources and a long time for specific update deployment); it effectively solves the problem of poor recognition performance of existing speech recognition schemes for multi-channel transmission. Attached Figure Description

[0059] Figure 1 This is a schematic diagram of the speech recognition method according to an embodiment of the present invention;

[0060] Figure 2 This is a schematic diagram of a speech recognition method implementation system according to an embodiment of the present invention;

[0061] Figure 3 This is a schematic diagram of the structure of the speech recognition device according to an embodiment of the present invention;

[0062] Figure 4 This is a schematic diagram of the structure of a speech recognition device according to an embodiment of the present invention. Detailed Implementation

[0063] To make the technical problems, technical solutions and advantages of the present invention clearer, a detailed description will be given below in conjunction with the accompanying drawings and specific embodiments.

[0064] This invention addresses the problem of poor recognition performance in existing speech recognition schemes for multi-channel transmission by providing a speech recognition method, such as... Figure 1 As shown, it includes:

[0065] Step 11: Encode the speech data to be recognized using at least two existing channel-corresponding autoencoders to obtain the encoded data corresponding to each existing channel;

[0066] Step 12: Using a self-attention mechanism, obtain at least two abstract spatial features of the speech data based on the encoded data;

[0067] Step 13: Utilize the self-attention mechanism to obtain the intrinsic relationship information between the various abstract space features;

[0068] Step 14: Decode the speech data based on the inherent connection information to obtain the recognition result of the speech data.

[0069] After step 14, the recognition results can also be returned to the other device (such as the other terminal), which is not limited here. Specifically, "intrinsic relationship information" can refer to the implicit relationships between the speech frames that need to be recognized, established using self-attention.

[0070] The speech recognition method provided in this invention encodes the speech data to be recognized using autoencoders corresponding to at least two existing channels, obtaining encoded data corresponding to each existing channel; it then uses a self-attention mechanism to obtain at least two abstract spatial features of the speech data based on the encoded data; furthermore, it uses the self-attention mechanism to obtain the intrinsic relationship information between each of the abstract spatial features; and finally, it decodes the speech data based on the intrinsic relationship information to obtain the recognition result of the speech data. This method enables an integrated robust speech recognition scheme that integrates multiple transmission channels. It utilizes a dual attention mechanism to learn the intrinsic relationships between different channels, automatically determining the proportion of learned information under each transmission channel, thereby achieving good recognition of multiple channels transmitted to the ASR port. Specifically, it can replace the Mel features used in the prior art by using embeddings generated by a multi-transmission channel autoencoder (i.e., a pre-trained autoencoder structure), which to some extent eliminates the influence of existing filter bank-based feature acquisition methods on channel difference sensitivity. Furthermore, it can utilize a dual attention mechanism to achieve autonomous learning of the intrinsic relationships between different channels without relying on prior channel information (i.e., prior channel feature labels), thus realizing an integrated robust speech recognition scheme. Furthermore, besides improving the speech recognition accuracy for N (N greater than 1) transmission channels known during model training, for newly added transmission channels, the use of N pre-trained auto-encoders (which can be understood as channel encoders) can effectively simulate the embedding features of the new channels (specifically, for the speech data of the new transmission channels, the different contributions of the existing channels can be generated through the channel self-attention mechanism to obtain the fused channel information expression), eliminating the need for multiple retraining and system updates, thus reducing memory and computational resource consumption (specifically, reducing the data accumulation required when the transmission channels change, the computational resource consumption and time consumption caused by model retraining, reducing the impact of parameter differences caused by the channel, and improving recognition accuracy; that is, it does not require a large amount of corresponding data for training, avoiding the need for model retraining to consume a large amount of computational resources and a long time for specific update deployment); it effectively solves the problem of poor recognition performance of existing speech recognition schemes for multi-channel transmission.

[0071] Furthermore, the speech recognition method further includes: using a self-attention mechanism to obtain intrinsic correlation information between the existing channels based on the encoded data; the step of using a self-attention mechanism to obtain at least two abstract spatial features of the speech data based on the encoded data includes: when the source channel of the speech data does not belong to the existing channels, using a self-attention mechanism to obtain at least two abstract spatial features of the speech data based on the encoded data and intrinsic correlation information.

[0072] This enables the recognition of speech data from newly added channels. The "intrinsic correlation information" specifically refers to the correlation between channel features of different channels. More specifically, "intrinsic correlation information" can be the correlation information established by using self-attention on the embedding information of existing channels in this scheme, which can be used to fit the feature representation of the new channel.

[0073] The step of using a self-attention mechanism to obtain the intrinsic correlation information between the existing channels based on the encoded data includes: using a self-attention mechanism to obtain the spatial response function correlation information between the existing channels based on the encoded data.

[0074] This allows for the accurate determination of the inherent relationships between different channels.

[0075] In this embodiment of the invention, the step of using a self-attention mechanism to obtain at least two abstract spatial features of the speech data based on the encoded data includes: using Formula 1, employing a self-attention mechanism to obtain at least two abstract spatial features of the speech data based on the encoded data; wherein, Formula 1 is: C out =∑G(en) j )×F(en j ), j = 1···N; C out Representing the at least two abstract space features, G(en) j F(en) represents the weighted weight matrix for self-attention. j ) represents the channel response abstract feature output matrix of the autoencoder, and N represents the number of existing channels.

[0076] This allows for the accurate acquisition of the abstract spatial features of speech data.

[0077] The following is an example of the speech recognition method provided in the embodiments of the present invention, with a terminal as an example of the speech recognition device.

[0078] To address the aforementioned technical problems, this invention provides a speech recognition method, specifically an integrated robust speech recognition scheme that integrates multiple transmission channels. It utilizes a dual attention mechanism to learn the inherent relationships between different channels, automatically determining the proportion of learned information for each transmission channel, thereby achieving good recognition of multi-channel transmissions to the ASR port. The end-to-end speech recognition architecture involved in this scheme mainly includes three parts: an encoder, a self-attention mechanism module, and a decoder. Specifically, this scheme adds a multi-transmission-channel auto-encoder and a channel self-attention module to the common end-to-end recognition model; it utilizes a dual attention mechanism to learn the inherent relationships between different channels without relying on prior channel information, thus establishing an integrated robust speech recognition scheme. In addition to improving the speech recognition accuracy of the system for the N known transmission channels during model training, this solution can effectively simulate the embedding features of newly added transmission channels using the channel-wise self-attention module with N pre-trained auto-encoders, eliminating the need for multiple retraining and system updates and reducing memory and computational resource consumption.

[0079] Specifically, this solution can adopt... Figure 2 The system shown (also known as an end-to-end speech recognition system with dual self-attention mechanism) is implemented, and the system mainly includes the following parts (in this case, N=4):

[0080] 1. Voice input module: Inputs the collected voice data into the recognition system;

[0081] 2. Multi-transmission channel autoencoder module: In speech recognition scenarios involving multiple different transmission channels, it utilizes the existing accumulated database of N channels ( Figure 2 (N=4) Train their respective encoding and decoding systems to obtain pre-trained auto-encoder modules, and then integrate them together; for example... Figure 2 The encoder module consists of channel 1 encoder, channel 2 encoder, channel 3 encoder and channel 4 encoder.

[0082] 3. Channel-wise self-attention module (e.g., Figure 2Channel attention in the encoding: Specifically, this module differs from a typical encoder. During system training, this module is only used to extract corresponding features from the learning abstract space (e.g., environmental features of sound). This scheme uses this module to obtain the intrinsic correlation of each channel information in the embedding generated by the autoencoder (corresponding to the aforementioned intrinsic correlation information, which can cover newly added channels). Specifically, this can be achieved through... Figure 2 It is obtained using the Embedding Concat(function).

[0083] Specifically, the (channel) output of this module is: C out =∑G(en) j )×F(en j ), j = 1,..,4;

[0084] Among them, C out The output (corresponding to at least two of the above abstract space features), F(en) j ) represents the channel response abstract feature output matrix of each channel encoder (corresponding to the channel response abstract feature output matrix of the autoencoder described above), G(en) j ) is the weighted weight matrix for self-attention (corresponding to the weighted weight matrix for self-attention mentioned above). Specifically, by obtaining C out The formula yields a set of deep-dimensional features. The distribution of these features can cover the characteristics of the new channel. The system can be further trained for the new channel without any restrictions.

[0085] Furthermore, based on the input of this module (i.e., the output of the multi-transmission channel autoencoder module, corresponding to the above-mentioned encoded data), the spatial response function correlation between different acquisition devices (channels) can be captured through the self-attention mechanism (corresponding to the above-mentioned spatial response function correlation information), and the abstract spatial feature representation of the input speech signal at the current time (corresponding to the above-mentioned speech data to be recognized) can be selected (corresponding to the above-mentioned acquisition of at least two abstract spatial features of the speech data based on the encoded data using the self-attention mechanism).

[0086] 4. Identify self-attention modules (corresponding to...) Figure 2 In the context of ASR attention, existing methods can be used to obtain and identify the intrinsic relationships between relevant features; however, this is not a limitation. This corresponds to the aforementioned use of self-attention mechanisms to obtain information about the intrinsic relationships between the various abstract spatial features.

[0087] 5. Decoder module (corresponding to...) Figure 2The Decoder in the code can be designed with a criterion function to obtain the recognition result output; corresponding to the above decoding based on the intrinsic connection information to obtain the recognition result of the speech data.

[0088] 6. Recognition Results: The final recognition results can be returned to the interactive device terminal (i.e., the terminal opposite to this terminal) for subsequent processing operations such as Natural Language Processing (NLP).

[0089] As can be seen from the above, the solution provided by the embodiments of the present invention can construct an integrated robust speech recognition system that integrates multiple transmission channels. It fully utilizes existing accumulated data from each channel and adds a multi-transmission channel auto-encoder and a channel-wise self-attention module. Specifically, a dual attention mechanism can be used to learn the inherent relationships between different channels without relying on prior channel information, thus achieving an integrated robust speech recognition solution. Furthermore, it can improve recognition performance for both known and unknown transmission channels during training. In addition, the channel-wise self-attention module can use a pre-trained auto-encoder to encode information, effectively simulating the embedding features of the input signal without requiring multiple retraining and system updates, thereby reducing memory and computational resource consumption.

[0090] In summary, the solution provided by the embodiments of the present invention has the following advantages:

[0091] 1. The embedding generated by the pre-trained autoencoder structure is used to replace the Mel features used in the existing technology, which to some extent eliminates the impact of this traditional filter bank feature acquisition method on the sensitivity of channel differences;

[0092] 2. Utilizing a dual self-attention mechanism, it can autonomously learn the intrinsic relationships between channel information without relying on prior channel feature labels;

[0093] 3. An integrated robust recognition system can generate different contribution levels of existing channels for voice data from newly added transmission channels through a channel self-attention module, thereby obtaining a fused channel information representation (specifically, the relevant parameter information of existing channels can be fused according to a weight matrix).

[0094] 4. It reduces the data accumulation required when the transmission channel changes, the computational resource consumption and time consumption caused by model retraining, and reduces the impact of parameter differences caused by the channel, which can improve the recognition accuracy; that is, it does not require a large amount of corresponding data for training, and avoids the need for a large amount of computational resources and a long time for specific update deployments when retraining the model.

[0095] In summary, the solution provided by the embodiments of the present invention has improvements over existing solutions in terms of channel sensitivity of feature extraction, system robustness, recognition accuracy, and computational resource and time consumption.

[0096] This invention also provides a voice recognition device, such as... Figure 3 As shown, it includes:

[0097] The first encoding module 31 is used to encode the speech data to be recognized using at least two existing channel-corresponding autoencoders to obtain the encoded data corresponding to each existing channel.

[0098] The first acquisition module 32 is used to acquire at least two abstract spatial features of the speech data based on the encoded data using a self-attention mechanism.

[0099] The second acquisition module 33 is used to acquire the intrinsic relationship information between the various abstract space features by utilizing a self-attention mechanism;

[0100] The first decoding module 34 is used to decode the speech data based on the inherent connection information to obtain the recognition result of the speech data.

[0101] The speech recognition device provided in this embodiment of the invention encodes the speech data to be recognized using autoencoders corresponding to at least two existing channels to obtain encoded data corresponding to each existing channel; it then uses a self-attention mechanism to obtain at least two abstract spatial features of the speech data based on the encoded data; it further uses the self-attention mechanism to obtain the intrinsic relationship information between each of the abstract spatial features; and finally, it decodes the speech data based on the intrinsic relationship information to obtain the recognition result of the speech data. This enables an integrated robust speech recognition scheme that integrates multiple transmission channels. It utilizes a dual attention mechanism to learn the intrinsic relationships between different channels, automatically determining the proportion of learned information under each transmission channel, thereby achieving good recognition of multiple channels transmitted to the ASR port. Specifically, it can replace the Mel features used in the prior art by using embeddings generated by a multi-transmission channel autoencoder (i.e., a pre-trained autoencoder structure), which to some extent eliminates the influence of existing filter bank-based feature acquisition methods on channel difference sensitivity. Furthermore, it can utilize a dual attention mechanism to achieve autonomous learning of the intrinsic relationships between different channels without relying on prior channel information (i.e., prior channel feature labels), thus realizing an integrated robust speech recognition scheme. Furthermore, besides improving the speech recognition accuracy for N (N greater than 1) transmission channels known during model training, for newly added transmission channels, the use of N pre-trained auto-encoders (which can be understood as channel encoders) can effectively simulate the embedding features of the new channels (specifically, for the speech data of the new transmission channels, the different contributions of the existing channels can be generated through the channel self-attention mechanism to obtain the fused channel information expression), eliminating the need for multiple retraining and system updates, thus reducing memory and computational resource consumption (specifically, reducing the data accumulation required when the transmission channels change, the computational resource consumption and time consumption caused by model retraining, reducing the impact of parameter differences caused by the channel, and improving recognition accuracy; that is, it does not require a large amount of corresponding data for training, avoiding the need for model retraining to consume a large amount of computational resources and a long time for specific update deployment); it effectively solves the problem of poor recognition performance of existing speech recognition schemes for multi-channel transmission.

[0102] Furthermore, the speech recognition device further includes: a third acquisition module, used to acquire, using a self-attention mechanism, intrinsic correlation information between the existing channels based on the encoded data; the acquisition of at least two abstract spatial features of the speech data based on the encoded data using the self-attention mechanism includes: when the source channel of the speech data does not belong to the existing channels, acquiring at least two abstract spatial features of the speech data based on the encoded data and intrinsic correlation information using the self-attention mechanism.

[0103] The step of using a self-attention mechanism to obtain the intrinsic correlation information between the existing channels based on the encoded data includes: using a self-attention mechanism to obtain the spatial response function correlation information between the existing channels based on the encoded data.

[0104] In this embodiment of the invention, the step of using a self-attention mechanism to obtain at least two abstract spatial features of the speech data based on the encoded data includes: using Formula 1, employing a self-attention mechanism to obtain at least two abstract spatial features of the speech data based on the encoded data; wherein, Formula 1 is: C out =∑G(en) j )×F(en j ), j = 1···N; C out Representing the at least two abstract space features, G(en) j F(en) represents the weighted weight matrix for self-attention. j ) represents the channel response abstract feature output matrix of the autoencoder, and N represents the number of existing channels.

[0105] The implementation embodiments of the above-described speech recognition method are all applicable to the embodiments of the speech recognition device and can achieve the same technical effect.

[0106] This invention also provides a voice recognition device, such as... Figure 4 As shown, it includes: processor 41 and transceiver 42;

[0107] The processor 41 is used to encode the speech data to be recognized using at least two existing channel-corresponding autoencoders to obtain encoded data corresponding to each existing channel.

[0108] Using a self-attention mechanism, at least two abstract spatial features of the speech data are obtained based on the encoded data;

[0109] By utilizing the self-attention mechanism, the intrinsic relationship information between the various abstract space features is obtained;

[0110] The recognition result of the speech data is obtained by decoding based on the inherent connection information.

[0111] The speech recognition device provided in this embodiment of the invention encodes the speech data to be recognized using autoencoders corresponding to at least two existing channels, obtaining encoded data corresponding to each existing channel; it then uses a self-attention mechanism to obtain at least two abstract spatial features of the speech data based on the encoded data; furthermore, it uses the self-attention mechanism to obtain the intrinsic relationship information between each of the abstract spatial features; and finally, it decodes the speech data based on the intrinsic relationship information to obtain the recognition result of the speech data. This enables an integrated robust speech recognition scheme that integrates multiple transmission channels. It utilizes a dual attention mechanism to learn the intrinsic relationships between different channels, automatically determining the proportion of learned information under each transmission channel, thereby achieving good recognition of multiple channels transmitted to the ASR port. Specifically, it can replace the Mel features used in the prior art by using embeddings generated by a multi-transmission channel autoencoder (i.e., a pre-trained autoencoder structure), which to some extent eliminates the influence of existing filter bank-based feature acquisition methods on channel difference sensitivity. Furthermore, it can utilize a dual attention mechanism to achieve autonomous learning of the intrinsic relationships between different channels without relying on prior channel information (i.e., prior channel feature labels), thus realizing an integrated robust speech recognition scheme. Furthermore, besides improving the speech recognition accuracy for N (N greater than 1) transmission channels known during model training, for newly added transmission channels, the use of N pre-trained auto-encoders (which can be understood as channel encoders) can effectively simulate the embedding features of the new channels (specifically, for the speech data of the new transmission channels, the different contributions of the existing channels can be generated through the channel self-attention mechanism to obtain the fused channel information expression), eliminating the need for multiple retraining and system updates, thus reducing memory and computational resource consumption (specifically, reducing the data accumulation required when the transmission channels change, the computational resource consumption and time consumption caused by model retraining, reducing the impact of parameter differences caused by the channel, and improving recognition accuracy; that is, it does not require a large amount of corresponding data for training, avoiding the need for model retraining to consume a large amount of computational resources and a long time for specific update deployment); it effectively solves the problem of poor recognition performance of existing speech recognition schemes for multi-channel transmission.

[0112] Furthermore, the processor is also configured to: utilize a self-attention mechanism to obtain intrinsic correlation information between the existing channels based on the encoded data; the step of utilizing a self-attention mechanism to obtain at least two abstract spatial features of the speech data based on the encoded data includes: when the source channel of the speech data does not belong to the existing channels, utilizing a self-attention mechanism to obtain at least two abstract spatial features of the speech data based on the encoded data and intrinsic correlation information.

[0113] The step of using a self-attention mechanism to obtain the intrinsic correlation information between the existing channels based on the encoded data includes: using a self-attention mechanism to obtain the spatial response function correlation information between the existing channels based on the encoded data.

[0114] In this embodiment of the invention, the step of using a self-attention mechanism to obtain at least two abstract spatial features of the speech data based on the encoded data includes: using Formula 1, employing a self-attention mechanism to obtain at least two abstract spatial features of the speech data based on the encoded data; wherein, Formula 1 is: C out =∑G(en) j )×F(en j ), j = 1···N; C out Representing the at least two abstract space features, G(en) j F(en) represents the weighted weight matrix for self-attention. j ) represents the channel response abstract feature output matrix of the autoencoder, and N represents the number of existing channels.

[0115] The implementation embodiments of the above-described speech recognition method are all applicable to the embodiments of the speech recognition device and can achieve the same technical effect.

[0116] This invention also provides a voice recognition device, including a memory, a processor, and a program stored in the memory and executable on the processor; when the processor executes the program, it implements the above-described voice recognition method.

[0117] The implementation embodiments of the above-described speech recognition method are all applicable to the embodiments of the speech recognition device and can achieve the same technical effect.

[0118] This invention also provides a readable storage medium storing a program that, when executed by a processor, implements the steps in the above-described speech recognition method.

[0119] The implementation embodiments of the above-described speech recognition method are all applicable to the embodiments of the readable storage medium and can achieve the same technical effect.

[0120] It should be noted that many of the functional components described in this specification are referred to as modules in order to more specifically emphasize the independence of their implementation.

[0121] In this embodiment of the invention, the module can be implemented in software so that it can be executed by various types of processors. For example, an identified executable code module may include one or more physical or logical blocks of computer instructions, which may be constructed as objects, procedures, or functions. Nevertheless, the executable code of the identified module does not need to be physically located together, but may include different instructions stored in different bits, which, when logically combined, constitute the module and achieve the module's intended purpose.

[0122] In practice, an executable code module can be a single instruction or many instructions, and can even be distributed across multiple different code segments, different programs, and across multiple memory devices. Similarly, operational data can be identified within the module and can be implemented in any suitable form and organized within any suitable type of data structure. This operational data can be collected as a single dataset or distributed across different locations (including different storage devices), and can exist, at least in part, solely as electronic signals within the system or network.

[0123] When a module can be implemented using software, considering the current level of hardware technology, modules that can be implemented in software can be implemented using hardware circuits by those skilled in the art to achieve the corresponding functions, without considering cost. These hardware circuits include conventional very-large-scale integrated circuits (VLSI) or gate arrays, as well as existing semiconductors such as logic chips and transistors, or other discrete components. Modules can also be implemented using programmable hardware devices, such as field-programmable gate arrays, programmable array logic, and programmable logic devices.

[0124] The above describes the preferred embodiments of the present invention. It should be noted that those skilled in the art can make various improvements and modifications without departing from the principles of the present invention, and these improvements and modifications should also be considered within the scope of protection of the present invention.

Claims

1. A speech recognition method, characterized in that, include: The speech data to be recognized is encoded using at least two existing channel-corresponding autoencoders to obtain the encoded data corresponding to each existing channel. Using a self-attention mechanism, at least two abstract spatial features of the speech data are obtained based on the encoded data; By utilizing the self-attention mechanism, the intrinsic relationship information between the various abstract space features is obtained; The recognition result of the speech data is obtained by decoding based on the inherent connection information; The speech recognition method further includes: Using a self-attention mechanism, the inherent correlation information between the existing channels is obtained based on the encoded data; The method of utilizing a self-attention mechanism to obtain at least two abstract spatial features of the speech data based on the encoded data includes: When the source channel of the speech data does not belong to the existing channel, a self-attention mechanism is used to obtain at least two abstract spatial features of the speech data based on the encoded data and the inherent correlation information.

2. The speech recognition method according to claim 1, characterized in that, The step of using a self-attention mechanism to obtain the intrinsic correlation information between the existing channels based on the encoded data includes: Using a self-attention mechanism, spatial response function correlation information between the existing channels is obtained based on the encoded data.

3. The speech recognition method according to claim 1, characterized in that, The method of utilizing a self-attention mechanism to obtain at least two abstract spatial features of the speech data based on the encoded data includes: Using Formula 1 and a self-attention mechanism, at least two abstract spatial features of the speech data are obtained based on the encoded data. Formula one is as follows: C out =∑G(en j )×F(en j ),j=1,···,N; C out Representing the at least two abstract space features, G(en) j F(en) represents the weighted weight matrix for self-attention. j ) represents the channel response abstract feature output matrix of the autoencoder, and N represents the number of existing channels.

4. A voice recognition device, characterized in that, include: The first encoding module is used to encode the speech data to be recognized using at least two existing channel-corresponding autoencoders to obtain the encoded data corresponding to each existing channel. The first acquisition module is used to acquire at least two abstract spatial features of the speech data based on the encoded data using a self-attention mechanism. The second acquisition module is used to acquire the intrinsic relationship information between the various abstract space features by utilizing a self-attention mechanism; The first decoding module is used to decode the speech data based on the inherent connection information to obtain the recognition result of the speech data; The voice recognition device further includes: The third acquisition module is used to acquire the intrinsic correlation information between the existing channels based on the encoded data using a self-attention mechanism. The method of utilizing a self-attention mechanism to obtain at least two abstract spatial features of the speech data based on the encoded data includes: When the source channel of the speech data does not belong to the existing channel, a self-attention mechanism is used to obtain at least two abstract spatial features of the speech data based on the encoded data and the inherent correlation information.

5. The speech recognition device according to claim 4, characterized in that, The step of using a self-attention mechanism to obtain the intrinsic correlation information between the existing channels based on the encoded data includes: Using a self-attention mechanism, spatial response function correlation information between the existing channels is obtained based on the encoded data.

6. The speech recognition device according to claim 4, characterized in that, The method of utilizing a self-attention mechanism to obtain at least two abstract spatial features of the speech data based on the encoded data includes: Using Formula 1 and a self-attention mechanism, at least two abstract spatial features of the speech data are obtained based on the encoded data. Formula one is as follows: C out =∑G(en j )×F(en j ),j=1,···,N; C out Representing the at least two abstract space features, G(en) j F(en) represents the weighted weight matrix for self-attention. j ) represents the channel response abstract feature output matrix of the autoencoder, and N represents the number of existing channels.

7. A voice recognition device, characterized in that, include: Processor and transceiver; The processor is used to encode the speech data to be recognized using at least two existing channel-corresponding autoencoders to obtain encoded data corresponding to each existing channel. Using a self-attention mechanism, at least two abstract spatial features of the speech data are obtained based on the encoded data; By utilizing the self-attention mechanism, the intrinsic relationship information between the various abstract space features is obtained; The recognition result of the speech data is obtained by decoding based on the inherent connection information; The processor is further configured to: Using a self-attention mechanism, the inherent correlation information between the existing channels is obtained based on the encoded data; The method of utilizing a self-attention mechanism to obtain at least two abstract spatial features of the speech data based on the encoded data includes: When the source channel of the speech data does not belong to the existing channel, a self-attention mechanism is used to obtain at least two abstract spatial features of the speech data based on the encoded data and the inherent correlation information.

8. A voice recognition device, comprising a memory, a processor, and a program stored in the memory and executable on the processor; characterized in that, When the processor executes the program, it implements the speech recognition method as described in any one of claims 1 to 3.

9. A readable storage medium having a program stored thereon, characterized in that, When the program is executed by the processor, it implements the steps in the speech recognition method as described in any one of claims 1 to 3.

Citation Information

Patent Citations

  • End-to-end multi-channel speech recognition method adopting advanced feature fusion

    CN111524519A

  • End-to-end speech recognition method based on connection time sequence classification and self-attention mechanism

    CN112509564A