Speech recognition method, deep learning model training method, device, and equipment
By decoding speech features with initial recognition results as a priori information to achieve uniform word-level representations, the method improves speech recognition accuracy and efficiency, overcoming the limitations of inconsistent feature lengths in conventional methods.
Patent Information
- Application Number
- JP2024148016
- Authority / Receiving Office
- JP · JP
- Patent Type
- Patents
- Current Assignee / Owner
- Priority Date
- 2023-08-29
- Filing Date
- 2024-08-29
- Publication Date
- 2026-01-30
- Estimated Expiration
- 2044-08-29
AI Technical Summary
Conventional speech recognition methods face challenges due to inconsistent lengths of feature representations caused by varying speaking speeds and intonation, leading to reduced accuracy and computational efficiency.
The method involves obtaining first speech features with multiple segment features, decoding them to get an initial recognition result, extracting word-level audio features using the initial result as a priori information, and then decoding these features to achieve uniform representations of equal length, thereby improving accuracy and efficiency.
This approach enhances speech recognition accuracy and computational efficiency by addressing the inconsistency in feature lengths and reducing redundant features, enabling faster and more precise speech recognition.
Smart Images

Figure 0007809174000002 
Figure 0007809174000003 
Figure 0007809174000004
Abstract
Description
[Technical Field]
[0001] The present disclosure relates to the field of artificial intelligence technology, in particular to technical fields such as speech recognition and deep learning, and specifically to a speech recognition method, a method for training a deep learning model for speech recognition, a speech recognition device, an apparatus for training a deep learning model for speech recognition, an electronic device, a computer-readable storage medium, and a computer program product. [Background technology]
[0002] Artificial intelligence is a field that studies how computers can imitate some human thought processes and intelligent behaviors (e.g., learning, reasoning, thinking, planning, etc.), and includes both hardware and software technologies. AI hardware technology generally includes sensors, dedicated AI chips, cloud computing, distributed storage, and big data processing, while AI software technology mainly includes natural language processing technology, computer vision technology, speech recognition technology, and several major directions such as machine learning / deep learning, big data processing technology, and knowledge graph technology.
[0003] Automatic speech recognition (ASR) is a technology that automatically converts speech signals input via a computer into corresponding text. Deep research into deep learning technologies in the field of speech recognition, especially the development of end-to-end speech recognition technologies, has significantly improved the accuracy of speech recognition while reducing the modeling complexity. With the increasing popularity of various intelligent devices, large vocabulary online speech recognition systems are widely used in various scenarios, such as speech transcription, intelligent customer service, in-vehicle navigation, and smart homes. In these speech recognition tasks, users typically expect fast and accurate responses and feedback from the system after completing their voice input, which places significant demands on the accuracy and real-time factors of speech recognition models.
[0004] The approaches described in this section are not necessarily approaches that have been previously conceived or adopted. Unless otherwise noted, any approach described in this section should not be considered prior art merely because it is included in this section. Likewise, unless otherwise noted, the problems addressed in this section should not be considered to be acknowledged in the prior art. Summary of the Invention
[0005] The present disclosure provides a speech recognition method, a method for training a deep learning model for speech recognition, a speech recognition device, an device for training a deep learning model for speech recognition, an electronic device, a computer-readable storage medium, and a computer program product.
[0006] According to one aspect of the present disclosure, there is provided a speech recognition method, the speech recognition method including: acquiring first speech features of speech to be recognized, the first speech features including a plurality of speech segment features corresponding to a plurality of speech segments in the speech to be recognized; decoding the first speech features using a first decoder to obtain a plurality of first decoding results corresponding to a plurality of words in the speech to be recognized, the first decoding results indicating a first recognition result of the corresponding words; extracting and obtaining second speech features from the first speech features based on first a priori information, the first a priori information including the plurality of first decoding results, the second speech features including a plurality of first word-level audio features corresponding to the plurality of words; and decoding the second speech features using a second decoder to obtain a plurality of second decoding results corresponding to the plurality of words, the second decoding results indicating a second recognition result of the corresponding words.
[0007] According to another aspect of the present disclosure, there is provided a method for training a deep learning model for speech recognition, the deep learning model including a first decoder and a second decoder, and the training method includes: obtaining sample speech and actual recognition results of a plurality of words in the sample speech; obtaining first sample speech features of the sample speech, the first sample speech features including a plurality of sample speech segment features corresponding to a plurality of sample speech segments in the sample speech; decoding the first sample speech features using the first decoder to obtain a plurality of first sample decoding results corresponding to the plurality of words in the sample speech, the first sample decoding results indicating first recognition results of the corresponding words; The method includes extracting and obtaining second sample speech features from the first sample speech features based on sample a priori information, where the first sample a priori information includes a plurality of first sample decoding results, and the second sample speech features include a plurality of first sample word-level audio features corresponding to a plurality of words; decoding the second sample speech features using a second decoder to obtain a plurality of second sample decoding results corresponding to the plurality of words, where the second sample decoding results indicate second recognition results of the corresponding words; and adjusting parameters of the deep learning model based on the actual recognition results of the plurality of words, the first recognition results, and the second recognition results, to obtain a trained deep learning model.
[0008] According to another aspect of the present disclosure, there is provided a speech recognition apparatus, the apparatus including: an audio feature encoding module configured to acquire first audio features of audio to be recognized, the first audio features including a plurality of audio segment features corresponding to a plurality of audio segments in the audio to be recognized; a first decoder configured to decode the first audio features and acquire a plurality of first decoding results corresponding to a plurality of words in the audio to be recognized, the first decoding results indicating a first recognition result for the corresponding words; a word-level feature extraction module configured to extract and acquire second audio features from the first audio features based on first a priori information, the first a priori information including the plurality of first decoding results, the second audio features including a plurality of first word-level audio features corresponding to the plurality of words; and a second decoder configured to decode the second audio features and acquire a plurality of second decoding results corresponding to the plurality of words, the second decoding results indicating a second recognition result for the corresponding words.
[0009] According to another aspect of the present disclosure, there is provided an apparatus for training a deep learning model for speech recognition, the deep learning model including a first decoder and a second decoder, the training apparatus including: an acquisition module configured to acquire a sample speech and an actual recognition result of a plurality of words in the sample speech; an audio feature encoding module configured to acquire first sample speech features of the sample speech, the first sample speech features including a plurality of sample speech segment features corresponding to a plurality of sample speech segments in the sample speech; a first decoder configured to decode the first sample speech features and obtain a plurality of first sample decoding results corresponding to the plurality of words in the sample speech, the first sample decoding results indicating a first recognition result of the corresponding words; the first sample a priori information includes a plurality of first sample decoding results, and the second sample audio features include a plurality of first sample word-level audio features corresponding to a plurality of words; a second decoder configured to decode the second sample audio features and obtain a plurality of second sample decoding results corresponding to the plurality of words, wherein the second sample decoding results indicate second recognition results of the corresponding words; and a parameter adjustment module configured to adjust parameters of a deep learning model based on the actual recognition results of the plurality of words, the first recognition results, and the second recognition results, to obtain a trained deep learning model.
[0010] According to another aspect of the present disclosure, there is provided an electronic device including at least one processor and a memory communicatively coupled to the at least one processor, wherein the memory stores instructions executable by the at least one processor, the instructions being executed by the at least one processor to cause the at least one processor to perform the method set forth above.
[0011] According to another aspect of the present disclosure, there is provided a non-transitory computer-readable storage medium having stored thereon computer instructions for causing a computer to perform the method described above.
[0012] According to another aspect of the present disclosure, there is provided a computer program product including a computer program which, when executed by a processor, implements the above-mentioned method.
[0013] According to one or more embodiments of the present disclosure, the present disclosure obtains and decodes first speech features including multiple speech segment features of a speech to be recognized, obtains an initial recognition result for the speech to be recognized, further uses the initial recognition result to extract word-level audio features from the first speech features, and then decodes the word-level audio features to obtain a final recognition result.
[0014] Using the initial recognition result for the speech to be recognized as a priori, a unified audio feature representation of equal length at the word level is extracted and obtained from the speech feature information of unequal length in the frame-level audio information, and the word-level audio features are decoded to obtain the final recognition result, thereby solving the problem of inconsistent lengths of feature representations of speech framing in the conventional technology, improving the accuracy of speech recognition and increasing computational efficiency.
[0015] It should be understood that the contents described in this section are not intended to identify key or important features of the embodiments of the present disclosure, and are not intended to limit the scope of protection of the present disclosure. Other features of the present disclosure will be easily understood from the following description. [Brief explanation of the drawings]
[0016] The drawings illustratively illustrate examples, constitute a part of the specification, and together with the written description serve to explain exemplary embodiments of the examples. The illustrated examples are for illustrative purposes only and do not limit the scope of the claims. In all drawings, the same reference numerals refer to similar, but not necessarily identical, elements.
[0017] [Figure 1] FIG. 1 is a schematic diagram illustrating an example system capable of implementing various methods described herein, according to embodiments of the present disclosure. [Figure 2] FIG. 1 is a flow chart diagram illustrating a speech recognition method according to an embodiment of the present disclosure. [Figure 3] FIG. 1 is a flowchart diagram illustrating obtaining first speech features of speech to be recognized according to an embodiment of the present disclosure. [Figure 4] FIG. 1 is a schematic diagram illustrating a Conformer streaming multi-layer disconnected attention model based on historical feature abstraction, according to an embodiment of the present disclosure. [Figure 5] FIG. 10 is a flowchart diagram of extracting and obtaining a second audio feature from a first audio feature according to an embodiment of the present disclosure. [Figure 6] FIG. 1 is a flow chart diagram illustrating a speech recognition method according to an embodiment of the present disclosure. [Figure 7] FIG. 1 is a flow chart diagram illustrating a speech recognition method according to an embodiment of the present disclosure. [Figure 8] FIG. 1 is a schematic diagram illustrating an end-to-end speech large model according to an embodiment of the present disclosure. [Figure 9] FIG. 1 is a flowchart illustrating a method for training a deep learning model for speech recognition, according to an embodiment of the present disclosure. [Figure 10] 1 is a configuration block diagram illustrating a speech recognition device according to an embodiment of the present disclosure. [Figure 11] FIG. 1 is a block diagram illustrating an apparatus for training a deep learning model for speech recognition according to an embodiment of the present disclosure. [Figure 12] FIG. 1 is a block diagram illustrating an exemplary electronic device that can be used to implement embodiments of the present disclosure. DETAILED DESCRIPTION OF THE INVENTION
[0018]
[0023] The following description will be made in conjunction with the drawings to illustrate exemplary embodiments of the present disclosure. Various details of the embodiments of the present disclosure are included to facilitate understanding and should be considered merely exemplary. Therefore, those skilled in the art should recognize that various changes and modifications can be made to the embodiments described herein without departing from the scope and spirit of the present disclosure. Similarly, for the sake of clarity and conciseness, the following description omits descriptions of known functions and structures.
[0019] In this disclosure, unless otherwise specified, the terms "first," "second," and the like, used to describe various elements are not intended to limit the location, timing, or importance of these elements. Such terms are used only to distinguish one element from another. In some instances, a first element and a second element may refer to the same instance of an element, or in some cases, may refer to different instances based on the context.
[0020] The terms used in the description of various examples of the present disclosure are intended only to describe particular examples and are not intended to be limiting. Unless the context clearly indicates otherwise, an element may be one or more unless the number of elements is specifically limited. Furthermore, the term "and / or" as used in this disclosure covers any and all possible combinations of the listed items.
[0021] In the related art, some speech recognition methods use framing speech features to learn audio feature representations. However, the content information contained in speech constantly changes with the speaker's speaking speed, intonation, tone, etc., and different speakers will express the same content in completely different ways. Therefore, this feature representation method will result in inconsistencies in the length of the representations obtained using framing speech features, which will affect the accuracy of speech recognition. In addition, the obtained feature representation will contain a large amount of redundant features, resulting in low computational efficiency.
[0022] To solve the above problem, the present disclosure obtains and decodes first speech features, including multiple speech segment features, of speech to be recognized, to obtain an initial recognition result for the speech to be recognized, and then uses the initial recognition result to extract word-level audio features from the first speech features, and then decodes the word-level audio features to obtain a final recognition result. Using the initial recognition result for the speech to be recognized as a priori, extract and obtain uniform audio feature representations of equal length at the word level from speech feature information of unequal lengths in frame-level audio information, and decodes the word-level audio features to obtain a final recognition result, thereby solving the problem of inconsistent lengths of feature representations of speech framing in the conventional technology, and improving speech recognition accuracy and computational efficiency.
[0023] Hereinafter, embodiments of the present disclosure will be described in detail with reference to the drawings. 1 illustrates a schematic diagram of an exemplary system 100 in which various methods and apparatus described herein may be implemented, according to embodiments of the present disclosure. Referring to FIG. 1, the system 100 includes one or more client devices 101, 102, 103, 104, 105, and 106, a server 120, and one or more communication networks 110 coupling the one or more client devices to the server 120. The client devices 101, 102, 103, 104, 105, and 106 may be configured to run one or more applications.
[0024] In an embodiment of the present disclosure, the server 120 operates to execute one or more services or software applications of the speech recognition method and / or the method for training a deep learning model for speech recognition of the present disclosure. In one exemplary embodiment, a complete speech recognition system or an assembly of parts of a speech recognition system, such as a speech large model, can be deployed on the server.
[0025] In some embodiments, server 120 may also provide other services or software applications, which may include non-virtualized and virtualized environments. In some embodiments, these services may be provided as web-based or cloud services, for example, provided to users of client devices 101, 102, 103, 104, 105, and / or 106 in a Software as a Service (SaaS) model.
[0026] In the configuration shown in FIG. 1 , server 120 may include one or more assemblies that implement the functionality performed by server 120. These assemblies may include software assemblies, hardware assemblies, or a combination thereof, executable on one or more processors. Users operating client devices 101, 102, 103, 104, 105, and / or 106 may interact with server 120 using one or more client applications to utilize the services provided by these assemblies. It should be understood that a variety of different system configurations are possible and may differ from system 100. Thus, FIG. 1 is intended to be illustrative of an example system for implementing various methods described herein and is not intended to be limiting.
[0027] A user can input speech to be recognized using client devices 101, 102, 103, 104, 105, and / or 106. The client devices can provide an interface through which a user of the client device interacts with the client device. The client devices can also output information to the user through the interface, such as outputting speech recognition results to the user. Although only six client devices are shown in FIG. 1 , one skilled in the art will understand that the present disclosure can support any number of client devices.
[0028] Client devices 101, 102, 103, 104, 105, and / or 106 may include various types of computing devices, such as portable handheld devices, general-purpose computers (e.g., personal computers or laptops), workstation computers, wearable devices, smart screen devices, self-service terminal devices, service robots, gaming systems, thin clients, various messaging devices, sensors, or other sensing devices. These computing devices may run various types and versions of software applications and operating systems, such as Microsoft Windows, Apple iOS, UNIX-like operating systems, Linux or Linux-like operating systems (e.g., Google Chrome OS), or include various mobile operating systems, such as Microsoft Windows Mobile OS, iOS, Windows Phone, and Android. Portable handheld devices may include mobile phones, intelligent phones, tablets, personal digital assistants (PDAs), and the like. Wearable devices may include head-mounted displays (e.g., smart glasses) and other devices. Gaming systems may include various handheld gaming devices, Internet-enabled gaming devices, and the like. The client device may run a variety of applications, such as Internet-related applications, communication applications (eg, email applications), short message service (SMS) applications, and may use a variety of communication protocols.
[0029] Network 110 may be any type of network known to those skilled in the art, which may use any one of several available protocols to support data communications (including, but not limited to, TCP / IP, SNA, IPX, etc.) By way of example, one or more networks 110 may be a local area network (LAN), an Ethernet-based network, a token loop, a wide area network (WAN), the Internet, a virtual network, a virtual private network (VPN), an intranet, an extranet, a blockchain network, a public switched telephone network (PSTN), an infrared network, a wireless network (e.g., Bluetooth, WIFI), and / or any combination of these and / or other networks.
[0030] Server 120 may include one or more general-purpose computers, dedicated server computers (e.g., PC (personal computer) servers, UNIX servers, mid-range servers), blade servers, mainframes, server clusters, or any other suitable arrangement and / or combination. Server 120 may also include one or more virtual machines running virtual operating systems or other computing architectures involving virtualization (e.g., one or more flexible pools of virtualized logical storage devices to maintain the server's virtual storage). In various embodiments, server 120 may run one or more services or software applications that provide the functionality described below.
[0031] The computing units in server 120 may run one or more operating systems, including any of the operating systems listed above and any commercial server operating system. Server 120 may also run any one of a variety of additional server and / or middle-tier applications, including an HTTP server, an FTP server, a CGI server, a JAVA server, a database server, etc.
[0032] In some embodiments, server 120 may include one or more applications for analyzing and consolidating data feeds and / or event updates received from users of client devices 101, 102, 103, 104, 105, and / or 106. Server 120 may include one or more applications for displaying data feeds and / or real-time events via one or more display devices of client devices 101, 102, 103, 104, 105, and / or 106.
[0033] In some embodiments, server 120 may be a server in a distributed system or a server incorporating blockchain. Server 120 may be a cloud server, or an intelligent cloud computing server or intelligent cloud host equipped with artificial intelligence technology. A cloud server is a host product in a cloud computing service system that solves the drawbacks of traditional physical hosts and virtual private server (VPS) services, such as high management difficulty and poor business scalability.
[0034] System 100 may include one or more databases 130. In some embodiments, these databases may be used to store data or other information. For example, one or more of databases 130 may be used to store information such as audio files or video files. Databases 130 may be located in a variety of locations. For example, a database used by server 120 may be local to server 120 or may be remote from server 120 and in communication with server 120 over a network or dedicated connection. Databases 130 may be of a variety of types. In some embodiments, a database used by server 120 may be a database, such as a relational database. One or more of these databases may store, update, and retrieve data from the database in response to instructions.
[0035] In some embodiments, one or more of databases 130 may be used by an application to store data for the application. The databases used by the application may be various types of databases, such as key-value repositories, object repositories, or general-purpose repositories supported by a file system.
[0036] The system 100 of FIG. 1 can be configured and operated in a variety of ways to accommodate the various methods and apparatus described in accordance with this disclosure. According to one aspect of the present disclosure, there is provided a speech recognition method. As shown in Figure 2, the speech recognition method includes: step S201, acquiring first speech features of speech to be recognized, the first speech features including a plurality of speech segment features corresponding to a plurality of speech segments in the speech to be recognized; step S202, using a first decoder to decode the first speech features and obtain a plurality of first decoding results corresponding to a plurality of words in the speech to be recognized, the first decoding results indicating a first recognition result of the corresponding words; step S203, extracting and obtaining second speech features from the first speech features based on first a priori information, the first a priori information including the plurality of first decoding results, the second speech features including a plurality of first word-level audio features corresponding to the plurality of words; and step S204, using a second decoder to decode the second speech features and obtain a plurality of second decoding results corresponding to the plurality of words, the second decoding results indicating a second recognition result of the corresponding words.
[0037] This solves the problem of inconsistent lengths of feature representations of speech framing in the past, by using the initial recognition result for the speech to be recognized as a priori, extracting and obtaining unified audio feature representations of equal length at the word level from the speech feature information of unequal lengths in the frame-level audio information, and decoding the word-level audio features to obtain the final recognition result. This improves the accuracy of speech recognition and increases the computational efficiency.
[0038] To facilitate the description of the technical concept, the speech to be recognized in the embodiments of the present disclosure includes speech content corresponding to multiple words. In step S201, various existing audio feature extraction methods can be adopted to obtain first audio features of the speech to be recognized. The multiple speech segments may be obtained by segmenting the speech to be recognized into fixed lengths, or by other segmentation methods. Multiple speech segment features may correspond one-to-one to multiple speech segments, and the same speech segment may correspond to multiple speech segment features (as described below), but this is not limited thereto.
[0039] According to some embodiments, as shown in FIG. 3, step S201, obtaining first speech features of speech to be recognized, may include step S301, obtaining original speech features of speech to be recognized; step S302, determining multiple spikes in the speech to be recognized based on the original speech features; and step S303, cutting the original speech features to obtain multiple speech segment features that correspond one-to-one to the multiple spikes.
[0040] Since spike signals usually correspond to each word in the speech to be recognized, first, the spike signals of the speech to be recognized are obtained, and then multiple speech segment features that correspond one-to-one to the multiple spikes are obtained based on the spike information, so that the first decoder decodes the first speech features under the guidance of the spike information, and obtains an accurate initial recognition result.
[0041] In step S301, speech features are extracted for a plurality of speech frames included in the speech to be recognized, and original speech features including a plurality of speech frame features can be obtained.
[0042] In step S302, a binary CTC (Connectionist Temporal Classification) module modeled based on a Causal Conformer is used to process the original speech features, thereby obtaining CTC spike information and determining multiple spikes in the speech to be recognized. As can be understood, multiple spikes in the speech to be recognized can also be determined by other methods, which are not limited here.
[0043] In step S303, the cutting into original audio features may be to cut a plurality of audio frame features corresponding to a plurality of audio frames into a plurality of sets of audio frame features, and each set of audio frame / audio frame features constitutes one audio segment / audio segment feature.
[0044] According to some embodiments, step S303, cutting the original audio features and obtaining a plurality of audio segment features corresponding one-to-one to the plurality of spikes, may include cutting the original audio features according to a predetermined time length, and taking the audio segment feature of the audio segment in which each of the plurality of spikes exists as the audio segment feature corresponding to the spike. Thus, in this manner, the audio segment features corresponding to each spike have the same length. It should be noted that in this manner, if a audio segment contains more than one spike, the audio segment features of the audio segment may simultaneously correspond to each of the spikes.
[0045] As can be appreciated, the preset time length d can be set as needed. In the embodiment described in Figure 4, the preset time length d is five audio frames.
[0046] According to some embodiments, step S303, cutting the original audio features and obtaining a plurality of audio segment features corresponding one-to-one to the plurality of spikes, may include cutting the original audio features according to the plurality of spikes, and taking the audio segment feature between each two adjacent spikes as the audio segment feature corresponding to one of the spikes, so that the audio segment feature corresponding to each spike contains the complete audio information of the audio segment between the two adjacent spikes.
[0047] In some embodiments, the original speech features may be downsampled (eg, convolutionally downsampled) before being used (by the CTC module or initial speech recognition) on the original speech features.
[0048] According to some embodiments, the plurality of speech segment features may be sequentially obtained by streaming-cutting the original speech features. Step S202, decoding the first speech feature using the first decoder, may include streaming-decoding the plurality of speech segment features sequentially using the first decoder. In this way, by streaming-cutting the original speech features and streaming-decoding the first speech features, an initial recognition result for the speech to be recognized can be quickly obtained.
[0049] According to some embodiments, the speech segment features can be further encoded using a scheme based on historical feature abstraction, thereby enhancing the descriptive power of the speech segment features and improving the accuracy of the initial recognition results obtained after decoding the speech segment features. As shown in Figure 3, step S201, obtaining first speech features of the speech to be recognized, can include step S304, obtaining corresponding historical feature abstraction information for the currently obtained speech segment features, where the historical feature abstraction information is obtained by attention modeling the previous speech segment features using a first decoding result corresponding to the previous speech segment features; and step S305, using a first encoder to encode the currently obtained speech segment features in combination with the historical feature abstraction information to obtain the corresponding enhanced speech segment features.
[0050] In some embodiments, the historical feature abstraction information corresponding to the currently obtained speech segment feature includes historical feature abstraction information corresponding to each of a plurality of previous speech segment features, and the historical feature abstraction information of each previous speech segment feature is obtained by attention modeling the previous speech segment feature using a first decoding result corresponding to the previous speech segment feature. In one exemplary embodiment, the first decoding result is used as a query feature Q, and the previous speech segment features are used as key features K and value features V to perform calculation of an attention mechanism, thereby obtaining the historical feature abstraction information of the previous speech segment feature. The calculation process of the attention mechanism can be represented as follows:
[0051]
number
[0052] where d kis the dimension of the feature. It can be understood that the calculation of other feature acquisition and attention mechanisms based on the query features, key features, and value features in this disclosure can all refer to this formula. It should be noted that the number of features obtained by this method is the same as the number of features included in the query features.
[0053] According to some embodiments, step S305, using a first encoder to encode the currently obtained speech segment features in combination with the historical feature abstraction information to obtain corresponding enhanced speech segment features, may include obtaining the corresponding enhanced speech segment features output by the first encoder by using the currently obtained speech segment features as query features Q of the first encoder and splicing results of the historical feature abstraction information and the currently obtained speech segment features as key features K and value features V of the first encoder.
[0054] This enables the method to fully discover more timing and linguistic relationships in speech features, greatly improving the model's ability to abstract history, and also improving the accuracy of the decoding results for the enhanced speech segment features.
[0055] In some embodiments, the first encoder and the first decoder together may constitute a Streaming Multi-Layer Truncated Attention (SMLTA) model based on historical feature abstraction. As shown in FIG. 4, the Conformer SMLTA model 400 mainly includes two parts: a Streaming Truncated Conformer Encoder 402 (i.e., the first encoder) and a Transformer Decoder 404 (i.e., the first decoder). The Streaming Truncated Conformer Encoder includes N stacked Conformer modules, each of which includes a feedforward module 406, a multi-head self-attention module 408, a convolution module 410, and a feedforward module 412. The Conformer module encodes the speech segment features layer by layer and obtains corresponding hidden features (i.e., the speech segment features after enhancement). The Transformer decoder contains M stacked Transformer modules, and uses a streaming attention mechanism to screen the hidden features output by the encoder and output a first decoded result that indicates the initial recognition result.
[0056] Figure 4 further illustrates the principle of Conformer SMLTA based on historical feature abstraction. The input original speech features 414 are first divided into speech segment features of the same length, and then the streaming Conformer encoder performs feature encoding on each speech segment feature. The Transformer decoder counts the number of spikes contained in each audio segment according to the spike information 416 of the binary CTC model, and decodes and outputs the recognition result of the current segment according to the spike number. Finally, according to the decoding result of the current segment, the hidden features of each layer of the Conformer encoder are modeled using correlation attention to obtain the historical feature abstraction contained in the corresponding speech segment. The historical feature abstraction information obtained by the abstraction of each layer is spliced with the currently obtained speech segment feature to calculate the next segment.
[0057] According to some embodiments, as shown in FIG. 5 , step S203, extracting and obtaining a second audio feature from the first audio feature based on the first a priori information, may include step S501, for each word of the plurality of words, obtaining a first word-level audio feature corresponding to the word, which is output by the attention module, by taking the first decoding result corresponding to the word as a query feature Q of the attention module, and the first audio feature as a key feature K and a value feature V of the attention module.
[0058] In this way, by using the first decoding results corresponding to each of the multiple words as query features Q and the first speech features as key features K and value features V, the initial recognition results for the speech to be recognized can be effectively used as a priori information, and word-level audio features corresponding to each word can be obtained.
[0059] In some embodiments, the first word-level audio features output by the attention module can be obtained by substituting the corresponding Q, K, and V into the formula of the attention mechanism described above and calculating them.
[0060] According to some embodiments, as shown in FIG. 5 , step S203, extracting and obtaining second speech features from the first speech features based on the first a priori information, may include step S502, using a second encoder to globally encode the plurality of first word-level audio features corresponding to the plurality of words to obtain enhanced second speech features.
[0061] By globally encoding multiple first word-level audio features corresponding to multiple words, the first encoder effectively compensates for the inability to encode global feature information due to the need to meet streaming recognition requirements, and significantly improves the description capability of the isometric unified feature representation.
[0062] In some embodiments, the second encoder may be a conformer encoder that can include N layers of stacked conformer modules. The conformer modules simultaneously combine attention and convolutional models, effectively modeling both long-range and local relationships in audio features, greatly improving the model's descriptive power.
[0063] As can be understood, extracting and obtaining the second audio feature from the first audio feature based on the first a priori information can be achieved by methods other than the attention mechanism and the conformer encoder, but this is not limited thereto.
[0064] According to some embodiments, step S204, using a second decoder to decode the second audio features and obtain a plurality of second decoded results corresponding to the plurality of words, may include, for each word of the plurality of words, obtaining a second decoded result corresponding to the word output by the second decoder by using the first decoded result corresponding to the word as a query feature Q of the second decoder and the second audio features as a key feature K and a value feature V of the second decoder.
[0065] As a result, by using the first decoding results corresponding to each of the multiple words as query features Q and the second speech features as key features K and value features V, the initial recognition results for the speech to be recognized can be effectively used as a priori information, and the second decoding results corresponding to each word can be obtained.
[0066] Conventional encoder-decoder or decoder-only architectures face cache loading issues during decoding. While GPU computation speeds have improved significantly, the speed at which decoders load model parameters into the cache during computation remains limited by the development of computer hardware resources, severely limiting the decoding efficiency of speech recognition models. Whether a speech recognition model has an encoder-decoder architecture or a decoder-only architecture, it must rely on the previous decoding result before it can perform the next computation. This recursive computation method requires the model to be repeatedly loaded into the cache, resulting in a certain degree of computation delay. As the number of parameters in a large speech model increases, the computation delay caused by cache loading becomes even more pronounced, making it impossible to meet the real-time decoding requirements of online decoding. However, by using the first decoding results corresponding to each of the multiple words already obtained as query features for the second decoder, the final recognition result can be obtained by performing parallel calculations only once, thereby effectively solving the cache load problem faced by large-scale models.
[0067] According to some embodiments, the second decoder may include a forward decoder and a backward decoder, both of which may be configured, for each word of the plurality of words, to take the first decoding result of the word as the input query feature Q and the second audio features as the input key feature K and value feature V, and the forward decoder may be configured to time mask the input features from left to right, and the backward decoder may be configured to time mask the input features from right to left.
[0068] This enables language modeling in two different directions by installing a forward decoder that temporally masks input features from left to right and a backward decoder that temporally masks input features from right to left, realizing simultaneous modeling of language context and further improving the model's predictive ability.
[0069] In some embodiments, the forward decoder may be referred to as a left-right Transformer decoder and the backward decoder may be referred to as a right-left Transformer decoder. Both the forward and backward decoders may include K stacked time mask Transformer modules.
[0070] According to some embodiments, for each word of the plurality of words, obtaining a second decoding result corresponding to the word output by the second decoder by using the first decoding result of the word as a query feature Q of the second decoder and the second audio feature as a key feature K and a value feature V of the second decoder may include fusing a plurality of forward decoding features corresponding to the plurality of words output by the forward decoder and a plurality of backward decoding features corresponding to the plurality of words output by the backward decoder to obtain a plurality of fused features corresponding to the plurality of words; and obtaining a plurality of second decoding results based on the plurality of fused features.
[0071] In some embodiments, the forward decoded features and the backward decoded features can be directly added to obtain the corresponding fused features, which can then be processed, such as Softmax, to obtain the final recognition result.
[0072] After obtaining the second decoding result, the second decoding result can be used as a priori information for the recognition result to re-extract word-level audio features or reuse the second decoder for decoding.
[0073] According to some embodiments, as shown in Figure 6, the speech recognition method may further include: in step S605, for each word of the plurality of words, obtaining an (N+1)th decoding result corresponding to the word output by the second decoder by taking the Nth decoding result of the word as a query feature Q of the second decoder, and the second speech features as a key feature K and a value feature V of the second decoder, where N is an integer greater than or equal to 2. It can be understood that the operations of steps S601 to S604 in Figure 6 are similar to the operations of steps S201 to S204 in Figure 2, and therefore, description thereof will be omitted here.
[0074] Therefore, by using the second decoder to perform multiple iterative decoding, the accuracy rate of speech recognition can be improved. According to some embodiments, as shown in FIG. 7, the speech recognition method may include: step S705, extracting and obtaining third speech features from the first speech features based on second a priori information, where the second a priori information includes a plurality of second decoding results, and the third speech features include a plurality of second word-level audio features corresponding to a plurality of words; and step S706, using a second decoder to decode the third speech features and obtain a plurality of third decoding results corresponding to the plurality of words, where the third decoding results indicate third recognition results of the corresponding words.
[0075] Therefore, the accuracy of speech recognition can be further improved by using the second decoding result as a priori for the recognition result, re-extracting word-level audio features, and then using a second decoder to decode the new word-level audio features.
[0076] As can be understood, the operations of steps S701 to S704 in FIG. 7 are similar to the operations of steps S201 to S204 in FIG. 2, and therefore will not be described here.
[0077] According to some embodiments, the second decoder may be a speech large model. The model size of the second decoder can reach billions of parameters, which can fully extract linguistic information contained in speech and greatly improve the modeling ability. In some exemplary embodiments, the parameter amount of the speech large model, which is the second decoder, can be 2B, but can also be other parameter amounts of 1 billion or more.
[0078] In some embodiments, the model size of the first decoder (or the model formed by the first encoder and the first decoder) may be, for example, several hundred megabytes. Since its function is to stream out the initial recognition result for the speech to be recognized, large parameters are not required.
[0079] In some embodiments, as shown in FIG. 8 , a first encoder 810 (SMLTA2 Encoder), a first decoder 820 (SMLTA2 Decoder), an attention module 830 (Attention Module), a second encoder 840 (Conformer Encoder), and a second decoder 850 (including a forward decoder 860 (Left-Right Transformer Decoder) and a backward decoder 870 (Right-Left Transformer Decoder)) can together constitute an end-to-end speech large model 800.
[0080] According to another aspect of the present disclosure, there is provided a method for training a deep learning model for speech recognition. The deep learning model includes a first decoder and a second decoder. As shown in FIG. 9 , the training method includes: step S901: obtaining a sample speech and an actual recognition result of a plurality of words in the sample speech; step S902: obtaining first sample speech features of the sample speech, the first sample speech features including a plurality of sample speech segment features corresponding to a plurality of sample speech segments in the sample speech; step S903: using the first decoder to decode the first sample speech features to obtain a plurality of first sample decoding results corresponding to a plurality of words in the sample speech, the first sample decoding results indicating first recognition results of the corresponding words; and step S904: decoding the first sample speech features based on the first sample a priori information. 9 includes: extracting and obtaining second sample speech features from the features, where the first sample a priori information includes a plurality of first sample decoding results, and the second sample speech features include a plurality of first sample word-level audio features corresponding to a plurality of words; step S905; decoding the second sample speech features using a second decoder to obtain a plurality of second sample decoding results corresponding to the plurality of words, where the second sample decoding results indicate second recognition results for the corresponding words; and step S906; adjusting parameters of the deep learning model based on the actual recognition results of the plurality of words, the first recognition results, and the second recognition results to obtain a trained deep learning model. As can be understood, the operations of steps S902 to S905 in FIG. 9 are similar to the operations of steps S201 to S204 in FIG. 2, and therefore will not be described here.
[0081] As a result, the deep learning model trained by the above method takes the initial recognition result for the speech to be recognized as a priori, extracts and obtains uniform audio feature representations of equal length at the word level from the speech feature information of unequal lengths in the frame-level audio information, and decodes the word-level audio features to obtain the final recognition result, thereby solving the problem of inconsistent lengths of feature representations of speech framing in the conventional technology, improving the accuracy of speech recognition and increasing computational efficiency.
[0082] In some embodiments, the deep learning model may also include other modules related to the above-mentioned speech recognition method, such as a first encoder, a second encoder, an attention module, etc. The operation of each module in the deep learning model may refer to the operation of the corresponding module in the above-mentioned speech recognition method.
[0083] In some embodiments, in step S906, a first loss value may be determined based on the actual recognition result and the second recognition result, and parameters of the deep learning model may be adjusted based on the first loss value. In some embodiments, a second loss value may be determined based on the actual recognition result and the first recognition result, and parameters of the deep learning model may be adjusted based on the first loss value and the second loss value. In some embodiments, the second loss value may be used to adjust parameters of the first decoder (and the first encoder), and the first loss value may be used to adjust parameters of the second decoder (and the attention module, the second encoder), and these may be used to adjust parameters of the deep learning model end-to-end. Furthermore, some modules of the deep learning model may be individually trained or pre-trained in advance. As can be understood, parameters of the deep learning model may be adjusted by other methods, and this is not limited thereto.
[0084] As can be appreciated, the speech recognition method described above can be implemented using a deep learning model obtained by training according to the training method described above. According to another aspect of the present disclosure, there is provided a speech recognition apparatus. As shown in Figure 10, the apparatus 1000 includes: an audio feature encoding module 1010 configured to acquire first audio features of a speech to be recognized, the first audio features including a plurality of audio segment features corresponding to a plurality of audio segments in the speech to be recognized; a first decoder 1020 configured to decode the first audio features and acquire a plurality of first decoding results corresponding to a plurality of words in the speech to be recognized, the first decoding results indicating a first recognition result of the corresponding words; a word-level feature extraction module 1030 configured to extract and acquire second audio features from the first audio features based on first a priori information, the first a priori information including the plurality of first decoding results, the second audio features including a plurality of first word-level audio features corresponding to the plurality of words; and a second decoder 1040 configured to decode the second audio features and acquire a plurality of second decoding results corresponding to the plurality of words, the second decoding results indicating a second recognition result of the corresponding words. As can be seen, the operations of modules 1010 to 1040 of the device 1000 are similar to the operations of steps S201 to S204 in FIG. 2, and therefore will not be described here.
[0085] According to some embodiments, the audio feature encoding module 1010 may be configured to obtain original audio features of the audio to be recognized, determine a plurality of spikes in the audio to be recognized based on the original audio features, and cut the original audio features to obtain a plurality of audio segment features that correspond one-to-one to the plurality of spikes.
[0086] According to some embodiments, cutting the original audio features and obtaining a plurality of audio segment features that correspond one-to-one to the plurality of spikes may include cutting the original audio features based on a predetermined time length, and taking the audio segment feature of the audio segment in which each spike of the plurality of spikes exists as the audio segment feature corresponding to the spike.
[0087] According to some embodiments, cutting the original audio features and obtaining a plurality of audio segment features corresponding one-to-one to the plurality of spikes may include cutting the original audio features based on the plurality of spikes, and taking the audio segment feature between each two adjacent spikes as the audio segment feature corresponding to one of the spikes.
[0088] According to some embodiments, the plurality of audio segment features may be obtained in sequence by streaming truncating the original audio features, and the first decoder may be configured to stream-decode the plurality of audio segment features in sequence.
[0089] According to some embodiments, the speech feature encoding module may be configured to obtain corresponding historical feature abstraction information for the currently obtained speech segment feature, where the historical feature abstraction information is obtained by attention modeling the previous speech segment feature using a first decoding result corresponding to the previous speech segment feature. The speech feature encoding module may include a first encoder configured to encode the currently obtained speech segment feature in combination with the historical feature abstraction information and output the corresponding enhanced speech segment feature.
[0090] According to some embodiments, the first encoder may be configured to receive currently obtained speech segment features as query features of the first encoder, and output corresponding enhanced speech segment features by receiving splicing results of the historical feature abstraction information and the currently obtained speech segment features as key features and value features of the first encoder.
[0091] According to some embodiments, the word-level feature extraction module may include an attention module configured to, for each word of the plurality of words, receive a first decoding result corresponding to the word as a query feature of the attention module, and receive first audio features as key features and value features of the attention module, thereby outputting a first word-level audio feature corresponding to the word.
[0092] According to some embodiments, the word-level feature extraction module may include a second encoder configured to obtain enhanced second audio features by globally encoding a plurality of first word-level audio features corresponding to a plurality of words.
[0093] According to some embodiments, the second decoder may be configured to, for each word of the plurality of words, output a second decoding result corresponding to the word by receiving a first decoding result corresponding to the word as a query feature of the second decoder and receiving second audio features as key features and value features of the second decoder.
[0094] According to some embodiments, the second decoder may include a forward decoder and a backward decoder, both of which are configured to receive, for each word of the plurality of words, a first decoding result of the word as an input query feature and a second audio feature as an input key feature and value feature, and the forward decoder is configured to time mask the input features from left to right, and the backward decoder is configured to time mask the input features from right to left.
[0095] According to some embodiments, the second decoder may be configured to fuse a plurality of forward decoding features corresponding to a plurality of words output by the forward decoder and a plurality of backward decoding features corresponding to a plurality of words output by the backward decoder to obtain a plurality of fused features corresponding to the plurality of words, and to obtain a plurality of second decoding results based on the plurality of fused features.
[0096] According to some embodiments, the second decoder may be configured to, for each word of the plurality of words, output an N+1 th decoding result corresponding to the word by receiving an N th decoding result of the word as a query feature of the second decoder and receiving second audio features as key features and value features of the second decoder, where N is an integer greater than or equal to 2.
[0097] According to some embodiments, the word-level feature extraction module may be configured to extract and obtain third speech features from the first speech features based on second a priori information, where the second a priori information includes a plurality of second decoding results, where the third speech features include a plurality of second word-level audio features corresponding to a plurality of words. The second decoder may be configured to decode the third speech features and obtain a plurality of third decoding results corresponding to the plurality of words, where the third decoding results indicate third recognition results for the corresponding words.
[0098] According to some embodiments, the second decoder may be a speech large model. According to another aspect of the present disclosure, there is provided an apparatus for training a deep learning model for speech recognition. The deep learning model includes a first decoder and a second decoder. As shown in FIG. 11 , the training apparatus 1100 includes an acquisition module 1110 configured to acquire sample speech and actual recognition results of a plurality of words in the sample speech; an audio feature encoding module 1120 configured to acquire first sample speech features of the sample speech, the first sample speech features including a plurality of sample speech segment features corresponding to a plurality of sample speech segments in the sample speech; a first decoder 1130 configured to decode the first sample speech features and obtain a plurality of first sample decoding results corresponding to a plurality of words in the sample speech, the first sample decoding results indicating first recognition results of the corresponding words; and a second decoder configured to select a second sample speech feature from the first sample speech features based on first sample a priori information. The apparatus 1100 further includes a word-level feature extraction module 1140 configured to extract and obtain sample speech features, where the first sample a priori information includes a plurality of first sample decoding results, and the second sample speech features include a plurality of first sample word-level audio features corresponding to a plurality of words; a second decoder 1150 configured to decode the second sample speech features and obtain a plurality of second sample decoding results corresponding to the plurality of words, where the second sample decoding results indicate second recognition results of the corresponding words; and a parameter adjustment module 1160 configured to adjust parameters of a deep learning model based on the actual recognition results of the plurality of words, the first recognition results, and the second recognition results to obtain a trained deep learning model. As can be understood, the operations of modules 1110 to 1160 of the apparatus 1100 are similar to those of steps S901 to S906 of FIG. 9, and therefore will not be described here.
[0099] In the technical solution disclosed herein, the collection, storage, use, processing, transmission, provision and disclosure of relevant user personal information shall all comply with the provisions of relevant laws and regulations and shall not violate public order and morals.
[0100] According to embodiments of the present disclosure, an electronic device, a readable storage medium, and a computer program product are further provided. Referring to FIG. 12 , a block diagram of an electronic device 1200 that can operate as a server or client of the present disclosure, which is an example of a hardware device applicable to each aspect of the present disclosure, will be described. The electronic device may represent various forms of digital electronic computing devices, such as laptop computers, desktop computers, stage computers, personal digital assistants, servers, blade servers, mainframe computers, and other suitable computers. The electronic device may also represent various forms of mobile devices, such as personal digital processing devices, mobile phones, intelligent phones, wearable devices, and other similar computing devices. The components, their connections, and their functions shown herein are merely exemplary and do not limit the implementation of the present disclosure as described and / or claimed herein.
[0101] 12, electronic device 1200 includes a computing unit 1201, which can perform various appropriate operations and processes according to a computer program stored in a read-only memory (ROM) 1202 or loaded from a storage unit 1208 into a random access memory (RAM) 1203. RAM 1203 may further store various programs and data necessary for operating electronic device 1200. Computing unit 1201, ROM 1202, and RAM 1203 are connected to each other via a bus 1204. An input / output (I / O) interface 1205 is also connected to bus 1204.
[0102] Multiple components of the electronic device 1200, such as an input unit 1206, an output unit 1207, a storage unit 1208, and a communication unit 1209, are connected to an input / output (I / O) interface 1205. The input unit 1206 may be any type of device capable of inputting information into the electronic device 1200. The input unit 1206 can input numeric or character information and generate key signal inputs related to user settings and / or function control of the electronic device, and may include, but is not limited to, a mouse, a keyboard, a touchscreen, a trackboard, a trackball, a control lever, a microphone, and / or a remote control. The output unit 1207 may be any type of device capable of presenting information, and may include, but is not limited to, a display, a speaker, a video / audio output terminal, a vibrator, and / or a printer. The storage unit 1208 may include, but is not limited to, a magnetic disk or an optical disk. The communications unit 1209 enables the electronic device 1200 to exchange information / data with other devices via a computer network, e.g., the Internet, and / or various telecommunications networks, and may include, but is not limited to, a modem, a network card, an infrared communications device, a wireless communications transceiver, and / or a chipset, e.g., a Bluetooth™ device, an 802.11 device, a WiFi device, a WiMax device, a cellular communications device, and / or the like.
[0103] The computing unit 1201 may be any of a variety of general-purpose and / or special-purpose processing components having processing and computing capabilities. Some examples of the computing unit 1201 may include, but are not limited to, a central processing unit (CPU), a graphics processing unit (GPU), various dedicated artificial intelligence (AI) computing chips, various computing units that execute machine learning model algorithms, a digital signal processor (DSP), and any suitable processor, controller, microcontroller, etc. The computing unit 1201 executes the methods and processes described above, such as the speech recognition method and / or the method for training a deep learning model for speech recognition. For example, in some embodiments, the speech recognition method and / or the method for training a deep learning model for speech recognition may be implemented as a computer software program tangibly embodied in a machine-readable medium, such as the storage unit 1208. In some embodiments, some or all of the computer program may be loaded and / or installed into the electronic device 1200 via the ROM 1202 and / or the communication unit 1209. When the computer program is loaded into RAM 1203 and executed by computing unit 1201, it can perform one or more steps of the speech recognition method and / or the method for training a deep learning model for speech recognition described above. Alternatively, in other embodiments, computing unit 1201 may be configured to perform the speech recognition method and / or the method for training a deep learning model for speech recognition in any other suitable manner (e.g., by firmware).
[0104] Various embodiments of the systems and techniques described herein may be implemented in digital electronic circuitry systems, integrated circuit systems, field programmable gate arrays (FPGAs), application specific integrated circuits (ASICs), application specific standard products (ASSPs), systems on a chip (SOCs), complex programmable logic devices (CPLDs), software hardware, firmware, software, and / or combinations thereof. These various embodiments may be embodied in one or more computer programs that may be executed and / or interpreted by a programmable system including at least one programmable processor, which may be a special purpose or general purpose programmable processor, and may receive data and instructions from, and send data and instructions to, a storage system, at least one input device, and at least one output device.
[0105] Program code implementing the methods of the present disclosure may be written in any combination of one or more programming languages. These program codes may be provided to a processor or controller of a general-purpose computer, a special-purpose computer, or other programmable data processing apparatus, so that when executed by the processor or controller, the program code performs the functions / operations specified in the flowcharts and / or block diagrams. The program code may be entirely machine-executable, partially machine-executable, partially machine-executable and partially remote machine-executable as a separate software package, or entirely on a remote machine or server.
[0106] In the context of this disclosure, a machine-readable medium may be a tangible medium, and may include or store a program for use in or in conjunction with an instruction execution system, apparatus, or device. A machine-readable medium may be a machine-readable signal medium or a machine-readable storage medium. A machine-readable medium may include, but is not limited to, an electronic, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any suitable combination of the foregoing. More specific examples of machine-readable storage media include an electrical connection with one or more leads, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the foregoing.
[0107] To provide for user interaction, the systems and techniques described herein can be implemented on a computer that includes a display device (e.g., a cathode ray tube (CRT) or liquid crystal display (LCD) monitor) for displaying information to a user, and a keyboard and pointing device (e.g., a mouse or trackball) through which the user may provide input to the computer. Other types of devices can be used to provide for user interaction, for example, where feedback provided to the user can be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback) and input from the user can be received in any form (including acoustic, speech, or tactile input).
[0108] The systems and techniques described herein may be implemented in a computing system including backstage components (e.g., as a data server), middleware components (e.g., as an application server), front-end components (e.g., a user computer having a graphical user interface or web browser through which a user can interact with the system or technique implementation), or any combination of backstage components, middleware components, or front-end components. The components of the system may be interconnected by any form or medium of digital data communication (e.g., a communications network). Examples of communications networks include, for example, a local area network (LAN), a wide area network (WAN), the Internet, and a blockchain network.
[0109] The computer system may include a client and a server. The client and the server are generally remote from each other and usually interact via a communication network. The client-server relationship is created by running computer programs on corresponding computers. The server may be a cloud server, a server in a distributed system, or a server combined with a blockchain.
[0110] It should be understood that the various forms of flow described above may be used to rearrange the order, add or remove steps, etc. For example, the steps described in this disclosure may be performed in parallel, sequentially, or in a different order, as long as the technical solutions disclosed in this disclosure can achieve the desired results, and the present disclosure is not limited thereto.
[0111] Although embodiments or examples of the present disclosure have been described with reference to the drawings, it should be understood that the above-described methods, systems, and devices are merely illustrative embodiments or examples, and that the scope of the present invention is not limited by these embodiments or examples, but only by the appended claims and their equivalents. Various elements of the embodiments or examples may be omitted or replaced by their equivalents. Furthermore, steps may be performed in a different order than described in this disclosure. Furthermore, various elements of the embodiments or examples may be combined in various ways. Importantly, as technology evolves, many elements described herein may be replaced by equivalent elements that appear later in this disclosure.
Claims
1. 1. A method for speech recognition, comprising: acquiring first speech features of a speech to be recognized, the first speech features including a plurality of speech segment features corresponding to a plurality of speech segments in the speech to be recognized; Decoding the first speech features using a first decoder to obtain a plurality of first decoding results corresponding to a plurality of words in the speech to be recognized, wherein the first decoding results indicate first recognition results for the corresponding words; Extracting and acquiring second speech features from the first speech features based on first a priori information, wherein the first a priori information includes the plurality of first decoding results, and the second speech features include a plurality of first word-level audio features corresponding to the plurality of words; and decoding the second speech features using a second decoder to obtain a plurality of second decoding results corresponding to the plurality of words, the second decoding results indicating second recognition results for the corresponding words.
2. Extracting and acquiring the second speech feature from the first speech feature based on the first a priori information includes:
2. The method of claim 1, comprising: for each word of the plurality of words, obtaining the first word-level audio feature corresponding to the word output by the attention module by using the first decoding result corresponding to the word as a query feature of an attention module and the first speech feature as a key feature and a value feature of the attention module.
3. Extracting and acquiring the second speech feature from the first speech feature based on the first a priori information includes:
3. The method of claim 2, further comprising: utilizing a second encoder to globally encode the plurality of first word-level audio features corresponding to the plurality of words to obtain the enhanced second speech features.
4. Using the second decoder to decode the second speech feature and obtain the second decoding results corresponding to the words, 4. A method according to any one of claims 1 to 3, comprising, for each word of the plurality of words, obtaining the second decoding result corresponding to the word, which is output by the second decoder, by using the first decoding result corresponding to the word as a query feature of the second decoder and the second audio features as key features and value features of the second decoder.
5. 5. The method of claim 4, wherein the second decoder includes a forward decoder and a backward decoder, and both the forward decoder and the backward decoder are configured, for each word of the plurality of words, to take the first decoding result of the word as an input query feature and the second audio features as input key features and value features, the forward decoder is configured to time mask the input features from the past to the future, and the backward decoder is configured to time mask the input features from the future to the past.
6. obtaining, for each word of the plurality of words, the second decoding result corresponding to the word output by the second decoder by using the first decoding result of the word as a query feature of the second decoder and the second speech features as key features and value features of the second decoder; fusing a plurality of forward decoding features corresponding to the plurality of words output by the forward decoder and a plurality of backward decoding features corresponding to the plurality of words output by the backward decoder to obtain a plurality of fused features corresponding to the plurality of words; and obtaining the second decoding results based on the fusion features.
7. 5. The method of claim 4, further comprising: for each word of the plurality of words, obtaining an N+1th decoding result corresponding to the word output by the second decoder by using the Nth decoding result of the word as a query feature of the second decoder and the second audio features as key features and value features of the second decoder, wherein N is an integer greater than or equal to 2.
8. extracting and acquiring third speech features from the first speech features based on second a priori information, wherein the second a priori information includes the second decoding results, and the third speech features include second word-level audio features corresponding to the words; 4. The method of claim 1, further comprising: using the second decoder to decode the third speech feature to obtain a plurality of third decoding results corresponding to the plurality of words, wherein the third decoding results indicate third recognition results for the corresponding words.
9. Obtaining a first speech feature of a speech to be recognized includes: obtaining original speech features of the speech to be recognized; determining a plurality of spikes in the speech to be recognized based on the original speech features; truncating the original audio features to obtain the audio segment features that correspond one-to-one to the spikes.
10. The plurality of audio segment features are sequentially obtained by streaming truncating the original audio features, and decoding the first audio feature using a first decoder includes:
10. The method of claim 9, comprising utilizing the first decoder to stream-decode the plurality of audio segment features in sequence.
11. Obtaining a first speech feature of a speech to be recognized includes: Obtaining corresponding historical feature abstraction information for a currently obtained speech segment feature, wherein the historical feature abstraction information is obtained by attention modeling the previous speech segment feature using a first decoding result corresponding to the previous speech segment feature; and encoding the currently obtained speech segment features in combination with the historical feature abstraction information using a first encoder to obtain corresponding enhanced speech segment features.
12. obtaining corresponding enhanced speech segment features by encoding the currently obtained speech segment features in combination with the historical feature abstraction information using a first encoder; 12. The method of claim 11, further comprising obtaining the corresponding enhanced speech segment features output by the first encoder by using the currently obtained speech segment features as query features of the first encoder and splicing results of the historical feature abstraction information and the currently obtained speech segment features as key features and value features of the first encoder.
13. truncating the original audio features to obtain the audio segment features that correspond one-to-one to the spikes; 10. The method of claim 9, further comprising: cutting the original audio features based on a predetermined time length; and determining the audio segment feature of the audio segment in which each spike of the plurality of spikes exists as the audio segment feature corresponding to the spike.
14. truncating the original audio features to obtain the audio segment features that correspond one-to-one to the spikes; 10. The method of claim 9, further comprising: cutting the original audio features based on the plurality of spikes; and taking the audio segment features between each two adjacent spikes as the audio segment features corresponding to one of the spikes.
15. The method according to any one of claims 1 to 3, wherein the second decoder is a speech large model, the speech large model is a deep learning model used for speech recognition, and the number of parameters of the speech large model exceeds 1 billion.
16. 1. A method for training a deep learning model for speech recognition, the deep learning model including a first decoder and a second decoder, the training method comprising: obtaining a sample speech and actual recognition results for a plurality of words in the sample speech; obtaining first sample speech features of the sample speech, the first sample speech features including a plurality of sample speech segment features corresponding to a plurality of sample speech segments in the sample speech; using a first decoder to decode the first sample speech features to obtain a plurality of first sample decoding results corresponding to a plurality of words in the sample speech, the first sample decoding results indicating first recognition results for the corresponding words; Extracting and obtaining second sample speech features from the first sample speech features based on first sample a priori information, wherein the first sample a priori information includes the plurality of first sample decoding results, and the second sample speech features include a plurality of first sample word-level audio features corresponding to the plurality of words; using a second decoder to decode the second sample speech features to obtain a plurality of second sample decoding results corresponding to the plurality of words, the second sample decoding results indicating second recognition results for the corresponding words; and adjusting parameters of the deep learning model based on the actual recognition results of the plurality of words, the first recognition result, and the second recognition result to obtain a trained deep learning model.
17. A speech recognition device, a speech feature encoding module configured to acquire first speech features of a speech to be recognized, the first speech features including a plurality of speech segment features corresponding to a plurality of speech segments in the speech to be recognized; a first decoder configured to decode the first speech features and obtain a plurality of first decoding results corresponding to a plurality of words in the speech to be recognized, the first decoding results indicating first recognition results of the corresponding words; a word-level feature extraction module configured to extract and obtain second speech features from the first speech features based on first a priori information, wherein the first a priori information includes the plurality of first decoding results, and the second speech features include a plurality of first word-level audio features corresponding to the plurality of words; a second decoder configured to decode the second speech features to obtain a plurality of second decoding results corresponding to the plurality of words, the second decoding results indicating second recognition results for the corresponding words.
18. 18. The apparatus of claim 17, wherein the word-level feature extraction module includes an attention module configured to, for each word of the plurality of words, receive the first decoding result corresponding to the word as a query feature of the attention module and receive the first speech feature as a key feature and a value feature of the attention module, thereby outputting the first word-level audio feature corresponding to the word.
19. The word level feature extraction module:
20. The apparatus of claim 18, comprising: a second encoder configured to globally encode a plurality of first word-level audio features corresponding to the plurality of words to obtain enhanced second speech features.
20. 20. The apparatus of claim 17, wherein the second decoder is configured to, for each word of the plurality of words, receive a first decoding result corresponding to the word as a query feature of the second decoder, and receive the second audio features as key features and value features of the second decoder, thereby outputting a second decoding result corresponding to the word.
21. 21. The apparatus of claim 20, wherein the second decoder includes a forward decoder and a backward decoder, both of which are configured to, for each word of the plurality of words, receive a first decoding result of the word as an input query feature and receive the second audio features as input key features and value features, and wherein the forward decoder is configured to time mask the input features from the past to the future, and the backward decoder is configured to time mask the input features from the future to the past.
22. The second decoder comprises: fusing a plurality of forward decoded features corresponding to the plurality of words output by the forward decoder and a plurality of backward decoded features corresponding to the plurality of words output by the backward decoder to obtain a plurality of fused features corresponding to the plurality of words; The apparatus of claim 21 , configured to obtain the second decoding results based on the fusion features.
23. The second decoder comprises:
21. The device of claim 20, configured to, for each word of the plurality of words, receive an Nth decoding result of the word as a query feature of the second decoder and receive the second audio features as key features and value features of the second decoder, thereby outputting an N+1th decoding result corresponding to the word, where N is an integer greater than or equal to 2.
24. The word-level feature extraction module is configured to extract third speech features from the first speech features based on second a priori information, the second a priori information including the second decoding results, and the third speech features including a plurality of second word-level audio features corresponding to the words; 20. The apparatus of claim 17, wherein the second decoder is configured to decode the third speech features to obtain a plurality of third decoding results corresponding to the plurality of words, the third decoding results indicating third recognition results for the corresponding words.
25. The audio feature encoding module: Acquire original speech features of the speech to be recognized; determining a plurality of spikes in the speech to be recognized based on the original speech features; 20. The apparatus of any one of claims 17 to 19, configured to truncate the original audio features to obtain the audio segment features that correspond one-to-one to the spikes.
26. 26. The apparatus of claim 25, wherein the plurality of speech segment features are obtained in sequence by streaming truncating the original speech features, and the first decoder is configured to stream-decode the plurality of speech segment features in sequence.
27. The audio feature encoding module: The method is configured to obtain corresponding historical feature abstraction information for a currently obtained speech segment feature, wherein the historical feature abstraction information is obtained by attention modeling the previous speech segment feature using a first decoding result corresponding to the previous speech segment feature; wherein the speech feature encoding module:
27. The apparatus of claim 26, comprising a first encoder configured to encode currently obtained speech segment features in combination with the historical feature abstraction information to output corresponding enhanced speech segment features.
28. The first encoder comprises:
28. The apparatus of claim 27, configured to output the corresponding enhanced speech segment features by receiving the currently obtained speech segment features as query features of the first encoder and receiving splicing results of the historical feature abstraction information and the currently obtained speech segment features as key features and value features of the first encoder.
29. truncating the original audio features to obtain the audio segment features that correspond one-to-one to the spikes; 26. The apparatus of claim 25, further comprising: cutting the original audio features based on a predetermined length of time; and determining an audio segment feature of an audio segment in which each spike of the plurality of spikes exists as the audio segment feature corresponding to the spike.
30. truncating the original audio features to obtain the audio segment features that correspond one-to-one to the spikes; The apparatus of claim 25, further comprising: cutting the original audio features based on the plurality of spikes; and taking the audio segment features between each two adjacent spikes as the audio segment features corresponding to one of the spikes.
31. The apparatus according to any one of claims 17 to 19, wherein the second decoder is a speech large model, the speech large model is a deep learning model used for speech recognition, and the number of parameters of the speech large model exceeds 1 billion.
32. 1. An apparatus for training a deep learning model for speech recognition, the deep learning model including a first decoder and a second decoder, the training apparatus comprising: an acquisition module configured to acquire a sample speech and an actual recognition result of a plurality of words in the sample speech; a speech feature encoding module configured to obtain first sample speech features of the sample speech, the first sample speech features including a plurality of sample speech segment features corresponding to a plurality of sample speech segments in the sample speech; a first decoder configured to decode the first sample speech features to obtain a plurality of first sample decoding results corresponding to a plurality of words in the sample speech, the first sample decoding results indicating first recognition results for the corresponding words; a word-level feature extraction module configured to extract and obtain second sample speech features from the first sample speech features based on first sample a priori information, wherein the first sample a priori information includes the plurality of first sample decoding results, and the second sample speech features include a plurality of first sample word-level audio features corresponding to the plurality of words; a second decoder configured to decode the second sample speech features to obtain a plurality of second sample decoding results corresponding to the plurality of words, the second sample decoding results indicating second recognition results for the corresponding words; and a parameter adjustment module configured to adjust parameters of the deep learning model based on actual recognition results of the plurality of words, the first recognition result, and the second recognition result, to obtain a trained deep learning model.
33. An electronic device, the electronic device comprising: at least one processor; and a memory communicatively connected to the at least one processor, wherein: An electronic device, wherein the memory stores instructions executable by at least one processor, the instructions being executed by the at least one processor to enable the at least one processor to perform the method of any one of claims 1 to 3.
34. A non-transitory computer-readable storage medium having stored thereon computer instructions for causing a computer to carry out the method of any one of claims 1 to 3.
35. A computer program product comprising a computer program which, when executed by a processor, implements the method of any one of claims 1 to 3.
Citation Information
Patent Citations
System and method for end-to-end speech recognition with trigger door tension
JP2022522379A
Speech recognition method, codec method and apparatus, electronic device and storage medium
JP2023041610A
Derivation Model-Based Two-Pass End-to-End Speech Recognition
JP2023041867A
Predicting Word Boundaries for On-Device Batching of End-To-End Speech Recognition Models
US20230107493A1