Speech recognition method and related device
By encoding and mapping speech segments of a preset duration, combining them with character information from the previous decoding step, and dynamically acquiring acoustic features, a method for step-by-step decoding is used to solve the speech recognition problem of large language models in real-time scenarios, achieving streaming recognition and timely result output.
Patent Information
- Application Number
- CN202511465329.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-10-14
- Publication Date
- 2025-12-09
AI Technical Summary
Existing speech recognition solutions based on large language models cannot be applied to real-time scenarios and require waiting for the user to finish speaking before providing the recognition result.
By encoding each pre-set duration speech segment to obtain frame-level acoustic features, mapping them to the embedding space of a large language model, and combining them with the character information from the previous decoding step, the acoustic features of the character to be decoded in the current decoding step are dynamically obtained and gradually input into the large language model for decoding, thus achieving streaming recognition.
It enables speech recognition in real-time scenarios and can output recognition results in a timely manner, making it suitable for a variety of real-time application scenarios.
Smart Images

Figure CN121096342A_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the technical field of speech recognition, and in particular to a speech recognition method and related device. BACKGROUND
[0002] Automatic speech recognition technology, also known as speech recognition technology, is a technology that converts speech signals into text information that can be understood by humans through calculation. Speech recognition technology is widely used in mobile phone voice assistants, input method software, car navigation, and various artificial intelligence wearable devices, and has important application value.
[0003] Large language model, also known as large-scale language model, is an artificial intelligence model designed to understand and generate human language. In recent years, with the emergence of rich world knowledge and the expansion of model parameters, the capabilities of large language models have been greatly enhanced, and they have emerged with unprecedented capabilities in text understanding and generation tasks.
[0004] Current speech recognition schemes include large language model-based recognition schemes, but most current large language model-based recognition schemes are designed for non-real-time scenarios. This scheme requires the user to finish speaking the complete content before providing the speech recognition result, and it cannot be applied to real-time scenarios. SUMMARY
[0005] Therefore, the present application provides a speech recognition method and related device to solve the problem that the current large language model-based recognition scheme cannot be applied to real-time scenarios. The technical solution is as follows:
[0006] The first aspect of the present application provides a speech recognition method, comprising:
[0007] Encode each obtained speech segment of a preset time length to obtain frame-level acoustic features;
[0008] Map the obtained frame-level acoustic features to the embedding space of the large language model to obtain mapped acoustic features;
[0009] Obtain the acoustic features of the to-be-decoded character of the current decoding step according to the obtained mapped acoustic features and the character information of the previous decoding step;
[0010] Input the character decoded by the previous decoding step and the acoustic features of the to-be-decoded character of the current decoding step into the large language model for decoding to obtain the character decoded by the current decoding step.
[0011] In one possible implementation, the method for obtaining the acoustic features of the to-be-decoded character of the current decoding step according to the obtained mapped acoustic features and the character information of the previous decoding step comprises:
[0012] determine the acoustic feature boundary of the to-be-decoded character of the current decoding step according to the obtained mapped acoustic features and the character information of the previous decoding step;
[0013] obtain the acoustic features of the to-be-decoded character of the current decoding step from the obtained mapped acoustic features according to the acoustic feature boundary of the to-be-decoded character of the current decoding step.
[0014] In a possible implementation, the determining the acoustic feature boundary of the to-be-decoded character of the current decoding step according to the obtained mapped acoustic features and the character information of the previous decoding step comprises:
[0015] obtain the attention weights of the current decoding step on each frame acoustic feature of the target acoustic feature according to the obtained mapped acoustic features and the character state vector of the previous decoding step, wherein the target acoustic feature is the acoustic feature after the acoustic feature boundary determined by the previous decoding step;
[0016] determine the acoustic feature boundary of the to-be-decoded character of the current decoding step according to the attention weights of the current decoding step on each frame acoustic feature of the target acoustic feature.
[0017] In a possible implementation, the determining the acoustic feature boundary of the to-be-decoded character of the current decoding step according to the attention weights of the current decoding step on each frame acoustic feature of the target acoustic feature comprises:
[0018] accumulate the attention weights from the first frame acoustic feature of the target acoustic feature;
[0019] when the accumulated value of the attention weights is greater than a set threshold, determine the acoustic feature corresponding to the last attention weight participating in the accumulation as the acoustic feature boundary of the to-be-decoded character of the current decoding step.
[0020] In a possible implementation, the obtaining the acoustic features of the to-be-decoded character of the current decoding step according to the obtained mapped acoustic features and the character information of the previous decoding step comprises:
[0021] obtain the acoustic features of the to-be-decoded character of the current decoding step according to the obtained mapped acoustic features and the character information of the previous decoding step by using a monotonic attention adaptation model;
[0022] wherein the monotonic attention adaptation model and the large language model are jointly trained by using training speech labeled with text.
[0023] In a possible implementation, the joint training target of the monotonic attention adaptation model and the large language model comprises:
[0024] increase the attention weight of the monotonic attention adaptation model on the acoustic feature of the first frame to the acoustic feature of the b-th frame of the training speech in the first decoding step, b is the expected delay;
[0025] make the character predicted by the monotonic attention adaptation model according to the acoustic feature of the to-be-decoded character in each decoding step consistent with the corresponding labeled character;
[0026] make the character obtained by decoding the acoustic feature of the to-be-decoded character by the large language model in each decoding step consistent with the corresponding labeled character.
[0027] In a possible implementation, the joint training process of the monotonic attention adaptation model and the large language model includes:
[0028] divide the training speech into training speech segments of a preset time length, encode each training speech segment in turn to obtain frame-level acoustic features, and map the obtained frame-level acoustic features to an embedding space of the large language model to obtain mapped acoustic features;
[0029] In each decoding step: use the monotonic attention adaptation model to obtain the acoustic feature of the to-be-decoded character in the decoding step from the obtained mapped acoustic features, and predict the character according to the acoustic feature of the to-be-decoded character in the decoding step; input the character decoded in the previous decoding step and the acoustic feature of the to-be-decoded character in the decoding step into the large language model for decoding;
[0030] determine a speech recognition loss according to the decoding result of the large language model in each decoding step, the character prediction result of the monotonic attention adaptation model in each decoding step, and the text labeled by the training speech;
[0031] determine a first word delay loss according to the expected delay and the attention weight of the first decoding step on the mapped acoustic features, wherein the attention weight of the first decoding step on the mapped acoustic features is obtained in the process of obtaining the acoustic feature of the to-be-decoded character in the first decoding step;
[0032] update the parameters of the monotonic attention adaptation model and the large language model according to the speech recognition loss and the first word delay loss.
[0033] In a possible implementation, the determination of the first word delay loss according to the expected delay and the attention weight of the first decoding step on the mapped acoustic features includes:
[0034] sum the attention weights of the first decoding step on the acoustic features of the first frame to the acoustic features of the b-th frame to obtain an attention weight sum;
[0035] calculate the absolute value of the difference between the upper limit value of the attention weight sum and the attention weight sum to obtain the first word delay loss.
[0036] In a possible implementation, the step of predicting the character according to the acoustic features of the character to be decoded in the decoding step comprises:
[0037] determining a contribution weight corresponding to each frame acoustic feature of the character to be decoded in the decoding step according to the acoustic features of the character to be decoded in the decoding step, the character state vector of the previous decoding step, and the attention weight of the decoding step on each frame acoustic feature of the character to be decoded;
[0038] performing weighted summation on the frame acoustic features of the character to be decoded in the decoding step according to the contribution weights corresponding to the frame acoustic features of the character to be decoded in the decoding step, to obtain a context vector of the decoding step;
[0039] predicting the character according to the context vector of the decoding step.
[0040] The second aspect of the present application provides an electronic device, comprising at least one processor and a memory connected to the processor, wherein:
[0041] The memory is configured to store a computer program;
[0042] The processor is configured to execute the computer program, so that the electronic device can implement the steps of any one of the speech recognition methods.
[0043] The third aspect of the present application provides a computer storage medium, which carries one or more computer programs, and when the one or more computer programs are executed by an electronic device, the electronic device can implement the steps of any one of the speech recognition methods.
[0044] The fourth aspect of the present application provides a computer program product, comprising computer readable instructions, when the computer readable instructions run on an electronic device, the electronic device can implement the steps of any one of the speech recognition methods.
[0045] By the technical solution, the speech recognition method provided by the application is used to encode the speech segment of the preset time length to obtain frame-level acoustic features, map the obtained frame-level acoustic features to an embedding space of a large language model to obtain mapped acoustic features, obtain acoustic features of a to-be-decoded character of a current decoding step according to the obtained mapped acoustic features and character information of a previous decoding step, and input the character decoded by the previous decoding step and the acoustic features of the to-be-decoded character of the current decoding step into the large language model to obtain a recognized character of the current decoding step. The speech recognition method provided by the application is a streaming recognition method, which can realize streaming decoding and thus can be applied to a real-time scenario. In addition, the most refined and most relevant acoustic features (i.e., the acoustic features of the to-be-decoded character of the current decoding step) are provided to the large language model in each decoding step of the streaming decoding, so that the large language model can give a correct recognition result. BRIEF DESCRIPTION OF DRAWINGS
[0046] In order to more clearly illustrate the technical solutions of the embodiments of the present application or the prior art, the drawings needed to be used in the embodiments or the prior art description will be briefly introduced. Obviously, the drawings in the following description are only embodiments of the present application, and other drawings can be obtained by those skilled in the art without creative labor on the basis of the provided drawings.
[0047] Figure 1 The flowchart of the speech recognition method provided by the embodiment of the present application is shown in the figure.
[0048] Figure 2 The flowchart of obtaining acoustic features of a to-be-decoded character of a current decoding step according to obtained mapped acoustic features and character information of a previous decoding step provided by the embodiment of the present application is shown in the figure.
[0049] Figure 3 The schematic diagram of implementing streaming decoding based on a monotonic attention adaptation model and a large language model provided by the embodiment of the present application is shown in the figure.
[0050] Figure 4 The flowchart of jointly training a monotonic attention adaptation model and a large language model provided by the embodiment of the present application is shown in the figure.
[0051] Figure 5 The schematic diagram of mask operation provided by the embodiment of the present application is shown in the figure.
[0052] Figure 6 The schematic diagram of attention weight of the first character before and after introducing the first character delay loss provided by the embodiment of the present application is shown in the figure.
[0053] Figure 7A structural schematic diagram of a voice recognition device provided by an embodiment of the present application is shown. DETAILED DESCRIPTION
[0054] The embodiments of the present application are described below in conjunction with the accompanying drawings. The terms used in the embodiment part of the present application are only used to explain the specific embodiments of the present application, and are not intended to limit the present application.
[0055] The embodiments of the present application are described below in conjunction with the accompanying drawings. Those skilled in the art can know that, as technology develops and new scenarios appear, the technical solutions provided by the embodiments of the present application are also applicable to similar technical problems.
[0056] The terms "first", "second", and the like in the specification and claims of the present application and the above-mentioned drawings are used to distinguish similar objects, and do not necessarily indicate a specific order or sequence. It should be understood that the terms used in this way can be interchanged under appropriate circumstances, and this is only a distinguishing way used in the description of the embodiments of the present application to describe the objects with the same attributes. In addition, the terms "include" and "have" and any variations thereof are intended to cover non-exclusive inclusion, so that the processes, methods, systems, products or equipment containing a series of units do not necessarily limit to those units, but can include other units not clearly listed or inherent to these processes, methods, products or equipment.
[0057] At present, the process of implementing voice recognition based on a large language model voice recognition scheme is to obtain voice features of a voice to be recognized, then discretize the voice features of the voice to be recognized into tokens that can be received by the large language model, input all the tokens into the large language model, and after the large language model encodes the input tokens, start decoding token by token.
[0058] The above-mentioned scheme makes full use of the language knowledge and reasoning generation ability possessed by the large language model, and surpasses the traditional small model in effect evaluation. However, the above-mentioned scheme has a big problem, that is, the above-mentioned scheme does not have real-time capability, because the large language model must encode all the tokens before it can start generating tasks, which limits the application of the above-mentioned scheme in real-time scenarios. In view of this, the present application provides a voice recognition method applicable to real-time scenarios.
[0059] Next, the voice recognition method provided by the present application is introduced through the following embodiments.
[0060] Please refer to Figure 1, shows a flowchart of a speech recognition method provided by an embodiment of the present application. The method can be applied in real-time scenarios (such as real-time conference transcription, multi-language translation, speaker separation, etc. in the field of conferences and collaboration, live real-time subtitles, video subtitle generation, etc. in the field of audio and video and media production, online education real-time subtitles, pronunciation evaluation for language learning, etc. in the field of education, in-vehicle voice assistants, etc. in the field of vehicle-mounted and transportation, voice electronic medical records, medical consultation, etc. in the field of medical treatment, smart home voice control, wearable device voice interaction, etc. in the field of smart hardware and Internet of Things, voice-controlled game characters, real-time transcription of in-game voice chat, etc. in the field of games and entertainment, voice input method, automatic generation of conference minutes, etc. in the field of enterprise efficiency and tools), which can include:
[0061] Step S101: encode the preset length of the speech segment to obtain the frame-level acoustic feature.
[0062] In order to realize streaming speech recognition, each time a preset length of speech segment is obtained, it is encoded without waiting for the end of the sentence.
[0063] For example, continuously receive the audio stream, and after obtaining a 40ms (40 frames, 1 frame 10ms) speech segment, encode the obtained 40ms speech segment, that is, after obtaining the 1st frame of speech to the 40th frame of speech, encode the 1st frame of speech to the 40th frame of speech, after obtaining the 41st frame of speech to the 80th frame of speech, encode the 41st frame of speech to the 80th frame of speech, and so on.
[0064] Optionally, the preset length of the speech segment can be encoded based on a speech encoder (such as a Conformer encoder). Assuming that the speech segment is 40 frames, the speech encoder (such as the Conformer encoder) encodes only 40 frames of speech each time. The attention field of the self-attention module in the speech encoder (such as the Conformer encoder) is 40 frames.
[0065] Step S102: map the obtained frame-level acoustic feature to the embedding space of the large language model to obtain the mapped feature.
[0066] Considering that speech and text are two different modalities of data, after obtaining the frame-level acoustic feature, the frame-level acoustic feature is mapped to the embedding space of the large language model, that is, the frame-level acoustic feature is "translated" into a feature that the large language model can understand. The dimension of the mapped feature meets the input dimension of the large language model.
[0067] Optionally, the obtained frame-level acoustic feature can be mapped to the embedding space of the large language model using a linear mapping layer.
[0068] Step S103: obtaining the acoustic feature of the character to be decoded in the current decoding step according to the obtained mapped acoustic feature and the character information of the previous decoding step.
[0069] In streaming speech recognition, decoding is performed character by character, but the mapped acoustic feature comes in segments, so it is necessary to dynamically and step by step find the most relevant feature from the obtained mapped acoustic feature to the current decoding step.
[0070] In order to achieve accurate and efficient streaming decoding, the present application obtains the acoustic feature of the character to be decoded in the current decoding step from the obtained mapped acoustic feature according to the obtained mapped acoustic feature and the character information of the previous decoding step.
[0071] Step S104: inputting the character decoded from the previous decoding step and the acoustic feature of the character to be decoded in the current decoding step into the large language model for decoding to obtain the character decoded from the current decoding step.
[0072] After obtaining the acoustic feature of the character to be decoded in the current decoding step, the character decoded from the previous decoding step and the acoustic feature of the character to be decoded in the current decoding step are combined and input into the large language model for decoding to obtain the recognized character in the current decoding step. It should be noted that when the character decoded from the previous decoding step is input into the large language model, it needs to be converted into an embedding vector.
[0073] The present embodiment decodes step by step based on the large language model until the end character is decoded, and the recognized characters in each decoding step form the speech recognition result of the whole speech.
[0074] The large language model in the present embodiment can be a decoder-only architecture large language model. The decoder-only architecture large language model focuses on the generation task, making it perform exceptionally well in terms of text fluency, creativity and consistency.
[0075] The present embodiment ingeniously utilizes the autoregressive generation capability of the decoder-only architecture large language model to generate the speech recognition result.
[0076] The voice recognition method provided in the embodiments of the present application is a streaming recognition method, which can realize streaming decoding and thus can be applied to a real-time scenario. In addition, in each decoding step of the streaming decoding, the most refined and most relevant acoustic feature (i.e., the acoustic feature of the to-be-decoded character of the current decoding step) is provided to the large language model, so that the large language model can give a correct recognition result.
[0077] In some embodiments of the present application, the implementation process of "step S103: obtaining the acoustic feature of the to-be-decoded character of the current decoding step according to the obtained mapped acoustic feature and the character information of the previous decoding step" is introduced.
[0078] In a possible implementation manner, as shown in Figure 2 obtaining the acoustic feature of the to-be-decoded character of the current decoding step according to the obtained mapped acoustic feature and the character information of the previous decoding step can include the following steps.
[0079] Step S201: determining the acoustic feature boundary of the to-be-decoded character of the current decoding step according to the obtained mapped acoustic feature and the character information of the previous decoding step.
[0080] In a possible implementation manner, the process of determining the acoustic feature boundary of the to-be-decoded character of the current decoding step according to the obtained mapped acoustic feature and the character information of the previous decoding step can include the following steps.
[0081] Step S2011: obtaining the attention weight of the current decoding step on each frame acoustic feature of the target acoustic feature according to the mapped acoustic feature and the character information of the previous decoding step.
[0082] The target acoustic feature is the acoustic feature located after the acoustic feature boundary determined by the previous decoding step in the mapped acoustic feature.
[0083] In a possible implementation manner, the attention weight of the current decoding step on each frame acoustic feature of the mapped acoustic feature can be determined according to the mapped acoustic feature and the character information of the previous decoding step, and then the attention weight of the current decoding step on each frame acoustic feature of the target acoustic feature is obtained from the obtained attention weight.
[0084] In another possible implementation, the target acoustic features can be obtained from the mapped acoustic features, and then, based on the target acoustic features and the character information of the previous decoding step, the attention weight of the current decoding step on the acoustic features of each frame of the target acoustic features can be determined.
[0085] The following describes the process of determining the attention weights of the current decoding step on each frame of the mapped acoustic features based on the mapped acoustic features and the character information of the previous decoding step (the process of determining the attention weights of the current decoding step on each frame of the target acoustic features based on the target acoustic features and the character information of the previous decoding step is similar).
[0086] The process of determining the attention weights of the current decoding step on the acoustic features of each frame of the mapped acoustic features, based on the mapped acoustic features and the character information from the previous decoding step, may include:
[0087] Step a1: Based on the mapped acoustic features and the character state vector of the previous decoding step, determine the monotonic energy value of each frame of acoustic features for the current decoding step with respect to the mapped acoustic features.
[0088] Among them, the monotonic energy value of the acoustic features of a frame in the current decoding step represents the degree of attention the current decoding step pays to the acoustic features of that frame.
[0089] Assuming the current decoding step is the i-th decoding step, for the j-th frame acoustic feature h after mapping... j The i-th decoding step for acoustic feature h j The monotonic energy value e ij It can be represented as:
[0090] e ij =monotonicEnergy(s i-1 ,h j (1).
[0091] Among them, s i-1 This represents the character state vector at the (i-1)th decoding step, specifically the state vector corresponding to the character at the (i-1)th decoding step, which is generated at the (i-1)th decoding step. The ith decoding step, for acoustic feature h... j The monotonic energy value e ij Characterizing the acoustic feature h in the i-th decoding step j The level of attention given to it.
[0092] Step a2: Convert the monotonic energy values of each frame of acoustic features in the current decoding step into probabilities to obtain the selection probabilities corresponding to each frame of acoustic features in the mapped acoustic features.
[0093] Optionally, the monotone energy value of each frame of the mapped acoustic feature corresponding to the current decoding step can be converted into a probability based on a sigmoid activation function, to obtain a selection probability corresponding to each frame of the mapped acoustic feature.
[0094] Assuming that the current decoding step is the i-th decoding step, the selection probability p j corresponding to the j-th frame of the mapped acoustic feature h ij may be represented as:
[0095] p ij = σ (e ij ) (2).
[0096] wherein σ represents a sigmoid activation function.
[0097] Step a3, determining the attention weight of the current decoding step on each frame of the mapped acoustic feature according to the selection probability corresponding to each frame of the mapped acoustic feature.
[0098] Assuming that the current decoding step is the i-th decoding step, the selection probability p j corresponding to the j-th frame of the mapped acoustic feature h ij may be calculated by the following formula to obtain the attention weight a j of the current decoding step on the acoustic feature h ij :
[0099] (3).
[0100] Step S2012, determining the acoustic feature boundary of the character to be decoded in the current decoding step according to the attention weight of the current decoding step on each frame of the target acoustic feature.
[0101] In one possible implementation, the process of determining the acoustic feature boundary of the character to be decoded in the current decoding step according to the attention weight of the current decoding step on each frame of the target acoustic feature can include: starting from the 1st frame of the target acoustic feature, accumulating the attention weights, and when the accumulated value of the attention weights is greater than a set threshold (such as 0.5), determining the acoustic feature corresponding to the last attention weight participating in the accumulation as the acoustic feature boundary of the character to be decoded in the current decoding step.
[0102] It should be noted that the embodiment does not limit the determination of the acoustic feature boundary of the character to be decoded in the current decoding step according to the attention weight of the current decoding step on each frame of the target acoustic feature, and for example, the acoustic feature boundary of the character to be decoded in the current decoding step can also be directly determined according to the selection probability corresponding to each frame of the target acoustic feature.
[0103] Step S202: According to the acoustic feature boundary of the to-be-decoded character of the current decoding step, the acoustic feature of the to-be-decoded character of the current decoding step is obtained from the obtained mapped acoustic feature.
[0104] If the current decoding step is the first decoding step, and the acoustic feature boundary of the first decoding step is at the k1th frame acoustic feature of the mapped acoustic feature, then the acoustic features from the first frame to the k1th frame are obtained from the mapped acoustic feature as the acoustic features of the to-be-decoded character of the first decoding step. If the current decoding step is the second decoding step, and the acoustic feature boundary of the second decoding step is at the k2th frame acoustic feature of the mapped acoustic feature, then the acoustic features from the k1+1th frame to the k2th frame are obtained from the mapped acoustic feature as the acoustic features of the to-be-decoded character of the second decoding step. The other decoding steps are similar.
[0105] In some embodiments of the present application, a monotonic attention adaptation model can be used to obtain the acoustic features of the to-be-decoded character of the current decoding step according to the obtained mapped acoustic feature and the character decoded by the previous decoding step.
[0106] That is, using the monotonic attention adaptation model, the acoustic feature boundary of the to-be-decoded character of the current decoding step is determined according to the obtained mapped acoustic feature and the character information of the previous decoding step, and the acoustic features of the to-be-decoded character of the current decoding step are obtained from the obtained mapped acoustic feature according to the acoustic feature boundary of the to-be-decoded character of the current decoding step.
[0107] Further, the character decoded by the previous decoding step and the acoustic features of the to-be-decoded character of the current decoding step are input into a large language model for decoding to obtain the character decoded by the current decoding step.
[0108] The following will be described in combination with Figure 3The process of realizing streaming decoding based on the monotonic attention adaptation model and the large language model is described as follows: after obtaining a speech segment of a preset time length (such as 40 ms), the obtained speech segment is encoded based on an encoder (such as a Conformer encoder) to obtain frame-level acoustic features, the frame-level acoustic features are mapped to an embedding space of the large language model by using a linear mapping layer to obtain mapped acoustic features satisfying an input dimension of the large language model, in the first decoding step, the acoustic feature boundary of a to-be-decoded character in the first decoding step is determined according to state information corresponding to a starting character "sos" and the mapped acoustic features, the acoustic features of the to-be-decoded character in the first decoding step are obtained from the mapped acoustic features according to the acoustic feature boundary of the to-be-decoded character in the first decoding step, the starting character "sos" and the acoustic features of the to-be-decoded character in the first decoding step are input into the large language model for decoding, and the large language model outputs a recognized character "jin" in the first decoding step, in the second decoding step, the acoustic feature boundary of a to-be-decoded character in the second decoding step is determined according to state information corresponding to the character "jin" and the mapped acoustic features, the acoustic features of the to-be-decoded character in the second decoding step are obtained from the mapped acoustic features according to the acoustic feature boundary determined in the first decoding step and the acoustic feature boundary of the to-be-decoded character in the second decoding step, the recognized character "jin" in the first decoding step and the acoustic features of the to-be-decoded character in the second decoding step are input into the large language model for decoding, and the large language model outputs a recognized character "tian" in the second decoding step, and other decoding steps are performed in the same manner until the large language model outputs an ending character.
[0109] In a possible implementation, the monotonic attention adaptation model and the large language model are jointly trained by using training data in a training data set, where the training data set includes a plurality of training data, and each piece of training data is training speech annotated with text.
[0110] The joint training objectives of the monotonic attention adaptation model and the large language model include: (1) increasing the attention weight of the monotonic attention adaptation model for the first frame acoustic feature to the bth frame acoustic feature of the training speech in the first decoding step, where b is the expected delay; (2) making the character predicted by the monotonic attention adaptation model according to the acoustic features of the to-be-decoded character in each decoding step consistent with the corresponding annotated character; and (3) making the character decoded by the large language model according to the acoustic features of the to-be-decoded character in each decoding step consistent with the corresponding annotated character.
[0111] It should be noted that the joint training objectives of the monotonic attention adaptation model and the large language model are not limited to the above three objectives, for example, the joint training objectives can also include the two objectives of (1) and (3).
[0112] In some embodiments of the present application, the joint training process of the monotonic attention adaptation model and the large language model is introduced.
[0113] Please refer to Figure 4 , a flowchart showing the joint training process of the monotonic attention adaptation model and the large language model, which can include:
[0114] Step S401: obtaining training speech labeled with text from a training data set.
[0115] Step S402: dividing the obtained training speech into training speech segments of a preset duration, and sequentially encoding each training speech segment to obtain frame-level acoustic features.
[0116] After obtaining the training speech, the obtained training speech is divided into training speech segments of a preset duration (such as 40 ms). After obtaining the training speech segments, the training speech segments are sequentially input into a speech encoder (such as a Conformer encoder) for encoding to obtain frame-level acoustic features.
[0117] Step S403: mapping the obtained frame-level acoustic features to the embedding space of the large language model to obtain mapped acoustic features.
[0118] The obtained frame-level acoustic features can be mapped to the embedding space of the large language model using a linear mapping layer to obtain acoustic features that can be understood by the large language model and meet the input dimension of the large language model.
[0119] Step S404: at each decoding step, using the monotonic attention adaptation model to obtain the acoustic features of the to-be-decoded character of the decoding step from the obtained mapped acoustic features.
[0120] The process of obtaining the acoustic features of the to-be-decoded character of the decoding step can refer to the specific implementation process of the above-mentioned "S103: obtaining the acoustic features of the to-be-decoded character of the current decoding step according to the obtained mapped acoustic features and the character information of the previous decoding step". This embodiment will not be repeated here.
[0121] Step S405a: using the monotonic attention adaptation model to predict characters according to the acoustic features of the to-be-decoded character of the decoding step to obtain the character prediction result of the decoding step.
[0122] Specifically, the process of using the monotonic attention adaptation model to predict characters according to the acoustic features of the to-be-decoded character of the decoding step can include:
[0123] Step S405a-1, using the monotonic attention adaptation model to determine the contribution weight corresponding to each frame acoustic feature of the to-be-decoded character of the decoding step according to the acoustic features of the to-be-decoded character of the decoding step, the character state vector of the previous decoding step, and the attention weight of the decoding step on the acoustic features of the to-be-decoded character.
[0124] Assuming that the decoding step is the i-th decoding step, the mapped acoustic features are h, and the acoustic feature boundary of the i-th decoding step is at the j-th frame acoustic feature h j of the mapped acoustic features h
[0125] (4);
[0126] (5);
[0127] (6)。
[0128] Wherein, pos i-1 represents the acoustic feature boundary determined by the i-1-th decoding step, μ ij represents the chunk energy of the i-th decoding step on the acoustic features h j (chunk energy is the refined energy, and the calculation of this energy only focuses on the frames of acoustic features of the to-be-decoded character, and does not focus on other acoustic features), s i-1 represents the character state vector of the i-1-th decoding step, that is, the state vector corresponding to the character of the i-1-th decoding step, β ij is the acoustic feature h j of the to-be-decoded character of the i-th decoding step corresponding to the contribution weight.
[0129] Step S405a-2, using the monotonic attention adaptation model, weighting and summing the frames of acoustic features of the to-be-decoded character of the decoding step according to the contribution weight corresponding to each frame of acoustic feature of the to-be-decoded character of the decoding step, to obtain the context vector of the decoding step.
[0130] Assuming that the decoding step is the i-th decoding step, the context vector c i of the i-th decoding step is:
[0131] (7)。
[0132] Wherein, T represents the number of frames of the obtained mapped acoustic features, β i,m represents the contribution weight corresponding to the m-th frame acoustic feature h m in the mapped acoustic features h
[0133] Step S405a-3, using the monotonic attention adaptation model, predicting the character according to the context vector of the decoding step.
[0134] Wherein, the monotonic attention adaptation model includes a decoder, and the context vector of the decoding step can be input into the decoder to obtain a character prediction result.
[0135] Step S405b: input the character decoded by the previous decoding step and the acoustic feature of the character to be decoded in the decoding step into the large language model to obtain the decoding result of the decoding step.
[0136] The character decoded by the previous decoding step in this step refers to the character decoded by the large language model in the previous decoding step.
[0137] Step S406a: determine the speech recognition loss according to the decoding result of each decoding step of the large language model, the character prediction result of each decoding step of the monotonic attention adaptation model, and the text of the training speech annotation.
[0138] Specifically, the speech recognition loss of the large language model is determined according to the decoding result of each decoding step of the large language model and the text of the training speech annotation, and the speech recognition loss of the monotonic attention adaptation model is determined according to the character prediction result of each decoding step of the monotonic attention adaptation model and the text of the training speech annotation. It should be noted that the purpose of introducing the speech recognition loss of the monotonic attention adaptation model is to enable the monotonic attention adaptation model to have the alignment capability of word to acoustic feature for the speech recognition task.
[0139] Step S406b: determine the first word delay loss according to the expected delay and the attention weight of the 1st decoding step on the obtained mapped acoustic feature.
[0140] Considering that when the first word recognition is performed, the attention weight of the 1st decoding step on the mapped acoustic feature is usually small, and the small attention weight will lead to an increase in delay (from the 1st frame acoustic feature, the attention weight is accumulated, and the position where the accumulated value of the attention weight is first greater than a set threshold is found, that is, the acoustic feature boundary of the character to be decoded in the 1st decoding step, and the small attention weight will lead to difficulty in finding the acoustic feature boundary of the character to be decoded in the 1st decoding step), in view of this situation, the first word delay loss is introduced to solve the first word delay problem, thereby improving the user experience.
[0141] In a possible implementation manner, the process of determining the first word delay loss according to the expected delay and the attention weight of the 1st decoding step on the obtained mapped acoustic feature can include: summing the attention weight of the 1st decoding step on the 1st frame acoustic feature to the bth frame acoustic feature to obtain an attention weight sum; calculating the absolute value of the difference between the upper limit value (usually 1) of the attention weight sum and the attention weight sum to obtain the first word delay loss.
[0142] The calculation method of the first word delay loss is as follows:
[0143] (8)
[0144] wherein a0 represents the attention weight of the first decoding step on the obtained mapped acoustic features, mask(a0, b) represents setting the attention weight of the first decoding step on each acoustic feature after the bth acoustic feature to 0, and an example is shown in Figure 5 As shown in the left side of the figure, the attention weight of the first decoding step on the obtained mapped acoustic features is 0.04, 0.05, 0.05, 0.04, 0.15, 0.25, 0.13, 0.15, 0.05, 0.04, 0.02, 0.02, and 0.01 in sequence, b = 6, and mask(a0, 6) means setting the attention weight on the 7th acoustic feature to the 13th acoustic feature to 0. sum(mask(a0, b)) is the sum of the attention weight of the first decoding step on the first acoustic feature to the bth acoustic feature.
[0145] Step S407: updating the parameters of the monotonic attention adaptation model and the large language model according to the speech recognition loss and the first word delay loss.
[0146] The first word delay loss is introduced to optimize the monotonic attention adaptation model, so that the monotonic attention adaptation model can obtain a larger attention degree within the expected delay b when decoding the first word, as shown in Figure 6 Figure 6 The right side of the figure is the attention weight of the first word after introducing the first word delay loss, so that the acoustic feature boundary can be found more easily.
[0147] When updating the parameters of the large language model, the LoRA fine-tuning strategy can be used, that is, in the fine-tuning process, the Q, K, and V matrices of the self-attention module of the large language model are not directly updated, but a set of trainable LoRA adapters (LoRA parameters, that is, a pair of new, small, and trainable matrices) are added to these three matrices in parallel, and during the training process, the newly added LoRA parameters are updated.
[0148] In addition, in order to enable the model to have non-real-time decoding capability, when jointly training the monotonic attention adaptation model and the large language model, in addition to inputting the training speech into the speech encoder (such as the Conformer encoder) in the form of speech segments with a preset time length in sequence, the entire training speech can also be input into the speech encoder (for example, the two input methods can be alternately used).
[0149] Embodiments of the present application also provide a speech recognition device, as shown in Figure 7 The speech recognition device can include a speech encoding unit 701, a feature mapping unit 702, a character acoustic feature obtaining unit 703, and a character decoding unit 704.
[0150] The speech coding unit 701 is configured to code a preset length of speech segment to obtain frame-level acoustic features for each obtained preset length of speech segment.
[0151] The feature mapping unit 702 is configured to map the obtained frame-level acoustic features to an embedding space of a large language model to obtain mapped acoustic features.
[0152] The character acoustic feature obtaining unit 703 is configured to obtain acoustic features of a to-be-decoded character at a current decoding step according to the obtained mapped acoustic features and character information at a previous decoding step.
[0153] The character decoding unit 704 is configured to input the character decoded at the previous decoding step and the acoustic features of the to-be-decoded character at the current decoding step into the large language model to obtain a character decoded at the current decoding step.
[0154] In a possible implementation, the process in which the character acoustic feature obtaining unit 703 obtains the acoustic features of the to-be-decoded character at the current decoding step according to the obtained mapped acoustic features and the character information at the previous decoding step includes:
[0155] determining acoustic feature boundaries of the to-be-decoded character at the current decoding step according to the obtained mapped acoustic features and the character information at the previous decoding step;
[0156] obtaining the acoustic features of the to-be-decoded character at the current decoding step from the obtained mapped acoustic features according to the acoustic feature boundaries of the to-be-decoded character at the current decoding step.
[0157] In a possible implementation, the process in which the character acoustic feature obtaining unit 703 determines the acoustic feature boundaries of the to-be-decoded character at the current decoding step according to the obtained mapped acoustic features and the character information at the previous decoding step includes:
[0158] obtaining attention weights of the current decoding step on each frame acoustic feature of a target acoustic feature according to the obtained mapped acoustic features and a character state vector at the previous decoding step, wherein the target acoustic feature is an acoustic feature after the acoustic feature boundaries determined at the previous decoding step;
[0159] determining the acoustic feature boundaries of the to-be-decoded character at the current decoding step according to the attention weights of the current decoding step on each frame acoustic feature of the target acoustic feature.
[0160] In a possible implementation, the process in which the character acoustic feature obtaining unit 703 determines the acoustic feature boundaries of the to-be-decoded character at the current decoding step according to the attention weights of the current decoding step on each frame acoustic feature of the target acoustic feature includes:
[0161] accumulating the attention weights from the first frame acoustic feature of the target acoustic feature;
[0162] When the attention weight cumulative value is greater than the set threshold value, the acoustic feature corresponding to the last attention weight participating in the accumulation is determined as the acoustic feature boundary of the to-be-decoded character of the current decoding step.
[0163] In a possible implementation, the process of obtaining, by the character acoustic feature obtaining unit 703, the acoustic feature of the to-be-decoded character of the current decoding step according to the obtained mapped acoustic feature and the character information of the previous decoding step includes:
[0164] obtaining, by the monotonic attention adaptation model, the acoustic feature of the to-be-decoded character of the current decoding step according to the obtained mapped acoustic feature and the character information of the previous decoding step;
[0165] The monotonic attention adaptation model and the large language model are jointly trained by using the training speech annotated with text.
[0166] In a possible implementation, the joint training target of the monotonic attention adaptation model and the large language model includes:
[0167] increasing the attention weight of the monotonic attention adaptation model for the first frame acoustic feature to the b-th frame acoustic feature of the first frame acoustic feature of the training speech in the first decoding step, b being the expected delay;
[0168] making the character predicted by the monotonic attention adaptation model according to the acoustic feature of the to-be-decoded character in each decoding step consistent with the corresponding annotated character;
[0169] making the character obtained by the large language model by decoding the acoustic feature of the to-be-decoded character in each decoding step consistent with the corresponding annotated character.
[0170] In a possible implementation, the speech recognition apparatus can further include a model training unit.
[0171] The model training unit is configured to jointly train the monotonic attention adaptation model and the large language model by using the training speech annotated with text.
[0172] The process of jointly training, by the model training unit, the monotonic attention adaptation model and the large language model by using the training speech annotated with text includes:
[0173] dividing the training speech into training speech segments of a preset time length, sequentially encoding each training speech segment to obtain frame-level acoustic features, mapping the obtained frame-level acoustic features to an embedding space of the large language model to obtain mapped acoustic features;
[0174] In each decoding step: obtaining, by using the monotonic attention adaptation model, acoustic features of a character to be decoded in the decoding step from the obtained mapped acoustic features, and predicting the character according to the acoustic features of the character to be decoded in the decoding step; inputting the character decoded in a previous decoding step and the acoustic features of the character to be decoded in the decoding step into the large language model for decoding;
[0175] determining, according to a decoding result of the large language model in each decoding step, a character prediction result of the monotonic attention adaptation model in each decoding step, and the training text with speech annotation, a speech recognition loss;
[0176] determining, according to the expected delay and an attention weight of the 1st decoding step on the mapped acoustic features, a first word delay loss, wherein the attention weight of the 1st decoding step on the mapped acoustic features is obtained in a process of obtaining the acoustic features of the character to be decoded in the 1st decoding step;
[0177] updating parameters of the monotonic attention adaptation model and the large language model according to the speech recognition loss and the first word delay loss.
[0178] In a possible implementation, the process of determining, by the model training unit, the first word delay loss according to the expected delay and the attention weight of the 1st decoding step on the mapped acoustic features includes:
[0179] summing the attention weights of the 1st decoding step on the 1st frame acoustic feature to the bth frame acoustic feature to obtain an attention weight sum;
[0180] calculating an absolute value of a difference between the attention weight sum upper limit value and the attention weight sum to obtain the first word delay loss.
[0181] In a possible implementation, the process of predicting, by the model training unit, the character according to the acoustic features of the character to be decoded in the decoding step includes:
[0182] determining, according to the acoustic features of the character to be decoded in the decoding step, a character state vector of a previous decoding step, and the attention weight of the decoding step on each frame acoustic feature of the character to be decoded, a contribution weight corresponding to each frame acoustic feature of the character to be decoded in the decoding step;
[0183] weighting and summing, according to the contribution weight corresponding to each frame acoustic feature of the character to be decoded in the decoding step, each frame acoustic feature of the character to be decoded in the decoding step to obtain a context vector of the decoding step;
[0184] predicting the character according to the context vector of the decoding step.
[0185] The voice recognition apparatus provided by the embodiments of the present application can implement streaming decoding based on a large language model, and can be applied to a real-time scenario. In addition, the most refined and most relevant acoustic features (i.e., the acoustic features of the to-be-decoded character in the current decoding step) are provided to the large language model in each decoding step of the streaming decoding, so that the large language model can give a correct recognition result.
[0186] The embodiments of the present application also provide an electronic device, comprising at least one processor and a memory connected with the processor, wherein:
[0187] The memory is configured to store a computer program.
[0188] The processor is configured to execute the computer program, so that the electronic device can implement the steps of the voice recognition method provided by the above embodiments.
[0189] The present disclosure also provides a computer-readable storage medium having computer program instructions stored thereon, the computer program instructions being executed by a processor to implement the steps of the voice recognition method provided by the above embodiments.
[0190] The embodiments of the present application also provide a computer program product comprising computer readable instructions, which, when executed on an electronic device, cause the electronic device to implement the steps of the voice recognition method provided by the above embodiments.
[0191] In addition, it should be noted that the apparatus embodiments described above are merely illustrative, wherein the units described as separate components can or can not be physically separate, and the components displayed as units can or can not be physical units, i.e., they can be located in one place, or distributed on multiple network units. Part or all of the modules can be selected according to actual needs to achieve the purpose of the embodiments. In addition, in the apparatus embodiments provided by the present application, the connection relationship between the modules indicates that there is a communication connection between them, which can be implemented as one or more communication buses or signal lines.
[0192] Those skilled in the art can clearly understand that the application can be implemented by means of software plus necessary universal hardware, and of course can also be implemented by means of dedicated hardware including special integrated circuit, special CPU, special memory, special component, etc. Generally, any function completed by computer program can be easily implemented by corresponding hardware, and the specific hardware structure for implementing the same function can also be various, such as analog circuit, digital circuit or special circuit, etc. However, for the application, software program implementation is a better embodiment. Based on such understanding, the technical solution of the application or the part of the application which makes contribution to the prior art can be embodied in the form of software product, which is stored in readable storage medium, such as computer floppy disk, U disk, mobile hard disk, ROM, RAM, magnetic disk or optical disk, etc., and includes a plurality of instructions for making a computer device (which can be personal computer, training device or network device, etc.) execute the method described in various embodiments of the application.
[0193] In the above embodiments, the implementation can be achieved by software, hardware, firmware or any combination thereof, entirely or partially. When implemented by software, the implementation can be achieved in the form of a computer program product, entirely or partially.
[0194] The computer program product includes one or more computer instructions. When the computer program instructions are loaded and executed on a computer, the flow or function described in the embodiments of the application is generated entirely or partially. The computer can be a general-purpose computer, a special-purpose computer, a computer network or other programmable device. The computer instructions can be stored in a computer-readable storage medium or transmitted from one computer-readable storage medium to another, for example, the computer instructions can be transmitted from one website, computer, training device or data center to another website, computer, training device or data center through wired (such as coaxial cable, optical fiber, digital subscriber line (DSL)) or wireless (such as infrared, wireless, microwave, etc.) mode. The computer-readable storage medium can be any available medium that can be stored by the computer or a data storage device such as training device, data center, etc. integrated with one or more available media sets. The available medium can be magnetic medium (such as floppy disk, hard disk, magnetic tape), optical medium (such as DVD) or semiconductor medium (such as solid state disk (SSD)) etc.
Claims
1. A speech recognition method, characterized in that, include: For each speech segment of a preset duration obtained, the speech segment is encoded to obtain frame-level acoustic features; The obtained frame-level acoustic features are mapped to the embedding space of the large language model to obtain the mapped acoustic features. Based on the obtained mapped acoustic features and the character information from the previous decoding step, the acoustic features of the character to be decoded in the current decoding step are obtained. The acoustic features of the character decoded in the previous decoding step and the character to be decoded in the current decoding step are input into the large language model for decoding to obtain the character decoded in the current decoding step.
2. The speech recognition method according to claim 1, characterized in that, The step of obtaining the acoustic features of the character to be decoded in the current decoding step based on the obtained mapped acoustic features and the character information of the previous decoding step includes: Based on the obtained mapped acoustic features and the character information from the previous decoding step, determine the acoustic feature boundaries of the character to be decoded in the current decoding step. Based on the acoustic feature boundaries of the character to be decoded in the current decoding step, obtain the acoustic features of the character to be decoded in the current decoding step from the obtained mapped acoustic features.
3. The speech recognition method according to claim 2, characterized in that, The step of determining the acoustic feature boundary of the character to be decoded in the current decoding step based on the obtained mapped acoustic features and the character information from the previous decoding step includes: Based on the obtained mapped acoustic features and the character state vector of the previous decoding step, the attention weights of each frame acoustic features of the target acoustic features in the current decoding step are obtained, wherein the target acoustic features are the acoustic features after the acoustic feature boundary determined by the previous decoding step. Based on the attention weights of the current decoding step on the acoustic features of each frame of the target acoustic features, the acoustic feature boundaries of the character to be decoded in the current decoding step are determined.
4. The speech recognition method according to claim 3, characterized in that, The step of determining the acoustic feature boundary of the character to be decoded in the current decoding step based on the attention weights of the target acoustic features in each frame of the current decoding step includes: The attention weight is accumulated starting from the first frame of acoustic features of the target acoustic features; When the cumulative value of the attention weights exceeds the set threshold, the acoustic feature corresponding to the last attention weight that participated in the accumulation will be determined as the acoustic feature boundary of the character to be decoded in the current decoding step.
5. The speech recognition method according to any one of claims 1 to 4, characterized in that, The step of obtaining the acoustic features of the character to be decoded in the current decoding step based on the obtained mapped acoustic features and the character information of the previous decoding step includes: Using a monotonic attention adaptation model, the acoustic features of the character to be decoded in the current decoding step are obtained based on the mapped acoustic features and the character information of the previous decoding step. The monotonic attention adaptation model and the large language model are jointly trained using labeled training speech.
6. The speech recognition method according to claim 5, characterized in that, The joint training objectives of the monotonic attention adaptation model and the large language model include: Increase the attention weight of the monotonic attention adaptation model for the acoustic features of the training speech from the first frame to the bth frame in the first decoding step, where b is the expected delay; This makes the character predicted by the monotonic attention adaptation model in each decoding step based on the acoustic features of the character to be decoded tend to be consistent with the corresponding labeled character; This ensures that the character obtained by decoding the acoustic features of the character to be decoded in each decoding step of the large language model is consistent with the corresponding labeled character.
7. The speech recognition method according to claim 5, characterized in that, The joint training process of the monotonic attention adaptation model and the large language model includes: The training speech is segmented into training speech segments of a preset duration, and each training speech segment is encoded sequentially to obtain frame-level acoustic features. The obtained frame-level acoustic features are then mapped to the embedding space of a large language model to obtain mapped acoustic features. In each decoding step: using a monotonic attention adaptation model, the acoustic features of the character to be decoded in this decoding step are obtained from the obtained mapped acoustic features, and the character is predicted based on the acoustic features of the character to be decoded in this decoding step; the character decoded in the previous decoding step and the acoustic features of the character to be decoded in this decoding step are input into the large language model for decoding. The speech recognition loss is determined based on the decoding results of the large language model at each decoding step, the character prediction results of the monotonic attention adaptation model at each decoding step, and the text of the trained speech annotations. The first-character delay loss is determined based on the expected delay and the attention weight of the first decoding step on the mapped acoustic features. The attention weight of the first decoding step on the mapped acoustic features is obtained during the process of acquiring the acoustic features of the character to be decoded in the first decoding step. Based on the speech recognition loss and the first-word delay loss, the parameters of the monotonic attention adaptation model and the large language model are updated.
8. The speech recognition method according to claim 7, characterized in that, The determination of the first-word delay loss based on the expected delay and the attention weight of the first decoding step on the mapped acoustic features includes: The attention weights of the first decoding step on the acoustic features of the first frame to the acoustic features of the b frame are summed to obtain the attention weight sum. The absolute value of the difference between the sum of attention weights and the upper limit value and the sum of attention weights is used to obtain the first-word delay loss.
9. The speech recognition method according to claim 7, characterized in that, The prediction of characters based on the acoustic features of the character to be decoded in the decoding step includes: Based on the acoustic features of the character to be decoded in this decoding step, the character state vector of the previous decoding step, and the attention weight of this decoding step on the acoustic features of each frame of the character to be decoded, the contribution weight corresponding to the acoustic features of each frame of the character to be decoded in this decoding step is determined. Based on the contribution weights corresponding to the acoustic features of each frame of the character to be decoded in this decoding step, the acoustic features of each frame of the character to be decoded in this decoding step are weighted and summed to obtain the context vector of this decoding step. Predict the character based on the context vector of this decoding step.
10. An electronic device, characterized in that, It includes at least one processor and a memory connected to the processor, wherein: The memory is used to store computer programs; The processor is used to execute the computer program to enable the electronic device to implement the steps of the speech recognition method as described in any one of claims 1 to 9.
11. A computer storage medium, characterized in that, The storage medium carries one or more computer programs that, when executed by an electronic device, enable the electronic device to implement the steps of the speech recognition method as described in any one of claims 1 to 9.
12. A computer program product, characterized in that, It includes computer-readable instructions that, when executed on an electronic device, enable the electronic device to perform the steps of the speech recognition method as described in any one of claims 1 to 9.