Speech recognition error correction method and apparatus

CN121331103BActive Publication Date: 2026-09-11ANHUI IFLYTEK UNIVERSAL LANGUAGE TECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202511426635.2
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-09-30
Publication Date
2026-09-11
Estimated Expiration
2045-09-30

AI Technical Summary

Technical Problem

[0004]本发明提供一种语音识别纠错方法及装置,用以解决现有技术中如何有效的对语音识别结果进行后处理,以提高语音识别的准确率的问题

Benefits of technology

[0018] The speech recognition error correction method and apparatus provided by this invention, through a forced alignment method, generates frame-by-frame accurate supervision labels synchronized with the speech frames for the training data, thereby solving the core problem of the difficulty in training streaming error correction models. This clear training objective enables the model to efficiently learn how to locate and correct recognition errors. Combined with the powerful context awareness capabilities of a large language model, the resulting speech recognition error correction model not only responds quickly but also exhibits higher robustness and error correction accuracy when processing long and complex sentences. Furthermore, by directly feeding the frame-by-frame output of the speech recognition model into the speech recognition error correction model, end-to-end streaming processing of recognition and error correction is achieved without waiting for the entire sentence to end, greatly shortening the time for the first character to appear on screen and the time for correcting intermediate results, providing users with a smooth, low-latency real-time speech recognition experience.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121331103B_ABST
    Figure CN121331103B_ABST
Patent Text Reader

Abstract

The application provides a speech recognition error correction method and device, and relates to the technical field of data processing, which comprises the following steps: inputting a current audio frame in to-be-recognized user audio into a speech recognition model to obtain acoustic features of the current audio frame and first text characters; inputting the acoustic features of the current audio frame and the first text characters, and a historical correction text sequence into a speech recognition error correction model to obtain correction text characters corresponding to the current audio frame; wherein the historical correction text sequence is a historical audio frame before the current audio frame in the to-be-recognized user audio, and is a correction text sequence obtained by sequentially passing through the speech recognition model and the speech recognition error correction model; the speech recognition error correction model is obtained based on training samples carrying forced alignment text labels, and the forced alignment text labels are obtained by performing forced alignment on real text labels corresponding to the training samples and speech sample signals corresponding to the training samples.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of data processing technology, and in particular to a speech recognition error correction method and apparatus. Background Technology

[0002] Recognition accuracy is a key factor affecting the widespread application of Automatic Speech Recognition (ASR). Therefore, reducing the error rate of speech recognition during the recognition process is very important for ASR.

[0003] In addition to improving the accuracy of the speech recognition model itself, the recognition results can also be post-processed to further reduce the recognition error rate. Therefore, how to effectively post-process the speech recognition results to improve the accuracy of speech recognition has become an urgent problem to be solved in the industry. Summary of the Invention

[0004] This invention provides a speech recognition error correction method and apparatus to solve the problem in the prior art of how to effectively post-process the speech recognition results to improve the accuracy of speech recognition.

[0005] This invention provides a speech recognition error correction method, characterized in that it includes: Input the current audio frame in the user's audio to be identified into the speech recognition model to obtain the acoustic features and the first text character of the current audio frame; The acoustic features of the current audio frame, the first text character, and the historical corrected text sequence are input into the speech recognition error correction model to obtain the corrected text character corresponding to the current audio frame. The historical corrected text sequence is the corrected text sequence obtained by passing the historical audio frames in the user audio to be identified before the current audio frame through the speech recognition model and the speech recognition error correction model in sequence. The speech recognition error correction model is trained based on training samples carrying forced alignment text labels, which are obtained by forcibly aligning the real text labels corresponding to the training samples with the speech sample signals corresponding to the training samples.

[0006] According to the speech recognition error correction method provided in this application, before the step of inputting the acoustic features of the current audio frame, the first text character, and the historical corrected text sequence into the speech recognition error correction model to obtain the corrected text character corresponding to the current audio frame, the method further includes: Based on the target dictionary characters, the embedding layer and output layer in the initial large language model are initialized to obtain the initialized large language model; wherein, the target dictionary characters are determined based on the speech recognition dictionary and the large language model dictionary of the preset large language model; The acoustic feature sample and text character sequence sample corresponding to the speech sample signal are used as one training sample to obtain multiple training samples; The large language model after initialization is trained based on each training sample carrying a forced alignment text label to obtain a speech recognition error correction model.

[0007] According to the speech recognition error correction method provided in this application, before the step of training the initialized large language model based on training samples carrying forced alignment text labels to obtain the speech recognition error correction model, the method further includes: The speech sample signal is input into the speech recognition model to obtain the probability distribution of outputting any text character or blank mark for each frame of audio sample of the speech sample signal; Based on the probability distribution, an alignment path search space is constructed, wherein the horizontal dimension of the alignment path search space corresponds to the time frame sequence of the speech sample signal, and the vertical dimension corresponds to the character sequence of the real text label. In the alignment path search space, a dynamic programming algorithm is used to find the alignment path with the highest probability from the start position to the end position, wherein the start position corresponds to the zeroth frame and the zeroth character, and the end position corresponds to the last frame and the last character. Based on the alignment path with the highest probability, a forced alignment text label sequence is generated, wherein each frame position in the forced alignment text label sequence corresponds to a real text character or a blank mark.

[0008] According to the speech recognition error correction method provided in this application, in the alignment path search space, a dynamic programming algorithm is used to find the alignment path with the highest probability from the starting position to the ending position, including: Construct a state transition table, wherein the number of rows in the state transition table is equal to the number of frames of the speech sample plus one, and the number of columns in the state transition table is equal to the number of characters in the real text label plus one; For each state point in the state transition table, calculate the path score from the previous state to the current state, where the path includes a horizontal transition path that outputs a blank marker and a diagonal transition path that outputs text characters. Select the transition path with the highest probability as the optimal path to the current state point, and record the path selection of the optimal path in the backtracking pointer table; Starting from the end position of the table, trace back to the beginning position using the backtracking pointer table to reconstruct the alignment path with the highest probability from the beginning position to the end position.

[0009] According to the speech recognition error correction method provided in this application, the method for calculating the path score includes: Initialize the starting state of the state transition table, and set the cumulative score of the zeroth character position in the zeroth frame to zero; The cumulative score of each state point in the state transition table is calculated frame by frame, starting from the first frame, according to the time frame order. For each character position in the current frame, based on the cumulative state score of the previous frame and the probability distribution of the current frame, the maximum cumulative score to reach the character position is calculated to obtain the path score; While calculating the cumulative score, the source direction of the optimal predecessor state is recorded at the corresponding position in the backtracking pointer table.

[0010] According to the speech recognition error correction method provided in this application, the large language model after initialization is trained based on each training sample carrying a forced alignment text label to obtain a speech recognition error correction model, including: for each training sample carrying a forced alignment text label, inputting the training sample into the large language model and outputting the corrected text character sequence sample corresponding to the training sample; based on the corrected text character sequence sample and the forced alignment text label, calculating the connection temporal classification loss and the cross-entropy loss, wherein the cross-entropy loss is used to measure the difference between the prediction result of each audio frame sample and the forced alignment text label, and the connection temporal classification loss is used to measure the accuracy of the overall prediction; Based on the connection-time classification loss and the cross-entropy loss, the overall loss is determined, and the parameters of the large language model are updated through the backpropagation algorithm. If the total loss is less than a first preset threshold, training is stopped, and a trained speech recognition error correction model is obtained.

[0011] According to the speech recognition error correction method provided in this application, the step of inputting the training samples into the large language model further includes applying a causal masking mechanism. The causal mask matrix restricts the large language model to only use the current position, audio frame samples before the current position, and text labels before the current position when predicting the output at the current position.

[0012] According to the speech recognition error correction method provided in this application, the method further includes: Obtain the characters in the speech recognition dictionary that are the same as those in the large language model dictionary of the preset large language model, and get the first dictionary characters; Obtain dictionary characters that exist in the speech recognition dictionary but do not appear in the large language model dictionary to obtain the second dictionary characters; Each character in the second dictionary character set is encoded using a byte-level byte-pair encoding method to obtain the byte-pair encoding sequence corresponding to each second dictionary character, and a dictionary character mapping table is constructed; wherein, the dictionary character mapping table uses the second dictionary character as the key and the byte-pair encoding sequence corresponding to the second dictionary character as the value; Based on the first dictionary character and the dictionary character mapping table, the target dictionary character is obtained.

[0013] According to the speech recognition error correction method provided in this application, the initialization process of the embedding layer and output layer in the initial large language model based on the target dictionary characters to obtain the initialized large language model includes: Copy the embedding vector and output layer vector corresponding to the first dictionary character in the target dictionary characters into the speech recognition dictionary initialization parameters; For each second dictionary character in the dictionary character mapping table, the embedding vector and output layer vector of each word in the large language model dictionary in the byte pair encoding sequence corresponding to the second dictionary character are averaged to obtain an averaged vector, and the averaged vector is copied as the initialization vector of the second dictionary character into the speech recognition dictionary initialization parameters. Based on the initialization parameters of the speech recognition dictionary, the embedding layer and output layer in the initial large language model are initialized to obtain the initialized large language model.

[0014] This application also provides a speech recognition error correction device, including: The recognition module is used to input the current audio frame in the user's audio to be recognized into the speech recognition model to obtain the acoustic features and the first text character of the current audio frame; The error correction module is used to input the acoustic features of the current audio frame, the first text character, and the historical corrected text sequence into the speech recognition error correction model to obtain the corrected text character corresponding to the current audio frame. The historical corrected text sequence is a corrected text sequence obtained by sequentially passing the speech recognition model and the speech recognition error correction model through the historical audio frames in the user audio to be recognized before the current audio frame; the speech recognition error correction model is trained based on training samples carrying forced alignment text labels, and the forced alignment text labels are obtained by forcibly aligning the real text labels corresponding to the training samples with the speech sample signals corresponding to the training samples.

[0015] Using the acoustic feature samples and text character sequence samples corresponding to the speech sample signals as training samples, the present invention also provides an electronic device, including a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor executes the computer program to implement the speech recognition error correction method as described in any of the above claims.

[0016] The present invention also provides a non-transitory computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements the speech recognition error correction method as described in any of the preceding claims.

[0017] The present invention also provides a computer program product, including a computer program that, when executed by a processor, implements the speech recognition error correction method as described in any of the preceding claims.

[0018] The speech recognition error correction method and apparatus provided by this invention, through a forced alignment method, generates frame-by-frame accurate supervision labels synchronized with the speech frames for the training data, thereby solving the core problem of the difficulty in training streaming error correction models. This clear training objective enables the model to efficiently learn how to locate and correct recognition errors. Combined with the powerful context awareness capabilities of a large language model, the resulting speech recognition error correction model not only responds quickly but also exhibits higher robustness and error correction accuracy when processing long and complex sentences. Furthermore, by directly feeding the frame-by-frame output of the speech recognition model into the speech recognition error correction model, end-to-end streaming processing of recognition and error correction is achieved without waiting for the entire sentence to end, greatly shortening the time for the first character to appear on screen and the time for correcting intermediate results, providing users with a smooth, low-latency real-time speech recognition experience. Attached Figure Description

[0019] To more clearly illustrate the technical solutions in this invention or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are some embodiments of this invention. For those skilled in the art, other drawings can be obtained from these drawings without creative effort.

[0020] Figure 1 A schematic diagram of the speech recognition error correction method provided in this application; Figure 2 Example diagram of alignment path provided for this application; Figure 3 Example diagram of the dynamic programming alignment search algorithm provided in this application; Figure 4 This is a schematic diagram of the speech recognition dictionary adapted to the large language model in this invention; Figure 5 A schematic diagram of the speech recognition error correction model framework provided in this application; Figure 6 A schematic diagram of the disturbance provided by the present invention; Figure 7 This is a schematic diagram of the speech recognition error correction device provided by the present invention; Figure 8 This is a schematic diagram of the structure of the electronic device provided by the present invention. Detailed Implementation

[0021] To make the objectives, technical solutions, and advantages of this invention clearer, the technical solutions of this invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some, not all, of the embodiments of this invention. All other embodiments obtained by those skilled in the art based on the embodiments of this invention without creative effort are within the scope of protection of this invention.

[0022] This application provides a speech recognition error correction method. The subject executing this method can be any computing device capable of performing speech recognition and text processing, such as a personal computer, server, mobile terminal, such as a smartphone, tablet computer, or dedicated embedded device.

[0023] Figure 1 A schematic diagram of the speech recognition error correction method provided in this application is shown below. Figure 1 As shown, it includes: Step 110: Input the current audio frame in the user's audio to be identified into the speech recognition model to obtain the acoustic features and the first text character of the current audio frame; In this application, the user audio to be recognized can be a continuous speech stream that the user is typing and needs to be recognized as text in real time.

[0024] To enable streaming processing, the continuous audio stream is divided into a series of consecutive audio frames. Each frame is a short segment of the audio signal, with a length that can be 10 milliseconds, 20 milliseconds, or 30 milliseconds. The current audio frame refers to the audio frame that is currently being processed in the processing flow.

[0025] In this application, the speech recognition model converts audio signals into text. This speech recognition model can be a connection-based temporal classification model, a recurrent neural network transducer-based model, or an attention-based encoding / decoding model. In a preferred embodiment of the invention, considering the real-time requirements of streaming recognition, a recurrent neural network transducer model is employed.

[0026] When the current audio frame is input into the speech recognition model, the speech recognition model will output two items: the acoustic features and the first text character.

[0027] Acoustic features are deep, abstract mathematical representations of audio frames within the encoder portion of a speech recognition model, containing the speech information of that audio frame. The first text character represents the initial recognition result of the speech recognition model for the current audio frame.

[0028] It should be noted that in streaming models, the speech recognition model may not necessarily output a text character in each frame; it may also output a blank marker, indicating that the current audio frame does not correspond to any new character. Therefore, the first text character can be a specific text character or a blank marker.

[0029] Step 120: Input the acoustic features of the current audio frame, the first text character, and the historical corrected text sequence into the speech recognition error correction model to obtain the corrected text character corresponding to the current audio frame; The historical corrected text sequence is the corrected text sequence obtained by passing the historical audio frames in the user audio to be identified before the current audio frame through the speech recognition model and the speech recognition error correction model in sequence. The speech recognition error correction model is trained based on training samples carrying forced alignment text labels, which are obtained by forcibly aligning the real text labels corresponding to the training samples with the speech sample signals corresponding to the training samples.

[0030] In this application, the input to the speech recognition error correction model includes the acoustic features of the current audio frame, the first text character, and the historical correction text sequence.

[0031] The acoustic features of the current audio frame provide raw acoustic evidence for the speech recognition error correction model, helping it determine whether the first text character is accurate. The first text character represents the initial recognition result of the speech recognition model.

[0032] The historical corrected text sequence is the key to achieving context awareness and streaming error correction. The historical corrected text sequence refers to the corrected text sequence obtained after all historical audio frames before the current audio frame are processed, and then passed through the speech recognition model and the speech recognition error correction model.

[0033] For example, when processing frame t, the historical corrected text sequence is the final error-corrected output from frame 1 to frame t-1. By introducing the historical corrected text sequence, the speech recognition error correction model can utilize the already recognized and corrected contextual information to more accurately judge and correct the recognition result of the current audio frame.

[0034] Based on these inputs, the speech recognition error correction model outputs a corrected text character. This corrected text character is the recognition result of the speech recognition error correction model for the current audio frame after combining acoustic and contextual information. This result may be an affirmation of the first text character, a correction, or maintaining the whitespace marker.

[0035] More specifically, the speech recognition error correction model is trained on training samples carrying forced alignment text labels.

[0036] The training samples include a segment of speech sample signal and its corresponding real text labels. The forced alignment of the text labels is obtained by forcibly aligning the real text labels with the corresponding speech sample signal in the time dimension. The length of the aligned label sequence is strictly equal to the number of frames of the speech sample, where each frame corresponds to a real text character or a blank mark.

[0037] By training with training samples carrying forced alignment text labels, the speech recognition error correction model can learn the specific character or whitespace mark that should be output given the acoustic features and contextual input of each frame.

[0038] This embodiment constructs a cascaded recognition and error correction streaming framework. By utilizing a speech recognition error correction model trained with forced alignment data, and combining the acoustic features of the current audio frame, preliminary recognition results, and historical correction contextual information, it achieves real-time online error correction of speech recognition results. This avoids the high latency caused by traditional error correction methods relying on complete sentences. At the same time, through frame-level fine-grained processing, it significantly improves the accuracy of streaming speech recognition.

[0039] In this application, an innovative forced alignment method is used to cleverly generate frame-by-frame accurate supervision labels synchronized with the speech frames for the training data, thus solving the core challenge of training streaming error correction models. This clear training objective enables the model to efficiently learn how to locate and correct recognition errors. Combined with the powerful context awareness capabilities of a large language model, the resulting speech recognition error correction model not only responds quickly but also exhibits higher robustness and error correction accuracy when handling long and complex sentences. Furthermore, by directly feeding the frame-by-frame output of the speech recognition model into the speech recognition error correction model, end-to-end streaming processing of recognition and error correction is achieved without waiting for the entire sentence to end. This significantly shortens the time for the first character to appear on screen and the time for correcting intermediate results, providing users with a smooth, low-latency real-time speech recognition experience.

[0040] Optionally, before the step of inputting the acoustic features of the current audio frame, the first text character, and the historical corrected text sequence into the speech recognition error correction model to obtain the corrected text character corresponding to the current audio frame, the method further includes: Based on the target dictionary characters, initializing processing is performed on the embedding layer and the output layer of the initial large language model to obtain an initialized large language model; wherein the target dictionary characters are determined based on the speech recognition dictionary and the large language model dictionary of a preset large language model; Taking an acoustic feature sample corresponding to a speech sample signal and a text character sequence sample as one training sample, a plurality of training samples are obtained; According to each training sample carrying a forced alignment text label, the initialized large language model is trained to obtain an automatic speech recognition error correction model.

[0041] In the present application and the present invention, the speech recognition dictionary is a vocabulary set specially used for speech recognition tasks, which contains vocabularies that are expected to appear in speech signals, such as specific-domain terms and common vocabularies. For example, in the telephone speech recognition scenario, the dictionary contains a large number of vocabularies related to communication and calls, such as "call", "signal", "phone bill" and the like.

[0042] In the present invention, the large language model dictionary is a vocabulary set supported by the preset large language model, and the preset large language model may be a large language model the same as or different from the initial large language model. The preset large language model is trained with extensive data, and its dictionary contains rich vocabularies covering various topics and fields. For example, a general large language model dictionary may contain hundreds of thousands or even millions of vocabularies, ranging from common words such as "of", "is", "in" to professional terms such as "quantum computing", "gene editing" and the like.

[0043] In the present invention, the target dictionary characters are a character set determined by combining the requirements of speech recognition tasks and the capability of the preset large language model. It includes both vocabularies from the speech recognition dictionary and general language vocabularies from the preset large language model dictionary.

[0044] In the present invention, the embedding layer is the input layer of the initial large language model, and is responsible for converting discrete characters or vocabularies into continuous vector representations. For example, the character "I" may be represented as a 128-dimensional or higher-dimensional vector.

[0045] The output layer is the last layer of the initial large language model, and is responsible for converting the internal vector representation of the initial large language model into a prediction result. It usually outputs the probability of each possible character or vocabulary. For example, in the next character prediction task, the output layer may output that the probability of the occurrence of the character "is" is 0.6.

[0046] In this invention, the initialization process of the large language model can be specifically as follows: First, determine the first dictionary characters that are commonly contained in the speech recognition dictionary and the large language model dictionary. Then, determine the second dictionary field that only appears in the speech recognition dictionary but not in the large language model dictionary. Convert the second dictionary characters into a byte-level byte pair encoding form to obtain the third dictionary characters. The target dictionary characters can be determined based on the first dictionary characters and the third dictionary characters.

[0047] The vectors corresponding to the characters in the first dictionary are directly written into the speech recognition dictionary initialization parameters. Then, the vectors related to the characters in the third dictionary are averaged to obtain an averaged vector, which is also written into the speech recognition dictionary initialization parameters. Thus, based on the speech recognition dictionary initialization parameters, the initial large language model is initialized before training. Before training, the initial large language model has already initialized its embedding and output layers based on the target dictionary characters. This ensures that the initialized large language model possesses the basic vocabulary and feature representations suitable for speech recognition tasks at the start of training, enabling it to better handle speech recognition error correction issues.

[0048] In this invention, each training sample contains acoustic feature samples corresponding to the speech sample signal and text character sequence samples, which are frame-level acoustic feature sequences obtained after the speech sample signal is processed by the encoder of the speech recognition model.

[0049] Acoustic feature samples retain key acoustic information of the speech signal, providing acoustic input for the error correction model. Text character sequence samples are text sequences corresponding to the acoustic feature samples. These text character sequence samples can be preliminary recognition results from the speech recognition model or processed text sequences.

[0050] In practical applications, text character sequence samples in the training samples usually need to be aligned with acoustic feature samples in the time dimension. This alignment can be achieved through a forced alignment algorithm, so that each frame of acoustic features corresponds to a text character or whitespace marker.

[0051] By collecting and processing large amounts of speech data and corresponding text data, a training set containing multiple training samples can be constructed.

[0052] In this application, supervised learning is used for training. For each training sample, it is input into the initialized large language model, which outputs a predicted text sequence. Forced alignment of text labels serves as the training objective, providing the correct answer to be output for each frame. The loss function can be calculated by comparing the model's predicted output with the forced alignment of text labels.

[0053] During training, the gradient of the loss function with respect to each parameter of the model is calculated using the backpropagation algorithm. Then, the optimizer updates the model parameters based on the gradient. This process is repeated until the model's performance reaches the preset standard.

[0054] In this application, forced alignment of text labels provides explicit supervision signals for the acoustic features of each frame by precisely aligning real text labels with speech sample signals in the time dimension. This frame-level supervision enables the trained speech recognition error correction model to perform accurate streaming error correction, making accurate judgments and corrections when processing each frame.

[0055] When the training process is complete and the model's performance no longer improves or reaches the preset convergence condition, the trained speech recognition error correction model is obtained. The speech recognition error correction model inherits the language understanding ability of the large language model, and at the same time, through training on specific tasks, it acquires the ability to correct errors in speech recognition results.

[0056] In this application, a scheme for determining target dictionary characters and initializing the model is designed based on a pre-trained large language model, and training is performed using training samples with forced alignment text labels. This method fully utilizes the pre-trained knowledge of the large language model, avoiding the high cost of training a large-scale model from scratch. At the same time, the forced alignment training data ensures the streaming error correction capability of the model, providing an efficient and reliable error correction scheme for real-time speech recognition applications.

[0057] Optionally, before the step of training the initialized large language model based on each training sample carrying forced alignment text labels to obtain a speech recognition error correction model, the method further includes: The speech sample signal is input into the speech recognition model to obtain the probability distribution of outputting any text character or blank mark for each frame of audio sample of the speech sample signal; Based on the probability distribution, an alignment path search space is constructed, wherein the horizontal dimension of the alignment path search space corresponds to the time frame sequence of the speech sample signal, and the vertical dimension corresponds to the character sequence of the real text label. In the alignment path search space, a dynamic programming algorithm is used to find the alignment path with the highest probability from the start position to the end position, wherein the start position corresponds to the zeroth frame and the zeroth character, and the end position corresponds to the last frame and the last character. Based on the alignment path with the highest probability, a forced alignment text label sequence is generated, wherein each frame position in the forced alignment text label sequence corresponds to a real text character or a blank mark.

[0058] In this application, for a training data pair, the speech sample signal is first input into a pre-trained speech recognition model. This speech recognition model can be a model based on a recurrent neural network transducer (RNN-T). The model processes the input speech signal, and within the speech recognition model, for example, the outputs of the encoder and prediction networks generate a probability distribution for each audio frame sample of the speech signal.

[0059] This probability distribution describes the likelihood of any text character or a special whitespace marker appearing in the output vocabulary of the current frame and the current decoding step.

[0060] This output can be represented as a probability matrix, the dimensions of which are the number of time frames multiplied by the maximum number of decoding steps multiplied by the vocabulary size. This matrix forms the basis for subsequent alignment calculations. For each training data pair, we set the maximum number of decoding steps to the length of the actual text sequence of that training sample.

[0061] Based on the probability distribution obtained in the previous step, a two-dimensional alignment path search space can be constructed. This search space can be defined as a grid. The horizontal dimension of this grid corresponds to the time frame sequence of the speech sample signal, extending from frame zero to the last frame.

[0062] The vertical dimension of the grid corresponds to the character sequence of the actual text label, extending from the zeroth character to the last character. In this two-dimensional space, any path from the start position to the end position represents a possible alignment, i.e., how to correspond the audio frame sequence to the text character sequence.

[0063] In the constructed alignment path search space, there are multiple possible alignment paths. The next step is to determine the alignment path most likely to generate the actual text label, which can be efficiently solved using a dynamic programming algorithm.

[0064] Specifically, the dynamic programming algorithm starts from a starting position, corresponding to the zeroth frame and the zeroth character, i.e., before the audio input is fed into the RNN-T model, and progressively calculates the optimal path score to each point in the grid. The path score to a point is calculated based on the path score to its predecessor node and the transition probability from the predecessor node to the current node, i.e., the probability of outputting the corresponding character or blank marker in the corresponding time frame and decoding step.

[0065] The dynamic programming algorithm records the optimal predecessor node for each point. This process is iterated until the termination position is reached, which corresponds to the last frame and the last character. By backtracking from the termination position based on the recorded predecessor node information, the entire maximum probability alignment path can be reconstructed, and the final forced alignment text label sequence can be generated.

[0066] Specifically, starting from the last decoding step of the last frame, if the movement from the optimal predecessor node to the current node is horizontal, that is, the time frame increases but the corresponding text character index does not increase, then a blank mark is marked at the position of the forced alignment text label sequence corresponding to that time frame.

[0067] If the movement from the optimal predecessor node to the current node is diagonal, meaning both the time frame and the text character index increase simultaneously, then the position corresponding to that time frame is marked with the actual text character corresponding to the current node.

[0068] After this step, the final generated forced alignment text label sequence has a length that is exactly the same as the number of audio frames in the speech sample signal, and each position in the sequence clearly corresponds to a real text character or blank mark.

[0069] This embodiment addresses the temporal mismatch between speech signals and text labels in training data by introducing a dynamic programming-based forced alignment method. It identifies the most probable text character or whitespace marker for each frame of audio signal, generating high-quality, frame-level aligned labels. This forced alignment text label sequence provides a precise supervised learning target for subsequent training of the speech recognition error correction model, a crucial prerequisite for achieving high-performance streaming error correction.

[0070] Optionally, in the alignment path search space, a dynamic programming algorithm is used to find the alignment path with the highest probability from the starting position to the ending position, including: Construct a state transition table, wherein the number of rows in the state transition table is equal to the number of frames of the speech sample plus one, and the number of columns in the state transition table is equal to the number of characters in the real text label plus one; For each state point in the state transition table, calculate the path score from the previous state to the current state, where the path includes a horizontal transition path that outputs a blank marker and a diagonal transition path that outputs text characters. Select the transition path with the highest score as the optimal path to the current state point, and record this path selection in the backtracking pointer table; Starting from the end position of the table, trace back to the beginning position using the backtracking pointer table to reconstruct the alignment path with the highest probability from the beginning position to the end position.

[0071] In this application, a two-dimensional state transition table is first constructed. If the number of frames of the speech sample is T and the number of characters of the real text label is U, then the number of rows in the state transition table is equal to T plus one, and the number of columns is equal to U plus one.

[0072] Each state point (t, u) in the table represents an alignment state, meaning that when processing reaches the t-th time frame, u real text characters have been output.

[0073] At the same time, a backtracking pointer table of the same size as the state transition table is constructed to record the source of the optimal path.

[0074] When filling the state transition table, the optimal path score for each state point (t, u) in the table needs to be calculated one by one. For any state point, its optimal path score is determined by comparing the scores of different paths that transitioned to that point from the previous time step (i.e., frame t-1).

[0075] These paths mainly include two types: one is a horizontal transition path representing the output of blank markers, originating from state point (t-1, u); the other is a diagonal transition path representing the output of text characters, originating from state point (t-1, u-1). After calculating and comparing the scores of these two transition paths, the one with the higher score, i.e., the one with the highest probability, is taken as the cumulative score of the optimal path to the current state point (t, u) and stored in the state transition table.

[0076] At the same time, the source direction of the optimal path is recorded at the corresponding position (t, u) in the backtracking pointer table, i.e., whether it comes from a horizontal transfer or a diagonal transfer.

[0077] Once the entire state transition table and backtrack pointer table are filled, the path with the highest probability alignment can be reconstructed.

[0078] The reconstruction process begins at the end position of the table, i.e., state point (T, U), and traces backwards according to the source direction recorded in the backtracking pointer table, gradually backtracking to the beginning position of the table, i.e., state point (0, 0). The resulting path chain is the highest probability alignment path from the beginning position to the end position.

[0079] In this application, by constructing a state transition table and a backtracking pointer table, and utilizing the idea of ​​dynamic programming, the optimal transition path is calculated and selected point by point. Finally, through backtracking, the globally optimal alignment path can be found deterministically and efficiently. This method is the core of implementing the forced alignment algorithm, ensuring the accuracy and optimality of the generated forced alignment text label sequence.

[0080] Optionally, the method for calculating the path score includes: Initialize the starting state of the state transition table, and set the cumulative score of the zeroth character position in the zeroth frame to zero; The cumulative score of each state point in the state transition table is calculated frame by frame, starting from frame zero, according to the time frame order. For each character position in the current frame, based on the cumulative state score of the previous frame and the probability distribution of the current frame, the maximum cumulative score to reach the character position is calculated to obtain the path score; While calculating the cumulative score, the source direction of the optimal predecessor state is recorded at the corresponding position in the backtracking pointer table.

[0081] In this application, the state transition table is first initialized. Specifically, the initial state of the state transition table is initialized, that is, the cumulative score of the zeroth character position in the zeroth frame is set to zero. In the logarithmic probability domain, a score of zero is equivalent to a probability of 1, which means that the probability of being at the starting position is certain before processing begins.

[0082] Correspondingly, other initially unreachable state points in the table can be set to a minimum value, such as negative infinity, to ensure that they are not selected as the starting point of the optimal path.

[0083] After initialization, the cumulative score of each state point in the state transition table is calculated frame by frame, starting from the first frame, according to the time frame order. This calculation process is usually implemented through a nested loop, with the outer loop traversing time frame t (from 1 to T) and the inner loop traversing character position u (from 1 to U).

[0084] The calculation of the path score (i.e. the maximum cumulative score) for each character position (t, u) in the current frame requires the calculation of the state cumulative score of the previous frame (t-1) and the probability distribution of the current frame.

[0085] Specifically, the score for reaching (t, u) via the lateral transfer path is calculated: the score is the cumulative score of the state point (t-1, u), plus the log probability of outputting a blank marker at frame t.

[0086] Calculate the score for reaching (t, u) via the diagonal transition path: the score is the cumulative score of the state point (t-1, u-1), plus the log probability of the u-th character in the output real text label at frame t.

[0087] Then, compare the two scores and take the maximum value as the maximum cumulative score to reach the current state point (t, u), and update the corresponding position in the state transition table. This maximum cumulative score is the path score for the current state point.

[0088] Finally, while calculating the cumulative score, the source direction of the optimal predecessor state is recorded at the corresponding position in the backtracking pointer table. After determining whether the maximum cumulative score for reaching the state point (t, u) is contributed by a lateral or diagonal transition, a corresponding record is made at the (t, u) position in the backtracking pointer table.

[0089] For example, a marker can be recorded (such as 0 representing horizontal and 1 representing diagonal), or the coordinates of the predecessor state point can be stored directly.

[0090] In this application, the path score calculation method decomposes the global optimal path search problem into a series of locally optimal decisions at each time frame. By iteratively calculating the cumulative score frame by frame and recording the source direction, this method ensures that the dynamic programming algorithm can accurately and efficiently fill the entire state transition table, providing a solid computational foundation for finally finding the alignment path with the highest probability.

[0091] In one alternative embodiment, a dynamic programming-based forced alignment algorithm is used, with initial data including: a sequence of real label texts. The sequence does not contain whitespace and has a length of U. The probability matrix output by the RNN-T model. Its dimensions are Where B is the batch size, T is the number of audio frames, U is the maximum number of decoding steps, and V is the number of characters including whitespace. the vocabulary size. The alignment objects of the alignment algorithm are the audio frame sequence (T frames) and the ground-truth label text sequence (length U), and finally an LLM target sequence with a length strictly equal to T is generated , wherein , all B in Z are removed to obtain the ground-truth label text sequence Y.

[0092] For a data pair including an audio with T audio frames in one decoding step and the corresponding ground-truth label text sequence Y with length U, there may be multiple alignment paths, Figure 2 which is an exemplary diagram of an alignment path provided in the present application, Figure 2 showing an example of an alignment path where T=5 and U=2. Assuming the ground-truth label is the audio has a total of 5 frames, as shown by the solid arrows in the figure, starting from the starting point t=0, u=0 (representing that the initial state corresponds to the ASR output ), and ending at the end point t=5, u=2 (representing that ASR has output the last word "hǎo" before the last frame), there are multiple possible alignment paths, and after searching using the forced alignment algorithm based on dynamic programming, an optimal path can be obtained, for example, the path shown by the black solid arrow in Figure 2 .

[0093] Specifically, for any time t and decoding step u, whenever the RNN-T model completes one frame of inference, the frame number t is increased by 1, and there are only two possible decoding results, representing two possible transition paths: transitioning to the directly right position, indicating that ASR continues outputting without generating a new character; transitioning to the lower right position, indicating that ASR has generated a new character, and the decoding step u is increased by 1. To ensure that all the characters can be decoded and generated, when the number of remaining frames is less than or equal to the number of remaining characters, only transition to the lower right position is allowed.

[0094] To find the optimal alignment path, solving can be performed based on the dynamic programming algorithm. Construct a state table and a backpointer matrix to store intermediate results, so as to avoid repeated calculations.

[0095] represents the maximum score when the first t frames of audio are aligned to the first u non-empty labels in , when the algorithm is initialized, dp[0][0]=0.0 is set as the initial state, and dp[0][u] corresponding to u>0 is set to negative infinity. Starting from t=0, traversing to t=T, recursively update the state table dp and the backpointer matrix backpointers. During the state transition process, for each state (t,u), two possible transition paths are considered: one is transitioning from , which means the first t-1 frames of audio are aligned to In the first u non-empty labels, the RNN-T model outputs a whitespace character. The probability is Its candidate score is Another one is from This is a transfer, indicating that the audio of the first t-1 frames is aligned to... In the first u-1 non-empty labels, the RNN-T model outputs non-empty labels. Its candidate score is The algorithm selects the path with the higher score from the two paths mentioned above to update the current state, and records the selection result in the backtracking pointer matrix. 0 represents an output blank character. 1 indicates that a non-empty label will be output. If the scores of the two paths are equal, prioritize outputting a non-empty label. For the boundary case where u=0, since it can only originate from... It was transferred, therefore the candidate score is The above formula for calculating candidate scores is called the state transition equation. Based on this state transition equation, it can be deduced that the state at time t is only related to the state at time t-1. Related, and Given that, therefore, dp at the same time It can be computed in parallel.

[0096] After traversing to the final state (T, U) and filling in all the values ​​in dp and backpointers, backtracking begins from the final state (T, U'), and the frame-level whitespace-labeled sequence is gradually constructed based on the values ​​of the backpointers matrix. If it is 0, then update the value of the t-th bit of Z to a blank character. If the value is 1, then update the value of the t-th bit of Z to the u-th label of Y, repeating until backtracking to the initial state (0,0) to obtain the sequence Z corresponding to the optimal path. This design ensures both strict matching of sequence lengths and consistency of content, providing a reliable target for subsequent LLM training.

[0097] Figure 3 Example diagram of the dynamic programming alignment search algorithm provided in this application, such as Figure 3 As shown, the green table represents the filling process of dp, and the orange table represents the backpointers matrix. As mentioned above, dp[0][0]=0.0 is taken as the initial state, and the dp value at the boundary condition position at time t=0 is set to negative infinity (-inf). Then, starting from t=0, the system iterates through the state transition equation, calculates the scores corresponding to all decoding steps u at time t=1 in parallel, selects the updated state with the higher score, and records the selection result. For example, at t=1, when u=1, a non-empty character is generated. The path with the higher score is recorded as backpointers[1][1]=1. This process is repeated until t=5. In the example, let the blue transition equation in the graph correspond to the path with the higher score.

[0098] Then, starting from the endpoint backpointers[5][2], since the value is 1, the path is transferred from the top left corner t=4, u=1, and the position of Z at t=5 is... Update to non-empty character Next, we examine the values ​​corresponding to backpointers[4][1]. Since the value is 1, the path originates from the leftmost position t=3, u=1. Therefore, the position of Z at t=4 is... Update to whitespace By following this pattern, tracing back to the starting point u=0, t=0, the optimal path is obtained.

[0099] Optionally, the initialized large language model is trained based on each training sample carrying a forced alignment text label to obtain a speech recognition error correction model, including: for each training sample carrying a forced alignment text label, inputting the training sample into the large language model and outputting the corrected text character sequence sample corresponding to the training sample; and calculating the connection temporal classification loss and cross-entropy loss based on the corrected text character sequence sample and the forced alignment text label, wherein the cross-entropy loss is used to measure the difference between the prediction result of each audio frame sample and the forced alignment text label, and the connection temporal classification loss is used to measure the accuracy of the overall prediction. Based on the connection-time classification loss and the cross-entropy loss, the overall loss is determined, and the parameters of the large language model are updated through the backpropagation algorithm. If the total loss is less than a first preset threshold, training is stopped, and a trained speech recognition error correction model is obtained.

[0100] The step of inputting the training samples into the large language model also includes applying a causal masking mechanism. The causal mask matrix restricts the large language model to only use the current position, audio frame samples before the current position, and text labels before the current position when predicting the output at the current position.

[0101] In this application, for each training sample in the training set carrying a forced-alignment text label, the training sample is input into the large language model under the current parameter state. This process also includes the application of a causal masking mechanism. Specifically, when the large language model processes the input sequence, it applies a causal masking matrix. This matrix restricts the large language model to using only the current position and the audio frame samples and text label information preceding the current position when predicting the output at the current position. The large language model performs one forward propagation calculation and outputs a corrected text character sequence sample of the same length as the input audio frame sequence.

[0102] After obtaining the model's predicted output, the connection time classification loss and cross-entropy loss are calculated based on the corrected text character sequence samples and the forced alignment text labels.

[0103] The cross-entropy loss is used to measure the difference between the prediction result of each audio frame sample and the forced alignment text label, providing a fine and powerful supervision signal.

[0104] The connection-time classification loss is used to measure the accuracy of the overall prediction and provides sequence-level global supervision for the model.

[0105] Then, based on the connection-time classification loss and the cross-entropy loss, the overall loss is determined. The overall loss is typically a weighted sum of these two losses. After determining the overall loss, the gradient of the overall loss with respect to each trainable parameter of the large language model is calculated using the backpropagation algorithm, and the optimizer is used to update the parameters of the large language model based on this gradient.

[0106] Finally, repeat the forward propagation, loss calculation, and backpropagation parameter update process described above. Training stops when the overall loss is less than a first preset threshold, or when the validation set performance no longer improves. At this point, the trained large language model is the final speech recognition error correction model required.

[0107] This application ensures the streaming processing capability of the model by introducing a causal masking mechanism. Simultaneously, a dual-objective training strategy combining cross-entropy loss and connection-temporal classification loss provides comprehensive supervision from the frame level to the sequence level, enabling the final trained speech recognition error correction model to guarantee both the accuracy of the recognized text and achieve precise real-time error correction.

[0108] In an optional embodiment, after obtaining training-ready speech and label data pairs, we designed a dual-objective joint optimization framework to train the LLM. The output layer of the LLM head is designed as a shared structure, and its generated output simultaneously serves two training objectives: cross-entropy loss. Connectionist Temporal Classification (CTC) loss This allows the model to simultaneously learn forced alignment of text labels. Combine the real label text Y with the data to improve the robustness of the system.

[0109] First, we calculate the LLM output probability. With target sequence Cross-entropy loss : ); Secondly, due to the target sequence It is relatively sparse and contains a large number of whitespace characters. To help the model converge better, we further introduce the output sequence. With target sequence CTC loss: ) In its implementation, this scheme uses dynamic adjustment of coefficients. This allows for a gradual shift in training focus. The training objective loss function of LLM can be expressed as: In the early stages of training, set a lower [configuration]. The value is adjusted to focus on global path learning in CTC; it increases linearly as training progresses. The system progressively strengthens the accuracy requirements of frame-level alignment. This dynamic balancing mechanism effectively avoids conflicts between different optimization objectives. In final deployment, the system only needs simple probability sampling and whitespace filtering to obtain the recognition results.

[0110] In addition, a Text Encoder module is introduced, which significantly improves the context consistency of the speech recognition system during streaming processing by establishing a memory mechanism for historical outputs. This innovative design enables the language model to dynamically refer to previously recognized content, effectively improving the error propagation problem in long text recognition. In its implementation, we designed a lightweight text encoder that receives previously recognized historical text sequences as input and encodes them into a dense memory representation through a multi-layer Transformer structure. The memory fusion process is implemented using a cross-attention mechanism.

[0111] A cross-attention module is inserted into the intermediate layer of the language model. The memory information generated by the text encoder is used as a key / value cache, and the current hidden state of the language model is used as the query vector to calculate attention weights for interaction. To meet the requirements of streaming processing, a strict causal mask is applied during attention calculation based on the position of the text in Z, ensuring that the model can only access historical information prior to that time step when decoding each time step, fully conforming to the constraints of real-time speech recognition scenarios.

[0112] Optionally, the method further includes: Obtain the characters in the speech recognition dictionary that are the same as those in the large language model dictionary of the preset large language model, and get the first dictionary characters; Obtain dictionary characters that exist in the speech recognition dictionary but do not appear in the large language model dictionary to obtain the second dictionary characters; Encoding each character in the second dictionary characters through a byte-level byte pair encoding method to obtain a byte pair encoding sequence corresponding to each second dictionary character, and constructing a dictionary character mapping table; wherein, the dictionary character mapping table uses the second dictionary character as a key, and uses the byte pair encoding sequence corresponding to the second dictionary character as a value; Obtaining a target dictionary character based on the first dictionary character and the dictionary character mapping table.

[0113] In the present invention, the speech recognition dictionary is a vocabulary used by a speech recognition model, and includes a set of all characters or words that are expected to appear in speech signals.

[0114] The large language model dictionary is a vocabulary supported by a preset large language model. Since large language models are usually trained on large-scale text data, their dictionary is very extensive and covers vocabulary of various topics and fields.

[0115] In the present invention, each character in the speech recognition dictionary is compared with characters in the preset large language model dictionary, characters that appear in both of the two dictionaries are found out, and these characters are aggregated to form the first dictionary character.

[0116] For example, assuming that the characters in the speech recognition dictionary are "I, am, student, study, speech, recognition", and the characters in the large language model dictionary are "I, am, human, study, nature, speech, intelligence", then the first dictionary character obtained through comparison is "I, am, study, speech".

[0117] In the present invention, the large language model uses byte-encoded keys, while speech recognition uses single characters or letters, and their representation forms are different. Therefore, there may be dictionary characters that exist in the speech recognition dictionary but do not appear as a complete token in the large language model dictionary among the dictionary characters of the large language model.

[0118] The dictionary of a large language model usually encodes text in the form of BBPE, has a wider coverage, and encodes text into a variable-length byte subword sequence, which is different from the representation form of the dictionary used in speech recognition.

[0119] For example, the index of the character "I" in the speech recognition model dictionary is 62, while the large language model dictionary uses BBPE to encode "I" into a subword sequence corresponding to two indexes 1378 and 564. During initialization, the average of the word embedding vectors of the subwords corresponding to the two indexes 1378 and 564 in the large model is calculated, and the result is used as the initial value of the word embedding vector of the 62nd word in the new word embedding vector.

[0120] In other words, the speech recognition dictionary and the large language model dictionary may contain characters with different representations. Therefore, we obtain the dictionary characters that exist in the speech recognition dictionary but do not appear as a complete word in the large language model dictionary, i.e., the second dictionary characters.

[0121] Using a large language model and an encoding method based on byte-level byte-pair encoding, each second dictionary character is used as input for transformation, generating an ordered sequence of subwords or lexical units from a large language model vocabulary for each input character, i.e., a byte-pair encoding sequence.

[0122] In order to enable this conversion relationship to be accurately queried and used in subsequent steps, a dictionary character mapping table is constructed using the second dictionary character as the key and the byte pair encoding sequence corresponding to the second dictionary character as the value.

[0123] In the dictionary character mapping table, the original second dictionary character is used as a unique lookup key, while the encoded byte pair sequence obtained after encoding is stored as the value corresponding to that key. This establishes a bridge from unknown words in the speech recognition dictionary to known sub-word sequences in the large language model, providing a clear and unambiguous lookup basis for subsequent initialization parameter calculations.

[0124] Optionally, the initialization process of the embedding layer and output layer in the initial large language model based on the target dictionary characters to obtain the initialized large language model includes: Copy the embedding vector and output layer vector corresponding to the first dictionary character in the target dictionary characters into the speech recognition dictionary initialization parameters; For each second dictionary character in the dictionary character mapping table, the embedding vector and output layer vector of each word in the large language model dictionary in the byte pair encoding sequence corresponding to the second dictionary character are averaged to obtain an averaged vector, and the averaged vector is copied as the initialization vector of the second dictionary character into the speech recognition dictionary initialization parameters. Based on the initialization parameters of the speech recognition dictionary, the embedding layer and output layer in the initial large language model are initialized to obtain the initialized large language model.

[0125] In this invention, for the first dictionary character in the target dictionary characters, the embedding vector and output layer vector of the first dictionary character in the preset large language model are directly copied to the speech recognition dictionary initialization parameters, which retains the existing information of the preset large language model and provides good initial parameters for the model.

[0126] For any second dictionary character, its corresponding byte pair encoding sequence is first looked up through the dictionary character mapping table. This sequence consists of one or more token indices. Based on these indices, all corresponding embedding vector sets and output layer vector sets are obtained from the initial large language model.

[0127] Then, an averaging operation is performed on the obtained vector set. This averaging operation can be an arithmetic average, that is, summing each element of all vectors in the vector set, and then dividing each element of the resulting vector by the number of vectors in the set, thereby generating a synthesized average vector.

[0128] This average vector semantically constitutes an approximate representation of the characters in the original second dictionary, effectively integrating the semantic information of its various sub-word components.

[0129] Finally, the averaged vector is used as the initialization vector for the second dictionary character and loaded into the corresponding position in the speech recognition dictionary initialization parameter set.

[0130] In this application, the speech recognition dictionary initialization parameter set integrates the direct copy parameters of all characters in the first dictionary and the synthetic average parameters of all characters in the second dictionary, forming a complete parameter set perfectly aligned with the vocabulary size of the speech recognition dictionary. Based on this parameter set, the embedding and output layers of the initial large language model are initialized.

[0131] The sizes of the embedding and output layers are adjusted to match the vocabulary size of the speech recognition dictionary, and the weights are loaded using the speech recognition dictionary's initialization parameters. This results in a large language model with an adapted vocabulary and parameters that fully inherit prior knowledge after initialization.

[0132] In this invention, the embedding vector in the speech recognition dictionary initialization parameters is assigned to the vector at the corresponding position in the initial large language model embedding layer. In this way, each character has an initial vector representation in the embedding layer, thus completing the embedding layer initialization.

[0133] Similarly, the output layer vector from the speech recognition dictionary initialization parameters is assigned to the vector at the corresponding position in the output layer of the initial large language model. This ensures that each character has an initial probability distribution when the model outputs, thus completing the output layer initialization.

[0134] In this invention, after the above initialization steps, the embedding and output layers of the initial large language model have been initialized according to the speech recognition dictionary initialization parameters. At this point, the large language model is adapted to the needs of the speech recognition task and can better handle words and characters in speech recognition.

[0135] Optionally, after the step of initializing the embedding layer and output layer of the initial large language model based on the target dictionary characters to obtain the initialized large language model, the method further includes: Freeze the backbone transformer module in the large language model, and fine-tune the initialized embedding layer and output layer in the large language model based on multiple monolingual training samples; If the first preset condition is met, training is stopped, and a fine-tuned large language model is obtained.

[0136] In this invention, the step of fine-tuning the large language model can be performed after the initialization of the initial large language model is completed.

[0137] The backbone transformer module is the core part of a large language model. It is usually composed of multiple layers of transformers and is responsible for processing the input embedding vectors and generating context-sensitive representations.

[0138] In this invention, the freeze operation maintains the parameters of the backbone converter module unchanged during fine-tuning training. After the freeze operation, the parameters of the backbone converter module will not be updated during training.

[0139] In this invention, monolingual training samples are training data used to fine-tune a large language model, containing text data in a specific language. Each sample may include text sentence samples or paragraph samples.

[0140] The embedding and output layers of the large language model are further trained using monolingual training samples. This allows the large language model to better adapt to the data distribution of a specific language or task.

[0141] More specifically, monolingual training samples are input into a large language model, and forward propagation is performed through the embedding layer, backbone transformer module, and output layer of the large language model to obtain the prediction results.

[0142] Calculate the loss between the predicted results of the monolingual training samples and the corresponding ground truth labels, for example, using cross-entropy loss. Since the backbone transformer module is frozen, backpropagation is performed only on the parameters of the embedding and output layers to update the parameters of these layers to minimize the loss.

[0143] In this invention, the first preset condition can be various, such as training reaching a certain number of iterations, the loss value falling below a certain threshold, or performance on the validation set no longer improving. When the first preset condition is met, the fine-tuning training process of the large language model is stopped, resulting in the fine-tuned large language model.

[0144] In this invention, while maintaining the core capabilities of the large language model, the embedding and output layers are fine-tuned to adapt it to specific speech recognition tasks. Freezing the backbone transformer module reduces computational resource consumption while preserving the model's general language understanding capabilities. Through fine-tuning training, the fine-tuned large language model can better handle speech recognition data of specific languages, improving the accuracy and error correction capabilities of speech recognition.

[0145] Figure 4 This is a schematic diagram of the speech recognition dictionary adapted to the large language model in this invention, as shown below. Figure 4 As shown, the core Transformer modules of the Large Language Model (LLM) are frozen, meaning their parameters are not updated during training. The Transformer modules mainly consist of self-attention mechanisms, responsible for processing the input sequence and generating context-sensitive representations.

[0146] The embedding and output layers of the LLM are replaced with layers adapted to the ASR dictionary. The replacement process can be understood as the initialization operation mentioned in the above embodiments.

[0147] For characters that exist in the ASR dictionary and are also in the LLM original dictionary (such as...) <s>、< / s> and <unk>(Special markers, etc.) copy the embedding and output vectors corresponding to these characters in the original LLM dictionary to the speech recognition dictionary initialization parameters.

[0148] For characters present in the ASR dictionary but not in the original LLM dictionary (such as specific Chinese characters, English words, etc.), these characters are converted into byte-level byte-pair encoding. Then, the embedding and output vectors associated with these new characters in the original LLM dictionary are averaged and copied into the speech recognition dictionary initialization parameters. The embedding and output layers are then initialized according to the speech recognition dictionary initialization parameters.

[0149] For example, <unk>The vector corresponding to the character at index 15123 is copied 3 times; the vectors corresponding to the character "wo (I)" at indices 1378 and 564 are averaged, and 62 average vectors are generated.

[0150] In the present invention, the overall architecture and most parameters of the large language model remain unchanged, and only the embedding layer and output layer are adjusted to adapt to the new speech recognition task.

[0151] The initialization operation achieves the goal of enabling the large language model to adapt to the ASR task-specific vocabulary and characters, while retaining the parameters and structure of other parts of the model. This not only effectively utilizes the knowledge of the pre-trained model, but also improves the performance of the model on the speech recognition task through targeted initialization and fine-tuning.

[0152] Figure 5 is a schematic diagram of the framework of the speech recognition error correction model provided by the present application, as Figure 5 shown, first, the speech signal is input into an ASR encoder to extract the corresponding frame-level acoustic feature sequence.

[0153] the acoustic feature sequence is then sent to an ASR joint network. This network is the core part of the front-end speech recognition model. It performs preliminary speech recognition based on acoustic features and outputs a preliminary recognition result sequence containing errors and blank tokens, which is represented by the green square sequence.

[0154] For example, for the real text "Summer Palace is really beautiful", the ASR joint network may output an incorrect sequence, such as "B yi Yi B B He B Yuan B B B B Mei", where "B" represents a blank token, "Yi (the character Yi in Summer Palace)" is incorrectly recognized as "yi Yi", "He (the character He in Summer Palace)" is incorrectly recognized as "He (lotus)", and the character "Zhen (really)" is missing.

[0155] Next, it enters the core training part of the speech recognition error correction model. A fusion module (Fusion+MLP+EncStack) fuses the original acoustic features from the ASR encoder and the preliminary recognition results from the ASR joint network. Through structures such as multi-layer perceptron (MLP) and encoder stack (EncStack), the fusion module integrates acoustic information and preliminary text information into a richer mixed feature representation. The parameters of this fusion module are trainable (as shown by the flame icon in the figure).

[0156] the mixed feature representation is used as the main input and fed into a large language model (LLM) which serves as the core of back-end error correction. In order to provide stronger guidance for the large language model during the training process, the present embodiment also introduces a parallel text processing path.

[0157] The ground truth text label ("The Summer Palace is beautiful") is input into a Text Encoder to generate the corresponding text representation. The text representation is then stored in a Key / Value Memory.

[0158] Inside the large language model, through a Cross Attention mechanism, the mixed feature representation from the fusion module is used as Query, and the text representation from the Key / Value Memory is used as Key and Value. In this way, when performing error correction prediction, the large language model can directly "refer to" the representation of the correct answer, thereby learning a more accurate error correction mapping relationship.

[0159] To efficiently fine-tune the large language model, this embodiment adopts LORA (Low-Rank Adaptation) technology. Only a small number of LORA parameters inserted into the large language model are trained, while the main parameters of the large language model remain frozen. This parameter-efficient training method greatly reduces the computing resources and time required for training.

[0160] The final output of the large language model is a corrected text character sequence that has the same length as the input audio frame. During training, the target of this output is to be completely consistent with the force-aligned text label corresponding to the ground truth text label (represented by a sequence of purple squares in the figure, for example, "B B B Yi B B He Yuan B Zhen B B Mei").

[0161] Finally, based on the output of the model and the force-aligned text labels, cross-entropy loss is calculated. The loss compares the predicted output of the model with the force-aligned text labels frame by frame, aiming to optimize the prediction accuracy of the model on each frame.

[0162] Connectionist Temporal Classification loss: This loss measures the consistency between the overall prediction sequence of the model after removing blanks and repeated characters and the original ground truth text label - "The Summer Palace is beautiful", aiming to optimize the overall accuracy of the final output text. These two losses together constitute the total loss of training, which is used to update the trainable parameters of the fusion module and the LORA module through the back propagation algorithm.

[0163] In the present application, by designing a training framework that includes multiple inputs of acoustic features, preliminary recognition results and ground truth text representations, and combining LORA efficient fine-tuning technology and a dual-objective loss function, the powerful language capability of the pre-trained large language model can be efficiently transferred to the speech recognition error correction task, thereby training a high-performance streaming speech recognition error correction model.

[0164] Optionally, the method further comprises: For any character in the text character sequence sample of the training samples, perturb the character alignment information or the character content to generate multiple training samples.

[0165] In this invention, during speech recognition, the text character sequence needs to be aligned with the acoustic feature sequence of the speech signal to determine the speech frame position corresponding to each character. Character alignment information perturbation refers to making small adjustments to the alignment position of characters, enabling the model to adapt to minor changes in alignment position and improving the flexibility and robustness of alignment.

[0166] In this invention, character content perturbation refers to replacing character content while maintaining the alignment position essentially unchanged, thereby generating different text sequences. This can help the model better handle easily confused characters such as homophones or synonyms.

[0167] Figure 6 The disturbance diagram provided by the present invention is as follows: Figure 6 As shown, during the decoding process of RNNT, multiple candidate paths (such as beam1 to beam4) are generated. Among them, beam2 ({aecd}) is a candidate sequence of the same length as the actual label but is incorrect. Such a sequence is selected to increase the diversity of training samples.

[0168] In addition to candidate paths generated using RNNT, other encoder-decoder (ED) models are employed to generate candidate sequences of equal length and with errors, further enriching the training samples. For example, beam4({abcdf}) might be generated using other encoder-decoder models.

[0169] Using LLM to replace some homophones to generate new candidate sequences, such as LLM / ed({abgh}), can increase the diversity of character content in the training samples.

[0170] On the other hand, by perturbing the RNNT-aligned sequence, the alignment position of the speech frames can be varied within a certain range.

[0171] For example, each row represents a speech frame from the encoder, and each column represents the output recognition result. The first result could be either the output of the second frame (red circle) or the output of the third frame (green circle). This approach enriches the alignment of speech frames and improves the model's adaptability to changes in alignment position.

[0172] In this invention, character alignment information perturbation or character content perturbation generates multiple training samples, each containing different alignment information or character content. This diversity of training samples helps the model better learn various variations in speech recognition, improving the model's robustness and error correction capabilities.

[0173] Figure 7 This is a schematic diagram of the speech recognition error correction device provided by the present invention, as shown below. Figure 7 As shown, it includes: The recognition module 710 is used to input the current audio frame in the user's audio to be recognized into the speech recognition model to obtain the acoustic features and the first text character of the current audio frame; The error correction module 720 is used to input the acoustic features of the current audio frame, the first text character, and the historical corrected text sequence into the speech recognition error correction model to obtain the corrected text character corresponding to the current audio frame. The historical corrected text sequence is a corrected text sequence obtained by sequentially passing the speech recognition model and the speech recognition error correction model through the historical audio frames in the user audio to be recognized before the current audio frame; the speech recognition error correction model is trained based on training samples carrying forced alignment text labels, and the forced alignment text labels are obtained by forcibly aligning the real text labels corresponding to the training samples with the speech sample signals corresponding to the training samples.

[0174] In this application, a forced alignment method is used to generate frame-by-frame accurate supervisory labels synchronized with the speech frames for the training data, thus solving the core challenge of training streaming error correction models. This clear training objective enables the model to efficiently learn how to locate and correct recognition errors. Combined with the powerful context awareness capabilities of a large language model, the resulting speech recognition error correction model not only responds quickly but also exhibits higher robustness and error correction accuracy when handling long and complex sentences. Furthermore, by directly feeding the frame-by-frame output of the speech recognition model into the speech recognition error correction model, end-to-end streaming processing of recognition and error correction is achieved without waiting for the entire sentence to end. This significantly shortens the time for the first character to appear on screen and the time for correcting intermediate results, providing users with a smooth, low-latency real-time speech recognition experience.

[0175] Figure 8 This is a schematic diagram of the structure of the electronic device provided by the present invention, such as... Figure 8 As shown, the electronic device may include: a processor 810, a communications interface 820, a memory 830, and a communication bus 840, wherein the processor 810, the communications interface 820, and the memory 830 communicate with each other via the communication bus 840. The processor 810 can call logical instructions in the memory 830 to execute a speech recognition error correction method, which includes: Input the current audio frame in the user's audio to be identified into the speech recognition model to obtain the acoustic features and the first text character of the current audio frame; The acoustic features of the current audio frame, the first text character, and the historical corrected text sequence are input into the speech recognition error correction model to obtain the corrected text character corresponding to the current audio frame. The historical corrected text sequence is the corrected text sequence obtained by passing the historical audio frames in the user audio to be identified before the current audio frame through the speech recognition model and the speech recognition error correction model in sequence. The speech recognition error correction model is trained based on training samples carrying forced alignment text labels. The forced alignment text labels are obtained by forcibly aligning the real text labels corresponding to the training samples with the speech sample signals corresponding to the training samples.

[0176] Furthermore, the logical instructions in the aforementioned memory 830 can be implemented as software functional units and, when sold or used as independent products, can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present invention, essentially, or the part that contributes to the prior art, or a part of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute all or part of the steps of the methods described in the various embodiments of the present invention. The aforementioned storage medium includes various media capable of storing program code, such as USB flash drives, portable hard drives, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical disks.

[0177] On the other hand, the present invention also provides a computer program product, the computer program product comprising a computer program that can be stored on a non-transitory computer-readable storage medium, wherein when the computer program is executed by a processor, the computer is able to execute the speech recognition error correction method provided by the above methods, the method comprising: Input the current audio frame in the user's audio to be identified into the speech recognition model to obtain the acoustic features and the first text character of the current audio frame; The acoustic features of the current audio frame, the first text character, and the historical corrected text sequence are input into the speech recognition error correction model to obtain the corrected text character corresponding to the current audio frame. The historical corrected text sequence is the corrected text sequence obtained by passing the historical audio frames in the user audio to be identified before the current audio frame through the speech recognition model and the speech recognition error correction model in sequence. The speech recognition error correction model is trained based on training samples carrying forced alignment text labels. The forced alignment text labels are obtained by forcibly aligning the real text labels corresponding to the training samples with the speech sample signals corresponding to the training samples.

[0178] In another aspect, the present invention also provides a non-transitory computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, is implemented to perform the speech recognition error correction methods provided by the above methods, the method comprising: Input the current audio frame in the user's audio to be identified into the speech recognition model to obtain the acoustic features and the first text character of the current audio frame; The acoustic features of the current audio frame, the first text character, and the historical corrected text sequence are input into the speech recognition error correction model to obtain the corrected text character corresponding to the current audio frame. The historical corrected text sequence is the corrected text sequence obtained by passing the historical audio frames in the user audio to be identified before the current audio frame through the speech recognition model and the speech recognition error correction model in sequence. The speech recognition error correction model is trained based on training samples carrying forced alignment text labels. The forced alignment text labels are obtained by forcibly aligning the real text labels corresponding to the training samples with the speech sample signals corresponding to the training samples.

[0179] The device embodiments described above are merely illustrative. The units described as separate components may or may not be physically separate. The components shown as units may or may not be physical units; that is, they may be located in one place or distributed across multiple network units. Some or all of the modules can be selected to achieve the purpose of this embodiment according to actual needs. Those skilled in the art can understand and implement this without any creative effort.

[0180] Through the above description of the embodiments, those skilled in the art can clearly understand that each embodiment can be implemented by means of software plus necessary general-purpose hardware platforms, and of course, it can also be implemented by hardware. Based on this understanding, the above technical solutions, in essence or the part that contributes to the prior art, can be embodied in the form of a software product. This computer software product can be stored in a computer-readable storage medium, such as ROM / RAM, magnetic disk, optical disk, etc., and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute the methods described in the various embodiments or some parts of the embodiments.

[0181] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, and not to limit them; although the present invention has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some of the technical features; and these modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of the present invention.< / unk> < / unk>

Claims

1. A speech recognition error correction method, characterized in that, include: Input the current audio frame in the user's audio to be identified into the speech recognition model to obtain the acoustic features and the first text character of the current audio frame; Based on the target dictionary characters, the embedding layer and output layer in the initial large language model are initialized to obtain the initialized large language model; wherein, the target dictionary characters are determined based on the speech recognition dictionary and the large language model dictionary of the preset large language model; The acoustic feature sample and text character sequence sample corresponding to the speech sample signal are used as one training sample to obtain multiple training samples; The speech sample signal is input into the speech recognition model to obtain the probability distribution of outputting any text character or blank mark for each frame of audio sample of the speech sample signal; Based on the probability distribution, an alignment path search space is constructed. The horizontal dimension of the alignment path search space corresponds to the time frame sequence of the speech sample signal, and the vertical dimension corresponds to the character sequence of the real text label corresponding to the training sample. In the alignment path search space, a dynamic programming algorithm is used to find the alignment path with the highest probability from the start position to the end position, wherein the start position corresponds to the zeroth frame and the zeroth character, and the end position corresponds to the last frame and the last character. Based on the alignment path with the highest probability, a forced alignment text label sequence is generated, wherein each frame position in the forced alignment text label sequence corresponds to a real text character or a blank mark; The large language model after initialization is trained based on each training sample carrying a forced alignment text label to obtain a speech recognition error correction model. The acoustic features of the current audio frame, the first text character, and the historical corrected text sequence are input into the speech recognition error correction model to obtain the corrected text character corresponding to the current audio frame. The historical corrected text sequence is the corrected text sequence obtained by passing the historical audio frames in the user audio to be identified before the current audio frame through the speech recognition model and the speech recognition error correction model in sequence. The speech recognition error correction model is trained based on training samples carrying forced alignment text labels, which are obtained by forcibly aligning the real text labels corresponding to the training samples with the speech sample signals corresponding to the training samples.

2. The speech recognition error correction method according to claim 1, characterized in that, In the alignment path search space, a dynamic programming algorithm is used to find the alignment path with the highest probability from the starting position to the ending position, including: Construct a state transition table, wherein the number of rows in the state transition table is equal to the number of frames of the speech sample plus one, and the number of columns in the state transition table is equal to the number of characters in the real text label plus one; For each state point in the state transition table, calculate the path score from the previous state to the current state, where the path includes a horizontal transition path that outputs a blank marker and a diagonal transition path that outputs text characters. Select the transition path with the highest probability as the optimal path to the current state point, and record the path selection of the optimal path in the backtracking pointer table; Starting from the end position of the table, trace back to the beginning position using the backtracking pointer table to reconstruct the alignment path with the highest probability from the beginning position to the end position.

3. The speech recognition error correction method according to claim 2, characterized in that, The method for calculating the path score includes: Initialize the starting state of the state transition table, and set the cumulative score of the zeroth character position in the zeroth frame to zero; The cumulative score of each state point in the state transition table is calculated frame by frame, starting from the first frame, according to the time frame order. For each character position in the current frame, based on the cumulative state score of the previous frame and the probability distribution of the current frame, the maximum cumulative score to reach the character position is calculated to obtain the path score; While calculating the cumulative score, the source direction of the optimal predecessor state is recorded at the corresponding position in the backtracking pointer table.

4. The speech recognition error correction method according to claim 1, characterized in that, Based on the training samples carrying forced alignment text labels, the initialized large language model is trained to obtain a speech recognition error correction model, including: For each training sample carrying a forced alignment text label, the training sample is input into the initialized large language model, and the corrected text character sequence sample corresponding to the training sample is output. Based on the corrected text character sequence sample and the forced alignment text label, the connection temporal classification loss and cross-entropy loss are calculated, wherein the cross-entropy loss is used to measure the difference between the prediction result of each audio frame sample and the forced alignment text label, and the connection temporal classification loss is used to measure the accuracy of the overall prediction. Based on the connection-time classification loss and the cross-entropy loss, the overall loss is determined, and the parameters of the large language model after initialization are updated using the backpropagation algorithm. If the total loss is less than a first preset threshold, training is stopped, and a trained speech recognition error correction model is obtained.

5. The speech recognition error correction method according to claim 4, characterized in that, The step of inputting the training samples into the initialized large language model also includes applying a causal masking mechanism. The causal mask matrix restricts the large language model after initialization to only use the current position, the audio frame samples before the current position, and the text tags before the current position when predicting the output at the current position.

6. The speech recognition error correction method according to claim 1, characterized in that, The method further includes: Obtain the characters in the speech recognition dictionary that are the same as those in the large language model dictionary of the preset large language model, and get the first dictionary characters; Obtain dictionary characters that exist in the speech recognition dictionary but do not appear in the large language model dictionary to obtain the second dictionary characters; Each character in the second dictionary character set is encoded using a byte-level byte-pair encoding method to obtain the byte-pair encoding sequence corresponding to each second dictionary character, and a dictionary character mapping table is constructed; wherein, the dictionary character mapping table uses the second dictionary character as the key and the byte-pair encoding sequence corresponding to the second dictionary character as the value; Based on the first dictionary character and the dictionary character mapping table, the target dictionary character is obtained.

7. The speech recognition error correction method according to claim 6, characterized in that, The initialization process, based on the target dictionary characters, is performed on the embedding and output layers of the initial large language model to obtain the initialized large language model, including: Copy the embedding vector and output layer vector corresponding to the first dictionary character in the target dictionary characters into the speech recognition dictionary initialization parameters; For each second dictionary character in the dictionary character mapping table, the embedding vector and output layer vector of each word in the large language model dictionary in the byte pair encoding sequence corresponding to the second dictionary character are averaged to obtain an averaged vector, and the averaged vector is copied as the initialization vector of the second dictionary character into the speech recognition dictionary initialization parameters. Based on the initialization parameters of the speech recognition dictionary, the embedding layer and output layer in the initial large language model are initialized to obtain the initialized large language model.

8. A speech recognition error correction device, characterized in that, include: The recognition module is used to input the current audio frame in the user's audio to be recognized into the speech recognition model to obtain the acoustic features and the first text character of the current audio frame; Based on the target dictionary characters, the embedding layer and output layer in the initial large language model are initialized to obtain the initialized large language model; wherein, the target dictionary characters are determined based on the speech recognition dictionary and the large language model dictionary of the preset large language model; The acoustic feature sample and text character sequence sample corresponding to the speech sample signal are used as one training sample to obtain multiple training samples; The speech sample signal is input into the speech recognition model to obtain the probability distribution of outputting any text character or blank mark for each frame of audio sample of the speech sample signal; Based on the probability distribution, an alignment path search space is constructed. The horizontal dimension of the alignment path search space corresponds to the time frame sequence of the speech sample signal, and the vertical dimension corresponds to the character sequence of the real text label corresponding to the training sample. In the alignment path search space, a dynamic programming algorithm is used to find the alignment path with the highest probability from the start position to the end position, wherein the start position corresponds to the zeroth frame and the zeroth character, and the end position corresponds to the last frame and the last character. Based on the alignment path with the highest probability, a forced alignment text label sequence is generated, wherein each frame position in the forced alignment text label sequence corresponds to a real text character or a blank mark; The large language model after initialization is trained based on each training sample carrying a forced alignment text label to obtain a speech recognition error correction model. The error correction module is used to input the acoustic features of the current audio frame, the first text character, and the historical corrected text sequence into the speech recognition error correction model to obtain the corrected text character corresponding to the current audio frame. The historical corrected text sequence is a corrected text sequence obtained by sequentially passing the speech recognition model and the speech recognition error correction model through the historical audio frames in the user audio to be recognized before the current audio frame; the speech recognition error correction model is trained based on training samples carrying forced alignment text labels, and the forced alignment text labels are obtained by forcibly aligning the real text labels corresponding to the training samples with the speech sample signals corresponding to the training samples.

9. An electronic device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that, When the processor executes the computer program, it implements the speech recognition error correction method as described in any one of claims 1-7.

10. A non-transitory computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by the processor, it implements the speech recognition error correction method as described in any one of claims 1-7.

Citation Information

Patent Citations

  • Speech recognition error correction method and device

    CN106486126A

  • Training method of voice transfer text error correction model and computer equipment

    CN115293139A