Speech recognition pre-training model fine tuning method and system, and speech recognition method and system
By adding a low-rank adapter to the speech recognition model and using the ORPO algorithm, the problem of poor fine-tuning of the vertical field model is solved, and a more efficient speech recognition effect is achieved.
Patent Information
- Application Number
- CN202510596257.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-09
- Publication Date
- 2025-07-29
AI Technical Summary
The prior art speech recognition model fine-tuning in the vertical field is poor and lacks the applicability of professional knowledge.
By obtaining audio samples and their annotated samples, building training data, adding low-rank adapters to the pretrained model, using the Preference Alignment Optimization Training Algorithm (ORPO), the model is guided to learn the Preference Sample and avoiding the generation of non-preference samples.
The effect of the speech recognition model has been significantly improved, with the improvement range from 3.7% to 7.8%.
Smart Images

Figure CN120388566A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of speech recognition, and in particular, to a method and system for fine-tuning a pre-trained speech recognition model, and also to a speech recognition method and system. Background Art
[0002] As the core perception module of a human-computer interaction system, the recognition accuracy of automatic speech recognition (ASR) directly affects the reliability of downstream natural language understanding and dialogue generation.
[0003] Currently, the general models obtained by training on large-scale public datasets perform well in logical reasoning and language generation, but lack professional knowledge in vertical fields. In order to make these generative large models truly applicable to vertical fields, domain fine-tuning is usually required. Summary of the Invention
[0004] Aiming at the disadvantage of the poor effect of the fine-tuning scheme disclosed in the prior art, the present invention provides a method and system for fine-tuning a pre-trained speech recognition model, and a speech recognition method and system.
[0005] To solve the above technical problems, the present invention is solved by the following technical solutions:
[0006] In a first aspect, a method for fine-tuning a pre-trained speech recognition model is provided, including the following steps:
[0007] Obtain an audio sample and its annotation sample, and use the annotation sample as the corresponding preference sample;
[0008] Input the audio sample into the pre-trained model, perform speech recognition by the pre-trained model to obtain the corresponding first recognition sample, and use the first recognition sample as the corresponding non-preference sample;
[0009] Construct training data, each piece of training data including an audio sample, and the preference sample and non-preference sample corresponding to the audio sample;
[0010] Add a low-rank adapter to the pre-trained model to construct a target fine-tuning model;
[0011] Perform preference alignment optimization training on the target fine-tuning model based on the training data;
[0012] Determine a target speech recognition model based on the target fine-tuning model with optimization completed.
[0013] As an implementable manner, perform iterative training on the target fine-tuning model based on the training data, and the specific steps of each iteration are:
[0014] For each piece of training data, the first fine-tuning model generates corresponding positive sample recognition probabilities and negative sample recognition probabilities, where the positive sample recognition probability is the probability that the speech recognition result of the audio sample is the corresponding preferred sample, and the negative sample recognition probability is the probability that the speech recognition result of the audio sample is the corresponding non-preferred sample;
[0015] Calculate the negative log-likelihood loss based on the positive sample recognition probability to obtain the corresponding supervised fine-tuning loss;
[0016] Calculate the ratio of generating the preferred sample based on the positive sample recognition probability to obtain the first ratio;
[0017] Calculate the ratio of generating the non-preferred sample based on the negative sample recognition probability to obtain the second ratio;
[0018] Generate the corresponding relative ratio loss based on the first ratio and the second ratio;
[0019] Generate the target loss based on the supervised fine-tuning loss and the relative ratio loss corresponding to each piece of training data;
[0020] Perform LoRA fine-tuning on the model parameters of the first fine-tuning model based on the target loss to obtain the corresponding second fine-tuning model, and use the second fine-tuning model as the first fine-tuning model for the next iteration step.
[0021] Furthermore:
[0022] In each iteration step, the first fine-tuning model performs speech recognition on each audio sample to obtain the corresponding second recognition sample, and updates the corresponding non-preferred sample based on the obtained second recognition sample to obtain the updated training data for the next iteration step.
[0023] As an implementable manner:
[0024] The pre-trained model includes an encoder and a decoder;
[0025] Add an intermediate processing module to the pre-trained model, and add low-rank matrices to the encoder and the decoder to construct a target fine-tuning model. The intermediate processing module is used to copy the transcribed data output by the encoder to generate first transcribed data and second transcribed data; the first transcribed data is paired with the corresponding preferred sample, and the second transcribed data is paired with the corresponding non-preferred sample.
[0026] Furthermore:
[0027] Construct audio input data based on the audio samples corresponding to all training data in the current iteration step, and construct preference learning data based on the preferred samples and non-preferred samples corresponding to all training data;
[0028] Input the audio input data into an encoder, and the encoder outputs corresponding encoded data, where the encoded data includes transcription data corresponding one-to-one to the audio samples;
[0029] Input the encoded data into an intermediate processing module, which performs copying and sorting, and outputs encoded input data corresponding to the preference learning data;
[0030] Input the encoded input data and the preference learning data into a decoder, guiding the decoder to learn the preference samples in the preference learning data and penalize the non-preference samples.
[0031] Furthermore:
[0032] When a preset iteration termination condition is reached, determine a target speech recognition model based on the encoder and decoder in the obtained second fine-tuned model.
[0033] In a second aspect, the present invention proposes a system for fine-tuning a speech recognition pre-trained model, including:
[0034] A preparation module:
[0035] For obtaining audio samples and their annotation samples, and using the annotation samples as corresponding preference samples;
[0036] For inputting the audio samples into a pre-trained model, performing speech recognition by the pre-trained model to obtain corresponding first recognition samples, and using the first recognition samples as corresponding non-preference samples;
[0037] For constructing training data, each piece of training data includes an audio sample, and a preference sample and a non-preference sample corresponding to the audio sample;
[0038] A first construction module, for adding a low-rank adapter to the pre-trained model to construct a target fine-tuned model;
[0039] A training module, for performing preference alignment optimization training on the target fine-tuned model based on the training data to obtain an optimized target fine-tuned model;
[0040] A second construction module, for determining a target speech recognition model based on the optimized target fine-tuned model.
[0041] In a third aspect, the present invention proposes a speech recognition method, including the following steps:
[0042] Obtain an audio to be recognized;
[0043] Input the audio to be recognized into the speech recognition model obtained by the method for fine-tuning a speech recognition pre-trained model described in any one of the above, and the speech recognition model outputs corresponding speech recognition results.
[0044] Fourthly, the present invention proposes a speech recognition system, including the following steps:
[0045] An acquisition module, configured to acquire the audio to be recognized;
[0046] A recognition module, configured to input the audio to be recognized into the speech recognition model obtained by the method of fine-tuning any of the above-mentioned speech recognition pre-training models, and output a corresponding speech recognition result by the speech recognition model.
[0047] Fifthly, the present invention proposes a terminal device, including a memory, a processor, and a computer program stored in the memory and executable on the processor. When the processor executes the computer program, the method of fine-tuning any of the above-mentioned speech recognition pre-training models is implemented; or, when the processor executes the computer program, the above-mentioned speech recognition method is implemented.
[0048] The method for fine-tuning a speech recognition pre-training model proposed by the present invention overcomes the technical prejudice that the existing preference optimization algorithm can only be used for large prediction models, and creatively proposes to apply the preference optimization algorithm to the fine-tuning of the speech recognition pre-training model. The manually annotated annotation samples are used as preference samples, and the recognition results output by the pre-training model are used as non-preference samples to guide the model to learn the manually annotated preference samples and avoid the model from generating unsatisfactory recognition results, effectively improving the model effect. Description of the Drawings
[0049] In order to more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the following will briefly introduce the drawings required for the description of the embodiments or the prior art. Obviously, the following drawings are only some embodiments of the present invention. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.
[0050] Figure 1 It is a schematic flowchart of the method for fine-tuning a speech recognition pre-training model of the present invention;
[0051] Figure 2 It is a schematic flowchart of generating and updating training data;
[0052] Figure 3 It is a schematic diagram of data flow when directly fine-tuning the pre-training model;
[0053] Figure 4 It is a schematic diagram of data flow when fine-tuning the target fine-tuning model including an intermediate processing module. Detailed Embodiments
[0054] The present invention will be further described in detail below in conjunction with embodiments. The following embodiments are explanations of the present invention, and the present invention is not limited to the following embodiments.
[0055] An embodiment of the present application provides a method for fine-tuning a speech recognition pre-trained model. Refer to Figure 1 , including the following steps:
[0056] S100. Data preparation;
[0057] Construct training data, and each piece of training data includes an audio sample, a preference sample, and a non-preference sample;
[0058] Refer to Figure 2 , and the steps for constructing training data are as follows:
[0059] S110. Obtain an audio sample and its annotation sample, and use the annotation sample as the corresponding preference sample;
[0060] S120. Input the audio sample into the pre-trained model, perform speech recognition by the pre-trained model to obtain a corresponding first recognition sample, and use the first recognition sample as the corresponding non-preference sample;
[0061] S130. Construct training data, and each piece of training data includes an audio sample, and a preference sample and a non-preference sample corresponding to the audio sample.
[0062] S200. Construct a target fine-tuning model;
[0063] In this embodiment, the pre-trained model is a pre-trained general speech recognition model, and the pre-trained model includes an encoder and a decoder.
[0064] Those skilled in the art can select an existing publicly available network architecture with an encoder and a decoder for speech recognition according to actual needs. For example, an open-source end-to-end speech recognition and multilingual translation system based on Transformer can be used;
[0065] In this embodiment, the encoder is responsible for processing the input speech signal, usually using Mel-spectrogram as the input, and extracting features in the audio signal through multiple Transformer layers to obtain an encoded representation representing the context information of the audio sequence;
[0066] The decoder is responsible for generating the target text (such as the transcribed text or the translated sentence). The decoder obtains information from the context representation output by the encoder and generates a step-by-step output based on this information.
[0067] Those skilled in the art can pre-train the selected network model using a general dataset according to actual needs to obtain a corresponding pre-trained model, and this specification will not elaborate on the pre-training steps.
[0068] S210. Add a low-rank adapter to the pre-trained model to construct a target fine-tuning model;
[0069] In this embodiment, the LoRA technology is used for model fine-tuning;
[0070] Those skilled in the art can determine the layers for fine-tuning and updating network parameters according to actual needs. For example, select the attention mechanism layer and the feedforward network layer in the encoder and / or decoder;
[0071] The feedforward network layer in Transformer contains a large number of parameters. By applying LoRA to these layers, the computational cost and memory consumption during model fine-tuning can be effectively reduced while maintaining model performance;
[0072] Specifically:
[0073] Add a first low-rank matrix A and a second low-rank matrix B to the weight matrix of the pre-trained model. During the iterative training process, keep the weights of the pre-trained model fixed and only update the first low-rank matrix and the second low-rank matrix based on the obtained target loss.
[0074] In this embodiment, through the use of LoRA fine-tuning, a low-rank matrix is added to the pre-trained model, and it is not necessary to update all parameters when fine-tuning the pre-trained model, thereby reducing the computational and storage costs.
[0075] S300. Perform preference alignment optimization training on the target fine-tuning model based on the training data;
[0076] In this embodiment, the ORPO (Odds Ratio Preference Optimization algorithm) is adopted. Based on ORPO, the training data is used to perform iterative training on the target fine-tuning model. The specific steps for each iteration are as follows:
[0077] S310. For each piece of training data, generate corresponding positive sample recognition probabilities and negative sample recognition probabilities by the first fine-tuning model;
[0078] The first fine-tuning model is the target fine-tuning model corresponding to the current iteration step;
[0079] The positive sample recognition probability is the probability that the speech recognition result of the audio sample is the corresponding preference sample. In this embodiment, the positive sample recognition probability is denoted as P θ (y ω ∣x), that is, the recognition result of the audio sample x corresponds to the preference sample yω Probability;
[0080] The negative sample recognition probability is the probability that the result of speech recognition of the audio sample is the corresponding non-preferred sample. In this embodiment, the negative sample recognition probability is denoted as P θ (y l ∣x), that is, the probability that the recognition result corresponding to the audio sample x is the non-preferred sample y l Probability.
[0081] S320. Calculate the negative log-likelihood loss based on the positive sample recognition probability to obtain the corresponding supervised fine-tuning loss L SFT .
[0082] S330. Calculate the corresponding relative ratio loss based on the positive sample recognition probability and the negative sample recognition probability;
[0083] Specifically:
[0084] S331. Calculate the ratio of generating the preferred sample based on the positive sample recognition probability to obtain the first ratio, and the first ratio is denoted as odd θ (y ω ∣x);
[0085] In this embodiment, the calculation formula of the ratio is:
[0086]
[0087] Where:
[0088] odd θ (y∣x) represents the ratio that the input data x (audio sample) corresponds to the recognition result as the output data y (preferred sample or non-preferred sample);
[0089] P θ (y∣x) represents the probability that the input data x (audio sample) corresponds to the recognition result as the output data y (preferred sample or non-preferred sample);
[0090] θ is the model parameter corresponding to the target fine-tuning model;
[0091] S332. Calculate the ratio of generating the non-preferred sample based on the negative sample recognition probability to obtain the second ratio, and the first ratio is denoted as odd θ (y l ∣x);
[0092] S333. Generate the corresponding relative ratio loss based on the first ratio and the second ratio;
[0093] The relative ratio loss L OR The calculation formula is:
[0094]
[0095] where σ represents the sigmoid function.
[0096] S340. Generate a target loss based on the supervised fine-tuning loss and the relative ratio loss corresponding to each training data;
[0097] The target loss L ORPO is calculated by the formula:
[0098]
[0099] where λ is a weight, and those skilled in the art can set the value of λ by themselves.
[0100] S350. Perform LoRA fine-tuning on the model parameters of the first fine-tuning model based on the target loss to obtain a corresponding second fine-tuning model;
[0101] When the preset iteration completion condition is reached, output the second fine-tuning model as the target fine-tuning model with optimization completed. When the preset iteration completion condition is not reached, use the second fine-tuning model as the first fine-tuning model for the next iteration step.
[0102] As an implementable manner, in each iteration step, the first fine-tuning model performs speech recognition on each audio sample to obtain a corresponding second recognition sample, and updates the corresponding non-preferred sample based on the obtained second recognition sample to obtain updated training data for the next iteration step.
[0103] Refer to Figure 2 , new recognition results will be generated during the iterative training process. In this embodiment, the new recognition results replace the non-preferred samples, and the audio samples, annotation samples, and the corresponding new recognition results are assembled into a new piece of training data for the next model training;
[0104] With this design, the model can dynamically challenge the training data during iterative training, self-correct the non-preferred data and gradually optimize it as the model is trained, further improving the training efficiency.
[0105] S400. Determine a target speech recognition model based on the target fine-tuning model with optimization completed.
[0106] The ORPO algorithm was invented in the scenario of combining large language models with reinforcement learning. For example, in the question-and-answer scenario, the answers generated by the large prediction model are open-ended. They may be good answers that fit the user's intention or bad answers that may contain illegal content. To address this defect, preferred answers and non-preferred answers (answers containing illegal content) are prepared for the same question, guiding the large language model to learn the preferred answers while avoiding generating non-preferred answers containing illegal content;
[0107] However, for the speech recognition scenario, due to the clear correspondence between audio and transcribed text, the common training method is supervised fine-tuning to guide the model to learn the correct answers marked by humans;
[0108] In the actual training process, the model still outputs unsatisfactory recognition results, and different from the illegal content to be avoided in the large speech model, the unsatisfactory speech recognition results are diverse and unpredictable. To address this problem, those skilled in the art often add a scoring model to score the quality of the recognition results output by the model through a scoring mechanism and optimize the model based on the corresponding evaluation scores.
[0109] The characteristic of the clear and unique recognition result in the speech recognition scenario makes those skilled in the art focus on how to guide the model to learn the transcribed samples with annotations during model training, and it is difficult to think of applying the preference optimization technology to avoid generating illegal content in the large speech training scenario to the speech recognition scenario;
[0110] This application uses the marked annotation samples as preferred samples and the recognition results output by the model as non-preferred samples. During the model training process, it not only guides the model to learn the accurate recognition results manually annotated but also enables the model to perceive unsatisfactory recognition results through the design of non-preferred samples, guiding the model to avoid generating such results. In the actual experimental results, it is found that after using the ORPO strategy, compared with only using the supervised fine-tuning strategy, the improvement rate of the model effect has increased from 3.7% to 7.8%.
[0111] Furthermore, step S200 for constructing the target fine-tuning model further includes the step of adding an intermediate processing module to the pre-trained model;
[0112] That is, add an intermediate processing module to the pre-trained model and add a low-rank matrix to the pre-trained model to construct the target fine-tuning model.
[0113] At this time, the target fine-tuning model includes an encoder, an intermediate processing module, and a decoder.
[0114] For each piece of training data:
[0115] The encoder is used to transcribe the input audio samples and output the corresponding transcribed data.
[0116] The intermediate processing module is used to copy the transcription data output by the encoder to generate first transcription data and second transcription data;
[0117] The first transcription data is paired with the corresponding preference sample to form a preference sample pair;
[0118] The second transcription data is paired with the corresponding non-preference sample to form a non-preference sample pair.
[0119] The decoder outputs the probability of generating the preference sample and the probability of generating the non-preference sample based on the preference sample pair and the non-preference sample pair;
[0120] The decoder is also used to generate a second recognition sample corresponding to the first transcription data or the second transcription data for updating the non-preference sample.
[0121] When a preset iteration termination condition is reached, a target speech recognition model is determined based on the encoder and decoder in the obtained second fine-tuning model. That is, after fine-tuning, the intermediate processing module is removed, and the trained encoder and decoder are output as the target speech recognition model.
[0122] Those skilled in the art can set the iteration termination condition according to actual needs, such as presetting the maximum number of iterations.
[0123] When directly fine-tuning the pre-trained model based on the training data, in order to make the transcription data of the decoder match the corresponding preference samples and non-preference samples, the training data needs to be constructed as positive sample training data and negative sample training data, where the positive sample training data is the audio sample and its preference sample, and the negative sample training data is the audio sample and its non-preference sample;
[0124] During the training process, the audio sample in the positive sample training data or the negative sample training data is input into the encoder, and the preference sample or the non-preference sample is input into the decoder. The decoder outputs the probability that the recognition result is the corresponding preference sample or non-preference sample based on the transcription data output by the encoder;
[0125] Taking one piece of training data as an example, referring to Figure 3 , this training data includes the audio sample audio_0, the preference sample response_0_w, and the non-preference sample response_0_l. The input of the encoder is (audio_0, audio_0), and the output is (audio_0_hidden, audio_0_hidden), so that the data output by the encoder can be matched with the preference sample response_0_w and the non-preference sample response_0_l one by one. In this solution, the same audio sample audio_0 will be repeatedly encoded in the encoder.
[0126] In this embodiment, to address the problem of duplicate encoding, the model architecture of the pre-trained model to be fine-tuned is optimized. That is, an intermediate processing model is added between the encoder and the decoder. During the training process, the intermediate processing model copies and sorts the transcription data output by the encoder, eliminating the need for the encoder to perform duplicate encoding on the same audio sample. In this embodiment, the computational cost is reduced by 50%, significantly reducing the computational overhead of model training.
[0127] Taking two training data as an example, the first training data includes audio sample audio_0, preferred sample response_0_w, and non-preferred sample response_0_l, and the second training data includes audio sample audio_1, preferred sample response_1_w, and non-preferred sample response_1_l.
[0128] Refer to Figure 3 , the input of the encoder is (audio_0, audio_1), and the output is (audio_0_hidden, audio_1_hidden). The data output by the encoder is copied and sorted in the intermediate processor, and the intermediate processing module outputs (audio_0_hidden, audio_0_hidden, audio_1_hidden, audio_1_hidden). The transcription data output by the intermediate processing module is matched one by one with the corresponding preferred sample / non-preferred sample. In this solution, the encoder only encodes the audio sample of each training data once.
[0129] As an implementable manner, model training is performed based on batch processing.
[0130] Construct audio input data based on the audio samples corresponding to all training data in the current iteration step, and construct preference learning data from the preferred samples and non-preferred samples corresponding to all training data.
[0131] Input the audio input data into the encoder, and the encoder outputs corresponding encoded data, where the encoded data includes transcription data corresponding one by one to the audio samples.
[0132] Input the encoded data into the intermediate processing module, and the intermediate processing module copies and sorts it to output encoded input data corresponding to the preference learning data.
[0133] Input the encoded input data and the preference learning data into the decoder, and the decoder outputs the probabilities corresponding to each preferred sample or non-preferred sample in the preference learning data, guiding the decoder to learn the preferred samples in the preference learning data and penalize the non-preferred samples.
[0134] The embodiment of the present application also provides a system for fine-tuning a speech recognition pre-training model, including:
[0135] A preparation module:
[0136] Used to obtain audio samples and their annotation samples, and use the annotation samples as corresponding preference samples;
[0137] Used to input the audio samples into the pre-training model, perform speech recognition by the pre-training model to obtain corresponding first recognition samples, and use the first recognition samples as corresponding non-preference samples;
[0138] Used to construct training data, and each piece of training data includes an audio sample, and a preference sample and a non-preference sample corresponding to the audio sample.
[0139] A first construction module, used to add a low-rank adapter to the pre-training model to construct a target fine-tuning model;
[0140] As an implementable manner, the first construction module is further used to add an intermediate processing model to the pre-training model to construct a target fine-tuning model.
[0141] A training module, used to perform preference alignment optimization training on the target fine-tuning model based on the training data to obtain an optimized target fine-tuning model.
[0142] A second construction module, used to determine a target speech recognition model based on the optimized target fine-tuning model.
[0143] The embodiment of the present application also provides a speech recognition method, including the following steps:
[0144] Obtain the audio to be recognized;
[0145] Input the audio to be recognized into the speech recognition model obtained by the method for fine-tuning the speech recognition pre-training model described above, and the speech recognition model outputs corresponding speech recognition results.
[0146] The speech recognition model includes an encoder and a decoder. The encoder is used to generate transcription data based on the audio to be recognized, and the decoder is used to generate corresponding speech recognition results based on the patent data.
[0147] The embodiment of the present application also provides a speech recognition system, including:
[0148] An acquisition module, used to acquire the audio to be recognized;
[0149] A recognition module, used to input the audio to be recognized into the speech recognition model obtained by the method for fine-tuning the speech recognition pre-training model described above, and the speech recognition model outputs corresponding speech recognition results.
[0150] An embodiment of the present application further provides a terminal device, including a memory, a processor, and a computer program stored in the memory and executable on the processor;
[0151] When the processor executes the computer program, the method for fine-tuning the above-mentioned speech recognition pre-training model is implemented; or, when the processor executes the computer program, the above-mentioned speech recognition method is implemented.
[0152] For the device embodiment, since it is basically similar to the method embodiment, the description is relatively simple. For the relevant parts, please refer to the partial description of the method embodiment.
[0153] Each embodiment in this specification is described in a progressive manner. The key point of each embodiment is to illustrate the differences from other embodiments. The same or similar parts among the embodiments can be referred to each other.
[0154] Those skilled in the art should understand that the embodiments of the present invention can be provided as a method, a device, or a computer program product. Therefore, the present invention can be implemented in the form of a complete hardware embodiment, a complete software embodiment, or an embodiment combining software and hardware aspects. Moreover, the present invention can be implemented in the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk memories, CD-ROMs, optical memories, etc.) containing computer-usable program codes.
[0155] The present invention is described with reference to the flowcharts and / or block diagrams of methods, terminal devices (systems), and computer program products according to the present invention. It should be understood that each process and / or block in the flowchart and / or block diagram can be implemented by computer program instructions, and the combination of processes and / or blocks in the flowchart and / or block diagram can also be implemented by computer program instructions. These computer program instructions can be provided to the processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing terminal devices to generate a machine, so that the instructions executed by the processor of the computer or other programmable data processing terminal devices generate a device for implementing the functions specified in Figure 1 one process or multiple processes and / or blocks Figure 1 one block or multiple blocks.
[0156] These computer program instructions can also be stored in a computer-readable memory that can direct a computer or other programmable data processing terminal device to work in a specific manner, so that the instructions stored in the computer-readable memory generate a manufactured product including an instruction device, and the instruction device implements the functions specified in Figure 1 one process or multiple processes and / or blocks Figure 1 one block or multiple blocks.
[0157] These computer program instructions can also be loaded onto a computer or other programmable data processing terminal device, so that a series of operation steps are executed on the computer or other programmable terminal device to generate a computer-implemented process. Thus, the instructions executed on the computer or other programmable terminal device provide steps for implementing the functions specified in one process or multiple processes and / or blocks. Figure 1 One process or multiple processes and / or blocks Figure 1 Steps for implementing the functions specified in one block or multiple blocks.
[0158] It should be noted that:
[0159] As used in the specification, the phrase "one embodiment" or "an embodiment" means that a particular feature, structure, or characteristic described in connection with the embodiment is included in at least one embodiment of the present invention. Therefore, the phrases "one embodiment" or "an embodiment" that appear throughout the specification do not necessarily all refer to the same embodiment.
[0160] Although the preferred embodiments of the present invention have been described, those skilled in the art can make additional changes and modifications once they learn the basic creative concept. Therefore, the appended claims are intended to be construed to include the preferred embodiments and all changes and modifications that fall within the scope of the present invention.
[0161] In addition, it should be noted that the specific embodiments described in this specification may have different shapes, names, etc. for their components. Any equivalent or simple changes made to the structure, features, and principles described according to the inventive concept of the present invention are included in the protection scope of the present invention. Those skilled in the art to which the present invention pertains can make various modifications, supplements, or use similar means of substitution to the specific embodiments described, as long as they do not deviate from the structure of the present invention or exceed the scope defined by the claims, and should fall within the protection scope of the present invention.
Claims
1. A method for fine-tuning a pre-trained model for speech recognition, characterized in that, The following steps are involved: Obtaining an audio sample and its annotated sample, and using the annotated sample as a corresponding preferred sample; Inputting the audio sample into a pre-trained model, performing speech recognition by the pre-trained model to obtain a corresponding first recognition sample, and using the first recognition sample as a corresponding non-preferred sample; Constructing training data, each piece of training data includes an audio sample, and a preferred sample and a non-preferred sample corresponding to the audio sample; Add a low-rank adapter to the pre-trained model to build the target fine-tuned model; Performing preference alignment optimization training on the target fine-tuning model based on the training data; The target speech recognition model is determined based on the optimized target fine-tuning model.
2. The method for fine-tuning a pre-trained speech recognition model according to claim 1, wherein The target fine-tuning model is iteratively trained based on the training data. The specific steps of each iteration are: For each piece of training data, the first fine-tuning model generates corresponding positive sample recognition probabilities and negative sample recognition probabilities, where the positive sample recognition probability is the probability that the speech recognition result of the audio sample is the corresponding preferred sample, and the negative sample recognition probability is the probability that the speech recognition result of the audio sample is the corresponding non-preferred sample; Calculate the negative log-likelihood loss based on the positive sample recognition probability to obtain the corresponding supervised fine-tuning loss; Calculating the ratio of the preferred samples based on the positive sample recognition probability to obtain a first ratio; Calculating the ratio of the non-preferred samples based on the negative sample recognition probability to obtain a second ratio; generating a corresponding relative ratio loss based on the first ratio and the second ratio; Generate target loss based on the supervised fine-tuning loss and relative ratio loss corresponding to each training data; Based on the target loss, LoRA fine-tuning is performed on the model parameters of the first fine-tuning model to obtain a corresponding second fine-tuning model, and the second fine-tuning model is used as the first fine-tuning model of the next iteration step.
3. The method for fine-tuning a speech recognition pre-training model according to claim 2, characterized in that: In each iteration step, the first fine-tuning model performs speech recognition on each audio sample to obtain a corresponding second recognition sample, and updates the corresponding non-preferred sample based on the obtained second recognition sample to obtain updated training data for use in the next iteration step.
4. The method for fine-tuning a speech recognition pre-training model according to any one of claims 1 to 3, characterized in that: The pre-trained model includes an encoder and a decoder; An intermediate processing module is added to the pre-trained model, and a low-rank matrix is added to the encoder and decoder to construct a target fine-tuning model. The intermediate processing module is used to copy the transcription data output by the encoder to generate first transcription data and second transcription data; the first transcription data is paired with the corresponding preferred sample, and the second transcription data is paired with the corresponding non-preferred sample.
5. The method for fine-tuning a speech recognition pre-training model according to claim 4, characterized in that: Construct audio input data based on the audio samples corresponding to all training data in the current iteration step, and construct preference learning data based on the preference samples and non-preference samples corresponding to all training data; Inputting the audio input data into an encoder, and having the encoder output corresponding encoded data, the encoded data including transcription data corresponding one-to-one to the audio samples; Inputting the encoded data into an intermediate processing module, which copies and sorts the encoded data and outputs the encoded input data corresponding to the preference learning data; The encoded input data and the preference learning data are input into a decoder, guiding the decoder to learn the preferred samples in the preference learning data and punish the non-preferred samples.
6. The method for fine-tuning a speech recognition pre-training model according to claim 4, characterized in that: When a preset iteration termination condition is reached, a target speech recognition model is determined based on the encoder and decoder in the obtained second fine-tuning model.
7. A system for fine-tuning a pre-trained model for speech recognition, characterized in that, include: Prepare the module: Used to obtain audio samples and their annotated samples, and use the annotated samples as corresponding preferred samples; Inputting the audio sample into a pre-trained model, performing speech recognition by the pre-trained model, obtaining a corresponding first recognition sample, and using the first recognition sample as a corresponding non-preferred sample; Used to construct training data, each piece of training data includes an audio sample, and a preferred sample and a non-preferred sample corresponding to the audio sample; The first building block is used to add a low-rank adapter to the pre-trained model to build a target fine-tuning model; A training module, configured to perform preference alignment optimization training on the target fine-tuning model based on the training data to obtain an optimized target fine-tuning model; The second building module is used to determine the target speech recognition model based on the optimized target fine-tuning model.
8. A speech recognition method, characterized in that: The following steps are involved: Get the audio to be recognized; The audio to be recognized is input into a speech recognition model obtained by the method for fine-tuning a speech recognition pre-training model according to any one of claims 1 to 6, and the speech recognition model outputs a corresponding speech recognition result.
9. A voice recognition system, characterized in that, The following steps are involved: An acquisition module, used to acquire the audio to be recognized; A recognition module is used to input the audio to be recognized into a speech recognition model obtained by the method for fine-tuning the speech recognition pre-training model according to any one of claims 1 to 6, and the speech recognition model outputs a corresponding speech recognition result.
10. A terminal device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein: When the processor executes the computer program, it implements the method for fine-tuning a speech recognition pre-training model as described in any one of claims 1 to 6; or, when the processor executes the computer program, it implements the speech recognition method as described in claim 8.