Speech recognition model optimization method and device based on reinforcement learning, equipment and medium

Through the speech recognition model optimization method based on reinforcement learning, the problem of insufficient recognition accuracy and generalization capabilities of the ASR system in complex environments is solved, and higher recognition accuracy and cross-environment recognition capabilities are achieved, and the stability of policy updates is ensured.

CN120071933APending Publication Date: 2025-05-30SHANDONG INSPUR SCI RES INST CO LTD
View PDF 0 Cites 8 Cited by

Patent Information

Application Number
CN202510346896.7
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-24
Publication Date
2025-05-30

AI Technical Summary

Technical Problem

The existing automatic speech recognition (ASR) system has problems such as data distribution deviation, loss function optimization limitations and insufficient generalization capabilities in complex environments, resulting in poor recognition accuracy and generalization capabilities.

Method used

Using the speech recognition model optimization method based on reinforcement learning, the model learns different noise, accent and time frequency characteristics by randomly selecting the enhanced speech training data, reducing overfitting of specific scenarios and improving cross-environment recognition capabilities. At the same time, multiple ASR prediction texts are generated through the same input, and reward scale differences are eliminated through in-group normalization, training oscillations are reduced, strategy update stability is ensured, and target reward values ​​are calculated based on text matching degree and semantic similarity, and character-level matching and overall sentence semantics are optimized.

Benefits of technology

It improves the recognition accuracy and generalization ability of the automatic speech recognition model, reduces overfitting of specific scenarios, improves cross-environment recognition capabilities, and ensures the stability of strategy updates and the naturalness and practicality of the recognition results.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120071933A_ABST
    Figure CN120071933A_ABST
Patent Text Reader

Abstract

The invention discloses a voice recognition model optimization method and device based on reinforcement learning, equipment and a medium, and relates to the technical field of voice recognition, and the method comprises the steps: determining a current optimization strategy of a current time step; randomly selecting one piece of enhanced speech training data and corresponding real text training data from the data training set as target training data, and inputting the target training data into the current automatic speech recognition model for prediction to obtain a plurality of prediction texts output by the current automatic speech recognition model; determining a text matching degree and a semantic similarity between each prediction text and real text training data to obtain a group of target reward values, updating an optimization strategy at the current moment based on the target reward values, and skipping to execute the step of determining the current optimization strategy of the current time step to obtain a target reward value; and outputting the corresponding optimized automatic speech recognition model until the current automatic speech recognition model meets the optimization ending condition. And reasonable optimization of the automatic speech recognition model is realized.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of speech recognition, and particularly to a method, device, equipment and medium for optimizing a speech recognition model based on reinforcement learning. Background Art

[0002] Currently, the mainstream automatic speech recognition models mainly rely on supervised learning for training, such as end-to-end architectures based on neural networks such as Transformer and Conformer. However, the existing ASR systems have the following problems in complex environments:

[0003] (1) Data distribution bias: The training data does not match the real application data, resulting in a decline in the performance of ASR when there are large changes in accents, speech rates, and background noises.

[0004] (2) Limitations in loss function optimization: Traditional ASR training is usually based on cross-entropy or CTC (Connectionist Temporal Classification) loss. These methods mainly optimize character or word-level matching, but do not directly optimize the semantic accuracy of the overall sentence.

[0005] (3) Insufficient generalization ability: The model performs poorly on unseen speech data or in complex environments and is prone to recognition errors.

[0006] In summary, how to reasonably optimize the automatic speech recognition model and improve the recognition accuracy and generalization ability of the automatic speech recognition model is a technical problem to be solved in this field. Summary of the Invention

[0007] In view of this, the purpose of the present invention is to provide a method, device, equipment and medium for optimizing a speech recognition model based on reinforcement learning, which can reasonably optimize the automatic speech recognition model and improve the recognition accuracy and generalization ability of the automatic speech recognition model. The specific solutions are as follows:

[0008] In the first aspect, the present application discloses a method for optimizing a speech recognition model based on reinforcement learning, including:

[0009] Determine the current optimization strategy in the optimization action of the automatic speech recognition model at the current time step;

[0010] Randomly select any enhanced speech training data and the corresponding real text training data from the data training set as target training data, and input the target training data into the current automatic speech recognition model for prediction to obtain a number of predicted texts output by the current automatic speech recognition model; wherein, the current automatic speech recognition model is a model whose model parameters are updated based on the current optimization strategy;

[0011] Determine the text matching degree and semantic similarity between each of the predicted texts and the true text training data to obtain a set of target reward values, update the optimization strategy at the current moment based on the target reward values, and jump to execute the step of the current optimization strategy in the step of determining the optimization action of the automatic speech recognition model at the current time step until the current automatic speech recognition model meets the optimization end condition, and then output the corresponding optimized automatic speech recognition model.

[0012] Optionally, before randomly selecting any one of the augmented speech training data and the corresponding true text training data from the data training set as the target training data, it further includes:

[0013] Obtain the augmented speech training data after data augmentation processing; wherein, the augmented speech training data is the augmented speech data after background noise augmentation and speaker accent augmentation;

[0014] Input the augmented speech training data into the base model so that the base model transcribes the augmented speech training data to obtain the true text training data corresponding to the augmented speech training data; wherein, the base model is an audio language model used to process speech data into corresponding unique text data.

[0015] Construct a data training set based on each of the augmented speech training data and the corresponding unique true text training data.

[0016] Optionally, the obtaining of the augmented speech training data after data augmentation processing includes:

[0017] Perform content-level augmentation processing including background noise augmentation and speaker accent augmentation on the original speech data to obtain the first-process speech training data;

[0018] Perform feature-level augmentation processing including time stretching processing and frequency masking processing on the first-process speech training data to obtain the second-process speech training data;

[0019] Perform quality assessment processing including perceptual evaluation of speech quality, speech intelligibility evaluation in a noisy environment, and speech signal strength evaluation in a noisy environment on the second-process speech training data to screen out the augmented speech training data that meets the preset quality assessment conditions.

[0020] Optionally, the determining of the current optimization strategy in the optimization action of the speech recognition model at the current time step includes:

[0021] Perform reward value normalization processing on a set of target reward values at the previous time step to obtain relative reward values centered on the within-group reward mean and in units of the standard deviation;

[0022] Construct a policy update objective function by using the relative reward value, KL regularization term, the previous optimization policy at the previous time step, and the current optimization policy at the current time step;

[0023] Obtain the current optimization policy under the function maximum value of the policy update objective function.

[0024] Optionally, before inputting the target training data into the current automatic speech recognition model, it further includes:

[0025] Update the model parameters of the previous automatic speech recognition model at the previous time step through gradient ascent and based on the current optimization policy to obtain the current automatic speech recognition model at the current time step.

[0026] Optionally, determining the text matching degree and semantic similarity between each of the predicted texts and the true text training data to obtain a set of target reward values includes:

[0027] Calculate the character-level edit distance between each of the predicted texts and the true text training data to obtain a text matching degree score;

[0028] Calculate the semantic similarity between each of the predicted texts and the true text training data to obtain a semantic similarity score;

[0029] Obtain the target reward value through the text matching degree score, the semantic similarity score, and the corresponding weight coefficients;

[0030] Statistically calculate the target reward values between all the predicted texts and the true text training data to obtain a set of target reward values.

[0031] Optionally, the method for optimizing a speech recognition model based on reinforcement learning further includes:

[0032] Calculate the predicted text character error rate and predicted text semantic similarity score of the predicted text output by the current automatic speech recognition model to obtain corresponding evaluation results;

[0033] Adjust the weight coefficients in the calculation process of the target reward value according to the evaluation results, and dynamically adjust the hyperparameters of the reinforcement learning according to the model performance through an adaptive learning rate strategy.

[0034] In a second aspect, the present application discloses an apparatus for optimizing a speech recognition model based on reinforcement learning, including:

[0035] A policy determination module for determining the current optimization policy in the optimization action of the automatic speech recognition model at the current time step;

[0036] A prediction module, configured to randomly select any augmented speech training data and corresponding true text training data from a data training set as target training data, and input the target training data into the current automatic speech recognition model for prediction to obtain a plurality of predicted texts output by the current automatic speech recognition model; wherein, the current automatic speech recognition model is a model whose model parameters are updated based on the current optimization strategy.

[0037] A strategy adjustment module, configured to determine the text matching degree and semantic similarity between each of the predicted texts and the true text training data to obtain a set of target reward values, and update the optimization strategy at the current moment based on the target reward values, and jump to execute the step of the current optimization strategy in the step of determining the optimization action of the automatic speech recognition model at the current time step, until the current automatic speech recognition model meets the optimization end condition and outputs the corresponding optimized automatic speech recognition model.

[0038] In a third aspect, the present application discloses an electronic device, including:

[0039] A memory, configured to store a computer program;

[0040] A processor, configured to execute the computer program to implement the steps of the foregoing disclosed method for optimizing a speech recognition model based on reinforcement learning.

[0041] In a fourth aspect, the present application discloses a computer-readable storage medium, configured to store a computer program; wherein, when the computer program is executed by a processor, the steps of the foregoing disclosed method for optimizing a speech recognition model based on reinforcement learning are implemented.

[0042] It can be seen that the present application discloses an optimization method for a speech recognition model based on reinforcement learning, including: determining the current optimization strategy in the optimization action of the automatic speech recognition model at the current time step; randomly selecting any enhanced speech training data and the corresponding real text training data from the data training set as target training data, and inputting the target training data into the current automatic speech recognition model for prediction to obtain a number of predicted texts output by the current automatic speech recognition model; wherein, the current automatic speech recognition model is a model whose model parameters are updated based on the current optimization strategy; determining the text matching degree and semantic similarity between each predicted text and the real text training data to obtain a set of target reward values, and updating the optimization strategy at the current moment based on the target reward values, and jumping to execute the step of determining the current optimization strategy in the optimization action of the automatic speech recognition model at the current time step, until the current automatic speech recognition model meets the optimization end condition and outputs the corresponding optimized automatic speech recognition model. Thus, by randomly selecting enhanced speech training data, the model learns different noises, accents and time-frequency features, reduces overfitting to specific scenarios, and improves cross-environment recognition ability. By generating multiple ASR predicted texts from the same input, the reward scale difference is eliminated through within-group normalization, training oscillations are reduced, and the stability of policy updates is ensured. Moreover, in the process of calculating the target reward value, by combining the calculation of text matching degree and semantic similarity, not only the character-level matching is optimized, but also the overall semantics of the sentence is concerned, improving the naturalness and practicality of the recognition result. BRIEF DESCRIPTION OF THE DRAWINGS

[0043] In order to more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the following will briefly introduce the drawings required for use in the description of the embodiments or the prior art. Obviously, the drawings in the following description are only the embodiments of the present invention, and those of ordinary skill in the art can also obtain other drawings according to the provided drawings without creative efforts.

[0044] Figure 1 It is a flowchart of an optimization method for a speech recognition model based on reinforcement learning disclosed in the present application;

[0045] Figure 2 It is a schematic diagram of an ASR model optimization system based on reinforcement learning disclosed in the present application;

[0046] Figure 3 It is a schematic diagram of the structure of an optimization device for a speech recognition model based on reinforcement learning disclosed in the present application;

[0047] Figure 4 It is a structural diagram of an electronic device disclosed in the present application. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0048] Next, the technical solutions in the embodiments of the present application will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present application. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.

[0049] Current mainstream ASR models mainly rely on supervised learning for training, such as end-to-end architectures based on neural networks like Transformer and Conformer. However, the existing ASR systems have the following problems in complex environments:

[0050] (1) Data distribution bias: The training data does not match the real application data, resulting in a decline in the performance of ASR when there are significant changes in accent, speech rate, and background noise.

[0051] (2) Limitations in loss function optimization: Traditional ASR training usually relies on cross-entropy or CTC loss. These methods mainly optimize character or word-level matching and fail to directly optimize the semantic accuracy of the overall sentence.

[0052] (3) Insufficient generalization ability: The model performs poorly on unseen speech data or in complex environments and is prone to recognition errors.

[0053] To this end, the present invention provides an optimization scheme for a speech recognition model based on reinforcement learning, which can achieve reasonable optimization of the automatic speech recognition model and improve the recognition accuracy and generalization ability of the automatic speech recognition model.

[0054] Refer to Figure 1 As shown, an optimization method for a speech recognition model based on reinforcement learning is disclosed in an embodiment of the present invention, including:

[0055] Step S11: Determine the current optimization strategy in the optimization action of the automatic speech recognition model at the current time step.

[0056] In this embodiment, a group of target reward values at the previous time step are subjected to reward value normalization processing to obtain relative reward values centered on the within-group reward mean and with the standard deviation as the unit; a policy update objective function is constructed by using the relative reward values, the KL regularization term, the previous optimization policy at the previous time step, and the current optimization policy at the current time step; and the current optimization policy under the function maximum value of the policy update objective function is obtained. It can be understood that in order to reasonably optimize the automatic speech recognition model (ASR model), the GRPO (Group Relative Policy Optimization) method is used for training. It should be noted that compared with previous reinforcement learning (traditional Critic network), GRPO is based on within-group relative reward optimization. Specifically, for the enhanced speech training data q input in the previous time step, after obtaining multiple predicted texts, a group of target reward values are obtained through reward value calculation. Then relative reward calculation is performed. Specifically,

[0057] ;

[0058] wherein, represents the reward value of the th predicted text, represents the average value of a group of target reward values, represents the standard deviation of the target reward values within the same group, represents the relative reward value, that is, the relative advantage value, represents the total number of predicted texts in this group.

[0059] wherein, >0: the reward of the predicted text is higher than the within-group average level; <0: the reward of the predicted text is lower than the within-group average level; =0: the reward of the predicted text is equal to the within-group average level. In this way, positive encourages the policy to increase the probability of generating the corresponding predicted text, and negative promotes the policy to reduce the probability of generating the corresponding predicted text.

[0060] For example: Suppose a group of reward values is {3, 6, 9}, the mean mean = 6, and the standard deviation std = 3. Then the relative reward values are respectively:

[0061] , , ;

[0062] The conclusion is that the maximum relative advantage value is +1 (corresponding to a reward value of 9), but it is not the original maximum reward value, but the standardized result relative to other samples within the group.

[0063] Furthermore, a policy update objective function for reinforcement learning (GRPO optimization objective) is constructed. In this embodiment, the constructed policy update objective function is composed of the aforementioned relative reward value, KL regularization term, the optimized policy at the previous time step (previous optimized policy), and the optimized policy at the current time step (current optimized policy). The specific mathematical form is as follows:

[0064] ;

[0065] ;

[0066] Among them, represents the objective function value, represents the optimized policy parameter, represents the expected value of the enhanced speech training data q, represents the th predicted text, represents a set of predicted texts of the enhanced speech training data q of the previous optimized policy, represents the reference optimized policy, represents the current optimized policy of the th predicted text of the enhanced speech training data q, represents the previous optimized policy of the th predicted text of the enhanced speech training data q, represents the KL divergence regularization term, and are hyperparameters. Among them, is the coefficient of the KL divergence regularization term, specifically a hyperparameter that controls the amplitude of policy update.

[0067] In the policy update objective function, the key term of the function is the ratio term ( ), which is used to represent the probability ratio of the new policy to the old policy. The relative reward value serves as the direction signal for policy update, and the KL regularization term constrains the difference between the new policy and the reference policy to prevent performance collapse caused by excessive policy update amplitude.

[0068] For example: The relative advantage value of a certain set of predicted texts is the relative reward value , and the generation probability of the old policy is , and the generation probability of the new policy is , then the policy ratio term is , the policy gain term is 0.167. When = , then KL = 0.04 (the specific calculation is based on the probability distribution), and then the total target value = ; If = 0.1, then = 0.163, indicating that the policy update is effective.

[0069] During the process of optimizing the ASR model by reinforcement learning, the target reward value in the first iteration needs to be calculated based on multiple sets of prediction results generated by the initial policy, and the model parameters are adjusted accordingly. The following is the detailed process combined with the GRPO algorithm:

[0070] 1. Source of the initial policy: Before the first reinforcement learning optimization, the parameters of the ASR model are initialized by the supervised learning stage (such as training based on the CTC loss). At this time, the model already has basic speech recognition capabilities, but there are problems with data distribution deviation and limitations of the loss function.

[0071] 2. Process of the first reinforcement learning iteration:

[0072] Step 1: Generate multiple sets of prediction results.

[0073] Input: Randomly select a sample q (augmented speech + original transcription text) from the data-augmented training set.

[0074] Policy sampling: Use the initial policy (i.e., the model after supervised learning training) to generate G sets of predicted texts for the same speech q .

[0075] Sampling method: Increase diversity through random sampling (such as adjusting the output probability distribution with temperature parameters).

[0076] For example: Generate 10 different sets of predicted texts for the same speech, which may contain character errors or semantic biases.

[0077] Step 2: Calculate the reward value. Specifically, for each set of predicted texts, calculate the reward value for each group.

[0078] Step 3: Calculate the target reward value. Specifically, convert the reward value into a ranking relative to other samples within the group to reduce the impact of reward scale changes.

[0079] Step 4: Update the policy parameters. Based on the maximum value result of the function optimizing the GRPO objective, adjust the parameters. Specifically, maximize the optimization objective through gradient ascent, and at the same time use the KL divergence regularization term to constrain the amplitude of the policy update to prevent overfitting.

[0080] 3. Determination conditions for the relative reward value.

[0081] The calculation of the relative reward value is based on the following conditions:

[0082] Comparison within the same input group: Only relevant to other prediction results generated from the same speech q, rather than a global comparison.

[0083] Dynamic normalization: The mean and standard deviation of each group are calculated independently to adapt to the difference in the reward distribution of different inputs.

[0084] Hyperparameter adjustment: Control the weights of character accuracy and semantic similarity; group size Affects the stability of the relative reward, usually taking values from 10 to 100.

[0085] In this way, by maximizing the calculation of the above-mentioned constructed policy update objective function through gradient ascent, it shows that the policy is more inclined to generate prediction texts with high relative advantage values. By multiplying the policy ratio by the relative advantage value, it encourages the generation of prediction texts with high rewards. Through the KL divergence regularization term, the amplitude of the policy update is constrained to maintain stability. Finally, through gradient ascent optimization, the policy achieves a balance between relative advantage and policy stability. Therefore, the current optimized policy under the maximum value of the acquisition function is obtained.

[0086] Step S12: Randomly select any enhanced speech training data and the corresponding real text training data from the data training set as the target training data, and input the target training data into the current automatic speech recognition model for prediction to obtain a number of prediction texts output by the current automatic speech recognition model; wherein, the current automatic speech recognition model is a model whose model parameters are updated based on the current optimized policy.

[0087] In this embodiment, after determining the current optimized policy and before inputting the target training data into the current automatic speech recognition model, it further includes: updating the model parameters of the previous automatic speech recognition model at the previous time step through gradient ascent and based on the current optimized policy to obtain the current automatic speech recognition model at the current time step. It can be understood that the policy parameters are updated through gradient ascent based on the current optimized policy to further update the model parameters of the previous automatic speech recognition model at the previous time step and obtain the current automatic speech recognition model at the current time step.

[0088] In this embodiment, before randomly selecting any enhanced speech training data and the corresponding real text training data from the data training set as the target training data, the following steps are further included: obtaining the enhanced speech training data after data enhancement processing; wherein, the enhanced speech training data is enhanced speech data after background noise enhancement and speaker accent enhancement; inputting the enhanced speech training data into the base model so that the base model performs transcription processing on the enhanced speech training data to obtain the real text training data corresponding to the enhanced speech training data; wherein, the base model is an audio language model used to process speech data into corresponding unique text data; constructing a data training set based on each enhanced speech training data and the corresponding unique real text training data. It can be understood that first, the enhanced speech training data after data enhancement processing is obtained to simulate real environmental noise interference and accent variability. It should be noted that the speech processed through noise mixing, accent conversion, time-frequency masking, etc. has a large distribution difference from the original training data. Therefore, it needs to be processed once to obtain the predicted text. qwen2-audio-instruct is an end-to-end ASR model pre-trained on a large-scale speech data and has achieved a high recognition accuracy on the public data set. Therefore, further use a base model such as qwen2-audio-instruct to transcribe each enhanced speech sample once to obtain the unique predicted text. Moreover, the GRPO algorithm requires the initial policy to generate multiple sets of predicted texts as the optimization starting point. Therefore, the output of the base model is directly used as the initial policy, enabling the reinforcement learning optimization to start without having to start from scratch and reducing training oscillations.

[0089] Among them, by simulating different noise environments and accent changes, the generalization ability of the ASR model is improved, and data augmentation processing is performed on the data participating in model training. Specifically, content-level augmentation processing including background noise augmentation and speaker accent augmentation is performed on the original speech data to obtain first-process speech training data; feature-level augmentation processing including time stretching processing and frequency masking processing is performed on the first-process speech training data to obtain second-process speech training data; quality evaluation processing including perceptual evaluation of speech quality, speech intelligibility evaluation in a noise environment, and speech signal intensity evaluation in a noise environment is performed on the second-process speech training data to screen out enhanced speech training data that meets the preset quality evaluation conditions. It can be understood that noise is randomly selected from a public noise database (such as MUSAN, ESC-50) and mixed into the original speech data at different signal-to-noise ratios (SNRs) to obtain first-process speech training data after background noise augmentation; speech conversion technology (such as GAN-based speech style conversion) is used to generate speech data with different accent styles based on the original speech data to obtain first-process speech training data after accent augmentation, so as to improve the model's adaptability to accent changes. Further, the SpecAugment technology is used to perform time-frequency domain augmentation processing on the first-process speech training data through time stretching, frequency masking, etc. to obtain second-process speech training data, so as to enhance the diversity of training data, expand the scale of training data, and enable the model to learn richer speech change features. Quality evaluation is performed on the enhanced second-process speech training data. Specific quality evaluation indicators can include: PESQ (Perceptual Evaluation of Speech Quality), which is used to measure the intelligibility and naturalness of enhanced speech, with a range of [-0.5, 4.5], and a value below 1.5 is considered low quality. STOI (Short-Time Objective Intelligibility), which is used to evaluate the intelligibility of speech in a noise environment, with a range of [0, 1], and a value below 0.6 needs to be filtered. SNR (Signal-to-Noise Ratio), which is used to check the signal intensity after noise mixing to ensure that it is not lower than 5 dB (to avoid completely drowning out the original speech). Further, a quality evaluation condition is constructed, where PESQ < 1.5 or STOI < 0.6 → directly discard; SNR < 5 dB → readjust the noise mixing parameters. Based on the above quality evaluation indicators and corresponding quality evaluation conditions, enhanced speech training data that meets the quality evaluation conditions is screened from the second-process speech training data.

[0090] In this embodiment, any enhanced speech training data and the corresponding real text training data are randomly selected from the data training set, and then the enhanced speech training data is input into the current automatic speech recognition model, and a set of predicted texts is obtained by predicting it through the current automatic speech recognition model. , where, Note that during reinforcement learning training, a random sampling method is used to select actions (i.e., generate text). For example, the softmax probability distribution is used to select characters instead of the greedy decoding method. In this way, even if the input data is the same, the predicted text generated each time will be different. Moreover, by generating a set of samples for the same input to estimate the expected reward of the policy, the influence of noise can be reduced, making the update of the policy more stable.

[0091] Step S13: Determine the text matching degree and semantic similarity between each of the predicted texts and the true text training data to obtain a set of target reward values, and update the optimization policy at the current moment based on the target reward values, and then jump to execute the step of determining the current optimization policy in the optimization action of the automatic speech recognition model at the current time step until the current automatic speech recognition model meets the optimization end condition and then output the corresponding optimized automatic speech recognition model.

[0092] In this embodiment, the character-level edit distance between each of the predicted texts and the true text training data is calculated to obtain a text matching degree score; the semantic similarity between each of the predicted texts and the true text training data is calculated to obtain a semantic similarity score; the target reward value is obtained through the text matching degree score, the semantic similarity score, and the corresponding weight coefficient; the target reward values between all the predicted texts and the true text training data are statistically calculated to obtain a set of target reward values. It can be understood that the character-level edit distance (Levenshtein Distance) between each predicted text and the true text training data is calculated to measure the text matching degree, and then, based on the character edit distance, the character error rate is further calculated, and the corresponding calculation result is used as the text matching degree score. Among them, the character error rate (Character Error Rate, CER) is calculated based on the Levenshtein distance, which represents the minimum number of edit operations required to convert the predicted text into the reference text, including three operations: insertion, deletion, and substitution. The Sentence-BERT semantic model is used to calculate the semantic similarity Semantic_Score between the predicted text and the true text to obtain the semantic similarity score to ensure the rationality of the ASR result at the semantic level.

[0093] Combine the text matching degree score and the semantic similarity score to calculate the final target reward value in a weighted manner:

[0094] ;

[0095] Among them, and is a hyperparameter, CER represents the text matching score, and Semantic_Score represents the semantic similarity score, which is used to adjust the importance of text matching and semantic similarity.

[0096] Statistically calculate the target reward value between the same group of predicted texts and the true texts, and normalize the above-mentioned target reward values of the same group to obtain a group of normalized target reward values, ensuring that the rewards are within a reasonable range and preventing the reward scale from changing too much to affect the stability of reinforcement learning.

[0097] In this embodiment, during the overall optimization training process, calculate the prediction text character error rate and the prediction text semantic similarity score of the prediction text output by the current automatic speech recognition model to obtain the corresponding evaluation results; adjust the weight coefficient of the target reward value calculation process according to the evaluation results, and dynamically adjust the hyperparameters of reinforcement learning according to the model performance through an adaptive learning rate strategy. It can be understood that the reward strategy is dynamically adjusted according to the evaluation results. Calculate the common ASR evaluation metric character error rate (CER), calculate the semantic similarity score to measure the semantic rationality of the ASR text. Adjust the weight parameters in the reward calculation according to the evaluation results and , to optimize the reward strategy, and at the same time adopt an adaptive learning rate strategy to dynamically adjust the hyperparameters of reinforcement learning according to the model performance.

[0098] As Figure 2 shown, the present invention discloses a specific ASR model optimization system, which includes four modules, namely: a data augmentation module, a reinforcement learning optimization module, a reward calculation module, and a model evaluation module. Among them, the data augmentation module enhances the data through operations such as adding background noise, simulating accent changes, speed changes, and sample screening and filtering, and at the same time processes the data with the "qwen2-audio-instruct" automatic speech recognition technology. The processed data enters the data preprocessing link. Reward calculation module: Calculate and process the rewards between the predicted text output after the ASR model prediction and the true text, "character-level accuracy reward" and "semantic similarity reward", to provide a basis for model optimization. Reinforcement learning optimization module: Based on the reward calculation results, implement model optimization through "relative ranking reward optimization" and "KL divergence regularization" methods. Model evaluation module: Use evaluation metrics such as WER, CER, BLEU, etc. to evaluate the effect of the optimized model, and then adjust the reward weights according to the evaluation results to form a feedback loop of the optimization strategy.

[0099] It can be seen that the present application discloses an optimization method for a speech recognition model based on reinforcement learning, including: determining the current optimization strategy in the optimization action of the automatic speech recognition model at the current time step; randomly selecting any augmented speech training data and the corresponding real text training data from the data training set as target training data, and inputting the target training data into the current automatic speech recognition model for prediction to obtain a plurality of predicted texts output by the current automatic speech recognition model; wherein, the current automatic speech recognition model is a model whose model parameters are updated based on the current optimization strategy; determining the text matching degree and semantic similarity between each predicted text and the real text training data to obtain a set of target reward values, and updating the optimization strategy at the current moment based on the target reward values, and jumping to execute the step of determining the current optimization strategy in the optimization action of the automatic speech recognition model at the current time step, until the current automatic speech recognition model meets the optimization end condition, and outputting the corresponding optimized automatic speech recognition model. Thus, by randomly selecting augmented speech training data, the model learns different noises, accents and time-frequency features, reduces overfitting to a specific scenario, and improves cross-environment recognition ability. By generating multiple ASR predicted texts from the same input, the reward scale difference is eliminated through within-group normalization, training oscillations are reduced, and the stability of policy updates is ensured. Moreover, in the process of calculating the target reward value, by combining the calculation of text matching degree and semantic similarity, not only the character-level matching is optimized, but also the overall semantics of the sentence is concerned, improving the naturalness and practicality of the recognition result.

[0100] Referring to Figure 3 as shown, the present invention also correspondingly discloses an optimization device for a speech recognition model based on reinforcement learning, including:

[0101] A policy determination module 11, configured to determine the current optimization strategy in the optimization action of the automatic speech recognition model at the current time step;

[0102] A prediction module 12, configured to randomly select any augmented speech training data and the corresponding real text training data from the data training set as target training data, and input the target training data into the current automatic speech recognition model for prediction to obtain a plurality of predicted texts output by the current automatic speech recognition model; wherein, the current automatic speech recognition model is a model whose model parameters are updated based on the current optimization strategy;

[0103] The policy adjustment module 13 is configured to determine the text matching degree and semantic similarity between each of the predicted texts and the true text training data, so as to obtain a set of target reward values, update the optimization policy at the current moment based on the target reward values, and jump to execute the step of the current optimization policy in the step of determining the optimization action of the automatic speech recognition model at the current time step, until the current automatic speech recognition model meets the optimization end condition, and then output the corresponding optimized automatic speech recognition model.

[0104] It can be seen that this application discloses determining the current optimization policy in the optimization action of the automatic speech recognition model at the current time step; randomly selecting any enhanced speech training data and the corresponding true text training data from the data training set as the target training data, and inputting the target training data into the current automatic speech recognition model for prediction, so as to obtain several predicted texts output by the current automatic speech recognition model; wherein, the current automatic speech recognition model is a model whose model parameters are updated based on the current optimization policy; determining the text matching degree and semantic similarity between each of the predicted texts and the true text training data, so as to obtain a set of target reward values, updating the optimization policy at the current moment based on the target reward values, and jumping to execute the step of the current optimization policy in the step of determining the optimization action of the automatic speech recognition model at the current time step, until the current automatic speech recognition model meets the optimization end condition, and then outputting the corresponding optimized automatic speech recognition model. Thus, by randomly selecting enhanced speech training data, the model learns different noise, accents, and time-frequency characteristics, reduces overfitting to specific scenarios, and improves cross-environment recognition ability. By generating multiple ASR predicted texts from the same input, the reward scale difference is eliminated through within-group normalization, training oscillations are reduced, and the stability of policy updates is ensured. Moreover, in the process of calculating the target reward values, by combining the calculation of text matching degree and semantic similarity, not only the character-level matching is optimized, but also the overall semantics of the sentence is concerned, improving the naturalness and practicality of the recognition results.

[0105] Furthermore, the embodiment of this application also discloses an electronic device Figure 4 It is a structural diagram of an electronic device 20 shown according to an exemplary embodiment, and the content in the figure should not be regarded as any limitation on the scope of use of this application.

[0106] Figure 4Schematic diagram of the structure of an electronic device 20 provided by an embodiment of the present application. The electronic device 20 may specifically include: at least one processor 21, at least one memory 22, a power supply 23, a communication interface 24, an input / output interface 25, and a communication bus 26. Among them, the memory 22 is used to store a computer program, and the computer program is loaded and executed by the processor 21 to implement the relevant steps in the voice recognition model optimization method based on reinforcement learning disclosed in any of the foregoing embodiments. In addition, the electronic device 20 in this embodiment may specifically be an electronic computer.

[0107] In this embodiment, the power supply 23 is used to provide operating voltage for each hardware device on the electronic device 20; the communication interface 24 can create a data transmission channel between the electronic device 20 and external devices, and the communication protocol it follows is any communication protocol applicable to the technical solution of the present application, and no specific limitation is imposed thereon here; the input / output interface 25 is used to obtain external input data or output data to the outside, and the specific interface type thereof can be selected according to specific application needs, and no specific limitation is made here.

[0108] Among them, the processor 21 may include one or more processing cores, such as a 4-core processor, an 8-core processor, etc. The processor 21 may be implemented in at least one of the following hardware forms: DSP (Digital Signal Processing), FPGA (Field-Programmable Gate Array), and PLA (Programmable Logic Array). The processor 21 may also include a main processor and a coprocessor. The main processor is a processor used to process data in the wake state, also known as the CPU (Central Processing Unit); the coprocessor is a low-power processor used to process data in the standby state. In some embodiments, the processor 21 may be integrated with a GPU (Graphics Processing Unit), and the GPU is responsible for rendering and drawing the content to be displayed on the display screen. In some embodiments, the processor 21 may also include an AI (Artificial Intelligence) processor, and the AI processor is used to process computational operations related to machine learning.

[0109] In addition, as a carrier for resource storage, the memory 22 may be a read-only memory, a random access memory, a disk, or an optical disc, etc. The resources stored thereon may include an operating system 221, a computer program 222, etc., and the storage method may be short-term storage or permanent storage.

[0110] Among them, the operating system 221 is used to manage and control each hardware device on the electronic device 20 and the computer program 222, so as to implement the operation and processing of the massive data 223 in the memory 22 by the processor 21. It can be Windows Server, Netware, Unix, Linux, etc. The computer program 222 can further include computer programs capable of completing other specific tasks in addition to the computer program capable of implementing the method for optimizing the speech recognition model based on reinforcement learning executed by the electronic device 20 disclosed in any of the foregoing embodiments. The data 223 can include not only the data transmitted by external devices received by the electronic device, but also the data collected by its own input / output interface 25, etc.

[0111] Furthermore, the present application also discloses a computer-readable storage medium for storing a computer program; wherein, when the computer program is executed by a processor, the method for optimizing the speech recognition model based on reinforcement learning disclosed above is implemented. For the specific steps of this method, reference can be made to the corresponding content disclosed in the foregoing embodiments, and details will not be repeated here.

[0112] The various embodiments in this specification are described in a progressive manner. Each embodiment focuses on the differences from other embodiments. The same or similar parts between the various embodiments can be referred to each other. For the devices disclosed in the embodiments, since they correspond to the methods disclosed in the embodiments, the description is relatively simple, and the relevant parts can be referred to the description of the method part.

[0113] Those skilled in the art may further realize that the units and algorithm steps of each example described in connection with the embodiments disclosed herein can be implemented by electronic hardware, computer software, or a combination of the two. To clearly illustrate the interchangeability of hardware and software, the composition and steps of each example have been generally described according to functions in the above description. Whether these functions are executed in a hardware or software manner depends on the specific application and design constraints of the technical solution. Those skilled in the art can use different methods to implement the described functions for each specific application, but such implementation should not be considered to exceed the scope of this application. The steps of the methods or algorithms described in connection with the embodiments disclosed herein can be directly implemented by hardware, software modules executed by a processor, or a combination of the two. The software modules can be placed in a random access memory (RAM), internal memory, read-only memory (ROM), electrically programmable read-only memory (EPROM), electrically erasable programmable read-only memory (EEPROM), registers, hard disks, removable disks, CD-ROM (Compact Disc-Read Only Memory), or any other form of storage medium known in the technical field.

[0114] Finally, it should also be noted that in this document, relational terms such as "first" and "second" are only used to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any actual relationship or order between these entities or operations. Moreover, the term "comprising", "including" or any other variant thereof is intended to cover non-exclusive inclusion, so that a process, method, article or device comprising a series of elements not only includes those elements, but also includes other elements not expressly listed, or elements inherent to such process, method, article or device. Without further limitation, an element defined by the statement "comprising an..." does not exclude the existence of additional identical elements in the process, method, article or device comprising the element.

[0115] The above has introduced the solution provided by the present invention in detail. Specific examples are used herein to illustrate the principle and implementation manner of the present invention. The description of the above embodiments is only used to help understand the method and its core idea of the present invention; at the same time, for those of ordinary skill in the art, according to the idea of the present invention, there will be changes in the specific implementation manner and application scope. In summary, the content of this specification should not be construed as a limitation on the present invention.

Claims

1. A speech recognition model optimization method based on reinforcement learning, characterized in that: include: Determining a current optimization strategy in an automatic speech recognition model optimization action at a current time step; Randomly select any enhanced speech training data and corresponding real text training data from the data training set as target training data, and input the target training data into the current automatic speech recognition model for prediction, so as to obtain a number of predicted texts output by the current automatic speech recognition model; wherein the current automatic speech recognition model is a model for updating model parameters based on the current optimization strategy; Determine the text matching degree and semantic similarity between each of the predicted texts and the real text training data to obtain a set of target reward values, and update the optimization strategy at the current moment based on the target reward values, and jump to execute the step of determining the current optimization strategy in the automatic speech recognition model optimization action at the current time step, until the current automatic speech recognition model meets the optimization end condition and outputs the corresponding optimized automatic speech recognition model.

2. The method for optimizing a speech recognition model based on reinforcement learning according to claim 1, characterized in that: Before randomly selecting any enhanced speech training data and corresponding real text training data from the data training set as target training data, the method further includes: Acquire enhanced speech training data after data enhancement processing; wherein the enhanced speech training data is enhanced speech data after background noise enhancement and speaker accent enhancement; Inputting the enhanced speech training data into a base model so as to perform transcription processing on the enhanced speech training data through the base model to obtain real text training data corresponding to the enhanced speech training data; wherein the base model is an audio language model for processing speech data into corresponding unique text data; A data training set is constructed based on each of the enhanced speech training data and the corresponding unique real text training data.

3. The method for optimizing a speech recognition model based on reinforcement learning according to claim 2, characterized in that: The step of obtaining the enhanced speech training data after data enhancement processing includes: Performing content level enhancement processing including background noise enhancement and speaker accent enhancement processing on the original speech data to obtain first process speech training data; Performing a feature level enhancement process including a time expansion process and a frequency masking process on the first process speech training data to obtain a second process speech training data; The second-process speech training data is subjected to quality assessment processing including speech quality perception assessment, speech intelligibility assessment in a noisy environment, and speech signal strength assessment in a noisy environment, so as to screen out enhanced speech training data that meets preset quality assessment conditions.

4. The method for optimizing a speech recognition model based on reinforcement learning according to claim 1, characterized in that: The determining of the current optimization strategy in the automatic speech recognition model optimization action at the current time step includes: Standardize the reward values ​​of a group of target reward values ​​at the previous time step to obtain a relative reward value centered on the mean reward value within the group and with the standard deviation as the unit; Constructing a strategy update objective function using the relative reward value, the KL regularization term, the previous optimization strategy at the previous time step, and the current optimization strategy at the current time step; Obtain the current optimization strategy at the function maximum value of the strategy update objective function.

5. The method for optimizing a speech recognition model based on reinforcement learning according to claim 4, characterized in that: Before inputting the target training data into the current automatic speech recognition model, the method further includes: The model parameters of the previous automatic speech recognition model at the previous time step are updated by gradient ascent based on the current optimization strategy to obtain the current automatic speech recognition model at the current time step.

6. The method for optimizing a speech recognition model based on reinforcement learning according to claim 1, characterized in that: Determining the text matching degree and semantic similarity between each of the predicted texts and the real text training data to obtain a set of target reward values ​​includes: Calculating the character-level edit distance between each of the predicted texts and the real text training data to obtain a text matching score; Calculating the semantic similarity between each of the predicted texts and the real text training data to obtain a semantic similarity score; Obtaining a target reward value through the text matching score, the semantic similarity score and the corresponding weight coefficient; The target reward values ​​between all the predicted texts and the real text training data are counted to obtain a set of target reward values.

7. The method for optimizing a speech recognition model based on reinforcement learning according to any one of claims 1 to 6, characterized in that: Also includes: Calculate the predicted text character error rate and the predicted text semantic similarity score of the predicted text output by the current automatic speech recognition model to obtain corresponding evaluation results; The weight coefficient of the target reward value calculation process is adjusted according to the evaluation results, and the hyperparameters of the reinforcement learning are dynamically adjusted according to the model performance through an adaptive learning rate strategy.

8. A speech recognition model optimization device based on reinforcement learning, characterized in that: include: A strategy determination module, used to determine a current optimization strategy in an automatic speech recognition model optimization action at a current time step; A prediction module, used for randomly selecting any enhanced speech training data and corresponding real text training data from the data training set as target training data, and inputting the target training data into the current automatic speech recognition model for prediction, so as to obtain a number of predicted texts output by the current automatic speech recognition model; wherein the current automatic speech recognition model is a model for updating model parameters based on the current optimization strategy; The strategy adjustment module is used to determine the text matching degree and semantic similarity between each of the predicted texts and the real text training data to obtain a set of target reward values, and update the optimization strategy at the current moment based on the target reward values, and jump to execute the steps of determining the current optimization strategy in the automatic speech recognition model optimization action at the current time step, until the current automatic speech recognition model meets the optimization end conditions and outputs the corresponding optimized automatic speech recognition model.

9. An electronic device, characterized in that: include: Memory, used to store computer programs; A processor, configured to execute the computer program to implement the steps of the reinforcement learning-based speech recognition model optimization method as described in any one of claims 1 to 7.

10. A computer-readable storage medium, characterized in that: Used to store computer programs; wherein, when the computer program is executed by a processor, the steps of the reinforcement learning-based speech recognition model optimization method as described in any one of claims 1 to 7 are implemented.

Citation Information

Cited By

  • Analog circuit parameter determination method and device, medium and product

    CN120337840A

  • Explainable electric power image-text large model training method and system

    CN121303337A

  • Immersive man-machine interaction method and system based on voice recognition

    CN121415769A

  • Immersive Human-Computer Interaction Method and System Based on Speech Recognition

    CN121415769B

  • Voice text matching method and device based on reinforcement learning, equipment and medium

    CN122024705A