Speech recognition model training method and device and readable storage medium
Through the multi-source information fusion and improved Whisper model construction method, the shortcomings of the weakly supervised speech recognition model in the inaccurate labeling information and multimodal data fusion are solved, and higher recognition accuracy and robustness are achieved, and complex and changeable speech environments and multilingual multi-task scenarios are adapted.
Patent Information
- Application Number
- CN202510585574.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-07
- Publication Date
- 2025-08-08
AI Technical Summary
The existing weakly supervised speech recognition model has shortcomings in the inaccurate annotation information and the fusion of multimodal data feature representations and semantic information, resulting in insufficient performance and robustness, making it difficult to adapt to complex and changeable speech environments and multilingual multitasking scenarios.
Multi-source information fusion correction and feature extraction are carried out through an external knowledge base, and improved Whisper encoder, discriminator network and multimodal encoder are built, and adversarial learning and cross-modal attention mechanism are adopted, combined with distributed training and model compression technology to perform model training and evaluation.
It improves the robustness and generalization capabilities of the model, improves the recognition accuracy and performance in multimodal tasks, and adapts to different voice scenarios and user groups.
Smart Images

Figure CN120452426A_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of artificial intelligence technology, and in particular to a speech recognition model training method, device, and readable storage medium. Background Art
[0002] Early speech recognition technology was primarily based on statistical methods, such as dynamic time warping and hidden Markov models. With the development of deep learning, neural network models have been widely used in speech recognition, significantly improving recognition performance. Deep learning models typically require large amounts of labeled data for training, and this demand has become a bottleneck restricting their further development. The data labeling process is costly, time-consuming, and labor-intensive.
[0003] Weakly supervised learning provides a more cost-effective solution. It does not require each training sample to have a precise label. By utilizing incompletely, imprecisely, or inaccurately labeled data, it can still train an effective model, thereby reducing the cost of data labeling and alleviating the problem of data scarcity to a certain extent.
[0004] With the widespread global adoption of voice interaction, such as various voice assistants and multilingual customer service, a massive amount of multilingual and multi-task voice data has been generated. Weakly supervised learning speech recognition models can fully tap into and leverage the value of this data, improving the model's performance and adaptability in multilingual and multi-task scenarios. In areas such as smart homes, smart security, and autonomous driving, speech recognition technology must be able to operate accurately and stably in complex and changing environments and for diverse user groups. Weakly supervised learning speech recognition models can better meet the performance, robustness, and adaptability requirements of these practical applications, promoting the widespread application and implementation of speech recognition technology in more fields.
[0005] Therefore, how to improve the performance of weakly supervised models in the field of speech recognition has become a problem that needs to be solved. Summary of the Invention
[0006] The technical problem to be solved by this application is to provide a speech recognition model training method, device and readable storage medium to solve the problems existing in the prior art.
[0007] In a first aspect, the present application provides a method for training a speech recognition model, the method comprising:
[0008] S1. Data preparation: Multi-source information fusion and feature extraction are performed through an external knowledge base to obtain training data;
[0009] S2. Model construction, including: improved Whisper encoder, discriminator network, improved Whisper decoder and multimodal encoder;
[0010] The Transformer layer within the improved Whisper encoder includes a multi-level fusion module for fusing different features contained in the training data and adjusting the model's weight distribution for different features;
[0011] The discriminator network and the improved Whisper encoder form an adversarial structure;
[0012] The improved Whisper decoder internally includes a bidirectional cross-modal attention layer for performing data attention calculation, and the bidirectional cross-modal attention layer includes a dynamic weight adjuster for dynamically adjusting the attention weight in the bidirectional fusion process;
[0013] The multimodal encoder is used to encode different features into a shared semantic space, wherein the shared semantic space includes a semantic information interaction module for fusing neighboring data features;
[0014] S3. Model training: Train the constructed model using training data to obtain a training model;
[0015] S4. Model evaluation: Evaluate the training model through multi-dimensional indicators to obtain evaluation results.
[0016] In some embodiments, S1 includes:
[0017] Verify and correct the initial text data and audio data through an external knowledge base to obtain corrected data;
[0018] Feature extraction processing is performed on the corrected data to obtain the phoneme features, prosodic features and vocal tract features of the speech signal.
[0019] In some embodiments, in S2, the input of the discriminator network is the feature vector output by the improved Whisper encoder, and the discriminator network is used to determine the data category of the feature vector, wherein the data category of the feature vector includes a long-tail data category and a mainstream data category.
[0020] In some embodiments, before S3, the process further includes:
[0021] A magnitude-based pruning algorithm is used to sparsify the Transformer architecture of the Whisper model.
[0022] In some embodiments, S3 includes at least one of the following:
[0023] In a distributed training environment, training data is dynamically allocated to different computing nodes through an adaptive data partitioner to perform distributed training of the model;
[0024] During the transmission of model parameters between different computing nodes, model compression technology is used to compress the parameters in real time;
[0025] During the model training process, the model is compressed and evaluated, and the compression strategy is dynamically adjusted based on the compressed model performance and resource consumption.
[0026] In some embodiments, S3 further includes:
[0027] After the model training iteration, the model stability is evaluated based on the uncertainty index, and the model hyperparameters are optimized according to the evaluation results, wherein the uncertainty index includes at least one of the prediction entropy and the prediction variance.
[0028] In some embodiments, S4 includes:
[0029] The training model is evaluated using a multi-dimensional indicator through a multi-indicator evaluator to obtain an evaluation result, wherein the multi-dimensional indicator includes at least one of speech recognition accuracy, recall rate, model calculation efficiency, memory usage, and robustness.
[0030] In a second aspect, the present application provides a speech recognition model training device, the device comprising:
[0031] The data preparation module is configured to prepare data by performing multi-source information fusion correction and feature extraction processing through an external knowledge base to obtain training data;
[0032] A model building module configured to build a model including: an improved whisper encoder, a discriminator network, an improved whisper decoder, and a multimodal encoder;
[0033] The Transformer layer within the improved Whisper encoder includes a multi-level fusion module for fusing different features contained in the training data and adjusting the model's weight distribution for different features;
[0034] The discriminator network and the improved Whisper encoder form an adversarial structure;
[0035] The improved Whisper decoder internally includes a bidirectional cross-modal attention layer for performing data attention calculation, and the bidirectional cross-modal attention layer includes a dynamic weight adjuster for dynamically adjusting the attention weight in the bidirectional fusion process;
[0036] The multimodal encoder is used to encode different features into a shared semantic space, wherein the shared semantic space includes a semantic information interaction module for fusing neighboring data features;
[0037] A model training module is configured to perform model training: the constructed model is trained using training data to obtain a training model;
[0038] The model evaluation module is configured to perform model evaluation: the training model is evaluated through multi-dimensional indicators to obtain evaluation results.
[0039] In a third aspect, the present application provides a speech recognition model training device, comprising a memory and a processor, wherein the memory stores a computer program, and the processor is configured to run the computer program to implement the speech recognition model training method described in the first aspect above.
[0040] In a fourth aspect, the present application provides a computer-readable storage medium having a computer program stored thereon, and when the computer program is executed by a processor, the speech recognition model training method described in the first aspect above is implemented.
[0041] The speech recognition model training method, device and readable storage medium provided in the present application include: data preparation: multi-source information fusion correction and feature extraction processing are performed through an external knowledge base to obtain training data; model construction, including: an improved Whisper encoder, a discriminator network, an improved Whisper decoder and a multimodal encoder; wherein the Transformer layer inside the improved Whisper encoder includes a multi-level fusion module for fusing different features contained in the training data and adjusting the model's weight distribution for different features; the discriminator network and the improved Whisper encoder form an adversarial structure; the improved Whisper decoder includes a bidirectional cross-modal attention layer for performing data attention calculation, and the bidirectional cross-modal attention layer includes a dynamic weight adjuster for dynamically adjusting the attention weight in the bidirectional fusion process; the multimodal encoder is used to encode different features into a shared semantic space, and the shared semantic space includes a semantic information interaction module for fusing neighboring data features; model training: training the constructed model with the training data to obtain a training model; model evaluation: evaluating the training model with multi-dimensional indicators to obtain an evaluation result. This application can improve the robustness and generalization ability of the model, improve recognition accuracy, improve the overall performance of the model, and improve the performance of the model in multimodal tasks. BRIEF DESCRIPTION OF THE DRAWINGS
[0042] The accompanying drawings, which are incorporated in and constitute a part of this specification, illustrate embodiments consistent with the present application and, together with the description, serve to explain the principles of the present application.
[0043] Figure 1 A schematic diagram of a speech recognition model training method provided in an embodiment of the present application;
[0044] Figure 2 A schematic diagram of the overall structure of the speech recognition model provided in the embodiment of the present application;
[0045] Figure 3 A schematic diagram of an improved Whisper encoder provided in an embodiment of the present application;
[0046] Figure 4 A schematic diagram of an improved Whisper decoder provided in an embodiment of the present application;
[0047] Figure 5 A schematic diagram of the dynamic model structure sparsification provided in an embodiment of the present application;
[0048] Figure 6 A schematic diagram of hyperparameter optimization based on uncertainty sampling provided in an embodiment of the present application;
[0049] Figure 7 A schematic diagram of the structure of a speech recognition model training device provided in an embodiment of the present application;
[0050] Figure 8 A schematic structural diagram of another speech recognition model training device provided in an embodiment of the present application.
[0051] The above drawings illustrate specific embodiments of the present application, which will be described in more detail below. These drawings and the textual description are not intended to limit the scope of the present application in any way, but rather to illustrate the concepts of the present application to those skilled in the art by reference to specific embodiments. DETAILED DESCRIPTION
[0052] In order to enable those skilled in the art to better understand the technical solution of the present application, the implementation methods of the present application will be further described in detail below with reference to the accompanying drawings.
[0053] It should be understood that the specific embodiments and drawings described herein are only used to explain the present application, rather than to limit the present application.
[0054] It can be understood that, in the absence of conflict, the various embodiments and features in the embodiments of the present application can be combined with each other.
[0055] It will be understood that, for the sake of ease of description, the drawings of this application only show the parts related to this application, while the parts not related to this application are not shown in the drawings.
[0056] It can be understood that each unit and module involved in the embodiments of the present application may correspond to only one physical structure, or may be composed of multiple physical structures, or multiple units and modules may be integrated into one physical structure.
[0057] It can be understood that the terms "first", "second", etc. in the embodiments of the present application are used to distinguish different objects, or to distinguish different processing of the same object, rather than to describe a specific order of objects.
[0058] It is understandable that, in the absence of conflict, the functions and steps marked in the flowcharts and block diagrams of the present application may occur in an order different from that marked in the drawings.
[0059] It is understood that the flowcharts and block diagrams of the present application illustrate the possible architectures, functions, and operations of the systems, devices, equipment, and methods according to the various embodiments of the present application. Each box in the flowchart or block diagram may represent a unit, module, program segment, or code, which contains executable instructions for implementing the specified functions. Moreover, each box or combination of boxes in the block diagram and flowchart may be implemented by a hardware-based system that implements the specified functions, or by a combination of hardware and computer instructions.
[0060] It can be understood that the units and modules involved in the embodiments of the present application can be implemented by software or hardware, for example, the units and modules can be located in a processor.
[0061] The shortcomings of existing weakly supervised models are as follows:
[0062] 1. Incomplete, inaccurate, or noisy annotation information causes the model to learn limited or incorrect information, affecting the model's performance and generalization ability.
[0063] 2. Traditional models are relatively weak at learning complex features and semantic information, making it difficult to fully explore the underlying patterns in the data. They also struggle to discover and leverage the underlying connections and shared features between different long-tail data categories. Traditional speech recognition models primarily focus on basic features such as phonemes, while underutilizing other important features such as prosody and the vocal tract. Different types of speech data have varying degrees of dependency on features, making fixed feature extraction methods difficult to adapt to diverse speech scenarios.
[0064] 3. Traditional evaluation indicators are difficult to accurately measure the performance of weakly supervised learning models, and the model's hyperparameters are difficult to adjust, making it difficult to find the optimal hyperparameter combination.
[0065] 4. In multimodal weakly supervised learning, the feature representations and semantic information of different modal data differ, making effective alignment and fusion difficult, resulting in the model being unable to fully utilize multimodal information to improve performance.
[0066] The technical problems to be solved by this application are as follows:
[0067] 1. Reduce model learning errors caused by inaccurate labeling and improve the robustness and generalization ability of the model.
[0068] 2. Add richer details to the labels so that the model can learn more comprehensive speech information, thereby improving recognition accuracy.
[0069] 3. Propose new model evaluation indicators to more effectively adjust hyperparameters to improve the overall performance of the model.
[0070] 4. Fusion features to fully utilize the complementarity of multimodal data and improve the performance of the model in multimodal tasks.
[0071] The following specific embodiments describe in detail the technical solution of the present application and how the technical solution of the present application solves the above-mentioned technical problems. The following specific embodiments can be combined with each other, and the same or similar concepts or processes may not be repeated in some embodiments. The embodiments of the present application will be described below in conjunction with the accompanying drawings.
[0072] The present application provides a speech recognition model training method, the working process of which can be implemented by electronic devices, such as computers, handheld smart terminals, etc. For the convenience of explanation, the embodiments of the present application are described with the method execution subject being a computer.
[0073] Figure 1 A schematic diagram of a speech recognition model training method provided in an embodiment of the present application, Figure 2 The overall structure diagram of the speech recognition model provided in the embodiment of the present application is as follows: Figure 1 as well as Figure 2 As shown, the present application provides a speech recognition model training method, which includes S1-S4, as follows:
[0074] S1. Data preparation: Multi-source information fusion and feature extraction are performed through an external knowledge base to obtain training data;
[0075] In some embodiments, S1 includes:
[0076] Verify and correct the initial text data and audio data through an external knowledge base to obtain corrected data;
[0077] Feature extraction processing is performed on the corrected data to obtain the phoneme features, prosodic features and vocal tract features of the speech signal.
[0078] In this application, multi-source information fusion is used to correct labels: the initial annotated phoneme labels are corrected using external knowledge bases to ensure the quality of training data.
[0079] Specifically, before model training, an external knowledge fusion module is constructed. During the initial labeling phase, this multi-source information is used to perform preliminary verification and correction of the annotated phoneme labels. When users provide extensive feedback on a particular speech recognition result, the labels for that data are prioritized for review and correction, and the corrected labeled data is used as training data input. Furthermore, a feedback mechanism is incorporated to enable the model to dynamically adjust its parameters based on the discrepancy between its own predictions and the corrected labels, in order to adapt to more accurate label information and reduce the interference of incorrect labels on model learning.
[0080] S2. Model construction, including: improved Whisper encoder, discriminator network, improved Whisper decoder and multimodal encoder;
[0081] The Transformer layer within the improved Whisper encoder includes a multi-level fusion module for fusing different features contained in the training data and adjusting the model's weight distribution for different features;
[0082] Specifically, Figure 3 A schematic diagram of the improved Whisper encoder provided in an embodiment of the present application is shown in FIG. Figure 3 As shown in the figure, in the Transformer layer inside the encoder, a multi-level fusion module is designed to fuse phoneme features, prosodic features, and vocal tract features in Transformer blocks at different levels, automatically adjusting the model's weight distribution for different features. In the lower-level Transformer blocks, phoneme features are first fused with prosodic features, and then fused with vocal tract features at a higher level, enabling the model to gradually learn more comprehensive and richer speech feature representations.
[0083] In the model constructed in this application, the discriminator network and the improved Whisper encoder form an adversarial structure;
[0084] In some embodiments, in S2, the input of the discriminator network is the feature vector output by the improved Whisper encoder, and the discriminator network is used to determine the data category of the feature vector, wherein the data category of the feature vector includes a long-tail data category and a mainstream data category.
[0085] In this application, adversarial learning is used to enhance long-tail data: a discriminator network is constructed and adversarial training is performed with the encoder to improve the model's ability to recognize long-tail data.
[0086] Specifically, this application constructs a discriminator network, which forms an adversarial structure with the encoder part of the Whisper model. The input of the discriminator network is the feature vector output by the encoder, and its task is to determine whether the feature vector belongs to the long-tail data category or the mainstream data category. During the training process, the encoder attempts to generate feature vectors that can deceive the discriminator, that is, to make it difficult for the discriminator to distinguish the category to which it belongs; the discriminator strives to improve its discrimination ability. Through this adversarial training, the encoder learns more generalized feature representations, especially for long-tail data categories.
[0087] At the same time, for long-tail data categories, a group adversarial training strategy is adopted. Based on the similarity or correlation between categories, the long-tail data is divided into several groups. Different adversarial training parameters and strategies are used for data in different groups. This promotes feature sharing and transfer between categories and improves the model's ability to recognize long-tail data.
[0088] In the model constructed in this application, the improved Whisper decoder includes a bidirectional cross-modal attention layer for performing data attention calculation, and the bidirectional cross-modal attention layer includes a dynamic weight adjuster for dynamically adjusting the attention weight in the bidirectional fusion process;
[0089] In this application, a bidirectional fusion of cross-modal attention mechanism is adopted: a bidirectional cross-modal attention layer is used to fuse speech features and text features, and the attention weights are dynamically adjusted according to real-time feature changes.
[0090] Specifically, Figure 4 A schematic diagram of the improved Whisper decoder provided in an embodiment of the present application is shown in FIG. Figure 4 As shown in the figure, based on the Whisper model, this application adds a bidirectional cross-modal attention layer. This layer receives speech features and text features as input and calculates speech-to-text attention and text-to-speech attention, respectively. Within the bidirectional cross-modal attention layer, a dynamic weight adjuster is designed. This adjuster dynamically adjusts the attention weights during the bidirectional fusion process based on real-time feature changes in the speech and text data.
[0091] In the model constructed in this application, the multimodal encoder is used to encode different features into a shared semantic space, and the shared semantic space includes a semantic information interaction module for fusing neighboring data features;
[0092] In this application, a multimodal shared semantic space is constructed: speech features and text features are encoded into a shared semantic space, and neighbor data features are fused through a semantic information interaction module.
[0093] Specifically, refer to Figure 4, this application constructs a multimodal encoder that encodes speech features and text features into a shared semantic space respectively. The encoder consists of two sub-encoders, one for speech feature encoding and the other for text feature encoding. During the encoding process, by sharing parameters or constraints, the vector representations of speech and text features in the shared semantic space are made comparable, and a semantic information interaction module is designed in the shared semantic space. This module uses the nearest neighbor search algorithm in the semantic space to find text data that is semantically similar to the speech data, and fuses the feature information of these neighboring data into the feature representation of the original data.
[0094] In some embodiments, before S3, the method further includes: performing sparsification processing on the Transformer architecture of the Whisper model using an amplitude-based pruning algorithm.
[0095] In this application, dynamic model structure sparsification is adopted: pruning algorithm and sparse regularization technology are used to make model parameters more sparse and effective, thereby improving computational efficiency.
[0096] Specifically, Figure 5 A schematic diagram of the dynamic model structure sparsification provided in the embodiment of the present application is shown as follows: Figure 5 As shown in the figure, before model training, the Whisper model's Transformer architecture is sparsified using a magnitude-based pruning algorithm. The importance of neurons or connections is regularly assessed, and sparse regularization techniques are used to add a sparsity penalty term to the loss function, enabling the model to automatically learn a sparser but effective parameter representation during training. A dynamic computation module is added after the model's input layer. This module is responsible for determining whether to skip certain complex computation layers or modules based on complexity metrics of the input speech data (such as speech duration, speaking rate, and vocabulary diversity).
[0097] S3. Model training: Train the constructed model using training data to obtain a training model;
[0098] In some embodiments, S3 includes at least one of the following:
[0099] In a distributed training environment, training data is dynamically allocated to different computing nodes through an adaptive data partitioner to perform distributed training of the model;
[0100] During the transmission of model parameters between different computing nodes, model compression technology is used to compress the parameters in real time;
[0101] During the model training process, the model is compressed and evaluated, and the compression strategy is dynamically adjusted based on the compressed model performance and resource consumption.
[0102] In this application, distributed training model compression collaborative optimization is adopted: adaptive data partitioner and model compression technology are used to improve distributed training efficiency and reduce resource consumption.
[0103] Specifically, this application adopts a distributed training model compression collaborative optimization model: in a distributed training environment, an adaptive data partitioner is designed. This partitioner dynamically allocates training data based on the performance of different computing nodes (such as computing speed, memory size, etc.) and network bandwidth. This ensures balanced computing load on each node and efficient data transmission, improving the overall efficiency of distributed training.
[0104] In addition, during the model parameter transmission process, model compression technologies such as quantization and Huffman coding are used to compress the parameters in real time, reducing the amount of data transmitted over the network. After the computing nodes receive the parameters, they are decompressed and calculated.
[0105] In addition, during the model training process, the model is compressed and evaluated regularly, and the compression strategy is dynamically adjusted based on the performance and resource consumption of the compressed model to minimize resource consumption and improve the efficiency of distributed training while ensuring model performance.
[0106] In some embodiments, S3 further includes:
[0107] After the model training iteration, the model stability is evaluated based on the uncertainty index, and the model hyperparameters are optimized according to the evaluation results, wherein the uncertainty index includes at least one of the prediction entropy and the prediction variance.
[0108] In this application, hyperparameter optimization based on uncertainty sampling is performed: uncertainty indicators such as prediction entropy and prediction variance are used to evaluate model stability, and hyperparameters are adjusted through Bayesian optimization.
[0109] Specifically, Figure 6 A schematic diagram of hyperparameter optimization based on uncertainty sampling provided in an embodiment of the present application is shown in FIG. Figure 6 As shown, this application adopts a hyperparameter optimization strategy based on uncertainty sampling: after each training iteration, the model is used to make multiple predictions on a small amount of validation data, and the predicted entropy and prediction variance of the prediction results are calculated. Based on these uncertainty indicators, the stability and generalization ability of the model under the current hyperparameter settings are evaluated. The uncertainty reduction term is added to the acquisition function of Bayesian optimization, so that hyperparameter adjustment not only considers improving model performance but also reducing model uncertainty.
[0110] S4. Model evaluation: Evaluate the training model through multi-dimensional indicators to obtain evaluation results.
[0111] In some embodiments, S4 includes:
[0112] The training model is evaluated using a multi-dimensional indicator through a multi-indicator evaluator to obtain an evaluation result, wherein the multi-dimensional indicator includes at least one of speech recognition accuracy, recall rate, model calculation efficiency, memory usage, and robustness.
[0113] In this application, a multi-index comprehensive evaluation is adopted: in addition to indicators such as recognition accuracy, multi-dimensional indicators such as the model's computational efficiency, memory usage, and robustness are also evaluated to ensure the comprehensive performance of the model.
[0114] Specifically, we built a multi-metric evaluator that, in addition to traditional speech recognition metrics like accuracy and recall, comprehensively considers multiple dimensions, including the model's computational efficiency, memory usage, robustness to varying noise levels and speaker diversity. During model training, we regularly use the multi-metric evaluator to evaluate the model and collect performance data on various metrics.
[0115] For ease of understanding, the following is a detailed description of the formulas for each part of this application model:
[0116] 1. Multi-source information fusion correction label and multi-level feature adaptive fusion
[0117] Let the input speech data be x, and its corresponding original annotated phoneme label be y original The rule set in the linguistic rule base is denoted as R = {r1, r2, ..., r m}, each rule r i It can be expressed as a function r i (x), the output is the correct phoneme label corresponding to the rule or a Boolean value indicating whether the rule is matched. The speech corpus metadata is denoted as M, which can include speech feature statistics, etc. Let the prediction function of the phoneme label based on the metadata be P M (y|x) represents the probability of the phoneme label y corresponding to the speech data x according to the metadata M. The user feedback information is recorded as F, and the function that determines whether the label needs to be revised based on the user feedback is C F (y original ,x), if the user feedback indicates that the label needs to be revised, the output is 1, otherwise it is 0. For the linguistic rule base R, traverse each rule r i To check whether the speech data x matches the rules, if it matches, the label is modified according to the rules. The modified phoneme label y rule for:
[0118]
[0119] in, Indicates that the rule does not output a valid correction label (i.e., the speech data x does not match the correction condition of the rule). If there is a rule that can output a valid correction label, the label corrected by the rule is used, otherwise the original label remains unchanged. The label is corrected using the speech corpus metadata M, and the probability P of each phoneme label obtained based on the metadata is calculated. M (y|x), take the phoneme label with the highest probability as the corrected label y meta :
[0120]
[0121] According to the user feedback information F, it is determined whether the original label y original Make corrections, if C F (y original ,x)=1, the label can be modified by combining the metadata method, and the final modified label is y feedback :
[0122]
[0123] Based on the above labels modified based on linguistic rules, speech corpus metadata and user feedback information, the corresponding weights λ1, λ2, and λ3 are assigned to each modification result (satisfying λ1+λ2+λ3=1), and the final modified phoneme label y is obtained. final for:
[0124] y final =λ1y rule +λ2y feedback +λ3y original
[0125] Let the phoneme feature be F p , the rhythmic feature is F r , the vocal tract characteristics are F v , the fusion calculation in the l-th layer Transformer block is:
[0126]
[0127] in is the learnable weight parameter at layer l, It is the feature representation after corresponding pre-processing and preliminary transformation.
[0128] 2. Adversarial Learning Enhances Long-Tail Data and Sparsification of Dynamic Model Structures
[0129] Let the feature extractor be G and the discriminator be D. For the input speech data x, the feature extractor outputs the feature z=G(x), and the discriminator outputs the category prediction y d =D(z). The adversarial loss function is:
[0130] L adv =logD(G(x))+log(1-D(G(x perturbed )))
[0131] where x perturbed is a sample after a slight perturbation of x, which is used to enhance the robustness of the discriminator. At the same time, combined with the conventional speech recognition loss function L asr , the total loss function L = L asr +αL adv , α is the weight parameter to balance the two losses.
[0132] Let the weight of neuron i be w i , define the importance measure i i =|w i During the pruning process, neurons whose importance is greater than the threshold τ are retained, that is, if I i >τ, neuron i is retained, otherwise it is removed from the model. For dynamic calculation, let the speech complexity index be C(x), the computational adjustment function be F(C(x)), and the actual computational amount of the model is:
[0133] M=M full ×F(C(x))
[0134] Among them, M full is the complete computational cost of the model.
[0135] 3. Co-optimization of distributed training model compression and hyperparameter optimization based on uncertainty sampling
[0136] Assume that the total amount of training data is N, the number of computing nodes is K, and the performance index of node k is P k , computing speed, memory size, and network bandwidth are B k,j , the bandwidth between node k and other nodes j.
[0137] The amount of data allocated to node k:
[0138]
[0139] where Δn k It is a correction item adjusted according to network bandwidth
[0140]
[0141] Used to compensate for data transmission delays caused by differences in network bandwidth.
[0142] For the hyperparameter combination θ, calculate the corresponding uncertainty reduction U(θ) and add the uncertainty reduction term to the acquisition function of Bayesian optimization:
[0143] a(θ)=βE[max(MM best ,0)|θ]+(1-β)U(θ)
[0144] Where β is a weight parameter that balances the improvement of model performance and the reduction of uncertainty.
[0145] Assume that the model is validating data x v The predicted probability distribution is p(y|x v ,θ), predicted entropy:
[0146] H(p)=∑ y -p(y|x v ,θ)logp(y|x v ,θ)
[0147] For the hyperparameter combination θ1 and θ2, the uncertainty reduction is:
[0148]
[0149] in and are the predicted probability distributions when the hyperparameters are θ1 and θ2, respectively.
[0150] 4. Comprehensive evaluation of multiple indicators
[0151] Define a comprehensive evaluation index I = α1A+α2R+α3E+α4M+α5Rb+…, where A is the accuracy, R is the recall, E is the computational efficiency, N is the memory usage, Rb is the robustness index, etc., α i is the weight coefficient of the corresponding indicator. Assume that the accuracy Recall Computational efficiency (T is the model training or inference time), memory usage M is measured in bytes, and the robustness index Rb can be calculated by the average accuracy on different noise levels and speaker sets. Comprehensive evaluation indicators:
[0152]
[0153] Among them A i is the accuracy under the i-th noise level or speaker condition, and N is the total number of test conditions.
[0154] 5. Bidirectional Fusion of Cross-Modal Attention Mechanism and Construction of Multimodal Shared Semantic Space
[0155] Assume that the speech-to-text attention is calculated as:
[0156]
[0157] The text-to-speech attention is calculated as:
[0158]
[0159] where Q s , K s 、V s is the linear transformation function of speech features, Q t , K t 、V t is the linear transformation function of text features, and d is the feature dimension. The fused speech features are:
[0160] F s,fused =W s1 F s +W s2 Attention t2s (F t ,F s )
[0161] The fused text features:
[0162] F t,fused =W t1 F t +W t2 Attention s2t (F s ,F t )
[0163] Where W s1 、W s2 、W t1 、W t2 is a learnable fusion weight parameter. Assume that the speech encoder is E s , the text encoder is E t , for speech data x s and text data x t , whose vector representations in the shared semantic space are v s =E s (x s ) and v t =E t (x t ). Define semantic similarity measure In the information interaction and enhancement stage, for the voice data x s , its enhanced feature representation:
[0164]
[0165] in is x s Text data with similar semantics The characteristics of ωi Based on semantic similarity The calculated weights are:
[0166]
[0167] The following are specific application examples of the technical solution of this application:
[0168] 1. Collect voice data and preprocess data
[0169] Voice samples were collected from multiple sources, including public speech datasets (such as the LibriSpeech dataset, which contains a large number of English speech samples from different speakers and accents), domain-specific voice recordings (such as doctor-patient conversations in the medical field and customer service consultations), and open source voice resources available online. In total, approximately 1,000 hours of voice samples were collected, covering a variety of language styles, scenarios, and speaker groups. This voice data was initially organized and stored by speaker, language type, duration, and other characteristics for subsequent processing and use.
[0170] For some speech data (about 100 hours), professional speech annotators were asked to perform phoneme-level annotation as the initial supervised annotation data. The annotation process follows the international phonetic symbol standard to ensure the accuracy of the annotation. At the same time, we collected external multi-source information, integrated authoritative linguistic grammar books on common languages such as English and Chinese, and pronunciation rule documents published by language research institutions, to form a linguistic rule library covering speech connection, tone change, stress rules, etc. Statistical information is extracted from the existing large-scale speech corpus, such as the frequency of different phonemes in various contexts, the mean and variance of acoustic features of different speech segments, etc., and stored as metadata to assist in judging the rationality of phoneme labels.
[0171] For each speech data entry, the Whisper model's standard audio preprocessing process is first performed, resampling the speech data to 16,000 Hz. An 80-channel log-mel spectrogram representation is then computed over a 25-millisecond window with a 10-millisecond step, yielding the audio's spectral features. During the initial annotation phase, the annotated phoneme labels are verified and corrected using the previously constructed multi-source information fusion and label correction module.
[0172] 2. Encoder parameter adjustment
[0173] Multi-level feature fusion is performed within the Transformer layers within the encoder. At lower Transformer layers, phoneme features and prosodic features are first fused using a weighted summation method. The weights are initially set to 0.5 each and dynamically adjusted during training using an adaptive adjustment mechanism. At higher layers, these fused features are further fused with vocal tract features. The adaptive adjustment mechanism automatically increases the weight of prosodic features to 0.6, phoneme features to 0.3, and vocal tract features to 0.1.
[0174] Long-tail data is divided into groups based on similarity in speech features, including vowel pronunciation characteristics and the number of tones. During training, the discriminator network is constructed to form an adversarial structure with the encoder. The feature vectors output by the encoder are fed into the discriminator. Before model training, the Whisper model's Transformer architecture is sparsified using an amplitude-based pruning algorithm. An initial pruning threshold is set based on the absolute value of the parameter, and neuronal connections with an absolute value less than 0.01 are considered unimportant and pruned. During training, the importance of neurons or connections is reassessed after every 10 training epochs. A sparsity penalty is added to the loss function, proportional to the number of pruned parameters. For the dynamic computation module, speech complexity metrics are calculated for each input of speech data. Speech tasks are classified as simple if they are less than 3 seconds long, spoken at a fast rate, and judged using simple vocabulary statistics. Speech tasks are classified as simple if they are longer than 10 seconds long, spoken at a moderate rate, and have a rich vocabulary.
[0175] 3. Model training part
[0176] In a distributed training environment, training data is first allocated to each node based on its performance and network bandwidth using an adaptive data partitioner. Node 1, with its faster computational speed but average network bandwidth, receives approximately 40% of the total data volume. Node 2, with its higher network bandwidth but relatively smaller memory footprint, receives approximately 30% of the total data volume, and so on. This ensures balanced computational load and efficient data transmission across each node. During model parameter transmission, quantization technology is used to convert model parameters from 32-bit floating-point representation to 8-bit integer representation for compression, while Huffman coding is used to further reduce the amount of data. For example, for the weight parameters of a Transformer layer, quantization reduces the amount of data to approximately one-quarter of its original size, and Huffman coding further reduces the amount of data transmitted.
[0177] During model training, approximately 10% of the total data set was used as a validation set. The model was used to perform multiple predictions on this validation data, with five predictions performed after each training epoch. The predicted entropy of the predictions was calculated using an initial learning rate of 0.001 and a batch size of 32.
[0178] Add uncertainty reduction terms to the Bayesian optimization acquisition function, for example, by setting a weight parameter that balances model performance improvement with uncertainty reduction. Based on the acquisition function, select the next hyperparameter combination that is likely to improve performance and reduce uncertainty for evaluation. Adjusting the learning rate to 0.0008 and the batch size to 64 both improves the model's accuracy on the validation set and reduces prediction entropy, thereby reducing the model's uncertainty. Continue training with this new set of hyperparameters, and through continuous iterative optimization, improve the model's overall performance.
[0179] 4. Decoder parameter adjustment
[0180] A bidirectional cross-modal attention mechanism is enabled in an intermediate layer of the decoder. This emphasis is reflected when text features focus on intonation changes in speech through the text-to-speech attention mechanism, making the generated text more consistent with the semantics and emotion of the speech. During training, a dynamic weight adjuster automatically adjusts the attention weights in the bidirectional fusion process based on real-time feature changes in the speech and text data. In the multimodal encoder, the speech feature encoder converts speech data features into vector representations in a shared semantic space, and the text feature encoder also maps word vector representations of text into the same shared semantic space. In the semantic information interaction module, a nearest neighbor search algorithm is used to identify semantically similar text data when processing speech data, and the feature information of this text data is integrated into the speech features.
[0181] This application provides a speech recognition model training method with the following characteristics and
[0182] Beneficial effects:
[0183] 1. Improved Whisper encoder
[0184] The encoder front-end adds a prosody and vocal tract feature extraction branch, and the Transformer layer integrates phoneme, prosody, and vocal tract features in a layered manner. Phoneme and prosody features are first integrated at the lower layers, and then at the higher layers, vocal tract features are combined. Weights are adaptively adjusted based on speech characteristics to construct an accurate speech representation.
[0185] 2. Internal structure of the training module
[0186] Distributed training adaptively partitions data based on node performance and bandwidth, quantizing and compressing data during transmission, and dynamically optimizing Huffman encoding parameters during training. Hyperparameter optimization utilizes uncertainty metrics such as prediction entropy, comprehensively considering performance and uncertainty in Bayesian optimization to select optimal hyperparameters.
[0187] 3. Comprehensive evaluation of multiple indicators
[0188] It covers multi-dimensional indicators such as accuracy, recall rate, computational efficiency, memory usage, noise robustness, and speaker diversity adaptability, and is regularly evaluated during training to provide a comprehensive basis for model optimization.
[0189] 4. Improved Whisper decoder
[0190] A bidirectional cross-modal attention layer enables two-way attention between speech and text, and a dynamic weight adjuster updates weights based on data changes. A multimodal encoder constructs a shared semantic space, and a semantic information interaction module fuses semantic information through nearest neighbor search, improving multimodal task performance.
[0191] It should be understood that, although the various steps in the flowcharts of the above embodiments are shown in sequence as indicated by the arrows, these steps are not necessarily performed in the order indicated by the arrows. Unless otherwise specified herein, there is no strict order restriction on the execution of these steps, and they may be performed in other orders. Moreover, at least a portion of the steps in the figure may include multiple sub-steps or multiple stages, and these sub-steps or stages are not necessarily performed at the same time, but may be performed at different times, and their execution order is not necessarily sequential, but may be performed in turn or alternately with other steps or at least a portion of sub-steps or stages of other steps.
[0192] Figure 7 A schematic diagram of a speech recognition model training device provided in an embodiment of the present application is shown in FIG. Figure 7 As shown, the present application provides a speech recognition model training device, the device comprising:
[0193] The data preparation module 11 is configured to prepare data by performing multi-source information fusion correction and feature extraction processing through an external knowledge base to obtain training data;
[0194] A model building module 12, configured to build a model, including: an improved Whisper encoder, a discriminator network, an improved Whisper decoder, and a multimodal encoder;
[0195] The Transformer layer within the improved Whisper encoder includes a multi-level fusion module for fusing different features contained in the training data and adjusting the model's weight distribution for different features;
[0196] The discriminator network and the improved Whisper encoder form an adversarial structure;
[0197] The improved Whisper decoder internally includes a bidirectional cross-modal attention layer for performing data attention calculation, and the bidirectional cross-modal attention layer includes a dynamic weight adjuster for dynamically adjusting the attention weight in the bidirectional fusion process;
[0198] The multimodal encoder is used to encode different features into a shared semantic space, wherein the shared semantic space includes a semantic information interaction module for fusing neighboring data features;
[0199] The model training module 13 is configured to perform model training: training the constructed model using training data to obtain a training model;
[0200] The model evaluation module 14 is configured to perform model evaluation: the training model is evaluated through multi-dimensional indicators to obtain evaluation results.
[0201] Regarding the limitation of the speech recognition model training device, please refer to the limitation of the speech recognition model training method in the above embodiments of this application, which will not be repeated in this embodiment.
[0202] Figure 8 Another schematic diagram of the speech recognition model training device provided in the embodiment of the present application is as follows Figure 8 As shown, the device includes a memory 22 and a processor 21, the memory stores a computer program, and the processor is configured to run the computer program to execute the methods in the above embodiments of the present application.
[0203] The memory is connected to the processor, the memory may be a flash memory, a read-only memory or other memory, and the processor may be a central processing unit or a single-chip microcomputer.
[0204] In some embodiments, the present application provides a computer-readable storage medium having a computer program stored thereon. When the computer program is executed by a processor, the methods in the above embodiments of the present application are implemented.
[0205] The computer-readable storage medium includes volatile or non-volatile, removable or non-removable media implemented in any method or technology for storing information (such as computer-readable instructions, data structures, computer program modules or other data). Computer-readable storage media include, but are not limited to, RAM (Random Access Memory), ROM (Read-Only Memory), EEPROM (Electrically Erasable Programmable read only memory), flash memory or other memory technology, CD-ROM (Compact Disc Read-Only Memory), digital versatile disks (DVD) or other optical disk storage, magnetic cassettes, magnetic tape, magnetic disk storage or other magnetic storage devices, or any other medium that can be used to store the desired information and can be accessed by a computer.
[0206] It is understood that the above embodiments are merely exemplary embodiments for illustrating the principles of the present application, and the present application is not limited thereto. Those skilled in the art may make various modifications and improvements without departing from the spirit and substance of the present application, and such modifications and improvements are also considered to be within the scope of protection of the present application.
Claims
1. A speech recognition model training method, characterized in that: The method comprises: S1. Data preparation: Multi-source information fusion and feature extraction are performed through an external knowledge base to obtain training data; S2. Model construction, including: improved Whisper encoder, discriminator network, improved Whisper decoder and multimodal encoder; The Transformer layer within the improved Whisper encoder includes a multi-level fusion module for fusing different features contained in the training data and adjusting the model's weight distribution for different features; The discriminator network and the improved Whisper encoder form an adversarial structure; The improved Whisper decoder internally includes a bidirectional cross-modal attention layer for performing data attention calculation, and the bidirectional cross-modal attention layer includes a dynamic weight adjuster for dynamically adjusting the attention weight in the bidirectional fusion process; The multimodal encoder is used to encode different features into a shared semantic space, wherein the shared semantic space includes a semantic information interaction module for fusing neighboring data features; S3. Model training: Train the constructed model using training data to obtain a training model; S4. Model evaluation: Evaluate the training model through multi-dimensional indicators to obtain evaluation results.
2. The speech recognition model training method according to claim 1, characterized in that S1, including: Verify and correct the initial text data and audio data through an external knowledge base to obtain corrected data; Feature extraction processing is performed on the corrected data to obtain the phoneme features, prosodic features and vocal tract features of the speech signal.
3. The speech recognition model training method according to claim 1, characterized in that In S2, the input of the discriminator network is the feature vector output by the improved Whisper encoder, and the discriminator network is used to determine the data category of the feature vector, wherein the data category of the feature vector includes a long-tail data category and a mainstream data category.
4. The speech recognition model training method according to claim 1, characterized in that Before S3, it also included: A magnitude-based pruning algorithm is used to sparsify the Transformer architecture of the Whisper model.
5. The speech recognition model training method according to claim 1, characterized in that: S3, including at least one of the following: In a distributed training environment, training data is dynamically allocated to different computing nodes through an adaptive data partitioner to perform distributed training of the model; During the transmission of model parameters between different computing nodes, model compression technology is used to compress the parameters in real time; During the model training process, the model is compressed and evaluated, and the compression strategy is dynamically adjusted based on the compressed model performance and resource consumption.
6. The speech recognition model training method according to claim 5, characterized in that: S3, also includes: After the model training iteration, the model stability is evaluated based on the uncertainty index, and the model hyperparameters are optimized according to the evaluation results, wherein the uncertainty index includes at least one of the prediction entropy and the prediction variance.
7. The speech recognition model training method according to claim 1, characterized in that: S4, including: The training model is evaluated using a multi-dimensional indicator through a multi-indicator evaluator to obtain an evaluation result, wherein the multi-dimensional indicator includes at least one of speech recognition accuracy, recall rate, model calculation efficiency, memory usage, and robustness.
8. A speech recognition model training device, characterized in that: The device comprises: The data preparation module is configured to prepare data by performing multi-source information fusion correction and feature extraction processing through an external knowledge base to obtain training data; A model building module configured to build a model including: an improved Whisper encoder, a discriminator network, an improved Whisper decoder, and a multimodal encoder; The Transformer layer within the improved Whisper encoder includes a multi-level fusion module for fusing different features contained in the training data and adjusting the model's weight distribution for different features; The discriminator network and the improved Whisper encoder form an adversarial structure; The improved Whisper decoder internally includes a bidirectional cross-modal attention layer for performing data attention calculation, and the bidirectional cross-modal attention layer includes a dynamic weight adjuster for dynamically adjusting the attention weight in the bidirectional fusion process; The multimodal encoder is used to encode different features into a shared semantic space, wherein the shared semantic space includes a semantic information interaction module for fusing neighboring data features; A model training module is configured to perform model training: the constructed model is trained using training data to obtain a training model; The model evaluation module is configured to perform model evaluation: the training model is evaluated through multi-dimensional indicators to obtain evaluation results.
9. A speech recognition model training device, characterized in that: It includes a memory and a processor, wherein the memory stores a computer program, and the processor is configured to run the computer program to implement the speech recognition model training method as described in any one of claims 1 to 7.
10. A computer-readable storage medium, characterized in that The computer-readable storage medium stores a computer program, which, when executed by a processor, implements the speech recognition model training method according to any one of claims 1 to 7.
Citation Information
Cited By
Intelligent questioning and answering method and equipment for legal knowledge and medium
CN121029961A