Voice large model training method and device, equipment and medium
By training a large speech model and utilizing entropy-weighted cross-entropy loss function and perturbation reinforcement learning, the large language model module was optimized, solving the problem of poor speech recognition performance in teaching scenarios and improving the accuracy and stability of teaching quality assessment.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-12-29
- Publication Date
- 2026-04-10
AI Technical Summary
How to improve speech recognition performance, especially in teaching scenarios such as flipped classrooms and teaching quality inspection, by performing speech recognition on teaching speech data to improve teaching quality.
The large speech model is used for training, including a text encoder, an audio encoder, a bridge, and a large language model module. Through multiple rounds of training, the parameters of the large language model module and speech recognition labels are optimized using the entropy-weighted cross-entropy loss function and perturbation reinforcement learning, thereby improving the accuracy of the predicted probability distribution.
It improves the accuracy and stability of speech recognition, enhances the effectiveness of teaching quality assessment, and is suitable for speech recognition applications in electronic devices such as personal computers and mobile terminals.
Smart Images

Figure CN121438813B_ABST
Abstract
Description
TECHNICAL FIELD
[0001] One or more embodiments of the present disclosure relate to a speech large model training method, a speech large model training apparatus, an electronic device, and a non-transitory computer-readable storage medium. BACKGROUND
[0002] In some scenarios, there is a demand for speech recognition, for example, in flipped classroom, teaching quality inspection and other teaching scenarios, by performing speech recognition on teaching speech data, combined with the speech recognition result, the teaching quality can be improved accordingly.
[0003] How to improve the speech recognition effect becomes a problem to be solved. SUMMARY
[0004] This summary is provided to introduce a selection of concepts, which are further described below in the detailed description. This summary is not intended to identify key or essential features of the claimed subject matter, nor is it used to limit the scope of the claimed subject matter's scope.
[0005] At least one embodiment of the present disclosure provides a speech large model training method, the speech large model comprising a text encoder, an audio encoder, a bridge and a large language model module, the text encoder is used to generate text tokens based on the text corresponding to the speech data, the audio encoder is used to generate audio tokens based on the speech data, the bridge is used to generate fusion tokens based on the text tokens and the audio tokens, and the large language model module is used to generate a speech recognition result based on the fusion tokens, the method comprising: obtaining a training speech data set, wherein the training speech data set comprises a plurality of training speech data, and the plurality of training speech data correspond to speech recognition labels respectively; performing a plurality of rounds of training processes on the large language model module, each round of the training process comprising the following steps: determining a training speech data subset of the training process from the training speech data set, and determining a true probability distribution of the training speech data subset according to the speech recognition labels of the training speech data subset; inputting the training speech data subset into the speech large model to obtain a predicted probability distribution of the training speech data subset output by the large language model module in the speech large model; determining an entropy-weighted cross-entropy loss function according to the predicted probability distribution and the true probability distribution, wherein the entropy weight of the entropy-weighted cross-entropy loss function is determined based on the distribution entropy of the predicted probability distribution at the current time; updating the parameters of the large language model module and the speech recognition labels of the training speech data subset with the aim of minimizing the entropy-weighted cross-entropy loss function.
[0006] The training device of the speech large model provided in at least one embodiment of the present disclosure includes a text encoder, an audio encoder, a bridge, and a large language model module. The text encoder is configured to generate text tokens based on text corresponding to speech data. The audio encoder is configured to generate audio tokens based on the speech data. The bridge is configured to generate fusion tokens based on the text tokens and the audio tokens. The large language model module is configured to generate a speech recognition result based on the fusion tokens. The training device includes an obtaining module configured to obtain a training speech data set. The training speech data set includes a plurality of training speech data, and the plurality of training speech data correspond to speech recognition labels respectively. The training device also includes a training module configured to perform a plurality of training processes on the large language model module. Each training process includes the following steps: determining a training speech data subset for the training process from the training speech data set, and determining a true probability distribution of the training speech data subset based on speech recognition labels of the training speech data subset. The training speech data subset is input into the speech large model, and a predicted probability distribution of the training speech data subset output by the large language model module in the speech large model is obtained. An entropy-weighted cross-entropy loss function is determined based on the predicted probability distribution and the true probability distribution. The entropy weight of the entropy-weighted cross-entropy loss function is determined based on a distribution entropy of the predicted probability distribution at a current time. The parameters of the large language model module and the speech recognition labels of the training speech data subset are updated to minimize the entropy-weighted cross-entropy loss function.
[0007] The electronic device provided in at least one embodiment of the present disclosure includes at least one processor and a memory connected to the at least one processor in communication. The memory stores instructions executable by the at least one processor. The instructions are executed by the at least one processor to enable the at least one processor to perform the training method of the speech large model provided in at least one embodiment of the present disclosure.
[0008] The non-transitory computer-readable storage medium provided in at least one embodiment of the present disclosure stores computer instructions for causing a computer to perform the training method of the speech large model provided in at least one embodiment of the present disclosure.
[0009] The computer program product provided in at least one embodiment of the present disclosure includes a computer program. When the computer program is executed by a processor, the training method of the speech large model provided in at least one embodiment of the present disclosure is implemented. BRIEF DESCRIPTION OF DRAWINGS
[0010] In order to more clearly illustrate the technical solutions of one or more embodiments of the present disclosure, the drawings of the embodiments will be briefly introduced as follows. Obviously, the drawings below only relate to some embodiments of the present disclosure and not limit the present disclosure.
[0011] Figure 1 A structural schematic diagram of a voice large model provided by at least one embodiment of the present disclosure is illustratively shown;
[0012] Figure 2 A schematic diagram of a training process of a voice large model provided by at least one embodiment of the present disclosure is illustratively shown;
[0013] Figure 3 A flowchart of a training method of a voice large model provided by at least one embodiment of the present disclosure is illustratively shown;
[0014] Figure 4 A flowchart of determining a contrastive learning loss function provided by at least one embodiment of the present disclosure is illustratively shown;
[0015] Figure 5 A schematic diagram of a search tree provided by at least one embodiment of the present disclosure is illustratively shown;
[0016] Figure 6 A structural schematic diagram of a training device of a voice large model provided by at least one embodiment of the present disclosure is illustratively shown; and
[0017] Figure 7 A structural schematic diagram of an electronic device for implementing a training function of a voice large model provided by at least one embodiment of the present disclosure is illustratively shown. DETAILED DESCRIPTION
[0018] In order to make the purpose, technical solutions and advantages of one or more embodiments of the present disclosure clearer, the technical solutions of one or more embodiments of the present disclosure will be described clearly and completely below with reference to the drawings of one or more embodiments of the present disclosure. Obviously, the described embodiments are part of the embodiments of the present disclosure, rather than all the embodiments. Based on the described one or more embodiments of the present disclosure, all other embodiments obtained by those of ordinary skill in the art without any creative effort fall within the scope of protection of the present disclosure.
[0019] Unless otherwise defined, technical terms or scientific terms used in the present disclosure shall have the same meaning as commonly understood by one of ordinary skill in the art to which the present disclosure belongs. The terms "first", "second", and similar terms in the present disclosure do not denote any order, quantity, or importance, but are used to distinguish different components. The terms "include", "contain", and similar terms mean that the elements or objects before the terms encompass the elements or objects listed after the terms and their equivalents, and do not exclude other elements or objects. The terms "connect" or "connected" and similar terms are not limited to physical or mechanical connections, but can include electrical connections, whether direct or indirect. The terms "upper", "lower", "left", "right", and the like are used only to indicate relative positional relationships, and when the absolute positions of the described objects change, the relative positional relationships can also change accordingly.
[0020] In order to keep the following description of one or more embodiments of the present disclosure clear and concise, the present disclosure omits detailed descriptions of some known functions and known components.
[0021] In flipped classroom, teaching quality inspection and other teaching scenarios, there is often a need for speech recognition, for example, by converting speech data generated in a teaching scenario into a text recognition result, using the text recognition result as subtitle information, adding it to a teaching video, or testing the teaching effect according to the text recognition result to improve the teaching quality.
[0022] In some examples, speech recognition can be performed using a speech large model, and it is particularly important to train the speech large model to improve the speech recognition effect.
[0023] In view of this, at least one embodiment of the present disclosure provides a speech large model training method. The speech large model includes a text encoder, an audio encoder, a bridge, and a large language model module. The text encoder is configured to generate text tokens based on text corresponding to speech data. The audio encoder is configured to generate audio tokens based on the speech data. The bridge is configured to generate fusion tokens based on the text tokens and the audio tokens. The large language model module is configured to generate a speech recognition result based on the fusion tokens. The speech large model training method includes: obtaining a training speech data set, the training speech data set including a plurality of training speech data, the plurality of training speech data each corresponding to a speech recognition label; performing a plurality of training processes on the large language model module, each training process including the following steps: determining a training speech data subset for the training process from the training speech data set, and determining a true probability distribution of the training speech data subset based on the speech recognition labels of the training speech data subset; inputting the training speech data subset into the speech large model; obtaining a predicted probability distribution of the training speech data subset output by the large language model module in the speech large model; determining an entropy-weighted cross-entropy loss function based on the predicted probability distribution and the true probability distribution, an entropy weight of the entropy-weighted cross-entropy loss function being determined based on a distribution entropy of the predicted probability distribution at a current time; and updating parameters of the large language model module and the speech recognition labels of the training speech data subset, with the goal of minimizing the entropy-weighted cross-entropy loss function.
[0024] In one or more embodiments of the present disclosure, the large language model module in the speech large model is trained. In each training process, a predicted probability distribution is generated using the speech large model, an entropy-weighted cross-entropy loss function is determined based on a true probability distribution and the predicted probability distribution, with the distribution entropy of the predicted probability distribution at a current time, to update parameters of the large language model module and speech recognition labels generated by the speech large model. By adding an entropy weight in the entropy-weighted cross-entropy loss function, the uncertainty problem of the predicted probability distribution is solved to some extent, and a self-feedback signal based on the distribution entropy of the predicted probability distribution is used for iteration to continuously improve the speech recognition effect of the speech large model.
[0025] One or more embodiments of the present disclosure also provide a speech large model training apparatus, an electronic device, a computer-readable storage medium, and a computer program product. The speech large model training method described above can be applied to the speech large model training apparatus provided by one or more embodiments of the present disclosure, which can be configured on an electronic device. The electronic device can be a personal computer, a mobile terminal, etc., and the mobile terminal can be a mobile phone, a headset, a tablet computer, a vehicle-mounted device, etc.
[0026] One or more embodiments of the present disclosure will be described in detail below with reference to the accompanying drawings, but the present disclosure is not limited to these specific embodiments.
[0027] Figure 1 A structural diagram of a speech large model is shown.
[0028] In one or more embodiments of the present disclosure, the speech large model can include at least a text encoder, an audio encoder, a bridge, and a large language model (LLM) module. The text encoder can be configured to generate text tokens based on text corresponding to speech data. The audio encoder can be configured to generate audio tokens based on the speech data. The bridge can be configured to generate fusion tokens based on the text tokens and the audio tokens. The LLM module can be configured to generate speech recognition results based on the fusion tokens.
[0029] Referring to FIG. 1, a speech large model can include a text encoder 102, an audio encoder 104, a bridge 106, and a large language model (LLM) module 108. Figure 1 In some embodiments, an unlabeled speech training set 101, for example, a set of audio files, is input into the text encoder 102 and the audio encoder 104 of the speech large model. The text encoder 102 encodes the text of the unlabeled speech training set 101 to obtain text tokens 103. The audio encoder 104 encodes the unlabeled speech training set 101 to obtain audio tokens 105, for example, acoustic feature sequences. The text encoder 102 encodes the text of the unlabeled speech training set 101 to obtain text tokens 103. The audio encoder 104 encodes the unlabeled speech training set 101 to obtain audio tokens 105, for example, acoustic feature sequences. The text encoder 102 encodes the text of the unlabeled speech training set 101 to obtain text tokens 103. The audio encoder 104 encodes the unlabeled speech training set 101 to obtain audio tokens 105, for example, acoustic feature sequences. The text encoder 102 encodes the text of the unlabeled speech training set 101 to obtain text tokens 103. The audio encoder 104 encodes the unlabeled speech training set 101 to obtain audio tokens 105, for example, acoustic feature sequences.
[0030] The bridge 106 fuses the text tokens 103 and the audio tokens 105 to obtain fusion tokens 107. The LLM module 108 processes the fusion tokens to obtain speech recognition results, i.e., a labeled speech training set 109.
[0031] The one or more embodiments of the present disclosure do not limit the manner of obtaining the initial parameters of the speech large model. For example, the audio encoder 104 can adopt an open-source whisper encoder architecture, perform unsupervised training on the audio encoder 104, freeze the parameters of other modules in the speech large model, perform scene training on the audio encoder 104 using a specific speech data set in a specific scene (for example, a teaching scene), so that the audio encoder 104 can better extract audio features related to the task in the specific scene, improve the representation ability of the specific speech data in the specific scene, and obtain the initial parameters of the audio encoder 104; for another example, the bridge 106 can use a 2-layer multilayer perceptron (MLP) neural network, freeze the parameters of other modules in the speech large model, and only train the bridge 106 to obtain the initial parameters of the bridge 106, so that the bridge 106 realizes alignment between the audio token and the text token, and learns to effectively convert the speech representation of the audio encoder 104 into a form suitable for processing by the large language model module 108; for another example, the initial parameters of the large language model module 108 can be a third-party open-source large language model.
[0032] It should be noted that in one or more embodiments of the present disclosure, the above-mentioned text token, audio token and fusion token are in the form of embedding.
[0033] Figure 2 An illustrative diagram of a training process of a speech large model provided by at least one embodiment of the present disclosure is shown.
[0034] As shown in Figure 2 , the training process of the speech large model can include multiple rounds of training processes, and each round of training process can include three stages. Stage 1: probability control sampling 202 is performed on a training speech data set 201, for example, a training speech data set 201, to obtain a training speech data subset 203, for example, a training speech data subset 203. , , representing training speech data and corresponding speech recognition labels; stage 2: self-feedback signal-based supervised iterative training 204; stage 3: perturbation reinforcement learning sampling information-based supervised training 205.
[0035] After completing the multiple rounds of training processes, a small amount of scene data-based supervised fine-tuning training 206 can also be performed using a labeled scene data set, for example, a labeled scene data set 206.
[0036] The following will introduce each step in the training process of the speech large model.
[0037] Figure 3 A flowchart schematically showing a method for training a large speech model according to at least one embodiment of the present disclosure is shown in FIG. 3, which specifically includes the following steps. Figure 3
[0038] Step S301: Obtain a training speech data set.
[0039] In one or more embodiments of the present disclosure, the training speech data set can include a plurality of training speech data, and each of the plurality of training speech data corresponds to a speech recognition label.
[0040] For example, as previously described in relation to Figure 1 , the training speech data set without labels is input into the initial large speech model, and the initial large speech model is used to generate speech recognition labels to obtain a training speech data set with labels.
[0041] Since the speech recognition labels are generated by the initial large speech model, the speech recognition labels can have inaccuracy problems, and therefore the speech recognition labels can also be understood as pseudo labels. In the subsequent multiple rounds of training process, the speech recognition labels can be used as "real" labels for supervised training, and the accuracy of the pseudo labels can be gradually improved in the multiple rounds of training process.
[0042] Step S302: Perform a multiple-round training process on the large language model module.
[0043] The multiple-round training process performed on the large language model module corresponds to the probability control sampling 202 and the supervised iterative training based on the self-feedback signal 204 in Figure 2 In some embodiments, the multiple-round training process performed on the large language model module can also correspond to the supervised training based on perturbation reinforcement learning sampling information 205, which will be described below.
[0044] In each round of training process, steps S3021 to S3024 are included.
[0045] Step S3021: From the training speech data set, determine a training speech data subset for the training process, and according to the speech recognition labels of the training speech data subset, determine a real probability distribution of the training speech data subset.
[0046] In each round of training process, the training speech data subset is determined, which can be understood as the training data in this round of training process.
[0047] In some possible implementations, the training speech data set is sampled to determine the training speech data subset for the training process, and the intersection probability between the training speech data subset and the training speech data subset in the historical training process in the training speech data set is greater than a first threshold.
[0048] In other words, as mentioned above... Figure 2 As described in at least one embodiment of this disclosure, probability-controlled sampling 202 can be performed, in which the sampling probability is adjusted in each round of training so that the intersection probability between the newly sampled subset of training speech data and the historically sampled subset of training speech data is greater than a first threshold.
[0049] In this way, during each training round, a subset of training speech data is randomly sampled for optimization, while ensuring cross-reference with historical training speech data subsets. This effectively ensures that the iterative effect of the training speech data subset does not decrease during the current training round.
[0050] For example, the training speech dataset is The training speech data subset in each round of training is .
[0051] Because probabilistic control sampling follows a sequential principle (i.e., sampling is performed in chronological order), the subset of training speech data sampled in the current training round differs from a subset of training speech data sampled in a previous round. The probability that there is no intersection between them is: Therefore, the probability that the subset of training speech data sampled in the current training round intersects with at least one subset of training speech data sampled in the past is: .
[0052] Thus, the sampling step can be: calculating the union of subsets of historically sampled training speech data. ,from Sampling is performed using probability p, and the data is divided from the training speech dataset. The remaining set is sampled with probability 1-p, where p satisfies the following condition:
[0053]
[0054] By adjusting p to control the intersection probability, a subset of the training speech data is obtained. This makes the intersection probability greater than the first threshold.
[0055] In the first round of training, the speech recognition labels of the training speech data subset can be the speech recognition results generated by the initial large speech model. In subsequent training rounds, the speech recognition labels of the training speech data subset can be the updated speech recognition labels from the previous training round.
[0056] The true probability distribution of a subset of training speech data can be understood as the true probability value assigned to each speech recognition label in the subset. For example, the true probability distribution of a single training speech data point at a given time step can be: V represents the total number of probability distribution tags (i.e., tokens) output by the large language model module.
[0057] Step S3022: Input the training speech data subset into the speech model and obtain the predicted probability distribution of the training speech data subset output by the large language model module in the speech model.
[0058] In each training round, each piece of training speech data from the training speech data subset is input back into the speech big model. The current speech big model (i.e., the speech big model updated after the training process before this round) is used to perform model inference and obtain the predicted probability distribution of the training speech data subset output by the big language model module.
[0059] The prediction probability distribution of a subset of training speech data can be understood as the prediction probability value assigned to each speech recognition label in the subset. For example, the prediction probability distribution of a single training speech data point at a given time step can be: .
[0060] Step S3023: Determine the entropy-weighted cross-entropy loss function based on the predicted probability distribution and the true probability distribution.
[0061] In the cross-entropy loss function without added entropy weights, the cross-entropy loss function can be:
[0062]
[0063] Where M represents a subset of training speech data The amount of training speech data, This indicates the speech recognition label of category v for the i-th training speech data. For a large speech model, predict the probability that the i-th training speech data belongs to class v. When the speech recognition labels are one-hot encoded, the cross-entropy loss function can be simplified to:
[0064]
[0065] in, It is the speech recognition label of the i-th training speech data.
[0066] However, in the multi-round training process, by analyzing the prediction probability distribution output by the large language model module in the speech large model at each moment, it is found that at different moments, even different prediction probability distributions will lead to the same cross-entropy loss function described above, but the learning difficulty of the large language model module is obviously different, that is, the cross-entropy loss function described above ignores the prediction uncertainty problem and does not consider the existence of probability unimodal distribution, probability multimodal distribution, probability average distribution and the like of the prediction probability distribution output by the large language model module in the speech large model.
[0067] Therefore, in one or more embodiments of the present disclosure, the cross-entropy loss function is improved and the entropy weight is increased. The entropy weight of the entropy-weighted cross-entropy loss function can be determined based on the distribution entropy of the prediction probability distribution at the current moment.
[0068] That is, for each moment, the distribution entropy of the prediction probability distribution at the current moment is considered to determine the corresponding entropy weight at the moment, and the entropy-weighted cross-entropy loss function is dynamically adjusted. In the training process, the prediction uncertainty problem is considered.
[0069] For example, the entropy-weighted cross-entropy loss function can be:
[0070]
[0071] wherein, is the entropy weight.
[0072] In some embodiments, the value of the entropy weight of the entropy-weighted cross-entropy function is determined according to the unimodality ( ) of the prediction probability distribution at the current moment and the probability ( ) of the target token to be predicted For example, the unimodality can be calculated using kurtosis or Gini sparsity coefficient. For example, the formula for calculating the unimodality based on kurtosis is:
[0073]
[0074] represents the probability of the i-th token in the prediction probability distribution, is the mean of the token probability in the prediction probability distribution, is the standard deviation of the token probability in the prediction probability distribution. The minus 3 in the above formula for calculating the unimodality based on the peak value is to make the kurtosis of the normal distribution 0.
[0075] For example, the formula for calculating the unimodality based on the Gini sparsity coefficient is:
[0076]
[0077] Both of the above two calculation methods need to be further normalized to Based on the above measurable calculation of unimodality based on kurtosis or Gini sparse coefficient, the weight calculation formula of cross entropy loss function is:
[0078]
[0079] k is a manually set correction coefficient, for example, , to ensure , , , so that , so that the accuracy is close to 0.5.
[0080] In some embodiments, according to the prediction probability distribution of the current moment, the distribution entropy of the prediction probability distribution of the current moment is determined, and in response to the distribution entropy of the prediction probability distribution of the current moment being less than a first entropy threshold, the entropy weight of the entropy-weighted cross entropy loss function is determined to be smaller, and the entropy weight of the entropy-weighted cross entropy loss function is determined to be larger; in response to the distribution entropy of the prediction probability distribution of the current moment being greater than or equal to the first entropy threshold, the entropy weight of the entropy-weighted cross entropy loss function is determined to be larger, and the entropy weight of the entropy-weighted cross entropy loss function is determined to be larger.
[0081] And the entropy weight of the entropy-weighted cross entropy loss function when the distribution entropy of the prediction probability distribution of the current moment is less than the first entropy threshold is greater than the entropy weight of the entropy-weighted cross entropy loss function when the distribution entropy of the prediction probability distribution of the current moment is greater than or equal to the first entropy threshold.
[0082] That is, when the distribution entropy of the prediction probability distribution of the current moment is small, the entropy weight is large, and decreases with the increase of the distribution entropy; when the distribution entropy of the prediction probability distribution of the current moment is large, the entropy weight is small, and increases with the increase of the distribution entropy.
[0083] For example, the distribution entropy of the prediction probability distribution of the current moment can be , and the entropy weight of the entropy-weighted cross entropy loss function can be:
[0084]
[0085] wherein, is an adjustable factor, is the distribution entropy threshold of the current moment, and , is the first entropy threshold.
[0086] In some embodiments, the distribution entropy of the prediction probability distribution of the current time is determined according to the prediction probability distribution of the current time, the entropy change of the prediction probability distribution of the current time and the adjacent time is determined according to the distribution entropy of the prediction probability distribution of the current time, the static entropy of the current time is determined according to the distribution entropy of the prediction probability distribution of the current time, the dynamic entropy of the current time is determined according to the entropy change of the prediction probability distribution of the current time and the adjacent time, and the entropy weight of the entropy-weighted cross-entropy loss function is determined according to the static entropy of the current time and the dynamic entropy of the current time.
[0087] The static entropy of the current time can be used to measure the uncertainty of the probability distribution of the current time, and the dynamic entropy of the current time can be used to measure the stability of the probability distribution of the current time.
[0088] That is, the entropy weight of the entropy-weighted cross-entropy loss function can be determined by the static entropy and the dynamic entropy of the current time, which combines the stability and uncertainty of the prediction probability distribution to determine the entropy weight, and then realizes the iteration of different intensities.
[0089] For example, the entropy weight of the entropy-weighted cross-entropy loss function can be:
[0090]
[0091] wherein, is the entropy change of the prediction probability distribution of the current time and the adjacent time, is the dynamic entropy of the current time, is the static entropy of the current time.
[0092] When , it indicates that the static entropy of the current time has high certainty, and the learning information is expanded; when , it indicates that the static entropy of the current time has high chaos, and the learning noise is compressed.
[0093] The greater the value is, the more intense the fluctuation of the prediction probability distribution is, and vice versa, if , it indicates that the training tends to be stable, , it indicates that the fluctuation is intense.
[0094] Step S3024: updating the parameters of the large language model module and the speech recognition labels of the training speech data subset with the goal of minimizing the entropy-weighted cross-entropy loss function.
[0095] The parameters of the large language model module and the speech recognition labels of the training speech data subset are iterated once by minimizing the entropy-weighted cross-entropy loss function. In the subsequent round of training, the model after iteration is used to generate a prediction probability distribution, and a self-feedback signal based on distribution entropy is used to correct the prediction probability distribution cross-entropy training, which has a positive effect on model updating.
[0096] The steps S3022 to S3024 correspond to Figure 2 The self-feedback signal based supervision iterative training 204 in the middle, after updating the parameters of the large language model module and the speech recognition labels of the training speech data subset by minimizing the entropy-weighted cross-entropy loss function, can also perform supervision training 205 based on perturbation reinforcement learning sampling information.
[0097] In at least one embodiment of the present disclosure, a perturbation speech subset of the training speech data subset is constructed, the training speech data subset and the perturbation speech subset are input into the speech large model, the prediction token sequence of the training speech data subset and the prediction token sequence of the perturbation speech subset output by the large language model module in the speech large model are obtained, and a contrastive learning loss function is determined according to the prediction token sequence of the training speech data subset and the prediction token sequence of the perturbation speech subset. The parameters of the large language model module and the speech recognition labels of the training speech data subset are updated by minimizing the contrastive learning loss function.
[0098] That is, considering that the better prediction result exists in other paths with lower path probability in the multiple sampling paths predicted and output by the large language model module in the speech large model, more sampling paths of the large language model module output are obtained to obtain a better prediction token sequence, and the speech large model is further optimized to realize supervision training based on multiple sampling information of perturbation input.
[0099] In some possible implementations, a search tree is constructed according to the prediction token sequence of the training speech data subset and the prediction token sequence of the perturbation speech subset, the scores of the token nodes in the search tree are determined according to the occurrence frequencies of the tokens in the prediction token sequence of the training speech data subset and the prediction token sequence of the perturbation speech subset, the optimal token path and the worst token path are determined according to the scores of the token nodes in the search tree, and the anchor token path is determined according to the prediction token sequence of the training speech data subset. The contrastive learning loss function is determined according to the distance between the optimal token path and the anchor token path and the distance between the worst token path and the anchor token path.
[0100] In other words, by inputting a subset of training speech data and a subset of perturbed speech data into the large speech model and sampling multiple outputs, the multiple outputs (i.e. multiple word sequences) are merged into a search tree, and a contrastive learning loss function is designed based on the word paths in the search tree to optimize the large speech model.
[0101] Figure 4 The illustration shows a flowchart of determining a contrastive learning loss function provided by at least one embodiment of the present disclosure.
[0102] like Figure 4 As shown, for the training speech data subset 401, a perturbation speech subset 402 is constructed. For example, by adjusting the volume, adjusting the speech rate, performing high-pass and low-pass filtering, and indoor / outdoor reverberation, the perturbation speech subset 402 is constructed. The training speech data subset 401 and the perturbation speech subset 402 are input into the large speech model, respectively. After passing through the audio encoder 104 and the large language model module 108, a predicted word sequence 405 is generated. The predicted word sequence 405 may include the predicted word sequence of the training speech data subset. Predicted lexical sequences of perturbed speech subsets A search tree 406 is constructed. The optimal lexical path 4072, the worst lexical path 4071, and the anchor lexical path 4073 are determined from the search tree 406. Then, based on the optimal lexical path 4072, the worst lexical path 4071, and the anchor lexical path 4073, the contrastive learning loss function 408 is determined. The gradient of the large language model module is updated using the contrastive learning loss function 408, completing another iteration of the training process.
[0103] The process of determining the optimal and worst lexical paths is explained below.
[0104] In some possible implementations, the word nodes in the search tree can be individual speech recognition tags (i.e., tokens) in the predicted word sequences of either the training speech data subset or the perturbed speech data subset. For each word node in the search tree, the number of times the word node appears in both the predicted word sequences of the training speech data subset and the perturbed speech data subset is determined, and this number of appearances is defined as the word node's score.
[0105] In other words, for each word node in the search tree, the number of times the word node appears in multiple predicted word sequences of the training speech data subset and multiple predicted word sequences of the perturbed speech subset is calculated. When the word node appears at the corresponding position in a predicted word sequence, the score of the word node is increased by 1, and the scores of each word node are calculated cumulatively.
[0106] Figure 5 The illustration shows a schematic diagram of a search tree provided by at least one embodiment of the present disclosure.
[0107] As shown in Figure 5 the prediction word sequence of the training speech data subset and the prediction word sequence of the perturbed speech subset have 16 in common, the search tree includes word node 1 to word node 8, the score of word node 1 is 11, the score of word node 2 is 5, indicating that among the 16 prediction word sequences, 11 prediction word sequences have the first speech recognition label as word node 1, and the other 5 prediction word sequences have the first speech recognition label as word node 2; similarly, the score of word node 3 is 11, the score of word node 4 is 5, the score of word node 5 is 4, the score of word node 6 is 2, and the score of word node 7 is 5.
[0108] That is, among the 16 prediction word sequences, 4 are “word node 1, word node 3, and word node 5”, 2 are “word node 1, word node 3, and word node 6”, 5 are “word node 1, word node 3, and word node 7”, and 5 are “word node 2, word node 4, and word node 8”.
[0109] In this way, by merging multiple prediction word sequences to form a search tree, and using the number of times each word node appears in multiple prediction word sequences to determine the accuracy of each word node in the search tree, the search tree can not only represent the prediction paths of multiple prediction word sequences, but also represent the confidence of different prediction paths.
[0110] In one or more embodiments of the present disclosure, the optimal word path 4072 can be understood as a positive sample in the contrastive learning loss function, the anchor word path 4073 can be understood as an anchor or benchmark in the contrastive learning loss function, and the worst word path 4071 can be understood as a negative sample in the contrastive learning loss function.
[0111] For example, the optimal word path 4072 can be the prediction path with the highest cumulative score of word nodes in the search tree, the worst word path 4071 can be the prediction path with the lowest cumulative score of word nodes in the search tree, and the anchor word path 4073 can be the prediction path composed of each word node of the prediction word sequence of the training speech data subset in the search tree.
[0112] In this case, the parameters of the large language model module and the speech recognition labels of the training speech data subset are updated to reduce the distance between the optimal word path and the anchor word path and increase the distance between the worst word path and the anchor word path.
[0113] Since the parameters of the large language model module and the speech recognition labels of the training speech data subset need to be updated in a round of training process, by narrowing the distance between the anchor point and the positive sample and pushing away the distance between the anchor point and the negative sample, the speech recognition labels of the training speech data subset can be close to the optimal word piece path with higher accuracy, and the inference ability of the large language model module is closer to the optimal word piece path with higher output accuracy.
[0114] For example, the training speech data subset includes 8 training speeches, each training speech in the training speech data subset is physically disturbed to obtain a disturbed speech subset including 8 disturbed speeches, and the training speech data subset and the disturbed speech subset are input into the speech large model respectively to obtain the predicted word piece sequence of the training speech data subset output by the large language model module in the speech large model and the predicted word piece sequence of the disturbed speech subset , that is, 16 different predicted output sampling word piece sequences are obtained.
[0115] The and are merged to obtain a search tree, and in the merging process, the occurrence times of each word piece node in each prediction path in the search tree in the predicted word piece sequence are calculated to determine the scores of each word piece node. Then, the search tree is searched based on the cumulative scores to determine that the prediction path with the highest cumulative score is the optimal word piece path , the prediction path with the lowest cumulative score is the worst word piece path , and the anchor point word piece path corresponding to the predicted word piece sequence of the training speech data subset is determined from the search tree .
[0116] The triplet is used to determine the contrastive learning loss function. For example, the pre-trained BERT model is used to encode the features of each element in the triplet to obtain high-dimensional vector representations: the feature vector of , the feature vector of , the feature vector of , there are , so that when calculating the distance (such as cosine similarity) subsequently, it can be simplified as vector dot product.
[0117] Since the core of the contrastive learning loss function is that the similarity between the anchor point and the positive sample is as high as possible, and the similarity between the anchor point and the negative sample is as low as possible, the contrastive learning loss function can be:
[0118]
[0119] wherein, For the temperature parameter, the discrimination of the similarity is adjusted, that is, the lower the temperature parameter is, the more sensitive the degree of punishment for similar or dissimilar is. In the above-mentioned contrast learning loss function, when the anchor point is more similar to the positive sample (that is, the is larger) and the anchor point is less similar to the negative sample (that is, the is smaller), the is smaller, which meets the optimization goal of "pulling the positive sample closer and pushing the negative sample away".
[0120] In some embodiments, there can be multiple worst token paths, for example, the , at this time, the contrast learning loss function can be:
[0121]
[0122] Since the denominator of the above-mentioned contrast learning loss function contains the similarity term of "anchor point and all negative samples", the distance between the anchor point and each negative sample can be sufficiently pushed away, the discrimination is enhanced, and through "probabilistic similarity comparison", the large language model module is effectively guided to learn the feature representation of "anchor point and positive sample semantic proximity, and negative sample semantic distance".
[0123] In this way, in one round of training process, the Figure 2 supervised iteration training 204 based on self-feedback signal and supervised training 205 based on perturbation reinforcement learning sampling information are performed, two times of supervised iteration are completed, after the round of training process is completed, the training speech data subset is put back into the training speech data set, so as to continue sampling in the next round of training process.
[0124] One or more embodiments of the present disclosure do not limit the condition for ending the multiple rounds of training process, for example, when the difference between the speech recognition labels of the training speech data subset in the two rounds of training process is less than a set threshold, or when the number of rounds of training process reaches a set value, the training of the large language model in the speech large model stops.
[0125] As described above, after the multiple rounds of training process are ended, the supervised fine-tuning training 206 based on a small amount of scene data can also be performed, so that the speech large model can effectively align the scene demand (for example, the speech recognition demand in the teaching scene).
[0126] In some possible implementations, first, the annotated scene data set is constructed , is the scene data, is the corresponding annotated label, is the number of scene data. Then, the cross-entropy loss function is defined:
[0127]
[0128] wherein, is the model parameter of the speech large model, is the predicted probability distribution output by the speech large model, can be determined by the following way:
[0129]
[0130] wherein, is the original output value output by the last layer (usually a fully connected layer) of the speech large model without normalization processing.
[0131] The ADAM optimization algorithm is used to update the model parameters of the speech large model, for example:
[0132]
[0133] wherein, is the learning rate, and are the first-order and second-order moment estimates after bias correction, respectively.
[0134] When the accuracy of the validation set no longer improves for multiple rounds in a row, or the loss value of the cross-entropy loss function is lower than the preset threshold, stop the supervised fine-tuning training 206 based on a small amount of scene data.
[0135] In this way, the speech large model obtained after the above training has good speech recognition capability and can be widely used in speech recognition, speech recognition and other scenarios, reducing the labeling cost and improving the model adaptability.
[0136] It can be understood that before using the technical solutions disclosed in the embodiments of the present disclosure, the type, use range, use scenario, etc. of the personal information (for example, user portrait features, historical dialogue text, etc.) involved in the present disclosure should be informed to the user and the authorization of the user should be obtained through appropriate means in accordance with relevant laws and regulations. For example, in response to receiving the active request of the user, a prompt information is sent to the user to explicitly prompt the user that the operation requested to be executed will need to obtain and use the personal information of the user. Thus, the user can voluntarily choose whether to provide personal information to the software or hardware such as electronic device, application program, server or storage medium, etc. that executes the technical solutions of the present disclosure according to the prompt information.
[0137] As an optional but non-limiting implementation manner, in response to receiving the active request of the user, the way of sending prompt information to the user may be, for example, the way of pop-up window, and the prompt information can be presented in the form of text in the pop-up window. In addition, the pop-up window can also carry selection controls for the user to select "agree" or "disagree" to provide personal information to the electronic device.
[0138] It can be understood that the above notification and user authorization process is only illustrative and does not limit the implementation of the present disclosure, and other ways that meet the relevant laws and regulations can also be applied to the implementation of the present disclosure.
[0139] It can be understood that the data involved in the technical solutions of the present disclosure (including but not limited to the data itself, the acquisition or use of the data) should comply with the requirements of the relevant laws and regulations and the relevant provisions.
[0140] The above Figures 1 to 5 The training method of the voice large model provided by one or more embodiments of the present disclosure is described in detail, and the device and electronic equipment provided by one or more embodiments of the present disclosure will be introduced below with reference to the accompanying drawings. Figure 6 The structure diagram of the training device of the voice large model provided by at least one embodiment of the present disclosure.
[0141] As Figure 6 indicated, in the voice large model training device 600 of the embodiment, the voice large model includes a text encoder, an audio encoder, a bridge, and a large language model module, the text encoder is configured to generate text tokens based on text corresponding to voice data, the audio encoder is configured to generate audio tokens based on voice data, the bridge is configured to generate fusion tokens based on text tokens and audio tokens, and the large language model module is configured to generate speech recognition results based on the fusion tokens. The voice large model training device 600 includes an acquisition module 601 and a training module 602. For example, these units or modules can be implemented by hardware (such as circuit) modules or software modules, and the following embodiments are the same, and will not be repeated. For example, these units or modules can be implemented by a central processing unit (CPU), a general-purpose graphics processing unit (GPGPU), a graphics processing unit (GPU), a tensor processing unit (TPU), a field programmable gate array (FPGA), or other forms of processing units with data processing capabilities and / or instruction execution capabilities, and corresponding computer instructions.
[0142] The acquisition module 601 is configured to acquire a training voice data set, wherein the training voice data set includes a plurality of training voice data, and the plurality of training voice data respectively correspond to voice recognition labels;
[0143] The training module 602 is configured to perform a plurality of rounds of training processes on the large language model module, each round of the training process including the following steps: determining a training speech data subset for the training process from the training speech data set, and determining a true probability distribution of the training speech data subset according to a speech recognition label of the training speech data subset; inputting the training speech data subset into the speech large model to obtain a predicted probability distribution of the training speech data subset output by the large language model module in the speech large model; determining an entropy-weighted cross-entropy loss function according to the predicted probability distribution and the true probability distribution, wherein an entropy weight of the entropy-weighted cross-entropy loss function is determined based on a distribution entropy of the predicted probability distribution at a current time; and updating parameters of the large language model module and the speech recognition label of the training speech data subset with a target of minimizing the entropy-weighted cross-entropy loss function.
[0144] For example, the obtaining module 601 can be configured to perform the step S301 described above, and the specific implementation principle can be referred to the related description of the step S301. The training module 602 can be configured to perform the step S302 described above, and the specific implementation principle can be referred to the related description of the step S302, which will not be described here again.
[0145] In at least one embodiment of the present disclosure, the training module 602 is further configured to sample the training speech data set to determine the training speech data subset for the training process, wherein an intersection probability between the training speech data subset and a training speech data subset in a historical training process in the training speech data set is greater than a first threshold.
[0146] In at least one embodiment of the present disclosure, the training module 602 is further configured to determine a distribution entropy of the predicted probability distribution at the current time according to the predicted probability distribution at the current time; in response to the distribution entropy of the predicted probability distribution at the current time being less than a first entropy threshold, it is determined that the greater the distribution entropy of the predicted probability distribution at the current time, the smaller the entropy weight of the entropy-weighted cross-entropy loss function; in response to the distribution entropy of the predicted probability distribution at the current time being greater than or equal to the first entropy threshold, it is determined that the greater the distribution entropy of the predicted probability distribution at the current time, the greater the entropy weight of the entropy-weighted cross-entropy loss function; wherein the entropy weight of the entropy-weighted cross-entropy loss function when the distribution entropy of the predicted probability distribution at the current time is less than the first entropy threshold is greater than the entropy weight of the entropy-weighted cross-entropy loss function when the distribution entropy of the predicted probability distribution at the current time is greater than or equal to the first entropy threshold.
[0147] In at least one embodiment of the present disclosure, the training module 602 is further configured to: determine a distribution entropy of the predicted probability distribution of the current time according to the predicted probability distribution of the current time; determine an entropy change of the predicted probability distribution of the current time and an adjacent time according to the distribution entropy of the predicted probability distribution of the current time; determine a static entropy of the current time according to the distribution entropy of the predicted probability distribution of the current time, wherein the static entropy of the current time is used to measure the uncertainty of the probability distribution of the current time; determine a dynamic entropy of the current time according to the entropy change of the predicted probability distribution of the current time and the adjacent time, wherein the dynamic entropy of the current time is used to measure the stability of the probability distribution of the current time; and determine an entropy weight of the entropy-weighted cross-entropy loss function according to the static entropy of the current time and the dynamic entropy of the current time.
[0148] In at least one embodiment of the present disclosure, the training module 602 is further configured to: construct a perturbed speech subset of the training speech data subset; input the training speech data subset and the perturbed speech subset into the speech large model to obtain a predicted token sequence of the training speech data subset and a predicted token sequence of the perturbed speech subset output by the large language model module in the speech large model; determine a contrastive learning loss function according to the predicted token sequence of the training speech data subset and the predicted token sequence of the perturbed speech subset; and update the parameters of the large language model module and the speech recognition labels of the training speech data subset to minimize the contrastive learning loss function.
[0149] In at least one embodiment of the present disclosure, the training module 602 is further configured to: construct a search tree according to the predicted token sequence of the training speech data subset and the predicted token sequence of the perturbed speech subset; determine scores of various token nodes in the search tree according to the number of occurrences of various tokens in the predicted token sequence of the training speech data subset and the predicted token sequence of the perturbed speech subset; determine an optimal token path and a worst token path according to the scores of various token nodes in the search tree, and determine an anchor token path according to the predicted token sequence of the training speech data subset; and determine a contrastive learning loss function according to the distance between the optimal token path and the anchor token path and the distance between the worst token path and the anchor token path.
[0150] In at least one embodiment of the present disclosure, the training module 602 is further configured to: update the parameters of the large language model module and the speech recognition labels of the training speech data subset to reduce the distance between the optimal token path and the anchor token path and increase the distance between the worst token path and the anchor token path.
[0151] In at least one embodiment of the present disclosure, the training module 602 is further configured to: for each of the wordpiece nodes in the search tree, determine a number of occurrences of the wordpiece node in the predicted wordpiece sequence of the training speech data subset and the predicted wordpiece sequence of the perturbed speech subset, and determine the number of occurrences as a score of the wordpiece node.
[0152] It should be noted that, for the sake of clarity and simplicity, one or more embodiments of the present disclosure do not give all the constituent units of the speech large model training apparatus 600. To achieve the necessary functions of the speech large model training apparatus, those skilled in the art can provide and set other constituent units not shown according to specific needs, and one or more embodiments of the present disclosure do not limit this.
[0153] One or more embodiments of the present disclosure also provide an electronic device. The electronic device is specifically used to implement the functions of the speech large model training apparatus 600 in the embodiments as Figure 6
[0154] Figure 7 A structural schematic diagram of an electronic device 700 is provided, as shown in Figure 7 The electronic device 700 includes a bus 701, a processor 702, a communication interface 703 and a memory 704. The processor 702, the memory 704 and the communication interface 703 communicate through the bus 701.
[0155] The bus 701 can be a peripheral component interconnect (PCI) bus or an extended industry standard architecture (EISA) bus, etc. The bus can be divided into an address bus, a data bus, a control bus, etc. For the sake of representation, Figure 7 only one thick line is used in the figure, but it does not mean that there is only one bus or only one type of bus.
[0156] The processor 702 can be any one or more of a central processing unit (CPU), a graphics processing unit (GPU), a microprocessor (MP), or a digital signal processor (DSP).
[0157] The communication interface 703 is used for external communication. For example, the communication interface 703 can be used for communication with a terminal.
[0158] Memory 704 may include volatile memory, such as random access memory (RAM). Memory 704 may also include non-volatile memory, such as read-only memory (ROM), flash memory, hard disk drive (HDD), or solid state drive (SSD).
[0159] The memory 704 stores executable code, which the processor 702 executes to perform the aforementioned training method for the large speech model.
[0160] Specifically, in achieving Figure 6 In the case of the illustrated embodiment, and Figure 6 When the modules or units of the speech model training device 600 described in the embodiment are implemented in software, the following steps are executed: Figure 6 The software or program code required for the functions of each module / unit can be partially or entirely stored in memory 704. Processor 702 executes the program code corresponding to each unit stored in memory 704 and executes the aforementioned training method for the large speech model.
[0161] One or more embodiments of this disclosure also provide a computer-readable storage medium. The computer-readable storage medium can be any available medium capable of being stored by a computing device or a data storage device such as a data center containing one or more available media. The available medium can be a magnetic medium (e.g., floppy disk, hard disk, magnetic tape), an optical medium (e.g., DVD), or a semiconductor medium (e.g., solid-state drive). The computer-readable storage medium includes instructions that instruct the computing device to execute the training method for the large speech model applied to the training apparatus 600 for the large speech model described above.
[0162] One or more embodiments of this disclosure also provide a computer program product containing instructions. The computer program product may be a software or program product containing instructions capable of running on a computing device or stored on any available medium. When the computer program product is run on at least one computing device, it causes the at least one computing device to perform a method for training a large speech model.
[0163] One or more embodiments of the present disclosure also provide a computer readable storage medium. The computer readable storage medium can be any available medium or data storage device that can be accessed by a computing device and includes one or more available media or data storage devices. The available medium can be a magnetic medium, (e.g., a floppy disk, a hard disk, a magnetic tape), an optical medium, (e.g., a DVD), or a semiconductor medium, (e.g., a solid state hard drive), etc. The computer readable storage medium includes instructions that instruct the computing device to perform the training method of the large language model.
[0164] The above description is merely exemplary of the disclosure and the application of the principles thereof and the scope of the disclosure is not limited to the specific embodiments described herein, but only by the claims that follow. It will be readily apparent to those skilled in the art that varying substitutions and modifications can be made to the embodiments described herein without departing from the scope and spirit of the disclosure. For example, features described herein can be combined, swapped, or eliminated in any combination.
[0165] In addition, while operations can be depicted in the drawings in a particular, serial order, this should not be understood as a requirement that such operations be performed in the particular order shown, or in sequential order, or that all illustrated operations be performed, to achieve desirable results. In certain circumstances, multitasking and parallel processing can be advantageous. Likewise, while several specific implementation details are contained in the above discussion, these should not be construed as limitations on the scope of the disclosure, but rather as exemplification of the application. Certain features that are described in the context of separate embodiments can also be implemented in combination in a single embodiment. Conversely, various features that are described in the context of a single embodiment can also be implemented in multiple embodiments separately or in any suitable sub-combination.
[0166] Although the subject matter has been described in language specific to structural features and / or methodological acts, it is to be understood that the subject matter defined in the appended claims is not necessarily limited to the specific features or acts described above. Rather, the specific features and acts described above are disclosed as example forms of implementing the claims.
Claims
1. A method for training a large speech model, wherein, The speech large model comprises a text encoder, an audio encoder, a bridge and a large language model module, the text encoder is used to generate text tokens based on text corresponding to speech data, the audio encoder is used to generate audio tokens based on the speech data, the bridge is used to generate fusion tokens based on the text tokens and the audio tokens, and the large language model module is used to generate a speech recognition result based on the fusion tokens, The method comprises: obtaining a training speech data set, wherein the training speech data set comprises a plurality of training speech data, and each of the plurality of training speech data corresponds to a speech recognition label; performing a plurality of rounds of training processes on the large language model module, and each round of the training process comprises the following steps: determining a training speech data subset of the training process from the training speech data set, and determining a true probability distribution of the training speech data subset according to the speech recognition label of the training speech data subset; inputting the training speech data subset into the speech large model to obtain a predicted probability distribution of the training speech data subset output by the large language model module in the speech large model; determining an entropy-weighted cross-entropy loss function according to the predicted probability distribution and the true probability distribution, wherein an entropy weight of the entropy-weighted cross-entropy loss function is determined based on a distribution entropy of the predicted probability distribution at a current moment, and the operation of determining the entropy weight of the entropy-weighted cross-entropy loss function comprises: determining the distribution entropy of the predicted probability distribution at the current moment according to the predicted probability distribution at the current moment; in response to the distribution entropy of the predicted probability distribution at the current moment being less than a first entropy threshold, determining that the greater the distribution entropy of the predicted probability distribution at the current moment, the smaller the entropy weight of the entropy-weighted cross-entropy loss function; in response to the distribution entropy of the predicted probability distribution at the current moment being greater than or equal to the first entropy threshold, determining that the greater the distribution entropy of the predicted probability distribution at the current moment, the greater the entropy weight of the entropy-weighted cross-entropy loss function; wherein the entropy weight of the entropy-weighted cross-entropy loss function when the distribution entropy of the predicted probability distribution at the current moment is less than the first entropy threshold is greater than the entropy weight of the entropy-weighted cross-entropy loss function when the distribution entropy of the predicted probability distribution at the current moment is greater than or equal to the first entropy threshold; updating parameters of the large language model module and the speech recognition label of the training speech data subset with the aim of minimizing the entropy-weighted cross-entropy loss function.
2. The method of claim 1, wherein, The determination of the training speech data subset of the training process from the training speech data set comprises: sampling the training speech data set to determine the training speech data subset of the training process, wherein an intersection probability between the training speech data subset and a training speech data subset in a historical training process in the training speech data set is greater than a first threshold.
3. The method of claim 1, wherein, The operation of determining the entropy weight of the entropy-weighted cross-entropy loss function comprises: determining the distribution entropy of the predicted probability distribution at the current moment according to the predicted probability distribution at the current moment; determine an entropy change of the prediction probability distribution of the current time and a neighboring time according to a distribution entropy of the prediction probability distribution of the current time; determine a static entropy of the current time according to the distribution entropy of the prediction probability distribution of the current time, wherein the static entropy of the current time is used to measure uncertainty of the probability distribution of the current time; determine a dynamic entropy of the current time according to the entropy change of the prediction probability distribution of the current time and the neighboring time, wherein the dynamic entropy of the current time is used to measure stability of the probability distribution of the current time; determine an entropy weight of the entropy-weighted cross-entropy loss function according to the static entropy of the current time and the dynamic entropy of the current time.
4. The method according to any one of claims 1 to 3, wherein, After the parameters of the large language model module and the speech recognition labels of the training speech data subset are updated with the objective of minimizing the entropy-weighted cross-entropy loss function, the training process of each round further includes the following steps: construct a perturbed speech subset of the training speech data subset; input the training speech data subset and the perturbed speech subset into the speech large model to obtain prediction token sequences of the training speech data subset and prediction token sequences of the perturbed speech subset output by the large language model module in the speech large model; determine a contrastive learning loss function according to the prediction token sequences of the training speech data subset and the prediction token sequences of the perturbed speech subset; update the parameters of the large language model module and the speech recognition labels of the training speech data subset with the objective of minimizing the contrastive learning loss function.
5. The method of claim 4, wherein, The determination of the contrastive learning loss function according to the prediction token sequences of the training speech data subset and the prediction token sequences of the perturbed speech subset includes: construct a search tree according to the prediction token sequences of the training speech data subset and the prediction token sequences of the perturbed speech subset; determine scores of each token node in the search tree according to the number of occurrences of each token in the prediction token sequences of the training speech data subset and the prediction token sequences of the perturbed speech subset; determine an optimal token path and a worst token path according to the scores of each token node in the search tree, and determine an anchor token path according to the prediction token sequences of the training speech data subset; determine the contrastive learning loss function according to the distance between the optimal token path and the anchor token path and the distance between the worst token path and the anchor token path.
6. The method of claim 5, wherein, The updating of the parameters of the large language model module and the speech recognition labels of the training speech data subset with the objective of minimizing the contrastive learning loss function includes: update the parameters of the large language model module and the speech recognition labels of the training speech data subset with the objective of reducing the distance between the optimal token path and the anchor token path and increasing the distance between the worst token path and the anchor token path.
7. The method of claim 5, wherein, The number of occurrences of each word element in the predicted word element sequence of the training speech data subset and the predicted word element sequence of the perturbed speech subset is determined, and the score of each word element node in the search tree is determined based on the number of occurrences. For each word element node in the search tree, the number of occurrences of the word element node in the predicted word element sequence of the training speech data subset and the predicted word element sequence of the perturbed speech subset is determined, and the number of occurrences is determined as the score of the word element node.
8. A device for training a large language model, comprising: The speech large model includes a text encoder, an audio encoder, a bridge, and a large language model module. The text encoder is configured to generate text word elements based on the text corresponding to the speech data. The audio encoder is configured to generate audio word elements based on the speech data. The bridge is configured to generate fusion word elements based on the text word elements and the audio word elements. The large language model module is configured to generate speech recognition results based on the fusion word elements. The device includes: An acquisition module configured to acquire a training speech data set, wherein the training speech data set includes a plurality of training speech data, and each of the plurality of training speech data corresponds to a speech recognition label; A training module configured to perform a plurality of rounds of training processes on the large language model module, wherein each round of the training process includes the following steps: From the training speech data set, a training speech data subset for the training process is determined, and a true probability distribution of the training speech data subset is determined based on the speech recognition label of the training speech data subset. The training speech data subset is input into the speech large model, and a predicted probability distribution of the training speech data subset output by the large language model module in the speech large model is acquired. An entropy-weighted cross-entropy loss function is determined based on the predicted probability distribution and the true probability distribution, wherein the entropy weight of the entropy-weighted cross-entropy loss function is determined based on the distribution entropy of the predicted probability distribution at the current moment. The operation of determining the entropy weight of the entropy-weighted cross-entropy loss function includes: determining the distribution entropy of the predicted probability distribution at the current moment based on the predicted probability distribution at the current moment; in response to the distribution entropy of the predicted probability distribution at the current moment being less than a first entropy threshold, determining that the greater the distribution entropy of the predicted probability distribution at the current moment, the smaller the entropy weight of the entropy-weighted cross-entropy loss function; in response to the distribution entropy of the predicted probability distribution at the current moment being greater than or equal to the first entropy threshold, determining that the greater the distribution entropy of the predicted probability distribution at the current moment, the greater the entropy weight of the entropy-weighted cross-entropy loss function; wherein the entropy weight of the entropy-weighted cross-entropy loss function when the distribution entropy of the predicted probability distribution at the current moment is less than the first entropy threshold is greater than the entropy weight of the entropy-weighted cross-entropy loss function when the distribution entropy of the predicted probability distribution at the current moment is greater than or equal to the first entropy threshold. The parameters of the large language model module and the speech recognition label of the training speech data subset are updated to minimize the entropy-weighted cross-entropy loss function.
9. An electronic device, comprising: at least one processor; and a memory communicatively connected to the at least one processor; wherein the memory stores instructions executable by the at least one processor, the instructions being executed by the at least one processor to enable the at least one processor to perform the method of any one of claims 1 to 7.
10. A non-transitory computer readable storage medium having stored thereon computer instructions, wherein, the computer instructions are for causing a computer to perform the method of any one of claims 1 to 7.
Citation Information
Patent Citations
Language model training method, device and system based on adaptive cross entropy and medium
CN119047324A
Speech recognition method and device, electronic equipment and storage medium
CN119600995A