Large language model training method and device
By inputting intermediate features in the vector mapping form of speech recognition model into the large language model for training, the problems of speech recognition errors and acoustic information extraction are solved, and the better understanding effect and user experience of the large language model are achieved.
Patent Information
- Application Number
- CN202311618472.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2023-11-29
- Publication Date
- 2025-05-30
AI Technical Summary
There may be recognition errors in the speech recognition results, which affects the use of large language models, and the acoustic information in the speech signal cannot be extracted from the recognized text, affecting the experience effect of the large model.
The voice data of the target user is collected, feature extraction is performed, audio features are obtained, and input it into the speech recognition model for speech recognition processing, and intermediate features in the form of vector mapping are obtained. This feature contains both acoustic information and semantic information. The feature fusion model is weighted and summed to obtain the fused intermediate features in the form of vector maps, which are used to train large language models.
Through the joint training of the speech recognition model and the large language model, the two models can be better tuned, system performance can be improved, and acoustic information can be input into the large language model to improve their understanding effect and user experience.
Smart Images

Figure CN120071897A_ABST
Abstract
Description
Technical Field
[0001] The present disclosure relates to the technical field of natural language processing, and in particular, to a method and apparatus for training a large language model. Background Art
[0002] As one of the important interfaces for human-computer interaction, the speech recognition module can conveniently connect people to a large language model (LLM). The accuracy of its recognition is particularly important for the understanding of the large language model. In related technologies, after the input audio is recognized into text using a speech recognition model, the text is directly input into the large language model module.
[0003] However, the speech recognition result may have recognition errors, which will affect the use of the large language model. In addition to semantic information, the speech signal usually also contains acoustic information such as emotion, intonation, accent, age, gender, and background environment. These information cannot be extracted from the recognized text. If the text of the speech recognition is directly input into the large model for training, it will affect the experience effect of the large model. Summary of the Invention
[0004] The present disclosure provides a method, apparatus, electronic device, and storage medium for training a large language model.
[0005] According to a first aspect of the present disclosure, there is provided a method for training a large language model, the method including: collecting speech data and a first label of a target user, and performing feature extraction on the speech data to obtain audio features, where the first label is a task label of the large language model corresponding to the speech data; inputting the audio features into a speech recognition model for speech recognition processing to obtain intermediate features in the form of vector mapping, where the intermediate features in the form of vector mapping include acoustic information and semantic information of the speech data; inputting the intermediate features in the form of vector mapping into a feature fusion model for weighted summation processing to obtain the fused intermediate features in the form of vector mapping; and training the large language model through the fused intermediate features in the form of vector mapping and the first label.
[0006] In some embodiments, the speech recognition model includes an encoder and a decoder. Inputting the audio features into the speech recognition model for speech recognition processing to obtain intermediate features in the form of vector mapping includes: inputting the audio features into the encoder, and processing the audio features through the encoder to obtain speech features in the form of a tensor, where the speech features in the form of a tensor only include acoustic information; inputting the intermediate features in the form of a tensor into the decoder, and processing the intermediate features in the form of a tensor through the decoder to obtain the intermediate features in the form of vector mapping output by the decoder, and the intermediate features in the form of vector mapping include acoustic information and semantic information.
[0007] In some embodiments, the encoder includes N encoder modules, where N is a positive integer greater than or equal to 2. The audio features are input into the encoder, and the encoder processes the audio features to obtain intermediate features in tensor form, including: traversing the N encoder modules, inputting the audio features into the first encoder module among the N encoder modules, and successively processing them through the first feed-forward network module, multi-head self-attention module, convolutional module, second feed-forward network module, and normalization module of the first encoder module to obtain the first intermediate feature in tensor form; inputting the first intermediate feature in tensor form into the second encoder module among the N encoder modules, and successively processing them through the first feed-forward network module, multi-head self-attention module, convolutional module, second feed-forward network module, and normalization module of the second encoder module to obtain the second intermediate feature in tensor form; traversing to the Nth encoder module to obtain the Nth intermediate feature in tensor form.
[0008] In some embodiments, the decoder module includes at least the Nth decoder module, where N is a positive integer greater than or equal to 2. The first processing result is input into the decoder, and the decoder processes the first processing result to obtain intermediate features in vector mapping form, including: traversing the N decoder modules, inputting the Nth intermediate feature in tensor form into the first decoder module among the N decoder modules, and processing them through the multi-head self-attention module and cross-attention module of the first decoder module to obtain the first intermediate feature in vector mapping form; inputting the first intermediate feature in vector mapping form into the feed-forward network module and normalization module of the first decoder module for processing to obtain the output result of the first decoder module, where the output result of the first decoder module is the initial recognition text corresponding to the speech data; inputting the first character output by the decoder and the output result of the first decoder module into the second decoder module among the N decoder modules, and processing them through the multi-head self-attention module and cross-attention module of the second decoder module to obtain the second intermediate feature in vector mapping form; traversing to the Nth decoder module to obtain the third to Nth intermediate features in vector mapping form.
[0009] In some embodiments, the intermediate features in vector mapping form are input into the feature fusion model for weighted summation processing to obtain the fused intermediate features in vector mapping form, including: inputting the first to Nth intermediate features in vector mapping form into the feature fusion model for weighted summation processing to obtain the fused intermediate features in vector mapping form.
[0010] In some embodiments, the method for training a large language model further includes: collecting plain text data and a second label, where the second label is the task label of the large language model corresponding to the plain text data; training the large language model with the intermediate features in the form of fused vector mapping and the first label, including: determining the intermediate features in the form of fused vector mapping and the first label as the first training data, where the first label is the task label of the large language model corresponding to the speech data; determining the plain text data and the second label as the second training data; mixing the first training data and the second training data according to a preset ratio to obtain a training dataset; and training the large language model with the training dataset.
[0011] According to a second aspect of the present disclosure, there is provided a large language model, which is trained by using the method in the first aspect of the present disclosure.
[0012] According to a third aspect of the present disclosure, there is provided a training device for a large language model, the device including: a collection unit, configured to collect speech data of a target user and a first label, and extract features from the speech data to obtain audio features, where the first label is the task label of the large language model corresponding to the speech data; a processing unit, configured to input the audio features into a speech recognition model for speech recognition processing to obtain intermediate features in the form of vector mapping, where the intermediate features in the form of vector mapping include acoustic information and semantic information of the speech data; a fusion unit, configured to input the intermediate features in the form of vector mapping into a feature fusion model for weighted summation processing to obtain intermediate features in the form of fused vector mapping; and a training unit, configured to train the large language model with the intermediate features in the form of fused vector mapping and the first label.
[0013] According to a fourth aspect of the present disclosure, there is provided an electronic device, including: at least one processor; and a memory communicatively connected to the at least one processor; where the memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor to enable the at least one processor to execute the method described in the first aspect of the present disclosure.
[0014] According to a fifth aspect of the present disclosure, there is provided a non-transitory computer-readable storage medium storing computer instructions, where the computer instructions are used to cause a computer to execute the method described in the embodiments of the first aspect of the present disclosure.
[0015] According to a sixth aspect of the present disclosure, there is provided a computer program product, including a computer program, where the computer program, when executed by a processor, implements the method described in the embodiments of the first aspect of the present disclosure.
[0016] Embodiments of the present disclosure provide a method for training a large language model. The method includes: collecting voice data and a first label of a target user, and performing feature extraction on the voice data to obtain audio features, where the first label is a task label of the large language model corresponding to the voice data; inputting the audio features into a speech recognition model for speech recognition processing to obtain intermediate features in the form of vector mapping, where the intermediate features in the form of vector mapping include acoustic information and semantic information of the voice data; inputting the intermediate features in the form of vector mapping into a feature fusion model for weighted summation processing to obtain the fused intermediate features in the form of vector mapping; and training the large language model through the fused intermediate features in the form of vector mapping and the first label. By inputting the vector mapping of speech recognition into the large language model, the present disclosure can prevent possible errors in speech recognition from affecting the use of the large language model, and the vector mapping contains both semantic information and acoustic information, which can better improve the understanding effect of the large language model and enhance the user experience.
[0017] It should be understood that the content described in this part is not intended to identify the key or important features of the embodiments of the present application, nor is it used to limit the scope of the present application. Other features of the present application will become easily understood through the following description. BRIEF DESCRIPTION OF THE DRAWINGS
[0018] The drawings are used to better understand the solution and do not constitute a limitation to the present disclosure. Among them:
[0019] Figure 1 is a schematic flowchart of a method for training a large language model provided by an embodiment of the present disclosure;
[0020] Figure 2 is a schematic diagram of a method for training a large language model provided by an embodiment of the present disclosure;
[0021] Figure 3 is a schematic flowchart of a method for training a large language model provided by an embodiment of the present disclosure;
[0022] Figure 4 is an example diagram of an encoder module provided by an embodiment of the present disclosure;
[0023] Figure 5 is an example diagram of a decoder module provided by an embodiment of the present disclosure;
[0024] Figure 6 is an example diagram of a method for training a large language model provided by an embodiment of the present disclosure;
[0025] Figure 7 is a schematic structural diagram of a training device for a large language model provided by an embodiment of the present disclosure;
[0026] Figure 8Schematic block diagram of an exemplary electronic device provided by an embodiment of the present disclosure. Detailed implementation manners
[0027] The following describes exemplary embodiments of the present disclosure with reference to the accompanying drawings. Various details of the embodiments of the present disclosure are included to facilitate understanding, and they should be considered merely exemplary. Therefore, those of ordinary skill in the art should recognize that various changes and modifications can be made to the embodiments described herein without departing from the scope and spirit of the present disclosure. Similarly, descriptions of well-known functions and structures are omitted in the following description for clarity and conciseness.
[0028] In recent years, large language model technology has developed rapidly. Its powerful language understanding ability has led to more and more applications in text generation and text understanding. On the other hand, as one of the important interfaces for human-computer interaction, the speech recognition module can conveniently connect people to the large language model, so its accuracy is particularly important for the understanding of the large language model. Currently, the general approach for large language models with speech as the interface is to directly use a speech recognition model to recognize the input audio into text and then directly input it into the large language model module. This solution has the following problems: on the one hand, the speech recognition result may contain recognition errors, which will affect the use of the large language model. And since the speech recognition model and the large language model are trained separately, the speech recognition result will not directly reflect the effect of the large language model. On the other hand, in addition to the semantic information contained in the speech, the speech signal usually also contains acoustic information such as emotion, intonation, accent, age, gender, background environment, etc. These information cannot be extracted from the recognized text. If the text recognized by speech is directly input into the large model for training, it will affect the experience effect of the large model.
[0029] For this reason, the present disclosure proposes a training method for a large language model. Instead of directly inputting the speech recognition result into the large language model, the intermediate features in vector form obtained by speech recognition are input into the large language model. The intermediate features in vector form contain both acoustic information such as emotion, accent, and age, and also semantic information. Directly inputting the intermediate features in vector form into the large language model, on the one hand, can realize the joint training of the speech recognition model and the large language model, which can better optimize the two models simultaneously and make the entire system achieve the best performance. On the other hand, acoustic information such as emotion, intonation, accent, age, gender, background noise, etc. can be input into the large language model to better improve the understanding effect of the large language model and enhance the user experience.
[0030] The following describes in detail a training method, device, and electronic device for a large language model proposed by the present disclosure with reference to the accompanying drawings. For ease of understanding, Figure 2 A schematic diagram of a training method for a large language model is shown.
[0031] Figure 1 The flowchart of a training method for a large language model provided by an embodiment of the present disclosure is shown as Figure 1 follows. The method includes steps 101-104:
[0032] Step 101, collect the voice data and the first label of the target user, and perform feature extraction on the voice data to obtain audio features.
[0033] Among them, the first label is the task label of the large language model corresponding to the voice data.
[0034] In an embodiment of the present disclosure, the audio feature is the result of numerically representing a certain attribute of the audio signal.
[0035] In some embodiments, by collecting the voice data of the target user, performing feature extraction on the voice data to obtain audio features, and based on the downstream task tags of the large language model, the task tags of the corresponding large language model are marked for the voice data and the text of the voice data.
[0036] In some embodiments, the downstream tasks of the large language model include one or more of translation, text generation, and sentiment analysis.
[0037] In some embodiments, the audio features may include time-domain features (such as duration, energy, etc.), frequency-domain features (such as spectrum, power spectrum, etc.), Mel spectrum, MFCC (Mel-scale Frequency Cepstral Coefficients), etc., which are not limited in the present disclosure.
[0038] In some embodiments, the audio features can also be directly obtained from the existing large language model training database, which is not limited in the present disclosure.
[0039] Step 102, input the audio features into a speech recognition model for speech recognition processing to obtain intermediate features in the form of vector mapping.
[0040] Among them, the intermediate features in the form of vector mapping include the acoustic information and semantic information of the voice data.
[0041] In an embodiment of the present disclosure, different from the result of speech recognition, the intermediate features refer to the features extracted during the speech recognition processing of the audio features. The result of speech recognition only contains semantic information, while the intermediate features contain both semantic information and acoustic information such as emotion, accent, age, gender, background noise, etc.
[0042] In an embodiment of the present disclosure, the vector mapping form is an embedding form. The intermediate feature in the vector mapping form can be understood as reducing the audio feature data to a feature representation of a fixed size, which is convenient for processing and calculation. Essentially, the intermediate feature in the vector mapping form is an N-dimensional real-valued vector, which can be used to represent text, music, video, etc.
[0043] In some embodiments of the present disclosure, the speech recognition model adopted is an Encoder-Decoder (decoder-encoder) structure. The Encoder is composed of multiple Encoder Blocks (decoder modules), and the Decoder is composed of multiple Decoder Blocks (decoder modules).
[0044] Furthermore, in specific implementations, neither the encoder nor the decoder is fixed. CNN (Convolutional Neural Network), RNN (Recurrent Neural Network), BiRNN (Bidirectional Recurrent Neural Network), etc. can be selected, and the present disclosure does not limit this.
[0045] In some embodiments of the present disclosure, as Figure 2 shown, in the speech recognition model, the encoder is composed of 12 encoder modules, the decoder is composed of 8 decoder modules, and the output of each decoder module includes the intermediate feature in the vector mapping form.
[0046] Step 103: Input the intermediate feature in the vector mapping form into the feature fusion model for weighted summation processing to obtain the fused intermediate feature in the vector mapping form.
[0047] In some embodiments of the present disclosure, if the decoder is composed of 8 decoder modules, and the output of each decoder module includes the intermediate feature in the vector mapping form, that is, 8 intermediate features in the vector mapping form can be obtained. Input these 8 intermediate features in the vector mapping form into the feature fusion model for weighted summation processing to obtain 1 fused intermediate feature in the vector mapping form.
[0048] Step 104: Train the large language model with the fused intermediate feature in the vector mapping form and the first label.
[0049] In an embodiment of the present disclosure, the large speech model is a deep learning model trained with a large amount of text, which can generate natural language text or understand the meaning of language text.
[0050] In an embodiment of the present disclosure, the intermediate features in the form of vector mapping generated during the speech recognition process are input into a large language model, and the large language model is trained through the intermediate features in the form of vector mapping. Since the intermediate features in the form of vector mapping include acoustic information such as emotion, accent, age, gender, background noise, etc. in addition to the speech recognition result, it is equivalent to inputting acoustic information such as emotion, accent, age, gender, background noise, etc. into the large language model for training at the same time, which can better improve the understanding effect of the large language model and be used for human-computer interaction to improve the user experience.
[0051] It can be understood that in this solution, the training labels of the speech recognition model and the large language model are the same, realizing the joint training of the speech recognition model and the large language model, and the training objectives are all the downstream task objectives of the large language model, which can improve the training effect.
[0052] In summary, according to the embodiment of the present disclosure, the method includes: collecting the speech data and the first label of the target user, and performing feature extraction on the speech data to obtain audio features, where the first label is the task label of the large language model corresponding to the speech data; inputting the audio features into the speech recognition model for speech recognition processing to obtain intermediate features in the form of vector mapping, where the intermediate features in the form of vector mapping include the acoustic information and semantic information of the speech data; inputting the intermediate features in the form of vector mapping into the feature fusion model for weighted summation processing to obtain the fused intermediate features in the form of vector mapping; training the large language model through the fused intermediate features in the form of vector mapping and the first label. The present disclosure inputs the vector mapping of speech recognition into the large language model, which can prevent the possible errors in speech recognition from affecting the use of the large language model, and the vector mapping contains both semantic information and acoustic information, which can better improve the understanding effect of the large language model and improve the user experience.
[0053] Figure 3 The flowchart of a training method for a large language model provided by an embodiment of the present disclosure is a further disclosure of the Figure 1 corresponding embodiment. For ease of understanding, Figure 4 a schematic diagram of a training method for a large language model is shown.
[0054] As Figure 3 shown, the method includes steps 301-306.
[0055] Step 301, collect the speech data and the first label of the target user, and perform feature extraction on the speech data to obtain audio features.
[0056] Among them, the first label is the task label of the large language model corresponding to the speech data.
[0057] In some embodiments of the present disclosure, the target user is the user targeted by the large language model. For example, the large language model is used in the human-machine interaction system of a vehicle, supporting functions such as voice chat and translation, and is targeted at the group using the vehicle.
[0058] In some embodiments of the present disclosure, the voice data of the user can be collected through a microphone and the collected voice data can be labeled with the tags of the corresponding large language model.
[0059] In the embodiments of the present disclosure, acoustic feature extraction is performed on the collected audio. The manner of feature extraction is not limited in the present disclosure. The extracted audio features may include time-domain features (such as duration, energy, etc.), frequency-domain features (such as spectrum, power spectrum, etc.), Mel spectrum, MFCC, etc.
[0060] In some embodiments of the present disclosure, the extracted audio features are 80-dimensional fbank (FilterBank) features.
[0061] In some embodiments, the speech recognition model includes an encoder and a decoder. The audio features are input into the speech recognition model for speech recognition processing, and the intermediate features in the form of vector mapping are obtained, including: step 302 and step 303.
[0062] Step 302: Input the audio features into the encoder, and process the audio features through the encoder to obtain speech features in the form of tensors.
[0063] Among them, the speech features in the form of tensors only include acoustic information.
[0064] In some embodiments of the present disclosure, the speech features in the form of tensors refer to the speech features in the form of Tensor.
[0065] In some embodiments of the present disclosure, the encoder includes N encoder modules, where N is a positive integer greater than or equal to 2. Step 302 includes: traversing the N encoder modules, inputting the audio features into the first encoder module among the N encoder modules, and successively passing through the first feed-forward network module, multi-head self-attention module, convolutional module, second feed-forward network module, and normalization module of the first encoder module to obtain the first intermediate feature in the form of a tensor; inputting the first intermediate feature in the form of a tensor into the second encoder module among the N encoder modules, and successively passing through the first feed-forward network module, multi-head self-attention module, convolutional module, second feed-forward network module, and normalization module of the second encoder module to obtain the second intermediate feature in the form of a tensor; traversing to the Nth encoder module to obtain the Nth intermediate feature in the form of a tensor.
[0066] In some embodiments of the present disclosure, such as Figure 2As shown, the encoder includes 12 encoder modules. The 12 encoder modules are connected in series to form the encoder. The output of the first encoder module is the input of the second encoder module, the output of the second encoder module is the input of the third encoder module, and so on until the eighth encoder module. The output of the eighth encoder module is the final output of the encoder. If the output of the first encoder module is a tensor of size (T, F), where T represents the time dimension and F represents the feature dimension, the output size of each subsequent encoder module in series will be (T, F).
[0067] In some embodiments of the present disclosure, the structure of each encoding module is as Figure 4 shown. A dual grid structure, namely the conformer structure, is adopted, which includes a first feed forward network (FFN) module, a multi-head self-attention module, a convolutional module, a second feed forward network module, and a layernorm module. The first feed forward network module, the multi-head self-attention module, the convolutional module, the second feed forward network module, and the layernorm module are connected in series in sequence.
[0068] Step 303: Input the first processing result into the decoder, and process the first processing result through the decoder to obtain intermediate features in the form of vector mapping.
[0069] Among them, the intermediate features in the form of vector mapping include acoustic information and semantic information.
[0070] In some embodiments, the decoder module includes at least the Nth decoder module, where N is a positive integer greater than or equal to 2. As Figure 2 shown, the decoder includes 8 decoder modules. The 8 decoder modules are connected in series to form the decoder.
[0071] In some embodiments of the present disclosure, as Figure 5 shown, each decoder module adopts the transformer structure, which includes a multi-head self-attention module, a cross attention module, a feed forward network module, and a layernorm module. The input of each layer of the decoder module is the first character output by the decoder and the output of the previous layer of the decoder module. Among them, the output of the cross attention module in each layer of the decoder module includes intermediate features in the form of vector mapping, namely embedding.
[0072] Further, in some embodiments, step 303 includes: traversing N decoder modules, inputting the Nth intermediate feature in tensor form into the first decoder module among the N decoder modules, and obtaining the first intermediate feature in vector mapping form after being processed by the multi-head self-attention module and the cross-attention module of the first decoder module; inputting the first intermediate feature in vector mapping form into the feed-forward network module and the normalization module of the first decoder module for processing to obtain the output result of the first decoder module, where the output result of the first decoder module is the initial recognition text corresponding to the speech data; inputting the first character output by the decoder and the output result of the first decoder module into the second decoder module among the N decoder modules, and obtaining the second intermediate feature in vector mapping form after being processed by the multi-head self-attention module and the cross-attention module of the second decoder module; traversing to the Nth decoder module to obtain the third to Nth intermediate features in vector mapping form.
[0073] In some embodiments of the present disclosure, as Figure 6 shown, if the decoder includes 8 decoder modules, input the output of the encoder into the decoder, and after being processed by the 8 decoder modules, the output of the cross-attention module in each layer of the decoder module can be obtained, including the intermediate features in vector mapping form output by the decoder modules from the 1st layer to the 8th layer.
[0074] Step 304: Input the intermediate feature in vector mapping form into the feature fusion model for weighted summation processing to obtain the fused intermediate feature in vector mapping form.
[0075] In some embodiments of the present disclosure, step 304 includes: inputting the first to Nth intermediate features in vector mapping form into the feature fusion model for weighted summation processing to obtain the fused intermediate feature in vector mapping form.
[0076] In some embodiments of the present disclosure, as Figure 6 shown, the audio feature is processed by the speech recognition encoder module and the decoder module once to obtain the intermediate features in vector mapping form of the 1st to 8th layers of the decoder module. Through the feature fusion model, weighted summation is performed on the intermediate features in vector mapping form of the 1st to 8th layers of the decoder module and integrated into an intermediate feature of vector mapping.
[0077] Step 305: Train the large language model with the fused intermediate feature in vector mapping form and the first label.
[0078] In an embodiment of the present disclosure, input the fused intermediate feature in vector mapping form and the first label into the downstream tasks of the large language model to train the large language model, and the downstream tasks of the large language model include text summarization, translation, sentiment analysis, etc.
[0079] In one embodiment of the present disclosure, the method for training the large language model of the present disclosure further includes: collecting plain text data and a second label, where the second label is the task label of the large language model corresponding to the plain text data. Further, step 305 includes: determining the intermediate features and the first label in the form of a fused vector as the first training data, where the first label is the task label of the large language model corresponding to the speech data; determining the plain text data and the second label as the second training data; mixing the first training data and the second training data according to a preset ratio to obtain a training data set; and training the large language model with the training data set.
[0080] In one embodiment of the present disclosure, for example, if the downstream task of the large language model is a translation task, then the speech data is the user speech to be translated, the first label is the translation label corresponding to the text of the speech, the plain text data is the text to be translated, and the second label is the translation label corresponding to the text.
[0081] Further, in some embodiments of the present disclosure, the training of the large language model not only includes training the large language model with the first training data, but also includes mixing the first training data and the second training data according to a preset ratio to obtain a training data set, and training the large language model with the training data set.
[0082] In some embodiments of the present disclosure, an Automatic Speech Recognition (ASR) model is used for speech recognition. This method can be used to train the automatic speech model, the feature fusion model, and the large language model simultaneously. The training methods include using the first training data and / or the second training data to first fix the automatic speech model and only train the feature fusion model and the large language model, and also include using the first training data and / or the second training data to simultaneously train the automatic speech model, the feature fusion model, and the large language model.
[0083] In one implementation manner of the present disclosure, both the automatic speech model and the large language model are initialized with pre-trained models, and the feature fusion module is randomly initialized. Therefore, at the beginning of training, the parameters of the automatic speech model need to be fixed to prevent the gradient from being too large when the reverse gradient is transmitted to the automatic speech model. After training for a certain period of time, the three models of the automatic speech model, the feature fusion model, and the large language model can be trained and updated together.
[0084] In summary, the present disclosure proposes a method for training a large language model. Instead of directly inputting the speech recognition result into the large language model, the vector mapping form of speech recognition is input into the large language model. The vector mapping form includes both acoustic information such as emotion, accent, and age, as well as semantic information. By directly inputting the vector mapping form into the large language model, on the one hand, the joint training of the recognition model and the large language model can better optimize the two models simultaneously, enabling the entire system to achieve the best performance; on the other hand, this method inputs acoustic information such as emotion, intonation, accent, age, gender, and background noise into the large language model, better improving the understanding effect of the large language model and enhancing the user experience.
[0085] Corresponding to the above method for training a large language model, the present disclosure proposes a training device for a large language model. Figure 7 FIG. is a schematic structural diagram of a training device 300 for a large language model provided by an embodiment of the present disclosure. As Figure 7 shown, the device includes: an acquisition unit 310, configured to acquire speech data and a first label of a target user, and perform feature extraction on the speech data to obtain audio features, where the first label is a task label of the large language model corresponding to the speech data; a processing unit 320, configured to input the audio features into a speech recognition model for speech recognition processing to obtain intermediate features in a vector mapping form, where the intermediate features in the vector mapping form include acoustic information and semantic information of the speech data; a fusion unit 330, configured to input the intermediate features in the vector mapping form into a feature fusion model for weighted summation processing to obtain fused intermediate features in the vector mapping form; and a training unit 340, configured to train the large language model through the fused intermediate features in the vector mapping form and the first label.
[0086] In some embodiments, the speech recognition model includes an encoder and a decoder. The processing unit 320 is specifically configured to: input the audio features into the encoder, and process the audio features through the encoder to obtain speech features in a tensor form, where the speech features in the tensor form only include acoustic information; input the intermediate features in the tensor form into the decoder, and process the intermediate features in the tensor form through the decoder to obtain intermediate features in a vector mapping form output by the decoder, and the intermediate features in the vector mapping form include acoustic information and semantic information.
[0087] In some embodiments, the encoder includes N encoder modules, where N is a positive integer greater than or equal to 2. When the audio features are input into the encoder, the processing unit 320 is specifically configured to: traverse the N encoder modules, input the audio features into the first encoder module among the N encoder modules, and successively process them through the first feed-forward network module, the multi-head self-attention module, the convolutional module, the second feed-forward network module, and the normalization module of the first encoder module to obtain the intermediate features in the form of a first tensor; input the intermediate features in the form of the first tensor into the second encoder module among the N encoder modules, and successively process them through the first feed-forward network module, the multi-head self-attention module, the convolutional module, the second feed-forward network module, and the normalization module of the second encoder module to obtain the intermediate features in the form of a second tensor; traverse to the Nth encoder module to obtain the intermediate features in the form of the Nth tensor.
[0088] In some embodiments, the decoder module includes at least the Nth decoder module, where N is a positive integer greater than or equal to 2. The processing unit 320 is specifically configured to: traverse the N decoder modules, input the intermediate features in the form of the Nth tensor into the first decoder module among the N decoder modules, and process them through the multi-head self-attention module and the cross-attention module of the first decoder module to obtain the intermediate features in the form of a first vector mapping; input the intermediate features in the form of the first vector mapping into the feed-forward network module and the normalization module of the first decoder module for processing to obtain the output result of the first decoder module, where the output result of the first decoder module is the initial recognition text corresponding to the speech data; input the first character output by the decoder and the output result of the first decoder module into the second decoder module among the N decoder modules, and process them through the multi-head self-attention module and the cross-attention module of the second decoder module to obtain the intermediate features in the form of a second vector mapping; traverse to the Nth decoder module to obtain the intermediate features in the form of the third to Nth vector mappings.
[0089] In some embodiments, the fusion unit 330 is specifically configured to: input the intermediate features in the form of the first to Nth vector mappings into the feature fusion model for weighted summation processing to obtain the intermediate features in the form of a fused vector mapping.
[0090] In some embodiments, the acquisition unit 310 is further configured to: acquire plain text data and a second tag, where the second tag is a task tag of a large language model corresponding to the plain text data; at this time, the training unit 340 is specifically configured to: determine the intermediate feature in the form of a fused vector mapping and the first tag as the first training data, where the first tag is a task tag of the large language model corresponding to the speech data; determine the plain text data and the second tag as the second training data; mix the first training data and the second training data according to a preset ratio to obtain a training data set; train the large language model through the training data set. The training unit 340 is further configured to: determine the intermediate feature in the form of a fused vector mapping and the first tag as the first training data; acquire plain text data and a second tag, and determine the plain text data and the second tag as the second training data, where the second tag is a task tag of the large language model corresponding to the plain text data; mix the first training data and the second training data according to a preset ratio to obtain a training data set; train the large language model through the training data set.
[0091] In summary, according to the embodiments of the present disclosure, the device uses the intermediate feature of the deep vector mapping of speech recognition instead of the recognized text result as the input to the large language model through the processing unit and the training unit, which can end-to-end train the speech recognition model and the large language model. At the same time, since the intermediate feature of the vector mapping of the language recognition deep neural network model contains both acoustic information such as emotion, accent, and age, as well as semantic information, the present invention enables the large language model to learn acoustic information such as emotion, intonation, accent, age, gender, and background noise, thereby improving the effect of the large language model.
[0092] It should be noted that since the device embodiments of the present disclosure correspond to the above method embodiments, the foregoing explanations of the method embodiments also apply to the devices in this embodiment. The principles are the same. For the details not disclosed in the device embodiments, reference may be made to the above method embodiments, and no further elaboration will be provided in the present disclosure.
[0093] According to the embodiments of the present disclosure, the present disclosure also provides an electronic device, a readable storage medium, and a computer program product.
[0094] Figure 8FIG. shows a schematic block diagram of an exemplary electronic device that can be used to implement embodiments of the present disclosure. The electronic device is intended to represent various forms of digital computers, such as, laptop computers, desktop computers, workstations, personal digital assistants, servers, blade servers, mainframe computers, and other suitable computers. The electronic device can also represent various forms of mobile devices, such as, personal digital processors, cellular telephones, smart phones, wearable devices, and other similar computing devices. The components shown herein, their connections and relationships, and their functions are merely exemplary and are not intended to limit the implementation of the present disclosure described and / or claimed herein.
[0095] As Figure 8 shown, the device 400 includes a computing unit 401 that can perform various appropriate actions and processes according to a computer program stored in a ROM (Read-Only Memory) 402 or a computer program loaded from a storage unit 408 into a RAM (Random Access Memory) 403. In the RAM 403, various programs and data required for the operation of the device 400 can also be stored. The computing unit 401, the ROM 402, and the RAM 403 are connected to each other via a bus 404. An I / O (Input / Output) interface 405 is also connected to the bus 404.
[0096] A plurality of components in the device 400 are connected to the I / O interface 405, including: an input unit 406, such as a keyboard, a mouse, etc.; an output unit 407, such as various types of displays, speakers, etc.; a storage unit 408, such as a magnetic disk, an optical disk, etc.; and a communication unit 409, such as a network card, a modem, a wireless communication transceiver, etc. The communication unit 409 allows the device 400 to exchange information / data with other devices via a computer network such as the Internet and / or various telecommunication networks.
[0097] The computing unit 401 can be various general-purpose and / or special-purpose processing components with processing and computing capabilities. Some examples of the computing unit 401 include, but are not limited to, CPU (Central Processing Unit), GPU (Graphic Processing Units), various dedicated AI (Artificial Intelligence) computing chips, various computing units running machine learning model algorithms, DSP (Digital Signal Processor), and any suitable processor, controller, microcontroller, etc. The computing unit 401 executes the various methods and processes described above, such as the training method of the large language model. For example, in some embodiments, the training method of the large language model can be implemented as a computer software program, which is tangibly included in a machine-readable medium, such as the storage unit 408. In some embodiments, part or all of the computer program can be loaded and / or installed onto the device 400 via the ROM 402 and / or the communication unit 409. When the computer program is loaded into the RAM 403 and executed by the computing unit 401, one or more steps of the method described above can be executed. Alternatively, in other embodiments, the computing unit 401 can be configured to execute the foregoing communication method by any other suitable means (e.g., by means of firmware).
[0098] Various embodiments of the systems and techniques described above in this document can be implemented in digital electronic circuit systems, integrated circuit systems, FPGAs (Field Programmable Gate Arrays), ASICs (Application-Specific Integrated Circuits), ASSPs (Application Specific Standard Products), SOCs (System On Chip), CPLDs (Complex Programmable Logic Devices), computer hardware, firmware, software, and / or combinations thereof. These various embodiments can include: being implemented in one or more computer programs that can be executed and / or interpreted on a programmable system including at least one programmable processor, which can be a special or general programmable processor, receiving data and instructions from a storage system, at least one input device, and at least one output device, and transmitting the data and instructions to the storage system, the at least one input device, and the at least one output device.
[0099] The program code for implementing the methods of the present disclosure can be written in any combination of one or more programming languages. These program codes can be provided to a processor or controller of a general-purpose computer, a special-purpose computer, or other programmable data processing devices, such that when the program codes are executed by the processor or controller, the functions / operations specified in the flowcharts and / or block diagrams are implemented. The program codes can be executed entirely on the machine, partially on the machine, executed partially on the machine as an independent software package and partially on a remote machine, or executed entirely on a remote machine or server.
[0100] In the context of the present disclosure, a machine-readable medium can be a tangible medium that can contain or store a program for use by or in connection with an instruction execution system, apparatus, or device. A machine-readable medium can be a machine-readable signal medium or a machine-readable storage medium. A machine-readable medium can include, but is not limited to, electronic, magnetic, optical, electromagnetic, infrared, or semiconductor systems, apparatus, or devices, or any suitable combination of the foregoing. More specific examples of machine-readable storage media would include electrical connections based on one or more wires, portable computer disks, hard disks, RAM, ROM, EPROM (Electrically Programmable Read-Only-Memory), or flash memory, optical fibers, CD-ROM (Compact Disc Read-Only Memory), optical storage devices, magnetic storage devices, or any suitable combination of the foregoing.
[0101] In order to provide interaction with a user, the systems and techniques described herein can be implemented on a computer having: a display device for displaying information to the user (e.g., a CRT (Cathode-Ray Tube) or an LCD (Liquid Crystal Display) monitor); and a keyboard and a pointing device (e.g., a mouse or a trackball) through which the user can provide input to the computer. Other kinds of devices can also be used to provide interaction with the user; for example, the feedback provided to the user can be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and input from the user can be received in any form (including acoustic input, voice input, or tactile input).
[0102] The systems and techniques described herein can be implemented in a computing system including backend components (e.g., as a data server), or a computing system including middleware components (e.g., an application server), or a computing system including frontend components (e.g., a user computer having a graphical user interface or a web browser through which a user can interact with an implementation of the systems and techniques described herein), or a computing system including any combination of such backend components, middleware components, or frontend components. The components of the system can be interconnected to each other by digital data communication in any form or medium (e.g., a communication network). Examples of communication networks include: LAN (Local Area Network), WAN (Wide Area Network), the Internet, and blockchain networks.
[0103] A computer system can include a client and a server. The client and the server are generally far from each other and usually interact through a communication network. The client-server relationship is created by computer programs running on respective computers and having a client-server relationship with each other. The server can be a cloud server, also known as a cloud computing server or a cloud host, which is a host product in the cloud computing service system, solving the defects of difficult management and weak business scalability existing in traditional physical hosts and VPS services ("Virtual Private Server", or simply "VPS"). The server can also be a server of a distributed system, or a server combined with blockchain.
[0104] It should be noted that artificial intelligence is a discipline that studies how to make a computer simulate certain thinking processes and intelligent behaviors of humans (such as learning, reasoning, thinking, planning, etc.), and it has both hardware-level technologies and software-level technologies. Artificial intelligence hardware technologies generally include technologies such as sensors, dedicated artificial intelligence chips, cloud computing, distributed storage, and big data processing; artificial intelligence software technologies mainly include several major directions such as computer vision technology, speech recognition technology, natural language processing technology, and machine learning / deep learning, big data processing technology, and knowledge graph technology.
[0105] It should be understood that various forms of processes shown above can be used, with steps reordered, added, or deleted. For example, the steps recorded in this disclosure can be executed in parallel, sequentially, or in a different order, as long as the desired results of the technical solutions disclosed in this disclosure can be achieved, and no limitation is imposed herein.
[0106] The above specific embodiments do not constitute a limitation on the protection scope of the present disclosure. Those skilled in the art should understand that various modifications, combinations, sub-combinations and substitutions can be made according to design requirements and other factors. Any modifications, equivalent substitutions and improvements made within the spirit and principle of the present disclosure shall be included within the protection scope of the present disclosure.
Claims
1. A training method for a large language model, characterized in that, the method includes: Collecting the speech data and the first label of the target user, and extracting features from the speech data to obtain audio features, where the first label is the task label of the large language model corresponding to the speech data; Inputting the audio features into a speech recognition model for speech recognition processing to obtain intermediate features in vector mapping form, where the intermediate features in vector mapping form include the acoustic information and semantic information of the speech data; Inputting the intermediate features in vector mapping form into a feature fusion model for weighted summation processing to obtain the fused intermediate features in vector mapping form; Training the large language model with the fused intermediate features in vector mapping form and the first label.
2. The method according to claim 1, characterized in that, the speech recognition model includes an encoder and a decoder, and the inputting the audio features into the speech recognition model for speech recognition processing to obtain intermediate features in vector mapping form includes: Inputting the audio features into the encoder, and processing the audio features through the encoder to obtain speech features in tensor form, where the speech features in tensor form only include acoustic information; Inputting the intermediate features in tensor form into the decoder, and processing the intermediate features in tensor form through the decoder to obtain the intermediate features in vector mapping form output by the decoder, and the intermediate features in vector mapping form include acoustic information and semantic information.
3. The method according to claim 2, characterized in that, the encoder includes N encoder modules, N being a positive integer greater than or equal to 2, and the inputting the audio features into the encoder and processing the audio features through the encoder to obtain intermediate features in tensor form includes: Traversing the N encoder modules, inputting the audio features into the first encoder module among the N encoder modules, and successively processing through the first feed-forward network module, multi-head self-attention module, convolutional module, second feed-forward network module and normalization module of the first encoder module to obtain the first intermediate features in tensor form; Inputting the first intermediate features in tensor form into the second encoder module among the N encoder modules, and successively processing through the first feed-forward network module, multi-head self-attention module, convolutional module, second feed-forward network module and normalization module of the second encoder module to obtain the second intermediate features in tensor form; Traversing to the Nth encoder module to obtain the Nth intermediate features in tensor form.
4. The method according to claim 3, characterized in that, the decoder module includes at least the Nth decoder module, N being a positive integer greater than or equal to 2, and the inputting the first processing result into the decoder and processing the first processing result through the decoder to obtain intermediate features in vector mapping form includes: Traverse the N decoder modules, input the intermediate feature in the form of the Nth tensor into the first decoder module among the N decoder modules, and after being processed by the multi-head self-attention module and the cross-attention module of the first decoder module, obtain the intermediate feature in the form of the first vector mapping; Input the intermediate feature in the form of the first vector mapping into the feed-forward network module and the normalization module of the first decoder module for processing to obtain the output result of the first decoder module, where the output result of the first decoder module is the initial recognition text corresponding to the speech data; Input the first character output by the decoder and the output result of the first decoder module into the second decoder module among the N decoder modules, and after being processed by the multi-head self-attention module and the cross-attention module of the second decoder module, obtain the intermediate feature in the form of the second vector mapping; Traverse to the Nth decoder module to obtain the intermediate features in the form of the third to the Nth vector mappings.
5. The method according to claim 4, wherein, the inputting the intermediate feature in the form of the vector mapping into the feature fusion model for weighted summation processing to obtain the intermediate feature in the form of the fused vector mapping includes: Input the intermediate features in the form of the first to the Nth vector mappings into the feature fusion model for weighted summation processing to obtain the intermediate feature in the form of the fused vector mapping.
6. The method according to any one of claims 1-5, wherein, the training method of the large language model further includes: collecting plain text data and a second label, where the second label is the task label of the large language model corresponding to the plain text data; the training of the large language model by using the intermediate feature in the form of the fused vector mapping and the first label includes: Determine the intermediate feature in the form of the fused vector mapping and the first label as the first training data, where the first label is the task label of the large language model corresponding to the speech data; Determine the plain text data and the second label as the second training data; Mix the first training data and the second training data according to a preset ratio to obtain a training data set; Train the large language model by using the training data set.
7. A large language model, wherein, the large language model is trained by using the training method of the large language model according to any one of claims 1-6.
8. A training device for a large language model, wherein, the device includes: An acquisition unit, configured to acquire the speech data and the first label of the target user, and perform feature extraction on the speech data to obtain audio features, where the first label is the task label of the large language model corresponding to the speech data; A processing unit, configured to input the audio features into a speech recognition model for speech recognition processing to obtain intermediate features in the form of vector mappings, where the intermediate features in the form of vector mappings include the acoustic information and semantic information of the speech data; A fusion unit, configured to input the intermediate feature in the form of a vector mapping into a feature fusion model for weighted summation processing to obtain a fused intermediate feature in the form of a vector mapping; A training unit, configured to train the large language model by using the fused intermediate feature in the form of a vector mapping and the first label.
9. An electronic device, comprising: at least one processor; and a memory communicatively connected to the at least one processor; wherein, the memory stores instructions executable by the at least one processor, and when the instructions are executed by the at least one processor, the at least one processor is enabled to execute the method according to any one of claims 1-6.
10. A non-transitory computer-readable storage medium storing computer instructions, wherein, the computer instructions are used to cause the computer to execute the method according to any one of claims 1-6.
11. A computer program product, comprising a computer program which, when executed by a processor, implements the method according to any one of claims 1-6.
Citation Information
Cited By
Audio understanding model training method and device, audio understanding method and device, storage medium and program product
CN120356465A