Large language model coding training method and device
Through partial dimensional coding and multi-layer target coding superimposed training for training data, the problems of low dimensional feature compression and efficiency in traditional model training are solved, and more efficient model training effects are achieved.
Patent Information
- Application Number
- CN202510547869.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2024-10-21
- Publication Date
- 2025-07-25
- Estimated Expiration
- 2044-10-21
AI Technical Summary
Traditional model training provides equally compressed features of each dimension of the training data, with high information density, high training difficulty and low training efficiency, and cannot meet higher performance requirements in certain or certain aspects.
The training data extracts some dimensions to encode, generates multi-layer target encoding, and superimposes them for large language model training to generate target models.
The training effect of large language models in certain dimensions is strengthened, and the training efficiency and output quality of the model are improved.
Smart Images

Figure CN120373393A_ABST
Abstract
Description
[0001] This application is a divisional application. The application number of the original application is 202411470413.6, the application date is October 21, 2024, and the invention title is "A Method and Device for Encoding and Training of Large Language Models". Technical Field
[0002] The present invention relates to the technical field of machine learning, and particularly to a method and device for encoding and training of large language models. Background Art
[0003] Traditional model training encodes the entire training data uniformly, such that the features of each dimension of the training data are equally compressed, with a high information density, high training difficulty, and low training efficiency; the models trained in this way have balanced performance in all aspects and cannot meet the higher performance requirements for one or some aspects or dimensions of the model. Summary of the Invention
[0004] The present invention provides a method and device for encoding and training of large language models, which can strengthen the training effect of large language models in certain dimensions and obtain better training effects of large language models.
[0005] According to one aspect of the present invention, there is provided a method for encoding and training of large language models, including:
[0006] Extracting partial training data of at least one dimension for at least one type of training data;
[0007] Encoding the partial training data of at least one dimension separately to generate at least one layer of enhanced target encoding;
[0008] Determining the N-layer target encoding of the training data for at least one type of training data in each training sample data group; where N is an integer greater than or equal to 2;
[0009] Stacking at least one layer of the enhanced target encoding and the N-layer target encoding to train a large language model and generate a target model.
[0010] According to another aspect of the present invention, there is provided a device for encoding and training of large language models, including:
[0011] A partial training data extraction module for extracting partial training data of at least one dimension for at least one type of training data;
[0012] An enhanced target encoding generation module for encoding the partial training data of at least one dimension separately to generate at least one layer of enhanced target encoding;
[0013] An N-layer target encoding generation module, configured to determine the N-layer target encoding of at least one type of training data in each of the training sample data groups; where N is an integer greater than or equal to 2;
[0014] A target model generation module, configured to stack at least one layer of the enhanced target encoding and the N-layer target encoding, train a large language model, and generate a target model.
[0015] According to another aspect of the present invention, there is provided an electronic device, which includes:
[0016] At least one processor; and
[0017] A memory communicatively connected to the at least one processor; where
[0018] The memory stores a computer program executable by the at least one processor, and when the computer program is executed by the at least one processor, the at least one processor is enabled to execute the large language model encoding training method according to any embodiment of the present invention.
[0019] According to another aspect of the present invention, there is provided a computer-readable storage medium storing computer instructions for enabling a processor to implement the large language model encoding training method according to any embodiment of the present invention when executed.
[0020] The large language model encoding training solution of the embodiments of the present invention extracts partial training data of at least one dimension for at least one type of training data; encodes the partial training data of at least one dimension respectively to generate at least one layer of enhanced target encoding; determines the N-layer target encoding of at least one type of training data in each training sample data group; where N is an integer greater than or equal to 2; stacks at least one layer of the enhanced target encoding and the N-layer target encoding, trains a large language model, and generates a target model. The technical solution provided by the embodiments of the present invention can strengthen the training effect of the large language model in certain dimensions and obtain a better large language model training effect.
[0021] It should be understood that the content described in this part is not intended to identify the key or important features of the embodiments of the present invention, nor is it used to limit the scope of the present invention. Other features of the present invention will become easily understood through the following description. Description of the Drawings
[0022] To more clearly illustrate the technical solutions in the embodiments of the present invention, the following will briefly introduce the accompanying drawings required for the description of the embodiments. Obviously, the accompanying drawings in the following description are only some embodiments of the present invention. For those of ordinary skill in the art, without creative efforts, other accompanying drawings can be obtained based on these drawings.
[0023] Figure 1 It is a flowchart of a large language model encoding training method provided in Embodiment 1 of the present invention;
[0024] Figure 2 It is a flowchart of a large language model encoding training method provided in Embodiment 2 of the present invention;
[0025] Figure 3 It is a schematic structural diagram of a large language model encoding training device provided in Embodiment 3 of the present invention;
[0026] Figure 4 It is a schematic structural diagram of an electronic device for implementing the large language model encoding training method of the embodiments of the present invention. Detailed implementation manners
[0027] In order to enable those skilled in the art to better understand the solutions of the present invention, the following will clearly and completely describe the technical solutions in the embodiments of the present invention in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only some of the embodiments of the present invention, rather than all of them. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of the present invention.
[0028] It should be noted that the terms "first", "second", etc. in the specification and claims of the present invention and the above accompanying drawings are used to distinguish similar objects, and do not necessarily need to describe a specific order or sequence. It should be understood that such data can be interchanged under appropriate circumstances so that the embodiments of the present invention described here can be implemented in an order other than those illustrated or described here. In addition, the terms "include" and "have" and any variations thereof are intended to cover non-exclusive inclusion. For example, a process, method, system, product or device that includes a series of steps or units does not necessarily have to be limited to those clearly listed steps or units, but may include other steps or units not clearly listed or inherent to these processes, methods, products or devices.
[0029] Embodiment 1
[0030] Figure 1FIG. 0 is a flowchart of a method for encoding and training a large language model according to Embodiment 1 of the present invention. This embodiment is applicable to the situation of encoding and training a large language model. The method can be executed by a large language model encoding and training device, which can be implemented in the form of hardware and / or software, and the large language model encoding and training device can be configured in an electronic device.
[0031] As Figure 1 shown, the method includes:
[0032] S110. Obtain a training data set; wherein, the training data set includes at least two training sample data groups, and each training sample data group is composed of at least two corresponding types of training data.
[0033] In the embodiment of the present invention, a training data set is obtained. The training data set includes multiple training sample data groups, and each training sample data group is composed of at least two corresponding types of data. For example, a training sample data group may include input data and fitting data; another example is that a training sample data group may include reference data, input data, and fitting data. Among them, the input data is the data that needs to output corresponding generated content according to the input content during the subsequent training of the large language model, such as the questions in the question-and-answer data, and the lyrics part in the music generation data; the reference data is the data that needs to refer to some of the information in it to output corresponding generated content during the subsequent training of the large language model, such as the reference pictures in the text-to-image data, the reference songs in the music generation data, or the reference accompaniment, reference singing style, etc. in the reference songs, which is used to make the large language model generate output data with the same style; the fitting data is the standard answer data, which is used to compare with the output result of the large language model one by one during the subsequent training of the large language model, calculate the loss value, and thus correct the probability of the next output result of the large model. It is the ideal data that the large language model needs to fit. The training data set is a data set that matches the application scenario of the large language model to be used for subsequent training. For example, if the large language model to be trained subsequently is a music generation model, each training sample data group in the training data set may include fitting music and target lyrics, or may also refer to music, fitting music, and target lyrics.
[0034] It should be noted that the embodiment of the present invention does not limit the number of input data and the number of reference data included in the training sample data group. Exemplarily, when training a large language model for music generation, the input data in the training sample data group can be a piece of lyric content or multiple pieces of lyrics, and the reference data in the training sample data group can be a certain reference song, or reference accompaniment and reference singing style, etc. in the reference song.
[0035] S120. For at least one type of training data in each of the training sample data groups, determine the first-layer target encoding of the training data.
[0036] In an embodiment of the present invention, for each training sample data group in the training dataset, determine the first-layer target encoding of at least one type of training data in the training sample data group. It can be understood that the first-layer target encoding of any one or more or all of the training data in the training sample data group can be determined, where the types of training data for which the first-layer target encoding is determined in each training sample data group in the training dataset can be the same or different. Exemplarily, an encoder (Encoder) can be used to encode at least one type of training data in the training sample data group to determine the first-layer target encoding, where the encoder can include a MERT encoder (MERT Encoder) and a Mel encoder (Mel Encoder).
[0037] Optionally, before determining the first-layer target encoding of at least one type of training data in each of the training sample data groups, it further includes: performing discretization processing on the at least one type of training data to generate a plurality of minimum discrete units; determining the first-layer target encoding of the training data for at least one type of training data in each of the training sample data groups includes: encoding each minimum discrete unit of at least one type of training data in each of the training sample data groups respectively to determine the first-layer target encoding of the training data. Exemplarily, perform discretization processing on at least one type of training data in the training sample data group to generate at least two minimum discrete units, and then use the encoder to encode each of the at least two minimum discrete units respectively to determine the corresponding first target encoding. It can be understood that the number of first target encodings is the same as the number of minimum discrete units.
[0038] Taking the large language model trained subsequently as the music generation model, and the training sample data set consisting of reference music, fitting music and target lyrics as an example, an exemplary description is given. The fitting music is discretized to generate at least two frames of fitting audio data. Exemplarily, the fitting music can be discretized into 10 frames of fitting audio data per second. For example, if the fitting music is a 10-second piece of music, the fitting music can be discretized into 100 frames of fitting audio data. Additionally, the fitting music can also be discretized as a whole. For example, the fitting music is discretized into 500 frames of fitting audio data. Among them, the number of frames of the discretized fitting audio data can be determined according to the length of the reference music. The longer the length of the fitting music, the more frames of the discretized fitting audio data. Each frame of fitting audio data is encoded by an encoder to generate the corresponding first-layer target encoding, where the number of the first-layer target encodings is the same as the number of frames of the fitting audio data. Optionally, the same encoder can be used to encode each frame of fitting audio data, or different encoders can be used to encode each frame of fitting audio data. Among them, the encoder can be an open-source encoder, such as the MERT encoder or the Mel encoder. In the embodiment of the present invention, the discretization process of the target lyrics may include: performing word segmentation on the target lyrics based on a preset word segmentation algorithm to divide the target lyrics into multiple word segmentation units, where the word segmentation unit can be each character, each word, or each sentence. Each word segmentation unit in the target lyrics is encoded by an encoder to generate the first-layer target encoding corresponding to each word segmentation unit. The method for determining the first-layer target encoding of the reference music can be similar to that of the fitting music, which will not be elaborated here.
[0039] Optionally, determining the first-layer target encoding of the training data includes: determining the first-layer initial encoding of the training data; comparing the first-layer initial encoding with each first encoding in a preset encoding codebook or the first standard encoding corresponding to the first encoding to determine the first target encoding corresponding to the first encoding in the encoding codebook that has the highest similarity to the first-layer initial encoding; and determining the first-layer target encoding based on the first target encoding. Exemplarily, at least one type of training data in a training sample data group is encoded by an encoder, and the obtained encoding information is used as the first-layer initial encoding of the training data. A preset encoding codebook is obtained, where the encoding codebook consists of multiple first encodings. Optionally, the first encoding in the encoding codebook can be a standard encoding (also referred to as a special encoding), or the standard encoding can be replaced with special characters. For example, the first standard encoding is used as encoding 0, the second standard encoding is used as encoding 1, and so on, to generate an encoding codebook such as [0, 1, 2,..., 1023]. This can greatly simplify the complexity of the codebook and at the same time isolate the encoding (i.e., special characters such as 0, 1, 2, etc.) and the standard encoding, facilitating future adjustment of the format and / or content of the standard encoding. For example, the first standard encoding can be a multi-dimensional vector, or an audio encoding, a graphic encoding, a character encoding, etc., any one or more; or the first standard encoding was originally a 256-dimensional encoding vector, and if it is adjusted to a 1024-dimensional encoding vector later, the codebook is still [0, 1, 2,..., 1023].
[0040] The first-layer initial encoding is successively compared with each first encoding in the encoding codebook or the first standard encoding corresponding to the first encoding to calculate the similarity between the two. The first encoding in the encoding codebook that has the highest similarity to the first-layer initial encoding is used as the first target encoding, and the first-layer target encoding of the training data is determined based on the first target encoding. Among them, the first target encoding in the encoding codebook can be directly used as the first-layer target encoding, or the first standard encoding corresponding to the first target encoding in the encoding codebook can be used as the first-layer target encoding, or the first target encoding in the encoding codebook can be encoded based on a preset encoding strategy to generate the first-layer target encoding corresponding to the standard data.
[0041] S130. Use the first-layer target encoding as the current-layer target encoding, determine the current-layer encoding deviation of the current-layer target encoding, and determine the next-layer target encoding based on the current-layer encoding deviation.
[0042] In an embodiment of the present invention, the first-layer target encoding of the training data is updated to the current-layer target encoding, and the current-layer encoding deviation of the current-layer target encoding is determined, where the current-layer encoding deviation may be the difference between the current-layer initial encoding and the current-layer target encoding. The next-layer target encoding of the training data is determined according to the current-layer encoding deviation. For example, the current-layer encoding deviation may be encoded based on a preset encoding algorithm to generate the next-layer target encoding of the training data. It can be understood that the first-layer encoding deviation of the first-layer target encoding of the training data is determined, where the residual between the first-layer initial encoding and the first-layer target encoding of the training data may be used as the first-layer encoding deviation. The second-layer target encoding of the training data is determined according to the first-layer encoding deviation.
[0043] Optionally, determining the current-layer encoding deviation of the current-layer target encoding and determining the next-layer target encoding based on the current-layer encoding deviation includes: calculating the residual between the current-layer initial encoding and the current-layer target encoding, and using the residual as the current-layer encoding deviation of the current-layer encoding; comparing the current-layer encoding deviation with each second encoding or the second standard encoding corresponding to the second encoding in a preset next-layer deviation codebook, determining the second target encoding corresponding to the second encoding with the highest similarity to the current-layer encoding deviation in the next-layer deviation codebook, and determining the next-layer target encoding based on the second target encoding.
[0044] Exemplarily, the residual between the current-layer initial encoding and the current-layer target encoding is used as the current-layer encoding deviation of the current-layer encoding. A next-layer deviation codebook is obtained, where the next-layer deviation codebook is composed of multiple second encodings, and the second encoding can be understood as a residual encoding. Optionally, the second encoding in the deviation codebook may be a standard residual encoding (also referred to as a special residual encoding), or the standard residual encoding may be replaced with a special character. For example, the first standard residual encoding is used as encoding 0, the second standard residual encoding is used as encoding 1, and so on, to generate a deviation codebook such as [0, 1, 2,..., 1023]. This can greatly simplify the complexity of the deviation codebook, and at the same time, the encoding (i.e., special characters such as 0, 1, 2, etc.) and the standard residual encoding can be isolated, facilitating future adjustment of the format and / or content of the standard residual encoding. For example, the first standard residual encoding may be a multi-dimensional vector, or any one or more of an audio encoding, a graphic encoding, a character encoding, etc.; or the first standard residual encoding was originally a 256-dimensional encoding vector and was subsequently adjusted to a 1024-dimensional encoding vector. At this time, the deviation codebook is still [0, 1, 2,..., 1023].
[0045] Compare the current layer coding deviation with each second coding in the next layer deviation codebook or the second standard coding corresponding to the second coding in turn, calculate the similarity between the two, use the second coding with the highest similarity to the current layer coding deviation in the next layer deviation codebook as the second target coding, and determine the next layer target coding of the training data based on the second target coding. Among them, the second target coding in the next layer deviation codebook can be directly used as the next layer target coding, or the second standard coding corresponding to the second target coding in the next layer deviation codebook can be used as the next layer target coding, or the second target coding in the next layer deviation codebook can be encoded based on a preset coding strategy to generate the next layer target coding corresponding to the standard data.
[0046] S140. Update the next layer target coding to the current layer target coding, and return to execute the current layer coding deviation for determining the current layer target coding to determine the Nth layer target coding; where N is an integer greater than or equal to 2.
[0047] In the embodiment of the present invention, compare the first layer coding deviation of the first layer target coding of the training data with the second standard coding corresponding to each second coding in the second layer deviation codebook, determine the second target coding corresponding to the second coding with the highest similarity to the first layer coding deviation in the second layer deviation codebook, and determine the second layer target coding based on the second target coding. Use the first layer coding deviation as the second layer initial coding, calculate the residual between the second initial coding and the second layer target coding, and use this residual as the second layer coding deviation, that is, use the residual between the first layer coding deviation and the second layer target coding as the second layer coding deviation. For the convenience of description, the second coding in the ith layer deviation codebook can also be called the ith coding, and the standard coding corresponding to the ith coding can be called the ith standard coding. Therefore, compare the second layer coding deviation with the third standard coding corresponding to each third coding in the third layer deviation codebook, determine the third target coding corresponding to the third coding with the highest similarity to the second layer coding deviation in the third layer deviation codebook, and determine the third layer target coding based on the third target coding. Use the second layer coding deviation as the third layer initial coding, calculate the residual between the third initial coding and the third layer target coding, and use this residual as the third layer coding deviation, that is, use the residual between the second layer coding deviation and the third layer target coding as the third layer coding deviation. Keep looping according to the above method until the Nth layer target coding of the training data is determined, where N is an integer greater than or equal to 2.
[0048] It can be understood that according to S120 - S130, the Nth layer target coding corresponding to at least one kind of training data in the training sample data group can be determined. It should be noted that the number of layers of the target coding of at least one kind of training data in each training sample data group in the training data set can be the same or different.
[0049] S150. Train a large language model based on the N-layer target encoding to generate a target model.
[0050] In the embodiments of the present invention, the N-layer target encoding corresponding to at least one type of training data in each training sample data group in the training data set is input into the large language model to train the large language model and generate a target model. Among them, the large language model can be any open-source large model, such as GPT, etc. It should be noted that the embodiments of the present invention do not limit the application scenarios of the target model. Exemplarily, the target model can be a music generation model. At this time, each training sample data group in the training data set used to train the target model may include reference music, sample lyrics, and fitting music; for another example, the target model can be a text-to-image model, that is, a model that generates images according to text. At this time, each training sample data group in the training data set used to train the target model may include sample text and fitting images.
[0051] The large language model encoding training method according to the embodiments of the present invention includes: obtaining a training data set; where the training data set contains at least two training sample data groups, and each training sample data group is composed of at least two corresponding types of training data; for at least one type of training data in each training sample data group, determining the first-layer target encoding of the training data; using the first-layer target encoding as the current-layer target encoding, determining the current-layer encoding deviation of the current-layer target encoding, and determining the next-layer target encoding based on the current-layer encoding deviation; updating the next-layer target encoding to the current-layer target encoding, and returning to execute determining the current-layer encoding deviation of the current-layer target encoding to determine the N-layer target encoding; where N is an integer greater than or equal to 2; training a large language model based on the N-layer target encoding to generate a target model. The technical solution provided by the embodiments of the present invention can provide richer data information by encoding at least one type of training data in the training sample data group into multi-layer target encoding, improve the refinement degree of data discretization and fitting, and thus improve the quality of the data output by the trained large language model.
[0052] In some embodiments, before determining the first-layer target encoding of at least one type of training data in each of the training sample data groups, it further includes: extracting partial training data of at least one dimension for the at least one type of training data; encoding the partial training data of at least one dimension respectively to generate at least one layer of enhanced target encoding; training a large language model based on N layers of the target encoding to generate a target model, including: superimposing at least one layer of the enhanced target encoding and the N layers of the target encoding to train the large language model to generate a target model. The advantage of this setting is that it can enhance the training effect of the large language model in certain dimensions and obtain a better training effect of the large language model.
[0053] In the embodiments of the present invention, for at least one type of training data in each training sample data group, partial training data of at least one dimension is extracted from the training data, and the partial training data of each dimension is encoded respectively to generate at least one layer of enhanced target encoding. Among them, the partial training data of each dimension can be encoded as a whole to generate at least one layer of enhanced target encoding, or the partial training data of each dimension can be discretized, and each generated minimum discrete unit is encoded to generate at least one layer of enhanced target encoding. It should be noted that when the enhanced target encoding is a multi-layer encoding, the determination method of the multi-layer enhanced target encoding is the same as the method for determining the N-layer target encoding of the training data in the above embodiments, and will not be elaborated here.
[0054] Exemplarily, if the training sample data group includes reference music, fitting music, and target lyrics, a splitting operation of a preset number of dimensions can be performed on the fitting music. For example, at least one dimension of the fitting music part (such as the fitting music parts of the accompaniment part and the vocal part in two dimensions) can be split from the fitting music. The fitting music part of each dimension is discretized to generate at least two frames of fitting audio data. Each frame of fitting audio data is encoded by an encoder to generate the corresponding at least one layer of enhanced target encoding. Similarly, a splitting operation of a preset number of dimensions is performed on the reference sample music. For example, at least one dimension of the reference music part (such as the reference music parts of the accompaniment part and the vocal part in two dimensions) is split from the reference music. The at least one dimension of the reference music part that is split is uniformly encoded as a whole to generate the corresponding at least one layer of enhanced target encoding. Alternatively, the at least one dimension of the reference music part that is split is discretized to generate at least two frames of reference audio data. Each frame of reference audio data is encoded by an encoder to generate the corresponding at least one layer of enhanced target encoding. The target lyrics are encoded to generate the corresponding at least one layer of enhanced target encoding. Also exemplarily, the bass, midrange, and treble in a frame of audio data are encoded in different layers to generate the corresponding enhanced target encoding, or different instruments and different vocals in a frame of audio data are separated and individually encoded in different layers to generate the corresponding enhanced target encoding.
[0055] In an embodiment of the present invention, at least one layer of enhanced target encoding is stacked with N layers of target encoding, and the stacked encoding is input into a large language model to train the large language model to generate a target model. Optionally, the stacking of at least one layer of the enhanced target encoding with N layers of the target encoding and training the large language model includes: placing at least one layer of the enhanced target encoding before the N layers of the target encoding, and at least a part of the enhanced target encoding is input into the large language model prior to the N layers of the target encoding for training the large language model.
[0056] Exemplarily, taking the fitting music in the training sample data group as an example, the enhanced target encoding of the vocal part extracted from the fitting music is used as the first layer, and the 4 layers of target encoding corresponding to the fitting music are used as the 2nd, 3rd, 4th, and 5th layers respectively. The enhanced target encoding of the vocal part of the fitting music is preferentially input into the large language model as the first layer encoding for prediction output, and then the 4 layers of target encoding corresponding to the fitting music are sequentially input into the large language model to continue training the large language model based on the 4 layers of target encoding to generate a target model. The advantage of such a setting is that the large language model can first perform fitting output with the vocal part of the simple fitting music, and then perform fitting training on the complex song, obtaining a clearer fitting training of the vocal part than simply performing song fitting, thereby achieving a better training effect of the large language model.
[0057] Example 2
[0058] Figure 2 The following is a flowchart of a large language model encoding training method provided in Example 2 of the present invention. As Figure 2 shown, the method includes:
[0059] S210. Obtain a training data set; wherein, the training data set contains at least two training sample data groups, each of the training sample data groups consists of at least two corresponding types of training data, each of the training samples includes input data and fitting data, or each of the training sample data groups includes reference data, input data and fitting data.
[0060] S220. For at least one type of training data in each of the training sample data groups, determine the first-layer target encoding of the training data.
[0061] S230. Take the first-layer target encoding as the current-layer target encoding, determine the current-layer encoding deviation of the current-layer target encoding, and determine the next-layer target encoding based on the current-layer encoding deviation.
[0062] S240. Update the next-layer target encoding to the current-layer target encoding, and return to execute determining the current-layer encoding deviation of the current-layer target encoding to determine the Nth-layer target encoding; wherein, N is an integer greater than or equal to 2.
[0063] S250. Input the input encoding of the non-fitting data and the first-layer target encoding of the fitting data into the large language model to determine the first output encoding of the large language model; wherein, the non-fitting data includes the input data, or the non-fitting data includes the reference data and the input data.
[0064] Wherein, each training sample data may include input data and fitting data, or may include reference data, input data and fitting data. When the training sample data includes input data and fitting data, the input data is used as the non-fitting data; when the training sample data includes reference data, input data and fitting data, the reference data and the input data are used as the non-fitting data.
[0065] In the embodiment of the present invention, the Nth-layer target encoding of the non-fitting data is used as the input encoding of the non-fitting data, and the input encoding of the non-fitting data and the first-layer target encoding of the fitting data are input into the large language model to train the large language model to obtain the first output encoding of the large language model. It can be understood that the first output encoding is the encoding output by the large language model after learning the input encoding of the non-fitting data and the first-layer target encoding of the fitting data.
[0066] S260. Determine the second output encoding of the large language model based on the second-layer target encoding of all the encodings input previously and the fitting data input; and so on until the Nth output encoding of the large language model is determined based on the Nth-layer target encoding of all the encodings input previously and the fitting data input to the large language model.
[0067] In an embodiment of the present invention, the second output encoding of the large language model is determined based on the second-layer target encoding of all the encodings input previously and the fitting data input to the large language model. Similarly, the third output encoding of the large language model is determined based on the third-layer target encoding of all the encodings input previously and the fitting data input to the large language model; the fourth output encoding of the large language model is determined based on the fourth-layer target encoding of all the encodings input previously and the fitting data input to the large language model. And so on until the Nth output encoding of the large language model is determined based on the Nth-layer target encoding of all the encodings input previously and the fitting data input to the large language model.
[0068] S270. Iteratively optimize the large language model based on at least one Nth-layer target encoding of the fitting data and the corresponding number of output encodings to generate a target model.
[0069] In an embodiment of the present invention, the Nth-layer target encoding of the fitting data may be one or multiple. When the fitting data corresponds to multiple Nth-layer target encodings, the output encodings corresponding to each Nth-layer target encoding are obtained through S260 - S270. The large language model is continuously iteratively optimized based on at least one Nth-layer target encoding of the fitting data and the corresponding number of output encodings to generate a target model.
[0070] Optionally, the fitting data corresponds to M Nth-layer target encodings, where M is an integer greater than 1. The method further includes: inputting the input encoding of the non-fitting data and the first first-layer target encoding of the fitting data into the large language model to determine the first output encoding of the large language model; determining the second output encoding of the large language model based on all the encodings input previously and the first second-layer target encoding and the second first-layer target encoding of the fitting data, where the first second-layer target encoding and the second first-layer target encoding of the fitting data are input in a manner of superposition summation or splicing; and so on until the (M + N - 1)th output encoding of the large language model is determined based on all the encodings input previously and the Mth Nth-layer target encoding of the fitting data; combining the (M + N - 1) output encodings of the large language model into M Nth-layer final output encodings; and iteratively optimizing the large language model based on at least M Nth-layer target encodings and M Nth-layer final output encodings of the fitting data to generate a target model.
[0071] In an embodiment of the present invention, since the output data of the large language model depends on the input data of the large language model corresponding to the previous fitting data, therefore, the M N-layer target encodings corresponding to the fitting data are input into the large language model in a delayed manner. Specifically, there are M N-layer target encodings corresponding to the fitting data. All N-layer target encodings of the non-fitting data are used as the input encodings of the non-fitting data. The input encodings of the non-fitting data and the first first-layer target encoding of the fitting data are input into the large language model to enable the large language model to make a prediction and obtain the first output encoding of the large language model. It can be understood that the first output encoding is the encoding output by the large language model after learning the input encodings of the non-fitting data and the first first-layer target encoding of the fitting data. Based on all the encodings previously input by the large language model, the first second-layer target encoding and the second first-layer target encoding of the fitting data input into the large language model, the second output encoding of the large language model is determined. Specifically, after the first second-layer target encoding and the second first-layer target encoding of the fitting data are superimposed and summed or concatenated, the second output encoding of the large language model is determined based on the superimposed and summed or concatenated encoding and all the encodings previously input. By analogy, until the M+N-1th output encoding of the large language model is determined based on all the encodings previously input by the large language model and the Mth N-layer target encoding of the fitting data input.
[0072] Exemplarily, M = 4 and N = 4. Table 1 is a table for superimposed input in a delayed manner of 4 4-layer target encodings corresponding to the fitting data provided by an embodiment of the present invention:
[0073] Table 1
[0074] t(14) t(24) t(34) t(44) t(13) t(23) t(33) t(43) t(12) t(22) t(32) t(42) t(11) t(21) t(31) t(41)
[0075] After inputting the input encoding of the non-fitted data and the target encoding t(11) of the first first layer of the fitted data into the large language model, the first output encoding of the large language model is determined. Then, the first second layer target encoding t(12) and the second first layer target encoding t(21) are superimposed and summed and then input into the large language model. The large language model determines the second output encoding based on all the encodings input previously and the encoding after the superposition and summation of t(12) and t(21). The first third layer target encoding t(13), the second second layer target encoding t(22), and the third first layer target encoding t(31) of the fitted data are superimposed and summed and then input into the large language model. The large language model determines the third output encoding based on all the encodings input previously and the encoding after the superposition and summation of t(13), t(22), and t(31). The first fourth layer target encoding t(14), the second third layer target encoding t(23), the third second layer target encoding t(32), and the fourth first layer target encoding t(41) of the fitted data are superimposed and summed and then input into the large language model. The large language model determines the fourth output encoding based on all the encodings input previously and the encoding after the superposition and summation of t(14), t(23), t(32), and t(41). And so on. In the manner described in Table 1, the 7th output encoding of the large language model is obtained. It can be understood that inputting the M N-layer target encodings of the fitted data into the large language model in a delayed superposition manner can enable the large language model to first obtain the target encoding of each unit of the low-level layer and then obtain the target encoding of the high-level layer, which is beneficial to the training of the large language model.
[0076] Since the output data of the large language model depends on all the inputs of the non-fitted data (input data, or reference data and input data), therefore, the N-layer target encodings corresponding to the non-fitted data can be input into the large language model by using a conventional superposition or splicing method, that is, all the N-layer target encodings of the non-fitted data are input into the large language model at one time, which can enable the large model to obtain all the N-layer target encodings of the non-fitted data at one time, thereby making the training of the large language model faster and more efficient. Optionally, the N-layer target encodings of the non-fitted data can also be input into the large language model by using the DELAY method for superposition to train the large language model. Exemplarily, Table 2 is a table of the conventional method for superposition input of 4 4-layer target encodings corresponding to the non-fitted data provided by an embodiment of the present invention:
[0077] Table 2
[0078] t(14) t(24) t(34) t(44) t(13) t(23) t(33) t(43) t(12) t(22) t(32) t(42) t(11) t(21) t(31) t(41)
[0079] Among them, when the N-layer target encoding of non-fitted data is stacked and input into the large language model in a conventional manner, the multi-dimensional encoding vectors obtained by adding or concatenating the N-layer target encoding of each unit can be used. For example, adding the 256-dimensional t(11), t(12), t(13), and t(14) results in a 256-dimensional vector, or concatenating them results in a 1024-dimensional vector. Adding the 256-dimensional t(21), t(22), t(23), and t(24) results in a 256-dimensional vector, which is still a 256-dimensional vector, or concatenating them results in a 1024-dimensional vector. The multi-layer encoding vectors are input into the large language model. Among them, the output of the large language model is also a multi-layer encoding vector, thereby improving the accuracy of the output of the large language model.
[0080] Combine the M + N - 1 output encodings of the large language model into M N-layer final output encodings, and then iteratively optimize the large language model based on at least M N-layer target encodings and M N-layer final output encodings of the fitted data to generate the target model.
[0081] The technical solution provided by the embodiments of the present invention can provide richer data information by encoding at least one type of training data in the training sample data group into multi-layer target encodings, improve the refinement degree of data discretization and fitting, and thus improve the quality of the data output by the trained large language model.
[0082] Embodiment III
[0083] Figure 3 It is a schematic structural diagram of a large language model encoding training device provided by Embodiment III of the present invention. As Figure 3 shown, the device includes:
[0084] A training data set acquisition module 310 for acquiring a training data set; wherein, the training data set contains at least two training sample data groups, and each training sample data group is composed of at least two corresponding types of training data;
[0085] A first-layer target encoding determination module 320 for determining the first-layer target encoding of the training data for at least one type of training data in each training sample data group;
[0086] A next-layer target encoding determination module 330 for using the first-layer target encoding as the current-layer target encoding, determining the current-layer encoding deviation of the current-layer target encoding, and determining the next-layer target encoding based on the current-layer encoding deviation;
[0087] An encoding deviation cycle determination module 340 is configured to update the target encoding of the next layer to the target encoding of the current layer, and return the current layer encoding deviation for determining the target encoding of the current layer, and determine the target encoding of the Nth layer; where N is an integer greater than or equal to 2;
[0088] A target model generation module 350 is configured to train a large language model based on the N-layer target encoding to generate a target model.
[0089] Optionally, the apparatus further includes:
[0090] A partial training data extraction module is configured to extract partial training data of at least one dimension for at least one type of training data in each training sample data group before determining the first layer target encoding of the training data;
[0091] A reinforced target encoding generation module is configured to encode the partial training data of at least one dimension respectively to generate at least one layer of reinforced target encoding;
[0092] The target model generation module includes:
[0093] A target model generation unit is configured to stack at least one layer of the reinforced target encoding and the N-layer target encoding to train the large language model to generate a target model.
[0094] Optionally, the target model generation unit is configured to:
[0095] Place at least one layer of the reinforced target encoding before the N-layer target encoding, and input at least part of the reinforced target encoding into the large language model prior to the N-layer target encoding for training the large language model.
[0096] Optionally, the apparatus further includes:
[0097] A minimum discrete unit generation module is configured to perform discretization processing on at least one type of training data to generate a plurality of minimum discrete units before determining the first layer target encoding of the training data in each training sample data group;
[0098] The first layer target encoding determination module is configured to:
[0099] Encode each minimum discrete unit of at least one type of training data in each training sample data group respectively to determine the first layer target encoding of the training data.
[0100] Optionally, the first layer target encoding determination module is configured to:
[0101] Determine the initial encoding of the first layer of the training data;
[0102] Compare the initial encoding of the first layer with each first encoding in the preset encoding codebook or the first standard encoding corresponding to the first encoding, and determine the first target encoding corresponding to the first encoding in the encoding codebook with the highest similarity to the initial encoding of the first layer;
[0103] Determine the target encoding of the first layer based on the first target encoding.
[0104] Optionally, the next layer target encoding determination module is used for:
[0105] Calculate the residual between the initial encoding of the current layer and the target encoding of the current layer, and use the residual as the current layer encoding deviation of the current layer encoding;
[0106] Compare the current layer encoding deviation with each second encoding in the preset next layer deviation codebook or the second standard encoding corresponding to the second encoding, determine the second target encoding corresponding to the second encoding in the next layer deviation codebook with the highest similarity to the current layer encoding deviation, and determine the next layer target encoding based on the second target encoding.
[0107] Optionally, determining the corresponding target encoding based on the encoding with the highest similarity in the codebook includes:
[0108] Directly determine the encoding with the highest similarity in the codebook as the target encoding; or,
[0109] Encode the encoding with the highest similarity in the codebook based on a preset encoding strategy to generate the target encoding.
[0110] Optionally, each training sample data includes input data and fitting data, or each training sample data group includes reference data, input data and fitting data;
[0111] The target model generation module is used for:
[0112] Input the input encoding of the non-fitting data and the first layer target encoding of the fitting data into the large language model to determine the first output encoding of the large language model; wherein, the non-fitting data includes the input data, or the non-fitting data includes the reference data and the input data;
[0113] Determine the second output encoding of the large language model based on all the previously input encodings and the second layer target encoding of the input fitting data; until the Nth output encoding of the large language model is determined based on all the previously input encodings of the large language model and the Nth layer target encoding of the input fitting data;
[0114] Iteratively optimize the large language model based on at least one N-layer target encoding of the fitting data and the corresponding number of output encodings to generate a target model.
[0115] Optionally, the fitting data corresponds to M N-layer target encodings, where M is an integer greater than 1;
[0116] The target model generation module is configured to:
[0117] Input the input encoding of the non-fitting data and the first first-layer target encoding of the fitting data into the large language model to determine the first output encoding of the large language model;
[0118] Based on all the previously input encodings and the first second-layer target encoding of the input fitting data and the second first-layer target encoding, determine the second output encoding of the large language model, where the first second-layer target encoding and the second first-layer target encoding of the fitting data are input in a superimposed summation or concatenation manner; until the (M + N - 1)-th output encoding of the large language model is determined based on all the previously input encodings of the large language model and the M-th N-layer target encoding of the input fitting data;
[0119] Combine the M + N - 1 output encodings of the large language model into M N-layer final output encodings;
[0120] Iteratively optimize the large language model based on at least M N-layer target encodings of the fitting data and the M N-layer final output encodings to generate a target model.
[0121] The large language model encoding training device provided by the embodiments of the present invention can execute the large language model encoding training method provided by any embodiment of the present invention, and has the corresponding functional modules and beneficial effects for executing the method.
[0122] Embodiment 4
[0123] Figure 4 FIG. shows a schematic structural diagram of an electronic device 10 that can be used to implement the embodiments of the present invention. The electronic device is intended to represent various forms of digital computers, such as, laptop computers, desktop computers, workstations, personal digital assistants, servers, blade servers, mainframe computers, and other suitable computers. The electronic device can also represent various forms of mobile devices, such as, personal digital processors, cellular phones, smart phones, wearable devices (such as helmets, glasses, watches, etc.) and other similar computing devices. The components shown herein, their connections and relationships, and their functions are only for illustration and are not intended to limit the implementation of the present invention described herein and / or claimed.
[0124] AsFigure 4 As shown in Figure 4 , the electronic device 10 includes at least one processor 11 and a memory communicatively connected to the at least one processor 11, such as a read-only memory (ROM) 12, a random access memory (RAM) 13, etc. The memory stores a computer program executable by the at least one processor. The processor 11 can execute various appropriate actions and processes according to the computer program stored in the read-only memory (ROM) 12 or the computer program loaded from the storage unit 18 into the random access memory (RAM) 13. In the RAM 13, various programs and data required for the operation of the electronic device 10 can also be stored. The processor 11, the ROM 12, and the RAM 13 are connected to each other via a bus 14. An input / output (I / O) interface 15 is also connected to the bus 14.
[0125] Multiple components in the electronic device 10 are connected to the I / O interface 15, including: an input unit 16, such as a keyboard, a mouse, etc.; an output unit 17, such as various types of displays, speakers, etc.; a storage unit 18, such as a magnetic disk, an optical disc, etc.; and a communication unit 19, such as a network card, a modem, a wireless communication transceiver, etc. The communication unit 19 allows the electronic device 10 to exchange information / data with other devices via a computer network such as the Internet and / or various telecommunication networks.
[0126] The processor 11 can be various general-purpose and / or special-purpose processing components with processing and computing capabilities. Some examples of the processor 11 include but are not limited to a central processing unit (CPU), a graphics processing unit (GPU), various dedicated artificial intelligence (AI) computing chips, various processors running machine learning model algorithms, a digital signal processor (DSP), and any appropriate processor, controller, microcontroller, etc. The processor 11 executes the various methods and processes described above, such as the large language model encoding training method.
[0127] In some embodiments, the large language model encoding training method can be implemented as a computer program tangibly embodied in a computer-readable storage medium, such as the storage unit 18. In some embodiments, part or all of the computer program can be loaded and / or installed onto the electronic device 10 via the ROM 12 and / or the communication unit 19. When the computer program is loaded into the RAM 13 and executed by the processor 11, one or more steps of the large language model encoding training method described above can be executed. Alternatively, in other embodiments, the processor 11 can be configured to execute the large language model encoding training method by any other appropriate means (e.g., by means of firmware).
[0128] The various embodiments of the systems and techniques described above in this specification can be implemented in digital electronic circuitry, integrated circuit systems, field programmable gate arrays (FPGAs), application specific integrated circuits (ASICs), application specific standard products (ASSPs), systems-on-chip (SOCs), complex programmable logic devices (CPLDs), computer hardware, firmware, software, and / or combinations thereof. These various embodiments can include: being implemented in one or more computer programs that are executable and / or interpretable on a programmable system including at least one programmable processor, which can be a special-purpose or general-purpose programmable processor that receives data and instructions from, and transmits data and instructions to, a storage system, at least one input device, and at least one output device.
[0129] The computer program for implementing the methods of the present invention can be written in any combination of one or more programming languages. These computer programs can be provided to a processor of a general purpose computer, special purpose computer, or other programmable data processing apparatus, such that the computer programs, when executed by the processor, cause the functions / operations specified in the flowchart and / or block diagram to be implemented. The computer program can be executed entirely on the machine, partly on the machine, as a stand-alone software package partly on the machine and partly on a remote machine or entirely on the remote machine or server.
[0130] In the context of the present invention, a computer-readable storage medium can be a tangible medium that can contain or store a computer program for use by or in connection with an instruction execution system, apparatus, or device. The computer-readable storage medium can include, but is not limited to, electronic, magnetic, optical, electromagnetic, infrared, or semiconductor systems, apparatus, or devices, or any suitable combination of the foregoing. Alternatively, the computer-readable storage medium can be a machine-readable signal medium. More specific examples of the machine-readable storage medium would include an electrical connection based on one or more wires, a portable computer diskette, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or Flash memory), an optical fiber, a portable compact disc read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the foregoing.
[0131] To provide interaction with a user, the systems and techniques described herein can be implemented on an electronic device having: a display device for displaying information to the user (e.g., a CRT (cathode ray tube) or LCD (liquid crystal display) monitor); and a keyboard and a pointing device (e.g., a mouse or a trackball) through which the user can provide input to the electronic device. Other kinds of devices can also be used to provide interaction with the user; for example, the feedback provided to the user can be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and input from the user can be received in any form (including acoustic input, voice input, or tactile input).
[0132] The systems and techniques described herein can be implemented in a computing system including backend components (e.g., as a data server), or a computing system including middleware components (e.g., an application server), or a computing system including frontend components (e.g., a user computer having a graphical user interface or a web browser through which the user can interact with an implementation of the systems and techniques described herein), or a computing system including any combination of such backend components, middleware components, or frontend components. The components of the system can be interconnected to each other by digital data communication in any form or medium (e.g., a communication network). Examples of communication networks include: local area network (LAN), wide area network (WAN), blockchain network, and the Internet.
[0133] The computing system can include a client and a server. The client and the server are generally far from each other and usually interact through a communication network. The client-server relationship is created by computer programs running on the respective computers and having a client-server relationship with each other. The server can be a cloud server, also known as a cloud computing server or a cloud host, which is a host product in the cloud computing service system and solves the defects of difficult management and weak business scalability existing in traditional physical hosts and VPS services.
[0134] It should be understood that various forms of the processes shown above can be used, with steps reordered, added, or deleted. For example, the steps recited in the present invention can be executed in parallel, sequentially, or in a different order, as long as the desired results of the technical solution of the present invention can be achieved, and no limitation is imposed herein.
[0135] The above specific embodiments do not constitute a limitation on the protection scope of the present invention. Those skilled in the art should understand that various modifications, combinations, sub-combinations, and substitutions can be made according to design requirements and other factors. Any modifications, equivalent substitutions, and improvements made within the spirit and principle of the present invention shall be included within the protection scope of the present invention.
Claims
1. A large language model encoding training method, characterized in that, Including: Extracting partial training data of at least one dimension for at least one type of training data; Encoding the partial training data of at least one dimension respectively to generate at least one layer of reinforcement target encoding; Determining the N-layer target encoding of the training data for at least one type of training data in each training sample data group; where N is an integer greater than or equal to 2; Stacking at least one layer of the reinforcement target encoding with the N-layer target encoding and training a large language model to generate a target model.
2. The method according to claim 1, wherein Determining the N-layer target encoding of the training data for at least one type of training data in each training sample data group includes: Determining the first-layer target encoding of the training data for at least one type of training data in each training sample data group; Taking the first-layer target encoding as the current-layer target encoding, determining the current-layer encoding deviation of the current-layer target encoding, and determining the next-layer target encoding based on the current-layer encoding deviation; where the current-layer encoding deviation is the difference between the current-layer initial encoding and the current-layer target encoding; Updating the next-layer target encoding to the current-layer target encoding and returning to execute determining the current-layer encoding deviation of the current-layer target encoding to determine the Nth-layer target encoding.
3. The method according to claim 1, wherein The step of stacking at least one layer of the reinforcement target encoding with the N-layer target encoding and training a large language model includes: Placing at least one layer of the reinforcement target encoding before the N-layer target encoding, and inputting at least part of the reinforcement target encoding prior to the N-layer target encoding into the large language model for training the large language model.
4. The method according to claim 2, wherein Before determining the first-layer target encoding of the training data for at least one type of training data in each training sample data group, it further includes: Performing discretization processing on the at least one type of training data to generate a plurality of minimum discrete units; Determining the first-layer target encoding of the training data for at least one type of training data in each training sample data group includes: Encoding each minimum discrete unit of the at least one type of training data in each training sample data group respectively to determine the first-layer target encoding of the training data.
5. The method according to claim 2, wherein Determining the first-layer target encoding of the training data includes: Determining the first-layer initial encoding of the training data; Comparing the first-layer initial encoding with each first encoding in a pre-set encoding codebook or the first standard encoding corresponding to the first encoding, and determining the first target encoding corresponding to the first encoding in the encoding codebook with the highest similarity to the first-layer initial encoding; Determining the first-layer target encoding based on the first target encoding.
6. The method according to claim 2, wherein Determining the current-layer encoding deviation of the current-layer target encoding and determining the next-layer target encoding based on the current-layer encoding deviation includes: Calculating the residual between the current-layer initial encoding and the current-layer target encoding and taking the residual as the current-layer encoding deviation of the current-layer encoding; Compare the current layer encoding deviation with each second encoding in the preset next layer deviation codebook or the second standard encoding corresponding to the second encoding, determine the second target encoding corresponding to the second encoding in the next layer deviation codebook with the highest similarity to the current layer encoding deviation, and determine the next layer target encoding based on the second target encoding.
7. The method according to claim 5 or 6, characterized in that, Determining the corresponding target encoding based on the encoding with the highest similarity in the codebook includes: Directly determining the encoding with the highest similarity in the codebook as the target encoding; or, Encoding the encoding with the highest similarity in the codebook based on a preset encoding strategy to generate the target encoding.
8. The method according to claim 1, wherein It further includes: Obtain a training data set; wherein, the training data set contains at least two training sample data groups, and each training sample data group consists of at least two corresponding types of training data; each training sample data includes input data and fitting data, or each training sample data group includes reference data, input data and fitting data; Training the large language model to generate a target model includes: Input the input encoding of the non-fitting data and the first layer target encoding of the fitting data into the large language model to determine the first output encoding of the large language model; wherein, the non-fitting data includes the input data, or the non-fitting data includes the reference data and the input data; Determine the second output encoding of the large language model based on all the previously input encodings and the second layer target encoding of the input fitting data; until the Nth output encoding of the large language model is determined based on all the previously input encodings of the large language model and the Nth layer target encoding of the input fitting data; Iteratively optimize the large language model based on at least one N-layer target encoding of the fitting data and the corresponding number of output encodings to generate a target model.
9. The method according to claim 8, wherein The fitting data corresponds to M N-layer target encodings, where M is an integer greater than 1, and the method further includes: Input the input encoding of the non-fitting data and the first first layer target encoding of the fitting data into the large language model to determine the first output encoding of the large language model; Determine the second output encoding of the large language model based on all the previously input encodings and the first second layer target encoding and the second first layer target encoding of the input fitting data, where the first second layer target encoding and the second first layer target encoding of the fitting data are input in a way of superposition summation or splicing; until the (M + N - 1)th output encoding of the large language model is determined based on all the previously input encodings of the large language model and the Mth Nth layer target encoding of the input fitting data; Combine the (M + N - 1) output encodings of the large language model into M N-layer final output encodings; Iteratively optimize the large language model based on at least M N-layer target encodings of the fitting data and M N-layer final output encodings to generate a target model.
10. A large language model encoding training device, characterized in that, It includes: A partial training data extraction module for extracting partial training data of at least one dimension for at least one type of training data; The enhanced target encoding generation module is used to encode at least part of the training data in at least one dimension respectively to generate at least one layer of enhanced target encoding; The N-layer target encoding generation module is used to determine the N-layer target encoding of at least one type of training data in each training sample data group; where N is an integer greater than or equal to 2; The target model generation module is used to stack at least one layer of the enhanced target encoding and the N-layer target encoding, train the large language model, and generate a target model.
Citation Information
Patent Citations
Language generation method and device, computer equipment and storage medium
CN111274764A
Language model training method and device and electronic equipment
CN115374780A
Text semantic representation method and system based on entity enhancement
CN116662480A
Rhythm annotation model training method, audio generation method, device and equipment
CN117316141A
Counterfeit voice detection method, system and device fused with large language model, and medium
CN117577119A