Coding training method and apparatus for large language model, and electronic device and storage medium
By employing a multi-layer target encoding training method, the problem of insufficient encoding fit in training large models with information-rich data is solved, thereby improving the quality of the output data.
Patent Information
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- SHANGHAI XIYU JIZHI TECH CO LTD
- Filing Date
- 2025-08-26
- Publication Date
- 2026-04-30
AI Technical Summary
When training large existing models, the encoding representation cannot effectively fit a large amount of useful information when faced with data rich in information, resulting in coarse output data that cannot meet user needs.
A multi-layer target encoding training method is adopted. By obtaining the training dataset, the multi-layer target encoding is determined, and the target model is generated based on the encoding bias.
It improves the fineness of data discretization and fitting, and enhances the quality of output data from large language models.
Smart Images

Figure CN2025116925_30042026_PF_FP_ABST
Abstract
Description
A method, apparatus, electronic device and storage medium for training large language model encoding
[0001] Cross-reference to related applications
[0002] This disclosure claims priority to Chinese Patent Application No. 2024114704136, filed on October 21, 2024, entitled “A Large Language Model Encoding Training Method and Apparatus”, the entire contents of which are incorporated herein by reference. Technical Field
[0003] This disclosure relates to the field of machine learning technology, and in particular to a method, apparatus, electronic device, and storage medium for encoding and training large language models. Background Technology
[0004] When training existing large-scale models, the training data is typically converted into encoded inputs. For example, each smallest unit of word segment (which could be a single character, word, or radical) in a text is converted into an encoding, thus transforming the text into a set of at least one encoding input to the large-scale model. While existing encoding methods are sufficient for relatively simple data, they are inadequate for data rich in information, such as music frames (each smallest unit of word segment) containing numerous instruments and / or vocals. Each frame of music possesses a wealth of key information such as pitch, timbre, range, etc. Fitting the data using a single encoding would result in the loss of significant useful information, leading to coarse output data from the trained large-scale model that falls far short of user needs. Summary of the Invention
[0005] This disclosure provides a method, apparatus, electronic device, and storage medium for training large language model encoding, which can provide richer data information, improve the fineness of data discretization and fitting, and thus improve the quality of the output data of the trained large language model.
[0006] According to one aspect of this disclosure, a method for training large language model encodings is provided, comprising:
[0007] Obtain a training dataset; wherein the training dataset contains at least two training sample data groups, and each training sample data group consists of at least two corresponding training data;
[0008] For each group of training sample data, a first-layer target encoding of the training data is determined;
[0009] The first layer target code is used as the current layer target code, the current layer code deviation of the current layer target code is determined, and the next layer target code is determined based on the current layer code deviation;
[0010] The next layer target code is updated to the current layer target code, and the current layer code deviation for determining the current layer target code is returned to determine the Nth layer target code; where N is an integer greater than or equal to 2;
[0011] The target model is trained based on the target encoding described in the N layers to generate the target model.
[0012] According to another aspect of this disclosure, a large language model encoding training apparatus is provided, comprising:
[0013] The training dataset acquisition module is configured to acquire a training dataset; wherein the training dataset contains at least two training sample data groups, and each training sample data group consists of at least two corresponding training data.
[0014] The first-layer target encoding determination module is configured to determine the first-layer target encoding of the training data for at least one type of training data in each training sample data group.
[0015] The next-layer target encoding determination module is configured to take the first-layer target encoding as the current-layer target encoding, determine the current-layer encoding deviation of the current-layer target encoding, and determine the next-layer target encoding based on the current-layer encoding deviation;
[0016] The encoding deviation cyclic determination module is configured to update the next layer target encoding to the current layer target encoding and return the current layer encoding deviation for determining the current layer target encoding, thereby determining the Nth layer target encoding; where N is an integer greater than or equal to 2;
[0017] The target model generation module is configured to train the large language model based on the target encoding described in the N layers to generate the target model.
[0018] According to another aspect of this disclosure, an electronic device is provided, the electronic device comprising:
[0019] At least one processor; and
[0020] A memory communicatively connected to the at least one processor; wherein,
[0021] The memory stores a computer program that can be executed by the at least one processor, the computer program being executed by the at least one processor to enable the at least one processor to perform the large language model encoding training method according to any embodiment of this disclosure.
[0022] According to another aspect of this disclosure, a computer-readable storage medium is provided that stores computer instructions configured to cause a processor to execute and implement the large language model encoding training method according to any embodiment of this disclosure.
[0023] This disclosure discloses a large language model encoding training scheme, which involves acquiring a training dataset. The training dataset contains at least two training sample data groups, each consisting of at least two corresponding training data types. For at least one training data type in each training sample data group, a first-layer target encoding is determined. This first-layer target encoding is used as the current-layer target encoding. A current-layer encoding deviation is determined for this current-layer target encoding, and a next-layer target encoding is determined based on the current-layer encoding deviation. The next-layer target encoding is updated to the current-layer target encoding, and the process returns to determine the current-layer encoding deviation, thus determining the Nth-layer target encoding. Here, N is an integer greater than or equal to 2. The large language model is trained based on the N layers of target encodings to generate a target model. The technical solution provided by this disclosure, by encoding at least one training data type in a training sample data group into multi-layer target encodings, can provide richer data information, improve the fineness of data discretization and fitting, and thus improve the quality of the output data of the trained large language model.
[0024] It should be understood that the description in this section is not intended to identify key or essential features of the embodiments of this disclosure, nor is it intended to limit the scope of this disclosure. Other features of this disclosure will become readily apparent from the following description. Attached Figure Description
[0025] To more clearly illustrate the technical solutions in the embodiments of this disclosure, the accompanying drawings used in the description of the embodiments will be briefly introduced below. Obviously, the accompanying drawings described below are only some embodiments of this disclosure. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0026] Figure 1 is a flowchart of a large language model encoding training method provided in an embodiment of this disclosure;
[0027] Figure 2 is a flowchart of a large language model encoding training method provided in an embodiment of this disclosure;
[0028] Figure 3 is a flowchart of a large language model encoding training method provided in an embodiment of this disclosure;
[0029] Figure 4 is a schematic diagram of the structure of a large language model coding training device provided in an embodiment of this disclosure;
[0030] Figure 5 is a schematic diagram of the structure of an electronic device that implements the large language model encoding training method of the present disclosure. Detailed Implementation
[0031] To enable those skilled in the art to better understand the present disclosure, the technical solutions of the present disclosure will be clearly and completely described below with reference to the accompanying drawings of the embodiments. Obviously, the described embodiments are only some embodiments of the present disclosure, and not all embodiments. Based on the embodiments of the present disclosure, all other embodiments obtained by those skilled in the art without creative effort should fall within the scope of protection of the present disclosure.
[0032] It should be noted that the terms "first," "second," etc., in the specification, claims, and accompanying drawings of this disclosure are used to distinguish similar objects and are not necessarily used to describe a specific order or sequence. It should be understood that such data can be interchanged where appropriate so that the embodiments of this disclosure described herein can be implemented in orders other than those illustrated or described herein. Furthermore, the terms "comprising" and "having," and any variations thereof, are intended to cover non-exclusive inclusion; for example, a process, method, system, product, or apparatus that comprises a series of steps or units is not necessarily limited to those steps or units explicitly listed, but may include other steps or units not explicitly listed or inherent to such processes, methods, products, or apparatus.
[0033] This disclosure provides a method for training large language model encoding. Figure 1 is a flowchart of a method for training large language model encoding provided in this disclosure. This embodiment is applicable to the encoding training of large language models. This method can be executed by a large language model encoding training device, which can be implemented in hardware and / or software and can be configured in an electronic device. As shown in Figure 1, the method includes:
[0034] S110. Obtain the training dataset; wherein the training dataset contains at least two training sample data groups, and each training sample data group consists of at least two corresponding training data.
[0035] In this embodiment, a training dataset is obtained, which contains multiple training sample data groups, each consisting of at least two types of data. For example, a training sample data group may include input data and fitted data; or, alternatively, a training sample data group may include reference data, input data, and fitted data. The input data is the data that needs to be used to output corresponding generated content based on the input content during subsequent large language model training, such as questions in question-and-answer data or lyrics in music generation data. The reference data is the data that needs to be used to output corresponding generated content based on some information during subsequent large language model training, such as reference images in text-to-image data, reference songs in music generation data, or reference accompaniment, reference singing style, etc., in reference songs, configured to enable the large language model to generate output data with the same style. The fitted data is the standard answer data, configured to be compared one-to-one with the output results of the large language model during subsequent large language model training, calculating the loss value, thereby correcting the probability of the next output result of the large model, and is the ideal data that the large language model needs to fit. The training dataset is a dataset that matches the application scenario of the large language model that will be trained later. For example, if the large language model to be trained later is a music generation model, each training sample data set in the training dataset can include fitted music and target lyrics, or it can refer to music, fitted music and target lyrics.
[0036] It should be noted that the embodiments of this disclosure do not limit the number of input data and the number of reference data included in the training sample data set. For example, when training a large language model for music generation, the input data in the training sample data set can be a piece of lyrics or multiple pieces of lyrics, and the reference data in the training sample data set can be a reference song, or reference accompaniment and reference singing style in a reference song, etc.
[0037] S120. For at least one training data in each of the training sample data groups, determine the first layer target encoding of the training data.
[0038] In this embodiment of the disclosure, for each training sample data group in the training dataset, a first-layer target encoding for at least one training data in the training sample data group is determined. It is understood that a first-layer target encoding for any one or more or all training data in the training sample data group can be determined, wherein the types of training data for the determined first-layer target encoding in each training sample data group in the training dataset can be the same or different. For example, an encoder can be used to encode at least one training data in the training sample data group to determine the first-layer target encoding, wherein the encoder may include a MERT encoder and a Mel encoder.
[0039] Optionally, before determining the first-layer target encoding of the training data for at least one training data in each training sample data group, the method further includes: discretizing the at least one training data to generate multiple smallest discrete units; determining the first-layer target encoding of the training data for at least one training data in each training sample data group includes: encoding each smallest discrete unit of the at least one training data in each training sample data group to determine the first-layer target encoding of the training data. For example, at least one training data in the training sample data group is discretized to generate at least two smallest discrete units, and then each of the at least two smallest discrete units is encoded by an encoder to determine the corresponding first target encoding. It is understood that the number of first target encodings is the same as the number of smallest discrete units.
[0040] Using a large language model trained subsequently as the music generation model, and taking a training sample data set consisting of reference music, fitted music, and target lyrics as an example, the process is illustrated below. The fitted music is discretized to generate at least two frames of fitted audio data. For example, the fitted music can be discretized into 10 frames per second; for instance, if the fitted music is a 10-second segment, it can be discretized into 100 frames. Alternatively, the fitted music can be discretized as a whole, for example, into 500 frames. The number of frames in the discretized fitted audio data can be determined based on the length of the reference music; the longer the fitted music, the more frames are needed. Each frame of fitted audio data is encoded using an encoder to generate a corresponding first-layer target code, where the number of first-layer target codes is the same as the number of frames in the fitted audio data. Optionally, the same encoder can be used to encode each frame of fitted audio data, or different encoders can be used. The encoder can be an open-source encoder, such as the MERT encoder or the Mel encoder. In this embodiment of the disclosure, the discretization processing of the target lyrics may include: segmenting the target lyrics into multiple segmentation units based on a preset word segmentation algorithm, wherein each segmentation unit may be a single character, a single word, or a single sentence. Each segmentation unit in the target lyrics is then encoded using an encoder to generate a first-layer target code corresponding to each segmentation unit. The method for determining the first-layer target code of the reference music can be similar to the method for fitting the first-layer target code of the music, and will not be elaborated further here.
[0041] In some embodiments, determining the first-layer target encoding of the training data includes: determining the first-layer initial encoding of the training data; comparing the first-layer initial encoding with each first encoding in a pre-set encoding codebook or a first standard encoding corresponding to the first encoding, and determining the first target encoding corresponding to the first encoding in the encoding codebook with the highest similarity to the first-layer initial encoding; and determining the first-layer target encoding based on the first target encoding.
[0042] For example, at least one training data in the training sample data set is encoded by an encoder, and the resulting encoded information is used as the first layer initial encoding of the training data. A pre-set encoding codebook is obtained, wherein the encoding codebook consists of multiple first codes. Optionally, the first codes in the encoding codebook can be standard codes (also known as special codes), or the standard codes can be replaced with special characters as the first codes, such as using the first standard code as code 0, the second standard code as code 1, and so on, to generate an encoding codebook such as [0,1,2,...,1023]. This can greatly simplify the complexity of the codebook, and at the same time, it can isolate the codes (i.e., special characters such as 0, 1, 2) and the standard codes, which is convenient for adjusting the format and / or content of the standard codes in the future. For example, the first standard code can be a multi-dimensional vector, or it can be any one or more of audio codes, image codes, character codes, etc.; or the first standard code was originally a 256-dimensional encoding vector, and if it is subsequently adjusted to be converted into a 1024-dimensional encoding vector, the codebook is still [0,1,2,...,1023].
[0043] The initial first-layer code is compared sequentially with each first code in the encoding codebook or the first standard code corresponding to each first code, and the similarity between the two is calculated. The first code in the encoding codebook with the highest similarity to the initial first-layer code is taken as the first target code. Alternatively, the similarity between the initial first-layer code and each first code in the encoding codebook can be calculated directly, and the first code in the encoding codebook with the highest similarity to the initial first-layer code is taken as the first target code; or, the similarity between the initial first-layer code and the first standard code corresponding to each first code in the encoding codebook can be calculated, and the first code in the encoding codebook corresponding to the first standard code with the highest similarity to the initial first-layer code is taken as the first target code. The first-layer target code for the training data is then determined based on the first target code. Specifically, the first target code in the encoding codebook can be directly used as the first-layer target code, or the first standard code corresponding to the first target code in the encoding codebook can be used as the first-layer target code, or the first target code in the encoding codebook can be encoded based on a preset encoding strategy to generate the first-layer target code corresponding to the first target code.
[0044] S130. Take the first layer target code as the current layer target code, determine the current layer code deviation of the current layer target code, and determine the next layer target code based on the current layer code deviation.
[0045] In this embodiment, the first-layer target encoding of the training data is updated to the current-layer target encoding, and the current-layer encoding deviation of the current-layer target encoding is determined. The current-layer encoding deviation can be the difference between the current-layer initial encoding and the current-layer target encoding. The next-layer target encoding of the training data is determined based on the current-layer encoding deviation. For example, the current-layer encoding deviation can be encoded based on a preset encoding algorithm to generate the next-layer target encoding of the training data. It is understood that the first-layer encoding deviation of the first-layer target encoding of the training data can be determined by using the residual between the first-layer initial encoding and the first-layer target encoding of the training data as the first-layer encoding deviation. The second-layer target encoding of the training data is determined based on the first-layer encoding deviation.
[0046] Optionally, determining the current layer coding deviation of the current layer target code and determining the next layer target code based on the current layer coding deviation includes: calculating the residual between the current layer initial code and the current layer target code, and using the residual as the current layer coding deviation of the current layer code; comparing the current layer coding deviation with each second code or the second standard code corresponding to the second code in a pre-set next layer deviation codebook, determining the second target code corresponding to the second code with the highest similarity to the current layer coding deviation in the next layer deviation codebook, and determining the next layer target code based on the second target code.
[0047] For example, the residual between the initial encoding and the target encoding of the current layer is used as the current layer encoding deviation. The next layer deviation cipherbook is obtained, where the next layer deviation cipherbook consists of multiple second encodings, which can be understood as residual encodings. Optionally, the second encodings in the deviation cipherbook can be standard residual encodings (also called special residual encodings), or the standard residual encodings can be replaced with special characters as the second encodings, such as using the first standard residual encoding as encoding 0, the second standard residual encoding as encoding 1, and so on, generating a deviation cipherbook such as [0,1,2,...,1023]. This greatly simplifies the complexity of the biased codebook and isolates the encoding (i.e., special characters such as 0, 1, 2) from the standard residual encoding, making it easier to adjust the format and / or content of the standard residual encoding in the future. For example, the first standard residual encoding can be a multi-dimensional vector, or it can be any one or more of audio encoding, image encoding, character encoding, etc.; or the first standard residual encoding is originally a 256-dimensional encoding vector, which is later adjusted to be a 1024-dimensional encoding vector. At this time, the biased codebook is still [0,1,2,...,1023].
[0048] The current layer encoding bias is compared sequentially with each second code or its corresponding second standard code in the next layer bias codebook, and the similarity between them is calculated. The second code in the next layer bias codebook with the highest similarity to the current layer encoding bias is taken as the second target code. This can be achieved by directly calculating the similarity between the current layer encoding bias and each second code in the next layer bias codebook, and taking the second code in the next layer bias codebook with the highest similarity to the current layer encoding bias as the second target code; or by calculating the similarity between the current layer encoding bias and the second standard code corresponding to each second code in the next layer bias codebook, and taking the second code in the next layer bias codebook corresponding to the second standard code with the highest similarity to the current layer encoding bias as the second target code. The next layer target code for the training data is then determined based on the second target code. Alternatively, the second target code in the next layer bias codebook can be directly used as the next layer target code, or the second standard code corresponding to the second target code in the next layer bias codebook can be used as the next layer target code, or the second target code in the next layer bias codebook can be encoded based on a preset encoding strategy to generate the next layer target code corresponding to the second target code.
[0049] S140. Update the next layer target code to the current layer target code, and return to the current layer code deviation for determining the current layer target code, thereby determining the Nth layer target code; where N is an integer greater than or equal to 2.
[0050] In this embodiment, the first-layer encoding deviation of the first-layer target code in the training data is compared with the second standard code corresponding to each second code in the second-layer deviation codebook. The second target code corresponding to the second code in the second-layer deviation codebook with the highest similarity to the first-layer encoding deviation is determined, and the second-layer target code is determined based on the second target code. The first-layer encoding deviation is used as the second-layer initial code, and the residual between the second initial code and the second-layer target code is calculated. This residual is used as the second-layer encoding deviation, i.e., the residual between the first-layer encoding deviation and the second-layer target code is used as the second-layer encoding deviation. For ease of description, the second code in the i-th layer deviation codebook can also be called the i-th code, and the standard code corresponding to the i-th code can be called the i-th standard code. Therefore, the second-layer encoding deviation is compared with the third standard code corresponding to each third code in the third-layer deviation codebook. The third target code corresponding to the third code in the third-layer deviation codebook with the highest similarity to the second-layer encoding deviation is determined, and the third-layer target code is determined based on the third target code. The second-layer coding bias is used as the initial coding of the third layer. The residual between the third initial coding and the third-layer target coding is calculated, and this residual is used as the third-layer coding bias. In other words, the residual between the second-layer coding bias and the third-layer target coding is used as the third-layer coding bias. This process is repeated until the Nth-layer target coding of the training data is determined, where N is an integer greater than or equal to 2.
[0051] It is understandable that, based on S120-S130, the N-layer target encoding corresponding to at least one training data point in the training sample data group can be determined. It should be noted that the number of layers in the target encoding of at least one training data point in each training sample data group in the training dataset can be the same or different.
[0052] S150. Train the large language model based on the target encoding described in the N layers to generate the target model.
[0053] In this embodiment, at least one N-layer target encoding corresponding to training data from each training sample data group in the training dataset is input into a large language model to train the large language model and generate a target model. The large language model can be any open-source large model, such as GPT. It should be noted that this embodiment does not limit the application scenario of the target model. For example, the target model can be a music generation model; in this case, each training sample data group in the training dataset configured to train the target model can include reference music, sample lyrics, and fitted music. Alternatively, the target model can be a text-to-image model, i.e., a model that generates images from text; in this case, each training sample data group in the training dataset configured to train the target model can include sample text and fitted images.
[0054] The large language model encoding training method of this disclosure includes obtaining a training dataset; wherein the training dataset contains at least two training sample data groups, each training sample data group consisting of at least two corresponding training data; for at least one training data in each training sample data group, determining a first-layer target encoding of the training data; using the first-layer target encoding as the current-layer target encoding, determining the current-layer encoding deviation of the current-layer target encoding, and determining the next-layer target encoding based on the current-layer encoding deviation; updating the next-layer target encoding to the current-layer target encoding, and returning to determine the current-layer encoding deviation of the current-layer target encoding to determine the Nth-layer target encoding; wherein N is an integer greater than or equal to 2; training the large language model based on the N layers of target encoding to generate a target model. The technical solution provided by this disclosure, by encoding at least one training data in the training sample data group into multi-layer target encodings, can provide richer data information, improve the fineness of data discretization and fitting, thereby improving the quality of the output data of the trained large language model.
[0055] In some embodiments, before determining the first layer target encoding for at least one training data in each training sample data group, the method further includes: extracting partial training data of at least one dimension for the at least one training data; encoding the partial training data of at least one dimension respectively to generate at least one layer of enhanced target encoding; and training a large language model based on the N layers of target encoding to generate a target model, including: superimposing the at least one layer of enhanced target encoding with the N layers of target encoding to train the large language model and generate the target model. The advantage of this configuration is that it can enhance the training effect of the large language model in certain dimensions, resulting in better training performance of the large language model.
[0056] In this embodiment of the disclosure, for at least one training data in each training sample data group, partial training data of at least one dimension is extracted from the training data, and the partial training data of each dimension is encoded to generate at least one layer of reinforcement target encoding. Alternatively, the partial training data of each dimension can be encoded as a whole to generate at least one layer of reinforcement target encoding, or the partial training data of each dimension can be discretized, and each generated smallest discrete unit is encoded to generate at least one layer of reinforcement target encoding. It should be noted that when the reinforcement target encoding is a multi-layer encoding, the method for determining the multi-layer reinforcement target encoding is the same as the method for determining the N-layer target encoding of the training data in the above embodiments, and will not be repeated here.
[0057] For example, if the training sample data set includes reference music, fitted music, and target lyrics, the fitted music can be split into a predetermined number of dimensions, such as splitting it into at least one dimension (e.g., a fitted music part consisting of an accompaniment and vocals). Each dimension of the fitted music part is then discretized to generate at least two frames of fitted audio data. Each frame of fitted audio data is then encoded using an encoder to generate at least one layer of enhancement target code. Similarly, the reference sample music is split into a predetermined number of dimensions, such as splitting it into at least one dimension (e.g., a reference music part consisting of an accompaniment and vocals). Each of these at least one dimension of reference music part is then uniformly encoded to generate at least one layer of enhancement target code. Alternatively, the at least one dimension of the split reference music part can be discretized to generate at least two frames of reference audio data. Each frame of reference audio data is then encoded using an encoder to generate at least one layer of enhancement target code. Finally, the target lyrics are encoded to generate at least one layer of enhancement target code. For example, the bass, midrange, and treble frequencies in a frame of audio data can be encoded at different layers to generate corresponding enhancement target codes, or different instruments and vocals in a frame of audio data can be separated and encoded at different layers to generate corresponding enhancement target codes.
[0058] In this embodiment of the disclosure, at least one layer of enhanced target encoding is superimposed with N layers of target encoding, and the superimposed encoding is input into a large language model to train the large language model and generate a target model. Optionally, superimposing at least one layer of enhanced target encoding with N layers of target encoding to train the large language model includes: placing at least one layer of enhanced target encoding before the N layers of target encoding, and inputting at least a portion of the enhanced target encoding into the large language model before the N layers of target encoding for training the large language model.
[0059] For example, taking the fitted music in the training sample data set as an example, the enhanced target encoding of the vocal part extracted from the fitted music is used as the first layer, and the four target encodings corresponding to the fitted music are used as the second, third, fourth, and fifth layers, respectively. The enhanced target encoding of the vocal part of the fitted music is first input into the large language model for prediction output. Then, the four target encodings corresponding to the fitted music are sequentially input into the large language model to continue training based on the four target encodings, generating the target model. The advantage of this setup is that it allows the large language model to first perform fitting output from the simple vocal part corresponding to the fitted music, and then perform fitting training for complex songs, obtaining a clearer fitting of the vocal part than simply fitting songs, thus achieving a better training effect for the large language model.
[0060] This disclosure also provides a large language model encoding training method. Figure 2 is a flowchart of a large language model encoding training method provided in this disclosure. As shown in Figure 2, the method includes:
[0061] S210. Obtain a training dataset; wherein the training dataset contains at least two training sample data groups, each training sample data group consists of at least two corresponding training data, and each training sample data includes input data and fitted data, or each training sample data group includes reference data, input data and fitted data.
[0062] S220. For at least one training data in each of the training sample data groups, determine the first layer target encoding of the training data.
[0063] S230. Take the first layer target code as the current layer target code, determine the current layer code deviation of the current layer target code, and determine the next layer target code based on the current layer code deviation.
[0064] S240. Update the next layer target code to the current layer target code, and return to the current layer code deviation for determining the current layer target code, thereby determining the Nth layer target code; where N is an integer greater than or equal to 2.
[0065] S250. Input the input encoding of the unfitted data and the first-layer target encoding of the fitted data into the large language model to determine the first output encoding of the large language model; wherein, the unfitted data includes the input data, or the unfitted data includes the reference data and the input data.
[0066] Each training sample data set can include input data and fitted data, or it can include reference data, input data, and fitted data. When the training sample data includes input data and fitted data, the input data is considered as unfitted data; when the training sample data includes reference data, input data, and fitted data, the reference data and input data are considered as unfitted data.
[0067] In this embodiment of the disclosure, the N-layer target encoding of the unfitted data is used as the input encoding of the unfitted data. The input encoding of the unfitted data and the first-layer target encoding of the fitted data are input into a large language model to train the large language model and obtain the first output encoding of the large language model. It can be understood that the first output encoding is the encoding output by the large language model after learning the input encoding of the unfitted data and the first-layer target encoding of the fitted data.
[0068] S260. Based on all previously input codes and the second-layer target code of the input fitted data, determine the second output code of the large language model; until the Nth output code of the large language model is determined based on all previously input codes of the large language model and the Nth-layer target code of the input fitted data.
[0069] In this embodiment of the disclosure, a second output code of the large language model is determined based on all previous input codes of the large language model and the second-layer target code of the fitted data input to the large language model. Similarly, a third output code of the large language model is determined based on all previous input codes of the large language model and the third-layer target code of the fitted data input to the large language model; a fourth-layer target code of the large language model is also determined based on all previous input codes of the large language model and the fourth-layer target code of the fitted data input to the large language model. This process continues until the Nth output code of the large language model is determined based on all previous input codes of the large language model and the Nth-layer target code of the fitted data input to the large language model.
[0070] S270. Based on at least one N-layer target encoding and the corresponding number of output encodings of the fitted data, the large language model is iteratively optimized to generate a target model.
[0071] In this embodiment of the disclosure, the N-layer target encoding of the fitted data may be one or more. When the fitted data corresponds to multiple N-layer target encodings, the output encoding corresponding to each N-layer target encoding is obtained through S260-S270. Based on at least one N-layer target encoding and the corresponding number of output encodings of the fitted data, the large language model is continuously iteratively optimized to generate the target model.
[0072] Optionally, the fitted data corresponds to M N-layer target codes, where M is an integer greater than 1. The method further includes: inputting the input codes of the unfitted data and the first first-layer target code of the fitted data into a large language model to determine the first output code of the large language model; determining the second output code of the large language model based on all previously input codes and the first second-layer target code and the second first-layer target code of the input fitted data, wherein the first second-layer target code and the second first-layer target code of the fitted data are input by superposition and summation or concatenation; until all previously input codes of the large language model and The Mth Nth layer target code of the input fitted data is used to determine the M+N-1th output code of the large language model; based on the M+N-1 fitted input codes obtained by superimposing, summing or concatenating the M Nth layer target codes of the fitted data and the corresponding M+N-1 output codes of the large language model, the large language model is iteratively optimized to generate a target model; or, the M+N-1 output codes of the large language model are combined into M Nth layer final output codes; based on at least M Nth layer target codes and M Nth layer final output codes of the fitted data, the large language model is iteratively optimized to generate a target model.
[0073] In this embodiment of the disclosure, the M N-layer target codes of the fitted data are superimposed or concatenated in a delayed manner and then input into the large language model.
[0074] Specifically, since the output data of the large language model depends on the input data of the large language model corresponding to the previous fitted data, a delay method is used to superimpose the M N-layer target codes corresponding to the fitted data into the large language model.
[0075] Specifically, the fitted data corresponds to M N-layer target codes. All codes of the unfitted data are used as input codes for the unfitted data. The input codes of the unfitted data and the first first-layer target code of the fitted data are input into the large language model to enable the large language model to make predictions and obtain the first output code of the large language model. This can be understood as the code output by the large language model after learning the input codes of the unfitted data and the first first-layer target code of the fitted data. Based on all previously input codes of the large language model, the first second-layer target code and the second first-layer target code of the fitted data, the second output code of the large language model is determined. Specifically, the first second-layer target code and the second first-layer target code of the fitted data are summed or concatenated. Based on this summed or concatenated code and all previously input codes, the second output code of the large language model is determined. This process continues until the (M+N-1)th output code of the large language model is determined based on all previously input codes of the large language model and the Mth Nth layer target code of the input fitted data.
[0076] For example, M=4, N=4, Table 1 is an input table for the delayed overlay of four 4-layer target codes corresponding to fitted data provided in an embodiment of this disclosure:
[0077] Table 1
[0078] After inputting the input codes of the unfitted data and the first target code t(11) of the fitted data into the large language model, the first output code of the large language model is determined. Then, the first target code t(12) of the second layer and the second target code t(21) of the first layer are superimposed and summed and input into the large language model. The large language model determines the second output code based on all the previously input codes and the code after superimposing and summing t(12) and t(21). The first third-layer target code t(13), the second second-layer target code t(22), and the third first-layer target code t(31) of the fitted data are summed and then input into the large language model. The large language model determines the third output code based on all previously input codes and the summed code of t(13), t(22), and t(31). The first fourth-layer target code t(14), the second third-layer target code t(23), the third second-layer target code t(32), and the fourth first-layer target code t(41) of the fitted data are summed and then input into the large language model. The large language model determines the fourth output code based on all previously input codes and the summed code of t(14), t(23), t(32), and t(41). This process is repeated to obtain the seventh output code of the large language model in the manner described in Table 1. It is understandable that by superimposing the M N-layer target codes of the fitted data into the large language model in a delayed manner, the large language model can first obtain a relatively coarse low-level target code for each coding unit in the fitted data, and then obtain a more refined high-level target code, which is beneficial to the prediction and training of the large language model.
[0079] In another embodiment, Figure 3 is a flowchart of a large language model encoding training method provided in this embodiment. As shown in Figure 3, the method includes:
[0080] S310. Obtain a training dataset; wherein the training dataset contains at least two training sample data groups, each training sample data group consists of at least two corresponding training data, and each training sample data includes input data and fitted data, or each training sample data group includes reference data, input data and fitted data.
[0081] S320. For at least one training data in each of the training sample data groups, determine the first layer target encoding of the training data.
[0082] S330. Take the first layer target code as the current layer target code, determine the current layer code deviation of the current layer target code, and determine the next layer target code based on the current layer code deviation.
[0083] S340. Update the next layer target code to the current layer target code, and return to the current layer code deviation for determining the current layer target code, thereby determining the Nth layer target code; where N is an integer greater than or equal to 2.
[0084] S350. Input the input encoding and the starting output encoding of the non-fit data into the large language model, and determine the first output encoding of the large language model; wherein, the non-fit data includes the input data, or the non-fit data includes the reference data and the input data.
[0085] Each training sample data set can include input data and fitted data, or it can include reference data, input data, and fitted data. When the training sample data includes input data and fitted data, the input data is considered as unfitted data; when the training sample data includes reference data, input data, and fitted data, the reference data and input data are considered as unfitted data.
[0086] In this embodiment, the N-layer target encoding of the unfitted data is used as the input encoding of the unfitted data. The input encoding of the unfitted data and the initial output encoding (a special encoding) are input into the large language model so that the large language model can make predictions and obtain the first output encoding of the large language model. It can be understood that the first output encoding is the first encoding output by the large language model after learning the input encoding of the unfitted data.
[0087] S360. Based on all previously input codes and the first-layer target code of the input fitted data, input the large language model and determine the second output code of the large language model.
[0088] In this embodiment of the disclosure, the second output code of the large language model is determined based on all the codes previously input to the large language model and the first-layer target code of the fitted data input to the large language model.
[0089] S370. Based on all previously input codes and the second-layer target code of the input fitted data, determine the third output code of the large language model.
[0090] In this embodiment of the disclosure, the third output code of the large language model is determined based on all the codes previously input to the large language model and the second-layer target code of the fitted data input to the large language model.
[0091] S380. Based on all previously input codes and the (N-1)th layer target code of the input fitted data, determine the Nth output code of the large language model; until based on all previously input codes of the large language model and the Nth layer target code of the input fitted data, determine the final output code of the large language model.
[0092] In this embodiment of the disclosure, the fourth output code of the large language model is determined based on all the codes previously input to the large language model and the third-layer target code of the fitted data input to the large language model; and so on, until the Nth output code of the large language model is determined based on all the codes previously input to the large language model and the (N-1)th-layer target code of the fitted data input to the large language model; and the final output code is determined based on all the codes previously input to the large language model and the Nth-layer target code of the fitted data input to the large language model.
[0093] S390. Based on at least one N-layer target encoding of the fitted data and the corresponding number of output encodings other than the final output encoding, the large language model is iteratively optimized to generate a target model.
[0094] In this embodiment of the disclosure, the N-layer target encoding of the fitted data may be one or more. When the fitted data corresponds to multiple N-layer target encodings, each N-layer target encoding is obtained through S350-S380 respectively. Based on at least one N-layer target encoding of the fitted data and the corresponding number of output encodings excluding the final output encoding, the large language model is continuously iteratively optimized to generate the target model.
[0095] Optionally, the fitted data corresponds to M N-layer target codes, where M is an integer greater than 1. The method further includes: inputting the input codes and starting output codes of the unfitted data into a large language model to determine the first output code of the large language model; determining the second output code of the large language model based on all previously input codes and the first first-layer target code of the input fitted data; determining the third output code of the large language model based on all previously input codes and the first second-layer and second first-layer target codes of the input fitted data, wherein the first second-layer and second first-layer target codes of the fitted data are input by superposition and summation or concatenation; and determining the third output code of the large language model based on all previously input codes of the large language model and the (M-1)th and (M-1)th N-layer target codes of the input fitted data. The target encoding is determined by identifying the (M+N-1)th output encoding of the large language model. This continues until the final output encoding of the large language model is determined, based on all previous input encodings of the large language model and the Mth Nth layer target encoding of the input fitted data. Based on the M+N-1 fitted input encodings obtained by superimposing or concatenating the M Nth layer target encodings of the fitted data and the M+N-1 output encodings of the large language model (excluding the final output encoding), the large language model is iteratively optimized to generate a target model. Alternatively, the M+N-1 output encodings of the large language model (excluding the final output encoding) are combined into M Nth layer final output encodings. Based on the M Nth layer target encodings and the M Nth layer final output encodings of the fitted data, the large language model is iteratively optimized to generate a target model.
[0096] Specifically, the fitted data corresponds to M N-layer target codes. All codes of the unfitted data are used as input codes for the unfitted data. The input codes of the unfitted data and the initial output code (a special type of code) are input into the large language model to enable the large language model to make predictions and obtain the first output code of the large language model. This first output code can be understood as the first code output by the large language model after learning the input codes of the unfitted data. Based on all previously input codes of the large language model and the first first-layer target code of the fitted data input into the large language model, the second output code of the large language model is determined. Based on all previously input codes of the large language model, and the first and second first-layer target codes of the fitted data input into the large language model, the third output code of the large language model is determined. Specifically, the first second-layer target code and the second first-layer target code of the fitted data are summed or concatenated. Based on this summed or concatenated code and all previously input codes, the third output code of the large language model is determined. Similarly, based on all the previous input codes of the large language model and the M-1th and M-1th layer target codes of the input fitted data, the M+N-1th output code of the large language model is determined, until the final output code of the large language model (another special code) is determined based on all the previous input codes of the large language model and the M-1th layer target code of the input fitted data.
[0097] For example, M=4, N=4, as shown in Table 1, the input codes of the unfitted data and the initial output codes are input into the large language model so that the large language model can make predictions and determine the first output code of the large language model. The input codes of the unfitted data and the first target code t(11) of the fitted data are input into the large language model to determine the second output code of the large language model. Then, the first target code t(12) of the fitted data and the second target code t(21) of the fitted data are summed and input into the large language model. The large language model determines the third output code based on all the previously input codes and the code after summing t(12) and t(21). The first third-layer target code t(13), the second second-layer target code t(22), and the third first-layer target code t(31) of the fitted data are summed and then input into the large language model. The large language model determines the fourth output code based on all the previously input codes and the code obtained by summing t(13), t(22), and t(31). The first fourth-layer target code t(14), the second third-layer target code t(23), the third second-layer target code t(32), and the fourth first-layer target code t(41) of the fitted data are summed and then input into the large language model. The large language model determines the fifth output code based on all the previously input codes and the code obtained by summing t(14), t(23), t(32), and t(41). Following this pattern, and in accordance with the method described in Table 1, the third fourth-layer target code t(34) and the fourth third-layer target code t(43) of the fitted data are summed and then input into the large language model. Based on all the previously input codes and the code obtained by summing t(34) and t(43), the large language model determines its seventh output code. This process continues until the final output code of the large language model is obtained, based on all the previously input codes and the fourth fourth-layer target code t(44) of the fitted data. It can be understood that inputting the M N-layer target codes of the fitted data into the large language model in a delayed manner allows the large language model to first obtain a relatively coarse low-level target code for each coding unit in the fitted data, and then obtain a more refined high-level target code, which is beneficial for the prediction and training of the large language model.
[0098] Optionally, the N-layer target codes of the unfitted data can be directly added or concatenated and then input into the large language model.
[0099] Since the output data of a large language model depends on all inputs of the non-fit data (input data, or reference data and input data), a conventional overlay or concatenation method can be used to input the N-layer target codes corresponding to the non-fit data into the large language model. This means inputting all N-layer target codes of the non-fit data into the large language model at once, allowing the large model to acquire all N-layer target codes of the non-fit data simultaneously, thus making the training of the large language model faster and more efficient. Optionally, a DELAY method can also be used to overlay the N-layer target codes of the non-fit data into the large language model for training. For example, Table 2 is a conventional overlay input table for four 4-layer target codes corresponding to non-fit data provided in an embodiment of this disclosure:
[0100] Table 2
[0101] When the N-layer target codes of the unfitted data are superimposed into the large language model in a conventional manner, the multi-dimensional encoding vectors obtained by adding or concatenating the N-layer target codes of each unit can be obtained. For example, adding 256-dimensional t(11), t(12), t(13), and t(14) results in a 256-dimensional vector, or concatenating them to obtain a 1024-dimensional vector. Similarly, adding 256-dimensional t(21), t(22), t(23), and t(24) results in a 256-dimensional vector, which is still a 256-dimensional vector, or concatenating them to obtain a 1024-dimensional vector. The multi-layer encoding vectors are then input into the large language model, and the output of the large language model is also a multi-layer encoding vector, thereby improving the accuracy of the output of the large language model.
[0102] The M+N-1 output codes of the large language model are combined into M N-layer final output codes. Then, based on at least M N-layer target codes and M N-layer final output codes of the fitted data, the large language model is iteratively optimized to generate the target model.
[0103] The technical solution provided in this disclosure, by encoding at least one training data in the training sample data group into a multi-layer target code, can provide richer data information, improve the fineness of data discretization and fitting, and thus improve the quality of the output data of the trained large language model.
[0104] This disclosure also provides a large language model coding training device. Figure 4 is a schematic diagram of the structure of a large language model coding training device provided in this disclosure. As shown in Figure 4, the device includes:
[0105] The training dataset acquisition module 310 is configured to acquire a training dataset; wherein the training dataset contains at least two training sample data groups, and each training sample data group consists of at least two corresponding training data.
[0106] The first-layer target encoding determination module 320 is configured to determine the first-layer target encoding of the training data for at least one training data in each training sample data group.
[0107] The next-layer target encoding determination module 330 is configured to take the first-layer target encoding as the current-layer target encoding, determine the current-layer encoding deviation of the current-layer target encoding, and determine the next-layer target encoding based on the current-layer encoding deviation;
[0108] The encoding deviation cyclic determination module 340 is configured to update the next layer target encoding to the current layer target encoding, and return to determine the current layer encoding deviation of the current layer target encoding, thereby determining the Nth layer target encoding; where N is an integer greater than or equal to 2;
[0109] The target model generation module 350 is configured to train the large language model based on the target encoding of the N layers to generate the target model.
[0110] Optionally, the device further includes:
[0111] A partial training data extraction module is configured to extract partial training data of at least one dimension for each of the at least one training data in each training sample data group before determining the first-layer target encoding of the training data.
[0112] The target encoding generation module is configured to encode a portion of the training data in at least one dimension separately to generate at least one layer of target encoding.
[0113] The target model generation module includes:
[0114] The target model generation unit is configured to superimpose at least one layer of the enhanced target encoding with N layers of the target encoding to train a large language model and generate a target model.
[0115] Optionally, the target model generation unit is configured as follows:
[0116] At least one layer of the enhanced target encoding is placed before the N layers of target encoding, and at least a portion of the enhanced target encoding is input into the large language model before the N layers of target encoding for training the large language model.
[0117] Optionally, the device further includes:
[0118] The minimum discrete unit generation module is configured to discretize the at least one training data for each training sample data group before determining the first layer target encoding of the training data to generate multiple minimum discrete units.
[0119] The first-layer target encoding determination module is configured as follows:
[0120] Encode each smallest discrete unit of at least one training data in each training sample data group to determine the first layer target encoding of the training data.
[0121] Optionally, the first-layer target encoding determination module is configured as follows:
[0122] Determine the first-layer initial encoding of the training data;
[0123] The first layer initial code is compared with each first code or the first standard code corresponding to the first code in the pre-set encoding codebook to determine the first target code corresponding to the first code in the encoding codebook that has the highest similarity to the first layer initial code;
[0124] The first layer target code is determined based on the first target code.
[0125] Optionally, the next-layer target encoding determination module is configured as follows:
[0126] Calculate the residual between the initial encoding of the current layer and the target encoding of the current layer, and use the residual as the current layer encoding deviation of the current layer encoding;
[0127] The current layer coding deviation is compared with each second code or the second standard code corresponding to the second code in the pre-set next layer deviation codebook to determine the second target code corresponding to the second code with the highest similarity to the current layer coding deviation in the next layer deviation codebook, and the next layer target code is determined based on the second target code.
[0128] Optionally, determining the target code corresponding to the code in the codebook that has the highest similarity to the initial code or code deviation includes:
[0129] Calculate the similarity between the initial code or coding deviation and each code in the codebook, and take the code in the codebook with the highest similarity to the initial code or coding deviation as the target code; or...
[0130] Calculate the similarity between the initial code or coding deviation and the standard code corresponding to each code in the codebook, and take the code in the codebook corresponding to the standard code with the highest similarity to the initial code or coding deviation as the target code.
[0131] Optionally, the corresponding layer target code is determined based on the target code with the highest similarity in the codebook, including:
[0132] The target code with the highest similarity in the codebook is directly determined as the layer target code; or...
[0133] The standard code corresponding to the target code with the highest similarity in the codebook is used as the layer target code; or...
[0134] The target code with the highest similarity in the codebook is encoded based on a preset encoding strategy to generate the layer target code.
[0135] Optionally, each training sample data includes input data and fitted data, or each training sample data group includes reference data, input data and fitted data;
[0136] The target model generation module is configured as follows:
[0137] The input encoding of the non-fitted data and the first-layer target encoding of the fitted data are input into the large language model to determine the first output encoding of the large language model; wherein, the non-fitted data includes the input data, or the non-fitted data includes the reference data and the input data;
[0138] Based on all previously input codes and the second-layer target code of the input fitted data, the second output code of the large language model is determined; until the Nth output code of the large language model is determined based on all previously input codes of the large language model and the Nth-layer target code of the input fitted data.
[0139] Based on at least one N-layer target encoding and the corresponding number of output encodings of the fitted data, the large language model is iteratively optimized to generate the target model.
[0140] Optionally, each training sample data includes input data and fitted data, or each training sample data group includes reference data, input data and fitted data;
[0141] The target model generation module is configured as follows:
[0142] The input encoding of the non-fit data and the initial output encoding are input into the large language model to determine the first output encoding of the large language model; wherein, the non-fit data includes the input data, or the non-fit data includes the reference data and the input data;
[0143] Based on all previously input codes and the first-layer target code of the input fitted data, input the large language model to determine the second output code of the large language model;
[0144] Based on all previously input codes and the second-layer target code of the input fitted data, the third output code of the large language model is determined;
[0145] Based on all previously input codes and the (N-1)th layer target code of the input fitted data, determine the Nth output code of the large language model; until based on all previously input codes of the large language model and the Nth layer target code of the input fitted data, determine the final output code of the large language model.
[0146] Based on at least one N-layer target encoding of the fitted data and the corresponding number of output encodings other than the final output encoding, the large language model is iteratively optimized to generate the target model.
[0147] Optionally, the fitted data corresponds to M N-layer target codes, where M is an integer greater than 1;
[0148] The target model generation module is configured as follows:
[0149] Based on the input encoding of the non-fitted data and the first first-layer target encoding of the fitted data, input the large language model to determine the first output encoding of the large language model;
[0150] Based on all previously input codes and the first and second first-layer target codes of the input fitted data, the second output code of the large language model is determined, wherein the first and second first-layer target codes of the fitted data are input by superposition and summation or concatenation; until the M+N-1th output code of the large language model is determined based on all previously input codes of the large language model and the Mth and Nth layer target codes of the input fitted data.
[0151] The M+N-1 output codes of the large language model are combined into M final output codes of N layers;
[0152] Based on at least M N-layer target codes and M N-layer final output codes of the fitted data, the large language model is iteratively optimized to generate the target model.
[0153] Optionally, the fitted data corresponds to M N-layer target codes, where M is an integer greater than 1;
[0154] The target model generation module is configured as follows:
[0155] Based on the input encoding and the initial output encoding of the nonfitted data, input the large language model to determine the first output encoding of the large language model;
[0156] Based on all previously input codes and the first target code of the first layer of the input fitted data, the second output code of the large language model is determined;
[0157] Based on all previously input codes and the first and second first-layer target codes of the input fitted data, the third output code of the large language model is determined, wherein the first and second first-layer target codes of the fitted data are input by superposition and summation or concatenation.
[0158] Based on all the previous input codes of the large language model and the (M-1)th and (M-1)th layer target codes of the input fitted data, the (M+N-1)th output code of the large language model is determined; until the final output code of the large language model is determined based on all the previous input codes of the large language model and the (M)th and (N)th layer target codes of the input fitted data.
[0159] The M+N-1 output codes output by the large language model, excluding the final output code, are combined into M N-layer final output codes;
[0160] Based on the M N-layer target codes and M N-layer final output codes of the fitted data, the large language model is iteratively optimized to generate the target model.
[0161] Optionally, the M N-layer target codes of the fitted data are superimposed or concatenated in a delayed manner and then input into the large language model.
[0162] Optionally, the N-layer target codes of the unfitted data can be directly added or concatenated and then input into the large language model.
[0163] The large language model coding training device provided in this disclosure can execute the large language model coding training method provided in any embodiment of this disclosure, and has the corresponding functional modules and beneficial effects of the execution method.
[0164] This disclosure provides an electronic device, and FIG5 shows a schematic diagram of the structure of an electronic device 10 that can be used to implement embodiments of this disclosure. The electronic device is intended to represent various forms of digital computers, such as laptop computers, desktop computers, workstations, personal digital assistants, servers, blade servers, mainframe computers, and other suitable computers. The electronic device may also represent various forms of mobile devices, such as personal digital processors, cellular phones, smartphones, wearable devices (such as helmets, glasses, watches, etc.), and other similar computing devices. The components shown herein, their connections and relationships, and their functions are merely examples and are not intended to limit the implementation of the disclosure described and / or claimed herein.
[0165] As shown in Figure 5, the electronic device 10 includes at least one processor 11 and a memory, such as a read-only memory (ROM) 12 or a random access memory (RAM) 13, communicatively connected to the at least one processor 11. The memory stores computer programs executable by the at least one processor. The processor 11 can perform various appropriate actions and processes based on the computer programs stored in the ROM 12 or loaded from storage unit 18 into the RAM 13. The RAM 13 can also store various programs and data required for the operation of the electronic device 10. The processor 11, ROM 12, and RAM 13 are interconnected via a bus 14. An input / output (I / O) interface 15 is also connected to the bus 14.
[0166] Multiple components in electronic device 10 are connected to I / O interface 15, including: input unit 16, such as keyboard, mouse, etc.; output unit 17, such as various types of displays, speakers, etc.; storage unit 18, such as disk, optical disk, etc.; and communication unit 19, such as network card, modem, wireless transceiver, etc. Communication unit 19 allows electronic device 10 to exchange information / data with other devices through computer networks such as the Internet and / or various telecommunications networks.
[0167] Processor 11 can be a variety of general-purpose and / or special-purpose processing components with processing and computing capabilities. Some examples of processor 11 include, but are not limited to, a central processing unit (CPU), a graphics processing unit (GPU), various special-purpose artificial intelligence (AI) computing chips, various processors running machine learning model algorithms, a digital signal processor (DSP), and any suitable processor, controller, microcontroller, etc. Processor 11 performs the various methods and processes described above, such as large language model encoding training methods.
[0168] In some embodiments, the large language model coding training method may be implemented as a computer program tangibly contained in a computer-readable storage medium, such as storage unit 18. In some embodiments, part or all of the computer program may be loaded and / or mounted on electronic device 10 via ROM 12 and / or communication unit 19. When the computer program is loaded into RAM 13 and executed by processor 11, one or more steps of the large language model coding training method described above may be performed. Alternatively, in other embodiments, processor 11 may be configured to perform the large language model coding training method by any other suitable means (e.g., by means of firmware).
[0169] Various embodiments of the systems and techniques described above herein can be implemented in digital electronic circuit systems, integrated circuit systems, field-programmable gate arrays (FPGAs), application-specific integrated circuits (ASICs), application-specific standard products (ASSPs), systems-on-a-chip (SoCs), payload-programmable logic devices (CPLDs), computer hardware, firmware, software, and / or combinations thereof. These various embodiments may include implementations in one or more computer programs that can be executed and / or interpreted on a programmable system including at least one programmable processor, which may be a dedicated or general-purpose programmable processor, capable of receiving data and instructions from a storage system, at least one input device, and at least one output device, and transmitting data and instructions to the storage system, the at least one input device, and the at least one output device.
[0170] Computer programs configured to implement the methods of this disclosure can be written in any combination of one or more programming languages. These computer programs can be provided to a processor of a general-purpose computer, a special-purpose computer, or other programmable data processing apparatus, such that when executed by the processor, the computer programs cause the functions / operations specified in the flowcharts and / or block diagrams to be performed. The computer programs can be executed entirely on a machine, partially on a machine, or as a standalone software package, partially on a machine and partially on a remote machine, or entirely on a remote machine or server.
[0171] In the context of this disclosure, a computer-readable storage medium can be a tangible medium that may contain or store a computer program for use by or in conjunction with an instruction execution system, apparatus, or device. A computer-readable storage medium can be, but is not limited to, electronic, magnetic, optical, electromagnetic, infrared, or semiconductor systems, apparatus, or devices, or any suitable combination of the foregoing. Alternatively, a computer-readable storage medium can be a machine-readable signal medium. More specific examples of machine-readable storage media include electrical connections based on one or more wires, portable computer disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fiber, portable compact disk read-only memory (CD-ROM), optical storage devices, magnetic storage devices, or any suitable combination of the foregoing.
[0172] To provide interaction with a user, the systems and techniques described herein can be implemented on an electronic device having: a display device configured to display information to a user (e.g., a CRT (cathode ray tube) or LCD (liquid crystal display) monitor); and a keyboard and pointing device (e.g., a mouse or trackball) through which the user provides input to the electronic device. Other types of devices can also be configured to provide interaction with a user; for example, the feedback provided to the user can be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and input from the user can be received in any form (including sound input, voice input, or tactile input).
[0173] The systems and technologies described herein can be implemented in computing systems that include backend components (e.g., as data servers), or computing systems that include middleware components (e.g., application servers), or computing systems that include frontend components (e.g., user computers with graphical user interfaces or web browsers through which users can interact with implementations of the systems and technologies described herein), or any combination of such backend, middleware, or frontend components. The components of the system can be interconnected via digital data communication of any form or medium (e.g., communication networks). Examples of communication networks include local area networks (LANs), wide area networks (WANs), blockchain networks, and the Internet.
[0174] A computing system can include clients and servers. Clients and servers are generally located far apart and typically interact through communication networks. The client-server relationship is created by computer programs running on the respective computers and having a client-server relationship with each other. The server can be a cloud server, also known as a cloud computing server or cloud host, which is a hosting product within the cloud computing service system to address the shortcomings of traditional physical hosts and VPS services, such as high management difficulty and weak business scalability.
[0175] It should be understood that the various forms of processes shown above can be used to rearrange, add, or delete steps. For example, the steps described in this disclosure can be executed in parallel, sequentially, or in different orders, as long as the desired result of the technical solution of this disclosure can be achieved, and this is not limited herein.
[0176] The specific embodiments described above do not constitute a limitation on the scope of protection of this disclosure. Those skilled in the art should understand that various modifications, combinations, sub-combinations, and substitutions can be made according to design requirements and other factors. Any modifications, equivalent substitutions, and improvements made within the spirit and principles of this disclosure should be included within the scope of protection of this disclosure. Industrial applicability
[0177] By adopting the above scheme, at least one training data in the training sample data group is encoded into a multi-layer target code, which can provide richer data information, improve the fineness of data discretization and fitting, and thus improve the quality of the output data of the large language model after training.
Claims
1. A method for training large language model encodings, characterized in that, include: Obtain a training dataset; wherein the training dataset contains at least two training sample data groups, and each training sample data group consists of at least two corresponding training data; For each group of training sample data, a first-layer target encoding of the training data is determined; The first layer target code is used as the current layer target code, the current layer code deviation of the current layer target code is determined, and the next layer target code is determined based on the current layer code deviation; The next layer target code is updated to the current layer target code, and the current layer code deviation for determining the current layer target code is returned to determine the Nth layer target code; where N is an integer greater than or equal to 2; The target model is trained based on the target encoding described in the N layers to generate the target model.
2. The method according to claim 1, characterized in that, Before determining the first-layer target encoding of the training data for at least one training data in each of the said training sample data groups, the method further includes: Extract partial training data of at least one dimension from the at least one of the training data; Encode at least one dimension of the training data separately to generate at least one layer of reinforcement target encoding.
3. The method according to claim 2, characterized in that, The large language model is trained based on the target encoding described in the N layers to generate the target model, including: The target model is generated by superimposing at least one layer of the enhanced target encoding with N layers of the target encoding and training the large language model.
4. The method according to claim 3, characterized in that, The step of superimposing at least one layer of the enhanced target encoding with N layers of the target encoding to train a large language model includes: At least one layer of the enhanced target encoding is placed before the N layers of target encoding, and at least a portion of the enhanced target encoding is input into the large language model before the N layers of target encoding for training the large language model.
5. The method according to any one of claims 1 to 4, characterized in that, Before determining the first-layer target encoding of the training data for at least one training data in each of the said training sample data groups, the method further includes: The at least one training data is discretized to generate multiple smallest discrete units; For at least one training data point in each of the said training sample data groups, determining the first-layer target encoding of the training data includes: Encode each smallest discrete unit of at least one training data in each training sample data group to determine the first layer target encoding of the training data.
6. The method according to any one of claims 1 to 5, characterized in that, Determining the first-layer target encoding of the training data includes: Determine the first-layer initial encoding of the training data; The first layer initial code is compared with each first code or the first standard code corresponding to the first code in the pre-set encoding codebook to determine the first target code corresponding to the first code in the encoding codebook that has the highest similarity to the first layer initial code; The first layer target code is determined based on the first target code.
7. The method according to claim 1, characterized in that, Determining the current layer coding deviation of the current layer target coding includes: Calculate the residual between the initial encoding of the current layer and the target encoding of the current layer, and use the residual as the current layer encoding deviation of the current layer encoding.
8. The method according to any one of claims 1 to 7, characterized in that, Determining the target encoding for the next layer based on the current layer encoding bias includes: The current layer coding deviation is compared with each second code or the second standard code corresponding to the second code in the pre-set next layer deviation codebook to determine the second target code corresponding to the second code with the highest similarity to the current layer coding deviation in the next layer deviation codebook, and the next layer target code is determined based on the second target code.
9. The method according to claim 6 or 8, characterized in that, The determination of the target code corresponding to the code in the codebook with the highest similarity to the initial code or the code deviation includes: Calculate the similarity between the initial code or coding deviation and each code in the codebook, and take the code in the codebook with the highest similarity to the initial code or coding deviation as the target code; or... Calculate the similarity between the initial code or coding deviation and the standard code corresponding to each code in the codebook, and take the code in the codebook corresponding to the standard code with the highest similarity to the initial code or coding deviation as the target code.
10. The method according to claim 6 or 8, characterized in that, The corresponding layer target code is determined based on the target code with the highest similarity in the codebook, including: The target code with the highest similarity in the codebook is directly determined as the layer target code; or... The standard code corresponding to the target code with the highest similarity in the codebook is used as the layer target code; or... The target code with the highest similarity in the codebook is encoded based on a preset encoding strategy to generate the layer target code.
11. The method according to claim 1, characterized in that, Each training sample data includes input data and fitted data, or each training sample data group includes reference data, input data, and fitted data. The step of training the large language model based on the N layers of target encoding to generate the target model includes: The input encoding of the non-fitted data and the first-layer target encoding of the fitted data are input into the large language model to determine the first output encoding of the large language model; wherein, the non-fitted data includes the input data, or the non-fitted data includes the reference data and the input data; Based on all previously input codes and the second-layer target code of the input fitted data, the second output code of the large language model is determined; until the Nth output code of the large language model is determined based on all previously input codes of the large language model and the Nth-layer target code of the input fitted data. Based on at least one N-layer target encoding and the corresponding number of output encodings of the fitted data, the large language model is iteratively optimized to generate the target model.
12. The method according to claim 1, characterized in that, Each training sample data includes input data and fitted data, or each training sample data group includes reference data, input data, and fitted data. The step of training the large language model based on the N layers of target encoding to generate the target model includes: The input encoding of the non-fit data and the initial output encoding are input into the large language model to determine the first output encoding of the large language model; wherein, the non-fit data includes the input data, or the non-fit data includes the reference data and the input data; Based on all previously input codes and the first-layer target code of the input fitted data, input the large language model to determine the second output code of the large language model; Based on all previously input codes and the second-layer target code of the input fitted data, the third output code of the large language model is determined; Based on all previously input codes and the (N-1)th layer target code of the input fitted data, determine the Nth output code of the large language model; until based on all previously input codes of the large language model and the Nth layer target code of the input fitted data, determine the final output code of the large language model; Based on at least one N-layer target encoding of the fitted data and the corresponding number of output encodings other than the final output encoding, the large language model is iteratively optimized to generate the target model.
13. The method according to claim 11, characterized in that, The fitted data corresponds to M N-layer target codes, where M is an integer greater than 1. The method further includes: Based on the input encoding of the non-fitted data and the first first-layer target encoding of the fitted data, input the large language model to determine the first output encoding of the large language model; Based on all previously input codes and the first and second first-layer target codes of the input fitted data, the second output code of the large language model is determined, wherein the first and second first-layer target codes of the fitted data are input by superposition and summation or concatenation; until the M+N-1th output code of the large language model is determined based on all previously input codes of the large language model and the Mth and Nth layer target codes of the input fitted data. Based on the M+N-1 fitted input codes obtained by superimposing and summing or concatenating the M N-layer target encoding inputs of the fitted data, and the M+N-1 output codes corresponding to the output of the large language model, the large language model is iteratively optimized to generate the target model; or, The M+N-1 output codes of the large language model are combined into M final output codes of N layers; Based on the M N-layer target codes and M N-layer final output codes of the fitted data, the large language model is iteratively optimized to generate the target model.
14. The method according to claim 12, characterized in that, The fitted data corresponds to M N-layer target codes, where M is an integer greater than 1. The method further includes: Based on the input encoding and the initial output encoding of the nonfitted data, input the large language model to determine the first output encoding of the large language model; Based on all previously input codes and the first target code of the first layer of the input fitted data, the second output code of the large language model is determined; Based on all previously input codes and the first and second first-layer target codes of the input fitted data, the third output code of the large language model is determined, wherein the first and second first-layer target codes of the fitted data are input by superposition and summation or concatenation. Based on all the previous input codes of the large language model and the (M-1)th and (M-1)th layer target codes of the input fitted data, the (M+N-1)th output code of the large language model is determined; until the final output code of the large language model is determined based on all the previous input codes of the large language model and the (M)th and (N)th layer target codes of the input fitted data. Based on the M+N-1 fitted input codes obtained by superimposing and summing or concatenating the M N-layer target encoding inputs of the fitted data, and the M+N-1 output codes of the large language model (excluding the final output code), the large language model is iteratively optimized to generate the target model; or, The M+N-1 output codes output by the large language model, excluding the final output code, are combined into M N-layer final output codes; Based on the M N-layer target codes and M N-layer final output codes of the fitted data, the large language model is iteratively optimized to generate the target model.
15. The method according to claim 13 or 14, characterized in that, The M N-layer target codes of the fitted data are superimposed or concatenated in a delayed manner and then input into the large language model.
16. The method according to claim 13 or 14, characterized in that, The N-layer target codes of the unfitted data are directly added or concatenated and then input into the large language model.
17. A large language model encoding training device, characterized in that, include: The training dataset acquisition module is configured to acquire a training dataset; wherein the training dataset contains at least two training sample data groups, and each training sample data group consists of at least two corresponding training data; The first-layer target encoding determination module is configured to determine the first-layer target encoding of the training data for at least one type of training data in each training sample data group. The next-layer target encoding determination module is configured to take the first-layer target encoding as the current-layer target encoding, determine the current-layer encoding deviation of the current-layer target encoding, and determine the next-layer target encoding based on the current-layer encoding deviation; The encoding deviation cyclic determination module is configured to update the next layer target encoding to the current layer target encoding and return the current layer encoding deviation used to determine the current layer target encoding, thereby determining the Nth layer target encoding; where N is an integer greater than or equal to 2; The target model generation module is configured to train the large language model based on the target encoding described in the N layers to generate the target model.
18. An electronic device, characterized in that, The electronic device includes: At least one processor; and A memory communicatively connected to the at least one processor; wherein, The memory stores a computer program that can be executed by the at least one processor, the computer program being executed by the at least one processor to enable the at least one processor to perform the large language model encoding training method as described in any one of claims 1 to 15.
19. A computer-readable storage medium, characterized in that... The computer-readable storage medium stores computer instructions configured to cause a processor to execute the large language model encoding training method as described in any one of claims 1 to 15.
Citation Information
Patent Citations
Training method and device of speech processing model
CN110503945A
Large language model training method and device, electronic equipment and storage medium
CN118211065A
Large language model coding training method and device
CN119204133A
Pre-trained language model fine-tuning method and apparatus and non-transitory computer-readable medium
EP4020305A1