Code generation method and device

Through the multi-layer target coding generation method, the problem of insufficient encoding representation in information-rich data training is solved, and the data is refined and fitted, and the output quality is improved.

CN120373392AActive Publication Date: 2025-07-25SHANGHAI XIYU JIZHI TECH CO LTD
View PDF 5 Cites 0 Cited by

Patent Information

Application Number
CN202510547064.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2024-10-21
Publication Date
2025-07-25
Estimated Expiration
2044-10-21

AI Technical Summary

Technical Problem

When training the existing large models, when facing data rich in information, the encoding representation method cannot meet the refined needs of output data, resulting in rough output data and unable to meet user needs.

Method used

Through training data encoding, the target encoding is determined layer by layer by layer, the encoding codebook and the deviation codebook are used to determine the target encoding layer by layer, calculate the encoding deviation, and generate richer data information.

Benefits of technology

It improves the degree of refinement of data discreteness and fitting, and improves the quality of data output by large language models.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120373392A_ABST
    Figure CN120373392A_ABST
Patent Text Reader

Abstract

The invention discloses a code generation method and device, and the method comprises the steps: comparing a first-layer initial code of training data with a first standard code in a preset coding password book, and determining a first target code corresponding to a first code with the highest similarity with the first-layer initial code in the coding password book; determining a first-layer target code based on the first target code; taking the first-layer target code as a current-layer target code, and taking a residual error between the current-layer initial code and the current-layer target code as a current-layer code deviation; and determining the Nth layer of target codes. According to the scheme, richer data information can be provided.
Need to check novelty before this filing date? Find Prior Art

Description

[0001] This application is a divisional application. The application number of the original application is 202411470413.6, the application date is October 21, 2024, and the invention title is "A Method and Device for Encoding Training of Large Language Models". Technical Field

[0002] The present invention relates to the technical field of machine learning, and in particular, to an encoding generation method and device. Background Art

[0003] When existing large models are trained, generally, training data needs to be converted into encodings and input into the large model. For example, each minimum word segmentation unit (which can be each character, each word, or each radical) in a piece of text is converted into an encoding, and then a piece of text is converted into a set of at least one encoding and input into the large model. The existing encoding representation methods are sufficient to meet the usage requirements after the large model is trained when facing data with relatively simple information content. However, when facing data with relatively rich information content, such as each frame (each minimum word segmentation unit) of music with numerous musical instruments and / or vocals, each frame of music has a large amount of key information such as pitch, timbre, pitch, and pitch range. If it is fitted with one encoding, a large amount of useful information will be lost, and as a result, the output data based on the trained large model will also become very rough, far from meeting the user's needs. Summary of the Invention

[0004] The present invention provides an encoding generation method and device. By encoding training data into multi-layer target encodings, more abundant data information can be provided, and the refinement degree of data discretization and fitting is improved.

[0005] According to one aspect of the present invention, an encoding generation method is provided, including:

[0006] Determine the first-layer initial encoding of the training data;

[0007] Compare the first-layer initial encoding with each first encoding in a pre-set encoding codebook or the first standard encoding corresponding to the first encoding, and determine the first target encoding corresponding to the first encoding in the encoding codebook that has the highest similarity to the first-layer initial encoding;

[0008] Determine the first-layer target encoding based on the first target encoding;

[0009] Take the first-layer target encoding as the current-layer target encoding, calculate the residual between the current-layer initial encoding and the current-layer target encoding, and use the residual as the current-layer encoding deviation of the current-layer encoding;

[0010] Compare the current layer coding deviation with each second coding in the preset next layer deviation codebook or the second standard coding corresponding to the second coding, determine the second target coding corresponding to the second coding with the highest similarity to the current layer coding deviation in the next layer deviation codebook, and determine the next layer target coding based on the second target coding;

[0011] Determine the target coding of the Nth layer; where N is an integer greater than or equal to 2.

[0012] According to another aspect of the present invention, there is provided an encoding generation device, including:

[0013] A first layer initial coding determination module, configured to determine the first layer initial coding of the training data;

[0014] A first target coding determination module, configured to compare the first layer initial coding with each first coding in the preset coding codebook or the first standard coding corresponding to the first coding, and determine the first target coding corresponding to the first coding with the highest similarity to the first layer initial coding in the coding codebook;

[0015] A first layer target coding determination module, configured to determine the first layer target coding based on the first target coding;

[0016] A current layer coding deviation determination module, configured to use the first layer target coding as the current layer target coding, calculate the residual between the current layer initial coding and the current layer target coding, and use the residual as the current layer coding deviation of the current layer coding;

[0017] A next layer target coding determination module, configured to compare the current layer coding deviation with each second coding in the preset next layer deviation codebook or the second standard coding corresponding to the second coding, determine the second target coding corresponding to the second coding with the highest similarity to the current layer coding deviation in the next layer deviation codebook, and determine the next layer target coding based on the second target coding;

[0018] An Nth layer target coding determination module, configured to determine the target coding of the Nth layer; where N is an integer greater than or equal to 2.

[0019] According to another aspect of the present invention, there is provided an electronic device, the electronic device includes:

[0020] At least one processor; and

[0021] A memory communicatively connected to the at least one processor; where

[0022] The memory stores a computer program executable by the at least one processor. When executed by the at least one processor, the computer program enables the at least one processor to execute the encoding generation method according to any embodiment of the present invention.

[0023] According to another aspect of the present invention, there is provided a computer-readable storage medium storing computer instructions for causing a processor to implement the encoding generation method according to any embodiment of the present invention when executed.

[0024] In the encoding generation scheme of the embodiments of the present invention, the first-layer initial encoding of the training data is determined; the first-layer initial encoding is compared with each first encoding in a preset encoding codebook or the first standard encoding corresponding to the first encoding to determine the first target encoding corresponding to the first encoding in the encoding codebook that has the highest similarity to the first-layer initial encoding; the first-layer target encoding is determined based on the first target encoding; the first-layer target encoding is used as the current-layer target encoding, the residual between the current-layer initial encoding and the current-layer target encoding is calculated, and the residual is used as the current-layer encoding deviation of the current-layer encoding; the current-layer encoding deviation is compared with each second encoding in a preset next-layer deviation codebook or the second standard encoding corresponding to the second encoding to determine the second target encoding corresponding to the second encoding in the next-layer deviation codebook that has the highest similarity to the current-layer encoding deviation, and the next-layer target encoding is determined based on the second target encoding; the Nth-layer target encoding is determined; where N is an integer greater than or equal to 2. The technical solution provided by the embodiments of the present invention can provide richer data information by encoding the training data into multi-layer target encodings, and improve the refinement degree of data discretization and fitting.

[0025] It should be understood that the content described in this part is not intended to identify the key or important features of the embodiments of the present invention, nor is it used to limit the scope of the present invention. Other features of the present invention will become easily understood through the following description. BRIEF DESCRIPTION OF THE DRAWINGS

[0026] To more clearly illustrate the technical solutions in the embodiments of the present invention, the following will briefly introduce the drawings required for the description of the embodiments. Obviously, the following drawings are only some embodiments of the present invention. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.

[0027] Figure 1 It is a flowchart of a large language model encoding training method provided in Embodiment 1 of the present invention;

[0028] Figure 2A flowchart of a large language model encoding training method provided in Embodiment 2 of the present invention;

[0029] Figure 3 A schematic structural diagram of a large language model encoding training device provided in Embodiment 3 of the present invention;

[0030] Figure 4 A schematic structural diagram of an electronic device for implementing the large language model encoding training method of the embodiments of the present invention. Detailed implementation manners

[0031] In order to enable those skilled in the art to better understand the solutions of the present invention, the technical solutions in the embodiments of the present invention will be clearly and completely described below in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of the present invention.

[0032] It should be noted that the terms "first", "second", etc. in the specification and claims of the present invention and the above drawings are used to distinguish similar objects, and do not necessarily need to describe a specific order or sequence. It should be understood that such data can be interchanged under appropriate circumstances so that the embodiments of the present invention described here can be implemented in an order other than those illustrated or described here. In addition, the terms "comprising" and "having" and any variations thereof are intended to cover non-exclusive inclusion. For example, a process, method, system, product or device comprising a series of steps or units does not necessarily have to be limited to those steps or units clearly listed, but may include other steps or units not clearly listed or inherent to these processes, methods, products or devices.

[0033] Embodiment 1

[0034] Figure 1 A flowchart of a large language model encoding training method is provided for Embodiment 1 of the present invention. This embodiment is applicable to the situation of encoding and training a large language model. This method can be executed by a large language model encoding training device, which can be implemented in the form of hardware and / or software, and the large language model encoding training device can be configured in an electronic device.

[0035] As Figure 1 shown, the method includes:

[0036] S110. Obtain a training data set; wherein, the training data set contains at least two training sample data groups, and each training sample data group is composed of at least two corresponding types of training data.

[0037] In an embodiment of the present invention, a training data set is obtained, where the training data set contains multiple training sample data groups, and each training sample data group is composed of at least two corresponding types of data. For example, a training sample data group may include input data and fitting data; or, a training sample data group may include reference data, input data, and fitting data. Among them, the input data is the data that needs to output corresponding generated content according to the input content during the subsequent training of the large language model, such as the questions in the question-and-answer data, and the lyrics part in the music generation data; the reference data is the data that needs to refer to some information in it to output corresponding generated content during the subsequent training of the large language model, such as the reference pictures in the text-to-image data, the reference songs in the music generation data, or the reference accompaniment, reference singing style, etc. in the reference songs, which is used to enable the large language model to generate output data with the same style; the fitting data is the standard answer data, which is used to compare with the output results of the large language model one by one during the subsequent training of the large language model, calculate the loss value, and thus correct the probability of the next output result of the large model. It is the ideal data that the large language model needs to fit. The training data set is a data set that matches the application scenario of the large language model to be used for subsequent training. For example, if the large language model to be trained subsequently is a music generation model, each training sample data group in the training data set may include fitting music and target lyrics, or may also refer to music, fitting music, and target lyrics.

[0038] It should be noted that the present invention embodiment does not limit the number of input data and the number of reference data included in the training sample data group. Exemplarily, when training a large language model for music generation, the input data in the training sample data group can be a piece of lyric content or multiple pieces of lyrics, and the reference data in the training sample data group can be a certain reference song, or reference accompaniment and reference singing style, etc. in the reference song.

[0039] S120. For at least one type of training data in each of the training sample data groups, determine the first-layer target encoding of the training data.

[0040] In an embodiment of the present invention, for each training sample data group in the training dataset, a first-layer target encoding of at least one type of training data in the training sample data group is determined. It can be understood that the first-layer target encoding of any one or more or all of the training data in the training sample data group can be determined, where the types of training data for which the first-layer target encoding is determined in each training sample data group in the training dataset can be the same or different. Exemplarily, an encoder (Encoder) can be used to encode at least one type of training data in the training sample data group to determine the first-layer target encoding, where the encoder can include a MERT encoder (MERT Encoder) and a Mel encoder (Mel Encoder).

[0041] Optionally, before determining the first-layer target encoding of the at least one type of training data for each of the training sample data groups, the method further includes: performing discretization processing on the at least one type of training data to generate a plurality of minimum discrete units; determining the first-layer target encoding of the at least one type of training data for each of the training sample data groups includes: encoding each of the minimum discrete units of the at least one type of training data for each of the training sample data groups respectively to determine the first-layer target encoding of the training data. Exemplarily, the at least one type of training data in the training sample data group is subjected to discretization processing to generate at least two minimum discrete units, and then the encoder encodes each of the at least two minimum discrete units respectively to determine the corresponding first target encoding. It can be understood that the number of first target encodings is the same as the number of minimum discrete units.

[0042] Taking the large language model after subsequent training as the music generation model, and the training sample data set consisting of reference music, fitting music and target lyrics as an example, an exemplary description is given. The fitting music is discretized to generate at least two frames of fitting audio data. Exemplarily, the fitting music can be discretized into 10 frames of fitting audio data per second. For example, if the fitting music is a 10-second piece of music, the fitting music can be discretized into 100 frames of fitting audio data. Additionally, the fitting music can be discretized as a whole. For example, the fitting music is discretized into 500 frames of fitting audio data. Among them, the number of frames of the discretized fitting audio data can be determined according to the length of the reference music. The longer the length of the fitting music, the more frames of the discretized fitting audio data. Each frame of fitting audio data is encoded by an encoder to generate the corresponding first-layer target encoding, where the number of the first-layer target encodings is the same as the number of frames of the fitting audio data. Optionally, the same encoder can be used to encode each frame of fitting audio data, or different encoders can be used to encode each frame of fitting audio data. Among them, the encoder can be an open-source encoder, such as the MERT encoder or the Mel encoder. In the embodiment of the present invention, the discretization process of the target lyrics may include: performing word segmentation on the target lyrics based on a preset word segmentation algorithm, and dividing the target lyrics into multiple word segmentation units, where the word segmentation units can be each character, each word, or each sentence. Each word segmentation unit in the target lyrics is encoded by the encoder to generate the first-layer target encoding corresponding to each word segmentation unit. The method for determining the first-layer target encoding of the reference music can be similar to that of the fitting music, which will not be elaborated here.

[0043] Optionally, determining the first-layer target encoding of the training data includes: determining the first-layer initial encoding of the training data; comparing the first-layer initial encoding with each first encoding in a preset encoding codebook or the first standard encoding corresponding to the first encoding to determine the first target encoding corresponding to the first encoding in the encoding codebook that has the highest similarity to the first-layer initial encoding; and determining the first-layer target encoding based on the first target encoding. Exemplarily, at least one type of training data in a training sample data group is encoded by an encoder, and the obtained encoding information is used as the first-layer initial encoding of the training data. A preset encoding codebook is obtained, where the encoding codebook is composed of multiple first encodings. Optionally, the first encoding in the encoding codebook can be a standard encoding (also referred to as a special encoding), or the standard encoding can be replaced with special characters. For example, the first standard encoding is used as encoding 0, the second standard encoding is used as encoding 1, and so on, to generate an encoding codebook such as [0, 1, 2,..., 1023]. This can greatly simplify the complexity of the codebook and at the same time isolate the encoding (i.e., special characters such as 0, 1, 2, etc.) and the standard encoding, facilitating future adjustment of the format and / or content of the standard encoding. For example, the first standard encoding can be a multi-dimensional vector, or an audio encoding, a graphic encoding, a character encoding, etc., either alone or in combination; or the first standard encoding was originally a 256-dimensional encoding vector, and if it is subsequently adjusted to a 1024-dimensional encoding vector, the codebook is still [0, 1, 2,..., 1023].

[0044] The first-layer initial encoding is successively compared with each first encoding in the encoding codebook or the first standard encoding corresponding to the first encoding to calculate the similarity between the two. The first encoding in the encoding codebook that has the highest similarity to the first-layer initial encoding is used as the first target encoding, and the first-layer target encoding of the training data is determined based on the first target encoding. Among them, the first target encoding in the encoding codebook can be directly used as the first-layer target encoding, or the first standard encoding corresponding to the first target encoding in the encoding codebook can be used as the first-layer target encoding, or the first target encoding in the encoding codebook can be encoded based on a preset encoding strategy to generate the first-layer target encoding corresponding to the standard data.

[0045] S130. Use the first-layer target encoding as the current-layer target encoding, determine the current-layer encoding deviation of the current-layer target encoding, and determine the next-layer target encoding based on the current-layer encoding deviation.

[0046] In an embodiment of the present invention, the first-layer target encoding of the training data is updated to the current-layer target encoding, and the current-layer encoding deviation of the current-layer target encoding is determined, where the current-layer encoding deviation may be the difference between the current-layer initial encoding and the current-layer target encoding. The next-layer target encoding of the training data is determined according to the current-layer encoding deviation. For example, the current-layer encoding deviation may be encoded based on a preset encoding algorithm to generate the next-layer target encoding of the training data. It can be understood that the first-layer encoding deviation of the first-layer target encoding of the training data is determined, where the residual between the first-layer initial encoding and the first-layer target encoding of the training data may be used as the first-layer encoding deviation. The second-layer target encoding of the training data is determined according to the first-layer encoding deviation.

[0047] Optionally, determining the current-layer encoding deviation of the current-layer target encoding and determining the next-layer target encoding based on the current-layer encoding deviation includes: calculating the residual between the current-layer initial encoding and the current-layer target encoding, and using the residual as the current-layer encoding deviation of the current-layer encoding; comparing the current-layer encoding deviation with each second encoding or the second standard encoding corresponding to the second encoding in a preset next-layer deviation codebook, determining the second target encoding corresponding to the second encoding with the highest similarity to the current-layer encoding deviation in the next-layer deviation codebook, and determining the next-layer target encoding based on the second target encoding.

[0048] Exemplarily, the residual between the current-layer initial encoding and the current-layer target encoding is used as the current-layer encoding deviation of the current-layer encoding. A next-layer deviation codebook is obtained, where the next-layer deviation codebook consists of multiple second encodings, and the second encoding can be understood as a residual encoding. Optionally, the second encoding in the deviation codebook may be a standard residual encoding (also referred to as a special residual encoding), or the standard residual encoding may be replaced with a special character. For example, the first standard residual encoding is used as encoding 0, the second standard residual encoding is used as encoding 1, and so on, to generate a deviation codebook such as [0, 1, 2,..., 1023]. This can greatly simplify the complexity of the deviation codebook, and at the same time, the encoding (i.e., special characters such as 0, 1, 2, etc.) and the standard residual encoding can be isolated, facilitating future adjustment of the format and / or content of the standard residual encoding. For example, the first standard residual encoding may be a multi-dimensional vector, or any one or more of an audio encoding, a graphic encoding, a character encoding, etc.; or the first standard residual encoding was originally a 256-dimensional encoding vector and was subsequently adjusted to a 1024-dimensional encoding vector. At this time, the deviation codebook is still [0, 1, 2,..., 1023].

[0049] Compare the current layer encoding deviation with each second encoding or the second standard encoding corresponding to the second encoding in the next layer deviation codebook in sequence, calculate the similarity between the two, take the second encoding with the highest similarity to the current layer encoding deviation in the next layer deviation codebook as the second target encoding, and determine the next layer target encoding of the training data based on the second target encoding. Among them, the second target encoding in the next layer deviation codebook can be directly used as the next layer target encoding, or the second standard encoding corresponding to the second target encoding in the next layer deviation codebook can be used as the next layer target encoding, or the second target encoding in the next layer deviation codebook can be encoded based on a preset encoding strategy to generate the next layer target encoding corresponding to the standard data.

[0050] S140. Update the next layer target encoding to the current layer target encoding, and return to execute the current layer encoding deviation for determining the current layer target encoding to determine the Nth layer target encoding; where N is an integer greater than or equal to 2.

[0051] In the embodiment of the present invention, compare the first layer encoding deviation of the first layer target encoding of the training data with the second standard encoding corresponding to each second encoding in the second layer deviation codebook, determine the second target encoding corresponding to the second encoding with the highest similarity to the first layer encoding deviation in the second layer deviation codebook, and determine the second layer target encoding based on the second target encoding. Take the first layer encoding deviation as the second layer initial encoding, calculate the residual between the second initial encoding and the second layer target encoding, and take this residual as the second layer encoding deviation, that is, take the residual between the first layer encoding deviation and the second layer target encoding as the second layer encoding deviation. For the convenience of description, the second encoding in the ith layer deviation codebook can also be called the ith encoding, and the standard encoding corresponding to the ith encoding can be called the ith standard encoding. Therefore, compare the second layer encoding deviation with the third standard encoding corresponding to each third encoding in the third layer deviation codebook, determine the third target encoding corresponding to the third encoding with the highest similarity to the second layer encoding deviation in the third layer deviation codebook, and determine the third layer target encoding based on the third target encoding. Take the second layer encoding deviation as the third layer initial encoding, calculate the residual between the third initial encoding and the third layer target encoding, and take this residual as the third layer encoding deviation, that is, take the residual between the second layer encoding deviation and the third layer target encoding as the third layer encoding deviation. Continuously loop according to the above method until the Nth layer target encoding of the training data is determined, where N is an integer greater than or equal to 2.

[0052] It can be understood that according to S120 - S130, the Nth layer target encoding corresponding to at least one kind of training data in the training sample data group can be determined. It should be noted that the number of layers of the target encoding of at least one kind of training data in each training sample data group in the training data set can be the same or different.

[0053] S150. Train a large language model based on the N - layer target encoding to generate a target model.

[0054] In an embodiment of the present invention, the N - layer target encoding corresponding to at least one type of training data in each training sample data group in the training data set is input into the large language model to train the large language model and generate a target model. Among them, the large language model can be any open - source large model, such as GPT, etc. It should be noted that the application scenario of the target model is not limited in the embodiments of the present invention. Exemplarily, the target model can be a music generation model. At this time, each training sample data group in the training data set used to train the target model may include reference music, sample lyrics, and fitting music; for another example, the target model can be a text - to - image model, that is, a model that generates images according to text. At this time, each training sample data group in the training data set used to train the target model may include sample text and fitting images.

[0055] The large language model encoding training method according to the embodiments of the present invention includes: obtaining a training data set; where the training data set contains at least two training sample data groups, and each training sample data group is composed of at least two types of corresponding training data; for at least one type of training data in each training sample data group, determining the first - layer target encoding of the training data; using the first - layer target encoding as the current - layer target encoding, determining the current - layer encoding deviation of the current - layer target encoding, and determining the next - layer target encoding based on the current - layer encoding deviation; updating the next - layer target encoding to the current - layer target encoding, and returning to execute determining the current - layer encoding deviation of the current - layer target encoding to determine the N - layer target encoding; where N is an integer greater than or equal to 2; training a large language model based on the N - layer target encoding to generate a target model. The technical solution provided by the embodiments of the present invention can provide richer data information by encoding at least one type of training data in the training sample data group into multi - layer target encoding, improve the refinement degree of data discretization and fitting, and thus improve the quality of the data output by the trained large language model.

[0056] In some embodiments, before determining the first-layer target encoding of at least one type of training data in each of the training sample data groups, the following steps are further included: extracting partial training data of at least one dimension from the at least one type of training data; encoding the partial training data of at least one dimension respectively to generate at least one layer of enhanced target encoding; training a large language model based on N layers of the target encoding to generate a target model, including: superimposing at least one layer of the enhanced target encoding and the N layers of the target encoding to train the large language model to generate a target model. The advantage of such a setting is that it can strengthen the training effect of the large language model in certain dimensions and obtain a better training effect of the large language model.

[0057] In the embodiments of the present invention, for at least one type of training data in each training sample data group, partial training data of at least one dimension is extracted from the training data, and the partial training data of each dimension is encoded respectively to generate at least one layer of enhanced target encoding. Among them, the partial training data of each dimension can be encoded as a whole to generate at least one layer of enhanced target encoding, or the partial training data of each dimension can be discretized, and each generated minimum discrete unit is encoded to generate at least one layer of enhanced target encoding. It should be noted that when the enhanced target encoding is multi-layer encoding, the determination method of the multi-layer enhanced target encoding is the same as the method for determining the N-layer target encoding of the training data in the above embodiments, and will not be elaborated here.

[0058] Exemplarily, if the training sample data group includes reference music, fitting music, and target lyrics, then a splitting operation of a preset number of dimensions can be performed on the fitting music. For example, at least one dimension of the fitting music part (such as the fitting music parts of the accompaniment part and the vocal part in two dimensions) can be split from the fitting music. The fitting music part of each dimension is discretized to generate at least two frames of fitting audio data, and each frame of fitting audio data is encoded by an encoder to generate at least one corresponding layer of enhanced target encoding. Similarly, a splitting operation of a preset number of dimensions is performed on the reference sample music. For example, at least one dimension of the reference music part (such as the reference music parts of the accompaniment part and the vocal part in two dimensions) is split from the reference music, and the whole of the at least one dimension of the split reference music part is uniformly encoded to generate at least one corresponding layer of enhanced target encoding. Or the at least one dimension of the split reference music part is discretized to generate at least two frames of reference audio data, and each frame of reference audio data is encoded by an encoder to generate at least one corresponding layer of enhanced target encoding. The target lyrics are encoded to generate at least one corresponding layer of enhanced target encoding. Also exemplarily, the bass, midrange, and treble in a frame of audio data are encoded in different layers to generate corresponding enhanced target encodings, or different instruments and different vocals in a frame of audio data are separated and individually encoded in different layers to generate corresponding enhanced target encodings.

[0059] In an embodiment of the present invention, at least one layer of enhanced target encoding is superimposed on N layers of target encoding, and the superimposed encoding is input into a large language model to train the large language model to generate a target model. Optionally, the superimposing at least one layer of the enhanced target encoding on N layers of the target encoding and training the large language model includes: placing at least one layer of the enhanced target encoding before N layers of the target encoding, and at least part of the enhanced target encoding is input into the large language model prior to the N layers of the target encoding for training the large language model.

[0060] Exemplarily, taking the fitting music in the training sample data group as an example, the enhanced target encoding of the vocal part extracted from the fitting music is used as the first layer, and the 4 layers of target encoding corresponding to the fitting music are used as the 2nd, 3rd, 4th, and 5th layers respectively. The enhanced target encoding of the vocal part of the fitting music is preferentially used as the first layer of encoding and input into the large language model for prediction output, and then the 4 layers of target encoding corresponding to the fitting music are sequentially input into the large language model to continue training the large language model based on the 4 layers of target encoding to generate a target model. The advantage of such a setting is that the large language model can first perform fitting output from the simple vocal part corresponding to the fitting music, and then perform fitting training on the complex song, obtaining a clearer fitting training of the vocal part than simply performing song fitting, thereby achieving a better training effect of the large language model.

[0061] Embodiment 2

[0062] Figure 2 The flowchart of a large language model encoding training method provided by Embodiment 2 of the present invention is as shown in Figure 2 and the method includes:

[0063] S210. Obtain a training data set; wherein, the training data set contains at least two training sample data groups, each of the training sample data groups consists of at least two corresponding types of training data, each of the training samples includes input data and fitting data, or each of the training sample data groups includes reference data, input data and fitting data.

[0064] S220. For at least one type of training data in each of the training sample data groups, determine the first-layer target encoding of the training data.

[0065] S230. Take the first-layer target encoding as the current-layer target encoding, determine the current-layer encoding deviation of the current-layer target encoding, and determine the next-layer target encoding based on the current-layer encoding deviation.

[0066] S240. Update the next-layer target encoding to the current-layer target encoding, and return to execute to determine the current-layer encoding deviation of the current-layer target encoding to determine the Nth-layer target encoding; wherein, N is an integer greater than or equal to 2.

[0067] S250. Input the input encoding of the non-fitting data and the first-layer target encoding of the fitting data into the large language model to determine the first output encoding of the large language model; wherein, the non-fitting data includes the input data, or the non-fitting data includes the reference data and the input data.

[0068] Each training sample data may include input data and fitting data, or may include reference data, input data and fitting data. When the training sample data includes input data and fitting data, the input data is used as the non-fitting data; when the training sample data includes reference data, input data and fitting data, the reference data and the input data are used as the non-fitting data.

[0069] In the embodiment of the present invention, the Nth-layer target encoding of the non-fitting data is used as the input encoding of the non-fitting data, and the input encoding of the non-fitting data and the first-layer target encoding of the fitting data are input into the large language model to train the large language model to obtain the first output encoding of the large language model. It can be understood that the first output encoding is the encoding output by the large language model after learning the input encoding of the non-fitting data and the first-layer target encoding of the fitting data.

[0070] S260. Determine the second output encoding of the large language model based on the second-layer target encoding of all the encodings input previously and the fitting data input; and so on until the Nth output encoding of the large language model is determined based on the Nth-layer target encoding of all the encodings input previously and the fitting data input to the large language model.

[0071] In an embodiment of the present invention, the second output encoding of the large language model is determined based on the second-layer target encoding of all the encodings input previously and the fitting data input to the large language model. Similarly, the third output encoding of the large language model is determined based on the third-layer target encoding of all the encodings input previously and the fitting data input to the large language model; the fourth output encoding of the large language model is determined based on the fourth-layer target encoding of all the encodings input previously and the fitting data input to the large language model. And so on until the Nth output encoding of the large language model is determined based on the Nth-layer target encoding of all the encodings input previously and the fitting data input to the large language model.

[0072] S270. Iteratively optimize the large language model based on at least one Nth-layer target encoding of the fitting data and the corresponding number of output encodings to generate a target model.

[0073] In an embodiment of the present invention, the number of Nth-layer target encodings of the fitting data may be one or multiple. When the fitting data corresponds to multiple Nth-layer target encodings, the output encodings corresponding to each Nth-layer target encoding are obtained through S260 - S270. Based on at least one Nth-layer target encoding of the fitting data and the corresponding number of output encodings, the large language model is continuously iteratively optimized to generate a target model.

[0074] Optionally, the fitting data corresponds to M Nth-layer target encodings, where M is an integer greater than 1. The method further includes: inputting the input encoding of the non-fitting data and the first first-layer target encoding of the fitting data into the large language model to determine the first output encoding of the large language model; determining the second output encoding of the large language model based on all the encodings input previously and the first second-layer target encoding and the second first-layer target encoding of the fitting data, where the first second-layer target encoding and the second first-layer target encoding of the fitting data are input in a way of superposition and summation or splicing; and so on until the (M + N - 1)th output encoding of the large language model is determined based on all the encodings input previously and the Mth Nth-layer target encoding of the fitting data; combining the (M + N - 1) output encodings of the large language model into M Nth-layer final output encodings; and iteratively optimizing the large language model based on at least M Nth-layer target encodings of the fitting data and the M Nth-layer final output encodings to generate a target model.

[0075] In the embodiments of the present invention, since the output data of the large language model depends on the input data of the large language model corresponding to the previous fitting data, therefore, the M N-layer target encodings corresponding to the fitting data are input into the large language model in a delayed manner. Specifically, there are M N-layer target encodings corresponding to the fitting data. All the N-layer target encodings of the non-fitting data are used as the input encodings of the non-fitting data. The input encodings of the non-fitting data and the first first-layer target encoding of the fitting data are input into the large language model to enable the large language model to make a prediction and obtain the first output encoding of the large language model. It can be understood that the first output encoding is the encoding output by the large language model after learning the input encodings of the non-fitting data and the first first-layer target encoding of the fitting data. Based on all the encodings previously input to the large language model, the first second-layer target encoding and the second first-layer target encoding of the fitting data input to the large language model, the second output encoding of the large language model is determined. Specifically, after the first second-layer target encoding and the second first-layer target encoding of the fitting data are superimposed and summed or concatenated, the second output encoding of the large language model is determined based on the superimposed and summed or concatenated encoding and all the encodings previously input. By analogy, until the M+N-1th output encoding of the large language model is determined based on all the encodings previously input to the large language model and the Mth N-layer target encoding of the fitting data input.

[0076] Exemplarily, M = 4 and N = 4. Table 1 is a table of the superimposed input in a delayed manner of 4 4-layer target encodings corresponding to the fitting data provided by the embodiments of the present invention:

[0077] Table 1

[0078] t(14) t(24) t(34) t(44) t(13) t(23) t(33) t(43) t(12) t(22) t(32) t(42) t(11) t(21) t(31) t(41)

[0079] After inputting the input encoding of the non-fitting data and the target encoding t(11) of the first first layer of the fitting data into the large language model, the first output encoding of the large language model is determined. Then, the sum of the first second layer target encoding t(12) and the second first layer target encoding t(21) is input into the large language model. Based on all the previously input encodings and the encoding after the superposition of t(12) and t(21), the large language model determines the second output encoding. The sum of the first third layer target encoding t(13), the second second layer target encoding t(22), and the third first layer target encoding t(31) of the fitting data is input into the large language model. Based on all the previously input encodings and the encoding after the superposition of t(13), t(22), and t(31), the large language model determines the third output encoding; the sum of the first fourth layer target encoding t(14), the second third layer target encoding t(23), the third second layer target encoding t(32), and the fourth first layer target encoding t(41) of the fitting data is input into the large language model. Based on all the previously input encodings and the encoding after the superposition of t(14), t(23), t(32), and t(41), the large language model determines the fourth output encoding. And so on, in the manner described in Table 1, the 7th output encoding of the large language model is obtained. It can be understood that inputting the M N-layer target encodings of the fitting data into the large language model in a delayed superposition manner can enable the large language model to first obtain the target encodings of each unit of the low-level layer and then obtain the target encodings of the high-level layer, which is beneficial to the training of the large language model.

[0080] Since the output data of the large language model depends on all the inputs of the non-fitting data (input data, or reference data and input data), therefore, the N-layer target encodings corresponding to the non-fitting data can be input into the large language model by using the conventional superposition or splicing method, that is, all the N-layer target encodings of the non-fitting data are input into the large language model at one time, which can enable the large model to obtain all the N-layer target encodings of the non-fitting data at one time, thus making the training of the large language model more efficient and fast. Optionally, the N-layer target encodings of the non-fitting data can also be input into the large language model by using the DELAY method for superposition to train the large language model. Exemplarily, Table 2 is a table for inputting the superposition of 4 four-layer target encodings corresponding to the non-fitting data provided by an embodiment of the present invention in a conventional manner:

[0081] Table 2

[0082] t(14) t(24) t(34) t(44) t(13) t(23) t(33) t(43) t(12) t(22) t(32) t(42) t(11) t(21) t(31) t(41)

[0083] Among them, when the N-layer target encoding of non-fitted data is stacked and input into the large language model in a conventional manner, the multi-dimensional encoding vectors obtained by adding or concatenating the N-layer target encodings of each unit can be used. For example, adding the 256-dimensional t(11), t(12), t(13), and t(14) results in a 256-dimensional vector, or concatenating them results in a 1024-dimensional vector. Adding the 256-dimensional t(21), t(22), t(23), and t(24) results in a 256-dimensional vector, which is still a 256-dimensional vector, or concatenating them results in a 1024-dimensional vector. The multi-layer encoding vectors are input into the large language model. Among them, the output of the large language model is also a multi-layer encoding vector, thereby improving the accuracy of the output of the large language model.

[0084] Combine the M + N - 1 output encodings of the large language model into M N-layer final output encodings, and then iteratively optimize the large language model based on at least M N-layer target encodings and M N-layer final output encodings of the fitted data to generate the target model.

[0085] The technical solution provided by the embodiments of the present invention can provide richer data information by encoding at least one type of training data in the training sample data group into multi-layer target encodings, improve the refinement degree of data discretization and fitting, and thus improve the quality of the output data of the trained large language model.

[0086] Embodiment III

[0087] Figure 3 It is a schematic structural diagram of a large language model encoding training device provided by Embodiment III of the present invention. As Figure 3 shown, the device includes:

[0088] A training dataset acquisition module 310, configured to acquire a training dataset; wherein, the training dataset contains at least two training sample data groups, and each training sample data group is composed of at least two corresponding types of training data;

[0089] A first-layer target encoding determination module 320, configured to determine the first-layer target encoding of the training data for at least one type of training data in each training sample data group;

[0090] A next-layer target encoding determination module 330, configured to use the first-layer target encoding as the current-layer target encoding, determine the current-layer encoding deviation of the current-layer target encoding, and determine the next-layer target encoding based on the current-layer encoding deviation;

[0091] The coding deviation loop determination module 340 is configured to update the target coding of the next layer to the target coding of the current layer, and return the coding deviation of the current layer for determining the target coding of the current layer, thereby determining the target coding of the Nth layer; where N is an integer greater than or equal to 2;

[0092] The target model generation module 350 is configured to train a large language model based on the N layers of the target coding to generate a target model.

[0093] Optionally, the device further includes:

[0094] The partial training data extraction module is configured to extract partial training data of at least one dimension for at least one type of training data in each training sample data group before determining the first-layer target coding of the training data;

[0095] The enhanced target coding generation module is configured to encode the partial training data of at least one dimension respectively to generate at least one layer of enhanced target coding;

[0096] The target model generation module includes:

[0097] The target model generation unit is configured to stack at least one layer of the enhanced target coding and the N layers of the target coding to train the large language model to generate a target model.

[0098] Optionally, the target model generation unit is configured to:

[0099] Place at least one layer of the enhanced target coding before the N layers of the target coding, and input at least part of the enhanced target coding into the large language model prior to the N layers of the target coding for training the large language model.

[0100] Optionally, the device further includes:

[0101] The minimum discrete unit generation module is configured to discretize at least one type of training data to generate a plurality of minimum discrete units before determining the first-layer target coding of the training data for each training sample data group;

[0102] The first-layer target coding determination module is configured to:

[0103] Encode each minimum discrete unit of at least one type of training data in each training sample data group respectively to determine the first-layer target coding of the training data.

[0104] Optionally, the first-layer target coding determination module is configured to:

[0105] Determine the initial encoding of the first layer of the training data;

[0106] Compare the initial encoding of the first layer with each first encoding in the preset encoding codebook or the first standard encoding corresponding to the first encoding, and determine the first target encoding corresponding to the first encoding with the highest similarity to the initial encoding of the first layer in the encoding codebook;

[0107] Determine the target encoding of the first layer based on the first target encoding.

[0108] Optionally, the next-layer target encoding determination module is configured to:

[0109] Calculate the residual between the initial encoding of the current layer and the target encoding of the current layer, and use the residual as the current-layer encoding deviation of the current layer encoding;

[0110] Compare the current-layer encoding deviation with each second encoding in the preset next-layer deviation codebook or the second standard encoding corresponding to the second encoding, determine the second target encoding corresponding to the second encoding with the highest similarity to the current-layer encoding deviation in the next-layer deviation codebook, and determine the next-layer target encoding based on the second target encoding.

[0111] Optionally, determining the corresponding target encoding based on the encoding with the highest similarity in the codebook includes:

[0112] Directly determine the encoding with the highest similarity in the codebook as the target encoding; or,

[0113] Encode the encoding with the highest similarity in the codebook based on a preset encoding strategy to generate the target encoding.

[0114] Optionally, each training sample data includes input data and fitting data, or each training sample data group includes reference data, input data, and fitting data;

[0115] The target model generation module is configured to:

[0116] Input the input encoding of the non-fitting data and the first-layer target encoding of the fitting data into the large language model to determine the first output encoding of the large language model; wherein, the non-fitting data includes the input data, or the non-fitting data includes the reference data and the input data;

[0117] Determine the second output encoding of the large language model based on all the previously input encodings and the second-layer target encoding of the input fitting data; until the Nth output encoding of the large language model is determined based on all the previously input encodings of the large language model and the Nth-layer target encoding of the input fitting data;

[0118] Iteratively optimize the large language model based on at least one N-layer target encoding of the fitting data and the corresponding number of output encodings to generate a target model.

[0119] Optionally, the fitting data corresponds to M N-layer target encodings, where M is an integer greater than 1;

[0120] The target model generation module is used for:

[0121] Input the input encoding of the non-fitting data and the first first-layer target encoding of the fitting data into the large language model to determine the first output encoding of the large language model;

[0122] Based on all the previously input encodings and the first second-layer target encoding of the input fitting data and the second first-layer target encoding, determine the second output encoding of the large language model, where the first second-layer target encoding and the second first-layer target encoding of the fitting data are input in a superposition summation or splicing manner; until the (M + N - 1)-th output encoding of the large language model is determined based on all the previously input encodings of the large language model and the M-th N-layer target encoding of the input fitting data;

[0123] Combine the M + N - 1 output encodings of the large language model into M N-layer final output encodings;

[0124] Iteratively optimize the large language model based on at least M N-layer target encodings of the fitting data and the M N-layer final output encodings to generate a target model.

[0125] The large language model encoding training device provided by the embodiments of the present invention can execute the large language model encoding training method provided by any embodiment of the present invention, and has the corresponding functional modules and beneficial effects for executing the method.

[0126] Embodiment 4

[0127] Figure 4 FIG. shows a schematic structural diagram of an electronic device 10 that can be used to implement the embodiments of the present invention. The electronic device is intended to represent various forms of digital computers, such as, laptop computers, desktop computers, workstations, personal digital assistants, servers, blade servers, mainframe computers, and other suitable computers. The electronic device can also represent various forms of mobile devices, such as, personal digital processors, cellular phones, smart phones, wearable devices (such as helmets, glasses, watches, etc.) and other similar computing devices. The components shown herein, their connections and relationships, and their functions are only examples and are not intended to limit the implementation of the present invention described and / or claimed herein.

[0128] AsFigure 4 As shown, the electronic device 10 includes at least one processor 11 and a memory communicatively connected to the at least one processor 11, such as a read-only memory (ROM) 12, a random access memory (RAM) 13, etc. Among them, the memory stores a computer program executable by the at least one processor. The processor 11 can execute various appropriate actions and processes according to the computer program stored in the read-only memory (ROM) 12 or the computer program loaded from the storage unit 18 into the random access memory (RAM) 13. In the RAM 13, various programs and data required for the operation of the electronic device 10 can also be stored. The processor 11, the ROM 12, and the RAM 13 are connected to each other through a bus 14. The input / output (I / O) interface 15 is also connected to the bus 14.

[0129] Multiple components in the electronic device 10 are connected to the I / O interface 15, including: an input unit 16, such as a keyboard, a mouse, etc.; an output unit 17, such as various types of displays, speakers, etc.; a storage unit 18, such as a disk, an optical disc, etc.; and a communication unit 19, such as a network card, a modem, a wireless communication transceiver, etc. The communication unit 19 allows the electronic device 10 to exchange information / data with other devices through a computer network such as the Internet and / or various telecommunication networks.

[0130] The processor 11 can be various general-purpose and / or special-purpose processing components with processing and computing capabilities. Some examples of the processor 11 include but are not limited to a central processing unit (CPU), a graphics processing unit (GPU), various dedicated artificial intelligence (AI) computing chips, various processors running machine learning model algorithms, a digital signal processor (DSP), and any appropriate processor, controller, microcontroller, etc. The processor 11 executes the various methods and processes described above, such as the large language model encoding training method.

[0131] In some embodiments, the large language model encoding training method can be implemented as a computer program, which is tangibly contained in a computer-readable storage medium, such as the storage unit 18. In some embodiments, part or all of the computer program can be loaded and / or installed onto the electronic device 10 via the ROM 12 and / or the communication unit 19. When the computer program is loaded into the RAM 13 and executed by the processor 11, one or more steps of the large language model encoding training method described above can be executed. Alternatively, in other embodiments, the processor 11 can be configured to execute the large language model encoding training method in any other appropriate way (e.g., by means of firmware).

[0132] The various embodiments of the systems and techniques described above in this specification can be implemented in digital electronic circuitry, integrated circuit systems, field programmable gate arrays (FPGAs), application specific integrated circuits (ASICs), application specific standard products (ASSPs), systems-on-chip (SOCs), complex programmable logic devices (CPLDs), computer hardware, firmware, software, and / or combinations thereof. These various embodiments can include: being implemented in one or more computer programs that are executable and / or interpretable on a programmable system including at least one programmable processor, which can be a special-purpose or general-purpose programmable processor that receives data and instructions from, and transmits data and instructions to, a storage system, at least one input device, and at least one output device.

[0133] The computer programs for implementing the methods of the present invention can be written in any combination of one or more programming languages. These computer programs can be provided to a processor of a general purpose computer, special purpose computer, or other programmable data processing apparatus, such that the computer programs, when executed by the processor, cause the functions / operations specified in the flowchart and / or block diagram to be implemented. The computer programs can be executed entirely on the machine, partly on the machine, as a stand-alone software package partly on the machine and partly on a remote machine or entirely on the remote machine or server.

[0134] In the context of the present invention, a computer-readable storage medium can be a tangible medium that can contain or store a computer program for use by or in connection with an instruction execution system, apparatus, or device. The computer-readable storage medium can include, but is not limited to, electronic, magnetic, optical, electromagnetic, infrared, or semiconductor systems, apparatus, or devices, or any suitable combination of the foregoing. Alternatively, the computer-readable storage medium can be a machine-readable signal medium. More specific examples of the machine-readable storage medium would include an electrical connection based on one or more wires, a portable computer diskette, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or Flash memory), an optical fiber, a portable compact disc read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the foregoing.

[0135] To provide interaction with a user, the systems and techniques described herein can be implemented on an electronic device having: a display device for displaying information to the user (e.g., a CRT (cathode ray tube) or LCD (liquid crystal display) monitor); and a keyboard and a pointing device (e.g., a mouse or a trackball) through which the user can provide input to the electronic device. Other kinds of devices can also be used to provide interaction with the user; for example, the feedback provided to the user can be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and input from the user can be received in any form (including acoustic input, voice input, or tactile input).

[0136] The systems and techniques described herein can be implemented in a computing system including backend components (e.g., as a data server), or a computing system including middleware components (e.g., an application server), or a computing system including frontend components (e.g., a user computer having a graphical user interface or a web browser through which the user can interact with an implementation of the systems and techniques described herein), or a computing system including any combination of such backend components, middleware components, or frontend components. The components of the system can be interconnected to each other by digital data communication in any form or medium (e.g., a communication network). Examples of communication networks include: local area network (LAN), wide area network (WAN), blockchain network, and the Internet.

[0137] The computing system can include a client and a server. The client and the server are generally far from each other and usually interact through a communication network. The client-server relationship is created by computer programs running on the respective computers and having a client-server relationship with each other. The server can be a cloud server, also known as a cloud computing server or a cloud host, which is a host product in the cloud computing service system and solves the defects of difficult management and weak business scalability existing in traditional physical hosts and VPS services.

[0138] It should be understood that the various forms of processes shown above can be used, with steps reordered, added, or deleted. For example, the steps recited in the present invention can be executed in parallel, sequentially, or in a different order, as long as the desired results of the technical solution of the present invention can be achieved, and no limitation is imposed herein.

[0139] The above specific embodiments do not constitute a limitation on the protection scope of the present invention. Those skilled in the art should understand that various modifications, combinations, sub-combinations, and substitutions can be made according to design requirements and other factors. Any modifications, equivalent substitutions, and improvements made within the spirit and principles of the present invention shall be included within the protection scope of the present invention.

Claims

1. A method for generating a code, characterized in that Including: Determine the initial encoding of the first layer of training data; Compare the initial encoding of the first layer with each first encoding in a preset encoding codebook or the first standard encoding corresponding to the first encoding, and determine the first target encoding corresponding to the first encoding with the highest similarity to the initial encoding of the first layer in the encoding codebook; Determine the target encoding of the first layer based on the first target encoding; Use the target encoding of the first layer as the current layer target encoding, calculate the residual between the current layer initial encoding and the current layer target encoding, and use the residual as the current layer encoding deviation of the current layer encoding; Compare the current layer encoding deviation with each second encoding in a preset next layer deviation codebook or the second standard encoding corresponding to the second encoding, determine the second target encoding corresponding to the second encoding with the highest similarity to the current layer encoding deviation in the next layer deviation codebook, and determine the next layer target encoding based on the second target encoding; Determine the target encoding of the Nth layer; where N is an integer greater than or equal to 2.

2. The method according to claim 1, wherein Before determining the initial encoding of the first layer of training data, it further includes: Obtain a training data set; where the training data set contains at least two training sample data groups, and each training sample data group consists of at least two corresponding types of training data; Determining the initial encoding of the first layer of training data includes: For at least one type of training data in each training sample data group, determine the initial encoding of the first layer of the training data; Determining the target encoding of the Nth layer includes: Update the next layer target encoding to the current layer target encoding, and return to execute to determine the current layer encoding deviation of the current layer target encoding, and then determine the target encoding of the Nth layer; After determining the target encoding of the Nth layer, it further includes: Train a large language model based on the target encodings of N layers to generate a target model.

3. The method according to claim 2, wherein Before determining the initial encoding of the first layer of training data for at least one type of training data in each training sample data group, it further includes: Extract partial training data of at least one dimension for the at least one type of training data; Encode the partial training data of at least one dimension respectively to generate at least one layer of enhanced target encoding; Training a large language model based on the target encodings of N layers to generate a target model includes: Superimpose at least one layer of the enhanced target encoding and the target encodings of N layers, and train the large language model to generate a target model.

4. The method according to claim 3, characterized in that The superimposing at least one layer of the enhanced target encoding and the target encodings of N layers and training the large language model includes: Place at least one layer of the enhanced target encoding before the target encodings of N layers, and at least part of the enhanced target encoding is input into the large language model prior to the target encodings of N layers for large language model training.

5. The method according to any one of claims 2 to 4, characterized in that, Before determining the initial encoding of the first layer of training data for at least one type of training data in each training sample data group, it further includes: Perform discretization processing on the at least one type of training data to generate multiple minimum discrete units; For at least one type of training data in each of the training sample data groups, determining a first-layer initial encoding of the training data, includes: Encoding each minimum discrete unit of at least one type of training data in each of the training sample data groups respectively, to determine the first-layer initial encoding of the training data.

6. The method according to claim 1, characterized in that Determining a corresponding target encoding based on the encoding with the highest similarity in the codebook, includes: Directly determining the encoding with the highest similarity in the codebook as the target encoding; or, Encoding the encoding with the highest similarity in the codebook based on a preset encoding strategy to generate the target encoding.

7. The method according to claim 2, wherein Each of the training sample data includes input data and fitting data, or each of the training sample data groups includes reference data, input data and fitting data. Training a large language model based on N layers of the target encodings to generate a target model, includes: Inputting the input encoding of the non-fitting data and the first-layer target encoding of the fitting data into the large language model to determine a first output encoding of the large language model; wherein, the non-fitting data includes the input data, or the non-fitting data includes the reference data and the input data; Determining a second output encoding of the large language model based on all the previously input encodings and the second-layer target encoding of the input fitting data; until determining an Nth output encoding of the large language model based on all the previously input encodings of the large language model and the Nth-layer target encoding of the input fitting data; Iteratively optimizing the large language model based on at least one Nth-layer target encoding of the fitting data and the corresponding number of output encodings to generate a target model.

8. The method according to claim 7, wherein The fitting data corresponds to M Nth-layer target encodings, where M is an integer greater than 1. The method further includes: Inputting the input encoding of the non-fitting data and the first first-layer target encoding of the fitting data into the large language model to determine a first output encoding of the large language model; Determining a second output encoding of the large language model based on all the previously input encodings and the first second-layer target encoding and the second first-layer target encoding of the input fitting data, where the first second-layer target encoding and the second first-layer target encoding of the fitting data are input in a manner of superposition summation or splicing; until determining an (M + N - 1)th output encoding of the large language model based on all the previously input encodings of the large language model and the Mth Nth-layer target encoding of the input fitting data; Combining the M + N - 1 output encodings of the large language model into M Nth-layer final output encodings; Iteratively optimizing the large language model based on at least M Nth-layer target encodings of the fitting data and the M Nth-layer final output encodings to generate a target model.

9. An encoding generation device, characterized in that, Includes: A first-layer initial encoding determination module, configured to determine the first-layer initial encoding of the training data; The first target encoding determination module is configured to compare the initial encoding of the first layer with each first encoding in a preset encoding codebook or a first standard encoding corresponding to the first encoding, and determine a first target encoding corresponding to the first encoding in the encoding codebook that has the highest similarity to the initial encoding of the first layer; The first-layer target encoding determination module is configured to determine a first-layer target encoding based on the first target encoding; The current-layer encoding deviation determination module is configured to use the first-layer target encoding as the current-layer target encoding, calculate the residual between the current-layer initial encoding and the current-layer target encoding, and use the residual as the current-layer encoding deviation of the current-layer encoding; The next-layer target encoding determination module is configured to compare the current-layer encoding deviation with each second encoding in a preset next-layer deviation codebook or a second standard encoding corresponding to the second encoding, determine a second target encoding corresponding to the second encoding in the next-layer deviation codebook that has the highest similarity to the current-layer encoding deviation, and determine a next-layer target encoding based on the second target encoding; Determine the target encoding of the Nth layer; where N is an integer greater than or equal to 2.

Citation Information

Patent Citations

  • Machine translation model training method and device, electronic equipment and storage medium

    CN115600613A

  • Point cloud attribute coding method, point cloud attribute decoding method and terminal

    CN116233386A

  • Determination method of image encoder and related device

    CN117011650A

  • Compressing audio waveforms using neural networks and vector quantizers

    US20230019128A1

  • Natural intelligence for natural language processing

    WO2024077002A2