A large language model encoding training method and device

By performing multi-layer object encoding processing on the training data and training the large language model based on these encodings, the problem of information loss when encoding information rich data in the prior art is solved, and the quality of the model output data is significantly improved.

CN119204133BActive Publication Date: 2025-05-06SHANGHAI XIYU JIZHI TECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202411470413.6
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-10-21
Publication Date
2025-05-06
Estimated Expiration
2044-10-21

AI Technical Summary

Technical Problem

When training, when existing large language models encode information-rich data (such as multi-instrument and/or vocal music), a large amount of useful information will be lost, resulting in rough output data and unable to meet user needs.

Method used

By encoding at least one training data in each training sample data set in the training data set into a multi-layer target encoding, the specific steps include determining the first layer target encoding, recursively determining the next layer target encoding based on the encoding bias, until determining the N-th layer target encoding, and training the large language model based on these target encodings.

Benefits of technology

It improves the degree of refinement of data discrete and fit, significantly improves the quality of data output by the trained large language model, and can better process and generate information-rich data.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119204133B_ABST
    Figure CN119204133B_ABST
Patent Text Reader

Abstract

The present invention discloses a large language model encoding training method and device. The method includes: obtaining a training data set; wherein the training data set includes at least two training sample data groups; for at least one training data in each training sample data group, determining the first layer target encoding of the training data; using the first layer target encoding as the current layer target encoding, determining the current layer encoding deviation of the current layer target encoding, and determining the next layer target encoding based on the current layer encoding deviation; updating the next layer target encoding to the current layer target encoding, and returning to execute the current layer encoding deviation of the current layer target encoding to determine the Nth layer target encoding; training the large language model based on the N layers of target encoding to generate a target model. This solution can provide richer data information and improve the quality of the output data of the trained large language model by encoding at least one training data in the training sample data group into multiple layers of target encoding.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of machine learning, and in particular to a large language model encoding training method and device. Background Art

[0002] When training an existing large model, it is generally necessary to convert the training data into a code and input it into the large model. For example, each minimum word segmentation unit in a paragraph of text (which can be each character, each word, or each radical) is converted into a code, and then the paragraph of text is converted into a group of at least one code and input into the large model. The existing coding representation method is sufficient to meet the use requirements of the large model after training when facing data with relatively simple information content, but when facing data with relatively rich information content, such as each frame (each minimum word segmentation unit) of music with many instruments and / or human voices, each frame of music has a large amount of key information such as tone, timbre, pitch, and range. If it is fitted through a code, a large amount of useful information will be lost, which will cause the output data based on the trained large model to become very rough and far from meeting user needs. Summary of the invention

[0003] The present invention provides a large language model encoding training method and device, which can provide richer data information, improve the degree of data discretization and fitting refinement, and thus improve the quality of the trained large language model output data.

[0004] According to one aspect of the present invention, a large language model encoding training method is provided, comprising:

[0005] Acquire a training data set; wherein the training data set includes at least two training sample data groups, and each of the training sample data groups is composed of corresponding at least two types of training data;

[0006] For at least one training data in each of the training sample data groups, determining a first-layer target encoding of the training data;

[0007] Taking the first layer target code as the current layer target code, determining a current layer code deviation of the current layer target code, and determining a next layer target code based on the current layer code deviation;

[0008] The next layer target code is updated to the current layer target code, and the current layer code deviation for determining the current layer target code is returned to determine the Nth layer target code; wherein N is an integer greater than or equal to 2;

[0009] The large language model is trained based on the N layers of target encoding to generate a target model.

[0010] According to another aspect of the present invention, a large language model encoding training device is provided, comprising:

[0011] A training data set acquisition module is used to acquire a training data set; wherein the training data set includes at least two training sample data groups, and each training sample data group is composed of corresponding at least two types of training data;

[0012] A first-layer target coding determination module, used for determining the first-layer target coding of the training data for at least one training data in each of the training sample data groups;

[0013] a next layer target coding determination module, configured to use the first layer target coding as the current layer target coding, determine the current layer coding deviation of the current layer target coding, and determine the next layer target coding based on the current layer coding deviation;

[0014] A coding deviation loop determination module, used for updating the next layer target coding to the current layer target coding, and returning to execute the current layer coding deviation for determining the current layer target coding, and determining the Nth layer target coding; wherein N is an integer greater than or equal to 2;

[0015] The target model generation module is used to train the large language model based on the N layers of target encoding to generate a target model.

[0016] According to another aspect of the present invention, an electronic device is provided, the electronic device comprising:

[0017] at least one processor; and

[0018] a memory communicatively connected to the at least one processor; wherein,

[0019] The memory stores a computer program that can be executed by the at least one processor, and the computer program is executed by the at least one processor so that the at least one processor can execute the large language model encoding training method described in any embodiment of the present invention.

[0020] According to another aspect of the present invention, a computer-readable storage medium is provided, wherein the computer-readable storage medium stores computer instructions, and the computer instructions are used to enable a processor to implement the large language model encoding training method described in any embodiment of the present invention when executed.

[0021] The large language model encoding training scheme of the embodiment of the present invention obtains a training data set; wherein the training data set contains at least two training sample data groups, and each of the training sample data groups is composed of at least two corresponding training data; for at least one training data in each of the training sample data groups, the first layer target encoding of the training data is determined; the first layer target encoding is used as the current layer target encoding, the current layer encoding deviation of the current layer target encoding is determined, and the next layer target encoding is determined based on the current layer encoding deviation; the next layer target encoding is updated to the current layer target encoding, and the current layer encoding deviation of the current layer target encoding is determined to determine the Nth layer target encoding; wherein N is an integer greater than or equal to 2; the large language model is trained based on the N layers of the target encoding to generate a target model. The technical solution provided by the embodiment of the present invention can provide richer data information by encoding at least one training data in the training sample data group into multiple layers of target encoding, improve the degree of refinement of data discretization and fitting, and thus improve the quality of the output data of the trained large language model.

[0022] It should be understood that the contents described in this section are not intended to identify the key or important features of the embodiments of the present invention, nor are they intended to limit the scope of the present invention. Other features of the present invention will become easily understood through the following description. BRIEF DESCRIPTION OF THE DRAWINGS

[0023] In order to more clearly illustrate the technical solutions in the embodiments of the present invention, the accompanying drawings required for use in the description of the embodiments will be briefly introduced below. Obviously, the accompanying drawings in the following description are only some embodiments of the present invention. For ordinary technicians in this field, other accompanying drawings can be obtained based on these accompanying drawings without paying creative work.

[0024] Figure 1 A flowchart of a large language model encoding training method provided in Embodiment 1 of the present invention;

[0025] Figure 2 A flowchart of a large language model encoding training method provided in Embodiment 2 of the present invention;

[0026] Figure 3 A schematic diagram of the structure of a large language model encoding training device provided in Embodiment 3 of the present invention;

[0027] Figure 4 A schematic diagram of the structure of an electronic device for implementing the large language model encoding training method according to an embodiment of the present invention. DETAILED DESCRIPTION

[0028] In order to enable those skilled in the art to better understand the scheme of the present invention, the technical scheme in the embodiments of the present invention will be clearly and completely described below in conjunction with the drawings in the embodiments of the present invention. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without creative work should fall within the scope of protection of the present invention.

[0029] It should be noted that the terms "first", "second", etc. in the specification and claims of the present invention and the above-mentioned drawings are used to distinguish similar objects, and are not necessarily used to describe a specific order or sequence. It should be understood that the data used in this way can be interchanged where appropriate, so that the embodiments of the present invention described herein can be implemented in an order other than those illustrated or described herein. In addition, the terms "including" and "having" and any variations thereof are intended to cover non-exclusive inclusions, for example, a process, method, system, product or device that includes a series of steps or units is not necessarily limited to those steps or units that are clearly listed, but may include other steps or units that are not clearly listed or inherent to these processes, methods, products or devices.

[0030] Embodiment 1

[0031] Figure 1 A flowchart of a large language model encoding training method is provided for Embodiment 1 of the present invention. This embodiment is applicable to the case of encoding training of a large language model. The method can be performed by a large language model encoding training device. The large language model encoding training device can be implemented in the form of hardware and / or software. The large language model encoding training device can be configured in an electronic device.

[0032] like Figure 1 As shown, the method includes:

[0033] S110, obtaining a training data set; wherein the training data set includes at least two training sample data groups, and each of the training sample data groups is composed of at least two corresponding training data.

[0034] In an embodiment of the present invention, a training data set is obtained, wherein the training data set includes a plurality of training sample data groups, and each training sample data group is composed of at least two corresponding data. For example, the training sample data group may include input data and fitting data; for another example, the training sample data group may include reference data, input data, and fitting data. Among them, the input data is data that needs to output corresponding generated content according to the input content when the large language model is subsequently trained, such as questions in question-and-answer data, and lyrics in music generation data; the reference data is data that needs to refer to part of the information in the subsequent large language model training to output corresponding generated content, such as reference pictures in literary image data, reference songs in music generation data, or reference accompaniment in reference songs, reference singing style and other reference data, which are used to enable the large language model to generate output data with the same style; the fitting data is standard answer data, which is used to compare one by one with the output results of the large language model in the subsequent large language model training, calculate the loss value, and thus correct the probability of the next output result of the large model, which is the ideal data that the large language model needs to fit. The training data set is a data set that matches the application scenario of the large language model used for subsequent training. For example, the large language model trained subsequently is a music generation model. Each training sample data group in the training data set may include fitted music and target lyrics, or may include reference music, fitted music, and target lyrics.

[0035] It should be noted that the embodiments of the present invention do not limit the number of input data and reference data contained in the training sample data set. For example, when training a large language model for music generation, the input data in the training sample data set can be a paragraph of lyrics or multiple paragraphs of lyrics, and the reference data in the training sample data set can be a reference song or a reference accompaniment and reference singing style in the reference song.

[0036] S120. For at least one training data in each of the training sample data groups, determine a first-layer target encoding of the training data.

[0037] In an embodiment of the present invention, for each training sample data group in a training data set, a first-level target coding of at least one training data in the training sample data group is determined. It is understandable that the first-level target coding of any one or more or all training data in the training sample data group can be determined, wherein the types of training data of the first-level target coding determined in each training sample data group in the training data set can be the same or different. Exemplarily, an encoder can be used to encode at least one training data in the training sample data group to determine the first-level target coding, wherein the encoder can include a MERT encoder and a Mel encoder.

[0038] Optionally, before determining the first-level target coding of at least one training data in each of the training sample data groups, the method further includes: discretizing the at least one training data to generate a plurality of minimum discrete units; determining the first-level target coding of the training data for at least one training data in each of the training sample data groups includes: encoding each minimum discrete unit of at least one training data in each of the training sample data groups to determine the first-level target coding of the training data. Exemplarily, discretizing the at least one training data in the training sample data groups to generate at least two minimum discrete units, and then encoding each of the at least two minimum discrete units by an encoder to determine the corresponding first target coding. It is understandable that the number of the first target coding is the same as the number of the minimum discrete units.

[0039] Taking the large language model for subsequent training as the music generation model, the training sample data group is composed of reference music, fitting music and target lyrics as an example for exemplary description. The fitting music is discretized to generate at least two frames of fitting audio data. Exemplarily, the fitting music can be discretized into 10 frames of fitting audio data per second. For example, if the fitting music is a 10-second music, the fitting music can be discretized into 100 frames of fitting audio data. Another exemplary method is to discretize the fitting music as a whole, for example, the fitting music is discretized into 500 frames of fitting audio data, wherein the number of frames of discrete fitting audio data can be determined according to the length of the reference music, and the larger the length of the fitting music, the more frames of discrete fitting audio data. Each frame of fitting audio data is encoded by an encoder to generate a corresponding first-layer target encoding, wherein the number of first-layer target encodings is the same as the number of frames of fitting audio data. Optionally, the same encoder can be used to encode each frame of fitting audio data, or different encoders can be used to encode each frame of fitting audio data. Among them, the encoder can be an open source encoder, such as a MERT encoder or a Mel encoder. In an embodiment of the present invention, the discretization processing of the target lyrics may include: performing word segmentation processing on the target lyrics based on a preset word segmentation algorithm, dividing the target lyrics into multiple word segmentation units, wherein the word segmentation unit may be each character, each word, or each sentence. Encoding each word segmentation unit in the target lyrics based on an encoder, generating a first-layer target encoding corresponding to each word segmentation unit. The method for determining the first-layer target encoding of the reference music may be similar to the method for determining the first-layer target encoding of the fitted music, which will not be repeated here.

[0040] Optionally, determining the first layer target coding of the training data includes: determining the first layer initial coding of the training data; comparing the first layer initial coding with each first coding or the first standard coding corresponding to the first coding in the pre-set coding code book, determining the first target coding corresponding to the first coding in the coding code book with the highest similarity to the first layer initial coding; and determining the first layer target coding based on the first target coding. Exemplarily, encoding at least one training data in the training sample data group by an encoder, and using the obtained coding information as the first layer initial coding of the training data. Obtaining a pre-set coding code book, wherein the coding code book consists of multiple first codings. Optionally, the first coding in the coding code book can be a standard coding (also called a special coding), or the standard coding can be replaced by a special character, such as using the first standard coding as coding 0, the second standard coding as coding 1, and so on, to generate a coding code book such as [0,1,2,...,1023]. This can greatly simplify the complexity of the codebook, and at the same time can isolate the codes (that is, special characters such as 0, 1, 2, etc.) and the standard codes, so as to facilitate the future adjustment of the format and / or content of the standard codes. For example, the first standard code can be a multidimensional vector, or any one or more of audio codes, graphic codes, character codes, etc.; or the first standard code is originally a 256-dimensional code vector, and if it is subsequently adjusted and converted into a 1024-dimensional code vector, the codebook is still [0,1,2,...,1023].

[0041] The first-layer initial code is sequentially compared with each first code in the coding code book or the first standard code corresponding to the first code, and the similarity between the two is calculated. The first code in the coding code book with the highest similarity to the first-layer initial code is used as the first target code, and the first-layer target code of the training data is determined based on the first target code. The first target code in the coding code book can be directly used as the first-layer target code, or the first standard code corresponding to the first target code in the coding code book can be used as the first-layer target code, and the first target code in the coding code book can also be encoded based on a preset coding strategy to generate the first-layer target code corresponding to the standard data.

[0042] S130. Use the first-layer target coding as the current-layer target coding, determine the current-layer coding deviation of the current-layer target coding, and determine the next-layer target coding based on the current-layer coding deviation.

[0043] In an embodiment of the present invention, the first layer target coding of the training data is updated to the current layer target coding, and the current layer coding deviation of the current layer target coding is determined, wherein the current layer coding deviation can be the difference between the current layer initial coding and the current layer target coding. The next layer target coding of the training data is determined based on the current layer coding deviation. For example, the current layer coding deviation can be encoded based on a preset coding algorithm to generate the next layer target coding of the training data. It can be understood that the first layer coding deviation of the first layer target coding of the training data is determined, wherein the residual between the first layer initial coding of the training data and the first layer target coding can be used as the first layer coding deviation. The second layer target coding of the training data is determined based on the first layer coding deviation.

[0044] Optionally, the current layer coding deviation of the current layer target coding is determined, and the next layer target coding is determined based on the current layer coding deviation, including: calculating the residual between the current layer initial coding and the current layer target coding, and using the residual as the current layer coding deviation of the current layer coding; comparing the current layer coding deviation with each second coding in a preset next layer deviation code book or a second standard coding corresponding to the second coding, determining the second target coding corresponding to the second coding in the next layer deviation code book with the highest similarity to the current layer coding deviation, and determining the next layer target coding based on the second target coding.

[0045] Exemplarily, the residual between the current layer initial code and the current layer target code is used as the current layer code deviation of the current layer code. Obtain the next layer deviation codebook, wherein the next layer deviation codebook is composed of multiple second codes, and the second code can be understood as a residual code. Optionally, the second code in the deviation codebook can be a standard residual code (also called a special residual code), or the standard residual code can be replaced by a special character, such as using the first standard residual code as code 0, the second standard residual code as code 1, and so on, to generate a deviation codebook such as [0,1,2,...,1023]. This can greatly simplify the complexity of the deviation codebook, and at the same time can isolate the codes (that is, special characters such as 0, 1, 2) and the standard residual codes, so as to facilitate the future adjustment of the format and / or content of the standard residual codes. For example, the first standard residual code can be a multi-dimensional vector, or any one or more of audio codes, graphic codes, character codes, etc.; or the first standard residual code is originally a 256-dimensional code vector, which is subsequently adjusted to be converted into a 1024-dimensional code vector. At this time, the deviation codebook is still [0,1,2,...,1023].

[0046] The current layer coding deviation is sequentially compared with each second coding in the next layer deviation code book or the second standard coding corresponding to the second coding, and the similarity between the two is calculated. The second coding in the next layer deviation code book with the highest similarity to the current layer coding deviation is used as the second target coding, and the next layer target coding of the training data is determined based on the second target coding. Among them, the second target coding in the next layer deviation code book can be directly used as the next layer target coding, or the second standard coding corresponding to the second target coding in the next layer deviation code book can be used as the next layer target coding, and the second target coding in the next layer deviation code book can also be encoded based on a preset coding strategy to generate the next layer target coding corresponding to the standard data.

[0047] S140, updating the next layer target coding to the current layer target coding, and returning to execute the current layer coding deviation for determining the current layer target coding, to determine the Nth layer target coding; wherein N is an integer greater than or equal to 2.

[0048] In an embodiment of the present invention, the first layer coding deviation of the first layer target coding of the training data is compared with the second standard coding corresponding to each second coding in the second layer deviation codebook, the second target coding corresponding to the second coding with the highest similarity to the first layer coding deviation in the second layer deviation codebook is determined, and the second layer target coding is determined based on the second target coding. The first layer coding deviation is used as the second layer initial coding, the residual between the second initial coding and the second layer target coding is calculated, and the residual is used as the second layer coding deviation, that is, the residual between the first layer coding deviation and the second layer target coding is used as the second layer coding deviation. For the convenience of description, the second coding in the i-th layer deviation codebook can also be called the i-th coding, and the standard coding corresponding to the i-th coding can be called the i-th standard coding. Therefore, the second layer coding deviation is compared with the third standard coding corresponding to each third coding in the third layer deviation codebook, the third target coding corresponding to the third coding with the highest similarity to the second layer coding deviation in the third layer deviation codebook is determined, and the third layer target coding is determined based on the third target coding. The second layer coding deviation is used as the third layer initial coding, the residual between the third initial coding and the third layer target coding is calculated, and the residual is used as the third layer coding deviation, that is, the residual between the second layer coding deviation and the third layer target coding is used as the third layer coding deviation. The above method is repeated until the Nth layer target coding of the training data is determined, where N is an integer greater than or equal to 2.

[0049] It is understandable that according to S120-S130, N layers of target coding corresponding to at least one training data in the training sample data set can be determined. It should be noted that the number of layers of target coding of at least one training data in each training sample data set in the training data set can be the same or different.

[0050] S150, training the large language model based on the N layers of target encoding to generate a target model.

[0051] In an embodiment of the present invention, the N-layer target encoding corresponding to at least one training data in each training sample data group in the training data set is input into the large language model to train the large language model and generate a target model. The large language model can be any open source large model, such as GPT. It should be noted that the embodiment of the present invention does not limit the application scenario of the target model. Exemplarily, the target model can be a music generation model. In this case, each training sample data group in the training data set used to train the target model can include reference music, sample lyrics and fitting music; for example, the target model can be a text-to-image model, that is, a model that generates pictures based on text. In this case, each training sample data group in the training data set used to train the target model can include sample text and fitting pictures.

[0052] The large language model encoding training method of the embodiment of the present invention obtains a training data set; wherein the training data set contains at least two training sample data groups, and each of the training sample data groups is composed of at least two corresponding training data; for at least one training data in each of the training sample data groups, the first layer target encoding of the training data is determined; the first layer target encoding is used as the current layer target encoding, the current layer encoding deviation of the current layer target encoding is determined, and the next layer target encoding is determined based on the current layer encoding deviation; the next layer target encoding is updated to the current layer target encoding, and the current layer encoding deviation of the current layer target encoding is determined to determine the Nth layer target encoding; wherein N is an integer greater than or equal to 2; the large language model is trained based on the N layers of the target encoding to generate a target model. The technical solution provided by the embodiment of the present invention can provide richer data information by encoding at least one training data in the training sample data group into multi-layer target encoding, improve the degree of refinement of data discretization and fitting, and thus improve the quality of the output data of the trained large language model.

[0053] In some embodiments, for at least one training data in each of the training sample data groups, before determining the first layer target coding of the training data, it also includes: extracting partial training data of at least one dimension for the at least one training data; encoding the partial training data of at least one dimension respectively to generate at least one layer of enhanced target coding; training the large language model based on the N layers of the target coding to generate the target model, including: superimposing at least one layer of the enhanced target coding with the N layers of the target coding, training the large language model to generate the target model. The advantage of such a setting is that the training effect of the large language model in certain dimensions can be enhanced to obtain a better training effect of the large language model.

[0054] In an embodiment of the present invention, for at least one type of training data in each training sample data group, partial training data of at least one dimension is extracted from the training data, and the partial training data of each dimension is encoded respectively to generate at least one layer of enhanced target code. Among them, the partial training data of each dimension can be encoded as a whole to generate at least one layer of enhanced target code, or the partial training data of each dimension can be discretized, and the generated minimum discrete units are encoded to generate at least one layer of enhanced target code. It should be noted that when the enhanced target code is a multi-layer code, the method for determining the multi-layer enhanced target code is the same as the method for determining the N-layer target code of the training data in the above embodiment, which will not be repeated here.

[0055] Exemplarily, the training sample data set includes reference music, fitting music and target lyrics, and the fitting music can be split into a preset number of dimensions, such as splitting at least one dimension of fitting music parts (such as fitting music parts of two dimensions of accompaniment part and vocal part) from the fitting music, discretizing the fitting music parts of each dimension, generating at least two frames of fitting audio data, encoding each frame of fitting audio data through an encoder, and generating at least one corresponding layer of enhanced target code. Similarly, the reference sample music is split into a preset number of dimensions, such as splitting at least one dimension of reference music parts (such as reference music parts of two dimensions of accompaniment part and vocal part) from the reference music, uniformly encoding the reference music parts of at least one dimension as a whole, and generating at least one corresponding layer of enhanced target code. Or the reference music parts of at least one dimension that are split are discretized to generate at least two frames of reference audio data, and each frame of reference audio data is encoded through an encoder to generate at least one corresponding layer of enhanced target code. Encode the target lyrics to generate at least one corresponding layer of enhanced target code. As another example, the bass, middle and treble in a frame of audio data are encoded at different layers to generate corresponding enhanced target codes, or different instruments and different voices in a frame of audio data are separated and encoded at different layers to generate corresponding enhanced target codes.

[0056] In an embodiment of the present invention, at least one layer of enhanced target coding is superimposed on N layers of target coding, and the superimposed coding is input into a large language model to train the large language model and generate a target model. Optionally, superimposing at least one layer of the enhanced target coding on N layers of the target coding to train the large language model includes: placing at least one layer of the enhanced target coding before the N layers of the target coding, and at least part of the enhanced target coding is input into the large language model in priority to the N layers of the target coding to train the large language model.

[0057] Exemplarily, taking the fitting music in the training sample data group as an example, the enhanced target code of the vocal part extracted from the fitting music is used as the first layer, and the 4 layers of target codes corresponding to the fitting music are used as the 2nd, 3rd, 4th, and 5th layers respectively. The enhanced target code of the vocal part of the fitting music is preferentially input as the first layer code to the large language model for prediction output, and then the 4 layers of target codes corresponding to the fitting music are sequentially input into the large language model, so as to continue to train the large language model based on the 4 layers of target codes to generate a target model. The advantage of such a setting is that the large language model can firstly fit and output the vocal part corresponding to the simple fitting music, and then perform fitting training for complex songs, so as to obtain a clearer fitting training of the vocal part than simply fitting the song, thereby achieving a better training effect of the large language model.

[0058] Embodiment 2

[0059] Figure 2 A flowchart of a large language model encoding training method provided in Embodiment 2 of the present invention is shown in FIG. Figure 2 As shown, the method includes:

[0060] S210, obtaining a training data set; wherein the training data set includes at least two training sample data groups, each of the training sample data groups is composed of at least two corresponding training data, each of the training sample data includes input data and fitting data, or each of the training sample data groups includes reference data, input data and fitting data.

[0061] S220. For at least one training data in each of the training sample data groups, determine a first-layer target encoding of the training data.

[0062] S230: Use the first-layer target coding as the current-layer target coding, determine the current-layer coding deviation of the current-layer target coding, and determine the next-layer target coding based on the current-layer coding deviation.

[0063] S240, updating the next layer target coding to the current layer target coding, and returning to execute the current layer coding deviation for determining the current layer target coding, to determine the Nth layer target coding; wherein N is an integer greater than or equal to 2.

[0064] S250, inputting the input code of the non-fitting data and the first layer target code of the fitting data into the large language model, and determining the first output code of the large language model; wherein the non-fitting data includes the input data, or the non-fitting data includes the reference data and the input data.

[0065] Each training sample data may include input data and fitting data, or may include reference data, input data, and fitting data. When the training sample data includes input data and fitting data, the input data is used as non-fitting data; when the training sample data includes reference data, input data, and fitting data, the reference data and input data are used as non-fitting data.

[0066] In the embodiment of the present invention, the N-layer target code of the non-fitting data is used as the input code of the non-fitting data, and the input code of the non-fitting data and the first-layer target code of the fitting data are input into the large language model to train the large language model and obtain the first output code of the large language model. It can be understood that the first output code is the code output by the large language model after learning the input code of the non-fitting data and the first-layer target code of the fitting data.

[0067] S260. Determine the second output code of the large language model based on all previously input codes and the second-level target code of the input fitting data; until the Nth output code of the large language model is determined based on the Nth-level target code of all previously input codes of the large language model and the input fitting data.

[0068] In the embodiment of the present invention, the second output code of the large language model is determined based on all the codes previously input to the large language model and the second layer target code of the fitting data input to the large language model. Similarly, the third output code of the large language model is determined based on all the codes previously input to the large language model and the third layer target code of the fitting data input to the large language model; the third output code of the large language model is determined based on all the codes previously input to the large language model and the fourth layer target code of the fitting data input to the large language model. And so on, until the Nth output code of the large language model is determined based on all the codes previously input to the large language model and the Nth layer target code of the fitting data input to the large language model.

[0069] S270. Based on at least one N-layer target encoding of the fitting data and a corresponding number of output encodings, iteratively optimize the large language model to generate a target model.

[0070] In the embodiment of the present invention, the N-layer target code of the fitting data may be one or more. When the fitting data corresponds to multiple N-layer target codes, the output code corresponding to each N-layer target code is obtained through S260-S270. Based on at least one N-layer target code of the fitting data and the corresponding number of output codes, the large language model is continuously iteratively optimized to generate a target model.

[0071] Optionally, the fitting data corresponds to M N-layer target codes, where M is an integer greater than 1, and the method further includes: inputting a large language model based on an input code of non-fitting data and a first first-layer target code of the fitting data to determine a first output code of the large language model; determining a second output code of the large language model based on all previously input codes and a first second-layer target code and a second first-layer target code of the input fitting data, wherein the first second-layer target code and the second first-layer target code of the fitting data are input by superposition, summation or concatenation; until the M+N-1th output code of the large language model is determined based on all previously input codes of the large language model and the Mth N-layer target code of the input fitting data; combining the M+N-1 output codes of the large language model into M N-layer final output codes; iteratively optimizing the large language model based on at least M N-layer target codes and M N-layer final output codes of the fitting data to generate a target model.

[0072] In the embodiment of the present invention, since the output data of the large language model depends on the input data of the large language model corresponding to the last fitting data, the M N-layer target codes corresponding to the fitting data are input to the large language model in a delayed manner. Specifically, the fitting data corresponds to M N-layer target codes, all N-layer target codes of the non-fitting data are used as the input codes of the non-fitting data, and the input codes of the non-fitting data and the first first-layer target codes of the fitting data are input to the large language model, so that the large language model performs prediction and obtains the first output code of the large language model. It can be understood that the first output code is the code output after the large language model learns the input code of the non-fitting data and the first first-layer target code of the fitting data. Based on all the codes previously input to the large language model, the first second-layer target code of the fitting data input to the large language model, and the second first-layer target code, the second output code of the large language model is determined. Specifically, after the first second-layer target code of the fitting data and the second first-layer target code are superimposed and summed or spliced, the second output code of the large language model is determined based on the superimposed, summed or spliced ​​code and all the codes previously input. And so on, until the M+N-1th output code of the large language model is determined based on all previously input codes of the large language model and the Mth-Nth layer target code of the input fitting data.

[0073] Exemplarily, M=4, N=4, Table 1 is a delayed superposition input table of four 4-layer target codes corresponding to fitting data provided in an embodiment of the present invention:

[0074] Table 1

[0075] t(14) t(24) t(34) t(44) t(13) t(23) t(33) t(43) t(12) t(22) t(32) t(42) t(11) t(21) t(31) t(41)

[0076] After the input code of the non-fitting data and the first first-layer target code t(11) of the fitting data are input into the large language model, the first output code of the large language model is determined. Then, the first second-layer target code t(12) and the second first-layer target code t(21) are superimposed and summed and input into the large language model. The large language model determines the second output code based on all previously input codes and the code after the superposition and summation of t(12) and t(21). The first third-layer target code t(13), the second second-layer target code t(22) and the third first-layer target code t(31) of the fitting data are superimposed and summed, and then input into the large language model. The large language model determines the third output code based on all the codes previously input and the codes after the superposition and summation of t(13), t(22) and t(31); the first fourth-layer target code t(14), the second third-layer target code t(23), the third second-layer target code t(32) and the fourth first-layer target code t(41) of the fitting data are superimposed and summed, and then input into the large language model. The large language model determines the fourth output code based on all the codes previously input and the codes after the superposition and summation of t(14), t(23), t(32) and t(41). And so on, according to the method described in Table 1, the seventh output code of the large language model is obtained. It can be understood that superimposing the M N-layer target codes of the fitting data into the large language model in a delayed manner can enable the large language model to first obtain the low-level target code of each unit and then obtain the high-level target code, which is beneficial to the training of the large language model.

[0077] Since the output data of the large language model depends on all the inputs of the non-fitting data (input data, or reference data and input data), the N-layer target codes corresponding to the non-fitting data can be input into the large language model in a conventional superposition or splicing manner, that is, all N-layer target codes of the non-fitting data are input into the large language model at one time, so that the large model can obtain all N-layer target codes of the non-fitting data at one time, thereby making the training of the large language model faster and more efficient. Optionally, the DELAY method can also be used to superimpose the N-layer target codes of the non-fitting data into the large language model to train the large language model. Exemplarily, Table 2 is a conventional superposition input table of 4 4-layer target codes corresponding to non-fitting data provided in an embodiment of the present invention:

[0078] Table 2

[0079] t(14) t(24) t(34) t(44) t(13) t(23) t(33) t(43) t(12) t(22) t(32) t(42) t(11) t(21) t(31) t(41)

[0080] Among them, when the N-layer target codes of the non-fitting data are superimposed and input into the large language model in a conventional manner, the multi-dimensional code vector obtained by adding or concatenating the N-layer target codes of each unit can be obtained. For example, the 256-dimensional t(11), t(12), t(13), and t(14) are added to obtain a 256-dimensional vector, or concatenated to obtain a 1024-dimensional vector. The 256-dimensional t(21), t(22), t(23), and t(24) are added to obtain a 256-dimensional vector, which is still a 256-dimensional vector, or concatenated to obtain a 1024-dimensional vector. The multi-layer code vector is input into the large language model, wherein the output of the large language model is also a multi-layer code vector, thereby improving the accuracy of the output of the large language model.

[0081] The M+N-1 output codes of the large language model are combined into M N-layer final output codes, and then the large language model is iteratively optimized based on at least M N-layer target codes and M N-layer final output codes of the fitting data to generate a target model.

[0082] The technical solution provided by the embodiment of the present invention can provide richer data information by encoding at least one training data in the training sample data group into a multi-layer target code, thereby improving the degree of refinement of data discretization and fitting, thereby improving the quality of the output data of the trained large language model.

[0083] Embodiment 3

[0084] Figure 3 This is a schematic diagram of the structure of a large language model encoding training device provided by Embodiment 3 of the present invention. Figure 3 As shown, the device comprises:

[0085] The training data set acquisition module 310 is used to acquire a training data set; wherein the training data set includes at least two training sample data groups, and each training sample data group is composed of corresponding at least two types of training data;

[0086] A first-layer target coding determination module 320, configured to determine a first-layer target coding of at least one training data in each of the training sample data groups;

[0087] A next layer target coding determination module 330, configured to use the first layer target coding as the current layer target coding, determine the current layer coding deviation of the current layer target coding, and determine the next layer target coding based on the current layer coding deviation;

[0088] The coding deviation loop determination module 340 is used to update the next layer target coding to the current layer target coding, and return to execute the current layer coding deviation for determining the current layer target coding, and determine the Nth layer target coding; wherein N is an integer greater than or equal to 2;

[0089] The target model generation module 350 is used to train the large language model based on the N layers of target encoding to generate a target model.

[0090] Optionally, the device further includes:

[0091] A partial training data extraction module, configured to extract partial training data of at least one dimension for at least one training data in each of the training sample data groups before determining the first layer target encoding of the training data;

[0092] An enhanced target code generation module, used to encode part of the training data of at least one dimension respectively to generate at least one layer of enhanced target code;

[0093] The target model generation module comprises:

[0094] The target model generating unit is used to superimpose at least one layer of the enhanced target encoding with N layers of the target encoding, train the large language model, and generate a target model.

[0095] Optionally, the target model generating unit is used to:

[0096] At least one layer of the enhanced target coding is placed before the N layers of the target coding, and at least part of the enhanced target coding is input into the large language model before the N layers of the target coding to perform the large language model training.

[0097] Optionally, the device further comprises:

[0098] A minimum discrete unit generation module is used to discretize the at least one training data in each of the training sample data groups before determining the first layer target encoding of the training data to generate a plurality of minimum discrete units;

[0099] The first-layer target coding determination module is used to:

[0100] Each minimum discrete unit of at least one training data in each of the training sample data groups is encoded respectively to determine the first layer target encoding of the training data.

[0101] Optionally, the first-layer target coding determination module is used to:

[0102] Determining a first layer initial encoding of the training data;

[0103] Compare the first layer initial code with each first code in a preset code book or a first standard code corresponding to the first code, and determine a first target code corresponding to the first code in the code book that has the highest similarity to the first layer initial code;

[0104] A first layer target coding is determined based on the first target coding.

[0105] Optionally, the next layer target coding determination module is used to:

[0106] Calculate the residual between the current layer initial code and the current layer target code, and use the residual as the current layer code deviation of the current layer code;

[0107] Compare the current layer coding deviation with each second code in the preset next layer deviation code book or the second standard code corresponding to the second code, determine the second target code corresponding to the second code in the next layer deviation code book with the highest similarity to the current layer coding deviation, and determine the next layer target code based on the second target code.

[0108] Optionally, determining a corresponding target code based on the code with the highest similarity in the codebook includes:

[0109] directly determining the code with the highest similarity in the codebook as the target code; or,

[0110] The codes with the highest similarity in the codebook are encoded based on a preset encoding strategy to generate the target code.

[0111] Optionally, each of the training sample data includes input data and fitting data, or each of the training sample data groups includes reference data, input data and fitting data;

[0112] The target model generation module is used to:

[0113] Inputting the input code of the non-fitting data and the first layer target code of the fitting data into the large language model, and determining the first output code of the large language model; wherein the non-fitting data includes the input data, or the non-fitting data includes the reference data and the input data;

[0114] Determine the second output code of the large language model based on all previously input codes and the second target code of the input fitting data; until determine the Nth output code of the large language model based on all previously input codes of the large language model and the Nth target code of the input fitting data;

[0115] Based on at least one N-layer target encoding of the fitting data and a corresponding number of output encodings, the large language model is iteratively optimized to generate a target model.

[0116] Optionally, the fitting data corresponds to M N-layer target codes, where M is an integer greater than 1;

[0117] The target model generation module is used to:

[0118] Determine a first output code of the large language model based on an input code of the non-fitting data and a first first-layer target code of the fitting data input into the large language model;

[0119] Determine the second output code of the large language model based on all previously input codes and the first second-layer target code and the second first-layer target code of the input fitting data, wherein the first second-layer target code and the second first-layer target code of the fitting data are input by superposition, summation or concatenation; until the M+N-1th output code of the large language model is determined based on all previously input codes of the large language model and the Mth-Nth-layer target code of the input fitting data;

[0120] Combining the M+N-1 output codes of the large language model into M N-layer final output codes;

[0121] Based on at least M N-layer target codes and M N-layer final output codes of the fitting data, the large language model is iteratively optimized to generate a target model.

[0122] The large language model encoding training device provided in the embodiment of the present invention can execute the large language model encoding training method provided in any embodiment of the present invention, and has the corresponding functional modules and beneficial effects of the execution method.

[0123] Embodiment 4

[0124] Figure 4 A schematic diagram of the structure of an electronic device 10 that can be used to implement an embodiment of the present invention is shown. The electronic device is intended to represent various forms of digital computers, such as laptop computers, desktop computers, workstations, personal digital assistants, servers, blade servers, mainframe computers, and other suitable computers. The electronic device can also represent various forms of mobile devices, such as personal digital processing, cellular phones, smart phones, wearable devices (such as helmets, glasses, watches, etc.) and other similar computing devices. The components shown herein, their connections and relationships, and their functions are merely examples and are not intended to limit the implementation of the present invention described and / or required herein.

[0125] like Figure 4 As shown, the electronic device 10 includes at least one processor 11, and a memory connected to the at least one processor 11, such as a read-only memory (ROM) 12, a random access memory (RAM) 13, etc., wherein the memory stores a computer program that can be executed by at least one processor, and the processor 11 can perform various appropriate actions and processes according to the computer program stored in the read-only memory (ROM) 12 or the computer program loaded from the storage unit 18 to the random access memory (RAM) 13. In the RAM 13, various programs and data required for the operation of the electronic device 10 can also be stored. The processor 11, the ROM 12, and the RAM 13 are connected to each other through a bus 14. An input / output (I / O) interface 15 is also connected to the bus 14.

[0126] A number of components in the electronic device 10 are connected to the I / O interface 15, including: an input unit 16, such as a keyboard, a mouse, etc.; an output unit 17, such as various types of displays, speakers, etc.; a storage unit 18, such as a disk, an optical disk, etc.; and a communication unit 19, such as a network card, a modem, a wireless communication transceiver, etc. The communication unit 19 allows the electronic device 10 to exchange information / data with other devices through a computer network such as the Internet and / or various telecommunication networks.

[0127] The processor 11 may be a variety of general and / or special processing components with processing and computing capabilities. Some examples of the processor 11 include, but are not limited to, a central processing unit (CPU), a graphics processing unit (GPU), various dedicated artificial intelligence (AI) computing chips, various processors running machine learning model algorithms, digital signal processors (DSPs), and any appropriate processors, controllers, microcontrollers, etc. The processor 11 executes the various methods and processes described above, such as a large language model encoding training method.

[0128] In some embodiments, the large language model encoding training method may be implemented as a computer program, which is tangibly contained in a computer-readable storage medium, such as a storage unit 18. In some embodiments, part or all of the computer program may be loaded and / or installed on the electronic device 10 via the ROM 12 and / or the communication unit 19. When the computer program is loaded into the RAM 13 and executed by the processor 11, one or more steps of the large language model encoding training method described above may be performed. Alternatively, in other embodiments, the processor 11 may be configured to execute the large language model encoding training method in any other appropriate manner (e.g., by means of firmware).

[0129] Various implementations of the systems and techniques described above herein can be implemented in digital electronic circuit systems, integrated circuit systems, field programmable gate arrays (FPGAs), application specific integrated circuits (ASICs), application specific standard products (ASSPs), systems on chips (SOCs), load programmable logic devices (CPLDs), computer hardware, firmware, software, and / or combinations thereof. These various implementations can include: being implemented in one or more computer programs that can be executed and / or interpreted on a programmable system including at least one programmable processor, which can be a special purpose or general purpose programmable processor that can receive data and instructions from a storage system, at least one input device, and at least one output device, and transmit data and instructions to the storage system, the at least one input device, and the at least one output device.

[0130] Computer programs for implementing the methods of the present invention may be written in any combination of one or more programming languages. These computer programs may be provided to a processor of a general-purpose computer, a special-purpose computer, or other programmable data processing device, so that when the computer program is executed by the processor, the functions / operations specified in the flow chart and / or block diagram are implemented. The computer program may be executed entirely on the machine, partially on the machine, partially on the machine and partially on a remote machine as a stand-alone software package, or entirely on a remote machine or server.

[0131] In the context of the present invention, a computer-readable storage medium may be a tangible medium that may contain or store a computer program for use by or in conjunction with an instruction execution system, device, or equipment. A computer-readable storage medium may include, but is not limited to, electronic, magnetic, optical, electromagnetic, infrared, or semiconductor systems, devices, or equipment, or any suitable combination of the foregoing. Alternatively, a computer-readable storage medium may be a machine-readable signal medium. A more specific example of a machine-readable storage medium may include an electrical connection based on one or more lines, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the foregoing.

[0132] To provide interaction with a user, the systems and techniques described herein may be implemented on an electronic device having: a display device (e.g., a CRT (cathode ray tube) or LCD (liquid crystal display) monitor) for displaying information to the user; and a keyboard and a pointing device (e.g., a mouse or trackball) through which the user can provide input to the electronic device. Other types of devices may also be used to provide interaction with the user; for example, the feedback provided to the user may be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and input from the user may be received in any form (including acoustic input, voice input, or tactile input).

[0133] The systems and techniques described herein may be implemented in a computing system that includes backend components (e.g., as a data server), or a computing system that includes middleware components (e.g., an application server), or a computing system that includes frontend components (e.g., a user computer with a graphical user interface or a web browser through which a user can interact with implementations of the systems and techniques described herein), or a computing system that includes any combination of such backend components, middleware components, or frontend components. The components of the system may be interconnected by any form or medium of digital data communication (e.g., a communication network). Examples of communication networks include: a local area network (LAN), a wide area network (WAN), a blockchain network, and the Internet.

[0134] A computing system may include a client and a server. The client and the server are generally remote from each other and usually interact through a communication network. The client and server relationship is generated by computer programs running on the corresponding computers and having a client-server relationship with each other. The server may be a cloud server, also known as a cloud computing server or cloud host, which is a host product in the cloud computing service system to solve the defects of difficult management and weak business scalability in traditional physical hosts and VPS services.

[0135] It should be understood that the various forms of processes shown above can be used to reorder, add or delete steps. For example, the steps described in the present invention can be executed in parallel, sequentially or in different orders, as long as the desired results of the technical solution of the present invention can be achieved, and this document does not limit this.

[0136] The above specific implementations do not constitute a limitation on the protection scope of the present invention. It should be understood by those skilled in the art that various modifications, combinations, sub-combinations and substitutions can be made according to design requirements and other factors. Any modification, equivalent substitution and improvement made within the spirit and principle of the present invention should be included in the protection scope of the present invention.

Claims

1. A large language model encoding training method, characterized in that: include: Acquire a training data set; wherein the training data set includes at least two training sample data groups, each of which is composed of at least two corresponding training data; each of which includes fitting music and target lyrics, or each of which includes reference music, fitting music and target lyrics; For at least one training data in each of the training sample data groups, determining a first-layer target encoding of the training data; Taking the first layer target code as the current layer target code, determining the current layer code deviation of the current layer target code, and determining the next layer target code based on the current layer code deviation; wherein the current layer code deviation is the difference between the current layer initial code and the current layer target code; The next layer target code is updated to the current layer target code, and the current layer code deviation for determining the current layer target code is returned to determine the Nth layer target code; wherein N is an integer greater than or equal to 2; The large language model is trained based on the N layers of target encoding to generate a target model.

2. The method according to claim 1, characterized in that For at least one training data in each of the training sample data groups, before determining the first layer target encoding of the training data, the method further includes: Extracting partial training data of at least one dimension for the at least one training data; Encode part of the training data of at least one dimension respectively to generate at least one layer of enhanced target encoding; The large language model is trained based on the target encoding of the N layers to generate a target model, including: At least one layer of the enhanced target code is superimposed on N layers of the target code, and the large language model is trained to generate a target model.

3. The method according to claim 2, characterized in that The step of superimposing at least one layer of the enhanced target code with N layers of the target code to train the large language model includes: At least one layer of the enhanced target coding is placed before the N layers of the target coding, and at least part of the enhanced target coding is input into the large language model before the N layers of the target coding to perform the large language model training.

4. The method according to any one of claims 1 to 3, characterized in that: Before determining the first-layer target encoding of at least one training data in each of the training sample data groups, the method further includes: Discretizing the at least one training data to generate a plurality of minimum discrete units; For at least one training data in each of the training sample data groups, determining a first-layer target encoding of the training data comprises: Each minimum discrete unit of at least one training data in each of the training sample data groups is encoded respectively to determine the first layer target encoding of the training data.

5. The method according to claim 1, characterized in that Determining a first layer target encoding of the training data includes: Determining a first layer initial encoding of the training data; Compare the first layer initial code with each first code in a preset code book or a first standard code corresponding to the first code, and determine a first target code corresponding to the first code in the code book that has the highest similarity to the first layer initial code; A first layer target coding is determined based on the first target coding.

6. The method according to claim 1, characterized in that Determining a current layer coding deviation of the current layer target coding, and determining a next layer target coding based on the current layer coding deviation, comprising: Calculate the residual between the current layer initial code and the current layer target code, and use the residual as the current layer code deviation of the current layer code; Compare the current layer coding deviation with each second code in the preset next layer deviation code book or the second standard code corresponding to the second code, determine the second target code corresponding to the second code in the next layer deviation code book with the highest similarity to the current layer coding deviation, and determine the next layer target code based on the second target code.

7. The method according to claim 5 or 6, characterized in that: Determine the corresponding target code based on the code with the highest similarity in the codebook, including: directly determining the code with the highest similarity in the codebook as the target code; or, The codes with the highest similarity in the codebook are encoded based on a preset encoding strategy to generate the target code.

8. The method according to claim 1, characterized in that Each of the training sample data includes input data and fitting data, or each of the training sample data groups includes reference data, input data and fitting data, and the large language model is trained based on the N layers of target encoding to generate a target model, including: Inputting the input code of the non-fitting data and the first layer target code of the fitting data into the large language model, and determining the first output code of the large language model; wherein the non-fitting data includes the input data, or the non-fitting data includes the reference data and the input data; Determine the second output code of the large language model based on all previously input codes and the second target code of the input fitting data; until determine the Nth output code of the large language model based on all previously input codes of the large language model and the Nth target code of the input fitting data; Based on at least one N-layer target encoding of the fitting data and a corresponding number of output encodings, the large language model is iteratively optimized to generate a target model.

9. The method according to claim 8, characterized in that The fitting data corresponds to M N-layer target codes, where M is an integer greater than 1, and the method further includes: Determine a first output code of the large language model based on an input code of the non-fitting data and a first first-layer target code of the fitting data input into the large language model; Determine the second output code of the large language model based on all previously input codes and the first second-layer target code and the second first-layer target code of the input fitting data, wherein the first second-layer target code and the second first-layer target code of the fitting data are input by superposition, summation or concatenation; until the M+N-1th output code of the large language model is determined based on all previously input codes of the large language model and the Mth-Nth-layer target code of the input fitting data; Combining the M+N-1 output codes of the large language model into M N-layer final output codes; Based on at least M N-layer target codes and M N-layer final output codes of the fitting data, the large language model is iteratively optimized to generate a target model.

10. A large language model encoding training device, characterized in that: include: A training data set acquisition module is used to acquire a training data set; wherein the training data set includes at least two training sample data groups, each of which is composed of at least two corresponding training data; each of which includes fitting music and target lyrics, or each of which includes reference music, fitting music and target lyrics; A first-layer target coding determination module, used for determining the first-layer target coding of the training data for at least one training data in each of the training sample data groups; a next layer target coding determination module, configured to use the first layer target coding as the current layer target coding, determine the current layer coding deviation of the current layer target coding, and determine the next layer target coding based on the current layer coding deviation; wherein the current layer coding deviation is the difference between the current layer initial coding and the current layer target coding; A coding deviation loop determination module, used for updating the next layer target coding to the current layer target coding, and returning to execute the current layer coding deviation for determining the current layer target coding, and determining the Nth layer target coding; wherein N is an integer greater than or equal to 2; The target model generation module is used to train the large language model based on the N layers of target encoding to generate a target model.

Citation Information

Patent Citations

  • Training device and method

    CN110909870A

  • Neural network architecture searching method and device

    CN118251679A