Model code generation method and device
By optimizing the encoding generation method of the big model, using the combination of inter-layer data and compression rate, the problems of low efficiency and insufficient integrity of the traditional big model generation are solved, and efficient and excellent content generation is achieved.
Patent Information
- Application Number
- CN202510780088.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-12
- Publication Date
- 2025-07-11
- Estimated Expiration
- 2045-06-12
AI Technical Summary
Traditional big models use autoregression to generate and encode for a long time, low efficiency, and insufficient integrity and quality of the generated content, especially lacking overall design and artistic conception in text, pictures and music generation.
By determining the input data of the target model at the current layer, based on the input data and output data encoded by the current layer, and combining the compression rate of the next layer, the parallel generation of the number of coded layers and the encoding information is optimized until the preset number of layers is reached and the target output information is decoded to obtain.
The efficiency and integrity of the model generated content is improved, the quality of the generated content is ensured, and the parallel generation of coded information is achieved and the overall style coordination is achieved.
Smart Images

Figure CN120301432A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of machine learning, and particularly to a method and device for encoding generation of a model. Background Art
[0002] Traditional large models all output codes sequentially in an autoregressive manner. The large model takes a long time to output codes, and the model generation efficiency is low. Moreover, since the large model generates one token at a time and finally splices multiple tokens into complete information for output, the integrity of the content generated by the large model is relatively poor. For example, when the large model outputs a piece of text, it generates one word segment at a time, and then combines multiple word segments into text. Although the generated text is very coherent, the content is empty and meaningless. When the model outputs a picture, it generates one pixel at a time, and then combines multiple pixels into a picture. Although the details of the generated picture are very rich, the overall style or content does not match the generation requirements, lacking overall design, composition, and artistic conception. When the large model outputs music, it outputs one frame of audio at a time, and then combines multiple audio frames into music. Although the music details of each frame are rich enough, the integrity of the output music is insufficient, and the distinction between the verse, chorus, prelude, climax, and ending is not obvious, and the overall quality of music generation is not high. Summary of the Invention
[0003] The present invention provides a method and device for encoding generation of a model, so as to greatly improve the efficiency of the content generated by the model on the basis of ensuring the integrity and quality of the content generated by the model.
[0004] According to one aspect of the present invention, there is provided a method for encoding generation of a model, the method comprising:
[0005] Determine the input data of the encoding at the current layer of the target model; the target model is a large language model;
[0006] Based on the input data of the encoding at the current layer, obtain the output data of the encoding at the current layer; the output data includes a first preset number of encoding information;
[0007] Based on the input data of the encoding at the current layer, the output data of the encoding at the current layer, and the compression ratio of the encoding at the next layer, determine the input data of the encoding at the next layer;
[0008] Until the number of encoding layers of the target model reaches the preset number of encoding layers, then form the output data of each layer of encoding into an encoding data set, so as to decode the encoding data set to obtain the target output information of the target model.
[0009] According to another aspect of the present invention, there is provided a device for encoding generation of a model, the device comprising:
[0010] A first data determination module, configured to determine input data encoded by a target model at a current layer; the target model is a large language model;
[0011] An encoding module, configured to obtain output data encoded by the current layer based on the input data encoded by the current layer; the output data includes a first preset number of encoding information;
[0012] A second data determination module, configured to determine input data for encoding in a next layer based on the input data encoded by the current layer, the output data encoded by the current layer, and a compression ratio of encoding in the next layer;
[0013] A determination module, configured to until the number of encoding layers of the target model reaches a preset number of encoding layers, form an encoded data set with the output data of each layer of encoding, so as to decode the encoded data set to obtain target output information of the target model.
[0014] According to another aspect of the present invention, there is provided an electronic device, the electronic device includes:
[0015] At least one processor; and
[0016] A memory communicatively connected to the at least one processor; wherein,
[0017] The memory stores a computer program executable by the at least one processor, and the computer program is executed by the at least one processor so that the at least one processor can execute the encoding generation method of the model according to any embodiment of the present invention.
[0018] According to another aspect of the present invention, there is provided a computer-readable storage medium, the computer-readable storage medium stores computer instructions, and the computer instructions are used to implement the encoding generation method of the model according to any embodiment of the present invention when executed by a processor.
[0019] The technical solution of the embodiment of the present invention determines the input data encoded by the target model at the current layer; the target model is a large language model; based on the input data encoded by the current layer, the output data encoded by the current layer is obtained; the output data includes the first preset number of encoded information, realizing that the encoded information of the same layer can be generated in parallel, greatly improving the parallel degree of the model to generate encoded information; further, based on the input data encoded by the current layer, the output data encoded by the current layer, and the compression rate of the next layer of encoding, the input data of the next layer of encoding is determined for the next layer of encoding to perform an inference operation based on the input data of the next layer of encoding; until the number of encoding layers of the target model reaches the preset number of encoding layers, the output data of each layer of encoding is formed into an encoded data set to decode the encoded data set to obtain the target output information of the target model, reducing the number of times the model outputs encoded information, and realizing the improvement of the efficiency of the model to generate content on the basis of ensuring the integrity and quality of the content generated by the model.
[0020] It should be understood that the content described in this part is not intended to identify the key or important features of the embodiments of the present invention, nor is it used to limit the scope of the present invention. Other features of the present invention will become easily understood through the following description. BRIEF DESCRIPTION OF THE DRAWINGS
[0021] In order to more clearly illustrate the technical solutions in the embodiments of the present invention, the following will briefly introduce the accompanying drawings required for the description of the embodiments. Obviously, the accompanying drawings in the following description are only some embodiments of the present invention. For those of ordinary skill in the art, other accompanying drawings can be obtained based on these drawings without creative efforts.
[0022] Figure 1 is a flowchart of a method for encoding and generating a model according to an embodiment of the present invention;
[0023] Figure 2 is a flowchart of another method for encoding and generating a model according to an embodiment of the present invention;
[0024] Figure 3 is a schematic structural diagram of an apparatus for encoding and generating a model according to an embodiment of the present invention;
[0025] Figure 4 is a schematic structural diagram of an electronic device for implementing the method for encoding and generating a model of the embodiment of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0026] To enable those skilled in the art to better understand the solution of the present invention, the following will clearly and completely describe the technical solutions in the embodiments of the present invention in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the scope of protection of the present invention.
[0027] It should be noted that the terms "first", "second", etc. in the description and claims of the present invention and the above-mentioned drawings are used to distinguish similar objects, and do not necessarily need to be used to describe a specific order or sequence. It should be understood that such data can be interchanged under appropriate circumstances so that the embodiments of the present invention described herein can be implemented in an order other than those illustrated or described herein. In addition, the terms "including" and "having" and any variations thereof are intended to cover non-exclusive inclusion. For example, a process, method, system, product, or device that includes a series of steps or units does not necessarily have to be limited to those steps or units clearly listed, but may include other steps or units not clearly listed or inherent to these processes, methods, products, or devices.
[0028] Embodiment 1
[0029] Figure 1 This is a flowchart of a method for generating encoding of a model provided in an embodiment of the present invention. This embodiment is applicable to the encoding generation situation in the process of generating content by a large language model in an autoregressive manner. This method can be executed by a model encoding generation device, which can be implemented in the form of hardware and / or software, and the model encoding generation device can be configured in any electronic device with network communication functions. As Figure 1 shown, the method for generating encoding of a model includes the following processes:
[0030] S110. Determine the input data for encoding at the current layer of the target model; the target model is a large language model.
[0031] Among them, the target model is pre-trained. Before training the target model, it is necessary to train the encoder-decoder of the target model to fix parameters such as the number of feature extraction layers of the encoder-decoder and the compression ratio of each layer of encoding, so as to reduce the number of encoding layers of the target model. Because the encoder-decoder of traditional large language models generally generates only one layer of encoded information and cannot perform hierarchical feature compression to output multiple encoded information. For example, if 256 encoded information need to be output, the traditional large language model needs to perform 256 times of encoding output. However, the fixing of the number of feature extraction layers of the encoder-decoder of the present invention and the compression ratio of each layer of encoding can make at least one encoded information be output for each layer of encoding, and the target model only needs to perform encoding output for the preset number of encoding layers. For example, compared with the 256 times of encoding output of the traditional large language model, the present invention does not need to perform 256 times of encoding output, but only needs to perform encoding output for a number of times roughly equivalent to the preset number of encoding layers, which greatly improves the efficiency of the model to generate content. Among them, the preset number of encoding layers is determined according to the number of feature extraction layers of the encoder-decoder in the target model, that is, the preset number of encoding layers is the same as the number of feature extraction layers of the encoder-decoder in the target model.
[0032] Specifically, after the target model is successfully trained, the first information can be input into the target model, and the target model internally performs an inference operation for the preset number of encoding layers to obtain the target output information of the target model. Further, in order to ensure the accuracy of the inference operation for the preset number of encoding layers, it is necessary to ensure the accuracy of the input data for each layer of encoding, that is, to determine the input data of the target model in the current layer of encoding, including steps A1-A2:
[0033] Step A1, if the current layer of encoding is the first layer of encoding, the input data of the current layer of encoding is the first information obtained in response to the input operation; the input operation is an input operation for inputting the first information into the target model triggered by the outside of the target model.
[0034] Among them, the first information can be understood as data that prompts the target model to generate the target output information and is data input from the outside into the target model; for example, the question in the question-and-answer data.
[0035] The target model can be a text large model, a visual large model or a sound large model. The text large model is used for text content analysis. For the text large model, the first information can be the proposed text document information; for example, "Arrange a 3-day travel itinerary to a certain place". The visual large model can be a model for image generation. For the visual large model, the first information can be information data describing the output image; for example, "Generate an image of a family celebrating the New Year". The sound large model can be a model for generating sound or music; for the sound large model, the first information can be sound data or music data; for example, "A piece of lyrics".
[0036] Step A2: If the current layer code is a code of other layers except the first layer code, the input data of the current layer code is determined based on the input data of the previous layer code, the output data of the previous layer code, and the compression ratio of the current layer code.
[0037] Among them, the compression ratio of the current layer code is the compression ratio between the current layer code and the previous layer code. The compression ratio is determined according to at least one of the convolution kernel size, the stride, the padding number, and the coding cross range between adjacent coding information. The compression ratio of each layer code can be the same or different, and they are all parameters determined during the training of the codec. For example, the compression ratio can be values such as 1.5, 2, 3, 4, etc. The smaller the compression ratio, the less information loss during the inference operation of the target model, and the higher the quality of the content generated by the target model; the larger the compression ratio, the faster the content generation rate of the target model and the higher the generation efficiency; therefore, the setting of the compression ratio is very important, that is, a suitable compression ratio can ensure both the quality of the content generated by the target model and the generation efficiency of the target model.
[0038] Among them, the adjacent coding information may not cross or may cross. The coding cross range can be adjusted by parameters such as the convolution kernel size and the stride, and then the compression ratio is adjusted through the coding cross range between adjacent coding information, so that the information extraction ranges of adjacent coding information cross. This can make the content connection between different coding information in the same layer code smoother and avoid the problem that the coding information of two adjacent frames is completely irrelevant or diametrically opposite; in addition, the information extraction ranges of two adjacent coding information cross, which can make the cross area refer to the features of multiple coding information during decoding, and then achieve a smoother transition through methods such as taking the average value.
[0039] Specifically, the input data of other layer codes except the first layer code can be obtained inside the target model. To ensure the quality of the output data between adjacent layer codes, the input data of each layer code can be determined by the input data of the previous layer code, the output data of the previous layer code, and the compression ratio of the current layer code, ensuring the accuracy of the input data of each layer code.
[0040] S120: Obtain the output data of the current layer code based on the input data of the current layer code; the output data includes a first preset number of coding information.
[0041] Specifically, the target model performs an inference operation on the input data encoded by the current layer and outputs the output data encoded by the current layer; and the output data includes a first preset number of encoded information, that is, the output data of each layer of encoding contains at least one encoded information. In addition, as the number of encoding layers increases, the more encoded information is included in the output data of the corresponding layer of encoding, that is, the more detailed information is included. The output data of the initial encoding layer more reflects the layout, style, design, etc. of the overall content. That is, the present invention not only ensures the integrity of the generated content but also controls the details of the generated content, ensuring higher quality of the generated content.
[0042] In an embodiment of the present invention, optionally, the output data includes a first preset number of encoded information, and the first preset number is related to the number of encoded information in the input data of the current layer of encoding. There is an association relationship between the number of encoded information in the input data of the current layer of encoding, the compression rate of the current layer of encoding, and the output data of the previous layer of encoding. That is, according to the association relationship between the compression rate of the current layer of encoding and the output data of the previous layer of encoding, the first preset number of the output data of the current layer of encoding can be accurately obtained. The setting of the compression rate can ensure that the number of generated encoded information is within a suitable range, ensuring the accuracy of the output data, that is, being able to accurately control the overall and details of the output content.
[0043] S130. Determine the input data of the next layer of encoding based on the input data of the current layer of encoding, the output data of the current layer of encoding, and the compression rate of the next layer of encoding.
[0044] Wherein, the compression rate of the next layer of encoding is the compression rate between the current layer of encoding and the next layer of encoding.
[0045] Specifically, each encoded information in the output data of the current layer of encoding is copied a preset number of copies to obtain first encoded data, and the preset number of copies is the product of the first preset number and the compression rate of the next layer of encoding. Further, the first encoded data and the input data of the current layer of encoding are combined to determine the input data of the next layer of encoding.
[0046] Optionally, combining the first encoded data and the input data of the current layer of encoding to determine the input data of the next layer of encoding may include: adjusting the number of the same encoded information in the input data of the current layer of encoding to a preset number of copies to obtain second encoded data; combining the first encoded data and the second encoded data to determine the input data of the next layer of encoding.
[0047] Exemplarily, the input data of the first - layer encoding is the first information a, and the output data of the first - layer encoding is the encoded information a1; the compression ratio of the second - layer encoding is 2, and the input data of the second - layer encoding is two first reference encoded information, and each first reference encoded information is the same. The first reference encoded information is the first information a and the encoded information a1, but the encoding position information corresponding to the two first reference encoded information is different. Further, the output data of the second - layer encoding obtained based on the input data of the second - layer encoding includes 2 encoded information, namely, encoded information b1 and encoded information b2; the compression ratio of the third - layer encoding is 2, then the preset number of copies is 2*2 = 4. The first encoded data of the third - layer encoding is encoded information b1, encoded information b1, encoded information b1, encoded information b1, encoded information b2, encoded information b2, encoded information b2, and encoded information b2; the second encoded data of the third - layer encoding is the first information a, the first information a, the first information a, the first information a, encoded information a1, encoded information a1, encoded information a1, and encoded information a1; the input data of the third - layer encoding is four second reference encoded data, and each second reference encoded data is the same. Each second reference encoded data is the first information a, encoded information a1, encoded information b1, and encoded information b2, but the encoding position information corresponding to each second reference encoded information can be different. And so on until the encoding layer of the target model reaches the preset encoding layer.
[0048] S140. Until the encoding layer of the target model reaches the preset encoding layer, the output data of each layer of encoding is formed into an encoded data set, so as to decode the encoded data set to obtain the target output information of the target model.
[0049] Specifically, until the encoding layer of the target model reaches the preset encoding layer, that is, the encoding times of the target model have been completed, and the output data of each layer of encoding has been saved in the encoded data set. In order to output the accurate target output information of the target model, it is necessary to decode the output data of each layer of encoding in the encoded data set to obtain the decoded data of the output data of each layer of encoding, and further combine the decoded data of each layer of encoding to obtain the target output information of the target model.
[0050] The technical solution of the embodiment of the present invention determines the input data encoded by the target model at the current layer; the target model is a large language model; based on the input data encoded by the current layer, the output data encoded by the current layer is obtained; the output data includes a first preset number of encoded information, realizing that the encoded information of the same layer can be generated in parallel, greatly improving the parallel degree of the model to generate encoded information; further, based on the input data encoded by the current layer, the output data encoded by the current layer, and the compression ratio of the next layer encoding, the input data of the next layer encoding is determined for the next layer encoding to perform an inference operation based on the input data of the next layer encoding; until the number of encoding layers of the target model reaches the preset number of encoding layers, the output data of each layer encoding is composed into an encoded data set to decode the encoded data set to obtain the target output information of the target model, reducing the number of times the model outputs encoded information, and realizing improving the efficiency of the model to generate content on the basis of ensuring the integrity and quality of the content generated by the model.
[0051] Embodiment 2
[0052] Figure 2 The flowchart of another method for encoding and generating a model provided by the embodiment of the present invention. The technical solution of this embodiment further optimizes the process of "determining the input data of the next layer encoding based on the input data encoded by the current layer, the output data encoded by the current layer, and the compression ratio of the next layer encoding" on the basis of the above embodiment. This embodiment can be combined with each optional solution in one or more of the above embodiments. As Figure 2 shown, the method for encoding and generating a model includes:
[0053] S210. Determine the input data encoded by the target model at the current layer; the target model is a large language model.
[0054] S220. Based on the input data encoded by the current layer, obtain the output data encoded by the current layer; the output data includes a first preset number of encoded information.
[0055] Among them, the input data encoded by the current layer further includes the second encoding position information of each encoded information. The second encoding position information can be understood as the position information corresponding to each encoded information to perform an inference operation.
[0056] Specifically, based on the autoregressive mechanism of the target model and the second encoding position information of each encoded information, an inference operation is performed on each encoded information in the input data encoded by the current layer to obtain the output data encoded by the current layer.
[0057] S230. Based on the input data encoded by the current layer and the output data encoded by the current layer, determine the data content of the input data of the next layer encoding.
[0058] Specifically, the input data encoded in the current layer and the output data encoded in the current layer are combined as the data content of the input data for the encoding of the next layer. By way of example, the input data encoded in the current layer are two first reference encoding messages, and each first reference encoding message is the same. The first reference encoding message is the first information a and the encoding message a1; the output data encoded in the current layer includes the encoding message b1 and the encoding message b2, corresponding to one input first reference encoding message respectively; then the data content of one input data for the encoding of the next layer is the first reference encoding message, the encoding message b1 and the encoding message b2.
[0059] The encoding in the current layer is the first-layer encoding, and the prediction length is greater than the preset length. The first preset quantity is 1. The input data for the first-layer encoding is the first information. The output data for the first-layer encoding includes 1 third encoding message, the second preset quantity of first encoding position messages, and the deviation values between the third encoding message and each fourth encoding message. The second preset quantity is the value obtained by rounding up the ratio between the prediction length and the preset length. The third encoding message can be the average value of the fourth encoding messages; the fourth encoding messages are the second preset quantity of encoding messages output one by one by the first-layer encoding of the target model based on the first information when the prediction length is greater than the preset length. There is a one-to-one correspondence between the fourth encoding messages and the first encoding position messages. The data content of the input data for the encoding of the next layer is the first information, 1 third encoding message, the second preset quantity of first encoding position messages, and the deviation values between the third encoding message and each fourth encoding message.
[0060] S240. Determine the number of encoding messages in the input data for the encoding of the next layer based on the product of the first preset quantity and the compression ratio of the encoding of the next layer; or, determine the number of encoding messages in the input data for the encoding of the next layer based on the product of the first preset quantity and the compression ratio of the encoding of the next layer and the first encoding position messages of the output data for the encoding of the current layer.
[0061] Specifically, if the encoding in the current layer is the first-layer encoding, and the prediction length is less than or equal to the preset length, and the first preset quantity is 1, then the number of encoding messages in the input data for the encoding of the next layer is the product of the first preset quantity and the compression ratio of the encoding of the next layer.
[0062] If the encoding in the current layer is the first-layer encoding, and the prediction length is greater than the preset length, and the first preset quantity is the value obtained by rounding up the ratio between the prediction length and the preset length, then the number of encoding messages in the input data for the encoding of the next layer is the product of the first preset quantity and the compression ratio of the encoding of the next layer.
[0063] If the current layer encoding is the first layer encoding, and the predicted length is greater than the preset length, and the first preset quantity is 1, then the number of encoded information in the input data of the next layer encoding is determined by the product of the first preset quantity and the compression ratio of the next layer encoding, and the first encoding position information of the output data of the current layer encoding. Specifically, based on the product of the first preset quantity and the compression ratio of the next layer encoding, and the first encoding position information of the output data of the current layer encoding, to determine the number of encoded information in the input data of the next layer encoding, it includes: the output data of the first layer encoding further includes a second preset quantity of first encoding position information, and the number of encoded information in the input data of the next layer encoding is the product of the first preset quantity, the compression ratio of the next layer encoding and the second preset quantity.
[0064] If the current layer encoding is other layer encodings except the first layer encoding, the number of encoded information in the input data of the next layer encoding is the product of the first preset quantity and the compression ratio of the next layer encoding.
[0065] S250. Determine the input data of the next layer encoding based on the data content of the input data of the next layer encoding and the number of encoded information in the input data of the next layer encoding.
[0066] Specifically, if the current layer encoding is the first layer encoding, and the number of encoded information in the input data of the next layer encoding is the product of the first preset quantity and the compression ratio of the next layer encoding, and the data content of the input data of the next layer encoding includes multiple encoded information, then adjust the number of encoded information in the data content of the input data of the next layer encoding to the number of encoded information in the input data of the next layer encoding, and combine the adjusted encoded information to obtain multiple identical encoded information as the input data of the next layer encoding. By way of example, the input data of the first layer encoding is the first information a, the output data of the first layer encoding is the encoded information a1, the compression ratio of the second layer encoding is 2, then the data content of the input data of the second layer encoding is the first information a and the encoded information a1, and the number of encoded information in the input data of the second layer encoding is 2; further, both the first information a and the encoded information a1 are made into two copies, and combined to obtain two identical encoded information, and each encoded information is a combination of the first information a and the encoded information a1, wherein, the position information corresponding to the two encoded information can be different.
[0067] If the current layer encoding is the first layer encoding and the predicted length is greater than the preset length, the first preset quantity is 1, the number of encoding information in the input data of the next layer encoding is the product of the first preset quantity and the compression ratio of the next layer encoding and the second preset quantity, and the data content of the input data of the next layer encoding is the first information, a third encoding information, the second preset quantity of first encoding position information, and the deviation values between the third encoding information and each fourth encoding information. Further, it is necessary to obtain the second preset quantity of fifth encoding information for the actual first layer encoding output through the third encoding information, the second preset quantity of first encoding position information, and the deviation values between the third encoding information and each fourth encoding information, and then combine the second preset quantity of fifth encoding information with the first information as the sixth encoding information, and the number of the sixth encoding information is the number of encoding information in the input data of the next layer encoding, and use multiple sixth encoding information as the input data of the next layer encoding.
[0068] Exemplarily, the data content of the input data of the next layer encoding is the first information a, a third encoding information A, 2 first encoding position information, and the deviation values between the third encoding information A and each fourth encoding information (the fourth encoding information a1, the fourth encoding information a2). The number of encoding information in the input data of the next layer encoding is 4. The fourth encoding information a1 and the fourth encoding information a2 are obtained through the deviation values between the third encoding information A and each fourth encoding information and the first encoding position information. Then, each sixth encoding information of the input data of the next layer encoding is the combination of the first information a, the fourth encoding information a1, and the fourth encoding information a2, and the input data of the next layer encoding includes 4 sixth encoding information.
[0069] If the current layer encoding is other layer encodings except the first layer encoding, the number of encoding information in the input data of the next layer encoding is the product of the first preset quantity and the compression ratio of the next layer encoding. If the data content of the input data of the next layer encoding includes multiple encoding information, then adjust the number of encoding information in the data content of the input data of the next layer encoding to the number of encoding information in the input data of the next layer encoding, and perform encoding combination on the adjusted encoding information to obtain multiple identical encoding information as the input data of the next layer encoding. The process of encoding combination can be understood as combining different encoding information included in the data content of the input data of the next layer encoding together as one encoding information included in the input data of the next layer encoding.
[0070] Exemplarily, the input data encoded at the current layer includes two pieces of first reference encoding information, where the first reference encoding information is a combination of first information a and encoding information a1. The output data encoded at the current layer includes encoding information b1 and encoding information b2. The data content of the input data for the next layer encoding includes two pieces of first reference encoding information, encoding information b1, and encoding information b2. The compression ratio of the next layer encoding is 2, and the number of encoding information in the input data for the next layer encoding is 2 * 2 = 4. Then, the two pieces of first reference encoding information are changed into four pieces of reference encoding information, the encoding information b1 is adjusted to 4 pieces of encoding information b1, and the encoding information b2 is adjusted to 4 pieces of encoding information b2. Further, the different encoding information is combined together as one piece of encoding information included in the input data for the next layer encoding, that is, the second reference encoding data, and the second reference encoding data is a combination of the first reference encoding information, encoding information b1, and encoding information b2. Then, the input data for the next layer encoding includes four pieces of second reference encoding data, where the position information corresponding to the four pieces of encoding information can be different.
[0071] S260. Until the number of encoding layers of the target model reaches the preset number of encoding layers, the output data of each layer of encoding is composed into an encoding data set to decode the encoding data set to obtain the target output information of the target model.
[0072] The technical solution of the embodiment of the present invention determines the input data encoded at the current layer of the target model; the target model is a large language model. Based on the input data encoded at the current layer, the output data encoded at the current layer is obtained; the output data includes a first preset number of encoding information. Based on the input data encoded at the current layer and the output data encoded at the current layer, the data content of the input data for the next layer encoding is determined, realizing the accurate determination of the data content, so as to facilitate the subsequent combination of the data content according to the number of the input data for the next layer encoding. At the same time, based on the product of the first preset number and the compression ratio of the next layer encoding, the number of encoding information in the input data for the next layer encoding is determined; or, based on the product of the first preset number and the compression ratio of the next layer encoding and the first encoding position information of the output data encoded at the current layer, the number of encoding information in the input data for the next layer encoding is determined; ensuring the accuracy of the number of encoding information in the input data for the next layer encoding, so as to accurately determine the input data for the next layer encoding based on the data content of the input data for the next layer encoding and the number of encoding information in the input data for the next layer encoding, for the next layer encoding to perform an inference operation, realizing that the encoding information of the same layer can be generated in parallel, greatly improving the parallel degree of the model generating encoding information; until the number of encoding layers of the target model reaches the preset number of encoding layers, the output data of each layer of encoding is composed into an encoding data set to decode the encoding data set to obtain the target output information of the target model, reducing the number of times of the model encoding information output, and realizing improving the efficiency of the model generating content on the basis of ensuring the integrity and quality of the model generating content.
[0073] Example 3
[0074] The technical solution of this embodiment further optimizes the process of "obtaining the output data of the current layer encoding based on the input data of the current layer encoding" in the foregoing embodiment on the basis of the foregoing embodiment. Optionally, the first preset quantity is related to the quantity of encoding information in the input data of the current layer encoding, and there is an association relationship between the quantity of encoding information in the input data of the current layer encoding, the compression ratio of the current layer encoding, and the output data of the previous layer encoding. This embodiment can be combined with each optional solution in one or more of the foregoing embodiments.
[0075] Specifically, the first preset quantity is related to the quantity of encoding information in the input data of the current layer encoding, and there is an association relationship between the quantity of encoding information in the input data of the current layer encoding, the compression ratio of the current layer encoding, and the output data of the previous layer encoding. Specifically, the following steps B1-B2 can be referred to:
[0076] Step B1: If the current layer encoding is the first layer encoding, the first preset quantity is determined based on the prediction length of the target output information; the prediction length is predicted by the target model based on the first information.
[0077] Among them, the pre-trained target model has the ability to predict the length of the target output information according to the length of the first information input to the target model, so that in the actual application process, when the first information is input to the target model, the prediction length can be accurately obtained. Further, according to the prediction length, the quantity of encoding information included in the output data of the first layer encoding, that is, the first preset quantity, is determined.
[0078] Specifically, when the prediction length is less than or equal to the preset length, the first preset quantity is 1; the preset length can be the maximum length of one encoding information. That is, when the prediction length is equal to the preset length, the output data of the current layer encoding includes 1 first encoding information, and the content in the first encoding information is all meaningful data; when the prediction length is less than the preset length, the output data of the current layer encoding includes 1 second encoding information, and the second encoding information includes meaningful data of the prediction length and meaningless data of the reference length; the reference length is the difference between the preset length and the prediction length; the meaningless data of the reference length is information filled in by special encoding, and when decoding, only the meaningful data of the prediction length can be decoded.
[0079] When the prediction length is greater than the preset length, the first preset quantity is the value obtained by rounding up the ratio between the prediction length and the preset length. By way of example, taking the voice large model to generate music as an example, the preset length is 1 second and the prediction length is 5 seconds, then the first preset quantity is 5, that is, the output information of the first layer encoding includes 5 encoding information.
[0080] When the prediction length is greater than the preset length, the first preset quantity is 1, that is, the output data of the first-layer encoding includes 1 piece of third encoding information. The output data of the first-layer encoding also includes a second preset quantity of first encoding position information, and the second preset quantity is the value obtained by rounding up the ratio between the prediction length and the preset length. The third encoding information can be the average value of the fourth encoding information; the fourth encoding information is the second preset quantity of encoding information output one by one based on the first-layer encoding of the first information target model when the prediction length is greater than the preset length, and there is a one-to-one correspondence between the fourth encoding information and the first encoding position information. The output data of the first-layer encoding also includes the deviation values between the third encoding information and each of the fourth encoding information.
[0081] Exemplarily, taking the sound large model to generate music as an example, the preset length is 1 second, the prediction length is 5 seconds, there are 5 pieces of fourth encoding information, the average value of the 5 pieces of fourth encoding information is obtained to get the third encoding information, and the third encoding information is used as the output data of the first-layer encoding. Further, the 5 pieces of fourth encoding information correspond to 5 pieces of first encoding position information and 5 deviation values, so as to accurately reflect the accuracy of the output data of the first-layer encoding obtained based on the first information when the prediction length is greater than the preset length and the output data of the first-layer encoding includes one piece of third encoding information. That is, the fourth encoding information corresponding to the first encoding position information can be accurately obtained through the first encoding position information, the deviation value corresponding to the first encoding position information, and the third encoding information. In this way, the encoding generation efficiency can be further improved.
[0082] Step B2: If the current-layer encoding is other than the first-layer encoding, the first preset quantity is determined based on the quantity of encoding information in the output data of the previous-layer encoding and the compression rate of the current-layer encoding.
[0083] Specifically, the first preset quantity can be the product of the quantity of encoding information in the output data of the previous-layer encoding and the compression rate of the current-layer encoding, and the first preset quantity is a positive integer.
[0084] In the embodiments of the present invention, by determining whether the current-layer encoding is the first-layer encoding, when the current-layer encoding is the first-layer encoding, the first preset quantity is accurately determined based on the prediction length of the target output information. When the current-layer encoding is other than the first-layer encoding, the product of the quantity of encoding information in the output data of the previous-layer encoding and the compression rate of the current-layer encoding is used as the first preset quantity, so as to accurately determine the quantity of encoding information in the output data of each layer of encoding. Moreover, the different quantities of encoding information of each layer of encoding with different quantities can also reflect whether the encoding information of this encoding layer reflects the overall information or the detailed information of the content generated by the target model.
[0085] Embodiment 4
[0086] Figure 3The figure is a schematic structural diagram of an encoding generation device for a model provided by an embodiment of the present invention. This embodiment is applicable to the encoding generation situation in the process of a large language model generating content in an autoregressive manner. The encoding generation device of the model can be implemented in the form of hardware and / or software, and the encoding generation device of the model can be configured in any electronic device with network communication functions. As Figure 3 shown, the encoding generation device of the model includes:
[0087] A first data determination module 310, configured to determine the input data of the target model encoded in the current layer; the target model is a large language model;
[0088] An encoding module 320, configured to obtain the output data of the current layer encoding based on the input data of the current layer encoding; the output data includes a first preset number of encoding information;
[0089] A second data determination module 330, configured to determine the input data of the next layer encoding based on the input data of the current layer encoding, the output data of the current layer encoding, and the compression rate of the next layer encoding;
[0090] A judgment module 340, configured to until the encoding layer of the target model reaches a preset encoding layer, form the output data of each layer encoding into an encoding data set, and decode the encoding data set to obtain the target output information of the target model.
[0091] Based on the above embodiment, optionally, the first data determination module is configured to: if the current layer encoding is the first layer encoding, the input data of the current layer encoding is the first information obtained in response to an input operation; the input operation is an input operation that triggers the input of the first information to the target model from outside the target model; if the current layer encoding is other layers of encoding except the first layer encoding, the input data of the current layer encoding is determined based on the input data of the previous layer encoding, the output data of the previous layer encoding, and the compression rate of the current layer encoding.
[0092] Based on the above embodiment, optionally, the encoding generation device of the model further includes a first quantity determination module, and the first quantity determination module is configured to: if the current layer encoding is the first layer encoding, the first preset quantity is determined based on the prediction length of the target output information; the prediction length is predicted by the target model based on the first information.
[0093] Based on the above embodiments, optionally, the first quantity determination module is further configured to: when the prediction length is less than or equal to the preset length, the first preset quantity is 1; when the prediction length is greater than the preset length, the first preset quantity is the value obtained by rounding up the ratio between the prediction length and the preset length; or, the first preset quantity is 1, and the output data further includes a second preset quantity of first coding position information, where the second preset quantity is the value obtained by rounding up the ratio between the prediction length and the preset length.
[0094] Based on the above embodiments, optionally, the encoding generation device of the model further includes a second quantity determination module, and the second quantity determination module is configured to: if the current layer encoding is other layer encodings except the first layer encoding, the first preset quantity is determined based on the quantity of encoding information in the output data of the previous layer encoding and the compression rate of the current layer encoding.
[0095] Based on the above embodiments, optionally, the second data determination module is configured to: determine the data content of the input data of the next layer encoding based on the input data of the current layer encoding and the output data of the current layer encoding; determine the quantity of encoding information in the input data of the next layer encoding based on the product of the first preset quantity and the compression rate of the next layer encoding; or, determine the quantity of encoding information in the input data of the next layer encoding based on the product of the first preset quantity and the compression rate of the next layer encoding and the first coding position information of the output data of the current layer encoding; determine the input data of the next layer encoding based on the data content of the input data of the next layer encoding and the quantity of encoding information in the input data of the next layer encoding.
[0096] Based on the above embodiments, optionally, the preset number of encoding layers is determined according to the number of feature extraction layers of the encoder-decoder in the target model.
[0097] Based on the above embodiments, optionally, the compression rate is determined according to at least one of the convolutional kernel size, the stride, the padding quantity, and the encoding cross range between adjacent encoding information.
[0098] Based on the above embodiments, optionally, the input data of the current layer encoding further includes second coding position information of each encoding information, and the encoding module is configured to: perform an inference operation on each encoding information in the input data of the current layer encoding based on the autoregressive mechanism of the target model and the second coding position information of each encoding information to obtain the output data of the current layer encoding.
[0099] The encoding generation device of the model provided by the embodiments of the present invention can execute the encoding generation method of the model provided by any embodiment of the present invention, and has the corresponding functional modules and beneficial effects for executing the method.
[0100] Embodiment V
[0101] According to an embodiment of the present disclosure, the present disclosure also provides an electronic device, a readable storage medium, and a computer program product.
[0102] Figure 4 The structural schematic diagram of an electronic device that can be used to implement the encoding generation method of the model according to an embodiment of the present invention is shown. The electronic device is intended to represent various forms of digital computers, such as, a laptop computer, a desktop computer, a workbench, a personal digital assistant, a server, a blade server, a mainframe computer, and other suitable computers. The electronic device can also represent various forms of mobile devices, such as, a personal digital processor, a cellular phone, a smart phone, a wearable device (such as a helmet, glasses, a watch, etc.) and other similar computing devices. The components shown herein, their connections and relationships, and their functions are only examples and are not intended to limit the implementation of the present invention described and / or claimed herein.
[0103] As Figure 4 shown, the electronic device 10 includes at least one processor 11, and a memory communicatively connected to the at least one processor 11, such as a read-only memory (ROM) 12, a random access memory (RAM) 13, etc. Among them, the memory stores a computer program executable by the at least one processor. The processor 11 can perform various appropriate actions and processes according to the computer program stored in the read-only memory (ROM) 12 or the computer program loaded from the storage unit 18 into the random access memory (RAM) 13. In the RAM 13, various programs and data required for the operation of the electronic device 10 can also be stored. The processor 11, the ROM 12, and the RAM 13 are connected to each other through a bus 14. The input / output (I / O) interface 15 is also connected to the bus 14.
[0104] Multiple components in the electronic device 10 are connected to the I / O interface 15, including: an input unit 16, such as a keyboard, a mouse, etc.; an output unit 17, such as various types of displays, speakers, etc.; a storage unit 18, such as a magnetic disk, an optical disc, etc.; and a communication unit 19, such as a network card, a modem, a wireless communication transceiver, etc. The communication unit 19 allows the electronic device 10 to exchange information / data with other devices through a computer network such as the Internet and / or various telecommunication networks.
[0105] The processor 11 can be various general-purpose and / or special-purpose processing components with processing and computing capabilities. Some examples of the processor 11 include, but are not limited to, a central processing unit (CPU), a graphics processing unit (GPU), various dedicated artificial intelligence (AI) computing chips, various processors running machine learning model algorithms, a digital signal processor (DSP), and any suitable processor, controller, microcontroller, etc. The processor 11 executes the various methods and processes described above, such as the method for encoding and generating a model.
[0106] In some embodiments, the method for encoding and generating a model can be implemented as a computer program, which is tangibly contained in a computer-readable storage medium, such as the storage unit 18. In some embodiments, part or all of the computer program can be loaded and / or installed onto the electronic device 10 via the ROM 12 and / or the communication unit 19. When the computer program is loaded into the RAM 13 and executed by the processor 11, one or more steps of the method for encoding and generating a model described above can be executed. Alternatively, in other embodiments, the processor 11 can be configured to execute the method for encoding and generating a model in any other suitable manner (e.g., by means of firmware).
[0107] Various embodiments of the systems and techniques described above herein can be implemented in digital electronic circuit systems, integrated circuit systems, field-programmable gate arrays (FPGA), application-specific integrated circuits (ASIC), application-specific standard products (ASSP), system-on-a-chip (SOC), complex programmable logic devices (CPLD), computer hardware, firmware, software, and / or combinations thereof. These various embodiments can include: being implemented in one or more computer programs, which can be executed and / or interpreted on a programmable system including at least one programmable processor, the programmable processor can be a special or general programmable processor, and can receive data and instructions from a storage system, at least one input device, and at least one output device, and transmit the data and instructions to the storage system, the at least one input device, and the at least one output device.
[0108] The computer program for implementing the method of the present invention can be written in any combination of one or more programming languages. These computer programs can be provided to the processor of a general-purpose computer, a special-purpose computer, or other programmable data processing devices, such that when the computer programs are executed by the processor, the functions / operations specified in the flowchart and / or block diagram are implemented. The computer programs can be executed entirely on the machine, partially on the machine, as an independent software package partially on the machine and partially on a remote machine, or entirely on a remote machine or server.
[0109] In the context of the present invention, a computer-readable storage medium can be a tangible medium that can contain or store a computer program for use by or in connection with an instruction execution system, apparatus, or device. The computer-readable storage medium can include, but is not limited to, electronic, magnetic, optical, electromagnetic, infrared, or semiconductor systems, apparatus, or devices, or any suitable combination of the foregoing. Alternatively, the computer-readable storage medium can be a machine-readable signal medium. More specific examples of the machine-readable storage medium would include an electrical connection based on one or more wires, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or Flash memory), an optical fiber, a portable compact disc read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the foregoing.
[0110] In order to provide interaction with a user, the systems and techniques described herein can be implemented on an electronic device having: a display device (e.g., a CRT (cathode ray tube) or LCD (liquid crystal display) monitor) for displaying information to the user; and a keyboard and a pointing device (e.g., a mouse or a trackball) by which the user can provide input to the electronic device. Other kinds of devices can also be used to provide interaction with the user; for example, the feedback provided to the user can be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and input from the user can be received in any form (including acoustic input, speech input, or tactile input).
[0111] The systems and techniques described herein can be implemented in a computing system that includes backend components (e.g., as a data server), or a computing system that includes middleware components (e.g., an application server), or a computing system that includes frontend components (e.g., a user computer having a graphical user interface or a web browser through which the user can interact with an implementation of the systems and techniques described herein), or a computing system that includes any combination of such backend components, middleware components, or frontend components. The components of the system can be interconnected by any form or medium of digital data communication (e.g., a communication network). Examples of communication networks include: local area network (LAN), wide area network (WAN), blockchain network, and the Internet.
[0112] A computing system may include a client and a server. The client and the server are generally far from each other and usually interact via a communication network. The client-server relationship is created by computer programs running on respective computers and having a client-server relationship with each other. The server can be a cloud server, also known as a cloud computing server or a cloud host, which is a host product in the cloud computing service system, solving the defects of difficult management and weak business scalability existing in traditional physical hosts and VPS services.
[0113] It should be understood that various forms of processes shown above can be used, with steps reordered, added, or deleted. For example, the steps described in the present invention can be executed in parallel, sequentially, or in a different order, as long as the desired results of the technical solution of the present invention can be achieved, and no limitation is made herein.
[0114] The above specific embodiments do not constitute a limitation to the protection scope of the present invention. Those skilled in the art should understand that various modifications, combinations, sub-combinations, and substitutions can be made according to design requirements and other factors. Any modifications, equivalent substitutions, and improvements made within the spirit and principle of the present invention shall be included within the protection scope of the present invention.
Claims
1. A method for encoding and generating a model, characterized in that, The method includes: Determine the input data encoded by the target model at the current layer; the target model is a large language model; Based on the input data encoded by the current layer, obtain the output data encoded by the current layer; the output data includes a first preset number of encoded information; Based on the input data encoded by the current layer, the output data encoded by the current layer, and the compression rate of the next layer encoding, determine the input data of the next layer encoding; Until the number of encoding layers of the target model reaches the preset number of encoding layers, the output data of each layer of encoding is formed into an encoded data set, and the target output information of the target model is obtained by decoding the encoded data set.
2. The method according to claim 1, wherein Determining the input data encoded by the target model at the current layer includes: If the current layer encoding is the first layer encoding, the input data of the current layer encoding is the first information obtained in response to the input operation; the input operation is an input operation that triggers the input of the first information to the target model from outside the target model; If the current layer encoding is other layer encodings except the first layer encoding, the input data of the current layer encoding is determined based on the input data of the previous layer encoding, the output data of the previous layer encoding, and the compression rate of the current layer encoding.
3. The method according to claim 2, wherein The method further includes: If the current layer encoding is the first layer encoding, the first preset number is determined based on the prediction length of the target output information; the prediction length is predicted by the target model based on the first information.
4. The method according to claim 3, characterized in that, If the current layer encoding is the first layer encoding, determining the first preset number based on the prediction length of the target output information includes: When the prediction length is less than or equal to the preset length, the first preset number is 1; When the prediction length is greater than the preset length, the first preset number is the value obtained by rounding up the ratio between the prediction length and the preset length; or, the first preset number is 1, and the output data further includes a second preset number of first encoding position information, and the second preset number is the value obtained by rounding up the ratio between the prediction length and the preset length.
5. The method according to claim 2, characterized in that, The method further includes: If the current layer encoding is other layer encodings except the first layer encoding, the first preset number is determined based on the number of encoded information in the output data of the previous layer encoding and the compression rate of the current layer encoding.
6. The method according to claim 1, wherein Based on the input data encoded by the current layer, the output data encoded by the current layer, and the compression rate of the next layer encoding, determining the input data of the next layer encoding includes: Based on the input data encoded by the current layer and the output data encoded by the current layer, determine the data content of the input data of the next layer encoding; Determine the number of encoded information in the input data of the next layer encoding based on the product of the first preset number and the compression rate of the next layer encoding; or, determine the number of encoded information in the input data of the next layer encoding based on the product of the first preset number and the compression rate of the next layer encoding and the first encoding position information of the output data of the current layer encoding; Determine the input data of the encoding in the next layer based on the data content of the input data of the encoding in the next layer and the number of encoding information in the input data of the encoding in the next layer.
7. The method according to claim 1, characterized in that The preset number of encoding layers is determined according to the number of feature extraction layers of the encoder-decoder in the target model.
8. The method according to claim 1, wherein The compression ratio is determined according to at least one of the convolutional kernel size, the stride, the padding number, and the encoding cross range between adjacent encoding information.
9. The method according to claim 1, characterized in that, The input data of the encoding in the current layer further includes the second encoding position information of each encoding information. Based on the input data of the encoding in the current layer, obtaining the output data of the encoding in the current layer includes: Based on the autoregressive mechanism of the target model and the second encoding position information of each encoding information, perform an inference operation on each encoding information in the input data of the encoding in the current layer to obtain the output data of the encoding in the current layer.
10. An encoding generation device for a model, characterized in that The device includes: A first data determination module, configured to determine the input data of the encoding in the current layer of the target model; the target model is a large language model; An encoding module, configured to obtain the output data of the encoding in the current layer based on the input data of the encoding in the current layer; the output data includes a first preset number of encoding information; A second data determination module, configured to determine the input data of the encoding in the next layer based on the input data of the encoding in the current layer, the output data of the encoding in the current layer, and the compression ratio of the encoding in the next layer; A judgment module, configured to until the number of encoding layers of the target model reaches the preset number of encoding layers, then form the output data of each layer of encoding into an encoding data set to decode the encoding data set to obtain the target output information of the target model.
Citation Information
Patent Citations
Data coding method and related equipment
CN114978189A
Data processing method and related equipment
CN116737895A
Layered hybrid structured compression method and device for large language model
CN120087419A
Training method and apparatus and recognition method and apparatus for text error correction model, and computer device
WO2022121178A1