Data processing method and related equipment
By selecting the target model that matches the input data attributes for pre-fill processing, and using the neural network model with the same model parameters for decoding, the pre-fill and decoding stages of the large language model are optimized, and the inefficiency problem in the existing technology is solved and faster inference time is achieved.
Patent Information
- Application Number
- CN202510316146.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-17
- Publication Date
- 2025-07-18
AI Technical Summary
In the inference process, the pre-filling and decoding stages of existing large language models are less efficient, resulting in an extended overall inference time.
The process of the pre-filling and decoding stages is optimized by selecting the target model that matches the input data attributes from multiple optional models for pre-filling processing, and using the neural network model with the same model parameters for decoding.
The processing efficiency of the pre-filling stage of the large language model and the processing time of the decoding stage are accelerated, and the overall inference efficiency is improved.
Smart Images

Figure CN120338095A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the technical field of data processing, and in particular, to a data processing method, a data processing device, and an electronic device. Background Art
[0002] A large language model (LLM) is a model used in the field of artificial intelligence and is usually used in inference scenarios. The inference process or stage of the large language model mainly includes a prefill stage and a decoding stage. Among them, the prefill stage is used to receive the content to be inferred input by the user. The decoding stage is used to infer the output expected by the user based on the content input by the user. It can be seen that how to improve the inference efficiency has become a technical problem to be solved urgently. Summary of the Invention
[0003] This application provides a data processing method, a data processing device, and an electronic device.
[0004] According to a first aspect of this application, a data processing method is provided, including:
[0005] Obtain first data, where the first data is the input data of the large language model;
[0006] Determine a first target model that matches the attributes of the first data from multiple first optional models, where each first optional model is obtained by compiling a program based on a different prefill length, and the first optional models have the same model parameters;
[0007] Use the first target model to perform prefill processing of the large language model on the first data to obtain a prefill result of the first data, where the prefill result includes the result of key-value caching of the input tokens of the first data and generating the first output token of the first data. Among them, the first output token generated by using the first target model is earlier than the first output token generated by using other first optional models except the first target model;
[0008] Based on the key-value caching result and the first output token, use a second target model to perform decoding processing of the large language model on the first data to obtain second data for replying to the first data.
[0009] In an implementable manner, the first data includes N first sub-data, where N is a positive integer greater than or equal to 2, and the method further includes:
[0010] Based on the attributes of each first sub-data, determine a first target model for each first sub-data from multiple first optional models, so as to perform prefill processing on each first sub-data by using each first target model;
[0011] Based on the priorities of the first sub - data, use the second target model to decode the pre - filling results of the first sub - data to obtain second sub - data, where the second sub - data is used to reply to the first sub - data, and the second data includes the second sub - data.
[0012] In an implementable manner, determining the first target model that matches the attribute of the first data from multiple first optional models includes:
[0013] Obtain the context information of the first data;
[0014] Obtain the relevance between the first data and the context information;
[0015] In response to the relevance satisfying a preset relevance condition, based on the attribute of the first data and the attribute of the context information, determine the first target model for the first data from multiple first optional models.
[0016] In an implementable manner, the method further includes:
[0017] Based on at least one of the attribute of the first data, the attribute of the first target model, and the pre - filling result of the first data, determine the second target model from multiple second optional models, where each second optional model is obtained by program compilation based at least on the attribute of the model.
[0018] In an implementable manner, the first data includes N first sub - data, where N is a positive integer greater than or equal to 2, and the method further includes:
[0019] Based on at least one of the attributes of the first sub - data, the attributes of the first target models determined for the first sub - data, and the pre - filling results of the first sub - data, determine the second target models for the first sub - data from multiple second optional models;
[0020] Use the second target models to decode the pre - filling results of the first sub - data to obtain second sub - data, where the second sub - data is used to reply to the first sub - data, and the second data includes the second sub - data.
[0021] In an implementable manner, the same model parameters are saved or stored in a parameter storage space, and the first target model is a neural network model that loads the same model parameters from the parameter storage space.
[0022] In an implementable manner, the first target model and the second target model are neural network models with the same model parameters;
[0023] The first target model and the second target model are neural network models that respectively load the same model parameters from the parameter storage space.
[0024] In one implementable manner, determining a first target model that matches the attributes of the first data from multiple first alternative models includes:
[0025] Obtain the attributes of the first data;
[0026] Obtain the target attributes of each first alternative model, where the target attributes are related to the attributes of the first data; determine, from the multiple first alternative models, the first alternative model whose target attributes match the attributes of the first data as the first target model.
[0027] According to a second aspect of the present application, there is provided a data processing device, including:
[0028] A first acquisition unit for acquiring first data, where the first data is input data of a large language model;
[0029] A first determination unit for determining a first target model that matches the attributes of the first data from multiple first alternative models, where each first alternative model is obtained by program compilation based on different pre-filled lengths, and the first alternative models have the same model parameters;
[0030] A first processing unit for performing pre-filling processing of the large language model on the first data by using the first target model to obtain a pre-filling result of the first data, where the pre-filling result includes a result of key-value caching of the input tokens of the first data and generating a first initial output token of the first data, and the first initial output token generated by using the first target model is earlier than the first initial output tokens generated by using other first alternative models except the first target model;
[0031] A second processing unit for performing decoding processing of the large language model on the first data by using a second target model based on the key-value caching result and the first initial output token to obtain second data for replying to the first data.
[0032] According to a third aspect of the present application, there is provided an electronic device, including:
[0033] At least one processor; and
[0034] A memory communicatively connected to the at least one processor; where
[0035] The memory stores instructions executable by the at least one processor, and when the instructions are executed by the at least one processor, the at least one processor is enabled to execute the method described in the present application.
[0036] According to a fourth aspect of the present application, there is provided a non-transitory computer-readable storage medium storing computer instructions for causing the computer to execute the method described in the present application.
[0037] According to a fifth aspect of the present application, there is provided a computer program product including a computer program or instructions, which when executed by a processor implement the method described in the present application.
[0038] In the present application, the pre-filling process of the large language model is implemented by using the determined first target model that matches the attributes of the first data, which can improve the processing efficiency of the pre-filling stage of the large language model and shorten the processing time of the pre-filling stage. The decoding process of the pre-filling result of the first data by using the second target model, due to the accelerated output of the pre-filling result in the pre-filling stage, also accelerates the decoding stage of the large language model, and can effectively improve the inference efficiency.
[0039] It should be understood that the content described in this part is not intended to identify the key or important features of the embodiments of the present application, nor is it used to limit the scope of the present application. Other features of the present application will become easily understood through the following description. BRIEF DESCRIPTION OF THE DRAWINGS
[0040] By referring to the drawings and reading the following detailed description, the above and other purposes, features and advantages of the exemplary embodiments of the present application will become easily understood. In the drawings, several embodiments of the present application are shown in an exemplary rather than restrictive manner, where:
[0041] In the drawings, the same or corresponding reference numerals represent the same or corresponding parts.
[0042] Figure 1 Shows the implementation process schematic of the data processing method in the embodiment of the present application Figure 1 ;
[0043] Figure 2 Shows the implementation process schematic of the data processing method in the embodiment of the present application Figure 2 ;
[0044] Figure 3 Shows the implementation block diagram of obtaining the first optional model and the second optional model in the embodiment of the present application;
[0045] Figure 4 Shows the schematic diagram of the computational flow difference between the pre-filling stage and the decoding stage in the embodiment of the present application;
[0046] Figure 5 Shows the schematic of the multi-core heterogeneous system in the embodiment of the present application Figure 1 ;
[0047] Figure 6 Shows a schematic diagram of a multi-core heterogeneous system in an embodiment of the present application Figure 2 ;
[0048] Figure 7 Shows a schematic diagram of a multi-core heterogeneous system in an embodiment of the present application Figure 3 ;
[0049] Figure 8 Shows a block diagram of the implementation of a data processing method in an embodiment of the present application;
[0050] Figure 9 Shows a schematic diagram of the composition structure of a data processing device in an embodiment of the present application;
[0051] Figure 10 Shows a schematic diagram of the composition structure of an electronic device in an embodiment of the present application. Detailed implementation manners
[0052] To make the objectives, features, and advantages of the present application more obvious and understandable, the technical solutions in the embodiments of the present application will be clearly and completely described below with reference to the accompanying drawings in the embodiments of the present application. Obviously, the described embodiments are only a part of the embodiments of the present application, rather than all the embodiments. Based on the embodiments in the present application, all other embodiments obtained by those skilled in the art without creative efforts belong to the scope of protection of the present application.
[0053] To make the objectives, technical solutions, and advantages of the present application clearer, the present application will be further described in detail below with reference to the accompanying drawings. The described embodiments should not be regarded as limitations on the present application. All other embodiments obtained by those of ordinary skill in the art without creative efforts belong to the scope of protection of the present application.
[0054] In the following description, "some embodiments" are involved, which describe a subset of all possible embodiments. However, it can be understood that "some embodiments" can be the same subset or different subsets of all possible embodiments, and can be combined with each other without conflict.
[0055] In the following description, the terms "first / second" involved are only used to distinguish similar objects and do not represent a specific order for the objects. It can be understood that "first / second" can be interchanged with a specific order or sequence when allowed, so that the embodiments of the present application described herein can be implemented in an order other than that illustrated or described herein.
[0056] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by those of ordinary skill in the technical field to which this application belongs. The terms used herein are only for the purpose of describing the embodiments of the present application and are not intended to limit the present application.
[0057] It should be understood that in various embodiments of the present application, the magnitude of the sequence numbers of each implementation process does not imply the sequence of execution order. The execution order of each process should be determined by its function and internal logic, and should not constitute any limitation to the implementation process of the embodiments of the present application.
[0058] In the technical solution of the present application, a first target model matching the attributes of the first data can be determined from multiple first optional models, and the pre-filling process of the large language model can be implemented by using the determined first target model matching the attributes of the first data, which can accelerate the processing efficiency of the pre-filling stage of the large language model and shorten the processing time of the pre-filling stage. By using the second target model to decode the pre-filling result of the first data, since the pre-filling result is output faster in the pre-filling stage, the decoding stage of the large language model is also accelerated. The acceleration of the pre-filling stage (pre-filling process) and the decoding stage (decoding process) can effectively shorten the inference time and improve the inference efficiency of the large language model.
[0059] The processing logic of the data processing method in the present application can be deployed in any reasonable device. For example, it can be deployed in terminal devices such as in-vehicle devices, mobile phones, computers, intelligent wearable devices, etc. Also, for example, it can be deployed in server devices such as cloud servers and ordinary servers. Since the processing logic of the data processing method in the present application involves the pre-filling process and decoding process logic of the large language model, the data processing logic of the present application can be deployed in any device that can apply the large language model.
[0060] Figure 1 Shows the implementation process schematic of the data processing method in the embodiments of the present application Figure 1 As Figure 1 shown, the method includes:
[0061] S101: Obtain first data, where the first data is the input data of the large language model.
[0062] In this step, the first data is any data that can be used for inference by the large language model (data to be inferred). The first data can be voice data or text data. In terms of content, it can be any data content that requires the large language model to perform inference. For example, "I'm going to place A, what's the driving route?", "Who are you?", "What's the weather like in the next three days?". It can be understood that the data input to the large language model is usually text data. If the first data is voice data, it is converted into text data. The text data is converted into a token sequence.
[0063] S102: Determine a first target model that matches the attributes of the first data from multiple first optional models, where each first optional model is obtained by compiling a program based on a different pre-filled length, and the first optional models have the same model parameters.
[0064] In this application, multiple first optional models that can be used to perform pre-filling processing are provided. The first optional model can be a neural network model for performing pre-filling processing. The first optional models can be neural network models with the same model parameters such as weight parameters. Since each first optional model is obtained by compiling a program based on a different pre-filled length, different first optional models correspond to different data lengths. Among them, the pre-filled length can be the length of the content that the user needs the large language model to reason about, such as the length of the content that needs to be reasoned about in the previous "I'm going to place A. What's the driving route?", "Who are you?", "What's the weather like in the next three days?", etc. The pre-filled length can also be the token length of the content that needs to be reasoned about, that is, the length of the token sequence converted from the above exemplary reasoning data.
[0065] Exemplarily, the pre-filled lengths are preset to be 2, 4, 8, 16, 32, etc., which are the lengths of 2 to the Nth power of characters, where N is a positive integer greater than or equal to 1. Using the data to be reasoned about with the length of 2 to the Nth power of the reasoning content, the existing preprocessing model is compiled to obtain (first) optional models corresponding to different data lengths. Among them, the pre-filled length can take other values, and no specific limitation is made.
[0066] It can be understood that the data that needs to be reasoned about by the large language model can be a question. The large language model performs reasoning through pre-filling processing and decoding processing to give a reply to the question. In practical applications, the questions can be various, and the lengths of the questions can be various. If the model that performs pre-filling processing in the large language model is to adapt to questions of all lengths, it is necessary for this model to be able to perform pre-filling processing on the longest questions. This model needs to perform a large amount of calculations to meet the pre-filling processing of long questions. Such a model still needs to perform a large amount of operations to obtain the pre-filled result when encountering short questions, and performing a large amount of operations on short questions is undoubtedly an ineffective calculation and wastes computing resources.
[0067] In this application, each first optional model obtained by compiling a program based on a different pre-filled length is a model that can perform pre-filling processing compiled for questions of different lengths. For example, for a pre-filled length of 2 4The (first) optional model 1 obtained by compiling the program with [[ID=]] = 16 can perform efficient pre-padding processing on problems with a length less than or equal to 16, without a large number of invalid operations. For a pre-padding length of 2 5 The (first) optional model 2 obtained by compiling the program with [[ID=]] = 32 can perform efficient pre-padding processing on problems with a length greater than or equal to 17 and less than or equal to 32, without a large number of invalid operations. For a pre-padding length of 2 6 The (first) optional model 3 obtained by compiling the program with [[ID=]] = 64 can perform efficient pre-padding processing on problems with a length greater than or equal to 33 and less than or equal to 64, without a large number of invalid operations.
[0068] In practical applications, each problem has a theme. For example, problem 1 "I am going to place A. What is the driving route?" is a route planning theme, and problem 2 "What is the weather condition in the next three days?" is a weather theme. For the same theme, the questions asked by users vary in length. Based on this, in this application, multiple first optional models for each theme are obtained by compiling the program based on different pre-padding lengths for each theme. For example, for the route planning theme, multiple first optional models for this theme are obtained by compiling the program based on different pre-padding lengths.
[0069] For the weather theme, multiple first optional models for this theme are obtained by compiling the program based on different pre-padding lengths. Taking the route planning theme as an example, for the (first) optional model 1 obtained by compiling the program with [[ID=]] = 16 4 It can perform efficient pre-padding processing on route planning problems with a length less than or equal to 16. The adaptation length of the optional model 1 is L1, and L1 is less than or equal to 16. For the (first) optional model 2 obtained by compiling the program with [[ID=]] = 32 5 It can perform efficient pre-padding processing on route planning problems with a length less than or equal to 32 and greater than or equal to 17. The adaptation length of the optional model 2 is L2, and L2 is less than or equal to 32 and greater than or equal to 17.
[0070] During implementation, obtain the attribute of the first data; obtain the target attribute of each first optional model, where the target attribute is related to the attribute of the first data; from the multiple first optional models, determine the first optional model whose target attribute matches the attribute of the first data as the first target model.
[0071] The attribute of the first data can be the length of the first data or the length of the first data converted into a token sequence. The target attribute can be the length of the problem adapted by the first optional model. Taking the attribute of the first data being the length of the first data as an example, in the case where a user generates a problem (the first data includes the problem), a first optional model that matches the length of the problem can be found from multiple first optional models as the first target model of the problem. Among the multiple first optional models, the model whose adapted problem length satisfies the following conditions: greater than the length of the problem and the difference between the adapted length of the model and the length of the problem is the smallest, is regarded as the first target model that matches the length of the problem.
[0072] Considering that the first data can be a problem, therefore, the attribute of the first data can be the length and theme of the problem, or the length and theme of the first data converted into a token sequence. The target attribute can be the length and theme of the problem adapted by the first optional model. Taking the attribute of the first data being the length and theme of the problem as an example,
[0073] In the case where a user generates a problem (the first data includes the problem), identify the length and theme of the problem. A model with the same theme as the problem can be found from the first optional models for each theme. From the found models, a model that matches the length of the problem is found and used as the first target model with matching attributes. Or, from the multiple first optional models for each theme, a model that matches the length of the problem is found, and then from the found models, a model with the same theme as the problem is found and used as the first target model with matching attributes.
[0074] S103: Use the first target model to perform pre-filling processing of the large language model on the first data to obtain a pre-filling result for the first data. The pre-filling result includes the result of key-value caching of the input tokens of the first data and generating the first character output token of the first data. Among them, the first character output token generated using the first target model is earlier than the first character output token generated using other first optional models except the first target model.
[0075] In this step, the token of the first data (input token) can be input into the first target model to enable the first target model to perform pre-filling processing on the input token. Among them, the pre-filling processing of the large language model includes: generating the key matrix (Key matrix) and value matrix (Value matrix) of each token in the input token, and caching the generated Key matrix and Value matrix in the form of KV cache; and generating the first output token through the Key matrix and Value matrix (KV matrix). The first output token can be the token of the first character in the answer obtained by the large language model through reasoning to reply to the user's question.
[0076] It can be understood that since the first target model is the model among many first alternative models that matches the attributes of the first data in terms of target attributes, therefore, for a question, using a model that is adapted to it in terms of target attributes for pre-filling processing will not generate a large number of invalid operations, thus accelerating the processing flow in the pre-filling stage and shortening the time to generate the first output token. The first output token generated by this model is earlier than the first output token generated by other first alternative models except the first target model. For the pre-filling processing process and decoding processing process of the large language model, please refer to the relevant descriptions and will not be elaborated.
[0077] S104: Based on the key-value cache result and the first output token, use the second target model to perform decoding processing of the large language model on the first data to obtain the second data for replying to the first data.
[0078] In this step, the cached Key matrix and Value matrix are used as the key-value cache result, and the key-value cache result and the first output token are input into the second target model. The second target model realizes the reasoning of the question based on the key-value cache result and the first output token to obtain the second data for replying to the first data. If the first data is a question, the second data can be the reply to this question given by the large language model through reasoning.
[0079] In S101 - S104, using the determined first target model that matches the attributes of the first data to perform pre-filling processing of the large language model can improve the processing efficiency of the pre-filling stage of the large language model and shorten the processing time of the pre-filling stage. Using the second target model to perform decoding processing on the pre-filling result of the first data, due to the accelerated output of the pre-filling result in the pre-filling stage, the decoding stage of the large language model is also accelerated. The acceleration of the pre-filling stage (pre-filling processing) and the decoding stage (decoding processing) can effectively shorten the inference time and improve the inference efficiency of the large language model.
[0080] In some embodiments, the first data may be a single text data for asking a question. Such as the aforementioned "Who are you?" or "I am going to place A. What is the driving route?" Thus, according to the length of the question, or the length of the question and the theme, a model that is adapted to its length, or a model with the same length and theme, can be selected from multiple first optional models as the first target model that matches the attribute of the first data for this question.
[0081] In some embodiments, the first data includes N first sub-data, where N is a positive integer greater than or equal to 2. As Figure 2 shown, the data processing method in the present application further includes:
[0082] S201: Based on the attributes of each first sub-data, determine a first target model for each first sub-data from multiple first optional models, so as to perform pre-filling processing on each first sub-data by using each first target model.
[0083] S202: Based on the priorities of each first sub-data, use the second target model to decode the pre-filled results of each first sub-data to obtain each second sub-data, where each second sub-data is used to reply to each first sub-data, and the second data includes each second sub-data.
[0084] In the solutions shown in S201 to S202 above, the first data may be multiple text data for asking multiple questions, such as each text data asking a question. For example, three text data sequentially ask three questions: "Who are you?" (Question 1), "Please plan a route. I am going to place A." (Question 2), "After arriving at place A today, I will have a meeting, and after the meeting, I will have dinner with the client. According to my itinerary today, what is the earliest possible time to return tomorrow?" (Question 3), etc. Each question can be regarded as a first sub-data. According to the length of each question, or the length and theme of each question, a model that is adapted to the length of the question, or a model that is adapted to the length of the question and has the same theme, is determined from multiple first optional models as the first target model determined for each question.
[0085] Exemplarily, for questions 1 to 3, the first target model (such as model 1) determined for question 1 is used to perform pre-filling processing on question 1 to obtain a pre-filling result, that is, the pre-filling result for question 1. The key-value cache result and the first-word output token obtained by model 1 for question 1 are input into the second target model, and the second target model is used to perform decoding processing to obtain reply data 1 for replying to question 1. The first target model (such as model 2) determined for question 2 is used to perform pre-filling processing on question 2 to obtain a pre-filling result, that is, the pre-filling result for question 2. The key-value cache result and the first-word output token obtained by model 2 for question 2 are input into the second target model, and the second target model is used to perform decoding processing to obtain reply data 2 for replying to question 2. The first target model (such as model 3) determined for question 3 is used to perform pre-filling processing on question 3 to obtain a pre-filling result, that is, the pre-filling result for question 3. The key-value cache result and the first-word output token obtained by model 3 for question 3 are input into the second target model, and the second target model is used to perform decoding processing to obtain reply data 3 for replying to question 3.
[0086] If the second target model for decoding questions 1 to 3 is the same model, then according to the generation order of questions 1, 2, and 3, the second target model sequentially performs decoding processing on the pre-filling results of each question 1. It is also possible to, according to the sequence of the pre-filling results generated by the first target model for performing pre-filling processing on each question, the second target model sequentially performs decoding processing on the pre-filling results of each question 1. Among them, the generation order of each question and / or the sequence of the pre-filling results generated by each question can be used as the priority of each first sub-data. The second target model sequentially performs decoding processing on the pre-filling results of each question, and the obtained decoding results are the second sub-data for replying to each question. For example, reply data 1 is the second sub-data for replying to question 1, reply data 2 is the second sub-data for replying to question 2, and reply data 3 is the second sub-data for replying to question 3. From the perspective of the user side, each reply data is used to reply to the corresponding question and is output below the corresponding question. For example, reply data 1 is output below question 1, and reply data 2 is output below question 2, giving a response in a question-and-answer manner.
[0087] Judging from the lengths of the previous questions 1 to 3, the three questions have different lengths. Determining models adapted to the lengths as their respective first target models for the three questions of different lengths can achieve the targeted determination or selection of the first target model, speed up the processing of the pre-filling stage, speed up the output of the first-word output token and the caching of the KV matrix, and improve the inference efficiency.
[0088] In some embodiments, the solution for determining a first target model that matches the attributes of the first data from multiple first optional models may be as follows: obtain the context information of the first data; obtain the relevance between the first data and the context information; in response to the relevance satisfying a preset relevance condition, based on the attributes of the first data and the attributes of the context information, determine a first target model for the first data from multiple first optional models. In this application, by combining the first data and its context information, the determination or selection of the first target model is achieved. This determination or selection method can improve the processing efficiency of the pre-filling stage while ensuring the accuracy of the inference result.
[0089] Taking the first data as the question "In that case, what time can I return at the earliest tomorrow?" (Question 1) as an example, the previous question of Question 1 is Question 2: "I need to go to Place A this morning. After arriving at Place A, I need to have a meeting, and after the meeting, I need to have dinner with a client. Please plan the itinerary and the route to Place A." Question 2 can be used as the context information of Question 1. For the user input Question 1, read the context information of this question, such as Question 2. Identify the relevance between Question 1 and Question 2. Considering that Question 2 is a question given based on the itinerary of Question 1, there is a relevance between Question 1 and Question 2, satisfying the preset relevance condition that the two are related. Taking the attribute as length, for example, the lengths of Question 1 and Question 2 can be added to obtain the total length, and from multiple first optional models, select the model whose adapted length is greater than the total length and the difference between the model adapted length and the total length is the smallest as the first target model determined for Question 1.
[0090] In a possible way, it is also possible to calculate the relevance strength between Question 1 and Question 2 when there is a relevance between Question 1 and Question 2. If the relevance strength is higher than a certain relevance threshold, it is considered to satisfy the preset relevance condition.
[0091] In practical applications, the context information of Question 1 may also include other previous questions in addition to Question 2. If Question 2 and other previous questions of Question 2 are regarded as the historical questions of Question 1, and Question 1 and its historical questions satisfy the preset relevance condition, then in this application, for the first data, the attributes of the historical questions of the first data can also be considered, and the attributes of the first data and the attributes of the historical questions of the first data can be combined to achieve the accurate selection or determination of the first target model. The pre-filling process of the accurately selected or determined first target model can not only accelerate the calculation of the first character output token and the KV matrix of this question, accelerating the output of the pre-filling process. Moreover, the pre-filling result given in combination with the historical questions is more accurate, improving the correctness of the pre-filling process.
[0092] In this application, for each problem, the same second target model can be used for the decoding process in the decoding stage. Alternatively, different second target models can be used for the decoding process for each problem. During implementation, one can: determine the second target model from multiple second optional models based on at least one of the attributes of the first data, the attributes of the first target model, and the pre-population result of the first data, where each second optional model is obtained by compiling a program based at least on the attributes of the model.
[0093] When the attribute of the first data is the length of the first data or the attribute of the first target model is the adaptation length of the first target model to the first data, the attribute of the model can be the length. A second optional model is obtained by compiling a program for a neural network model that decodes the first data with a length of L1. A second optional model is obtained by compiling a program for a neural network model that decodes the first data with a length of L2. And so on, each second optional model is obtained by compiling a program for each neural network model that decodes the first data with each length. When the attribute of the first data is the length of the first data and the topic of the first data, the attribute of the model can be the length and the topic. A second optional model is obtained by compiling a program for a neural network model that decodes the first data with a length of L1 and a topic of topic 1. A second optional model is obtained by compiling a program for a neural network model that decodes the first data with a length of L2 and a topic of topic 2. And so on, each second optional model is obtained by compiling a program for each neural network model that decodes the first data with each length and each topic. The adaptation lengths of the second optional models can be L1, L2, etc. In addition, considering that the pre-population result of the first data is the input data of the model used in the decoding stage, each neural network model can be compiled according to the data volume of various possible values of the input data to obtain each second optional model.
[0094] Based on this, during implementation, a model that adapts to the length of the first data or the adaptation length of the first target model can be determined or selected from multiple second optional models as the second target model. Herein, adapting to the length of the first data or the adaptation length of the first target model can be considered as being greater than the length of the first data (or the adaptation length of the first target model), and among all the second optional models, the difference between the adaptation length of the model and the length of the first data (or the adaptation length of the first target model) is the smallest. Alternatively, based on the data volume of the pre-population result of the first data, a model that adapts to the data volume of the pre-population result and has the smallest difference between the adaptation data volume of the second optional model and the data volume of the pre-population result can be selected from multiple second optional models as the second target model. Herein, adapting to the data volume of the pre-population result can be considered as having a data volume greater than the data volume of the pre-population result.
[0095] It can be seen that in this application, based on at least one of the attributes of the first data, the attributes of the first target model, and the pre-population result of the first data, targeted determination or selection of the second target model can be achieved. The targeted second target model can accelerate the decoding process and shorten the decoding time, thereby improving the inference efficiency.
[0096] In the solution where the second target model of this application can achieve targeted determination or selection, if the first data includes N first sub-data, this application further includes the following solution: based on at least one of the attributes of each first sub-data, the attributes of each first target model determined for the first sub-data, and the pre-population result of each first sub-data, determine each second target model for each first sub-data from multiple second optional models; use each second target model to decode the pre-population result of each first sub-data to obtain each second sub-data, and the second sub-data is used to reply to each first sub-data, where the second data includes each second sub-data.
[0097] Exemplarily, for question 1 (“Who are you”), question 2 (“Please plan a route as I am going to place A”), and question 3 (“After arriving at place A today, I will have a meeting and then have dinner with a client. According to my schedule today, what is the earliest possible time to return tomorrow”), among multiple second optional models, if model A is a model that is adapted to the first target model of question 1 in terms of adaptation length, then model A is taken as the second target model determined for question 1. If model B is a model that is adapted to the first target model of question 2 in terms of adaptation length, then model B is taken as the second target model determined for question 2. If model C is a model that is adapted to the first target model of question 3 in terms of adaptation length, then model C is taken as the second target model determined for question 3. The key-value cache result and the first-character output token obtained for question 1 are input into model A, and model A is used for decoding processing to obtain reply data 1 for replying to question 1. The key-value cache result and the first-character output token obtained for question 2 are input into model B, and model B is used for decoding processing to obtain reply data 2 for replying to question 2. The key-value cache result and the first-character output token obtained for question 3 are input into model C, and model C is used for decoding processing to obtain reply data 3 for replying to question 3.
[0098] As can be seen, in this application, not only can a targeted selection of the model (first target model) used in the pre-filling processing stage be achieved, but also a targeted selection of the model (second target model) used in the decoding processing node can be achieved. Among them, the targeted selection of the first target model can accelerate the pre-filling processing flow, and the targeted selection of the second target model can accelerate the decoding processing flow. The acceleration of the pre-filling processing flow and the decoding processing flow can effectively improve the inference efficiency.
[0099] In this application, the first alternative models have the same model parameters. Each of the first alternative models can be a neural network model, and its model parameters can be weight parameters. For convenient storage, the same model parameters used among the first alternative models can be saved or stored in a specified parameter storage space. Further, because the model parameters among the first alternative models are the same, only one copy of the model parameters needs to be stored in the parameter storage space, without the need to store multiple copies, so as to reduce the storage waste of repeated storage on space. When a first alternative model that matches the attributes of the first data is determined, the model parameters can be read from the storage space, and the first target model can be obtained by loading the model parameters into the first alternative model. That is, the first target model can be a neural network model loaded with the same model parameters from the parameter storage space. In this application, the solution for obtaining the first target model by loading the same model parameters from the parameter storage space can effectively avoid the problem of large storage occupancy caused by repeated storage of the same model parameters and save storage space. Regardless of which first alternative model is determined as the first target model, the first target model can be quickly obtained by loading the same model parameters into the determined first alternative model.
[0100] In this application, the first alternative models and the second alternative models have the same model parameters. Each of the second alternative models can be a neural network model, and its model parameters can be weight parameters. Thus, the model parameters stored in the specified parameter storage space are also the model parameters of the second alternative models. When a second alternative model is determined for the first data, the model parameters can be read from the storage space, and the second target model can be obtained by loading the model parameters into the second alternative model. That is, the second target model can be a neural network model loaded with the same model parameters from the parameter storage space. Thus, it can be considered that the first target model and the second target model are neural network models with the same model parameters. Further, the first target model and the second target model can be neural network models that respectively load the same model parameters from the parameter storage space. In this application, the second target model can be quickly obtained by loading the same model parameters into the second alternative model.
[0101] In this application, combined with Figure 3As shown, the original neural network model (model structure, model parameters, model configuration information) is selected, and various possible pre-padding lengths (L1, L2... Lm, where m is a positive integer greater than or equal to 2) are configured. Using a compilation tool, the original neural network model is compiled based on different pre-padding lengths to obtain each first optional model adapted to each length. For example, the first optional model adapted to the L1 length, the first optional model adapted to the L2 length. Different first optional models are obtained by compiling the same original neural network model, so their model parameters are the same. The model parameters of each first optional model can be stored in a specified parameter storage space, and the first target model can be obtained by loading the model parameters stored in the parameter storage space through each first optional model.
[0102] In this application, for the pre-padding results of different first data, if the same second target model is used for decoding, then the same second target model can be obtained by compiling the same original neural network model as the first optional model. For the pre-padding results of different first data, if different second target models are used for decoding, then each second optional model can be obtained by compiling the same original neural network model as the first optional model according to the model attributes. The second target model is obtained by loading the same model parameters into the determined second optional model. The first optional model and the second optional model are obtained through the compilation of the compilation tool, and the model parameters shared among the first optional models and / or the model parameters shared between the first optional model and the second optional model can also be obtained.
[0103] It can be understood that the first target model and the second target model are neural network models with the same model parameters, but they are different computational flows. This difference in the computational flow is reflected in: their inputs are different, their computational processes are different, and their outputs are also different. For example Figure 4 As shown, the input of the first target model includes: the token sequence (input_ids) of the first data, the position (pos_ids) of each token in the token sequence, whether the position of the token is output (e.g., mask = 0 means not output, mask = 1 means output), and the input size (shape) of the first data. The first target model performs pre-padding processing on the token sequence of the first data, calculates the K matrix and V matrix of each token in the token sequence, and caches the KV matrix, caching the K matrix and V matrix in the form of KV cache. According to the K matrix and V matrix of each token, the first output token is predicted. The first target model obtains the pre-padding results of the first output token and the KV matrix through pre-padding processing.
[0104] For example Figure 4As shown, the input of the second target model includes: the output token (input_ids) predicted through the previous iteration calculation, the iteration number (pos_ids), whether the output token is output (e.g., mask = 0 means not output, mask = 1 means output), the input size (shape), and the KV cache obtained through the previous iteration calculation. The second target model decodes the new token in each iteration. Further, based on the output token predicted through the previous iteration calculation, the K matrix and V matrix obtained in the previous iteration, the possible tokens (new tokens) in each iteration are predicted to obtain the prediction result of the new token. The K matrix and V matrix of the new token are calculated, and the KV cache obtained through the previous iteration calculation is updated through the K matrix and V matrix of the new token. The decoding process of the second target model is a multi-iteration process, mainly including the prediction of the new token and the update of the KV cache. If a prediction of a new output token (each token corresponding to each character in the reply data) can be obtained in one iteration, then the first output token and the new output tokens obtained in each iteration are arranged in the order of the calculation sequence to obtain the tokens corresponding to the reply data. The token is converted into a token-character to obtain the reply data for replying to the first data. Among them, for the first iteration of the second target model, the input_ids input to the second target model is the first output token obtained in the pre-filling stage, and the KV cache input to the second target model is the KV cache calculated in the pre-filling stage. The first iteration predicts the token corresponding to the second character in the reply data based on the first output token (the token corresponding to the first character in the reply data) and the KV cache calculated in the pre-filling stage. Based on the token corresponding to the second character, the KV cache is updated. The token obtained in the first iteration is used as the input_ids for the next iteration calculation, and the KV cache obtained through the first iteration is used as the KV cache for the next iteration calculation. And so on, until the second target model completes the prediction of the tokens corresponding to all characters in the reply data.
[0105] In this application, the logic deployment of the foregoing data processing method in a heterogeneous multi-core system is taken as an example. In the embodiments of this application, the heterogeneous multi-core system is a heterogeneous multi-core chip. A heterogeneous multi-core chip refers to a chip integrating two or more processor cores within a single chip. For example, a single SOC (system-on-chip) chip integrating two or more processor cores. Each processor core in the heterogeneous multi-core chip can serve as an independent processor, independently run the instructions required by each processor core, and implement the tasks required by each processor core. It can be understood that a heterogeneous multi-core chip is a chip with multi-core processors. Compared with a single-core processor chip, the independent operation of each core task can accelerate the running speed and improve the multi-task execution ability, thus bringing the advantage of high performance. Moreover, the multi-core processors are arranged on the same chip, having the advantage of low cost.
[0106] For example Figure 5 As shown in the figure, the heterogeneous multi-core chip includes multiple processor cores, and the multiple processor cores include a first processor core, a second processor core... an L-th processor core. L is a positive integer greater than or equal to 2, which is flexibly set according to the actual situation. Among the multiple processor cores, each processor core is equivalent to a computing engine, and its type and / or quantity can be different. Among them, the types of processor cores include cores with strong computing power and cores with strong real-time performance (fast computing). In practical applications, most of the multiple processor cores are processor cores of different types, and a small number of cores are processor cores of the same type. It is also possible that the multiple processor cores can be processor cores of different types. Thus, the heterogeneous multi-core chip is composed of two or more processor cores with different architectures. The difference in the type and / or quantity of processor cores can, to a certain extent, result in different architectures among the processor cores.
[0107] In practical applications, among all the processor cores, as long as there are two or more processor cores of different types, such processor cores can be called heterogeneous multi-core, and the chip including these processor cores can be regarded as a heterogeneous multi-core chip.
[0108] Exemplarily, since the embedded processor (ARM) has the advantages of low cost and low power consumption, the digital signal processor (DSP) has the advantage of digital dedicated processing, and the programmable logic array (FPGA) has the advantage of high-speed processing. Each type of processor is used as a processor core. These types of processors are designed on the same SOC chip, and a heterogeneous multi-core SOC chip can be obtained.
[0109] For example Figure 6As shown in the figure, each processor core and the hardware resources connected to each processor core, such as a clock controller, an interrupt controller, a memory space, etc., constitute each hardware domain. That is, a multi-core heterogeneous chip includes multiple hardware domains. In a multi-core heterogeneous chip, each hardware domain is a set of hardware resources. Different hardware domains are isolated from each other, and this isolation can be regarded as a physical isolation. For example, the hardware designs within the same hardware domain are located in close positions in the multi-core heterogeneous chip, while the hardware designs in different hardware domains are located in different positions in the multi-core heterogeneous chip to achieve isolation in terms of physical location. Of course, the mutual isolation between different hardware domains in the embodiments of the present application may not be a physical isolation but a logical isolation. This logical isolation can be reflected in that the hardware resources within the same hardware domain need to use the same communication identifier for access within this hardware domain. That is, different hardware resources within the same hardware domain can access each other based on the communication identifier within this hardware domain. The hardware resources in different hardware domains use different communication identifiers for access.
[0110] In practical applications, it is preferably that the mutual isolation between different hardware domains is a logical isolation, so that at least the chip space can be saved.
[0111] Such as Figure 7 As shown in the figure, in a multi-core heterogeneous chip, an operating system can be configured for each hardware domain. For example, a first operating system is configured for the first hardware domain, a second operating system is configured for the second hardware domain, etc. The operating systems configured for different hardware domains can be the same or different, and preferably different. For example, the first operating system configured for the first hardware domain is the Linux system, and the second operating system configured for the second hardware domain is the Android system. Among them, since the Linux system has the characteristic of high security and the Android system has the characteristic of lightness, tasks with high security requirements in the multi-core heterogeneous system can be handed over to Linux for execution, and tasks that need to run lightly in the multi-core heterogeneous system can be handed over to the Android system for execution. Therefore, different operating systems on different hardware domains in the multi-core heterogeneous system can be used to efficiently execute each task.
[0112] In a multi-core heterogeneous chip, according to actual needs, an operating system can also be configured for each hardware domain in most of the hardware domains, and no operating system is configured for a small part of the hardware domains, depending on the specific usage situation.
[0113] In the embodiments of the present application, there is also a communication requirement between hardware domains. When there is a communication requirement between different hardware domains, an inter-core communication mechanism can be used to achieve communication between hardware domains. Among them, the inter-core communication mechanism in a multi-core heterogeneous system includes a mailbox mechanism suitable for instruction transmission and a memory sharing mechanism suitable for data sharing. The inter-core communication within a single SOC chip can ensure the transmission of data within the same chip, ensuring data security and fast transmission.
[0114] Under normal circumstances, there are differences in hardware resources within different hardware domains. Such differences may be reflected in aspects such as hardware type, hardware model, and hardware quantity. To a certain extent, this kind of difference can reflect the heterogeneity of the heterogeneous multi-core system. As can be seen from the previous introduction, the heterogeneous multi-core in this application is a concept at the hardware level and has nothing to do with the software level.
[0115] As Figure 5 shown, the heterogeneous multi-core chip in the embodiment of this application further includes various types of control units. The various types of control units include but are not limited to: a power control unit, a non-volatile storage control unit, a volatile storage control unit, etc. Among them, the power control unit is used to control the power unit to realize the power supply to the heterogeneous multi-core chip. The non-volatile storage control unit is used to control at least one processor core in the heterogeneous multi-core chip to access the non-volatile storage unit. The volatile storage control unit is used to control at least one processor core in the heterogeneous multi-core chip to access the volatile storage unit.
[0116] Among them, the power unit, the non-volatile storage unit, and the volatile storage unit, as hardware resources outside the heterogeneous multi-core chip, can be called when the heterogeneous multi-core chip needs them. In addition, hardware such as an audio output unit (such as a speaker or a loudspeaker), an audio acquisition unit (such as a microphone), and a video output unit (such as a display screen), as hardware resources outside the heterogeneous multi-core chip, can also be called when the heterogeneous multi-core chip needs them to realize the normal output of audio and video.
[0117] The data processing method in the embodiment of this application is implemented on a heterogeneous multi-core system. The heterogeneous multi-core system involved in the data processing method of this application includes M hardware domains, where M is a positive integer greater than or equal to 2. The M hardware domains can be Figure 6 all or part of the L hardware domains shown, preferably all hardware domains. Each hardware domain is configured with an independent operating system, that is, each hardware domain corresponds to an operating system. For example, hardware domain 1 corresponds to the first operating system, hardware domain 2 corresponds to the second operating system... hardware domain M corresponds to the Mth operating system. Each hardware domain is composed of multiple processor cores with different architectures in the heterogeneous multi-core system and the hardware resources connected to each processor core. The hardware domains are isolated from each other, and this isolation can be physical isolation or logical isolation. Refer to the relevant descriptions above.
[0118] In this application, a hardware domain and the operating system corresponding to the hardware domain can form a domain system. The multi-core heterogeneous chip includes two or more such domain systems. Each domain system can access system resources through the system bus. System resources can include peripherals such as an integrated circuit bus (I2C), a universal asynchronous receiver / transmitter (UART), and an interface (IO), and can also include resources that can be shared among domain systems, such as speakers, microphones, and interrupt controllers.
[0119] The domain systems in the multi-core heterogeneous system can be roughly divided into two categories according to their roles: one is the application domain, and the other is the management domain. Among them, the number of management domains can be one, can be two or more, and preferably one. The number of application domains is usually two or more. Taking the application of the multi-core heterogeneous system in a vehicle as an example, the application domain is used to implement various functions of the vehicle, such as the instrument function and the in-vehicle infotainment function (IVI function). The management domain is used to manage each application domain. In actual applications, the instrument function and the in-vehicle infotainment function can be integrated into different application domains. In this way, in the multi-core heterogeneous system applied to a vehicle, its application domain includes at least an instrument domain and an IVI domain. Of course, in addition to the application domain and the management domain, according to actual application requirements, the domain system can be divided into different roles. For example, the multi-core heterogeneous system can also include a small system domain and / or a security domain. Among them, the security domain is used to ensure the security of the multi-core heterogeneous system, and the small system domain is used to assist other domain systems to implement the powerful functions of the multi-core heterogeneous system. For example, the small system domain can implement accelerated startup. In this application, the status or role of domain systems with different roles is the same.
[0120] The processing logic of the data processing method in this application can be deployed in a vehicle using a multi-core heterogeneous system, specifically in one of the domain systems of the multi-core heterogeneous system. Taking the deployment in one of the application domain systems as an example, taking the voice (input) data of "I'm going to place A, what's the driving route?" generated by the driver to the domain system as an example, the application domain system converts the language data into text data and converts the text data into a token sequence. According to the length of the text data or the length of the token sequence converted from the text data, from multiple first optional models, the model that matches this length is selected as the first target model. For example, if the length of the input token sequence is 15, the model that matches its length can be the first optional model with a length of 2 4 =16, and the first optional model is used to implement the pre-filling process of the large language model. Since the model for implementing the pre-filling process is a model whose length matches the length of the input token sequence, it can speed up the processing in the preprocessing stage and quickly obtain the first output token and the cached result of the KV matrix of the input token sequence. The second target model can perform decoding processing based on the quickly obtained first output token and the KV matrix, which can speed up the inference and obtain the response data.
[0121] Combined Figure 8 As shown, during implementation, the application (APP) of the application domain system receives the voice input data 1 generated by the user. The CPU of the application domain system converts the voice input data 1 into text data 1 and then converts the text data 1 into a token sequence 1. The CPU calls the NPU, and the NPU sends a call instruction to its module - runtime lib (runtime module) to achieve the call of the first target model with a matching adaptation length. The runtime lib responds to this instruction and selects, from multiple first optional models, the first optional model that matches the length based on the length of the token sequence 1. The runtime lib hands over the selected first optional model with the matching length to the NPU for driving. The NPU drives to load the model parameters of the first optional model from a specified parameter storage space such as memory to form the first target model, and loads the model parameters of the decoded version backbone to form the second target model. The NPU calls the first target model and uses the first target model to perform pre - filling processing on the token sequence 1. The NPU calls the second target model and uses the second target model to perform decoding processing on the token sequence 1 to infer the response data 1 for the voice input data 1.
[0122] If, after the user generates the aforementioned voice input data 1, voice input data 2 is generated, the APP of the application domain system receives the voice input data 2. The CPU of the application domain system converts the voice input data 2 into text data 2 and then converts the text data 2 into a token sequence 2. The NPU sends a call instruction to its module - runtime lib. The runtime lib responds to this instruction and selects, from multiple first optional models, the first optional model that matches the length based on the length of the token sequence 2. The runtime lib hands over the selected first optional model with the matching length for the token sequence 2 to the NPU for driving. The NPU drives to load the model parameters of the first optional model from a specified parameter storage space such as memory to form the first target model. The NPU calls the first target model and uses the first target model to perform pre - filling processing on the token sequence 2. The NPU uses the second target model to perform decoding processing on the token sequence 2 to infer the response data 2 for the voice input data 2.
[0123] For the voice input data successively generated by the user, through the foregoing solution, a response to each voice input data can be given to implement an interactive mode of question and answer between the user and the domain system. From the perspective of the user, the rapid output of the response data enables the user to obtain a response without waiting for a long time and enjoy a good user experience. From the perspective of the system side, the inference of the input data generated by the driver implemented by the data processing method of the present application can achieve the rapid output of the inference result such as the response data and improve the inference efficiency.
[0124] Combined with Figure 8 As shown, in the process of the NPU using two models to implement pre-filling processing and decoding processing, the data being calculated can be stored in the NPU to achieve fast calculation of the NPU. The data that is not being calculated and may be used later can be stored in the memory to avoid the impact of excessive storage of data in the NPU on the fast calculation of the NPU. In Figure 8 , the pre-filling version backbones of various lengths such as the pre-filling version backbone of length L1 and the pre-filling version backbone of length L2 can be considered as the first optional models adapted to lengths L1, L2, etc. The model parameters that can be shared among the first optional models are stored in the memory for each length of the first optional model to load the model parameters from the memory to form the first target model. If the same second target model is used for decoding processing for different first data, the decoded version backbone can be considered as the second optional model, and the model parameters of the second optional model are loaded from the memory to form the second target model.
[0125] It can be understood that if the two stages of the large language model: the pre-filling stage and the decoding stage, the pre-filling process and the decoding process implemented are handed over to two models for processing, then the first optional model or the first target model in the present application can be regarded as the model for executing the pre-filling process, and the second optional model or the second target model can be regarded as the model for executing the decoding process.
[0126] The present application provides a data processing device, as Figure 9 shown, including:
[0127] A first obtaining unit 901, configured to obtain first data, where the first data is the input data of the large language model;
[0128] A first determining unit 902, configured to determine a first target model that matches the attribute of the first data from multiple first optional models, where each first optional model is obtained by program compilation based on different pre-filling lengths, and the first optional models have the same model parameters;
[0129] The first processing unit 903 is configured to perform pre-filling processing of the large language model on the first data by using the first target model, so as to obtain a pre-filling result of the first data. The pre-filling result includes the result of key-value caching for the input tokens of the first data and generating the first initial output token of the first data. Among them, the first initial output token generated by using the first target model is earlier than the first initial output token generated by using other first alternative models except the first target model.
[0130] The second processing unit 904 is configured to perform decoding processing of the large language model on the first data by using the second target model based on the key-value caching result and the first initial output token, so as to obtain a second data for replying to the first data.
[0131] In some embodiments, the first data includes N first sub-data, where N is a positive integer greater than or equal to 2.
[0132] The first determination unit 902 is configured to determine, for each first sub-data, a first target model from multiple first alternative models based on the attributes of each first sub-data, so that the first processing unit 903 performs pre-filling processing on each first sub-data by using each first target model.
[0133] The second processing unit 904 is configured to perform decoding processing on the pre-filling results of each first sub-data by using the second target model based on the priorities of each first sub-data, so as to obtain each second sub-data, and each second sub-data is used to reply to each first sub-data, where the second data includes each second sub-data.
[0134] In some embodiments, the first determination unit 902 is configured to:
[0135] Obtain the context information of the first data;
[0136] Obtain the relevance between the first data and the context information;
[0137] In response to the relevance satisfying a preset relevance condition, determine, from multiple first alternative models, a first target model for the first data based on the attributes of the first data and the attributes of the context information.
[0138] In some embodiments, the data processing device further includes: a second determination unit, which is configured to:
[0139] Determine a second target model from multiple second alternative models based on at least one of the attributes of the first data, the attributes of the first target model, and the pre-filling result of the first data, where each second alternative model is obtained by at least performing program compilation based on the attributes of the model.
[0140] In some embodiments, the second determination unit is further configured to:
[0141] Based on at least one of the attributes of each first sub-data, the attributes of each first target model determined for the first sub-data, and the pre-population results of each first sub-data, determine, for each first sub-data, each second target model from multiple second optional models;
[0142] A second processing unit 904, configured to:
[0143] Perform decoding processing on the pre-population results of each first sub-data by using each second target model to obtain each second sub-data, where the second sub-data is used to reply to each first sub-data, and where the second data includes the second sub-data.
[0144] In some embodiments, the same model parameters are saved or stored in a parameter storage space, and the first target model is a neural network model that loads the same model parameters from the parameter storage space.
[0145] In some embodiments, the first target model and the second target model are neural network models with the same model parameters;
[0146] The first target model and the second target model are neural network models that respectively load the same model parameters from the parameter storage space.
[0147] In some embodiments, a first determination unit 902 is configured to:
[0148] Obtain the attribute of the first data;
[0149] Obtain the target attributes of each first optional model, where the target attributes are related to the attribute of the first data; determine, from multiple first optional models, the first optional model whose target attributes match the attribute of the first data as the first target model.
[0150] It should be noted that for the data processing device in the embodiments of the present application, since the principle of the data processing device for solving problems is similar to the foregoing data processing method, therefore, the implementation process and implementation principle of the data processing device can refer to the implementation process and implementation principle descriptions of the foregoing method, and the repeated parts will not be elaborated.
[0151] According to the embodiments of the present application, the present application also provides an electronic device and a readable storage medium.
[0152] Among them, the electronic device includes at least one processor; and a memory communicatively connected to the at least one processor; wherein the memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor to enable the at least one processor to execute the data processing method described in this application. The computer instructions are used to cause the computer to execute the data processing method described in this application.
[0153] This application also provides a computer program product, including a computer program / instructions, and when the computer program / instructions are executed by a processor, the data processing method of this application is implemented.
[0154] Figure 10 FIG. shows a schematic block diagram of an exemplary electronic device 800 that can be used to implement the embodiments of this application. The electronic device is intended to represent various forms of digital computers, such as, laptop computers, desktop computers, workstations, personal digital assistants, servers, blade servers, mainframe computers, and other suitable computers. The electronic device can also represent various forms of mobile devices, such as, personal digital processors, cellular phones, smart phones, wearable devices, and other similar computing devices. The components shown herein, their connections and relationships, and their functions are merely examples and are not intended to limit the implementation of this application described and / or claimed herein.
[0155] As Figure 10 shown, the device 800 includes a computing unit 801, which can execute various appropriate actions and processes according to the computer program stored in the read-only memory (ROM) 802 or the computer program loaded from the storage unit 808 into the random access memory (RAM) 803. In the RAM 803, various programs and data required for the operation of the device 800 can also be stored. The computing unit 801, the ROM 802, and the RAM 803 are connected to each other through a bus 804. The input / output (I / O) interface 805 is also connected to the bus 804.
[0156] Multiple components in the device 800 are connected to the I / O interface 805, including: an input unit 806, such as a keyboard, a mouse, etc.; an output unit 807, such as various types of displays, speakers, etc.; a storage unit 808, such as a magnetic disk, an optical disc, etc.; and a communication unit 809, such as a network card, a modem, a wireless communication transceiver, etc. The communication unit 809 allows the device 800 to exchange information / data with other devices through a computer network such as the Internet and / or various telecommunication networks.
[0157] The computing unit 801 can be various general-purpose and / or special-purpose processing components with processing and computing capabilities. Some examples of the computing unit 801 include, but are not limited to, a central processing unit (CPU), a graphics processing unit (GPU), various dedicated artificial intelligence (AI) computing chips, various computing units running machine learning model algorithms, a digital signal processor (DSP), and any suitable processor, controller, microcontroller, etc. The computing unit 801 executes the various methods and processes described above, such as the data processing method. For example, in some embodiments, the data processing method can be implemented as a computer software program that is tangibly embodied in a machine-readable medium, such as the storage unit 808. In some embodiments, part or all of the computer program can be loaded and / or installed onto the device 800 via the ROM 802 and / or the communication unit 809. When the computer program is loaded into the RAM 803 and executed by the computing unit 801, one or more steps of the data processing method described above can be executed. Alternatively, in other embodiments, the computing unit 801 can be configured to execute the data processing method in any other suitable manner (e.g., by means of firmware).
[0158] Various embodiments of the systems and techniques described above in this document can be implemented in digital electronic circuitry, integrated circuit systems, field-programmable gate arrays (FPGA), application-specific integrated circuits (ASIC), application-specific standard products (ASSP), systems-on-chip (SOC), complex programmable logic devices (CPLD), computer hardware, firmware, software, and / or combinations thereof. These various embodiments can include: being implemented in one or more computer programs that can be executed and / or interpreted on a programmable system including at least one programmable processor, which can be a dedicated or general-purpose programmable processor that can receive data and instructions from a storage system, at least one input device, and at least one output device, and transmit the data and instructions to the storage system, the at least one input device, and the at least one output device.
[0159] The program code for implementing the methods of this application can be written in any combination of one or more programming languages. These program codes can be provided to the processor or controller of a general-purpose computer, a special-purpose computer, or other programmable data processing device, such that when the program codes are executed by the processor or controller, the functions / operations specified in the flowcharts and / or block diagrams are implemented. The program codes can be executed entirely on the machine, partially on the machine, as a stand-alone software package partially on the machine and partially on a remote machine, or entirely on a remote machine or server.
[0160] In the context of this application, a machine-readable medium can be a tangible medium that can contain or store a program for use by or in connection with an instruction execution system, apparatus, or device. A machine-readable medium can be a machine-readable signal medium or a machine-readable storage medium. A machine-readable medium can include, but is not limited to, electronic, magnetic, optical, electromagnetic, infrared, or semiconductor systems, apparatus, or devices, or any suitable combination of the foregoing. More specific examples of a machine-readable storage medium would include an electrical connection based on one or more wires, a portable computer diskette, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or Flash memory), an optical fiber, a portable compact disc read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the foregoing.
[0161] To provide for interaction with a user, the systems and techniques described herein can be implemented on a computer having: a display device (e.g., a CRT (cathode ray tube) or LCD (liquid crystal display) monitor) for displaying information to the user; and a keyboard and a pointing device (e.g., a mouse or a trackball) by which the user can provide input to the computer. Other kinds of devices can also be used to provide for interaction with the user; for example, feedback provided to the user can be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and input from the user can be received in any form (including acoustic, speech, or tactile input).
[0162] The systems and techniques described herein can be implemented in a computing system that includes back-end components (e.g., as a data server), or a computing system that includes middleware components (e.g., an application server), or a computing system that includes front-end components (e.g., a user computer having a graphical user interface or a web browser through which the user can interact with an implementation of the systems and techniques described herein), or a computing system that includes any combination of such back-end components, middleware components, or front-end components. The components of the system can be interconnected by any form or medium of digital data communication (e.g., a communication network). Examples of communication networks include: a local area network (LAN), a wide area network (WAN), and the Internet.
[0163] A computer system can include a client and a server. The client and the server are generally remote from each other and typically interact through a communication network. The client-server relationship is created by computer programs running on the respective computers and having a client-server relationship to each other. The server can be a cloud server, a server of a distributed system, or a server incorporating a blockchain.
[0164] As described above, it is only the specific implementation manner of the present application, but the protection scope of the present application is not limited thereto. Any person skilled in the art within the technical scope disclosed by the present application can easily think of changes or substitutions, which should all be covered within the protection scope of the present application. Therefore, the protection scope of the present application shall be subject to the protection scope of the claimed rights.
Claims
1. A data processing method, characterized in that, Including: Obtain first data, where the first data is input data of a large language model; Determine a first target model that matches the attributes of the first data from multiple first optional models, where each first optional model is obtained by program compilation based on a different pre-filled length, and the first optional models have the same model parameters; Use the first target model to perform pre-filling processing of the large language model on the first data to obtain a pre-filling result for the first data, where the pre-filling result includes the result of key-value caching of the input tokens of the first data and generating the first output token of the first data, and the first output token generated using the first target model is earlier than the first output tokens generated using other first optional models except the first target model; Based on the key-value caching result and the first output token, use a second target model to perform decoding processing of the large language model on the first data to obtain second data for replying to the first data.
2. The method according to claim 1, characterized in that The first data includes N first sub-data, where N is a positive integer greater than or equal to 2, and the method further includes: Based on the attributes of each first sub-data, determine a first target model for each first sub-data from multiple first optional models to perform pre-filling processing on each first sub-data using each first target model; Based on the priorities of each first sub-data, use the second target model to perform decoding processing on the pre-filling results of each first sub-data to obtain each second sub-data, and each second sub-data is used to reply to each first sub-data, where the second data includes the second sub-data.
3. The method according to claim 1, characterized in that, The determining a first target model that matches the attributes of the first data from multiple first optional models includes: Obtain the context information of the first data; Obtain the relevance between the first data and the context information; In response to the relevance satisfying a preset relevance condition, based on the attributes of the first data and the attributes of the context information, determine a first target model for the first data from multiple first optional models.
4. The method according to claim 1, characterized in that, Also including: Based on at least one of the attributes of the first data, the attributes of the first target model, and the pre-filling result of the first data, determine a second target model from multiple second optional models, where each second optional model is obtained by program compilation based at least on the attributes of the model.
5. The method according to claim 4, wherein The first data includes N first sub-data, where N is a positive integer greater than or equal to 2, and the method further includes: Based on at least one of the attributes of each first sub-data, the attributes of each first target model determined for the first sub-data, and the pre-filling results of each first sub-data, determine each second target model for each first sub-data from multiple second optional models; Use each second target model to perform decoding processing on the pre-filling results of each first sub-data to obtain each second sub-data, and the second sub-data is used to reply to each first sub-data, where the second data includes the second sub-data.
6. The method according to claim 1, wherein The same model parameters are saved or stored in a parameter storage space, and the first target model is a neural network model that loads the same model parameters from the parameter storage space.
7. The method according to claim 1, wherein The first target model and the second target model are neural network models with the same model parameters; The first target model and the second target model are neural network models that respectively load the same model parameters from the parameter storage space.
8. The method according to claim 1, characterized in that, Determining the first target model that matches the attributes of the first data from multiple first alternative models includes: Obtaining the attributes of the first data; Obtaining the target attributes of each first alternative model, where the target attributes are related to the attributes of the first data; determining the first alternative model whose target attributes match the attributes of the first data from multiple first alternative models as the first target model.
9. A data processing device, characterized in that, Includes: A first obtaining unit for obtaining first data, where the first data is the input data of the large language model; A first determining unit for determining the first target model that matches the attributes of the first data from multiple first alternative models, where each first alternative model is obtained by compiling a program based on a different pre-padding length, and the first alternative models have the same model parameters; A first processing unit for performing pre-padding processing of the large language model on the first data using the first target model to obtain a pre-padding result for the first data, where the pre-padding result includes the result of key-value caching of the input tokens of the first data and generating the first character output token of the first data, and the first character output token generated using the first target model is earlier than the first character output token generated using other first alternative models except the first target model; A second processing unit for performing decoding processing of the large language model on the first data based on the key-value caching result and the first character output token to obtain a second data for replying to the first data.
10. An electronic device, characterized in that, The electronic device includes: At least one processor; and a memory communicatively connected to the at least one processor; wherein, The memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor so that the at least one processor can execute the method according to any one of claims 1-8.
Citation Information
Cited By
Data transmission method and device, storage medium and electronic equipment
CN120892163A