An Inference Optimization Method for a Language Model
Through segmented full-scale inference and mixed inference, combined with zero redundancy processing, the inference efficiency and video memory usage problems of large language models when processing long-text requests are solved, and more efficient inference services are achieved.
Patent Information
- Application Number
- CN202411076211.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-08-07
- Publication Date
- 2025-05-27
- Estimated Expiration
- 2044-08-07
AI Technical Summary
When large language models process long-term inference, the full inference time is long, resulting in the long-term waiting time of incremental inference, which reduces the efficiency of computational inference, and the use of video memory resources is tight, limiting the long-term inference process.
The input information is processed in segmented full-scale inference and mixed inference methods, and the video memory usage is optimized through zero redundancy processing to improve the model inference performance.
This reduces the video memory usage when inference of long-text information, and improves the throughput and computational inference efficiency of language model inference services.
Smart Images

Figure CN119090006B_ABST
Abstract
Description
Technical Field
[0001] This application relates to the technical field of large language models, and in particular, to a method for optimizing the inference of a language model. Background Art
[0002] The basic structure of large language models generally is the decoder part of the Transformer language model. The Transformer language model is an autoregressive language model. During the prediction process, it predicts the next text token according to the previous context, and uses the newly obtained text token as the input to predict the next text token, and continues to iterate like this until a token representing the end flag is generated to terminate; in terms of inference, it is divided into two stages. The first stage is the full inference (or the first inference, pre-filling) stage for the previous prompt input by the user, and the second stage is the incremental inference stage for the model prediction result.
[0003] Currently, for extremely long inputs in long text requests, the language model may take several seconds for full inference. The extremely long input cannot perform incremental inference simultaneously with short inputs, resulting in an overly long waiting time for incremental inference and reducing the efficiency of the language model during computational inference; in addition, in large model inference, storage resources have become an important competitive factor in high-performance computing. In the case of processing extremely long inputs, the occupation of video memory resources is very tight, restricting the process of long text inference of the language model. Summary of the Invention
[0004] In view of this, the purpose of this application is to provide a method for optimizing the inference of a language model. Based on the idea of continuous batch processing, it performs segmented full inference and hybrid inference on at least one request information, achieving less occupation of video memory when the language model infers long text information input, improving the throughput of the language model inference service in practical applications, and through zero redundancy processing in full inference and hybrid inference, optimizing the video memory occupation and model inference performance during the inference process, and improving the efficiency of the language model during computational inference.
[0005] The embodiment of this application provides a method for optimizing the inference of a language model, and the inference optimization method includes:
[0006] In response to obtaining at least one request information for inferring input information using a preset language model, dividing the first request information into multiple segments of request information according to a preset length;
[0007] Performing full-scale inference and hybrid inference on the multi-segment request information and the second request information using the language model to obtain multiple first inference results, where the second request information is at least one request information obtained simultaneously with the first request information or during the full-scale inference or hybrid inference of the multi-segment request information;
[0008] Performing incremental inference on the multiple first inference results using the language model to obtain the inference result output by the language model for the input information.
[0009] Further, the performing full-scale inference and hybrid inference on the multi-segment request information and the second request information using the language model to obtain multiple first inference results includes:
[0010] Performing full-scale inference, or full-scale inference and hybrid inference, on the first request fragment information in the multi-segment request information using the language model to obtain historical result information output by the full-scale inference, or the full-scale inference and the hybrid inference;
[0011] In response to obtaining the second request information for performing inference on the input information using a preset language model, taking the second request fragment information in the multi-segment request information and the second request information as a same batch of requests to obtain first request batch information;
[0012] Based on the historical result information and the first request batch information, performing hybrid inference using the language model to obtain multiple first inference results.
[0013] Further, the performing full-scale inference and hybrid inference on the multi-segment request information and the second request information using the language model to obtain multiple first inference results further includes:
[0014] In response to obtaining the third request information for performing inference on the input information using a preset language model, taking the multiple first inference results and the third request information as a same batch of requests to obtain second request batch information;
[0015] Based on the second request batch information, performing hybrid inference using the language model to obtain multiple first inference results.
[0016] Further, the performing full-scale inference and hybrid inference on the multi-segment request information and the second request information using the language model to obtain multiple first inference results includes:
[0017] Performing full inference, or full inference and hybrid inference, on the first request fragment information in the multiple segments of request information using the language model to obtain first historical result information output by the full inference, or the full inference and the hybrid inference;
[0018] In response to obtaining a second request information for performing inference on input information using a preset language model, combining the second request fragment information in the multiple segments of request information and the second request information as a same batch of requests to obtain first request batch information;
[0019] Based on the first historical result information and the first request batch information, performing hybrid inference using the language model to obtain at least one of the first inference results and second historical result information;
[0020] In response to obtaining a third request information for performing inference on input information using a preset language model, combining the first inference result, the third request fragment information in the multiple segments of request information, and the third request information as a same batch of requests to obtain third request batch information;
[0021] Based on the second historical result information and the third request batch information, performing hybrid inference using the language model to obtain multiple first inference results.
[0022] Further, the performing full inference and hybrid inference on the multiple segments of request information and the second request information using the language model to obtain multiple first inference results includes:
[0023] Batch - merging the multiple segments of request information and the second request information, and performing zero - redundant full inference and hybrid inference on the multiple segments of request information and the second request information using the language model to obtain multiple first inference results;
[0024] Further, the batch - merging the multiple segments of request information and the second request information, and performing zero - redundant full inference and hybrid inference on the multiple segments of request information and the second request information using the language model to obtain multiple first inference results includes:
[0025] Performing full inference, or full inference and hybrid inference, on the first request fragment information in the multiple segments of request information using the language model to obtain historical result information output by the full inference, or the full inference and the hybrid inference;
[0026] In response to obtaining the second request information for performing inference on input information using a preset language model, combining the second request fragment information in the multiple segments of request information and the second request information to obtain first combined information;
[0027] Based on the historical result information and the first merged information, perform zero-redundancy hybrid inference using the language model to obtain multiple first inference results.
[0028] Further, the performing zero-redundancy hybrid inference using the language model based on the historical result information and the first merged information to obtain multiple first inference results includes:
[0029] Perform token embedding and layer normalization on the first merged information and the historical result information respectively to obtain first merged input information;
[0030] Use the self-attention layer of the language model to perform vector linear transformation on the first merged input information to obtain first vector input information;
[0031] Based on the first vector input information, perform zero-redundancy full-scale operation using the self-attention layer to obtain a first output result;
[0032] Use the feed-forward layer of the language model to perform full-scale feed-forward processing on the first output result to obtain a first intermediate result;
[0033] Based on N pieces of the first merged input information, first vector input information, first output result, and first intermediate result, obtain multiple first inference results, where N is a natural number greater than 1.
[0034] Further, the performing zero-redundancy full-scale operation using the self-attention layer based on the first vector input information to obtain a first output result includes:
[0035] Disperse and input the first vector input information into multiple independent operators in the self-attention layer according to a preset partitioning strategy;
[0036] Use the multiple independent operators to perform full-scale operations on the dispersed and input first vector input information respectively to obtain multiple first operation results, and merge and write the multiple first operation results into the corresponding preset video memory of the language model;
[0037] Split the multiple first operation results merged and written into the preset video memory to obtain a first output result.
[0038] Further, the batch-merging the multi-segment request information and the second request information, and performing zero-redundancy full-scale inference and hybrid inference on the multi-segment request information and the second request information using the language model to obtain multiple first inference results further includes:
[0039] In response to a third request message for inferring input information using a preset language model, the first inference result and the third request message are merged to obtain a second merged message;
[0040] Based on the second merged message, zero-redundancy hybrid inference is performed using the language model to obtain multiple first inference results.
[0041] Further, the batch merging of the multiple segments of request messages and the second request message, and the zero-redundancy full-scale inference and hybrid inference of the multiple segments of request messages and the second request message using the language model to obtain multiple first inference results include:
[0042] Using the language model to perform full-scale inference, or full-scale inference and hybrid inference, on the first request fragment information in the multiple segments of request messages to obtain first historical result information output by the full-scale inference, or the full-scale inference and the hybrid inference;
[0043] In response to the second request message for inferring input information using a preset language model, the second request fragment information in the multiple segments of request messages and the second request message are merged to obtain a first merged message;
[0044] Based on the first historical result information and the first merged message, zero-redundancy hybrid inference is performed using the language model to obtain at least one of the first inference results and second historical result information;
[0045] In response to a third request message for inferring input information using a preset language model, the first inference result, the third request fragment information in the multiple segments of request messages, and the third request message are merged to obtain a third merged message;
[0046] Based on the second historical result information and the third merged message, zero-redundancy hybrid inference is performed using the language model to obtain multiple first inference results.
[0047] The embodiments of the present application further provide an inference optimization device for a language model, and the inference optimization device includes:
[0048] A segmentation processing module, configured to divide a first request message into multiple segments of request messages according to a preset length in response to at least one request message for inferring input information using a preset language model when the request message is obtained;
[0049] A hybrid inference module for performing full-scale inference and hybrid inference on the multi-segment request information and the second request information by using the language model to obtain multiple first inference results, where the second request information is at least one request information obtained simultaneously with the first request information or during the full-scale inference or hybrid inference of the multi-segment request information;
[0050] An incremental inference module for performing incremental inference on the multiple first inference results by using the language model to obtain the inference result output by the language model for the input information.
[0051] An embodiment of the present application further provides an electronic device, including: a processor, a memory, and a bus. The memory stores machine-readable instructions executable by the processor. When the electronic device runs, the processor communicates with the memory through the bus. When the machine-readable instructions are executed by the processor, the steps of the inference optimization method of the language model as described above are executed.
[0052] An embodiment of the present application further provides a computer-readable storage medium, on which a computer program is stored. When the computer program is run by a processor, the steps of the inference optimization method of the language model as described above are executed.
[0053] The inference optimization method of the language model provided by the embodiment of the present application includes: in response to obtaining at least one request information for performing inference on the input long text information by using a preset language model, dividing the first request information into multi-segment request information according to a preset length; performing full-scale inference and hybrid inference on the multi-segment request information and the second request information by using the language model to obtain multiple first inference results, where the second request information is at least one request information obtained simultaneously with the first request information or during the full-scale inference or hybrid inference of the multi-segment request information; performing incremental inference on the multiple first inference results by using the language model to obtain the inference result output by the language model for the input information.
[0054] Compared with the method for processing the ultra-long input of the long text request in the prior art, based on the idea of continuous batch processing, segmented full-scale inference and hybrid inference are performed on at least one request information, realizing less memory usage of the language model when performing inference on the input long text information, improving the throughput of the language model inference service in practical applications, and optimizing the memory usage and model inference performance during the inference process through zero redundancy processing in the full-scale inference and hybrid inference, improving the efficiency of the language model during computational inference.
[0055] To make the above - mentioned objects, features, and advantages of the present application more obvious and understandable, the following provides preferred embodiments in conjunction with the accompanying drawings and makes a detailed description as follows. Brief Description of the Drawings
[0056] To more clearly illustrate the technical solutions of the embodiments of the present application, the following briefly introduces the drawings required for the embodiments. It should be understood that the following drawings only show some embodiments of the present application and should not be regarded as limiting the scope. For those of ordinary skill in the art, without creative efforts, other related drawings can be obtained based on these drawings.
[0057] Figure 1 It is a hierarchical abstraction schematic diagram of an inference framework of a language model provided by an embodiment of the present application;
[0058] Figure 2 It is a flowchart of an inference optimization method of a language model provided by an embodiment of the present application;
[0059] Figure 3 It is a segmented schematic diagram of full - scale inference of a language model provided by an embodiment of the present application;
[0060] Figure 4 It is an inference schematic diagram of zero - redundancy processing of a language model provided by an embodiment of the present application;
[0061] Figure 5 It is a structural schematic diagram of an inference optimization device of a language model provided by an embodiment of the present application;
[0062] Figure 6 It is a structural schematic diagram of an electronic device provided by an embodiment of the present application. Detailed Embodiments
[0063] To make the objectives, technical solutions, and advantages of the embodiments of the present application clearer, the following will clearly and completely describe the technical solutions in the embodiments of the present application in conjunction with the drawings in the embodiments of the present application. Obviously, the described embodiments are only a part of the embodiments of the present application, rather than all the embodiments. Usually, the components of the embodiments of the present application described and shown in the drawings here can be arranged and designed in various different configurations. Therefore, the following detailed description of the embodiments of the present application provided in the drawings is not intended to limit the scope of the present application to be protected, but only represents the selected embodiments of the present application. Based on the embodiments of the present application, every other embodiment obtained by those of ordinary skill in the art without creative efforts belongs to the scope of protection of the present application.
[0064] Currently, the classical architecture of the Transformer model in language models includes three main components: the self-attention (Masked Multi-Head Attention) part, which is used to establish the dependencies between positions in the input sequence; the residual connection and normalization operation part (Add&Norm); the feed-forward neural network part (Feed Forward) and the multi-layer perceptron (MLP) part. Through the cooperation of these components, the Transformer model can effectively capture the key information in the input sequence and achieve powerful sequence modeling capabilities.
[0065] It has been found through research that the relatively large intermediate video memory part in the inference process based on the Transformer model structure mainly includes the input and output parts of the model, the output part of the self-attention layer, and the temporary video memory part required in the process of first increasing the dimension and then decreasing the dimension in the feed-forward layer; among them, the video memory occupancy is linearly related to the length of the text tokens processed by the language model. The size of the intermediate video memory part in the inference process of the language model has an important impact on processing long text inputs. In large model inferences, storage resources have become an important competitive factor in high-performance computing. In the case of processing extremely long inputs, the occupancy of video memory resources is very tight, restricting the process of long text inference by the language model.
[0066] Here, when the language model based on the Transformer model structure receives an extremely long input for a long text request, the language model may take several seconds for full-scale inference. The extremely long input cannot be incrementally inferred simultaneously with short inputs, resulting in an overly long waiting time for incremental inference and reducing the efficiency of the language model in computational inference; in addition, in large model inferences, storage resources have become an important competitive factor in high-performance computing. In the case of processing extremely long inputs, the occupancy of video memory resources is very tight, restricting the process of long text inference by the language model.
[0067] In view of the problems in at least one of the above aspects, the embodiments of the present application provide an inference optimization method for a language model.
[0068] In the embodiments of the present application, when the preset language model infers the input long text information, the overall service framework of the language model includes but is not limited to the application business layer, the inference service processor, and the inference engine, etc. Among them, the application business layer includes various application businesses. For example, it realizes application businesses such as dialogue chat, intelligent assistant, and creation; the inference service processor is responsible for processing multiple concurrent requests, such as requests for initialization of requests, invocation of the inference engine, and sampling of inference results; the inference engine can be classified according to the characteristics of inference, such as full-scale inference, incremental inference, and hybrid inference.
[0069] Figure 1A hierarchical abstraction schematic diagram of the inference framework of a language model provided by an embodiment of the present application.
[0070] As Figure 1 shown in the figure, in this example, it includes three abstraction layers. The first abstraction layer (the top layer) refers to model abstraction, which includes but is not limited to the full-scale inference process of the execution context, the incremental inference process of autoregressive execution, and the hybrid inference including full-scale inference and incremental inference. The second abstraction layer (the middle layer) refers to the model layer abstraction for performing operations on all the same layers, which mainly includes but is not limited to Self Attention (self-attention layer) and FeedForward (feed-forward layer). The third abstraction layer (the bottom layer) refers to operator abstraction, which includes but is not limited to parts such as FMHA (multi-head self-attention), linear transformation (Linear), add and normalization (Add&Norm), and softmax.
[0071] First, the full-scale inference, hybrid inference, and incremental inference involved in the embodiments of the present application are introduced:
[0072] Full-scale inference means that when a language model processes the input text data information, the language model will perform a one-time comprehensive analysis and processing through the entire dataset stored in the model to generate a complete output or result. Among them, the language model will consider all information and context related to the input text data information during the processing to ensure the accuracy and comprehensiveness of the output.
[0073] As an example of full-scale inference, assume that when a user inputs a complex question to a chatbot based on a language model, the chatbot can perform full-scale inference in the following way: First, the chatbot analyzes the syntactic structure of the question and understands what the user wants to ask. Then, the chatbot retrieves its knowledge base to find information and answers related to the question. Finally, the chatbot integrates this information to generate a complete, coherent, and accurate answer and presents it to the user.
[0074] Hybrid inference refers to a strategy that flexibly selects and combines methods between full-scale inference and incremental inference according to different parts or different stages of a complex problem when solving the complex problem.
[0075] Incremental inference is applicable to scenarios where the input data changes little or repetitive inference tasks need to be performed. When only local changes occur in the input data, only the changed part is re-inferred instead of repeating the inference on the entire dataset or input. Incremental inference can significantly reduce redundant calculations, reduce the amount of operations and power consumption, thereby improving the execution speed and efficiency of inference.
[0076] As an example of incremental reasoning, assume that in a video object detection task, there is a large amount of similar information between consecutive frames in the video. In this case, incremental reasoning can be performed only on regions with significant inter-frame differences, thereby reducing the repetitive processing of the entire frame image.
[0077] Please refer to Figure 2 , Figure 2 which is a flowchart of a method for optimizing inference of a language model provided by an embodiment of the present application.
[0078] As Figure 2 shown in
[0079] Step S100: In response to at least one request message for inferring input information using a preset language model obtained, divide the first request message into multiple segments of request messages according to a preset length.
[0080] In an embodiment of the present application, when inputting information into a preset language model and requesting the language model to perform inference on the input information, at least one request message can be obtained for the request. Here, the input information includes, but is not limited to, long text information, short text information, context information, metadata, feature representations, and other information. At least one request message refers to one or more pieces of information determined based on the information input into the language model, and the request message is used to instruct the language model to perform inference on the input information.
[0081] In addition, the first request message is the request message that needs to be split among at least one request message. The preset length can be specifically calibrated according to the performance of the preset language model and the length of the input long text information to set the division length.
[0082] In this step, in specific implementation, first, obtain at least one request message for inferring the input information input into the preset language model from an external request; then, in response to the acquisition of the at least one request message, obtain the preset length corresponding to the external request; finally, divide the first request message among the at least one request message into multiple segments of request messages according to the preset length.
[0083] For example, assume that the length of the first request message is 4000 and the preset length is 2000. At this time, the first request message with a length of 4000 can be divided into 2 segments of request messages according to the preset length of 2000, where the length of each segment of request message is 2000.
[0084] Step S200: Use the language model to perform full inference and hybrid inference on the multi-segment request information and the second request information to obtain multiple first inference results, where the second request information is at least one request information obtained simultaneously with the first request information or during the full inference or hybrid inference of the multi-segment request information.
[0085] Since it is necessary to perform inference on multi-segment request information and multi-batch request information in the embodiments of the present application, therefore, the processing methods for the multi-segment request information and the second request information may include two processing methods: batch parallel processing or merging processing.
[0086] One way is to process the information by batch parallel processing of the multi-segment request information and the second request information to improve the processing efficiency of the language model for the input information.
[0087] Another way is to process the information by merging the multi-segment request information and the second request information to achieve the effect of reducing resource occupancy while improving the processing efficiency of the language model for the input information.
[0088] Furthermore, considering that there are multiple partitioning situations for the multi-segment request information divided from the first request information, one situation is that the multi-segment request information is divided into N, and another situation is that the multi-segment request information is divided into N + 1. N is a natural number greater than 1.
[0089] Based on this, for each of the two processing methods of batch parallel processing or merging processing mentioned above, it respectively includes one situation where the multi-segment request information is divided into N and another situation where the multi-segment request information is divided into N + 1.
[0090] Next, two situations under the two processing methods in step S200 will be described with specific examples.
[0091] (1) The processing method is batch parallel processing:
[0092] In the first situation of the batch parallel processing method, for example, assuming that the multi-segment request information is divided into 2, use the language model to perform full inference and hybrid inference on the multi-segment request information and the second request information to obtain multiple first inference results, which may specifically include:
[0093] Step S211: Use the language model to perform full inference, or full inference and hybrid inference, on the first request fragment information in the multi-segment request information to obtain the historical result information output by the full inference, or the full inference and the hybrid inference.
[0094] In the embodiment of the present application, according to whether the obtained request information contains historical information and the input length of the request information, the specific inference method to be executed when the language model infers the request information is determined.
[0095] Specifically, when the request information does not contain historical information, full-scale inference is performed on the request information; when the request information contains historical information and the input length of the request information is 1, incremental inference is performed on the request information; when the request information contains historical information and the input length of the request information is not 1, hybrid inference is performed on the request information.
[0096] In this step, according to whether the first request fragment information in the divided multi-segment request information contains historical result information, and the input length of the first request fragment information, it is determined to perform full-scale inference on the first request fragment, or full-scale inference and hybrid inference; then, the historical result information output by the full-scale inference, or the full-scale inference and the hybrid inference is obtained.
[0097] For example, assume that the A request information does not contain historical information and the input length is 2000, then full-scale inference is performed on the A request information.
[0098] For another example, assume that the B request information and the C request information are obtained. The B request information contains historical information and the input length is 1, and incremental inference is performed on the B request information; the C request information contains historical information and the input length is not 1, and hybrid inference is performed on the C request information.
[0099] Here, the first request fragment information is not the first fragment in the physical sense, but the information that needs to be inferred after the first request information is segmented and before the second request information is received.
[0100] Specifically, in the process of performing full-scale inference and full-scale inference with historical information, both the fragment information for performing full-scale inference and the fragment information for full-scale inference with historical information (which can be called hybrid inference) are determined as the first request fragment information.
[0101] As an example of performing full-scale inference on the first request fragment information, the inference of the first request fragment information in the multi-segment request information is executed through the statement "model.run_gpt(input_tokens=split_1_2000)", where the length (tokens) of the first request information is 2000. Further, full-scale inference is performed on the first request fragment information through the statements "input_lengths=
[2000] ; resource_indices=[0]; infer_categories=[0]; history_lengths=[0]; is_first_finished_list=[0]", where the first request fragment information uses the resource with index 0 and the length of the historical result information is 0.
[0102] Figure 3 This is a segmented schematic diagram of full-scale inference of a language model provided by an embodiment of the present application.
[0103] As Figure 3 shown in [reference], when performing full-scale inference on request information by a language model, it is assumed that the first request information with a length of 8000 is divided into 4 request information with a length of 2000, and by adding full-scale inference with historical information, performing a full-scale inference with an input length of 8000 becomes performing 4 full-scale inferences with historical information; among them, the length of the historical information in the first full-scale inference is 0. In this way, since the size of the intermediate video memory is linearly related to the length of the processed text tokens (tokens), the segmented processing of full-scale inference can optimize the occupancy of the peak video memory.
[0104] In the example shown in Figure 3 [reference], performing full-scale inference by means of information segmentation reduces the occupancy of the system's intermediate video memory by three-quarters. Compared with the common method of optimizing video memory by offloading through a host device, performing full-scale inference by segmentation avoids the overhead of data copying between devices and the problems of complex heterogeneous scheduling implementation.
[0105] Step S212, in response to obtaining the second request information for inferring the input information using a preset language model, use the second request fragment information and the second request information in the multi-segment request information as a same batch of requests to obtain a first request batch information.
[0106] In an embodiment of the present application, during the process of performing full reasoning, or full reasoning and mixed reasoning on the first request fragment information in the first request information, the second request information requesting to use a preset language model to reason about the input information is obtained. At this time, the second request fragment information in the first request information can be merged with the second request information into the same batch for parallel processing.
[0107] For example, suppose that during the full inference, or the full inference and mixed inference process of the D1 request fragment information in the D request information, the E request information is obtained. At this time, the D2 request fragment information in the D request information needs to be merged with the E request information into the same batch request so as to process the same batch request in parallel.
[0108] Here, it should be understood that when at least one request information is obtained, the first request information needs to be immediately divided into multiple request information segments, and the second request fragment information in the multiple request information segments and the second request information need to be further merged into the same request batch information to process the request batch information in parallel.
[0109] The at least one request information is not necessarily acquired at one time, and may be second request information acquired during the processing of the first request information.
[0110] Step S213: Based on the historical result information and the first request batch information, hybrid reasoning is performed using the language model to obtain multiple first reasoning results.
[0111] In this step, during the specific implementation, the historical result information and the first request batch information are used as inputs of the language model. At this time, since the input contains historical result information and the input length is not 1, the language model can be used to perform mixed reasoning on the input to obtain multiple first reasoning results.
[0112] Optionally, in addition to using the language model to perform full reasoning and mixed reasoning on the multiple request information and the second request information to obtain multiple first reasoning results as described in the above steps, steps S214 to S215 are also included. Specifically, steps S214 to S215 are used to illustrate a reasoning method for obtaining the third request information after reasoning on the obtained second request information, and specifically include:
[0113] Step S214: In response to obtaining the third request information requesting to use a preset language model to infer the input information, the plurality of first inference results and the third request information are taken as the same batch request to obtain second request batch information.
[0114] In an embodiment of the present application, when the third request information is obtained after the second request information is obtained during the process of using a language model to reason about the first request fragment information in the first request information, the first reasoning result and the third request information can be used as requests in the same batch to obtain the second request batch information.
[0115] For example, assume that the D2 request fragment information in the D request information is merged with the E request information to form the first request batch information, and the historical result information obtained by reasoning about the D1 request fragment information in the D request information and this first request batch information are input into the language model for hybrid reasoning to obtain multiple first reasoning results S1. At this time, the F request information is obtained, and the first reasoning result S1 and the F request information are used as requests in the same batch to obtain the second request batch information.
[0116] Step S215: Based on the second request batch information, use the language model to perform hybrid reasoning to obtain multiple first reasoning results.
[0117] In this step, in specific implementation, the second request batch information is used as the input of the language model. At this time, since this input contains historical result information and the input length is not 1, the language model can be used to perform hybrid reasoning on this input to obtain multiple first reasoning results corresponding to the processing of the first request information, the second request information, and the third request information.
[0118] As an example of the first case in the way of batch parallel processing, first, the D request information obtained first is segmented to obtain the D1 request fragment information and the D2 request fragment information; then, the D1 request fragment information is reasoned to obtain the historical result information Z, and the E request information is obtained during this process; after that, the D2 request fragment information needs to be merged with the E request information to form the first request batch information, and the historical result information Z and this first request batch information are input into the language model for hybrid reasoning to obtain multiple first reasoning results S1; at this time, the F request information is obtained again, and the first reasoning result S1 and the F request information are used as requests in the same batch to obtain the second request batch information; finally, the second request batch information is input into the language model for hybrid reasoning to obtain multiple first reasoning results S2.
[0119] The second case in the way of batch parallel processing, for example, assume that when the multiple segments of request information are divided into 3 segments, use the language model to perform full-scale reasoning and hybrid reasoning on the multiple segments of request information and the second request information to obtain multiple first reasoning results, which may specifically include:
[0120] Step S221: Use the language model to perform full inference, or full inference and hybrid inference, on the first request fragment information in the multiple segments of request information, and obtain first historical result information output by the full inference, or the full inference and the hybrid inference.
[0121] Step S222: In response to the second request information for which a preset language model is used to perform inference on the input information upon receiving the request, use the second request fragment information and the second request information in the multiple segments of request information as a same batch of requests to obtain first request batch information.
[0122] Step S223: Based on the first historical result information and the first request batch information, use the language model to perform hybrid inference to obtain at least one of the first inference results and second historical result information.
[0123] Among them, the descriptions of steps S221 to S223 can refer to the descriptions of steps S211 to S213, and the same technical effects can be achieved, so details are not described herein.
[0124] Step S224: In response to the third request information for which a preset language model is used to perform inference on the input information upon receiving the request, use the first inference result, the third request fragment information in the multiple segments of request information, and the third request information as a same batch of requests to obtain third request batch information.
[0125] In the embodiment of the present application, since the third request fragment information is included in the multiple segments of request information, therefore, for the divided third request fragment information, the third request fragment information, the first inference result, and the third request information are used as a same batch of requests to obtain third request batch information for performing hybrid inference on the third request batch information.
[0126] Step S225: Based on the second historical result information and the third request batch information, use the language model to perform hybrid inference to obtain multiple first inference results.
[0127] In this step, the second historical result information and the third request batch information are used as the input of the language model, and the language model performs hybrid inference on this input to obtain multiple first inference results.
[0128] As an example of the second case in the way of batch parallel processing, first, segment the D request information obtained first to obtain D1 request segment information, D2 request segment information, and D3 request segment information; then, perform reasoning on the D1 request segment information to obtain the first historical result information Z1, and obtain E request information during this process. It is necessary to merge the D2 request segment information and the E request information into the first request batch information; after that, input the first historical result information Z1 and the first request batch information into the language model for hybrid reasoning to obtain at least one first reasoning result S1 and the second historical result information Z2; after that, obtain F request information, and use the first reasoning result S1, D3 request segment information, and F request information as the same batch of requests to obtain the third request batch information; finally, input the third request batch information and the second historical result information Z2 into the language model for hybrid reasoning to obtain multiple first reasoning results S2.
[0129] (2) The processing method is the merging processing method:
[0130] For the first case in the merging processing method, for example, assuming that when the multi-segment request information is divided into 2 segments, using the language model to perform full-scale reasoning and hybrid reasoning on the multi-segment request information and the second request information to obtain multiple first reasoning results, specifically, it may include:
[0131] Step S230: Batch and merge the multi-segment request information and the second request information, and use the language model to perform zero-redundancy full-scale reasoning and hybrid reasoning on the multi-segment request information and the second request information to obtain multiple first reasoning results.
[0132] In an implementation manner of the present application, in specific implementation, step S230 may include:
[0133] Step S231: Use the language model to perform full-scale reasoning, or full-scale reasoning and hybrid reasoning, on the first request segment information in the multi-segment request information to obtain the historical result information output by the full-scale reasoning, or the full-scale reasoning and the hybrid reasoning.
[0134] Among them, the description of S231 can refer to the description of S211, and the same technical effect can be achieved, so it will not be elaborated here.
[0135] Step S232: In response to obtaining the second request information for reasoning the input information using the preset language model, merge the second request segment information in the multi-segment request information and the second request information to obtain the first merged information.
[0136] In an embodiment of the present application, during the process of reasoning on the first request fragment information, the second request information is obtained, and the second request fragment information and the second request information can be merged and processed to reduce the resource occupation of the language model reasoning in a zero-redundancy reasoning manner.
[0137] In an example of the present application, the second request information is obtained through the statement "request_2 = GetLastRequest()", and the second request fragment information and the second request information are merged through the statement "input_tokens = Merge(split_2_2000, request_2)" to obtain the first merged information.
[0138] For example, assume that during the process of reasoning on the G1 request fragment information in the G request information, the H request information is obtained. At this time, the G1 request fragment information and the H request information need to be merged into the first merged information for zero-redundancy reasoning on the first merged information.
[0139] Step S233: Based on the historical result information and the first merged information, use the language model to perform zero-redundancy hybrid reasoning to obtain multiple first reasoning results.
[0140] In an embodiment of the present application, the language model runs reasoning in N layers during the reasoning process. There are multiple hidden layers in N layers. The multiple hidden layers are multi-level abstractions of the input features to linearly divide the data in the input layer into different types of data. N is a natural number greater than 1.
[0141] In an implementation manner of the present application, in specific implementation, in step S233, zero-redundancy hybrid reasoning is respectively performed on the vectors included in the hidden layer to obtain multiple first reasoning results, which specifically includes:
[0142] S2331: Perform token embedding and layer normalization processing on the first merged information and the historical result information respectively to obtain the first merged input information.
[0143] In this step, in specific implementation, first, perform model token embedding on the first merged information and the historical result information; then, calculate the data distribution characteristics included in the information after model token embedding, and adjust the data distribution characteristics to a specified range for normalization processing; finally, introduce a preset scaling factor (scale) and a translation factor (shift) into the data after normalization processing to obtain the first merged input information.
[0144] S2332: Use the self-attention layer of the language model to perform vector linear transformation on the first merged input information to obtain the first vector input information.
[0145] In this step, the self-attention layer of the language model is used to transform the first merged input information through three different linear transformation methods. Specifically, the first merged input information is multiplied by three different weight matrices respectively to obtain three vectors Q, K, and V, and these three vectors are determined as the obtained first vector input information.
[0146] Among them, the three different weight matrices can respectively correspond to the trainable parameter matrices W Q 、W K 、W V 。
[0147] S2333. Based on the first vector input information, use the self-attention layer to perform zero-redundancy full-scale operations to obtain a first output result.
[0148] In an implementation manner of the present application, in specific implementation, step S2333 may include:
[0149] Step S23331: Disperse the first vector input information into a plurality of independent operators in the self-attention layer according to a preset partitioning strategy.
[0150] In this step, based on the number of information in the first vector input information, the first vector input information is sequentially and dispersedly input into a plurality of independent operators in the self-attention layer.
[0151] For example, assume that the first vector input information includes information with lengths of 1000, 1, 1000, and 1 respectively. The first vector input information is respectively input into four independent operators for subsequent operator operations.
[0152] Step S23332: Use the plurality of independent operators to respectively perform full-scale operations on the first vector input information dispersedly input, obtain a plurality of first operation results, and merge and write the plurality of first operation results into the preset video memory corresponding to the language model.
[0153] In the embodiment of the present application, after the language model processes the entire input sequence, it will generate a plurality of first operation results at one time, and merge and write the plurality of first operation results into the preset video memory corresponding to the language model.
[0154] Step S23333: Split the plurality of first operation results merged and written into the preset video memory to obtain a first output result.
[0155] In this step, according to the input token information after request processing, the plurality of first operation results are split and processed inside the preset video memory to obtain a first output result.
[0156] Step S2334: Perform full-feed processing on the first output result using the feed-forward layer of the language model to obtain a first intermediate result.
[0157] In the embodiments of the present application, based on the ability of the feed-forward layer in the language model to process data, full-feed processing is performed on the first output result to obtain a first intermediate result, so as to perform operations on the first intermediate result through other processes included in the language model.
[0158] Step S2335: Based on N pieces of the first merged input information, the first vector input information, the first output result, and the first intermediate result, obtain multiple first inference results, where N is a natural number greater than 1.
[0159] In the embodiments of the present application, the hidden layer vector includes, but is not limited to, the first merged input information, the first vector input information, the first output result, and the first intermediate result.
[0160] Here, it should be understood that since no inference result is obtained after the full-feed processing of the first output result by the feed-forward layer, after the model runs through all layers, the language model performs other processes on the hidden layer vector to obtain multiple first inference results.
[0161] Figure 4 It is a schematic diagram of inference for zero-redundancy processing of a language model provided by the embodiments of the present application.
[0162] As an example of zero-redundancy inference, as Figure 4 shown, assume that the first merged information includes 4 pieces of input information, and their lengths are 1000, 1, 1000, and 1 in sequence, and the shape size of the first merged information input is 2002; after performing embedding and layer normalization processing on the first merged information respectively, since the above processing is for the text tokens of each input information, the resulting intermediate video memory shape size is [2002, H] (H represents the width of the hidden layer).
[0163] Further, during the inference process of the self-attention layer, since the attn operator for calculating the attention scores is for each input information in the first merged information, after performing a vector linear transformation (QKV linear) on the first merged input information, the first vector input information is scattered into 4 independent operators (attn1-4) for execution to internally split the obtained multiple first operation results and write them into the same preset video memory block; then, the feed forward layer of the language model is used to perform a full-feed forward process on the first output result to obtain a first intermediate result; finally, based on the first merged input information, the first vector input information, the first output result, and the first intermediate result, multiple first inference results are obtained through the processing of other layers of the model.
[0164] Through the above method of performing zero-redundancy processing at the input end of the language model, a separation and merging operation is achieved in the operator operation part, reducing the occupancy of the temporary video memory by the language model during inference.
[0165] Optionally, in addition to the full inference and hybrid inference for merging the multiple segment request information and the second request information using the language model as described in the above steps to obtain multiple first inference results, it further includes steps S231 to S233. Specifically, steps S234 to S235 are used to illustrate the inference method for obtaining a third request information after inferring the obtained second request information, and specifically include:
[0166] Step S234, in response to obtaining a third request information for inferring input information using a preset language model, merge the first inference result and the third request information to obtain second merged information.
[0167] In the embodiment of the present application, for the obtained third request information, the first inference result can be merged with the third request information to obtain second merged information for performing zero-redundancy hybrid inference on the second merged information.
[0168] Step S235, based on the second merged information, perform zero-redundancy hybrid inference using the language model to obtain multiple first inference results.
[0169] Among them, the description of step S235 can refer to the description of steps S2331 to S2335 and can achieve the same technical effect, which will not be elaborated here.
[0170] As an example of the first case in the merging process, first, the obtained G request information is segmented to obtain G1 request fragment information and G2 request fragment information; then, the G1 request fragment information is inferred to obtain historical result information Y, and H request information is obtained in this process; after that, the G2 request fragment information needs to be merged with the H request information to form a first merged information, and the historical result information Y and this first merged information are input into a language model for zero-redundancy hybrid inference to obtain multiple first inference results X1; at this time, I request information is obtained again, and the first inference results X1 and the I request information are merged to obtain a second merged information; finally, the second merged information is input into the language model for zero-redundancy hybrid inference to obtain multiple first inference results X2.
[0171] For the second case in the merging process, for example, assuming that the multi-segment request information is divided into 3 segments, the language model is used to perform full-scale inference and hybrid inference on the multi-segment request information and the second request information to obtain multiple first inference results. Specifically, it may include:
[0172] Step S241: Use the language model to perform full-scale inference, or full-scale inference and hybrid inference, on the first request fragment information in the multi-segment request information to obtain first historical result information output by the full-scale inference, or the full-scale inference and the hybrid inference.
[0173] Step S242: In response to obtaining the second request information for which the input information is inferred using a preset language model, merge the second request fragment information in the multi-segment request information and the second request information to obtain a first merged information.
[0174] Step S243: Based on the first historical result information and the first merged information, use the language model to perform zero-redundancy hybrid inference to obtain at least one of the first inference results and second historical result information.
[0175] Among them, the descriptions of steps S241 to S243 can refer to the descriptions of steps S231 to S233 and can achieve the same technical effects, which will not be elaborated here.
[0176] Step S244: In response to obtaining the third request information for which the input information is inferred using a preset language model, merge the first inference results, the third request fragment information in the multi-segment request information, and the third request information to obtain a third merged information.
[0177] In the embodiments of the present application, for the obtained third request information, the first inference results, the third request fragment information, and the third request information can be merged to obtain a third merged information, so as to perform zero-redundancy hybrid inference on the third merged information and the second historical result information.
[0178] Step S245: Based on the second historical result information and the third merging information, perform zero-redundancy hybrid reasoning using the language model to obtain multiple first reasoning results.
[0179] Among them, the description of step S245 can refer to the description of steps S2331 to S2335, and the same technical effects can be achieved, so it will not be elaborated here.
[0180] As an example of the second case in the way of batch parallel processing, first, segment the G request information obtained first to obtain G1 request fragment information, G2 request fragment information, and G3 request fragment information; then, perform reasoning on the G1 request fragment information to obtain the first historical result information Y1, and in this process, obtain the H request information. It is necessary to merge the G2 request fragment information and the H request information into the first merging information; then, input the first historical result information Y1 and the first merging information into the language model for zero-redundancy hybrid reasoning to obtain at least one first reasoning result X1 and the second historical result information Y2; then, obtain the I request information, and merge the first reasoning result X1, D3 request fragment information, and I request information into the third merging information; finally, input the second historical result information Y2 and the third merging information into the language model for zero-redundancy hybrid reasoning to obtain multiple first reasoning results X2.
[0181] Return for reference Figure 2 , step 300: Perform incremental reasoning on the multiple first reasoning results using the language model to obtain the reasoning result output by the language model for reasoning on the input information.
[0182] In this step, when processing the multi-segment request information and the second request information in the way of batch parallel processing, regard the multiple first reasoning results as the results of the same batch to obtain the first request batch result, and perform incremental reasoning on the first request batch result using the language model to obtain the reasoning result output by the language model for reasoning on the input information.
[0183] When processing the multi-segment request information and the second request information in the way of merging processing, this step may include: merging the multiple first reasoning results to obtain merging information, and performing zero-redundancy incremental reasoning on the merging information using the language model to obtain the merged reasoning result output by the language model for reasoning on the input information.
[0184] As an example of incremental inference on merged information, incremental inference on the merged information is performed through the statement "logits = model.run_gpt(input_tokens = input_tokens, input_lengths = [1, 1, 1], resource_indices = [0, 1, 2], infer_categories = [1, 1, 1], history_lengths = [4001, 2001, 1000], is_first_finished_list = [1, 1, 1])".
[0185] Among them, input_tokens represents a list of input text tokens (Tokens), whose shape is [total_length], and padding processing needs to be performed when used; input_lengths represents a list of input lengths, whose shape is [max_batch_size]; resource_indices represents a list of positions of computing resources, whose shape is [max_batch_size], and the value range is [0, max_batch_size - 1]; infer_categories represents a list of input inference categories, whose shape is [max_batch_size], and the parameter 0 represents full inference, and the parameter 1 represents incremental inference; history_lengths represents the input of the number of tokens that have been inferred, whose shape is [max_batch_size]; is_first_finished_list represents whether the input has been completed in the current full inference, that is, whether it is the last piece of input, whose shape is [max_batch_size], and the parameter -1 represents an invalid value (for incremental inference); the parameter 0 represents that the inference has not ended; the parameter 1 represents that the inference has ended.
[0186] Here, incremental inference is a full inference with historical information, and its input length is 1, that is, full inference with historical information naturally supports incremental inference; since in the input batch processing with a length of 1, the difficulty of inference lies in the IO bandwidth, therefore, short inputs with a length of 1 are merged with long inputs for inference, and the inference latency brought by short inputs can be ignored.
[0187] In this way, through the hybrid inference process of segmented batch processing, continuous requests can be efficiently processed, and while ensuring real-time performance, the advantages of batch processing can be fully utilized, dynamically increasing the batch size of incremental inference, so that the system can effectively utilize computing resources and improve the throughput and efficiency of inference.
[0188] The inference optimization method of the language model provided by the embodiment of the present application, based on the idea of continuous batch processing, performs segmented full-scale inference and hybrid inference on at least one request message, realizes less video memory occupancy when the language model infers long text information input, improves the throughput of the language model inference service in practical applications, and optimizes the video memory occupancy and model inference performance during the inference process through zero redundancy processing in full-scale inference and hybrid inference, thereby improving the efficiency of the language model during computational inference.
[0189] Please refer to Figure 5 , Figure 5 which is a schematic structural diagram of an inference optimization device for a language model provided by the embodiment of the present application. As Figure 5 shown in
[0190] The segmentation processing module 510 is configured to, in response to obtaining at least one request message for inferring input information by using a preset language model, divide the first request message into multiple segmented request messages according to a preset length;
[0191] The hybrid inference module 520 is configured to perform full-scale inference and hybrid inference on the multiple segmented request messages and the second request message by using the language model to obtain multiple first inference results, where the second request message is at least one request message obtained simultaneously with the first request message or during the full-scale inference or hybrid inference of the multiple segmented request messages;
[0192] The incremental inference module 530 is configured to perform incremental inference on the multiple first inference results by using the language model to obtain the inference result output by the language model for inferring the input information.
[0193] Further, when the hybrid inference module 520 is used to perform full-scale inference and hybrid inference on the multiple segmented request messages and the second request message by using the language model to obtain multiple first inference results, the hybrid inference module 520 is configured to:
[0194] Perform full-scale inference, or full-scale inference and hybrid inference, on the first request fragment information in the multiple segmented request messages by using the language model to obtain historical result information output by the full-scale inference, or the full-scale inference and the hybrid inference;
[0195] In response to obtaining the second request message for inferring input information by using a preset language model, use the second request fragment information in the multiple segmented request messages and the second request message as a same batch of requests to obtain a first request batch message;
[0196] Based on the historical result information and the first request batch information, perform hybrid reasoning using the language model to obtain multiple first reasoning results.
[0197] Further, when the hybrid reasoning module 520 is used to perform full-scale reasoning and hybrid reasoning on the multi-segment request information and the second request information using the language model to obtain multiple first reasoning results, the hybrid reasoning module 520 is further used to:
[0198] In response to obtaining a third request information for reasoning on input information using a preset language model, use the multiple first reasoning results and the third request information as a request of the same batch to obtain second request batch information;
[0199] Based on the second request batch information, perform hybrid reasoning using the language model to obtain multiple first reasoning results.
[0200] Further, when the hybrid reasoning module 520 is used to perform full-scale reasoning and hybrid reasoning on the multi-segment request information and the second request information using the language model to obtain multiple first reasoning results, the hybrid reasoning module 520 is further used to:
[0201] Perform full-scale reasoning, or full-scale reasoning and hybrid reasoning, on the first request fragment information in the multi-segment request information using the language model to obtain first historical result information output by the full-scale reasoning, or the full-scale reasoning and the hybrid reasoning;
[0202] In response to obtaining the second request information for reasoning on input information using a preset language model, use the second request fragment information in the multi-segment request information and the second request information as a request of the same batch to obtain first request batch information;
[0203] Based on the first historical result information and the first request batch information, perform hybrid reasoning using the language model to obtain at least one of the first reasoning results and second historical result information;
[0204] In response to obtaining a third request information for reasoning on input information using a preset language model, use the first reasoning result, the third request fragment information in the multi-segment request information, and the third request information as a request of the same batch to obtain third request batch information;
[0205] Based on the second historical result information and the third request batch information, perform hybrid reasoning using the language model to obtain multiple first reasoning results.
[0206] Further, when the hybrid inference module 520 is used to perform full-scale inference and hybrid inference on the multi-segment request information and the second request information using the language model to obtain multiple first inference results, the hybrid inference module 520 is further used for:
[0207] Batch-combine the multi-segment request information and the second request information, and use the language model to perform zero-redundancy full-scale inference and hybrid inference on the multi-segment request information and the second request information to obtain multiple first inference results;
[0208] Further, when the hybrid inference module 520 is used to batch-combine the multi-segment request information and the second request information, and use the language model to perform zero-redundancy full-scale inference and hybrid inference on the multi-segment request information and the second request information to obtain multiple first inference results, the hybrid inference module 520 is further used for:
[0209] Use the language model to perform full-scale inference, or full-scale inference and hybrid inference, on the first request fragment information in the multi-segment request information to obtain historical result information output by the full-scale inference, or the full-scale inference and the hybrid inference;
[0210] In response to the second request information for performing inference on the input information using a preset language model, merge the second request fragment information in the multi-segment request information and the second request information to obtain first merged information;
[0211] Based on the historical result information and the first merged information, use the language model to perform zero-redundancy hybrid inference to obtain multiple first inference results.
[0212] Further, when the hybrid inference module 520 is used to perform zero-redundancy hybrid inference based on the historical result information and the first merged information using the language model to obtain multiple first inference results, the hybrid inference module 520 is further used for:
[0213] Perform token embedding and layer normalization processing on the first merged information and the historical result information respectively to obtain first merged input information;
[0214] Use the self-attention layer of the language model to perform vector linear transformation on the first merged input information to obtain first vector input information;
[0215] Based on the first vector input information, use the self-attention layer to perform zero-redundancy full-scale operation to obtain a first output result;
[0216] Use the feed-forward layer of the language model to perform full-scale feed-forward processing on the first output result to obtain a first intermediate result;
[0217] Based on the N first merged input information, the first vector input information, the first output result, and the first intermediate result, multiple first inference results are obtained, where N is a natural number greater than 1.
[0218] Further, when the hybrid inference module 520 is used to perform zero-redundancy full-scale operations on the first vector input information using the self-attention layer to obtain the first output result, the hybrid inference module 520 is further used to:
[0219] Disperse the first vector input information into multiple independent operators in the self-attention layer according to a preset partitioning strategy;
[0220] Use the multiple independent operators to perform full-scale operations on the dispersed first vector input information respectively to obtain multiple first operation results, and merge and write the multiple first operation results into a preset video memory corresponding to the language model;
[0221] Split the multiple first operation results merged and written into the preset video memory to obtain the first output result.
[0222] Further, when the hybrid inference module 520 is used to batch merge the multi-segment request information and the second request information, and use the language model to perform zero-redundancy full-scale inference and hybrid inference on the multi-segment request information and the second request information to obtain multiple first inference results, the hybrid inference module 520 is further used to:
[0223] In response to obtaining a third request information for performing inference on input information using a preset language model, merge the first inference result and the third request information to obtain second merged information;
[0224] Based on the second merged information, perform zero-redundancy hybrid inference using the language model to obtain multiple first inference results.
[0225] Further, when the hybrid inference module 520 is used to batch merge the multi-segment request information and the second request information, and use the language model to perform zero-redundancy full-scale inference and hybrid inference on the multi-segment request information and the second request information to obtain multiple first inference results, the hybrid inference module 520 is further used to:
[0226] Use the language model to perform full-scale inference, or full-scale inference and hybrid inference, on the first request fragment information in the multi-segment request information to obtain first historical result information output by the full-scale inference, or the full-scale inference and the hybrid inference;
[0227] In response to obtaining the second request information requesting to use a preset language model to infer the input information, merging the second request fragment information in the multiple request information segments with the second request information to obtain first merged information;
[0228] Based on the first historical result information and the first merged information, using the language model to perform zero-redundancy hybrid reasoning to obtain at least one of the first reasoning results and the second historical result information;
[0229] In response to obtaining third request information requesting to use a preset language model to infer input information, merging the first inference result, the third request segment information in the multiple segments of request information, and the third request information to obtain third merged information;
[0230] Based on the second historical result information and the third merged information, zero-redundancy hybrid reasoning is performed using the language model to obtain multiple first reasoning results.
[0231] The language model reasoning optimization device provided in the embodiment of the present application performs segmented full reasoning and hybrid reasoning on at least one request information based on the idea of continuous batch processing, thereby achieving a smaller occupation of video memory when the language model performs reasoning on input long text information, improving the throughput of the language model reasoning service in practical applications, and optimizing the video memory occupancy and model reasoning performance during the reasoning process through zero-redundancy processing in full reasoning and hybrid reasoning, thereby improving the efficiency of the language model when performing computational reasoning.
[0232] See also Figure 6 , Figure 6 This is a schematic diagram of the structure of an electronic device provided in an embodiment of the present application. Figure 6 As shown in , the electronic device 600 includes a processor 610 , a memory 620 and a bus 630 .
[0233] The memory 620 stores machine-readable instructions executable by the processor 610. When the electronic device 600 is running, the processor 610 communicates with the memory 620 via the bus 630. When the machine-readable instructions are executed by the processor 610, the above-mentioned Figure 2 The steps of the method for optimizing the reasoning of the language model in the method embodiment shown, and the specific implementation method thereof can be found in the method embodiment, which will not be described in detail here.
[0234] The present application also provides a computer-readable storage medium on which a computer program is stored. When the computer program is executed by a processor, the computer program can execute the above-mentioned Figure 2For the steps of the inference optimization method of the language model in the method embodiments shown, the specific implementation manners can be referred to the method embodiments and will not be elaborated herein.
[0235] Those skilled in the art can clearly understand that for the sake of convenience and brevity of description, the specific working processes of the above-described systems, devices, and units can refer to the corresponding processes in the foregoing method embodiments and will not be elaborated herein.
[0236] In several embodiments provided in the present application, it should be understood that the disclosed systems, devices, and methods can be implemented in other ways. The device embodiments described above are merely illustrative. For example, the division of the units is only a logical function division, and there can be other division manners in actual implementation. For another example, multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the displayed or discussed couplings or direct couplings or communication connections to each other can be through some communication interfaces. The indirect couplings or communication connections of the devices or units can be in electrical, mechanical, or other forms.
[0237] The units described as separate components may or may not be physically separated. The components displayed as units may or may not be physical units, that is, they can be located in one place, or they can be distributed to multiple network units. Some or all of the units can be selected according to actual needs to achieve the purpose of the solution of this embodiment.
[0238] In addition, in each embodiment of the present application, the functional units can be integrated in a processing unit, or each unit can exist physically alone, or two or more units can be integrated in one unit.
[0239] If the function is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a non-volatile computer-readable storage medium executable by a processor. Based on such an understanding, the technical solution of the present application, in essence, or the part that contributes to the prior art, or a part of this technical solution, can be embodied in the form of a software product. The computer software product is stored in a storage medium and includes several instructions for causing a computer device (which can be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the methods described in each embodiment of the present application. The foregoing storage medium includes: various media such as USB flash drives, mobile hard disks, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical discs that can store program codes.
[0240] Finally, it should be noted that the above-described embodiments are only specific implementation manners of the present application, used to illustrate the technical solutions of the present application, rather than limiting it. The protection scope of the present application is not limited thereto. Although the present application has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that any person skilled in the art within the technical scope disclosed by the present application can still modify the technical solutions described in the foregoing embodiments or can easily think of changes, or perform equivalent replacements on some of the technical features; and these modifications, changes or replacements do not make the essence of the corresponding technical solutions deviate from the spirit and scope of the technical solutions of the embodiments of the present application, and should all be covered within the protection scope of the present application. Therefore, the protection scope of the present application shall be subject to the protection scope of the claims.
Claims
1. A method for optimizing reasoning of a language model, characterized in that: The reasoning optimization method comprises: In response to obtaining at least one request information requesting to use a preset language model to infer input information, dividing the first request information into a plurality of request information segments according to a preset length; Determine, according to whether the request information contains historical information and the input length of the request information, a first reasoning mode and a second reasoning mode of the language model for the request information; the first reasoning mode includes full reasoning and mixed reasoning, and the second reasoning mode includes incremental reasoning; Using the language model to perform full reasoning and mixed reasoning on the multiple segments of request information and the second request information, to obtain multiple first reasoning results, wherein the second request information is at least one request information obtained at the same time as the first request information or obtained in the process of performing full reasoning or mixed reasoning on the multiple segments of request information; The processing method of the multiple request information segments and the second request information at least includes batch parallel processing; when the processing method is batch parallel processing, the language model is used to perform full reasoning and mixed reasoning on the multiple request information segments and the second request information to obtain multiple first reasoning results, including: Using the language model, perform full reasoning, or full reasoning and mixed reasoning, on a first request segment information among the multiple request information segments to obtain historical result information output by the full reasoning, or the full reasoning and mixed reasoning; In response to obtaining the second request information requesting to use a preset language model to infer the input information, taking the second request fragment information in the multiple request information segments and the second request information as the same batch request to obtain first request batch information; Based on the historical result information and the first request batch information, using the language model to perform hybrid reasoning to obtain a plurality of first reasoning results; Incremental reasoning is performed on the plurality of first reasoning results using the language model to obtain a reasoning result output by the language model when reasoning on the input information.
2. The method according to claim 1, characterized in that: The method of using the language model to perform full reasoning and mixed reasoning on the multiple request information segments and the second request information to obtain multiple first reasoning results also includes: In response to obtaining third request information requesting to use a preset language model to infer input information, taking the plurality of first inference results and the third request information as a same batch request to obtain second request batch information; Based on the second request batch information, hybrid reasoning is performed using the language model to obtain multiple first reasoning results.
3. The method according to claim 1, characterized in that: The method of using the language model to perform full reasoning and mixed reasoning on the multiple request information segments and the second request information to obtain multiple first reasoning results includes: Using the language model, perform full reasoning, or full reasoning and mixed reasoning, on a first request segment information among the multiple request information segments to obtain first historical result information output by the full reasoning, or the full reasoning and mixed reasoning; In response to obtaining the second request information requesting to use a preset language model to infer the input information, taking the second request fragment information in the multiple request information segments and the second request information as the same batch request to obtain first request batch information; Based on the first historical result information and the first request batch information, hybrid reasoning is performed using the language model to obtain at least one of the first reasoning results and second historical result information; In response to obtaining third request information requesting to use a preset language model to infer input information, taking the first inference result, the third request segment information in the multiple request segments, and the third request information as a same batch request to obtain third request batch information; Based on the second historical result information and the third request batch information, hybrid reasoning is performed using the language model to obtain multiple first reasoning results.
4. The method according to claim 1, characterized in that: The method of using the language model to perform full reasoning and mixed reasoning on the multiple request information segments and the second request information to obtain multiple first reasoning results includes: The multiple segments of request information and the second request information are merged in batches, and zero-redundancy full reasoning and hybrid reasoning are performed on the multiple segments of request information and the second request information using the language model to obtain multiple first reasoning results.
5. The method according to claim 4, characterized in that The step of merging the multiple request information segments and the second request information in batches, and performing zero-redundancy full-volume reasoning and hybrid reasoning on the multiple request information segments and the second request information using the language model to obtain multiple first reasoning results, including: Using the language model, perform full reasoning, or full reasoning and mixed reasoning, on a first request segment information among the multiple request information segments to obtain historical result information output by the full reasoning, or the full reasoning and mixed reasoning; In response to obtaining the second request information requesting to use a preset language model to infer the input information, merging the second request fragment information in the multiple request information segments with the second request information to obtain first merged information; Based on the historical result information and the first merged information, zero-redundancy hybrid reasoning is performed using the language model to obtain multiple first reasoning results.
6. The method according to claim 5, characterized in that The method of performing zero-redundancy hybrid reasoning based on the historical result information and the first merged information using the language model to obtain a plurality of the first reasoning results includes: Performing label embedding and layer normalization processing on the first merged information and the historical result information respectively to obtain first merged input information; Performing a vector linear transformation on the first combined input information using the self-attention layer of the language model to obtain first vector input information; Based on the first vector input information, using the self-attention layer to perform zero-redundancy full-quantity operation to obtain a first output result; Performing full feed-forward processing on the first output result using the feed-forward layer of the language model to obtain a first intermediate result; Based on N first merged input information, first vector input information, first output result and first intermediate result, multiple first inference results are obtained, where N is a natural number greater than 1.
7. The method according to claim 6, characterized in that The method of performing a zero-redundancy full-quantity operation based on the first vector input information and obtaining a first output result by using the self-attention layer includes: Distribute the first vector input information into multiple independent operators in the self-attention layer according to a preset partitioning strategy; Using the multiple independent operators to respectively perform full operations on the first vector input information that is input in a scattered manner, to obtain multiple first operation results, and merging the multiple first operation results into a preset video memory corresponding to the language model; The multiple first operation results merged and written into the preset video memory are split to obtain a first output result.
8. The method according to claim 4, characterized in that The step of merging the multiple request information segments and the second request information in batches, and performing zero-redundancy full-volume reasoning and hybrid reasoning on the multiple request information segments and the second request information using the language model to obtain multiple first reasoning results further includes: In response to obtaining third request information requesting to use a preset language model to infer input information, merging the first inference result and the third request information to obtain second merged information; Based on the second merged information, zero-redundancy hybrid reasoning is performed using the language model to obtain multiple first reasoning results.
9. The method according to claim 4, characterized in that The step of merging the multiple request information segments and the second request information in batches, and performing zero-redundancy full-volume reasoning and hybrid reasoning on the multiple request information segments and the second request information using the language model to obtain multiple first reasoning results, including: Using the language model, perform full reasoning, or full reasoning and mixed reasoning, on a first request segment information among the multiple request information segments to obtain first historical result information output by the full reasoning, or the full reasoning and mixed reasoning; In response to obtaining the second request information requesting to use a preset language model to infer the input information, merging the second request fragment information in the multiple request information segments with the second request information to obtain first merged information; Based on the first historical result information and the first merged information, using the language model to perform zero-redundancy hybrid reasoning to obtain at least one of the first reasoning results and the second historical result information; In response to obtaining third request information requesting to use a preset language model to infer input information, merging the first inference result, the third request segment information in the multiple segments of request information, and the third request information to obtain third merged information; Based on the second historical result information and the third merged information, zero-redundancy hybrid reasoning is performed using the language model to obtain multiple first reasoning results.
Citation Information
Patent Citations
Large language model reasoning optimization method and device, computer equipment and storage medium
CN117194056A
Inference method and device and electronic equipment
CN118036741A