Large Model Inference Method and Server for User Requests
The method improves processing efficiency and user experience by caching historical prefix data in host memory for efficient pre-fill and recursive reasoning, addressing the issue of repeated reasoning in existing GPU memory clearing methods.
Patent Information
- Application Number
- CN202510222616.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-27
- Publication Date
- 2025-07-15
- Estimated Expiration
- 2045-02-27
AI Technical Summary
In the prior art, GPU videos have obtained the inference result of prefix data and cleared the prefix data, resulting in repeated inferences required for subsequent inferences, resulting in too long time and reducing user experience.
By matching the historical prefix key-value cache in host memory, pre-filling and recursive inference processing are performed, and duplicate inference data is avoided, and inference processing is performed using preset language models and knowledge bases.
Improve the inference efficiency and user experience of request text, and cache historical prefix key-value cache through host memory, reducing the time of repeated inference.
Smart Images

Figure CN119721261B_ABST
Abstract
Description
Technical Field
[0001] This application relates to the technical field of large model applications, and in particular to a large model inference method and server for user requests. Background Art
[0002] In recent years, with the rapid development of artificial intelligence and deep learning, users' demand for artificial intelligence large models has become increasingly urgent. A large model refers to a machine learning model with a large number of parameters and a complex computing structure. Especially with the explosion of AI intelligent question answering, the attention of language large models has soared.
[0003] In related technologies, the request text proposed by the user is inferred through the GPU video memory to obtain an inference result including prefix data. However, in related technologies, after obtaining the inference result including prefix data, the GPU video memory will clear the prefix data, resulting in the need to repeat the inference of the prefix data during subsequent inferences, leading to too long inference time and reducing the user experience. Summary of the Invention
[0004] This application provides a large model inference method and server for user requests, so as to at least solve the problem in related technologies that after obtaining the inference result including prefix data, the GPU video memory will clear the prefix data, resulting in the need to repeat the inference of the prefix data during subsequent inferences, leading to too long inference time and reducing the user experience.
[0005] This application provides a large model inference method for user requests, including:
[0006] Receiving a request text sent by a user terminal;
[0007] Performing vector conversion processing on the request text to obtain a corresponding request vector;
[0008] Determining a corresponding target statement from a pre-created knowledge base according to the request vector;
[0009] Inputting the request text and the target statement into a preset language model for inference processing to perform the following steps:
[0010] Judging whether the request text can match a pre-stored historical prefix key-value cache in the host memory, where the historical prefix key-value cache is pre-filled and inferred according to historical request texts;
[0011] If the request text can match the historical prefix key-value cache in the host memory, determining the historical prefix key-value cache as a preliminary pre-filled inference result;
[0012] Perform pre-filling inference processing based on the said request text, the said preliminary pre-filled inference result, and the said target statement to obtain a pre-filled inference result;
[0013] Perform recursive inference processing based on the said pre-filled inference result to obtain a recursive inference result;
[0014] Output the said recursive inference result to the said client.
[0015] This application also provides a server, including: a memory for storing a computer program; a processor for implementing the steps of any of the above-mentioned large model inference methods for user requests when executing the computer program.
[0016] This application also provides a computer-readable storage medium, in which a computer program is stored. Wherein, when the computer program is executed by a processor, it implements the steps of any of the above-mentioned large model inference methods for user requests.
[0017] This application also provides a computer program product, including a computer program, which implements the steps of any of the above-mentioned large model inference methods for user requests when executed by a processor.
[0018] The large model inference method and server for user requests provided by the embodiments of this application perform inference processing by inputting the obtained request text and target statement into a preset language model to execute the following steps: after determining that the request text can match the historical prefix key-value cache in the host memory, determining the historical prefix key-value cache as the preliminary pre-filled inference result; performing pre-filling inference processing based on the request text, the preliminary pre-filled inference result, and the target statement to obtain a pre-filled inference result; performing recursive inference processing based on the pre-filled inference result to obtain a recursive inference result; outputting the recursive inference result to the client. By storing the historical prefix key-value cache in the host memory and not needing to repeatedly infer the corresponding historical prefix key-value cache, the inference efficiency of the request text is improved, and the user experience is also improved. Brief Description of the Drawings
[0019] To more clearly illustrate the embodiments of this application, the following will briefly introduce the drawings required in the embodiments. Obviously, the drawings in the following description are only some embodiments of this application. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.
[0020] Figure 1 It is a schematic diagram of the scenario of the large model inference method for user requests provided by the embodiments of this application;
[0021] Figure 2Flow diagram of the large model inference method for user requests provided by the embodiments of the present application Figure 1 ;
[0022] Figure 3 Flow diagram of the large model inference method for user requests provided by the embodiments of the present application Figure 2 ;
[0023] Figure 4 Structural schematic diagram of the large model inference device for user requests provided by the embodiments of the present application;
[0024] Figure 5 Structural schematic diagram of the server provided by the embodiments of the present application. Detailed implementation manners
[0025] Next, the technical solutions in the embodiments of the present application will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present application. Obviously, the described embodiments are only a part of the embodiments of the present application, rather than all the embodiments. Based on the embodiments in the present application, all other embodiments obtained by those of ordinary skill in the art without creative efforts belong to the protection scope of the present application.
[0026] It should be noted that in the description of the present application, the terms "include", "comprise" or any other variant thereof are intended to cover a non-exclusive inclusion, so that a process, method, article or device including a series of elements not only includes those elements but also includes other elements not expressly listed, or further includes elements inherent to such process, method, article or device. The terms "first", "second", etc. in the present application are used to distinguish similar objects and are not used to describe a specific order or sequence.
[0027] In recent years, with the rapid development of artificial intelligence and deep learning, users' demands for artificial intelligence large models have become increasingly urgent. A large model refers to a machine learning model with a large number of parameters and a complex computational structure. These models are usually constructed by deep neural networks and have billions or even hundreds of billions of parameters, aiming to improve the model's expressive ability and prediction performance to handle more complex tasks and data. Especially with the explosion of popularity of AI intelligent question answering, the attention to language large models has soared. In related technologies, the request text proposed by the user is inferred through the GPU video memory to obtain an inference result including prefix data. However, in related technologies, after obtaining the inference result including prefix data, the GPU video memory will clear the prefix data, resulting in the need to repeatedly infer the prefix data during subsequent inferences, leading to too long inference time and reducing the user experience.
[0028] To solve the above technical problems, the embodiments of the present application propose the following technical concepts: The inventor considers the obtained request text and the corresponding target statement. After determining that the request text can match the historical prefix key-value cache in the host memory, based on a preset language model, pre-population inference processing and recursive inference processing are performed on the historical prefix key-value cache, the request text, and the corresponding target statement to obtain a pre-population inference result and a recursive inference result, and the pre-population inference result and the recursive inference result are output, thereby improving the inference efficiency of the request text and the user experience.
[0029] To enable those skilled in the art of the present technology to better understand the solution of the present application, the following further detailed description of the present application will be given in conjunction with the accompanying drawings and specific embodiments.
[0030] Combined with the specific application environment architecture or specific hardware architecture on which the execution of the large model inference method for user requests depends, the specific application environment architecture or specific hardware architecture will be described herein. Refer to Figure 1 , Figure 1 FIG. is a schematic diagram of the scenario of the large model inference method for user requests provided by the embodiments of the present application.
[0031] As Figure 1 shown, this scenario includes: a user terminal 101 and a server 102.
[0032] Among them, the user terminal 101 can be a mobile phone terminal or a personal computer terminal.
[0033] The server 102 can be an independent server or a cluster composed of multiple servers.
[0034] The user terminal 101 sends the request text to the server 102 through a wireless network. The server 102 performs vector conversion processing on the request text to obtain the corresponding request vector; according to the request vector, determines the corresponding target statement from the pre-created knowledge base; inputs the request text and the target statement into a preset language model for inference processing to perform the following steps: If it is determined that the request text can match the historical prefix key-value cache in the host memory, then the historical prefix key-value cache is determined as the preliminary pre-population inference result; pre-population inference processing is performed according to the request text, the preliminary pre-population inference result, and the target statement to obtain a pre-population inference result; recursive inference processing is performed according to the pre-population inference result to obtain a recursive inference result; and the recursive inference result is output to the user terminal 101.
[0035] The technical solutions of this application and how the technical solutions of this application solve the above technical problems will be described in detail below with specific embodiments. The following several specific embodiments can be combined with each other, and the same or similar concepts or processes may not be repeated in some embodiments. The embodiments of this application will be described below with reference to the accompanying drawings.
[0036] Figure 2 Flow schematic of the large model inference method for user requests provided by the embodiments of this application Figure 1 , such as Figure 2 shown, the embodiments of this application provide a large model inference method for user requests, and the method will be described in detail as follows:
[0037] S201: Receive the request text sent by the user terminal.
[0038] S202: Perform vector transformation processing on the request text to obtain the corresponding request vector.
[0039] Specifically, the request text is subjected to vector transformation processing according to the encoder to obtain the corresponding request vector.
[0040] In this embodiment, the request vector is a way to describe the object features. By transforming the request text into a structured vector, it is convenient for the large model to understand and process this data, and these vectors represent the features of the request text.
[0041] S203: Determine the corresponding target statement from the pre-created knowledge base according to the request vector.
[0042] Specifically, the creation process of the knowledge base in step S203 specifically includes steps a to b:
[0043] Step a: Receive one or more documents sent by the user terminal.
[0044] In this embodiment, the document can be company rules and regulations, notices or other documents.
[0045] Step b: Create a knowledge base according to one or more documents.
[0046] Specifically, step S203 specifically includes:
[0047] S2031: Determine the target document from the knowledge base according to the request vector.
[0048] S2032: Determine the corresponding target statement from the target document according to the request vector.
[0049] In this embodiment, the target statement can be a document fragment, a paragraph or other combined statements in the target document.
[0050] S204: Input the request text and the target statement into a preset language model for inference processing to perform the following steps:
[0051] S205: Determine whether the request text can match a pre-stored historical prefix key-value cache in the host memory, where the historical prefix key-value cache is pre-filled and inferred based on historical request texts.
[0052] Specifically, use the matching algorithm corresponding to the large model inference program to determine whether the request text can match the pre-stored historical prefix key-value cache in the host memory.
[0053] Among them, the L2 cache provides temporary space for the inference program, and the L2 cache belongs to the GPU video memory.
[0054] Among them, the L2 cache is the second-level cache of the CPU, located between the CPU and the main memory, and is usually composed of static random access memory. The capacity of the L2 cache is smaller than that of the memory but the exchange speed is faster. Its main function is to serve as a temporary memory between the CPU and the memory, storing data that the CPU may access in the short term to reduce the number of accesses to the memory, thereby improving the data access speed.
[0055] Among them, the CPU is the central processing unit.
[0056] Specifically, the storage process of the historical prefix key-value cache in step S205 specifically includes steps c to d:
[0057] Step c: Obtain multiple historical prefix key-value caches.
[0058] In this embodiment, the historical prefix key-value cache is the kv cache of the prefix.
[0059] Among them, the kv cache is a technology used to optimize the inference performance of the Transformer model, especially outstanding in autoregressive generation models. Its core function is to cache the Key and Value matrices, avoiding repeated calculation of these matrices when generating new tokens each time, thereby significantly improving the inference efficiency.
[0060] Step d: Store multiple historical prefix key-value caches in the host memory.
[0061] Specifically, step d specifically includes:
[0062] Step d1: Mark each historical prefix matrix data in a preset data form to obtain each marked historical prefix key-value cache.
[0063] In this embodiment, the preset data form can be length, type, keyword, and other data forms.
[0064] Step d2: Generate tree-structured data based on each labeled historical prefix key-value cache.
[0065] In this embodiment, each labeled historical prefix key-value cache corresponds to each tree node data in the tree-structured data.
[0066] Step d3: Store the tree-structured data into the host memory.
[0067] In this embodiment, the host memory is the server-side physical memory.
[0068] Specifically, step S205 specifically includes:
[0069] S2051: Obtain the tree-structured data in the host memory, where the tree-structured data includes one or more tree node data, and each tree node data corresponds to a pre-stored historical prefix key-value cache.
[0070] S2052: Obtain the identification information corresponding to each historical prefix key-value cache.
[0071] In this embodiment, the identification information is length, type, keyword, and other identification information.
[0072] S2053: Determine whether the request text can match the corresponding identification information.
[0073] Specifically, extract the key information of the request text, and determine whether it can match the corresponding identification information according to the key information.
[0074] Among them, the corresponding identification information is matched through real-time inference of the GPU.
[0075] Therefore, combining the above description, by storing the tree-structured data containing historical prefix key-value caches into the host memory, and by matching the identification information corresponding to the historical prefix key-value caches through real-time inference of the GPU, a corresponding hierarchical cache layout is formed.
[0076] Among them, the full name of GPU video memory is Graphics Processing Unit Memory or Video Memory, which is a high-speed memory specifically designed for the GPU, mainly used to store data related to graphics rendering, such as textures, vertex data, and frame buffers, etc. Video memory is usually used in graphics cards or integrated graphics solutions, and its design purpose is to meet the needs of large-scale parallel computing and provide a fast data access channel.
[0077] Among them, GPU is the Graphics Processing Unit.
[0078] S2054: If the request text can match the corresponding identification information, it is determined that the request text can match the historical prefix key-value cache in the host memory; if the request text cannot match the corresponding identification information, it is determined that the request text cannot match the historical prefix key-value cache in the host memory.
[0079] S206: If the request text can match the historical prefix key-value cache in the host memory, the historical prefix key-value cache is determined as the preliminary pre-filled inference result.
[0080] In addition, after step S206, steps e to j are further included:
[0081] Step e: Obtain the usage threshold of the host memory;
[0082] In this embodiment, the usage threshold can be any quantity among 10 times, 100 times, or 1000 times, or it can be other quantities.
[0083] Step f: Count the historical prefix key-value cache to obtain the current count;
[0084] Step g: In the host memory, obtain multiple historical counts;
[0085] Step h: Determine whether the total count of multiple historical counts and the current count is greater than the usage threshold;
[0086] Step i: If it is determined that the total count is greater than the usage threshold, determine the target tree node to be filtered out in the tree structure data;
[0087] In this embodiment, the target tree node to be filtered out can be the tree node with the lowest usage times, or it can be other tree nodes.
[0088] Step j: Filter out the target tree node from the host memory to update the host memory.
[0089] S207: Perform pre-filled inference processing based on the request text, the preliminary pre-filled inference result, and the target statement to obtain the pre-filled inference result.
[0090] Specifically, step S207 specifically includes:
[0091] S2071: Perform pre-filled inference processing based on the request text, the preliminary pre-filled inference result, and the target statement to obtain the prefix key-value cache and the initial token.
[0092] In this embodiment, the pre-filled inference is the inference in the Prefill stage.
[0093] Among them, the prefill stage inference refers to a preprocessing stage before the formal inference of the large model. The main purpose of this stage is to prefill or warm up the model to ensure that the model can respond to requests more quickly during the inference stage.
[0094] In this embodiment, the initial token is a token.
[0095] Among them, a token is a lexical unit, which refers to a symbol used to represent a word or phrase in the process of natural language processing; a token can be a single lexical unit or a sequence composed of multiple token symbols.
[0096] In addition, after step S2071, steps k to l are further included:
[0097] Step k: Decode the initial token to obtain the corresponding string.
[0098] Step l: Output the string to the user side.
[0099] S2072: Determine the prefix key-value cache and the initial token as the prefill inference result.
[0100] S208: Perform recursive inference processing based on the prefill inference result to obtain the recursive inference result.
[0101] Specifically, step S208 specifically includes:
[0102] S2081: Continue to perform inference based on the initial token, recursively infer until the preset end flag is inferred.
[0103] In this embodiment, the recursive inference is the Decode stage inference.
[0104] Among them, the Decode stage inference refers to the process in which, during the inference of the large model, after the model generates the first token, it continues to generate subsequent tokens autoregressively until the stop condition is met, such as generating a specific number of tokens or encountering a specific end marker.
[0105] In this embodiment, the preset end flag is eos_token_id.
[0106] Specifically, take the prefix key-value cache and the initial token as the input cache, continue to perform inference based on the initial token, recursively infer until the preset end flag is inferred.
[0107] Among them, the recursive inference refers to the process of using the previous inference token as the input for the next inference and repeating the inference.
[0108] S2082: Obtain multiple tokens before the last token.
[0109] In addition, decode multiple tokens and the last token to obtain corresponding strings.
[0110] S2083: Determine each token as the recursive inference result.
[0111] S209: Output the recursive inference result to the client.
[0112] Specifically, output the strings corresponding to the recursive inference result to the client.
[0113] In addition, when each string is generated, it can also be output to the client.
[0114] In summary, the large model inference method for user requests provided in this embodiment performs inference processing by inputting the obtained request text and target statement into a preset language model to execute the following steps: after determining that the request text can match the historical prefix key-value cache in the host memory, determine the historical prefix key-value cache as the preliminary pre-filled inference result; perform pre-filled inference processing based on the request text, preliminary pre-filled inference result, and target statement to obtain the pre-filled inference result; perform recursive inference processing based on the pre-filled inference result to obtain the recursive inference result; output the recursive inference result to the client, and store the historical prefix key-value cache in the host memory, without repeatedly inferring the corresponding historical prefix key-value cache, which improves the inference efficiency of the request text and improves the user experience.
[0115] In addition, the large model inference method for user requests provided in this embodiment stores the tree-structured data containing the historical prefix key-value cache in the host memory, and forms a corresponding hierarchical cache layout by real-time inferring and matching the identification information corresponding to the historical prefix key-value cache through the GPU, so that the prefix key-value cache can be cached for a long time without affecting the read and write speed of the GPU, and the arrangement of each storage unit is more reasonable for subsequent inference of request texts.
[0116] Figure 3 Schematic flow of the large model inference method for user requests provided in the embodiments of this application Figure 2 . In the embodiments of this application, based on the embodiments provided, a detailed description is given of the subsequent inference method when the request text cannot match the historical prefix key-value cache in the host memory after step S205. As Figure 2 shown, the method includes: Figure 3 shown, the method includes:
[0117] S301: Receive the request text sent by the client.
[0118] S302: Perform vector transformation processing on the request text to obtain the corresponding request vector.
[0119] In this embodiment, the discussion about the request vector has been described in detail in step S202, and will not be elaborated here.
[0120] S303: Determine the corresponding target statement from the pre-created knowledge base according to the request vector.
[0121] In this embodiment, the discussion about the pre-created knowledge base and the target statement has been described in detail in step S203, and will not be elaborated here.
[0122] S304: Input the request text and the target statement into a preset language model for inference processing to perform the following steps:
[0123] S305: Determine whether the request text can match the pre-stored historical prefix key-value cache in the host memory, where the historical prefix key-value cache is pre-filled and inferred according to the historical request text.
[0124] In this embodiment, the discussion about the host memory and the pre-stored historical prefix key-value cache has been described in detail in step S205, and will not be elaborated here.
[0125] S306: If the request text cannot match the historical prefix key-value cache in the host memory, input the request text and the target statement into the preset language model for pre-filling inference and recursive inference to obtain the corresponding inference result.
[0126] Specifically, step S306 specifically includes:
[0127] S3061: Input the request text and the target statement into the preset language model for pre-filling inference processing to obtain the prefix key-value cache and the first token.
[0128] In this embodiment, the discussion about pre-filling inference and tokens has been described in detail in step S207, and will not be elaborated here.
[0129] S3062: Determine the prefix key-value cache and the first token as the pre-filling inference result.
[0130] S3063: Input the first token into the preset language model for recursive inference until the preset end flag is inferred.
[0131] In this embodiment, the discussion about recursive inference has been described in detail in step S208, and will not be elaborated here.
[0132] S3064: Obtain multiple tokens before the last token.
[0133] S3065: Determine each token as the recursive inference result.
[0134] S3066: Determine the pre-filled inference result and the recursive inference result as the inference result.
[0135] In addition, after step S306, it further includes:
[0136] Integrate the prefix key-value cache corresponding to the request text into the tree-structured data from the inference result.
[0137] Specifically, integrate the prefix key-value cache corresponding to the request text into the tree-structured data from the pre-filled inference result of the inference result.
[0138] S307: Output the inference processing result to the client.
[0139] In summary, for the large model inference method for user requests provided in this embodiment, if the request text cannot match the historical prefix key-value cache in the host memory, the request text and the target statement are input into the preset language model for pre-filled inference and recursive inference to obtain the corresponding inference result; the inference processing result is output to the client, and the prefix key-value cache corresponding to the request text is integrated into the tree-structured data, so that the tree-structured data in the host memory can be supplemented in time for storage as the corresponding historical prefix key-value cache.
[0140] Through the description of the above embodiments, those skilled in the art can clearly understand that the method according to the above embodiments can be implemented by means of software plus the necessary general hardware platform, and of course, it can also be implemented by hardware, but in many cases, the former is a better implementation method.
[0141] Figure 4 It is a schematic structural diagram of the large model inference device for user requests provided in the embodiments of the present application. As Figure 4 shown, the embodiments of the present application also provide a large model inference device for user requests, including: a first receiving module 401, a first processing module 402, a first determining module 403, a second processing module 404, a first judging module 405, a second determining module 406, a third processing module 407, a fourth processing module 408, and a first output module 409.
[0142] The first receiving module 401 is configured to receive the request text sent by the client;
[0143] The first processing module 402 is configured to perform vector conversion processing on the request text to obtain the corresponding request vector;
[0144] The first determination module 403 is configured to determine a corresponding target statement from a pre-created knowledge base according to the request vector;
[0145] The second processing module 404 is configured to input the request text and the target statement into a preset language model for inference processing to perform the following steps:
[0146] The first judgment module 405 is configured to judge whether the request text can match a pre-stored historical prefix key-value cache in the host memory, where the historical prefix key-value cache is pre-filled and inferred according to historical request texts;
[0147] The second determination module 406 is configured to, if the request text can match a historical prefix key-value cache in the host memory, determine the historical prefix key-value cache as a preliminary pre-filled inference result;
[0148] The third processing module 407 is configured to perform pre-filling inference processing according to the request text, the preliminary pre-filled inference result, and the target statement to obtain a pre-filled inference result;
[0149] The fourth processing module 408 is configured to perform recursive inference processing according to the pre-filled inference result to obtain a recursive inference result;
[0150] The first output module 409 is configured to output the recursive inference result to the client.
[0151] In a possible implementation manner, the judgment module 405 specifically includes:
[0152] The first acquisition unit is configured to acquire tree structure data in the host memory, where the tree structure data includes one or more tree node data, and each tree node data corresponds to a pre-stored historical prefix key-value cache;
[0153] The second acquisition unit is configured to acquire identification information corresponding to each historical prefix key-value cache;
[0154] The judgment unit is configured to judge whether the request text can match the corresponding identification information;
[0155] The determination unit is configured to, if the request text can match the corresponding identification information, determine that the request text can match a historical prefix key-value cache in the host memory; if the request text cannot match the corresponding identification information, determine that the request text cannot match a historical prefix key-value cache in the host memory.
[0156] In a possible implementation manner, the third processing module 407 specifically includes:
[0157] A processing unit, configured to perform pre-filling inference processing according to the request text, the preliminary pre-filling inference result, and the target statement, so as to obtain a prefix key-value cache and initial tokens;
[0158] A determination unit, configured to determine the prefix key-value cache and the initial tokens as the pre-filling inference result.
[0159] In a possible implementation manner, the apparatus further includes:
[0160] A fifth processing module, configured to perform decoding processing on the initial tokens to obtain a corresponding string;
[0161] A second output module, configured to output the string to the client.
[0162] In a possible implementation manner, the fourth processing module 408 specifically includes:
[0163] An inference unit, configured to continue performing inference based on the initial tokens, recursively inferring until a preset end identifier is inferred;
[0164] An acquisition unit, configured to acquire multiple tokens before the last token;
[0165] A determination unit, configured to determine each token as the recursive inference result.
[0166] In a possible implementation manner, the apparatus further includes:
[0167] A first acquisition module, configured to acquire multiple historical prefix key-value caches;
[0168] A storage module, configured to store the multiple historical prefix key-value caches in the host memory.
[0169] In a possible implementation manner, the storage module specifically includes:
[0170] An annotation unit, configured to annotate each historical prefix matrix data in a preset data form to obtain each annotated historical prefix key-value cache;
[0171] A generation unit, configured to generate tree structure data according to the annotated historical prefix key-value caches;
[0172] A storage unit, configured to store the tree structure data in the host memory.
[0173] In a possible implementation manner, the apparatus further includes:
[0174] The sixth processing module is configured to, if the request text cannot match the historical prefix key-value cache in the host memory, input the request text and the target statement into the preset language model for pre-filling inference and recursive inference to obtain corresponding inference results;
[0175] The third output module is configured to output the inference processing result to the client.
[0176] In a possible implementation manner, the device further includes:
[0177] The integration module is configured to integrate the prefix key-value cache corresponding to the request text into the tree structure data from the inference results.
[0178] In a possible implementation manner, the device further includes:
[0179] The second acquisition module is configured to acquire the usage threshold of the host memory;
[0180] The counting module is configured to count the historical prefix key-value cache to obtain the current count;
[0181] The third acquisition module is configured to acquire multiple historical counts in the host memory;
[0182] The second judgment module is configured to judge whether the total count of the multiple historical counts and the current count is greater than the usage threshold;
[0183] The third determination module is configured to, if it is determined that the total count is greater than the usage threshold, determine the target tree node to be filtered in the tree structure data;
[0184] The filtering module is configured to filter the target tree node from the host memory to update the host memory.
[0185] In a possible implementation manner, the device further includes:
[0186] The second receiving module is configured to receive one or more documents sent by the client;
[0187] The creation module is configured to create a knowledge base according to the one or more documents.
[0188] In a possible implementation manner, the first determination module 403 specifically includes:
[0189] The first determination unit is configured to determine a target document from the knowledge base according to the request vector;
[0190] The second determination unit is configured to determine a corresponding target statement from the target document according to the request vector.
[0191] For the description of the features in the corresponding embodiments of the large model inference device for user requests, reference can be made to the relevant descriptions in the corresponding embodiments of the large model inference method for user requests, which will not be elaborated here one by one.
[0192] Figure 5 This is a schematic structural diagram of the server provided by this application. As Figure 5 shown, the server provided in this embodiment includes: at least one processor 501 and a memory 502. Optionally, the server further includes a communication component 503. Among them, the processor 501, the memory 502, and the communication component 503 are connected through a bus.
[0193] In the specific implementation process, at least one processor 501 executes the computer execution instructions stored in the memory 502, so that at least one processor 501 executes the above-mentioned embodiment of the large model inference method for user requests.
[0194] For the specific implementation process of the processor 501, reference can be made to the above method embodiment, and its implementation principle and technical effects are similar, which will not be elaborated here in this embodiment.
[0195] In the above embodiment, it should be understood that the processor may be a central processing unit (Central Processing Unit, abbreviated as: CPU), or other general-purpose processors, digital signal processors (Digital Signal Processor, abbreviated as: DSP), application specific integrated circuits (Application Specific Integrated Circuit, abbreviated as: ASIC), etc. The general-purpose processor may be a microprocessor or the processor may also be any conventional processor, etc. The steps of the method disclosed in combination with the application can be directly implemented by the hardware processor, or implemented by a combination of hardware and software modules in the processor.
[0196] The memory may include high-speed memory (Random Access Memory, RAM), and may also include non-volatile memory (Non-volatile Memory, NVM), such as at least one disk memory.
[0197] The bus can be an Industry Standard Architecture (ISA) bus, a Peripheral Component Interconnect (PCI) bus, an Extended Industry Standard Architecture (EISA) bus, etc. The bus can be divided into an address bus, a data bus, a control bus, etc. For ease of representation, the buses in the drawings of this application are not limited to only one bus or one type of bus.
[0198] An embodiment of the present application also provides a computer-readable storage medium, in which a computer program is stored. Wherein, the computer program is configured to execute the steps in any of the above embodiments of the large model inference method for user requests when running.
[0199] In an exemplary embodiment, the above computer-readable storage medium may include, but is not limited to: USB flash drives, read-only memories (ROMs), random access memories (RAMs), external hard drives, magnetic disks, or optical discs and other media that can store computer programs.
[0200] An embodiment of the present application also provides a computer program product. The above computer program product includes a computer program, and when the computer program is executed by a processor, it implements the steps in any of the above embodiments of the large model inference method for user requests.
[0201] An embodiment of the present application also provides another computer program product, including a non-volatile computer-readable storage medium. The non-volatile computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, it implements the steps in any of the above embodiments of the large model inference method for user requests.
[0202] Those skilled in the art can further realize that the units and algorithm steps of each example described in combination with the embodiments disclosed herein can be implemented by electronic hardware, computer software, or a combination of the two. To clearly illustrate the interchangeability of hardware and software, the composition and steps of each example have been generally described according to functions in the above description. Whether these functions are executed in a hardware or software manner depends on the specific application and design constraints of the technical solution. Skilled professionals can use different methods to implement the described functions for each specific application, but such implementation should not be considered to exceed the scope of this application.
[0203] The above has introduced in detail a large model inference method for user requests provided by this application. Specific examples are used in this article to elaborate on the principle and implementation manner of this application. The description of the above embodiments is only used to help understand the method of this application and its core idea. It should be noted that for those of ordinary skill in the art, without departing from the principle of this application, several improvements and modifications can be made to this application, and these improvements and modifications also fall within the protection scope of the claims of this application.
Claims
1. A large model inference method for user requests, characterized in that, Including: Receiving a request text sent by a user terminal; Performing vector transformation processing on the request text to obtain a corresponding request vector; Determining a corresponding target statement from a pre-created knowledge base according to the request vector; Inputting the request text and the target statement into a preset language model for inference processing to perform the following steps: Judging whether the request text can match a pre-stored historical prefix key-value cache in the host memory, where the historical prefix key-value cache is pre-filled and inferred according to historical request texts; the host memory stores tree-structured data generated from the labeled historical prefix key-value cache; If the request text can match the historical prefix key-value cache in the host memory, determining the historical prefix key-value cache as a preliminary pre-filled inference result; Performing pre-filled inference processing according to the request text, the preliminary pre-filled inference result, and the target statement to obtain a pre-filled inference result; Performing recursive inference processing according to the pre-filled inference result to obtain a recursive inference result; Outputting the recursive inference result to the user terminal; The judging whether the request text can match a pre-stored historical prefix key-value cache in the host memory includes: Obtaining the tree-structured data in the host memory, where the tree-structured data includes one or more tree node data, and each tree node data corresponds to a pre-stored historical prefix key-value cache; Obtaining the identification information corresponding to each historical prefix key-value cache; judging whether the request text can match the corresponding identification information.
2. The large model inference method for user requests according to claim 1, wherein The judging whether the request text can match a pre-stored historical prefix key-value cache in the host memory includes: If the request text can match the corresponding identification information, determining that the request text can match the historical prefix key-value cache in the host memory; if the request text cannot match the corresponding identification information, determining that the request text cannot match the historical prefix key-value cache in the host memory.
3. The large model inference method for user requests according to claim 1, wherein, The performing pre-filled inference processing according to the request text, the preliminary pre-filled inference result, and the target statement to obtain a pre-filled inference result includes: Performing pre-filled inference processing according to the request text, the preliminary pre-filled inference result, and the target statement to obtain a prefix key-value cache and an initial token; Determining the prefix key-value cache and the initial token as the pre-filled inference result.
4. The large model inference method for user requests according to claim 3, characterized in that, After the performing pre-filled inference processing according to the request text, the preliminary pre-filled inference result, and the target statement to obtain a prefix key-value cache and an initial token, it further includes: Performing decoding processing on the initial token to obtain a corresponding string; Outputting the string to the user terminal.
5. The large model inference method for user requests according to claim 3, wherein, The performing recursive inference processing according to the pre-filled inference result to obtain a recursive inference result includes: Continuing to perform inference and recursive inference according to the initial token until a preset end identifier is inferred; Obtaining multiple tokens before the last token; Determining each token as the recursive inference result.
6. The large model inference method for user requests according to claim 1, wherein The storage process of the historical prefix key-value cache includes: Obtain multiple historical prefix key-value caches; Store the multiple historical prefix key-value caches in the host memory.
7. The large model inference method for user requests according to claim 6, wherein The storing the multiple historical prefix key-value caches in the host memory includes: Label each historical prefix matrix data in a preset data form to obtain each labeled historical prefix key-value cache; Generate tree-structured data according to the labeled historical prefix key-value caches; Store the tree-structured data in the host memory.
8. The large model inference method for user requests according to claim 1, wherein, After determining whether the request text can match a pre-stored historical prefix key-value cache in the host memory, it further includes: If the request text cannot match a historical prefix key-value cache in the host memory, input the request text and the target statement into the preset language model for pre-filling inference and recursive inference to obtain corresponding inference results; Output the inference processing result to the client.
9. The large model inference method for user requests according to claim 8, wherein, After inputting the request text and the target statement into the preset language model for pre-filling inference and recursive inference to obtain corresponding inference results, it further includes: Integrate the prefix key-value cache corresponding to the request text into the tree-structured data from the inference results.
10. The large model inference method for user requests according to claim 1, wherein After determining the historical prefix key-value cache as the preliminary pre-filling inference result if the request text can match a historical prefix key-value cache in the host memory, it further includes: Obtain the usage threshold of the host memory; Count the historical prefix key-value cache to obtain the current count; Obtain multiple historical counts in the host memory; Judge whether the total count of the multiple historical counts and the current count is greater than the usage threshold; If it is determined that the total count is greater than the usage threshold, determine the target tree nodes to be filtered in the tree-structured data; Filter the target tree nodes from the host memory to update the host memory.
11. The large model inference method for user requests according to claim 1, wherein The creation process of the knowledge base includes: Receive one or more documents sent by the client; Create a knowledge base according to the one or more documents.
12. The large model inference method for user requests according to claim 11, wherein, The determining the corresponding target statement from the pre-created knowledge base according to the request vector includes: Determine the target document from the knowledge base according to the request vector; Determine the corresponding target statement from the target document according to the request vector.
13. A server, characterized in that, It includes: A memory for storing a computer program; A processor for implementing the steps of the large model inference method for user requests according to any one of claims 1 to 12 when executing the computer program.
14. A computer-readable storage medium, characterized in that, A computer program is stored in the computer-readable storage medium, wherein the computer program implements the steps of the large model inference method for user requests according to any one of claims 1 to 12 when executed by a processor.
15. A computer program product, comprising a computer program, characterized in that, The computer program implements the steps of the large model inference method for user requests according to any one of claims 1 to 12 when executed by a processor.
Citation Information
Patent Citations
Large language reasoning system and method
CN118052282A
Vertical domain model reasoning acceleration method and device
CN118246551A