Method and device for reasoning text data by using generative language model, equipment and storage medium

By introducing a wait queue and scheduler into the generative language model, it handles according to the priority of user requests, solving the problem of insufficient resource utilization when processing user requests, and improving the user experience.

CN120069071APending Publication Date: 2025-05-30SHANGHAI INFINIGENCE AI INTELLIGENT TECHNOLOGY CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510129874.5
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-02-05
Publication Date
2025-05-30

AI Technical Summary

Technical Problem

The generative language model lacks scheduling functions when processing user requests and cannot effectively prioritize requests, resulting in insufficient utilization of inference resources and affecting user experience.

Method used

By introducing a waiting queue and a scheduler, requests from multiple users are received, requests are prioritized based on the user's payment level, number of word elements of text data and waiting time, and requests are received and processed dynamically asynchronously, and requests with high priority are processed.

Benefits of technology

The generative language model is implemented to process according to the priority of user requests, efficiently utilize inference resources, and improve user experience.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120069071A_ABST
    Figure CN120069071A_ABST
Patent Text Reader

Abstract

The invention discloses a method and device for reasoning text data by using a generative language model, equipment and a storage medium. The method comprises: receiving a plurality of requests for reasoning from a plurality of users, and adding the plurality of requests to a waiting queue, each of the plurality of requests comprising text data; determining the priority of each request in the plurality of requests according to the payment level of the user of each request in the plurality of requests, the lexical element number of the text data and the waiting time; selecting one or more requests with the highest priority from the waiting queue as requests to be processed; a generative language model is caused to infer the textual data of the request to be processed.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of artificial intelligence technology, and specifically to a method, apparatus, device, and storage medium for inferring text data using a generative language model. Background Art

[0002] With the development of large language models, generative language models (GLMs) such as ChatGPT (Chat Generative Pre-trained Transformer) and Gemini have attracted more and more users. The surge in the number of users has led to the need to carefully consider computing power allocation and scheduling issues when the model provides services. The serving module for large language models (LLMs) such as generative language models has emerged to deploy and run the model to process user requests.

[0003] A generative language model is a model that can provide a dialogue mode service. It can receive the text input by the user and translate it into a vector that the model can understand, and generate text that the user can understand through inference. For example, users can use the generative language model to complete some question-and-answer tasks or search tasks.

[0004] The content described in this background art is only for facilitating the understanding of the relevant technologies in this field and is not regarded as an admission of the prior art. Summary of the Invention

[0005] However, the generative language models known to the inventors of this application usually process user requests in a single-batch or multi-batch inference manner, and do not have a scheduling function like an operating system, that is, they do not prioritize the received user requests and consider whether the user requests can be processed in parallel. Sometimes, the inference resources may not be efficiently utilized, affecting the user experience.

[0006] Therefore, this application intends to provide a method, apparatus, device, and storage medium for inferring text data using a generative language model, to provide an efficient and reasonable scheduling mechanism for the generative language model, so that the generative language model can process the corresponding text data according to the priority of the user request, efficiently utilize the inference resources, and improve the user experience.

[0007] In a first aspect, the present application provides a method for reasoning on text data using a generative language model, characterized in that the method includes: receiving a plurality of requests for reasoning from a plurality of users and adding the plurality of requests to a waiting queue, each request in the plurality of requests including text data; determining the priority of each request in the plurality of requests according to the payment level of the user of each request, the number of tokens of the text data, and the waiting time; selecting one or more requests with the highest priority from the waiting queue as the requests to be processed; and causing the generative language model to perform reasoning on the text data of the requests to be processed.

[0008] In addition, the method further includes: asynchronously receiving the plurality of requests dynamically, and adding each received request to the waiting queue; for each request, after the generative language model finishes reasoning on the text data of the request, removing the request from the waiting queue and returning the reasoning result of the request to the user, wherein the reception, reasoning, and return of the reasoning results of different requests are performed in parallel.

[0009] In addition, selecting one or more requests with the highest priority from the waiting queue as the requests to be processed includes: selecting one or more requests with the highest priority and received earlier from the waiting queue as the requests to be processed.

[0010] In addition, selecting one or more requests with the highest priority from the waiting queue as the requests to be processed includes: checking whether there are two or more requests in the request with the highest priority whose number of tokens of the text data is less than or equal to a first threshold and the difference in the number of tokens of the text data between them is less than or equal to a second threshold; and in the case where the result of the check is yes, splicing the two or more requests into a batch and selecting the batch as the request to be processed.

[0011] In addition, the reasoning sequentially includes a prefill step and a plurality of decoding steps, and the method further includes: for each request in the waiting queue, feeding back the reasoning state of the request, the reasoning state being prefill or decoding, indicating the prefill step or decoding step that the request is about to perform for reasoning; and when performing the splicing, preferentially splicing two or more requests with the reasoning state of prefill into a batch.

[0012] In addition, determining the priority of each request in the plurality of requests according to the payment level of the user of each request, the number of tokens of the text data, and the waiting time includes: calculating the priority of each request using the following formula:

[0013] P = g / l × t w

[0014] Wherein, P represents the priority of the request, g represents the payment level of the user making the request, l represents the number of tokens in the text data of the request, and t w represents the waiting time of the request.

[0015] In addition, after selecting one or more requests with the highest priority from the waiting queue as the requests to be processed, before enabling the generative language model to perform inference on the text data of the requests to be processed, the method further includes: according to the number of tokens in the text data of each request, allocating one or more storage spaces for KV caching to each request to be processed, and using a block table to record the numbers of the one or more storage spaces, wherein the one or more storage spaces can be discontinuous.

[0016] In a second aspect, the present application provides an apparatus for performing inference on text data using a generative language model, characterized in that the apparatus includes: a request receiver, configured to receive a plurality of requests for inference from a plurality of users and add the plurality of requests to a waiting queue managed by a scheduler, each request in the plurality of requests including text data; the scheduler, configured to determine the priority of each request in the plurality of requests according to the payment level of the user of each request, the number of tokens in the text data, and the waiting time, and select one or more requests with the highest priority from the waiting queue as the requests to be processed; and a generative language model, configured to perform inference on the text data of the requests to be processed.

[0017] In a third aspect, the present application further provides an electronic device, including a memory, a processor, and a computer program stored on the memory, characterized in that the processor executes the computer program to implement the method for performing inference on text data using a generative language model according to an embodiment of the present application.

[0018] In a fourth aspect, the present application further provides a computer-readable storage medium, on which a computer program is stored, and when the computer program is executed by a processor, it implements the method for performing inference on text data using a generative language model according to an embodiment of the present application.

[0019] In a fifth aspect, the present application further provides a computer program product, including computer instructions, and when the computer instructions are executed by a processor, it implements the method for performing inference on text data using a generative language model according to an embodiment of the present application.

[0020] In the method for reasoning text data using a generative language model according to an embodiment of the present application, a waiting queue for user requests is set, and the priority of a request is determined based on the user's payment level, the number of tokens in the text data, and the waiting time, enabling the generative language model to start reasoning from the request with the highest priority. This provides an efficient and reasonable scheduling mechanism for the generative language model, allowing it to process the corresponding text data according to the priority of user requests, efficiently utilize the reasoning resources, and enhance the user experience.

[0021] Some of the other optional features and technical effects of the embodiments of the present application are described below, and some can be understood by reading this text. BRIEF DESCRIPTION OF THE DRAWINGS

[0022] Hereinafter, embodiments of the present application will be described in detail with reference to the drawings. The elements shown are not limited by the scale shown in the drawings, and the same or similar reference numerals in the drawings represent the same or similar elements, where:

[0023] Figure 1 An exemplary flowchart of a method for reasoning text data using a generative language model according to an embodiment of the present application is shown;

[0024] Figure 2 A schematic diagram of storage space allocation according to an embodiment of the present application is shown;

[0025] Figure 3 An exemplary structural diagram of an apparatus for reasoning text data using a generative language model according to an embodiment of the present application is shown;

[0026] Figure 4 An exemplary structural diagram of a terminal device capable of implementing the method according to an embodiment of the present application is shown. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0027] To make the objectives, technical solutions, and advantages of the present application more clear and understandable, the present application will be further described in detail below with reference to the specific embodiments and the drawings. Here, the illustrative embodiments of the present application and their descriptions are used to explain the present application, but not to limit the present application.

[0028] The term "including" and its variants used herein mean open inclusion, that is, "including but not limited to". Unless otherwise stated, the term "or" means "and / or". The term "based on" means "at least partially based on". The term "an exemplary embodiment" and "an embodiment" mean "at least one exemplary embodiment". The term "another embodiment" means "at least one additional embodiment". The terms "first", "second", etc. may refer to different or the same objects. There may be other explicit and implicit definitions below.

[0029] To address the problem in the related art that the generative language model does not prioritize the received user requests and consider whether the user requests can be processed in parallel, the embodiments of the present application provide a method for reasoning about text data using a generative language model. As Figure 1 shown, the method may include the following steps.

[0030] Step S1: Receive multiple requests for reasoning from multiple users, and add the multiple requests to a waiting queue, where each request in the multiple requests includes text data.

[0031] The user may be, for example, a free user or a paid user of the generative language model. The user's request for reasoning may be input, for example, by entering a prompt in the graphical user interface or application at the front end of the generative language model. The prompt may be, for example, a question that the user wants to search through the generative language model, or a dialogue segment between the user and the generative language model. After the user enters the prompt, a request for reasoning can be generated for the user to request the generative language model to reason about the text data (the prompt entered by the user) included in the request to obtain the text output expected by the user, such as the answer to the question searched by the user or a dialogue segment with the user.

[0032] In some embodiments, for example, the request receiver 301 of the device for reasoning about text data using a generative language model as Figure 3 shown may be used to receive multiple requests for reasoning from multiple users and add the received multiple requests to the waiting queue. The reception of the requests and the addition to the waiting queue may be asynchronous, that is, dynamically and asynchronously receive multiple requests for reasoning from multiple users, and add each received request to the waiting queue. "Dynamically and asynchronously receive" is because user input is usually asynchronous, and in the case of real-time receiving of user input, the reception of requests corresponding to the user input is also dynamic and asynchronous. The waiting queue for requests may be managed, for example, by Figure 3 the scheduler 302 of the device for reasoning about text data using a generative language model as

[0033] After receiving a request for inference from a user, for example, the text data of the request can be tokenized to convert it into tokens. Exemplarily, for the text data "I love natural language processing", the tokens obtained after tokenization can be: "I", "love", "natural", "language", "processing". The request can be added to the waiting queue after tokenizing the text data, or the text data of the request can be tokenized after adding it to the waiting queue.

[0034] Step S2: Determine the priority of each request among the multiple requests according to the payment level of the user of each request, the number of tokens of the text data, and the waiting time.

[0035] In some embodiments, after adding a request to the waiting queue, the priority of the request can be determined. For example, it can be determined by Figure 3 scheduler 302 according to the payment level of the user, the number of tokens of the text data, and the waiting time, the priority of the requests in the waiting queue it manages. The payment level of the user can be the level of the fee paid by the user of the generative language model for the services provided by the generative language model. For example, different payment levels can be set according to user permissions or the number of uses. For example, level 0 without payment, level 1 with payment M1, level 2 with payment M2 (M2 > M1), etc. The higher the level, the higher the user permissions or the more the number of uses. The number of tokens of the text data is the number of tokens obtained after tokenizing the text data included in the request, indicating the length of the request. The waiting time is the time elapsed since the user entered the prompt or since the request receiver 301 received the request from the user. For example, scheduler 302 can calculate the priority of each request using the following formula:

[0036] P = g / l × t w

[0037] where P represents the priority of the request, g represents the payment level of the user of the request, l represents the number of tokens of the text data of the request, and t w represents the waiting time of the request. According to this formula, the higher the payment level of the user, the shorter the length of the request (the fewer the number of tokens of the text data), and the longer the waiting time, the higher the priority of the request and the more likely it is to be processed first.

[0038] Step S3: Select one or more requests with the highest priority from the waiting queue as the requests to be processed.

[0039] In some embodiments, this step may be performed, for example, by the scheduler 302. For example, the scheduler 302 may receive one or more previous requests from the requests with the highest priority selected from the waiting queue it manages as the requests to be processed. In this way, when the priorities are different, the requests with higher priorities are processed first, and when the priorities are the same, the requests received earlier are also preferentially processed. Thus, the requests in the waiting queue can be processed both according to the priority and the first-come-first-served criterion, improving the processing efficiency of the requests and enhancing the user experience.

[0040] In some embodiments, when selecting requests to be processed from the waiting queue, for example, it may be checked whether there are two or more requests in the requests with the highest priority (or also the earliest received at the same time) whose number of tokens of the text data is less than or equal to the first threshold and the difference in the number of tokens of the text data between them is less than or equal to the second threshold. If the check result is that there are such requests, the two or more requests that meet this condition are concatenated into a batch, and this batch is selected as the request to be processed. The purpose of this is to concatenate two or more shorter requests with approximate lengths into a batch so that the generative language model can process these requests simultaneously (start the inference for multiple requests in a batch at the same time), thereby reducing the idle computing power and efficiently utilizing the inference resources. For example, when the first threshold is set to 10 and the second threshold is set to 2, two requests in the request with the highest priority in the waiting list whose text data is less than or equal to 10 tokens and differ by within 2 tokens from each other can be concatenated into a batch. When concatenating, the shorter requests can also be padded to align them with the longest request in the batch. Exemplarily, in the waiting list, assume there is a request with 7 tokens of text data and a request with 9 tokens of text data among the requests with the highest priority. At this time, 2 tokens representing blanks (such as -1) can be padded at the end of the text data of the request with 7 tokens of text data to make it also have a length of 9 tokens, and these two requests are concatenated into a 2×9 batch and sent to the generative language model for inference together. For longer requests (such as those with the number of tokens of text data greater than the first threshold) or requests with a large difference in length (such as the difference in the number of tokens of text data between them being greater than the second threshold), considering the processing ability of the generative language model and the computational redundancy, they may not be concatenated and are treated as a batch separately.

[0041] In some embodiments, the selected requests to be processed can be marked by changing the processing status of the requests in the waiting queue. For example, the initial status of the requests in the waiting queue can be "waiting", while the processing status of the selected requests to be processed can be changed to "running". Each time the generative language model performs inference, it can obtain the requests with the processing status of "running" for inference. Note that each request does not require only one inference to output the result expected by the user, but rather requires multiple inferences by the generative language model to return the final result.

[0042] In some embodiments, to avoid wasting the storage space for the KV cache (key-value cache), after selecting one or more requests with the highest priority from the waiting queue as the requests to be processed, before enabling the generative language model to perform inference on the text data of the requests to be processed, for example, referring to the PagedAttention algorithm, according to the number of tokens in the text data of each request to be processed, allocate one or more consecutive or discontinuous blocks of storage space for KV cache for each request, and use a block table to record the numbers of the one or more allocated blocks of storage space. The KV cache is used to store the keys and values generated by the self-attention mechanism of each layer of the model during the inference process of the large language model so that the keys and values can be used for subsequent calculations of each layer. In the embodiments of the present application, allocating one or more consecutive or discontinuous blocks of storage space for KV cache for each request according to the number of tokens in the text data of each request to be processed before enabling the generative language model to perform inference on the text data of the requests to be processed can avoid wasting the storage space for the KV cache and more effectively manage and save the storage resources of the video memory.

[0043] For example, for each block of storage space in the video memory that can be used as a KV cache, an independent number can be assigned to this block of storage space for each layer of self-attention mechanism in the generative language model. After selecting one or more requests with the highest priority from the waiting queue as the requests to be processed, before enabling the generative language model to perform inference on the text data of the requests to be processed, according to the number of tokens in the text data of the requests, one or more continuous or discontinuous blocks of storage space for the KV cache are allocated to the requests. A block table record can be used to associatively record each request and the numbers of one or more blocks of storage space allocated to it. Since for a generative language model, the length of the input data for each layer is the same, the size of the KV cache for each layer is the same, and the block tables for the KV caches of each layer can also be the same. Thus, by using a single block table recorded before inference, the KV caches for the inference of each layer of the entire model can be managed. Since one or more blocks of storage space managed in the block table for each request can be discontinuous, the utilization rate of the video memory can be improved. The total size of the storage space for the KV cache that can be allocated to all requests to be processed by the model for one inference, or in other words, the total size of the storage space for the KV cache that can be allocated to each layer of the model, can be obtained by preloading the part of the model excluding the KV cache to calculate the remaining video memory, then multiplying by the expected utilization rate of the graphics processing unit (GPU), and then dividing by the number of model layers. The allocation of storage space and the management of the block table can be implemented, for example, by Figure 3 the unshown video memory manager of the device in

[0044] In one example, as Figure 2 shown, assume that the generative language model has 32 layers. The numbers of the storage space for the KV cache (KV_Cache) that can be allocated to all requests to be processed by the model for one inference are from 1 to 7004 (1 to 9 and 7001 to 7004 are shown in the figure). The size of each block of storage space, block_size, is 16, indicating that 16 tokens can be stored. The text data of user request request0 is 59 tokens, and 4 blocks of storage space numbered from 1 to 4 are allocated. Assume that the text data of user request request1 is 30 tokens, and 2 blocks of storage space numbered from 5 to 6 are allocated. Assume that the text data of user request request2 is 10 tokens, and 1 block of storage space numbered 7 is allocated. Assume that the text data of user request request3 is 25 tokens, and 2 blocks of storage space numbered from 7001 to 7002 are allocated. Figure 2 The block table block_table in

[0045] records the numbers of the storage space allocated to each request, where 0 indicates that the storage space is not allocated.

[0046] In some embodiments, enabling the generative language model to perform inference on the text data of the requests to be processed means enabling the generative language model to simultaneously start performing inference on the text data of one or more requests to be processed (or a batch formed by splicing two or more requests) selected through previous steps, that is, inputting the text data of the requests to be processed into the generative language model simultaneously. The generative language model can be, for example, Figure 3 the generative language model 303 as shown. The inference of the model includes a prefill step and multiple decode steps in sequence. For each request, the model generates one token each time it performs inference. The first inference to generate the first token is the prefill step of the inference, and each subsequent inference autoregressively generates more tokens one by one, which are the multiple decode steps of the inference. In some embodiments, for each request in the waiting queue, the inference state of the request can be fed back. The inference state is prefill or decode, indicating that the request is about to perform the prefill step or the decode step of the inference. For example, for an unprocessed request in the waiting list, first the prefill step to be performed in the inference, so its inference state is prefill. For a request that has entered the generative language model 303 for at least one inference and returned at least one token, the inference state can become decode, indicating that the request is about to perform the decode step of the inference. The initial inference state of the request added to the waiting queue can be set to prefill. Each time the generative language model 303 performs an inference, it feeds back the request processed in this inference to the scheduler 302, and the scheduler 302 can thereby set the inference state of the request processed in this inference to decode.

[0047] In this case, when splicing the requests to be processed as described above, two or more requests with the inference state of prefill can be preferentially spliced into a batch. Thus, it is possible to enable the generative language model to preferentially process requests that have not been inferred. As described above, for each request, multiple inferences of the generative language model are required to output the final inference result. Since the waiting queue is dynamically updated as described above, the requests to be processed may also be dynamically updated. The generative language model can obtain the requests to be processed from the waiting queue at the start of each inference and perform corresponding inference steps according to the inference state of the requests.

[0048] For each request, after the generative language model finishes reasoning on the text data of the request (i.e., finishes all multiple inferences for the request), the request can be removed from the waiting queue, and the inference result of the request can be returned to the user. Thus, whether the request is processed individually or spliced into a batch for processing, the generative language model can return its inference result in a timely manner after the inference for the request is completed, without waiting for the inferences of all requests in a batch to be completed before returning the inference results of the requests in that batch. For an entire apparatus that uses a generative language model to perform inference on text data, such as including a request receiver 301, a scheduler 302, and a generative language model 303, the reception, inference, and return of inference results for different requests are asynchronous and parallel. "Asynchronous" means that for different requests, they may not be received simultaneously due to different user input times; even if received simultaneously, they may not start processing simultaneously due to different priorities; even if they start processing simultaneously, they may not return inference results simultaneously due to different numbers of inferences. "Parallel" means that in the apparatus that uses a generative language model to perform inference on text data, at the same moment, requests are being received, the model is being inferred, and inference results are being returned. Thus, model scheduling can be carried out efficiently and reasonably, making full use of inference resources and enhancing the user experience.

[0049] Figure 3 Shows an exemplary structural diagram of an apparatus for performing inference on text data using a generative language model according to an embodiment of the present application. As mentioned above, Figure 3 The shown apparatus includes a request receiver 301, a scheduler 302, and a generative language model 303. The request receiver 301 is configured to receive multiple requests for inference from multiple users and add the multiple requests to a waiting queue managed by the scheduler, where each request in the multiple requests includes text data. The scheduler 302 is configured to determine the priority of each request in the multiple requests according to the payment level of the user of each request, the number of tokens of the text data, and the waiting time, and select one or more requests with the highest priority from the waiting queue as the requests to be processed. The generative language model 303 is configured to perform inference on the text data of the requests to be processed. In addition, the apparatus may further include a video memory manager (not shown) for allocating the storage space of the KV cache for the request and managing the block table. Note that the request receiver 301, the scheduler 302, and the video memory manager (not shown) can be implemented either as a computer program or an application or software module including a computer program, or as a hardware module including a computer program.

[0050] For a more detailed description of the functions of each module of the device, refer to the description of the method for reasoning on text data using a generative language model above. The features of the method of any embodiment can be combined, and vice versa, which will not be elaborated here.

[0051] In an embodiment of the present application, an electronic device is further provided, including: a memory, a processor, and a computer program stored on the memory, and the processor executes the computer program to implement the method for reasoning on text data using a generative language model according to the embodiments of the present application.

[0052] Figure 4 A schematic diagram showing a method that can implement the embodiments of the present application or a device 1000 that can implement the embodiments of the present application is shown. In some embodiments, there may be more or fewer devices than shown in the figure. In some embodiments, it can be implemented using a single or multiple devices. In some embodiments, it can be implemented using cloud or distributed devices.

[0053] As Figure 4 shown, the device 1000 includes a processor 1001, which can perform various appropriate operations and processes according to the programs and / or data stored in the read-only memory (ROM) 1002 or the programs and / or data loaded from the storage section 1008 into the random access memory (RAM) 1003. The processor 1001 can be a multi-core processor or can include multiple processors. In some embodiments, the processor 1001 can include a general main processor and one or more special coprocessors, such as a central processing unit (CPU), a graphics processing unit (GPU), a neural network processing unit (NPU), a digital signal processor (DSP), and so on. In the RAM 1003, various programs and data required for the operation of the device 1000 are also stored. The processor 1001, the ROM 1002, and the RAM 1003 are connected to each other through a bus 1004. The input / output (I / O) interface 1005 is also connected to the bus 1004.

[0054] The above-mentioned processor and memory are jointly used to execute the program stored in the memory, and when the program is executed by a computer, it can implement the methods, steps, or functions described in the above embodiments.

[0055] The following components are connected to the I / O interface 1005: an input section 1006 including a keyboard, a mouse, a touch screen, etc.; an output section 1007 including, for example, a cathode ray tube (CRT), a liquid crystal display (LCD), etc. and a speaker, etc.; a storage section 1008 including a hard disk, etc.; and a communication section 1009 including a network interface card such as a LAN card, a modem, etc. The communication section 1009 performs communication processing via a network such as the Internet. A drive 1010 is also connected to the I / O interface 1005 as needed. A removable medium 1011, such as a magnetic disk, an optical disk, a magneto-optical disk, a semiconductor memory, etc., is installed on the drive 1010 as needed so that a computer program read from it can be installed into the storage section 1008 as needed. Figure 4 Only some components are schematically shown, and it does not mean that the device 1000 only includes Figure 4 the components shown.

[0056] The systems, devices, modules or units illustrated in the above embodiments can be implemented by a computer or its associated components. The computer can be, for example, a mobile terminal, a smart phone, a personal computer, a laptop computer, an in-vehicle human-machine interaction device, a personal digital assistant, a media player, a navigation device, a game console, a tablet computer, a wearable device, a smart TV, an Internet of Things system, a smart home, an industrial computer, a server, or a combination thereof.

[0057] Although not shown, in the embodiments of the present application, a computer-readable storage medium is provided, on which a computer program is stored, and when the computer program is executed by a processor, the text data inference method described in the embodiments of the present application is implemented.

[0058] The storage medium in the embodiments of the present application includes items that are permanent and non-permanent, removable and non-removable, and can implement information storage by any method or technology. Examples of the storage medium include, but are not limited to, phase change memory (PRAM), static random access memory (SRAM), dynamic random access memory (DRAM), other types of random access memory (RAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory or other memory technologies, compact disc read-only memory (CD-ROM), digital versatile disc (DVD) or other optical storage, magnetic cassette tapes, magnetic tape magnetic disk storage or other magnetic storage devices, or any other non-transmission medium that can be used to store information that can be accessed by a computing device.

[0059] Although not shown, the embodiments of the present application also provide a computer program product, including: computer instructions, and when the computer instructions are executed by a processor, the text data inference method described in the embodiments of the present application is implemented.

[0060] The methods, programs, systems, devices, etc. of the embodiments of the present application can be executed or implemented in a single or multiple networked computers, and can also be practiced in a distributed computing environment. In the embodiments of this specification, in these distributed computing environments, tasks can be executed by remote processing devices connected through a communication network.

[0061] Those skilled in the art should understand that the embodiments of this specification can be provided as a method, a system, or a computer program product. Therefore, those skilled in the art can conceive that the implementation of the functional modules / units or controllers and related method steps illustrated in the above embodiments can be achieved in a software, hardware, or a combination of software and hardware manner.

[0062] Unless explicitly stated, the actions or steps of the methods and programs described according to the embodiments of the present application do not necessarily have to be executed in a specific order and can still achieve the desired results. In some embodiments, multitasking and parallel processing are also possible or may be advantageous.

[0063] In this document, multiple embodiments of the present application have been described. However, for the sake of brevity, the descriptions of each embodiment are not exhaustive, and the same or similar features or parts between various embodiments may be omitted. In this document, "one embodiment", "some embodiments", "example", "specific example", or "some examples" mean applicable to at least one embodiment or example according to the present application, rather than all embodiments. The above terms do not necessarily mean referring to the same embodiment or example. Without contradiction, those skilled in the art can combine and combine the different embodiments or examples described in this specification and the features of different embodiments or examples.

[0064] The exemplary systems and methods of the present application have been specifically shown and described with reference to the above embodiments, which are only examples of the best mode for implementing the systems and methods. Those skilled in the art can understand that various changes can be made to the embodiments of the systems and methods described herein when implementing the systems and / or methods without departing from the spirit and scope of the present application defined in the appended claims.

Claims

1. A method for reasoning about text data using a generative language model, characterized in that: The method comprises: receiving a plurality of requests for reasoning from a plurality of users, and adding the plurality of requests to a waiting queue, each of the plurality of requests including text data; Determining a priority of each of the plurality of requests according to a payment level of a user of each of the plurality of requests, a word unit number of the text data, and a waiting time; Select one or more requests with the highest priority from the waiting queue as requests to be processed; The generative language model is caused to perform reasoning on the text data of the request to be processed.

2. The method according to claim 1, characterized in that The method further comprises: Dynamically and asynchronously receiving the multiple requests, adding each received request to the waiting queue; For each request, after the generative language model completes the reasoning of the text data of the request, the request is removed from the waiting queue, and the reasoning result of the request is returned to the user. The reception, inference and return of inference results of different requests are carried out in parallel.

3. The method according to claim 2, characterized in that Selecting one or more requests with the highest priority from the waiting queue as requests to be processed includes: One or more requests with the highest priority and received previously are selected from the waiting queue as requests to be processed.

4. The method according to claim 2, characterized in that: Selecting one or more requests with the highest priority from the waiting queue as requests to be processed includes: Checking whether there are two or more requests in the highest priority request whose number of word units of text data is less than or equal to a first threshold and whose difference in the number of word units of text data is less than or equal to a second threshold; If the result of the checking is that there is a request, the two or more requests are combined into a batch, and the batch is selected as the request to be processed.

5. The method according to claim 4, characterized in that The reasoning includes, in sequence, a pre-filling step and a plurality of decoding steps, The method further comprises: For each request in the waiting queue, the inference state of the request is fed back, where the inference state is pre-filling or decoding, indicating that the request will perform a pre-filling step or a decoding step of inference; When performing the splicing, two or more requests whose inference status is pre-filled are preferentially spliced ​​into one batch.

6. The method according to any one of claims 1 to 5, characterized in that: Determining the priority of each of the multiple requests according to the payment level of the user of each of the multiple requests, the number of words of the text data, and the waiting time, including: The priority of each request is calculated using the following formula: P=g / l×t w Where P represents the priority of the request, g represents the payment level of the requesting user, l represents the number of tokens in the requested text data, and t w Indicates the waiting time for the request.

7. The method according to any one of claims 1 to 5, characterized in that: After selecting one or more requests with the highest priority from the waiting queue as requests to be processed, and before causing the generative language model to perform reasoning on the text data of the requests to be processed, the method further includes: According to the number of tokens of the text data of each request, one or more blocks of storage space for KV cache are allocated to each request to be processed, and the numbers of the one or more blocks of storage space are recorded using a block table. The one or more storage spaces may be discontinuous.

8. A device for reasoning about text data using a generative language model, characterized in that: The device comprises: A request receiver, configured to receive a plurality of requests for reasoning from a plurality of users, and add the plurality of requests to a waiting queue managed by a scheduler, each of the plurality of requests including text data; The scheduler is used to determine the priority of each of the multiple requests according to the payment level of the user of each of the multiple requests, the number of words in the text data, and the waiting time, and select one or more requests with the highest priority from the waiting queue as the request to be processed; A generative language model is used to perform reasoning on text data of the request to be processed.

9. An electronic device comprising a memory, a processor, and a computer program stored in the memory, characterized in that: The processor executes the computer program to implement the method for reasoning about text data using a generative language model according to any one of claims 1 to 7.

10. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the method for reasoning about text data using a generative language model according to any one of claims 1 to 7 is implemented.