Inference methods, platforms, computing devices, storage media, and program products

By performing batch sampling based on business priorities in the large language model inference service to construct target inference request batches, the problem of not being able to provide differentiated service quality in existing technologies is solved, and efficient priority request processing and improved hardware utilization are achieved.

CN122114170APending Publication Date: 2026-05-29BEIJING YUANLI WEILAI SCI & TECH CO LTD

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
BEIJING YUANLI WEILAI SCI & TECH CO LTD
Filing Date
2026-02-25
Publication Date
2026-05-29

AI Technical Summary

Technical Problem

Existing large language model inference services cannot provide differentiated service quality assurance while maintaining hardware utilization when handling mixed priority requests, resulting in high-priority requests being backlogged or computing resources being idle.

Method used

By acquiring the initial batch of inference requests, sampling is performed based on the business priority of each request to form the target batch of inference requests, and then processing is performed using inference hardware to ensure that high-priority requests are processed first, and low-priority requests enter the target batch in proportion, thereby achieving differentiated service quality.

Benefits of technology

It significantly reduced the processing time of high-priority requests, ensuring a good user experience for core businesses, while also reducing the waiting time of low-priority requests, thus balancing the needs of different priority requests.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122114170A_ABST
    Figure CN122114170A_ABST
Patent Text Reader

Abstract

Embodiments of the present specification provide a reasoning method, a platform, a computing device, a storage medium and a program product, wherein the reasoning method comprises: obtaining an initial reasoning request batch, wherein the initial reasoning request batch comprises a plurality of reasoning requests from different businesses, and the priority of each reasoning request is determined based on the business corresponding to each reasoning request; sampling the initial reasoning request batch based on the priority of each reasoning request in the initial reasoning request batch to obtain a target reasoning request batch; processing the target reasoning request batch by using reasoning hardware to obtain a target reasoning result batch corresponding to the target reasoning request batch. It can ensure that high-priority requests have a higher probability of entering the target reasoning request batch for priority processing, and on the premise of not affecting the processing of high-priority requests, the waiting time of low-priority reasoning requests is also reduced, balancing the needs of reasoning requests of different priorities.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This specification relates to the field of artificial intelligence technology, and in particular to a reasoning method, platform, computing device, storage medium, and program product. Background Technology

[0002] In existing technologies, current large language model inference services typically support multi-tenancy and multi-business scenarios. With fixed hardware resources, inference systems usually merge multiple user requests arriving at the same time into a single inference queue, generating the next token for each request in parallel during a single inference process. This significantly improves the utilization of inference hardware and reduces the unit inference cost. However, different inference requests have different priorities. High-priority inference requests require faster token output. When the number of inference requests in the same batch is too large, the token output speed decreases, slowing down the processing time of high-priority inference requests. Conversely, simply reducing the number of inference requests in the same batch slows down the initial response speed of lower-priority inference requests. Therefore, how to construct an inference method to balance the needs of inference requests with different priorities is a problem that needs to be solved. Summary of the Invention

[0003] In view of this, embodiments of this specification provide a reasoning method. One or more embodiments of this specification also relate to a reasoning platform, a computing device, a computer-readable storage medium, and a computer program product, to address the technical deficiencies existing in the prior art.

[0004] According to a first aspect of the embodiments of this specification, a reasoning method is provided, comprising: Obtain the initial inference request batch, which includes multiple inference requests from different services. The priority of each inference request is determined based on the service corresponding to each inference request. Based on the priority of each inference request in the initial inference request batch, the initial inference request batch is sampled to obtain the target inference request batch; Using inference hardware, the target inference request batch is processed to obtain the target inference result batch corresponding to the target inference request batch.

[0005] According to a second aspect of the embodiments of this specification, a question-and-answer method is provided, including: Obtain the initial batch of question and answer requests, which includes multiple question and answer requests from different question and answer services. The question and answer services include at least one of course type, subject type, and user information. The priority of each question and answer request is determined based on the question and answer service corresponding to each question and answer request. Based on the priority of each question and answer request in the initial question and answer request batch, the initial question and answer request batch is sampled to obtain the target question and answer request batch; Using inference hardware, batches of target question-and-answer requests are processed to obtain batches of target inference results corresponding to the batches of target question-and-answer requests.

[0006] According to a third aspect of the embodiments of this specification, an inference platform is provided, including an inference request interface and a response unit; The inference request interface is used to receive inference requests; A response unit is used to execute the steps of any of the above methods in response to a reasoning request.

[0007] According to a fourth aspect of the embodiments of this specification, a computing device is provided, comprising: Memory and processor; The memory is used to store computer-executable instructions, and the processor is used to execute the computer-executable instructions, which, when executed by the processor, implement the steps of any of the methods described above.

[0008] According to a fourth aspect of the embodiments of this specification, a computer-readable storage medium is provided that stores computer-executable instructions that, when executed by a processor, implement the steps of any of the methods described above.

[0009] According to a fifth aspect of the embodiments of this specification, a computer program product is provided, including a computer program / instructions that, when executed by a processor, implement the steps of any of the methods described above.

[0010] In an optional embodiment of this specification, an initial inference request batch is obtained, which includes multiple inference requests from different services, and the priority of each inference request is determined based on the service corresponding to each inference request. Based on the priority of each inference request in the initial inference request batch, the initial inference request batch is sampled to obtain a target inference request batch. The target inference request batch is then processed using inference hardware to obtain a target inference result batch corresponding to the target inference request batch. This method ensures that high-priority requests have a higher probability of entering the target inference request batch for priority processing, significantly reducing their processing time and ensuring the user experience of core services. Meanwhile, lower-priority inference requests will also enter the target inference request batch proportionally, reducing the waiting time of low-priority inference requests without affecting the processing of high-priority requests, thus balancing the needs of inference requests of different priorities. Attached Figure Description

[0011] Figure 1A flowchart of a reasoning method provided in one embodiment of this specification is shown; Figure 2 A flowchart of a question-and-answer method provided in one embodiment of this specification is shown; Figure 3 A flowchart illustrating a reasoning method applied to an inference platform according to an embodiment of this specification is shown; Figure 4 A schematic diagram of the architecture of an inference platform provided in one embodiment of this specification is shown;

[0012] Figure 5 A structural block diagram of a computing device provided in one embodiment of this specification is shown. Detailed Implementation

[0013] Many specific details are set forth in the following description to provide a full understanding of this specification. However, this specification can be implemented in many other ways than those described herein, and those skilled in the art can make similar extensions without departing from the spirit of this specification. Therefore, this specification is not limited to the specific implementations disclosed below.

[0014] The terminology used in one or more embodiments of this specification is for the purpose of describing particular embodiments only and is not intended to be limiting of the one or more embodiments of this specification. The singular forms “a” and “the” as used in one or more embodiments of this specification and the appended claims are also intended to include the plural forms, unless the context clearly indicates otherwise. It should also be understood that the term “and / or” as used in one or more embodiments of this specification refers to and includes any or all possible combinations of one or more associated listed items.

[0015] It should be understood that although the terms first, second, etc., may be used to describe various information in one or more embodiments of this specification, such information should not be limited to these terms. These terms are only used to distinguish information of the same type from one another. For example, first may also be referred to as second without departing from the scope of one or more embodiments of this specification, and similarly, second may also be referred to as first. Depending on the context, the word "if" as used herein may be interpreted as "when," "when," or "in response to a determination."

[0016] Furthermore, it should be noted that the user information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data used for analysis, stored data, displayed data, etc.) involved in one or more embodiments of this specification are all information and data authorized by the user or fully authorized by all parties. Moreover, the collection, use and processing of related data must comply with the relevant laws, regulations and standards of the relevant countries and regions, and corresponding operation entry points are provided for users to choose to authorize or refuse.

[0017] First, the terms and concepts used in one or more embodiments of this specification will be explained.

[0018] TPOT (Time Per Output Token) refers to the average time required for an inference system to generate one output token. This metric directly affects the system's response speed and user experience.

[0019] TTFT (Time To First Token) is the time required from when a user sends a request to when the user receives the first token generated by the model.

[0020] Throughput refers to the amount of inference tasks a system completes per unit of time.

[0021] This specification provides a reasoning method, and also relates to a question-and-answer method, a reasoning platform, a computing device, a computer-readable storage medium, and a computer program product, which will be described in detail in the following embodiments.

[0022] Traditional large language model inference services typically employ a uniform batch processing strategy when handling mixed-priority requests. This fails to maintain hardware utilization while providing differentiated quality of service (QoS) guarantees for high-priority requests, leading to a backlog of urgent requests or idle computing resources. The optional embodiments provided in this specification aim to construct a priority scheduling framework based on intra-batch sampling. This framework obtains an initial batch containing multi-service inference requests and binds each request to a priority identifier determined by the service. Within the same batch, low-priority requests are sampled proportionally based on priority, while all high-priority requests are retained, forming a target batch. The target batch is then delivered to the inference hardware for parallel processing. This design moves priority scheduling decisions from cross-batch queuing to within the batch, achieving QoS differentiation without disrupting batch processing continuity. It ensures deterministic resource guarantees for high-priority requests, while low-priority requests fill idle computing power with a controllable probability, balancing latency and throughput efficiency.

[0023] See Figure 1 , Figure 1 A flowchart of a reasoning method provided in one embodiment of this specification is shown, which specifically includes the following steps.

[0024] Step 102: Obtain the initial inference request batch, which includes multiple inference requests from different services. The priority of each inference request is determined based on the service corresponding to each inference request.

[0025] An inference request is a task unit submitted to a model with inference service capabilities. It typically includes input text, parameter configuration, and contextual information. The purpose of an inference request is to drive the model to perform a forward computation to generate output content. For example, inference requests include, but are not limited to, question answering requests from question-answering platforms, code writing requests from code assistants, and article continuation requests from content creation platforms.

[0026] Priority indicates the position of a reasoning request in the resource allocation and processing order, allowing the system to schedule requests differently based on business needs and ensure the service quality of high-value services. For example, priority can be an integer identifier (such as 0, 1, 2), where 0 represents the highest priority; for example, priority can be an enumerated value or a floating-point number used for dynamic weight allocation.

[0027] A business is a logical unit that invokes the inference service, such as different applications, tenants, product functions, or customer categories. Each business is typically associated with specific service level objectives, resource quotas, or billing policies. Businesses can serve as the basis for priority mapping, directly linking priority to the value corresponding to the business. For example, a business could be a "real-time question-answering robot," requiring extremely low latency; an business could be "offline batch document analysis," not sensitive to real-time requirements; or a business could be a "dedicated assistant for paid enterprise customers," enjoying higher processing priority than free users.

[0028] The initial inference request batch is a temporary set of multiple inference requests pre-extracted from the run queue or input buffer before priority sampling and hybrid inference are performed. It provides input for subsequent priority splitting and sampling selection and is the first-level data aggregation unit in the batch processing flow. For example, the initial inference request batch may be 64 requests taken from the head of the queue at once; for example, the initial inference request batch may be a group of requests aggregated according to a time window (such as requests arriving in the most recent 10 milliseconds).

[0029] The priority of each inference request is determined based on the business logic corresponding to that request; that is, priority can be assigned through mapping rules associated with the business identifier. For example, the server maintains a "business ID-priority level" mapping table, and when a request is received, the priority is obtained by looking up the table based on its business ID; for example, when an external API is called, a predefined priority parameter is directly carried in the request header.

[0030] To obtain the initial inference request batch, one optional implementation is that the inference service maintains a first-in-first-out (FIFO) running queue, and the scheduling thread pulls requests from the head of the queue in a loop with a fixed batch size (e.g., 64). Each time a batch is fully retrieved, it forms an initial inference request batch. Another optional implementation is that, in a distributed message queue scenario, the consumer group pulls messages from the topic subscription in micro-batches, aggregates them into batches based on the message arrival timestamp or offset, and temporarily stores them in local memory.

[0031] In one implementation of this specification, in an intelligent question-and-answer platform for education, student users upload challenging math and physics questions by taking photos. The platform's backend calls a large language model inference service to generate solutions and answers in real time. The platform supports three service types: online exam assistants require a first-character latency of less than 300 milliseconds and a character-by-character generation speed of less than 50 milliseconds, mapped to the highest priority; homework help is insensitive to second-level latency, mapped to medium priority; and batch test paper analysis is an offline task, tolerating latency of more than second-level latency, mapped to the lowest priority. The inference service entry point maintains a service-priority mapping table, assigning a corresponding priority value to each request based on the service type identifier carried in the request header. All requests then enter a global first-in-first-out (FIFO) queue. The scheduling thread cyclically pulls requests from the head of the queue in fixed batch sizes of 64. Each time 64 requests are pulled, they are aggregated and encapsulated into an initial inference request batch. This batch contains inference requests from different priority services, and each request carries a clear priority identifier, providing accurate input for subsequent priority-based sampling selection.

[0032] The purpose of step 102 is to establish clear data boundaries and starting states for the inference scheduling process, construct a large number of inference requests into an initial batch of inference requests, and provide a data foundation for the processing of subsequent target batches of inference requests and inference requests.

[0033] Step 104: Based on the priority of each inference request in the initial inference request batch, sample the initial inference request batch to obtain the target inference request batch.

[0034] Sampling is a process of selectively retaining or discarding inference requests within an initial batch of requests, based on their priority category. Its purpose is to control the proportion of requests with different priorities entering the final inference execution stage, thereby achieving differentiated service quality while maintaining batch processing efficiency. For example, sampling could involve randomly retaining half of the requests with priority 1; for example, sampling could involve sequentially selecting one-third of the requests with priority 2; and for example, sampling could involve temporarily excluding all requests with priority 3 or higher from the current batch.

[0035] The target inference request batch is the batch that is ultimately delivered to the inference hardware for inference execution, obtained by performing priority splitting on the initial inference request batch and sampling based on the priority.

[0036] Based on the priority of each inference request in the initial inference request batch, the initial inference request batch is sampled. One optional implementation is to traverse all requests in the initial batch, mark all requests with priority 0 as "reserved," and randomly select requests with priority greater than or equal to 1 according to a fixed sampling ratio corresponding to their priority. The selected requests are marked as "reserved," and the rest are marked as "not processed" and returned to the queue. Another optional implementation is to adopt a deterministic sampling strategy, filling the target batch sequentially in ascending order of priority until a preset target batch size is reached, with lower priority requests having a lower probability of being included.

[0037] To obtain the target inference request batch, one optional implementation is to extract the sampled requests from the initial batch and assemble them into a new list of requests, i.e., the target inference request batch, according to their original order in the batch or after reordering. Another optional implementation is, in an inference engine that supports dynamic batching, to directly fill the command buffer of the hardware computing unit with the retained requests, implicitly forming the target batch. Yet another optional implementation is to reinsert the sampled and excluded requests into the tail of the run queue or back them to a specific priority queue according to their priority, awaiting the next round of scheduling.

[0038] In one implementation of this specification, within a large-model inference service for an intelligent customer service system, the running queue simultaneously contains inquiry requests from member users (priority 0) and inquiries from ordinary users (priority 1). The scheduling thread pulls requests from the head of the queue in fixed batch sizes of 32, but the current batch only contains 12 priority 0 requests, insufficient to fill the hardware's optimal batch size. The system then enables a dynamic sampling strategy: based on the current service level target achievement, the sampling ratio of priority 1 requests is temporarily increased to 80%. The system proportionally extracts 16 requests from the 20 priority 1 requests in the batch and merges them with the 12 priority 0 requests to form a target inference request batch of 28 requests. This effectively improves batch processing utilization and system throughput while ensuring service quality for member users.

[0039] The purpose of step 104 is to split the unified initial inference request batch into processing paths with different priorities, and to selectively retain requests with different priorities through a sampling mechanism, thereby achieving differentiated service level goals under the constraint of shared batch processing resources, and achieving a configurable balance between batch processing efficiency, high-priority latency and low-priority throughput.

[0040] Step 106: Using inference hardware, process the target inference request batch to obtain the target inference result batch corresponding to the target inference request batch.

[0041] Inference hardware is a physical device or computing unit specifically designed to accelerate the forward computation process of neural network models. Inference hardware can provide high-throughput, low-latency matrix operation capabilities, enabling large language models to complete the iterative generation of word sequences within an acceptable timeframe. Exemplarily, inference hardware can be a graphics processing unit (GPU); exemplarily, inference hardware can be a neural network processor; exemplarily, inference hardware can be a field-programmable gate array (FPGA) or an application-specific integrated circuit (ASIC).

[0042] The target inference result batch is a set of terms generated by the inference hardware, corresponding one-to-one with each request in the target inference request batch. Its purpose is to serve as the output of this round of inference, either delivered to the caller to complete the response, or concatenated with the original request and re-entered into the scheduling process to generate subsequent terms. For example, the target inference result batch can be a list consisting of one term generated by each of the sixteen requests within the batch; for example, the target inference result batch can be a structured response object containing the generated terms, confidence scores, and intermediate state information for each request.

[0043] The process of processing target inference request batches using inference hardware involves inputting these batches into a loaded large language model, performing a complete forward propagation computation, and generating the output lexical for each request in the current iteration step. One optional implementation involves padding the input sequences of each request in the target inference request batch to a fixed length, performing lexicalization and positional encoding, copying them as tensors to the graphics processor's memory, calling the compiled model execution kernel to complete a single forward propagation, and sampling the next lexical for each request from the output probability matrix. Another optional implementation involves inserting the target inference request batches into the currently executing dynamic batches in an inference engine that supports continuous batch processing, utilizing idle computing units to process newly added requests in parallel.

[0044] To obtain the target inference result batch corresponding to the target inference request batch, one optional implementation is to assemble the tokens generated in this iteration of each request and their corresponding request identifiers into a response tuple, and organize all tuples into a list structure in batch order, and return it to the scheduling module as the target inference result batch; another optional implementation is to directly append the generated tokens to the output buffer of the request object, and determine whether the termination condition is met. If not, the request is re-inserted into the running queue to wait for the next round of scheduling.

[0045] In one implementation of this specification, for an educational intelligent question-and-answer platform, a batch of 28 target inference requests (12 online exam help requests and 16 homework help requests) is delivered to a neural network processor for inference computation. Some of the online exam help requests in this batch have generated 95 tokens, approaching the maximum length limit for a single request; some homework help requests have only generated about 20 tokens. After performing forward propagation, the neural network processor generates the token sequence for this iteration. The scheduling module identifies three online exam help requests that have generated terminating tokens, removes them from the processing flow, and pushes the complete 300-word solution to the student client in real time via the platform interface; the remaining 25 requests all generate non-terminating tokens, and the system appends the new tokens to the output buffer of each request and puts the updated requests back into the running queue. Simultaneously, the hardware computation time and token generation speed of this inference are recorded for dynamic adjustment of the sampling ratio in subsequent iterations.

[0046] The purpose of step 106 is to actually deliver the target inference request batches that have undergone priority screening and sampling control to high-performance computing hardware to complete the forward computation of the model and produce the generated lexical units for the current iteration step.

[0047] The optional embodiments of this specification can ensure that high-priority requests are more likely to enter the target inference request batch for priority processing, significantly reducing their processing time and ensuring the user experience of core businesses; while lower-priority inference requests will also enter the target inference request batch in proportion, reducing the waiting time of low-priority inference requests without affecting the processing of high-priority requests, thus balancing the needs of inference requests of different priorities.

[0048] Optionally, step 104 further includes: Based on the priority of each inference request in the initial inference request batch, the inference requests in the initial inference request batch are divided into high-priority inference requests and low-priority inference requests. Sample the low-priority inference requests in the initial inference request batch to obtain the sampled low-priority inference requests; The target inference request batch is constructed based on the sampled low-priority inference requests and the high-priority inference requests in the initial inference request batch.

[0049] High-priority inference requests are those identified as having the highest processing priority and requiring undifferentiated computing resource guarantees within the initial inference request batch. For example, a high-priority inference request can be any request with a priority value of 0; for example, a high-priority inference request can be any request marked as "high" in the priority enumeration value; for example, a high-priority inference request can be a request determined to require priority protection according to a business level agreement.

[0050] Low-priority inference requests are those with a priority level lower than the highest priority in the initial inference request batch, requiring sampling and filtering to control their resource consumption. Their purpose is to selectively retain them, maintaining batch processing efficiency while preventing excessive interference with high-priority requests. For example, low-priority inference requests can be all requests with a priority value greater than or equal to 1; for example, low-priority inference requests can be all requests marked as "medium" or "low" in the priority enumeration value; for example, low-priority inference requests can be requests determined to have tolerable latency according to the business level agreement.

[0051] The sampled low-priority inference requests are those low-priority requests selected and retained for subsequent inference stages after the initial batch of low-priority requests was sampled. Their purpose is to serve as the actual set of low-priority requests that receive service in this round of inference.

[0052] The initial inference request batch is divided into high-priority and low-priority inference requests. One possible implementation is to iterate through each request in the initial inference request batch, read its priority value from its metadata, and assign requests with a priority of 0 to the high-priority set, and requests with a priority greater than or equal to 1 to the low-priority set. Another possible implementation is to classify requests with priority values ​​less than a preset priority threshold (e.g., a threshold set to 1) as high-priority and requests with priority values ​​greater than or equal to the threshold as low-priority. Yet another possible implementation is to directly assign requests belonging to core businesses or paid tenants to high-priority based on a business type and priority mapping rule, and assign all other requests to low-priority.

[0053] The initial batch of inference requests samples low-priority inference requests to obtain a sampled set of low-priority inference requests. One possible implementation is to apply different fixed sampling ratios to the low-priority request set according to their respective priority levels; the sampling ratio decreases for each priority level increase of 1, and each request is independently and randomly selected for retention. Another possible implementation is to use overall quota control for the low-priority request set, dynamically calculating the upper limit of low-priority requests that the current batch can accommodate based on the current system load, queue backlog, and the number of high-priority requests, and then selecting requests sequentially from highest to lowest priority until the upper limit is reached. Yet another possible implementation is to sort low-priority requests by arrival time, sampling the system with a probability of the reciprocal of the priority level, prioritizing requests with longer waiting times to avoid starvation.

[0054] Based on the sampled low-priority inference requests and the high-priority inference requests in the initial inference request batch, a target inference request batch is constructed. One optional implementation is to merge the high-priority request set with the sampled low-priority request set, rearrange them according to their original order in the initial batch, and form a new request list as the target inference request batch. Another optional implementation is to place high-priority requests at the beginning of the batch to ensure they receive computing resources first, followed by the sampled low-priority requests, thus constituting the target inference request batch.

[0055] In one implementation of this specification, within a large-model inference service for an intelligent customer service system, the initial inference request batch contains 12 high-priority member user inquiry requests and 20 low-priority ordinary user inquiry requests. The system retains all 12 high-priority requests. However, the current batch contains too few high-priority requests to fill the optimal hardware batch size. Based on real-time monitoring of service level target achievement, the system dynamically increases the sampling ratio of low-priority requests to 80%, selecting 16 requests from the 20 ordinary user inquiry requests in sequence. The system then merges the 12 high-priority requests with the 16 sampled low-priority requests, forming a target inference request batch of 28 requests. This improves batch processing utilization while ensuring service quality for member users.

[0056] The optional embodiments provided in this specification simplify and solidify the priority scheduling logic by explicitly dividing the requests in the initial inference request batch into two categories: high priority and low priority. High priority requests are unconditionally all retained, while only low priority requests are sampled and selected. This effectively improves batch processing efficiency and hardware utilization while maintaining low latency for high priority requests.

[0057] Optionally, step 104 further includes: Based on the priority of each low-priority inference request, determine the sampling ratio of each low-priority inference request. Based on the sampling ratio, low-priority inference requests in the initial inference request batch are sampled to obtain the sampled low-priority inference requests.

[0058] The sampling ratio is a numerical rule used to determine the probability or number of requests or request types selected for retention when sampling low-priority inference requests. The sampling ratio can ensure that lower-priority requests receive less computational resource usage. For example, the sampling ratio can be a fixed ratio inversely proportional to the priority value; for example, the sampling ratio can be a percentage threshold pre-configured based on the priority level; for example, the sampling ratio can be a probability value dynamically adjusted based on system load.

[0059] Sampling according to a sampling ratio is a process that uses a sampling ratio as the basis for decision-making to retain or exclude low-priority inference requests. For example, sampling according to a sampling ratio could involve generating a random number independently for each request and comparing it with the sampling probability to determine whether to retain it; or it could involve calculating the number of requests that should be retained proportionally, and then selecting the corresponding number of requests in the order of arrival or a random order.

[0060] Based on the priority of each low-priority inference request, the sampling ratio of each low-priority inference request is determined. One possible implementation is to establish a sampling ratio function based on the priority value, with the sampling ratio decreasing accordingly by a preset attenuation coefficient for each priority level increase. Another possible implementation is to pre-configure an independent fixed sampling ratio for each priority level, with the system directly reading the configuration table to obtain the corresponding ratio value at runtime. Yet another possible implementation is to map priority values ​​to weight values, with the sampling ratio dynamically calculated by dividing the weight value of the current request by the sum of the weight values ​​of all low-priority requests in the same batch.

[0061] According to the sampling ratio, low-priority inference requests in the initial inference request batch are sampled to obtain the sampled low-priority inference requests. One optional implementation is to traverse the set of low-priority requests, generate a random number for each request based on the sampling probability determined by its priority, and retain requests whose random numbers are less than the probability, while leaving the remaining requests unprocessed. Another optional implementation is to first calculate the number of requests to be retained for each priority level based on the sampling ratio and the number of requests for that level, and then select the corresponding number of requests to be retained from each level in a first-in-first-out or random order. Another optional implementation is to use a reservoir sampling algorithm to dynamically maintain a retention set during the traversal of low-priority requests, ensuring that the final retention ratio is consistent with the preset sampling ratio.

[0062] In one implementation of this specification, a programming assignment tutoring platform for higher education institutions simultaneously serves three types of users: students taking timed online exams (priority 0, high priority), students submitting after-class programming assignments (priority 1), and teachers uploading past exam papers to generate a practice question bank (priority 2). The initial inference request batch contains 8 priority 0 requests, 40 priority 1 requests, and 16 priority 2 requests. The system assigns the 8 priority 0 requests to the high-priority set. For low-priority requests, the system configures a sampling ratio of 60% for priority 1 and 30% for priority 2. The system first calculates the number of requests to be retained for each level: 24 of the 40 priority 1 requests should be retained (60%), and 5 of the 16 priority 2 requests should be retained (30%). Subsequently, the system sequentially retrieves 24 requests from the head of the priority 1 queue and 5 requests from the head of the priority 2 queue, forming the sampled low-priority inference request set, which is then merged with the 8 high-priority requests to form the target inference request batch.

[0063] In an optional embodiment of this specification, by binding the sampling ratio to the priority level of low-priority requests, low-priority requests of different priorities can obtain differentiated resource occupation opportunities. Within the unified batch processing framework, non-urgent business can be finely scheduled. The lower the priority of a request, the lower the proportion that enters the inference stage, thereby releasing more computing resources to higher-value requests.

[0064] Optionally, in step 104, low priority includes multiple priority levels; Accordingly, based on the priority of each low-priority inference request, the sampling ratio of each low-priority inference request is determined, including: Based on the priority level to which each low-priority inference request belongs, determine the sampling ratio of multiple priority levels for each low-priority inference request. Priority levels are a sequence of multiple levels further divided based on business value or service level objectives. Each level corresponds to a specific priority value or enumeration category. Their purpose is to provide a more granular division of non-urgent business. For example, priority levels may include three levels: Priority 1, Priority 2, and Priority 3; for example, priority levels may be divided into second-level response layers, ten-second-level response layers, and minute-level response layers based on latency tolerance in the service level agreement.

[0065] Tiered sampling is an operation process that independently samples low-priority requests at different priority levels using their respective sampling ratios, and then aggregates the sampling results from each level. For example, tiered sampling could involve sampling priority 1 requests at a ratio of one-half, priority 2 requests at a ratio of one-third, and priority 3 requests at a ratio of one-quarter, with each level executing independently. Alternatively, tiered sampling could involve first selecting a corresponding number of requests according to a preset reservation quota for each level, and then merging them into a sampled set.

[0066] Based on the priority level to which each low-priority inference request belongs, the sampling ratio of multiple priority levels for each low-priority inference request is determined. One optional implementation is to pre-configure an independent fixed sampling ratio for each priority level, with the configured value being inversely proportional to the priority value. The sampling ratio decreases by a fixed attenuation factor for each increase in priority level. Another optional implementation is to map the priority level to a weight coefficient, and the sampling ratio is calculated by dividing the weight of the level by the sum of the weights of all low-priority levels, thereby achieving proportional allocation between levels.

[0067] Based on the sampling ratio, low-priority inference requests in the initial inference request batch are sampled to obtain the sampled low-priority inference requests, including: Based on the sampling ratio of multiple priority levels, the low-priority inference requests in the initial inference request batch are sampled in a stratified manner to obtain the sampled low-priority inference requests.

[0068] Based on the sampling ratio of multiple priority levels, the low-priority inference requests in the initial inference request batch are sampled in a stratified manner to obtain the sampled low-priority inference requests. One optional implementation is to split the low-priority requests into multiple independent subsets according to the priority level, and perform random sampling on each subset according to the sampling ratio corresponding to that level. The sampling process of each level is executed in parallel or serially, and finally the requests retained from all levels are merged. Another optional implementation is to pre-calculate the quota of requests to be retained for each level based on the sampling ratio of each level and the number of requests for that level, and then select the corresponding number of requests from the queue of each level in a first-in-first-out order to retain them. If the quota is insufficient, the actual number of requests is retained.

[0069] This specification describes an implementation process for an educational intelligent question-and-answer platform. The platform further divides low-priority services into two priority levels: homework assistance (priority level 1) with a fixed sampling ratio of 1 / 2; and batch test paper analysis (priority level 2) with a fixed sampling ratio of 1 / 3. The current initial inference request batch contains 30 homework assistance requests and 22 batch test paper analysis requests. The system performs tiered sampling: the 30 requests in priority level 1 are independently sampled randomly with a 1 / 2 probability, with 15 actually retained; the 22 requests in priority level 2 are independently sampled randomly with a 1 / 3 probability, with 7 actually retained. The sampling results of the two levels do not interfere with each other. The final sampled low-priority requests total 22, which are then combined with 12 high-priority requests to form the target inference request batch.

[0070] By further dividing low-priority requests into multiple priority levels and configuring the sampling ratio independently for each level, fine-grained hierarchical scheduling of non-urgent services is achieved, improving the resource scheduling flexibility and business adaptability in multi-service hybrid deployment scenarios.

[0071] Optionally, step 102 further includes: Retrieve multiple inference requests from different business units from the inference request queue; Based on the business corresponding to each inference request, determine the priority of each inference request; An initial batch of inference requests is constructed based on each inference request and its priority.

[0072] The inference request queue is a buffer data structure used to temporarily store inference requests that arrive at the inference service but have not yet been scheduled for processing. For example, the inference request queue can be a first-in, first-out (FIFO) memory queue; for example, the inference request queue can be multiple priority-based queues, with requests of different priorities stored in different queues.

[0073] To retrieve multiple inference requests from different business processes from the inference request queue, one possible implementation is that the scheduling thread pulls requests from the head of the global first-in-first-out queue in a fixed batch size, stopping once the specified number of requests is reached. Another possible implementation is that the system dynamically calculates the number of requests to be pulled based on the current queue length and a preset maximum waiting time using an adaptive algorithm, achieving a balance between latency constraints and batch processing efficiency.

[0074] Based on the business logic corresponding to each inference request, the priority of each inference request is determined. One possible implementation is that priority mapping is completed when the request arrives, and the priority value is attached to the request object as metadata. The scheduling thread directly reads this field from the request. Another possible implementation is that after obtaining the request, the scheduling thread queries the locally cached priority mapping table in real time based on the business identifier carried in the request, and writes the query result into the metadata field of the request.

[0075] Based on each inference request and its priority, an initial inference request batch is constructed. One optional implementation is to organize the acquired requests, along with their respective determined priority information, into a list structure according to their original order in the queue, as the initial inference request batch. Another optional implementation is to sort the requests within the batch from highest to lowest priority while constructing the batch, placing high-priority requests at the head of the batch to achieve better cache locality.

[0076] In one implementation of this specification, within an intelligent question-and-answer platform for smart education, the inference service maintains a global first-in-first-out (FIFO) inference request queue. All requests from students and teachers enter the tail of the queue in order of arrival time. The platform simultaneously supports three types of services: online exam assistants requiring real-time problem-solving guidance (mapped to priority 0); homework Q&A requiring detailed step-by-step explanations (mapped to priority 1); and batch test paper analysis for teachers to upload blank test papers and generate knowledge point tags (mapped to priority 2). The scheduling thread, with a fixed batch size of 64, continuously pulls requests from the head of the queue. When the accumulated requests at the head of the queue reach 64, the scheduling thread retrieves 64 requests at once, reads the priority field written by the front-end gateway from the metadata of each request, and organizes the 64 requests and their corresponding priority information into a list in their original order of retrieval, constructing the initial inference request batch, which is then delivered to the subsequent sampling module for processing.

[0077] The optional embodiments of this specification provide a complete and executable pre-implementation scheme for the entire priority sampling inference method.

[0078] Optionally, step 106 further includes: Using inference hardware, word-by-word processing is performed on the first inference request in the target inference request batch to obtain the current inference result in the first inference request, wherein the first inference request is any one of the inference requests in the target inference request batch; Based on the first inference request and the current inference result, determine the updated first inference request; Add the first inference request to the inference request queue, and return to execute the steps of retrieving multiple inference requests from different business processes from the inference request queue; Until the last word in the first inference request is processed, the target inference result in the first inference request is obtained; The process continues until the last word in each inference request is processed, at which point the target inference result batch corresponding to the target inference request batch is obtained.

[0079] The current inference result is the output token of this iteration, generated by the model after a single token-by-token processing of the first inference request. Its purpose is to serve as the direct output of this round of inference, used to determine whether to terminate or continue generation after concatenation. For example, the current inference result can be a token identifier in integer form; for example, the current inference result can be a string containing token text; for example, the current inference result can be a special terminating token indicating the end of a sentence.

[0080] The updated first inference request refers to a new inference request object formed by appending the new word to the end of the original first inference request's input sequence and updating the request state when the current inference result is not a terminating word. Its purpose is to preserve the generated context, allowing subsequent iterations to continue predicting subsequent content. For example, the updated first inference request can be a request structure that appends the new word to the end of the input word list; for example, the updated first inference request can be a request object carrying a portion of the generated text.

[0081] The target inference result is the output token obtained after performing word-by-word processing on the first inference request in this round, which is the token generated by the request in this inference scheduling. When the token is a terminating token, the request will not continue after this round of processing; when the token is a non-terminating token, the token will be used as the current inference result to construct the updated request and re-enqueued. For example, the target inference result can be a token "solution"; for example, the target inference result can be a terminating token; for example, the target inference result can be a special token representing a newline.

[0082] The target inference result batch is a collection of target inference results generated by each inference request in this round of inference scheduling, organized in the order of the requests. Its purpose is to serve as the output of this round of batch processing, delivered to the scheduling module for subsequent request status updates and queuing decisions. For example, the target inference result batch can be a list consisting of one lexical unit generated by each of the thirty-four requests; for example, the target inference result batch can be a set of lexical arrays corresponding one-to-one with the input batch.

[0083] Using inference hardware, the first inference request in the target inference request batch is processed word-by-word to obtain the current inference result in the first inference request. One optional implementation involves loading the target inference request batch into the graphics processor's memory as a tensor, calling the model execution kernel to perform parallel computation of the forward propagation of all requests within the batch, sampling the current word for each request from the output probability matrix, and extracting the word corresponding to the first inference request from the result tensor. Another optional implementation involves skipping the current round of computation and marking requests that have generated terminating words as completed.

[0084] Based on the first inference request and the current inference result, an updated first inference request is determined. One possible implementation is that if the current inference result is not a terminating term, it is appended to the end of the input term list of the first inference request, and the generated length is incremented by one to form the updated first inference request. Another possible implementation is that if the current inference result is a terminating term, no updated request is generated.

[0085] The first inference request is added to the inference request queue. The process then returns to retrieve multiple inference requests from different business processes from the queue. One possible implementation is to insert the updated first inference request at the end of the inference request queue according to its original priority, awaiting retrieval by the next scheduling thread. Another possible implementation is to not enqueue requests that have completed all lexical generation (i.e., generated the terminating lexical).

[0086] Until the processing of the last term in the first inference request is completed, and the target inference result of the first inference request is obtained, it means that when the current inference result generated by the first inference request in this round of processing is a terminating term, that terminating term is the target inference result of the request in this round, and the request is completed and will not be re-enqueued. Until the processing of the last term in each inference request is completed, and the target inference result batch corresponding to the target inference request batch is obtained, it means that when all requests in the target inference request batch have generated the target inference result of this round (regardless of whether it is a terminating term), the system aggregates the target inference results of all requests in batch order to form an output batch corresponding to the input batch.

[0087] This specification describes an intelligent question-and-answer platform for education during implementation. A target inference request batch contains 34 requests. The system fills and aligns the lexical input sequence of this batch of requests, copies it to the graphics processor's memory in tensor form, and performs one forward propagation calculation. The inference engine samples the current lexical for each request from the output probability distribution. For one online exam help request, the generated lexical is "solution," which is a non-terminal lexical. The system uses this as the current inference result and appends it to the end of the request input sequence, forming an updated request, which is then reinserted into the global inference request queue. For another online exam help request, the generated lexical is a terminal lexical. The system uses this terminal lexical as the target inference result for this request. This request is processed and not re-queued; its complete answer has already been generated in a previous round and pushed to the client. The system organizes all generated lexicals (including non-terminal and terminal lexicals) from all requests in this batch into a list according to the request order, forming a target inference result batch, which is then delivered to the scheduling module for subsequent request status updates and queuing decisions.

[0088] In one implementation of this specification, within an intelligent document collaboration platform designed for enterprise office scenarios, the target inference request batch contains 20 real-time meeting minutes generation requests. After the system performs word-by-word processing, 19 requests generate non-terminal tokens, and 1 request generates a terminal token. The system appends the 19 non-terminal tokens to the input sequence of their respective requests, forming updated requests, and re-inserts them into the priority queue. The request that generated the terminal token is marked as completed; this terminal token is the target inference result for that request in this round, and its complete meeting minutes text has already been generated and pushed to the client in previous rounds. The system aggregates the tokens generated by all 20 requests in this round into a target inference result batch, which is then used by the scheduling module to record the output of this inference.

[0089] By combining word-by-word meta-processing, request status updates, and re-queuing mechanisms, a complete closed loop of single-inference scheduling and continuous iterative generation is constructed, which improves the throughput efficiency and resource utilization of the inference service, while providing a unified implementation foundation for advanced features such as streaming output and dynamic priority adjustment.

[0090] Optionally, step 106 further includes: Combine the first inference request and the current inference result to obtain the updated first inference request.

[0091] The process involves concatenating the first inference request and the current inference result to obtain an updated first inference request. Specifically, this involves appending the output lexical corresponding to the current inference result to the end of the existing input sequence of the first inference request, forming a new request object that includes the original context and the content generated in this round. This allows the next round of inference to continue generating subsequent lexicals based on a more complete context. One optional implementation is to sequentially merge the input lexical list of the first inference request with the lexical identifier of the current inference result to form a new input lexical list, incrementing the generated length counter by one, while keeping other request attributes unchanged. Another optional implementation is that if the first inference request is stored in text format, the lexical text of the current inference result is directly appended to the end of the original text to generate a new text string as the updated request input.

[0092] This specification describes an implementation of an intelligent question-and-answer platform for education. In a target reasoning request batch, the current input sequence of an online exam assistance request is: "Given that a quadratic function passes through points (1,0) and (3,0), and the vertex's y-coordinate is -4, find the function's analytical expression. Solution steps:". The generated term for this round of reasoning is "Let". The system directly appends this term to the end of the original request input text, resulting in the updated request input: "Given that a quadratic function passes through points (1,0) and (3,0), and the vertex's y-coordinate is -4, find the function's analytical expression. Solution steps: Let". Simultaneously, the system increases the generated length of the request from thirty-two to thirty-three, while keeping other attributes (request identifier, business type, priority, etc.) unchanged, forming the updated first reasoning request, which is then reinserted into the end of the reasoning request queue, awaiting the next round of scheduling.

[0093] In one implementation of this specification, within an intelligent code-assisted generation platform for enterprise R&D scenarios, a developer requests the generation of implementation code for a sorting function. The input sequence of the current inference request includes the function declaration "def quick_sort(arr):" and the generated indentation and comments. The term generated in this round of inference is "if". The system appends this term as a term identifier to the end of the input term list maintained by the request object, forming a new input term sequence of length 112, and simultaneously updates the generated length counter. The updated request, carrying the complete generated code context, is reinserted into the priority queue for the next round of inference to continue generating subsequent code content.

[0094] By updating the inference request through splicing operations, the standardization and lightweighting of request state iteration are achieved, providing a simple and scalable implementation foundation for the entire iterative inference process.

[0095] Optionally, in step 106, the last lexical is a terminator lexical.

[0096] Termination lexical units are special lexical units used by large language models to mark the end of a complete output sequence during the autoregressive generation process.

[0097] The last terminator is the terminator terminator. Specifically, after performing word-by-word processing on the first inference request, the current inference result generated in this round is determined by the model to be the terminator. This terminator is both the output terminator of this round and the final terminator generated in the entire lifecycle of the request. For example, the last terminator can be an independently encoded terminator identifier; for example, the last terminator can be a specific text marker indicating the end of the answer; for example, the last terminator can be an empty sequence end marker.

[0098] One optional implementation involves comparing the generated current inference result with a pre-stored terminator token identifier after performing word-by-word processing. If they match, the current output is determined to be a terminator token, the request is processed, and no further updated requests are generated and enqueued. Another optional implementation involves including the probability distribution of the model output in the probability value corresponding to the terminator token. When this probability value is at its highest or the sample hits the terminator, the current generation is identified as terminated. Yet another optional implementation involves forcibly replacing the current output with a terminator token when the generated request length reaches a preset maximum length limit, thus preventing infinite generation.

[0099] By setting a terminator, a clear and executable task completion determination mechanism is provided for the iterative reasoning process, improving the overall efficiency and response accuracy of the reasoning service.

[0100] Optionally, the inference hardware is accelerated hardware for a graphics processor; Accordingly, step 106 also includes: Using inference hardware, the target inference request batch is processed to obtain the target inference result batch corresponding to the target inference request batch, including: By utilizing the accelerated hardware of the graphics processing unit, each inference request in the target inference request batch is processed in parallel to obtain the target inference result batch corresponding to the target inference request batch.

[0101] The accelerated hardware of a graphics processing unit (GPU) is a dedicated GPU computing unit for parallel processing of large-scale matrix operations. Its core architecture contains thousands of small computing cores, capable of performing a large number of identical arithmetic operations simultaneously. This provides high-throughput parallel computing capabilities for attention mechanisms and feedforward neural networks in large language models, reducing end-to-end latency in batch inference. For example, the accelerated hardware of a GPU can be a data center-class GPU for cloud inference services; for example, it can be a GPU accelerator card for edge computing.

[0102] By leveraging the accelerated hardware of the graphics processing unit (GPU), each inference request in a target inference request batch is processed in parallel to obtain the target inference result batch corresponding to the target inference request batch. This involves a multi-core architecture utilizing GPU-accelerated hardware, where forward propagation computation is performed simultaneously on all requests within the target inference request batch. Matrix operations of multiple requests within the batch are merged into one or more large-scale parallel computation tasks. One optional implementation involves filling the input lexical sequences of all requests in the target inference request batch to a uniform length, stacking them along the batch dimension into a three-dimensional tensor, copying it to the GPU's video memory, calling the compiled inference kernel to perform a single forward propagation, and outputting the tensor containing the probability distribution of the next lexical term corresponding to each request within the batch. After sampling, the current inference result for each request is obtained. Another optional implementation, for inference engines supporting dynamic batch processing, involves merging the target inference request batch into a continuously executing batch processing task on the GPU, utilizing idle computing resources to process newly arriving requests in parallel.

[0103] By setting the inference hardware as accelerated hardware for the graphics processor and limiting the processing method to parallel processing, the potential for batch inference to adapt to the hardware architecture is fully explored. By using intra-batch parallel processing, the marginal computation cost of low-priority requests is compressed to an extremely low level, achieving a dual optimization of service quality assurance and hardware resource efficiency.

[0104] See Figure 2 , Figure 2 A flowchart of a question-and-answer method provided in one embodiment of this specification is shown, which specifically includes the following steps.

[0105] Obtain the initial batch of question and answer requests, which includes multiple question and answer requests from different question and answer services. The question and answer services include at least one of course type, subject type, and user information. The priority of each question and answer request is determined based on the question and answer service corresponding to each question and answer request. Based on the priority of each question and answer request in the initial question and answer request batch, the initial question and answer request batch is sampled to obtain the target question and answer request batch; Using inference hardware, batches of target question-and-answer requests are processed to obtain batches of target inference results corresponding to the batches of target question-and-answer requests.

[0106] A question-and-answer request is a task unit submitted by a user to an intelligent problem-solving service through an educational application or webpage, containing the content of the question to be solved and related context. For example, a question-and-answer request may be a math multiple-choice question that includes a stem, options, and an image; for example, a question-and-answer request may be a physics calculation problem that requires step-by-step derivation.

[0107] The question-and-answer service is categorized into different service types based on educational scenarios, subject areas, or user attributes. Each service type is associated with specific service level objectives and resource allocation strategies. For example, the question-and-answer service could be categorized by course stage (elementary school mathematics, middle school physics, high school chemistry); it could be categorized by subject type (science problem-solving and humanities tutoring); or it could be categorized by user information (fast track for members and standard service for regular users).

[0108] User information refers to attribute data associated with the user account that initiated the question-and-answer request, including user level, usage duration, etc. Its purpose is to serve as one of the reference dimensions for determining priority, ensuring that high-value or highly active users receive a better service experience. For example, user information could be a membership level identifier; for example, user information could be statistics on the cumulative number of questions asked and the accuracy rate; for example, user information could be information about the user's school and grade level.

[0109] The initial question-and-answer request batch is a temporary set of multiple pending question-and-answer requests extracted all at once from the question-and-answer request queue. Its purpose is to provide a unified data input unit for subsequent priority sampling and batch processing inference.

[0110] The target question-and-answer request batch is a new batch that is ultimately delivered to the inference hardware for execution. This batch is formed by prioritizing and sampling low-priority questions from the initial question-and-answer request batch, merging all high-priority question requests with the sampled low-priority question requests. Its purpose is to ensure that high-value or time-sensitive questions are processed immediately, while introducing low-priority questions in a controlled proportion to maintain batch processing efficiency.

[0111] To obtain the initial batch of question-and-answer requests, one optional implementation is that the question-and-answer service platform maintains a first-in, first-out global request queue. A scheduling thread pulls question-and-answer requests from the head of the queue in a fixed batch size, assembling them into an initial batch each time a specified number of requests are retrieved. Another optional implementation is to dynamically calculate the optimal batch size based on the current number of online users and the backlog of questions, and then pull requests of different course types from the queue according to a weighted ratio to form the initial batch.

[0112] The priority of each question and answer request is determined based on the question and answer business corresponding to each request. One possible implementation is that the server maintains a business-priority mapping table and looks up the corresponding priority value according to the course type, subject type or user information carried in the request. Another possible implementation is that the user directly specifies the urgency of the question through parameters when initiating the request, and the server verifies it and uses it as the priority basis.

[0113] Based on the priority of each question and answer request in the initial question and answer request batch, the initial question and answer request batch is sampled to obtain the target question and answer request batch. One optional implementation is to directly retain all requests with priority 0 in the initial batch by placing them in a high-priority set, and to randomly sample requests with priority greater than or equal to 1 according to a fixed sampling ratio corresponding to their respective priorities. The retained requests are then merged with the high-priority requests to form the target question and answer request batch. Another optional implementation is to dynamically adjust the sampling ratio of each priority level according to the current system load and service level agreement (SLA) status, proactively reducing the sampling rate of low-priority questions when high-priority requests surge.

[0114] Using inference hardware, batches of target question-and-answer requests are processed to obtain corresponding batches of target inference results. One optional implementation involves concatenating the question text of each request in the batch with the generated solution steps into an input sequence, loading it as a tensor into the GPU memory, and then calling a large language model to perform parallel forward propagation computation to generate the next lexical unit for each request in the current iteration step. Another optional implementation involves removing requests with complete answers from the batch and returning the final answer to the user client via a platform interface.

[0115] In one implementation of this specification, in an embodiment of an intelligent problem-solving platform for primary and secondary school education, the platform simultaneously supports three types of question-and-answer services: challenging math problems submitted by members of the competition sprint camp, requiring extremely low latency and high precision, mapped to priority 0; homework problems submitted by ordinary members, tolerating second-level latency, mapped to priority 1; and practice problem correction requests submitted by free users, not sensitive to real-time performance, mapped to priority 2. The scheduling thread pulls question-and-answer requests from the head of the queue in a fixed batch size of 48. In this instance, 12 priority 0 requests, 20 priority 1 requests, and 16 priority 2 requests are obtained, forming the initial question-and-answer request batch. The system directly retains the 12 priority 0 requests by placing them in the high-priority set; randomly samples the 20 priority 1 requests at a 1 / 2 ratio, retaining 10; and randomly samples the 16 priority 2 requests at a 1 / 3 ratio, retaining 5. The 15 low-priority requests retained after sampling are merged with the 12 high-priority requests to form a target question-and-answer request batch of 27 requests, which are then delivered to the graphics processor for parallel inference computation.

[0116] In one implementation of this specification, within an intelligent question bank platform for professional qualification exam preparation, the platform determines request priorities based on user information and subject type. Requests submitted by users who registered indicating "taking the exam this month" are automatically marked as priority 0; requests submitted by users who registered indicating "taking the exam in three months" are marked as priority 1; and requests submitted by trial users who did not specify an exam plan are marked as priority 2. Simultaneously, all users' "Case Analysis" subjective questions (subject type identifier) ​​are prioritized one level above their base priority. The scheduling module aggregates requests every 200 milliseconds using a time window strategy. Within this window, 42 question-and-answer requests are collected. After assigning priorities based on user information and subject type, the system constructs an initial batch and executes subsequent sampling and inference processes.

[0117] By applying the priority-based batch sampling inference method to question-and-answer scenarios, the platform achieves multi-layered service quality assurance for educational businesses. It assigns precise priority tags to each question-and-answer request based on multiple factors such as course urgency, subject characteristics, and user value level, and performs differentiated sampling scheduling for mixed batches according to priority. This enables the educational service platform to serve both high-value paying users and a large number of free users effectively with limited computing power, achieving an optimal balance between business efficiency and user experience.

[0118] The following is in conjunction with the appendix Figure 3 This document uses the application of the reasoning method provided in this manual on a reasoning platform as an example to further illustrate the reasoning method. See also... Figure 3 , Figure 3 A flowchart of a reasoning method applied to an inference platform according to an embodiment of this specification is shown, which specifically includes the following steps.

[0119] Step 302: Receive inference requests and mark their priority; The inference service entry point receives inference requests from different businesses / tenants. Each request includes a business type and a priority identifier (which can be an integer n, where n=0 indicates the highest priority). Priority information can be obtained by attaching parameters when external business systems call the API (Application Programming Interface), or by the server querying the priority mapping table based on the business ID. Inference requests are then sent to the execution queue.

[0120] Step 304, build batch; Retrieve x requests from the run queue, where x is a preset parameter, such as 64 requests, and group the x requests into a batch.

[0121] Step 306, prioritize the splitting; In the constructed batch, based on the priority of the request metadata, the batch requests are divided into P0 class (high priority, all enter the inference execution stage) and Pn (n≥1) class (low priority, enter the sampling selection logic). Low priority business requests are sampled at a ratio of 1 / (n+1). If it is P1 class, the sampling ratio is one-half. If it is P2 class, the sampling ratio is one-third, and so on.

[0122] Step 308, inference execution.

[0123] The selected Pn request after sampling is merged with all P0 requests to form the final inference batch. The batch enters the GPU-accelerated hardware. Each request infers the next term. If the output term is "[Terminate]", the current request ends. Otherwise, the inference result and the request are concatenated and sent back to the running queue. Step 304 is repeated.

[0124] Within the same inference service and the same batch, the current requests to be inferred are divided into high-priority and low-priority services based on their service priorities. The low-priority requests are selected to enter the current batch for inference according to a pre-set or dynamically adjusted sampling ratio. This achieves differentiated TPOT control for different services while maintaining batch processing efficiency and a low TTFT.

[0125] Corresponding to the above method embodiments, this specification also provides inference platform embodiments. Figure 4 A schematic diagram of the architecture of an inference platform provided in one embodiment of this specification is shown. Figure 4 As shown, the platform includes: The inference request interface 402 is used to receive inference requests; Response unit 404 is used to execute the steps of any of the above methods in response to a reasoning request.

[0126] Figure 5 A structural block diagram of a computing device according to one embodiment of this specification is shown. The components of the computing device 500 include, but are not limited to, a memory 510 and a processor 520. The processor 520 is connected to the memory 510 via a bus 530, and a database 550 is used to store data.

[0127] The computing device 500 also includes an access device 540, which enables the computing device 500 to communicate via one or more networks 560. Examples of these networks include Public Switched Telephone Network (PSTN), Local Area Network (LAN), Wide Area Network (WAN), Personal Area Network (PAN), or combinations of communication networks such as the Internet. The access device 540 may include one or more of any type of wired or wireless network interface (e.g., a network interface card (NIC)), such as an IEEE 802.11 Wireless Local Area Network (WLAN) wireless interface, a Wi-MAX (Worldwide Interoperability for Microwave Access) interface, an Ethernet interface, a Universal Serial Bus (USB) interface, a cellular network interface, a Bluetooth interface, or a Near Field Communication (NFC) interface.

[0128] In one embodiment of this specification, the above-described components of the computing device 500 and Figure 5 Other components, not shown, can also be connected to each other, for example, via a bus. It should be understood that... Figure 5 The block diagram of the computing device shown is for illustrative purposes only and is not intended to limit the scope of this specification. Those skilled in the art can add or replace other components as needed.

[0129] The computing device 500 can be any type of stationary or mobile computing device, including mobile computers or mobile computing devices (e.g., tablet computers, personal digital assistants, laptop computers, notebook computers, netbooks, etc.), mobile phones (e.g., smartphones), wearable computing devices (e.g., smartwatches, smart glasses, etc.) or other types of mobile devices, or stationary computing devices such as desktop computers or personal computers (PCs). The computing device 500 can also be a mobile or stationary server.

[0130] The processor 520 is used to execute the following computer-executable instructions, which, when executed by the processor, implement the steps of the above-mentioned reasoning method and question-and-answer method.

[0131] The above is an illustrative scheme of a computing device according to this embodiment. It should be noted that the technical solution of this computing device and the technical solution of the above-described reasoning method belong to the same concept. For details not described in detail in the technical solution of the computing device, please refer to the description of the technical solutions of the above-described reasoning method and question-and-answer method.

[0132] An embodiment of this specification also provides a computer-readable storage medium storing computer-executable instructions that, when executed by a processor, implement the steps of the above-described reasoning method and question-and-answer method.

[0133] The above is an illustrative scheme of a computer-readable storage medium according to this embodiment. It should be noted that the technical solution of this storage medium belongs to the same concept as the technical solution of the reasoning method described above. For details not described in detail in the technical solution of the storage medium, please refer to the description of the technical solutions of the reasoning method and the question-and-answer method described above.

[0134] An embodiment of this specification also provides a computer program, wherein when the computer program is executed in a computer, it causes the computer to perform the steps of the above-described reasoning method and question-and-answer method.

[0135] The above is an illustrative example of a computer program according to this embodiment. It should be noted that the technical solution of this computer program and the technical solution of the above-described reasoning method belong to the same concept. For details not described in detail in the technical solution of the computer program, please refer to the descriptions of the technical solutions of the above-described reasoning method and question-and-answer method.

[0136] The foregoing has described specific embodiments of this specification. Other embodiments are within the scope of the appended claims. In some cases, the actions or steps recited in the claims may be performed in a different order than that shown in the embodiments and may still achieve the desired result. Furthermore, the processes depicted in the drawings do not necessarily require the specific or sequential order shown to achieve the desired result. In some embodiments, multitasking and parallel processing are possible or may be advantageous.

[0137] Computer instructions include computer program code, which can be in the form of source code, object code, executable files, or certain intermediate forms. Computer-readable media can include: any entity or device capable of carrying computer program code, recording media, USB flash drives, portable hard drives, magnetic disks, optical disks, computer memory, read-only memory (ROM), random access memory (RAM), electrical carrier signals, telecommunication signals, and software distribution media, etc. It should be noted that the content included in computer-readable media can be appropriately added or removed according to the requirements of patent practice. For example, in some regions, according to patent practice, computer-readable media do not include electrical carrier signals and telecommunication signals.

[0138] It should be noted that, for the sake of simplicity, the foregoing method embodiments are all described as a series of actions. However, those skilled in the art should understand that the embodiments in this specification are not limited to the described order of actions, because according to the embodiments in this specification, some steps can be performed in other orders or simultaneously. Furthermore, those skilled in the art should also understand that the embodiments described in this specification are all preferred embodiments, and the actions and modules involved are not necessarily essential to the embodiments in this specification.

[0139] In the above embodiments, the descriptions of each embodiment have different focuses. For parts not described in detail in a certain embodiment, please refer to the relevant descriptions of other embodiments.

[0140] The preferred embodiments disclosed above are merely illustrative of this specification. The optional embodiments do not exhaustively describe all details, nor do they limit the invention to specific implementations. Clearly, many modifications and variations can be made based on the embodiments described in this specification. These embodiments are selected and specifically described in this specification to better explain the principles and practical applications of the embodiments, thereby enabling those skilled in the art to better understand and utilize this specification. This specification is limited only by the claims and their full scope and equivalents.

Claims

1. A reasoning method, characterized in that, include: Obtain an initial inference request batch, wherein the initial inference request batch includes multiple inference requests from different services, and the priority of each inference request is determined based on the service corresponding to each inference request; Based on the priority of each inference request in the initial inference request batch, the initial inference request batch is sampled to obtain the target inference request batch; Using inference hardware, the target inference request batch is processed to obtain the target inference result batch corresponding to the target inference request batch.

2. The method according to claim 1, characterized in that, The step of sampling the initial inference request batch based on the priority of each inference request in the initial inference request batch to obtain the target inference request batch includes: Based on the priority of each inference request in the initial inference request batch, each inference request in the initial inference request batch is divided into high-priority inference requests and low-priority inference requests. The low-priority inference requests in the initial inference request batch are sampled to obtain the sampled low-priority inference requests; A target inference request batch is constructed based on the sampled low-priority inference requests and the high-priority inference requests in the initial inference request batch.

3. The method according to claim 2, characterized in that, The step of sampling low-priority inference requests in the initial inference request batch to obtain sampled low-priority inference requests includes: Based on the priority of each low-priority inference request, the sampling ratio of each low-priority inference request is determined. According to the sampling ratio, the low-priority inference requests in the initial inference request batch are sampled to obtain the sampled low-priority inference requests.

4. The method according to claim 3, characterized in that, The low priority includes multiple priority levels; The step of determining the sampling ratio of each low-priority inference request based on its priority includes: Based on the priority level to which each low-priority inference request belongs, determine the sampling ratio of multiple priority levels for each low-priority inference request; The step of sampling low-priority inference requests in the initial inference request batch according to the sampling ratio to obtain sampled low-priority inference requests includes: According to the sampling ratio of the multiple priority levels, the low-priority inference requests in the initial inference request batch are sampled in a stratified manner to obtain the sampled low-priority inference requests.

5. The method according to any one of claims 1-4, characterized in that, The process of obtaining the initial inference request batch includes: Retrieve multiple inference requests from different business units from the inference request queue; Based on the business corresponding to each inference request, determine the priority of each inference request; An initial batch of inference requests is constructed based on each inference request and its priority.

6. The method according to claim 5, characterized in that, The step of processing the target inference request batch using inference hardware to obtain the target inference result batch corresponding to the target inference request batch includes: Using inference hardware, word-by-word processing is performed on the first inference request in the target inference request batch to obtain the current inference result in the first inference request, wherein the first inference request is any one of the inference requests in the target inference request batch; Based on the first inference request and the current inference result, determine the updated first inference request; Add the first inference request to the inference request queue, and return to execute the step of obtaining multiple inference requests from different services from the inference request queue; Until the processing of the last word in the first inference request is completed, the target inference result in the first inference request is obtained; The process continues until the last word in each reasoning request is processed, thus obtaining the target reasoning result batch corresponding to the target reasoning request batch.

7. The method according to claim 6, characterized in that, Determining the updated first inference request based on the first inference request and the current inference result includes: By concatenating the first inference request and the current inference result, an updated first inference request is obtained.

8. The method according to claim 6, characterized in that, The last word is the terminator word.

9. The method according to claim 1, characterized in that, The inference hardware is acceleration hardware for a graphics processor; Using inference hardware, the target inference request batch is processed to obtain the target inference result batch corresponding to the target inference request batch, including: By utilizing the accelerated hardware of the graphics processor, each inference request in the target inference request batch is processed in parallel to obtain the target inference result batch corresponding to the target inference request batch.

10. A question-and-answer method, characterized in that, include: Obtain an initial batch of question and answer requests, wherein the initial batch of question and answer requests includes multiple question and answer requests from different question and answer services, wherein the question and answer services include at least one of course type, subject type and user information, and the priority of each question and answer request is determined based on the question and answer service corresponding to each question and answer request; Based on the priority of each question and answer request in the initial question and answer request batch, the initial question and answer request batch is sampled to obtain the target question and answer request batch; Using inference hardware, the batch of target question-and-answer requests is processed to obtain the batch of target inference results corresponding to the batch of target question-and-answer requests.

11. An inference platform, comprising an inference request interface and a response unit; The inference request interface is used to receive inference requests; The response unit is configured to perform the steps of the method according to any one of claims 1-9 in response to the reasoning request.

12. A computing device, characterized in that, include: Memory and processor; The memory is used to store computer-executable instructions, and the processor is used to execute the computer-executable instructions, which, when executed by the processor, implement the steps of the method according to any one of claims 1-10.

13. A computer-readable storage medium, characterized in that, It stores computer-executable instructions that, when executed by a processor, implement the steps of the method according to any one of claims 1-10.

14. A computer program product, characterized in that, Includes a computer program / instructions that, when executed by a processor, implement the steps of the method according to any one of claims 1-10.