Request processing method and device, electronic equipment and storage medium

By dynamically adjusting the ratio of pre-filling and decoding devices for large language models, the problem of hardware resource allocation mismatch in the existing technology is solved, hardware resource utilization and request processing efficiency are improved, and the computing requirements of different batches of requests are met.

CN120723461APending Publication Date: 2025-09-30KUNWANG (SHANGHAI) TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510899498.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-06-30
Publication Date
2025-09-30

AI Technical Summary

Technical Problem

Existing separate large language model deployment solutions fail to effectively and dynamically adjust the ratio of pre-population and decoding devices, resulting in mismatched hardware resource allocation when processing different batches of requests, affecting the achievement of throughput and service level objectives.

Method used

By dynamically determining the ratio of pre-filled devices and decoding devices based on the target scale information of pending requests, and optimizing device allocation using autoregressive moving average technology and performance database, dynamic scheduling of devices is achieved.

Benefits of technology

It improves the utilization of hardware resources, reduces resource waste, enhances request processing efficiency and user experience, and meets the computing needs of different batches of requests.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120723461A_ABST
    Figure CN120723461A_ABST
Patent Text Reader

Abstract

The invention discloses a request processing method and device, electronic equipment and a storage medium. The invention provides a request processing method, and relates to the technical field of artificial intelligence, in particular to the technical field of large models, artificial intelligence acceleration cards and chips. According to the specific implementation scheme, a target proportion is determined according to target scale information used for at least one to-be-processed request, and the target scale information is determined according to historical scale information of historical requests before the to-be-processed requests; and determining at least one target pre-filling device and at least one target decoding device for the at least one to-be-processed request in the plurality of candidate devices, wherein the quantity proportion between the quantity of the target pre-filling devices and the quantity of the target decoding devices is consistent with the target proportion. The invention further provides a request processing device, electronic equipment and a storage medium.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present disclosure relates to the field of artificial intelligence technology, and in particular to the fields of large models, artificial intelligence acceleration cards, and chip technologies. More specifically, the present disclosure provides a request processing method, apparatus, electronic device, and storage medium. Background Art

[0002] With the development of artificial intelligence technology, the application of large language models (LLMs) is increasing. Large language models have a large number of parameters and can handle text-based requests as well as multimodal requests. Summary of the Invention

[0003] The present disclosure provides a request processing method, apparatus, device, and storage medium.

[0004] According to one aspect of the present disclosure, a request processing method is provided, the method comprising: determining a target ratio based on target scale information for at least one pending request, the target scale information being determined based on historical scale information of historical requests prior to the pending request; and determining at least one target pre-population device and at least one target decoding device for the at least one pending request from a plurality of candidate devices, the quantity ratio between the number of target pre-population devices and the number of target decoding devices being consistent with the target ratio.

[0005] According to another aspect of the present disclosure, a request processing device is provided, comprising: a first determination module for determining a target ratio based on target scale information for at least one pending request, where the target scale information is determined based on historical scale information of historical requests before the pending request; and a second determination module for determining at least one target pre-filled device and at least one target decoding device for the at least one pending request from a plurality of candidate devices, where a quantity ratio between the number of target pre-filled devices and the number of target decoding devices is consistent with the target ratio.

[0006] According to another aspect of the present disclosure, an electronic device is provided, comprising: at least one processor; and a memory communicatively connected to the at least one processor; wherein the memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor to enable the at least one processor to execute the method provided according to the present disclosure.

[0007] According to another aspect of the present disclosure, a non-transitory computer-readable storage medium storing computer instructions is provided. The computer instructions are used to cause a computer to execute the method provided according to the present disclosure.

[0008] According to another aspect of the present disclosure, a computer program product is provided, including a computer program, which implements the method provided according to the present disclosure when executed by a processor.

[0009] It should be understood that the contents described in this section are not intended to identify the key or important features of the embodiments of the present disclosure, nor are they intended to limit the scope of the present disclosure. Other features of the present disclosure will become readily understood through the following description. BRIEF DESCRIPTION OF THE DRAWINGS

[0010] The accompanying drawings are provided to facilitate a better understanding of the present invention and do not constitute a limitation of the present disclosure.

[0011] Figure 1 is a flowchart of a request processing method according to an embodiment of the present disclosure;

[0012] Figure 2 is a schematic diagram of a system architecture to which a request processing method according to an embodiment of the present disclosure can be applied;

[0013] Figure 3 is a schematic flow chart of determining a performance database according to one embodiment of the present disclosure;

[0014] Figure 4 is a schematic flow chart of a request processing method according to an embodiment of the present disclosure;

[0015] Figure 5 is a block diagram of a request processing apparatus according to an embodiment of the present disclosure; and

[0016] Figure 6 is a block diagram of an electronic device to which a request processing method according to an embodiment of the present disclosure can be applied. DETAILED DESCRIPTION

[0017] The following description of exemplary embodiments of the present disclosure is made in conjunction with the accompanying drawings, including various details of the embodiments of the present disclosure to facilitate understanding. These details should be considered as merely exemplary. Therefore, those skilled in the art will recognize that various changes and modifications may be made to the embodiments described herein without departing from the scope and spirit of the present disclosure. Similarly, for the sake of clarity and conciseness, descriptions of well-known functions and structures are omitted in the following description.

[0018] After the launch of large conversational models, large language models have rapidly permeated every aspect of people's lives and learning. With the continuous iterative development of large language models, and the increasing number of model parameters, user requirements for model responsiveness are also increasing. To maximize the utilization of AI accelerator cards, improve overall server cluster throughput, and reduce latency, various effective inference optimization methods can be employed to optimize model results, key-value cache (KVCache) management, operators, and other aspects. These inference optimization methods include quantization, dynamic batching, and paged attention.

[0019] Large language models are primarily based on the Transformer architecture. Due to significant differences in computational characteristics, the request processing phases of large language models can be divided into the prefill phase and the decode phase. Prefill is a compute-bound operation. When processing requests with small batch sizes or long requests, the prefill calculation can utilize all computing resources to meet the Time To First Token (TTFT) requirement. A token can be a label. Decoding is a memory-bound operation. Decoding requires larger batch sizes to increase computational intensity and meet the Time Per Output Token (TPOT) requirement. A request can be a user-provided query. This query can consist of one or more characters or data in various modalities, such as images, audio, text, and video.

[0020] In some large language model service architectures, the computations for the pre-population and decoding phases can be combined. However, these two phases must meet different service level objectives (SLOs). Combining the computations can cause the two phases to interfere with each other, limiting the overall throughput of the service.

[0021] In other large language model service architectures, the pre-population stage can be separated from the decoding stage to achieve a separate deployment. Furthermore, different numbers of AI chips and model parallel strategies can be configured for the pre-population stage and the decoding stage, respectively, to resolve the issue of mutual interference between the two stages caused by merged computing. Given the same hardware resources, separate deployment can significantly improve throughput and service level objectives compared to merged computing. Separate deployment can achieve higher service levels with lower hardware costs. It is understood that AI chips can include various chips, such as general-purpose graphics processing units (GPGPUs), tensor processing units (TPUs), and neural network processing units (NPUs).

[0022] However, decentralized deployments are often implemented using small clusters of several servers. In real-world business scenarios, the scale and characteristics of requests are complex and varied. For a batch of requests with long inputs and short outputs, more pre-population nodes and fewer decoding nodes are required. Conversely, if the overall characteristics of requests over a certain period are short inputs and long outputs, fewer pre-population nodes and more decoding nodes are required. It can be understood that a pre-population node can be a server that performs pre-population calculations, while a decoding node can be a server that performs decoding calculations. Requests with long inputs and short outputs can be requests that require a large model to determine the correctness of an argument. Requests with short inputs and long outputs can be requests that require a large model to generate an article.

[0023] The pre-filling stage and the decoding stage have different computational characteristics. If the same AI server is used to perform both the pre-filling and decoding stages, it will inevitably lead to the sharing of hardware resources and parallelism settings between the two stages. However, both stages have their own unique computational characteristics and latency requirements, and different resource allocations are required to achieve better service level objectives and throughput. In some separate deployment solutions, most of them focus on the scheduling of key-value caches. By allocating the pre-filling and decoding stages to different servers to avoid mutual interference, and combining features such as prefix caching, prefix sharing between different pre-filling nodes, and elastic memory pools, latency can be further reduced and throughput improved.

[0024] In a decoupled deployment scenario, a key-value cache-centric disaggregated architecture can be employed to effectively separate the pre-population and decoding clusters. This disaggregated architecture fully utilizes the hardware resources of the GPU cluster to build an efficient key-value cache system and significantly improves overall system performance by optimizing resource allocation. The core component of this disaggregated architecture is a key-value cache-centric scheduler, designed to maximize overall effective throughput while strictly adhering to latency-related service level objectives. Contrary to the assumption that all requests are processed, this disaggregated architecture employs a predictive early rejection strategy to mitigate the impact of overload on service quality. This disaggregated architecture excels in long-context processing scenarios. While maintaining service level objectives, it significantly improves overall throughput. However, this disaggregated architecture does not account for the performance differences in the hardware of heterogeneous servers in the distributed cluster, nor does it consider the varying phases of requests, which can lead to significantly different overall request volumes at different times. Consequently, it cannot dynamically adjust the number or ratio of pre-population or decoding nodes in the cluster.

[0025] In other separate deployment schemes, optimization can be performed based on different directions, such as chunk attention, prefix cache sharing between nodes, length-based scheduling modules, and elastic memory pools.

[0026] However, most of these separate deployment solutions are static designs. Before the service is running, the ratio of pre-filled nodes to decoding nodes is fixed based on factors such as service level goals and server performance. These solutions do not take into account the periodic differences in user request characteristics and cannot dynamically adjust the ratio of pre-filled nodes to decoding nodes based on real-time request scale characteristics and service level goals.

[0027] In a few separate deployment schemes, when new requests exceed a threshold, servers are gradually allocated based on the number of requests in the pre-population phase, without pre-scheduling based on global considerations.

[0028] Therefore, in order to process requests more efficiently, the present disclosure provides a request processing method, which will be described below.

[0029] Figure 1 is a flowchart of a request processing method according to an embodiment of the present disclosure.

[0030] like Figure 1 As shown, the method 100 may include operations S110 to S120.

[0031] In operation S110 , a target ratio is determined based on target size information for at least one pending request.

[0032] In an embodiment of the present disclosure, when processing the i-th batch of requests, scale information for the i-th batch of requests may be determined. The i-th batch of requests may serve as at least one pending request. The scale information for the i-th batch of requests may serve as target scale information for the at least one pending request. i may be an integer greater than 1. The time period for processing the i-th batch of requests may be the t-th time period. t may be an integer greater than 1.

[0033] In an embodiment of the present disclosure, the target scale information may indicate the scale of requests to be processed.

[0034] In the disclosed embodiments, target scale information is determined based on historical scale information of requests prior to the pending request. The historical scale information may indicate the scale of one or more historical requests in the same historical batch. Multiple batches of historical requests may have been processed prior to one or more pending requests, and the target scale information may be determined based on one, multiple, or all of the multiple historical scale information for all historical batches. For example, for requests in batch i, requests in batch i-1 may be considered as requests in the same historical batch.

[0035] In an embodiment of the present disclosure, the target ratio may be a ratio between the number of pre-populated devices and the number of inference devices.

[0036] In operation S120 , at least one target pre-filling device and at least one target decoding device for at least one pending request are determined among a plurality of candidate devices.

[0037] In the embodiments of the present disclosure, multiple candidate devices may have different AI chips. For example, one or more candidate devices may have a general-purpose graphics processing unit (GPU). One or more candidate devices may have a neural network processing unit (NPU).

[0038] In the disclosed embodiment, the ratio between the target number of pre-populated devices and the target number of decoding devices is consistent with the target ratio. For example, if the target ratio is the number of pre-populated devices: the number of decoding devices, the ratio can be the target number of pre-populated devices: the target number of decoding devices.

[0039] Through the disclosed embodiments, before processing pending requests, the ratio of pre-population devices to decoding devices is determined based on the scale of pending requests, enabling dynamic device scheduling when processing requests in different time periods and batches. Based on the scale of pending requests, a more appropriate number of pre-population devices and decoding devices can be determined, enabling adaptive and efficient processing of different batches of requests, improving request processing efficiency and effectively enhancing the user experience with lower hardware resource overhead.

[0040] It can be understood that the above describes the method of the present disclosure, and the following will further describe the system architecture of the present disclosure.

[0041] Figure 2 FIG. 4 is a schematic diagram of a system architecture to which a request processing method according to an embodiment of the present disclosure can be applied.

[0042] like Figure 2 As shown, the system architecture may include a global scheduler gs20 , a request analysis unit ra20 , a prefill scheduler ps20 , and a decode scheduler ds20 .

[0043] The global scheduling unit GS20 may receive one batch of requests during each time period. For example, during the t-1 period, the global scheduling unit GS20 may receive a request for the i-1 batch. The global scheduling unit GS20 may also provide the ratio of pre-filling devices to decoding devices for the i-1 batch to the pre-fill scheduling unit PS20 and the decoding scheduling unit DS20. Based on the ratio of pre-filling devices to decoding devices for the i-1 batch, the pre-fill scheduling unit PS20 may determine one or more pre-filling devices to perform pre-fill calculations corresponding to the i-1 batch requests, and the decoding scheduling unit DS20 may determine one or more decoding devices to perform decoding calculations corresponding to the i-1 batch requests.

[0044] The request analysis unit ra20 can determine the scale of requests in batch i-1. This scale information can include parameters such as service level objectives, input size, output size, and queries per second (QPS). The input size can be the number of characters or tokens in the request provided by the user. The output size can be the number of characters or tokens in the request processing result provided by the model. The QPS can be the number of questions received per second during the period.

[0045] The request analysis unit ra20 may also predict the scale information of the requests in the i-th batch. The i-th batch of requests may include one or more requests received during the t-th time period. The scale information of the i-th batch of requests may serve as target scale information. The following describes how the request analysis unit ra20 determines the target scale information.

[0046] In some embodiments, the target scale information may be determined by performing the following operations based on the historical scale information of historical requests before the pending request: performing weighted fusion on multiple historical scale information based on multiple autoregressive coefficients for multiple historical scale information to obtain fused scale information. Performing weighted fusion on multiple initial noise information based on multiple moving average coefficients for multiple initial noise information to obtain fused noise information. Determine the target scale information based on preset parameters, fused scale information, and fused noise information. For example, the requests of the i-th batch may include one or more pending requests. The historical requests before the requests of the i-th batch may include the requests of the i-1-th batch. For the requests of the i-th batch, the scale information of the requests of the i-1-th batch may be used as historical scale information. The target scale information may be determined by the following formula:

[0047] (Formula 1)

[0048] The scale information of the i-th batch of requests received in the t-th time period may be used as the target scale information. Can be preset parameters. The scale information of the ijth batch of requests received in the tjth period may be used as the historical scale information of the tjth period. It can be the autoregressive coefficient of the historical scale information for the tj-th period. It can be the initial noise information of the tjth period. It can be the moving average coefficient of the initial noise information for the tj-th period. It can be used to integrate scale information. It can be used to fuse noise information.

[0049] j can be an integer greater than or equal to 1. p and q can be integers less than or equal to i-1. Through the disclosed embodiments, based on historical scale information, the Autoregressive-Moving Average (ARMA) technique can be used to accurately determine target scale information, thereby efficiently implementing dynamic scheduling of hardware resources for processing requests. This allows for more efficient and accurate allocation of sufficient and appropriate hardware resources to pending requests.

[0050] After determining the target scale information, the request analysis unit ra20 may further determine a target ratio for at least one pending request based on the performance database db20, as will be described below.

[0051] In some embodiments, in some implementations of the above operation S110, determining the target ratio based on the target scale information for at least one pending request includes: determining, from multiple candidate scale information, at least one filtered scale information whose scale similarity with the target scale information satisfies a preset similarity condition. The multiple candidate scale information pieces correspond one-to-one with multiple preset ratios. The target ratio can be determined from the at least one preset ratio corresponding to the at least one filtered scale information piece. For example, performance database db20 may include multiple candidate scale information pieces and multiple preset ratios for the multiple candidate scale information pieces. A higher throughput can be achieved based on the ratio of pre-filled devices to decoding devices indicated by the preset ratios. Multiple scale similarities between the target scale information and the multiple candidate scale information pieces can be determined using various methods. The scale similarity can include cosine similarity, Euclidean distance, and so on. The preset similarity condition can indicate selecting the candidate scale information with the highest scale similarity. Based on the preset similarity condition, the candidate scale information with the highest scale similarity can be selected as the filtered scale information piece. The preset ratio corresponding to the filtered scale information piece can be used as the target ratio. Next, based on the target ratio, the plurality of candidate pre-population devices among the plurality of candidate devices, and the plurality of candidate decoding devices, the above operation S120 may be performed.

[0052] In some embodiments, in some implementations of the above operation S120, at least one target pre-filled device for at least one pending request is determined from a plurality of candidate pre-filled devices. Figure 2 As shown, the pre-fill scheduling unit ps20 can determine, according to the target ratio, a target pre-fill device pd201, a target pre-fill device pd202, a target pre-fill device pd203, and a target pre-fill device pd204 among a plurality of candidate pre-fill devices.

[0053] In some embodiments, in some implementations of the above operation S120, a target decoding device for at least one pending request is determined among a plurality of candidate decoding devices. Figure 2 As shown, the decoding scheduling unit ds20 can determine, based on the target ratio, the target decoding device dd201 and the target decoding device dd202 from among multiple candidate decoding devices. The ratio between the number of target pre-populated devices and the number of target decoding devices is consistent with the target ratio. After the decoding scheduling unit ds20 determines the target decoding device, it can provide the identifier of the target decoding device to the pre-population scheduling unit ps20 so that the key-value cache data can be provided to the target decoding device.

[0054] After determining the target pre-filling device and the target decoding device for the pending request, the pending request may be processed using the target pre-filling device and the target decoding device, as will be described below.

[0055] In some embodiments, the method 100 may further include: processing pending requests using a target pre-population device to obtain pending cached data; providing the pending cached data to a target decoding device for the pending requests; and decoding the pending cached data using the target decoding device to obtain a processing result for the pending requests. For example, the target pre-population device PD201 may process one or more requests from the i-th batch of requests. After the target pre-population device PD201 completes processing of a request from the i-th batch of requests, the key-value cached data for the request may be obtained as the pending cached data for the request. For example, the target pre-population device PD201 may provide the requested key-value cached data to the target decoding device DD201 based on the Transmission Control Protocol (TCP) or Remote Direct Memory Access (RDMA). The target decoding device DD201 may have an artificial intelligence chip different from that of the target pre-population device PD201. A data normalizer DN201 may convert the key-value cached data into a format acceptable to the target decoding device DD201. Next, the target decoding device dd201 can decode the converted key-value cache data to obtain the processing result of the request. Through the disclosed embodiment, a batch of pending requests can be processed based on pre-populated devices and decoding devices that match the batch of pending requests, which can fully utilize the computing power and storage capacity of the hardware devices, reduce hardware resource waste, and effectively improve hardware resource utilization.

[0056] It is understood that the above describes the system architecture of the present disclosure, and the following will describe the method of determining the performance database.

[0057] Figure 3 is a schematic flow chart of determining a performance database according to one embodiment of the present disclosure.

[0058] like Figure 3 As shown, method 300 may include operations S301 to S307 performed before operation S210.

[0059] In operation S301 , a value range of a constraint factor is determined.

[0060] In some embodiments, a predetermined number of devices may be provided to establish a performance database. The predetermined number of devices may include devices with different chips. For example, in the process of establishing a performance database, one or more general-purpose graphics processing unit-based devices may be used, and one or more neural network processing unit-based devices may also be used. In one example, two general-purpose graphics processing unit-based devices and four neural network processing unit-based devices may be used.

[0061] In some embodiments, the constraint factors may be one or more. For example, taking multiple constraint factors as an example, the multiple constraint factors may include a service level target, input size, output size, number of queries per second, and an initial ratio. The initial ratio may be the ratio of pre-populated devices to decoding devices. Among the multiple constraint factors, the service level target, input size, output size, and number of queries per second may serve as the first independent variable, and the initial ratio may serve as the second independent variable. A numerical range may be set for each constraint factor. In addition, the number of high-quality processing results (goodputs) output after a predetermined number of devices process the request may be used as the dependent variable. Next, each constraint factor in the multiple constraint factors may be traversed.

[0062] In operation S302 , it is determined whether all optional constraint factors have been traversed.

[0063] In some embodiments, in response to determining that all optional constraint factors have not been traversed, operation S303 is performed. The optional constraint factors may be the one or more first independent variables mentioned above.

[0064] In operation S303 , variables to be tested are determined.

[0065] In some embodiments, a variable to be tested may be determined from one or more first independent variables. For example, a service level target may be used as a variable to be tested.

[0066] In operation S304 , it is determined whether all available variable values ​​are acquired from the numerical interval for the variable to be tested based on the preset step size for the variable to be tested.

[0067] In some embodiments, in response to determining that not all available variable values ​​have been obtained from the numerical interval for the variable to be tested based on the preset step size for the variable to be tested, operation S305 is performed. For example, the numerical interval for the service level target can be 90% to 99%. The preset step size for the service level target can be 1%. All available variable values ​​can include 90%, 91%, 92%, ..., 99%.

[0068] In operation S305 , based on a preset step size for the variable to be tested, a value of the variable to be tested is obtained from a numerical interval for the variable to be tested, and a target throughput indicator value corresponding to the value of the variable to be tested is tested.

[0069] For example, for a service level objective, 90% can be used as the value of the variable to be measured. Next, when the values ​​of the first independent variable, such as the input size and the output size, are preset values, multiple initial throughput index values ​​corresponding to the multiple initial ratios are determined. The largest initial throughput index value can be used as the target throughput index value. The initial ratio corresponding to the target throughput index value can be used as the preset ratio for the value of the variable to be measured.

[0070] In operation S306 , the test result corresponding to the value of the variable to be measured is stored.

[0071] For example, the value of the variable to be measured, the corresponding input size, output size, and other values ​​of the first independent variable, and the preset ratio for the value of the variable to be measured can be stored as a candidate scale information.

[0072] Next, the process may return to operation S304. For the service level target, there are still available variable values ​​such as 91%, ..., 99%, etc. Thus, operations S304 to S306 may be repeatedly performed until all available variable values ​​are obtained from the numerical interval for the service level target.

[0073] In some embodiments, in response to determining that all available variable values ​​have been obtained from the numerical interval for the variable to be tested based on the preset step size for the variable to be tested, the process returns to operation S302 to determine whether all constraints have been traversed. For example, after taking the service level target as the variable to be tested, it can be determined that constraints such as the input size and the output size have not been traversed. Next, operation S303 can be performed to determine the input size as the variable to be tested. Next, operations S304 to S305 can be repeated until all available variable values ​​have been obtained from the numerical interval for the input size. It can be understood that the description of the input size as the variable to be tested is the same or similar to the description of the service level target as the variable to be tested, and the present disclosure will not repeat them here.

[0074] It will be appreciated that the above description of the present disclosure uses the example of not traversing all optional constraints. However, the present disclosure is not limited thereto. After sequentially using the first independent variable, such as the input size and the output size, as the variable to be tested and completing the test, it can be determined that all optional constraints have been traversed.

[0075] In some embodiments, in response to traversing all optional constraints, operation S307 is performed, terminating the execution. After the execution is terminated, a performance database may be obtained. The performance database may include: multiple candidate scale information corresponding to one or more optional constraints, and multiple preset ratios for the multiple candidate scale information.

[0076] It can be understood that some methods of determining the performance database are described above, and other methods of determining scale similarity will be further described below.

[0077] In some embodiments, determining at least one selected scale information from multiple candidate scale information that satisfies a preset condition for scale similarity with the target scale information includes determining multiple scale similarities between the target scale vector and the multiple candidate scale vectors. The target scale vector is obtained based on the target scale information, and the candidate scale vectors are obtained based on the candidate scale information. For example, the target scale information and the multiple candidate scale information may be vectorized to obtain the target scale vector and the multiple candidate scale vectors. Subsequently, the scale similarities between the target scale vector and the candidate scale vectors may be calculated.

[0078] In some embodiments, the modulus of the target scale vector and the modulus of the candidate scale vectors can be determined. Next, a first scale vector and a second scale vector can be determined from the target scale vector and the candidate scale vectors. The first scale vector is the vector with the smaller modulus between the target scale vector and the candidate scale vector, and the second scale vector is the vector with the larger modulus between the target scale vector and the candidate scale vector.

[0079] In some embodiments, the length similarity may be obtained by dividing the first scale vector by the second scale vector, for example, by dividing the modulus of the first scale vector by the modulus of the second scale vector.

[0080] In some embodiments, the cosine similarity between the target scale vector and the candidate scale vector can also be determined. The scale similarity can be obtained by multiplying the cosine similarity and the length similarity. For example, the scale similarity S can be obtained by the following formula:

[0081] (Formula 2)

[0082] Can be a target scale vector. can be a candidate scale vector. A first scale vector among the target scale vector and the candidate scale vectors may be determined. A second scale vector among the target scale vector and the candidate scale vectors may be determined.

[0083] In some embodiments, among multiple candidate scale information, determining at least one filtered scale information whose scale similarity with the target scale information meets preset conditions includes: determining at least one filtered scale information that meets the preset similarity conditions among multiple candidate scale information based on multiple scale similarities. The preset similarity conditions include: selecting the maximum scale similarity; the scale similarity is greater than or equal to a preset scale similarity threshold. For example, the maximum scale similarity can be determined from multiple scale similarities. Based on the maximum scale similarity, the corresponding candidate scale information can be determined as the filtered scale information. Through the embodiment of the present disclosure, the scale similarity is determined by cosine similarity and length similarity, and accurate filtered scale information can be fully filtered based on the similarity of length, direction, and trend, and then the accurate ratio of pre-filled devices and decoding devices can be determined.

[0084] It can be understood that some methods of determining scale similarity are described above, and the method disclosed herein will be further described below.

[0085] Figure 4 is a schematic flowchart of a request processing method according to an embodiment of the present disclosure.

[0086] like Figure 4 As shown, at the t-1 time period, one or more requests req40 may be provided to the global scheduling unit gs40. After receiving the request, the global scheduling unit gs40 may perform operation S408.

[0087] In operation S408 , it is determined whether a target ratio of a target period is received.

[0088] In response to not receiving the target ratio for the target period, the request is provided to the pre-fill scheduling unit. For example, the target period may be a period subsequent to the current period. If the t-1 period is the current period, the ratio of pre-fill devices and decoding devices processing the request is already determined. Thus, the global scheduling unit GS40 may provide the request received in the t-1 period to the pre-fill scheduling unit PS40.

[0089] like Figure 4 As shown, the request analysis unit rs40 may determine scale information of one or more requests in the t-1 period. When the t-1 period is about to end, the request analysis unit rs40 may perform operation S409 based on the one or more historical scale information.

[0090] In operation S409, target scale information is determined. For example, the scale information of requests during period t can be determined as the target scale information. The one or more historical scale information can include the scale information of requests during period t-1 as the target scale information. It is understood that the scale information for period t can be determined using Formula 1 above, and this disclosure will not elaborate further here.

[0091] One or more requests in the tth time period may be used as one or more pending requests. After determining the target scale information, the request analysis unit ra40 may perform operation S410.

[0092] In operation S410, a target ratio is determined based on target scale information for at least one pending request. For example, the request analysis unit ra40 may determine the target ratio for one or more pending requests based on the performance database db40. It is understood that the performance database db40 may be constructed based on the aforementioned method 300. The above description regarding the request analysis unit ra20 determining the target ratio for at least one pending request also applies to the request analysis unit ra40 determining the target ratio for one or more pending requests, and is not further described herein.

[0093] like Figure 4 As shown, the target ratio can be provided to the global scheduling unit gs40, so that the global scheduling unit performs operation S408 to determine whether the target ratio of the target period is received. In response to receiving the target ratio of the target period, operation S421 is performed.

[0094] In operation S421, the target number of pre-populated devices and the target number of decoding devices are determined. For example, after receiving the target ratio for time period t, the target number of pre-populated devices and the target number of decoding devices can be determined based on the target ratio, the number of candidate pre-populated devices among the candidate devices, and the number of candidate decoding devices. The ratio between the target number of pre-populated devices and the target number of decoding devices is consistent with the target ratio. Next, after the end of time period t-1, time period t is entered.

[0095] The pre-fill scheduling unit ps40 may perform operation S422 based on the number of target pre-fill devices.

[0096] In operation S422 , at least one target pre-population device for the pending request is determined among a plurality of candidate pre-population devices.

[0097] In some embodiments, determining at least one target pre-population device for a pending request from a plurality of candidate pre-population devices includes: fusing a plurality of first scheduling weights for the plurality of candidate pre-population devices to obtain a first fused weight. The first scheduling weight is used to indicate the computing power of the candidate pre-population device. A first scheduling parameter that is less than or equal to the first fused weight is determined. Next, based on the first scheduling parameter, at least one target pre-population device that meets a first screening condition can be determined from the plurality of candidate pre-population devices. The first screening condition is that a first difference value for the candidate pre-population device is less than or equal to a first preset difference threshold. The first difference value is determined by the first scheduling parameter and at least one first scheduling weight for the at least one candidate pre-population device.

[0098] In some embodiments, determining, based on the first scheduling parameter, at least one target pre-filled device that satisfies a first screening condition from among a plurality of candidate pre-filled devices includes: determining, based on the first scheduling parameter, at least one screened pre-filled device that satisfies the first screening condition from among the plurality of candidate pre-filled devices; and determining, based on at least one first performance information indicating performance of the screened pre-filled device, at least one target pre-filled device from among the at least one screened pre-filled device.

[0099] For example, the multiple candidate pre-population devices include a general-purpose graphics processing unit (GPU)-based candidate pre-population device and a neural network processing unit (NPU)-based candidate pre-population device. The computing power of the general-purpose GPU is, for example, greater than that of the NPU. The first scheduling weight of the GPU-based candidate pre-population device may be greater than the first scheduling weight of the NPU-based candidate pre-population device. The multiple first scheduling weights are summed to obtain a first fusion weight. One or more random numbers less than or equal to the first fusion weight may be determined as one or more first scheduling parameters. Based on the first scheduling parameters, for example, multiple pre-population devices that meet a first screening condition may be determined. The first difference between the sum of the multiple first scheduling weights for the multiple pre-population devices and the corresponding first scheduling parameters is less than or equal to a first preset difference threshold. Next, metrics such as hardware resource utilization, power consumption, and temperature of the pre-population devices may be obtained to determine first performance information. Based on the first performance information, if a pre-population device among the multiple pre-population devices is determined to have lower hardware resource utilization, lower power consumption, and lower temperature, the pre-population device may be selected as the target pre-population device. It is understood that the first performance information may indicate the load of the pre-populated device. Lower hardware resource utilization, lower power consumption, and lower temperature indicate that the device load is lower. The device with the lower load is selected as the target pre-populated device.

[0100] In addition, before, simultaneously with, or after determining the target pre-population device, the decoding scheduling unit ds40 may perform operation S423 based on the number of target decoding devices.

[0101] In operation S423 , at least one target decoding device for the request to be processed is determined among a plurality of candidate decoding devices.

[0102] In some embodiments, determining at least one target decoding device for a pending request from a plurality of candidate decoding devices includes: fusing a plurality of second scheduling weights for the plurality of candidate decoding devices to obtain a second fused weight. The second scheduling weight is used to indicate the storage capacity of the candidate decoding device. A second scheduling parameter that is less than or equal to the second fused weight is determined. Based on the second scheduling parameter, at least one target decoding device that meets a second screening condition is determined from the plurality of candidate pre-populated devices. The second screening condition is that a second difference value for the candidate decoding device is less than or equal to a second preset difference threshold. The second difference value is determined by the second scheduling parameter and at least one second scheduling weight for the at least one candidate decoding device.

[0103] In some embodiments, determining at least one target decoding device that satisfies a second screening condition from among a plurality of candidate pre-populated devices based on the second scheduling parameter includes: determining at least one screened decoding device that satisfies the second screening condition from among the plurality of candidate decoding devices based on the second scheduling parameter, and determining at least one target decoding device from among the at least one screened decoding device based on at least one second performance information indicating performance of the screened decoding device.

[0104] For example, the multiple candidate decoding devices include a candidate decoding device based on a general-purpose graphics processing unit (GPU) and a candidate decoding device based on a neural network processing unit (NPU). The storage capacity of the general-purpose GPU may be smaller than that of the NPU. The second scheduling weight of the candidate decoding device based on the GPU may be smaller than the second scheduling weight of the candidate decoding device based on the NPU. The multiple second scheduling weights may be summed to obtain a second fusion weight. One or more random numbers less than or equal to the second fusion weight may be determined as one or more second scheduling parameters. Based on the second scheduling parameters, for example, multiple filtered decoding devices that meet a second screening condition may be determined. The second difference between the sum of the multiple second scheduling weights for the multiple filtered decoding devices and the corresponding second scheduling parameters may be less than or equal to a second preset difference threshold. Next, metrics such as hardware resource utilization, power consumption, and temperature of the filtered decoding devices may be obtained to determine second performance information. Based on the second performance information, if a filtered decoding device is determined to have lower hardware resource utilization, lower power consumption, and lower temperature among the multiple filtered decoding devices, the filtered decoding device may be selected as the target decoding device. It is understood that the second performance information can indicate the load of the decoding device. Low hardware resource utilization, low power consumption, and low temperature indicate that the device load is low. The device with the lower load is selected as the target decoding device. Through the embodiments of the present disclosure, candidate devices are screened based on scheduling weights. Even if different devices include different chips, sufficient and appropriate hardware resources can be accurately allocated for pre-filling and decoding. This can effectively improve hardware resource utilization, reduce hardware resource waste, and adapt to various practical scenarios.

[0105] Furthermore, after the target pre-fill device is determined, operation S431 may be performed.

[0106] In operation S431, the target pre-filling device processes the pending request, obtains the pending cache data, and provides the pending cache data to the target decoding device for the pending request. Figure 4 As shown, the target pre-fill device performs pre-fill calculations for the pending request, and the key-value cache data of the pending request can be obtained. The key-value cache data can be provided to the data conversion unit dn40 to convert the data format of the key-value cache data into a data format usable by the target decoding device. Next, operation S432 can be performed.

[0107] In operation S432, the target decoding device is used to decode the cached data to obtain a processing result of the request to be processed. For example, the target decoding device is used to perform a decoding calculation for the request to be processed to obtain a processing result of the request to be processed.

[0108] It is understood that during the process of processing the requests in the tth period, the scale information of the requests in the t+1th period and the ratio between the pre-filling devices and the decoding devices in the t+1th period can be determined. It is understood that the method for determining the scale information of the requests in the t+1th period and the ratio between the pre-filling devices and the decoding devices in the t+1th period is the same as or similar to the method for determining the scale information of the requests in the tth period and the ratio between the pre-filling devices and the decoding devices in the tth period, and the present disclosure will not be repeated here.

[0109] It can be understood that the method of the present disclosure is described above, and the device of the present disclosure will be described below.

[0110] Figure 5 is a block diagram of a request processing apparatus according to an embodiment of the present disclosure.

[0111] like Figure 5 As shown, the apparatus 500 may include a first determining module 510 and a second determining module 520 .

[0112] The first determining module 510 is configured to determine a target ratio based on target scale information for at least one pending request, wherein the target scale information is determined based on historical scale information of historical requests prior to the pending request.

[0113] The second determining module 520 is configured to determine at least one target pre-population device and at least one target decoding device for at least one pending request from a plurality of candidate devices, wherein the ratio between the number of target pre-population devices and the number of target decoding devices is consistent with a target ratio.

[0114] In some embodiments, the first determination module 510 includes: a first determination submodule for determining, from among a plurality of candidate scale information, at least one screened scale information whose scale similarity with the target scale information satisfies a preset similarity condition, wherein the plurality of candidate scale information corresponds one-to-one to a plurality of preset ratios; and a second determination submodule for determining a target ratio from among at least one preset ratio corresponding to the at least one screened scale information.

[0115] In some embodiments, the first determination submodule includes: a first determination unit configured to determine multiple scale similarities between a target scale vector and multiple candidate scale vectors, where the target scale vector is obtained based on target scale information, and the candidate scale vectors are obtained based on candidate scale information. A second determination unit configured to determine, based on the multiple scale similarities, at least one filtered scale information from the multiple candidate scale information that satisfies a preset similarity condition. The preset similarity condition includes: the scale similarity being greater than or equal to a preset scale similarity threshold.

[0116] In some embodiments, the first determination unit includes: a first determination subunit configured to determine the cosine similarity between the target scale vector and the candidate scale vector; a second determination subunit configured to divide the first scale vector by the second scale vector to obtain length similarity, where the first scale vector is the vector with the smaller modulus between the target scale vector and the candidate scale vector, and the second scale vector is the vector with the larger modulus between the target scale vector and the candidate scale vector; and a multiplication subunit configured to multiply the cosine similarity by the length similarity to obtain scale similarity.

[0117] In some embodiments, the target scale information is determined based on historical scale information of historical requests prior to the pending request by performing the following operations: performing weighted fusion of the multiple historical scale information based on multiple autoregressive coefficients for the multiple historical scale information to obtain fused scale information. Performing weighted fusion of the multiple initial noise information based on multiple moving average coefficients for the multiple initial noise information to obtain fused noise information. The target scale information is determined based on preset parameters, the fused scale information, and the fused noise information.

[0118] In some embodiments, the plurality of candidate devices includes a plurality of candidate pre-population devices and a plurality of candidate decoding devices. The second determination module 520 includes: a third determination submodule for determining at least one target pre-population device for at least one pending request from the plurality of candidate pre-population devices; and a fourth determination submodule for determining a target decoding device for at least one pending request from the plurality of candidate decoding devices.

[0119] In some embodiments, the third determination submodule includes: a first fusion unit configured to fuse multiple first scheduling weights for multiple candidate pre-population devices to obtain a first fused weight, where the first scheduling weight indicates the computing capacity of the candidate pre-population device; a third determination unit configured to determine a first scheduling parameter that is less than or equal to the first fused weight; and a fourth determination unit configured to determine, based on the first scheduling parameter, at least one target pre-population device from the multiple candidate pre-population devices that meets a first screening condition. The first screening condition is that a first difference value for the candidate pre-population device is less than or equal to a first preset difference threshold, where the first difference value is determined by the first scheduling parameter and at least one first scheduling weight for at least one candidate pre-population device.

[0120] In some embodiments, the fourth determination unit includes: a third determination subunit configured to determine, based on the first scheduling parameter, at least one screened pre-populated device from among the plurality of candidate pre-populated devices that satisfies a first screening condition; and a fourth determination subunit configured to determine, based on at least one first performance information, at least one target pre-populated device from among the at least one screened pre-populated device. The first performance information is configured to indicate performance of the screened pre-populated device.

[0121] In some embodiments, the fourth determination submodule includes: a second fusion unit for fusing multiple second scheduling weights for multiple candidate decoding devices to obtain a second fusion weight, where the second scheduling weight is used to indicate the storage capacity of the candidate decoding device. A fifth determination unit for determining a second scheduling parameter that is less than or equal to the second fusion weight. A sixth determination unit for determining, based on the second scheduling parameter, at least one target decoding device that meets a second screening condition among multiple candidate pre-filled devices. The second screening condition is: a second difference value for the candidate decoding device is less than or equal to a second preset difference threshold, where the second difference value is determined by the second scheduling parameter and at least one second scheduling weight for at least one candidate decoding device.

[0122] In some embodiments, the sixth determining unit includes: a fifth determining subunit configured to determine, based on the second scheduling parameter, at least one filtered decoding device from among the plurality of candidate decoding devices that satisfies a second filtering condition; and a sixth determining subunit configured to determine, based on at least one second performance information, at least one target decoding device from among the at least one filtered decoding device. The second performance information is configured to indicate performance of the filtered decoding device.

[0123] In some embodiments, 500 further includes: a processing module configured to process the pending request using the target pre-population device to obtain cached data to be processed; a providing module configured to provide the cached data to be processed to the target decoding device for the pending request; and a decoding module configured to decode the cached data to be processed using the target decoding device to obtain a processing result for the pending request.

[0124] In the technical solutions disclosed herein, the collection, storage, use, processing, transmission, provision and disclosure of user personal information involved comply with the provisions of relevant laws and regulations and do not violate public order and good morals.

[0125] According to an embodiment of the present disclosure, the present disclosure also provides an electronic device, a readable storage medium, and a computer program product.

[0126] Figure 6 A schematic block diagram of an example electronic device 600 that can be used to implement embodiments of the present disclosure is shown. The electronic device is intended to represent various forms of digital computers, such as laptop computers, desktop computers, workstations, personal digital assistants, servers, blade servers, mainframe computers, and other suitable computers. The electronic device can also represent various forms of mobile devices, such as personal digital assistants, cellular phones, smartphones, wearable devices, and other similar computing devices. The components shown herein, their connections and relationships, and their functions are provided as examples only and are not intended to limit the implementation of the present disclosure described and / or claimed herein.

[0127] like Figure 6 As shown, device 600 includes a computing unit 601, which can perform various appropriate actions and processes based on a computer program stored in a read-only memory (ROM) 602 or a computer program loaded from a storage unit 608 into a random access memory (RAM) 603. RAM 603 may also store various programs and data required for the operation of device 600. Computing unit 601, ROM 602, and RAM 603 are interconnected via a bus 604. An input / output (I / O) interface 605 is also connected to bus 604.

[0128] Various components in device 600 are connected to I / O interface 605, including an input unit 606, such as a keyboard, mouse, etc.; an output unit 607, such as various types of displays, speakers, etc.; a storage unit 608, such as a magnetic disk, optical disk, etc.; and a communication unit 609, such as a network card, modem, wireless communication transceiver, etc. The communication unit 609 allows device 600 to exchange information / data with other devices via a computer network such as the Internet and / or various telecommunication networks.

[0129] The computing unit 601 can be any general-purpose and / or specialized processing component with processing and computing capabilities. Some examples of the computing unit 601 include, but are not limited to, a central processing unit (CPU), a graphics processing unit (GPU), various specialized artificial intelligence (AI) computing chips, various computing units that run machine learning model algorithms, a digital signal processor (DSP), and any suitable processor, controller, microcontroller, etc. The computing unit 601 performs the various methods and processes described above, such as the request processing method. For example, in some embodiments, the request processing method can be implemented as a computer software program tangibly embodied in a machine-readable medium, such as the storage unit 608. In some embodiments, part or all of the computer program can be loaded and / or installed onto the device 600 via the ROM 602 and / or the communication unit 609. When the computer program is loaded into the RAM 603 and executed by the computing unit 601, one or more steps of the request processing method described above can be performed. Alternatively, in other embodiments, the computing unit 601 may be configured to execute the request processing method in any other appropriate manner (eg, by means of firmware).

[0130] Various embodiments of the systems and techniques described herein can be implemented in digital electronic circuit systems, integrated circuit systems, field programmable gate arrays (FPGAs), application specific integrated circuits (ASICs), application specific standard parts (ASSPs), system on chips (SOCs), complex programmable logic devices (CPLDs), computer hardware, firmware, software, and / or combinations thereof. These various embodiments can include being implemented in one or more computer programs that are executable and / or interpreted on a programmable system that includes at least one programmable processor, which can be a special purpose or general purpose programmable processor that can receive data and instructions from a storage system, at least one input device, and at least one output device, and transmit data and instructions to the storage system, the at least one input device, and the at least one output device.

[0131] The program code for implementing the method of the present disclosure can be written in any combination of one or more programming languages. These program codes can be provided to a processor or controller of a general-purpose computer, a special-purpose computer, or other programmable data processing device so that when the program code is executed by the processor or controller, the functions / operations specified in the flow chart and / or block diagram are implemented. The program code can be executed entirely on the machine, partially on the machine, as a stand-alone software package, partially on the machine and partially on a remote machine, or entirely on a remote machine or server.

[0132] In the context of the present disclosure, a machine-readable medium may be a tangible medium that may contain or store a program for use by or in conjunction with an instruction execution system, device, or apparatus. A machine-readable medium may be a machine-readable signal medium or a machine-readable storage medium. A machine-readable medium may include, but is not limited to, an electronic, magnetic, optical, electromagnetic, infrared, or semiconductor system, device, or apparatus, or any suitable combination of the foregoing. More specific examples of machine-readable storage media may include an electrical connection based on one or more wires, a portable computer disk, a hard disk, a random access memory, a read-only memory, an erasable programmable read-only memory (EPROM) or flash memory, an optical fiber, a portable compact disc read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the foregoing.

[0133] To provide interaction with a user, the systems and techniques described herein can be implemented on a computer having: a display device (e.g., a cathode ray tube (CRT) display or a liquid crystal display (LCD)) for displaying information to the user; and a keyboard and pointing device (e.g., a mouse or trackball) through which the user can provide input to the computer. Other types of devices can also be used to provide interaction with the user; for example, the feedback provided to the user can be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and input from the user can be received in any form (including acoustic input, voice input, or tactile input).

[0134] The systems and techniques described herein can be implemented in a computing system that includes back-end components (e.g., as a data server), or a computing system that includes middleware components (e.g., an application server), or a computing system that includes front-end components (e.g., a user computer with a graphical user interface or web browser through which a user can interact with implementations of the systems and techniques described herein), or a computing system that includes any combination of such back-end components, middleware components, or front-end components. The components of the system can be interconnected by any form or medium of digital data communication (e.g., a communication network). Examples of communication networks include a local area network (LAN), a wide area network (WAN), and the Internet.

[0135] Computer systems may include clients and servers. A client and server are generally remote from each other and typically interact through a communication network. The client and server relationship arises through computer programs running on the respective computers and having a client-server relationship to each other.

[0136] It should be understood that the various forms of the processes shown above can be used to reorder, add, or delete steps. For example, the steps described in this disclosure can be performed in parallel, sequentially, or in a different order, as long as the desired results of the technical solutions disclosed in this disclosure can be achieved. This is not limited herein.

[0137] The above specific embodiments do not constitute a limitation on the scope of protection of this disclosure. Those skilled in the art will appreciate that various modifications, combinations, sub-combinations, and substitutions may be made based on design requirements and other factors. Any modifications, equivalent substitutions, and improvements made within the spirit and principles of this disclosure shall be included within the scope of protection of this disclosure.

Claims

1. A request processing method, comprising: determining a target ratio based on target size information for at least one pending request, the target size information being determined based on historical size information of historical requests preceding the pending request; At least one target pre-filling device and at least one target decoding device for at least one of the pending requests are determined from a plurality of candidate devices, wherein a ratio between the number of the target pre-filling devices and the number of the target decoding devices is consistent with the target ratio.

2. The method according to claim 1, wherein Determining the target ratio based on target scale information for at least one pending request includes: Determine, from a plurality of candidate scale information, at least one screened scale information whose scale similarity with the target scale information satisfies a preset similarity condition, wherein the plurality of candidate scale information correspond one-to-one to a plurality of preset ratios; The target ratio is determined from at least one preset ratio corresponding to at least one of the filtered scale information.

3. The method according to claim 2, wherein: Determining at least one filtered scale information whose scale similarity with the target scale information satisfies a preset condition among the plurality of candidate scale information includes: determining a plurality of scale similarities between a target scale vector and a plurality of candidate scale vectors, the target scale vector being obtained based on the target scale information and the candidate scale vectors being obtained based on the candidate scale information; According to the plurality of scale similarities, at least one of the filtered scale information that meets the preset similarity condition is determined from the plurality of candidate scale information; The preset similarity condition includes: the scale similarity is greater than or equal to a preset scale similarity threshold.

4. The method according to claim 3, wherein: Determining the plurality of scale similarities between the target scale vector and the plurality of candidate scale vectors comprises: determining a cosine similarity between the target scale vector and the candidate scale vector; Dividing a first scale vector by a second scale vector to obtain a length similarity, where the first scale vector is a vector with a smaller modulus between the target scale vector and the candidate scale vectors, and the second scale vector is a vector with a larger modulus between the target scale vector and the candidate scale vectors; The cosine similarity and the length similarity are multiplied to obtain the scale similarity.

5. The method according to claim 1, wherein The target scale information is determined by performing the following operations based on historical scale information of historical requests before the pending request: performing weighted fusion on the plurality of historical scale information according to a plurality of autoregressive coefficients for the plurality of historical scale information to obtain fused scale information; performing weighted fusion on the multiple initial noise information according to multiple moving average coefficients for the multiple initial noise information to obtain fused noise information; The target scale information is determined according to preset parameters, the fusion scale information and the fusion noise information.

6. The method according to claim 1, wherein The plurality of candidate devices include a plurality of candidate pre-filled devices and a plurality of candidate decoding devices, The determining, from among a plurality of candidate devices, at least one target pre-population device and at least one target decoding device for at least one of the pending requests comprises: determining at least one target pre-filled device for at least one of the pending requests among a plurality of the candidate pre-filled devices; Among the plurality of candidate decoding devices, one target decoding device for at least one of the pending requests is determined.

7. The method according to claim 6, wherein: Determining at least one target pre-filled device for the pending request from among the plurality of candidate pre-filled devices comprises: fusing a plurality of first scheduling weights for a plurality of the candidate pre-filled devices to obtain a first fused weight, wherein the first scheduling weight is used to indicate a computing capability of the candidate pre-filled device; determining a first scheduling parameter that is less than or equal to the first fusion weight; determining, based on the first scheduling parameter, at least one target pre-filled device that meets a first screening condition among a plurality of candidate pre-filled devices; The first screening condition is that a first difference value for the candidate pre-filled device is less than or equal to a first preset difference threshold, and the first difference value is determined by the first scheduling parameter and at least one first scheduling weight for at least one of the candidate pre-filled devices.

8. The method according to claim 7, wherein: The determining, based on the first scheduling parameter, at least one target pre-filled device that satisfies a first screening condition from among the plurality of candidate pre-filled devices comprises: determining, based on the first scheduling parameter, at least one screened pre-filled device among the plurality of candidate pre-filled devices that satisfies the first screening condition; At least one target pre-filled device is determined from at least one of the screened pre-filled devices according to at least one first performance information, where the first performance information is used to indicate the performance of the screened pre-filled device.

9. The method according to claim 6, wherein: Determining at least one target decoding device for the pending request from among the plurality of candidate decoding devices comprises: fusing a plurality of second scheduling weights for a plurality of the candidate decoding devices to obtain a second fused weight, where the second scheduling weight is used to indicate a storage capacity of the candidate decoding device; determining a second scheduling parameter that is less than or equal to the second fusion weight; determining, based on the second scheduling parameter, at least one target decoding device that meets a second screening condition among a plurality of candidate pre-filled devices; The second screening condition is that the second difference value for the candidate decoding device is less than or equal to a second preset difference threshold, and the second difference value is determined by the second scheduling parameter and at least one second scheduling weight for at least one of the candidate decoding devices.

10. The method according to claim 9, wherein: The step of determining, based on the second scheduling parameter, at least one target decoding device that satisfies a second screening condition from among the plurality of candidate pre-populated devices comprises: determining, based on the second scheduling parameter, at least one screened decoding device that satisfies the second screening condition among the plurality of candidate decoding devices; At least one target decoding device is determined in at least one of the filtered decoding devices according to at least one second performance information, where the second performance information is used to indicate the performance of the filtered decoding device.

11. The method according to claim 1 , further comprising: Processing the pending request using the target pre-filling device to obtain pending cache data; providing the cached data to be processed to the target decoding device for the request to be processed; The target decoding device is used to perform decoding according to the cached data to be processed, to obtain a processing result of the request to be processed.

12. A request processing device, comprising: a first determining module, configured to determine a target ratio based on target scale information for at least one pending request, wherein the target scale information is determined based on historical scale information of historical requests prior to the pending request; The second determining module is configured to determine at least one target pre-filling device and at least one target decoding device for at least one of the pending requests from a plurality of candidate devices, wherein a quantity ratio between the quantity of the target pre-filling devices and the quantity of the target decoding devices is consistent with the target ratio.

13. An electronic device comprising: at least one processor; as well as a memory communicatively connected to the at least one processor; wherein, The memory stores instructions that can be executed by the at least one processor, and the instructions are executed by the at least one processor to enable the at least one processor to perform the method according to any one of claims 1 to 11.

14. A non-transitory computer-readable storage medium storing computer instructions, wherein: The computer instructions are used to cause the computer to execute the method according to any one of claims 1 to 11.

15. A computer program product comprising a computer program, which, when executed by a processor, implements the method according to any one of claims 1 to 11.