Request processing method and apparatus, computer device, and readable storage medium

By dynamically adjusting the bundle width of the sequence generation model and merging input operations according to the number of requests in the request queue, the trade-off between server throughput and generated sequence quality is resolved, thereby improving server throughput and computing power allocation efficiency.

CN115237618BActive Publication Date: 2026-09-11UBTECH ROBOTICS CORP LTD
1 Cites 0 Cited by

Patent Information

Application Number
CN202210805040.8
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-07-08
Publication Date
2026-09-11
Estimated Expiration
2042-07-08

AI Technical Summary

Technical Problem

When deploying a sequence generation model on a server, how can we improve server throughput while ensuring the quality of the generated sequences, especially the trade-off between bundle width and sequence length in terms of inference speed and computing power requirements?

Method used

By obtaining the real-time number of requests to be processed in the request queue, the bundle width of the sequence generation model is dynamically adjusted. When the number of requests is greater than or equal to a preset number, the preset bundle width is used. When the number is less than the preset number, the temporary bundle width is determined based on the real-time number of requests. The input sequence generation model is then merged for operation.

Benefits of technology

It achieves improved server throughput and optimized server computing power allocation to adapt to different request volumes while ensuring output quality and inference speed.

✦ Generated by Eureka AI based on patent content.
Patent Text Reader

Abstract

The request processing method and device, computer device and readable storage medium provided by the embodiments of the present application consider that the bundling width of the sequence generation model deployed on the server has a great influence on the time consumption and quality of the sequence generation operation, and a bundling width adaptive to the number of requests is selected to achieve a relatively optimized processing effect. For a case where the number of real-time requests is greater than or equal to a preset number of requests, the preset number of requests in the request queue are combined and input into the sequence generation model for sequence generation operation, while for a case where the number of real-time requests is less than the preset number of requests, a temporary bundling width is determined according to the number of real-time requests, and all the requests to be processed are combined and input into the sequence generation model for sequence generation operation. The temporary bundling width when the number of real-time requests to be processed is small is greater than the preset bundling width, which guarantees the output quality and inference speed, realizes the elastic allocation of server computing power, and improves the server throughput.
Need to check novelty before this filing date? Find Prior Art