A method, device and storage medium for reducing response delay of deep learning model

Through the load prediction method combined with sliding window sampling, polynomial regression model and LSTM model, the problem of poor resource allocation caused by load fluctuations in large-scale deep learning models is solved, and the accuracy and cost-effectiveness of load prediction are achieved, reducing response delay.

CN120123093BActive Publication Date: 2025-08-26HAITANGYUAN JING (TIANJIN) TECHNOLOGY CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510249305.4
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-03-04
Publication Date
2025-08-26
Estimated Expiration
2045-03-04

AI Technical Summary

Technical Problem

The existing technology cannot accurately predict workload burst fluctuations in large-scale deep learning models, resulting in insufficient resource allocation strategies and the inability to reduce the response delay of real-time inference services while ensuring cost-effectiveness.

Method used

Historical load data is obtained through sliding window sampling, combined with polynomial regression model and long and short-term memory network LSTM model for load prediction, obtain error compensation values, adjust the number of instances to match future needs, and optimize resource allocation.

Benefits of technology

Accurate load prediction of large-scale deep learning models is achieved, reducing response delays, improving the achievement rate of service level goals, and reducing service costs.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120123093B_ABST
    Figure CN120123093B_ABST
Patent Text Reader

Abstract

The present invention relates to a method, device and storage medium for reducing the response delay of a deep learning model, which is applied to the field of artificial intelligence technology, including: obtaining historical load through sliding window sampling, obtaining initial load prediction data using a dynamic joint prediction mechanism based on the historical load, obtaining an error compensation value through a corresponding load actual data sequence, and obtaining final load prediction data by performing error compensation on the initial load prediction data; determining the total number of instances required for a future period based on the final load prediction data sequence; and reducing the response delay of large-scale deep learning model inference work by adjusting the number of currently running instances to match the total number of instances required for a future period. The present application can effectively reduce the response delay of model inference through accurate workload prediction and resource scheduling, improve the achievement rate of service level goals, and reduce service costs.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of artificial intelligence technology, and specifically to a method, device, and storage medium for reducing the response delay of a deep learning model. Background Art

[0002] In cloud computing environments, large-scale deep learning models play a key role in providing advanced intelligent services. Due to their large number of parameters and complex structure, these models have extremely high demands on computing resources. In addition, since the inference service workload changes dynamically with user access volume and is difficult to predict, how to reduce the inference service latency of large-scale deep learning models in the real-time inference environment of cloud computing to meet service level objectives (SLO) is a complex and challenging problem.

[0003] Existing inference systems such as Cocktail and MArk have explored multi-dimensional optimization (such as using preemptible instances and serverless technology) to improve inference efficiency. However, these methods cannot accurately predict workloads, especially cannot handle sudden fluctuations in workloads. This leads to insufficiently refined resource allocation strategies and the inability to reduce the response latency of real-time inference services for large-scale deep learning models while ensuring cost-effectiveness. Summary of the Invention

[0004] In view of this, the purpose of the present invention is to provide a method, device and storage medium for reducing the response delay of deep learning models, so as to solve the problem in the prior art that the workload cannot be accurately predicted, especially the sudden fluctuations in the workload cannot be handled, resulting in insufficient resource allocation strategy and the inability to reduce the response delay of large-scale deep learning model real-time inference services while ensuring cost-effectiveness.

[0005] According to a first aspect of an embodiment of the present invention, a method for reducing response latency of a deep learning model is provided, the method comprising:

[0006] Acquire historical load data sequences for large-scale deep learning model inference tasks using a sliding window sampling method;

[0007] Based on the historical load data sequence, an initial load forecast data sequence for large-scale deep learning model inference work is obtained through a pre-built polynomial regression model and a long short-term memory network LSTM model;

[0008] Obtaining a load actual data sequence corresponding to an initial load prediction data sequence, and obtaining a workload fluctuation data sequence based on the initial load prediction data sequence and the corresponding load actual data sequence;

[0009] obtaining an error compensation value based on the workload fluctuation data sequence;

[0010] Performing error compensation on the initial load prediction data sequence using the error compensation value to obtain a final load prediction data sequence;

[0011] Determine the required computing throughput for a future period based on the final load prediction data sequence, and obtain the total number of instances required for the future period according to the throughput of the unit instance and the required computing throughput for the future period;

[0012] By adjusting the number of currently running instances to match the total number of instances required in the future, the response latency of large-scale deep learning model inference workloads can be reduced.

[0013] Preferably,

[0014] The initial load forecast data sequence for large-scale deep learning model inference work obtained by using the pre-built polynomial regression model and long short-term memory network LSTM model includes:

[0015] Establish a polynomial regression model and initialize the polynomial function; build a long short-term memory network LSTM model;

[0016] Fitting the historical load data sequence into a polynomial function using the least squares method, and updating the polynomial regression model using the fitted polynomial function;

[0017] Training and updating the long short-term memory network (LSTM) model using the historical load data sequence;

[0018] The workload of the large-scale deep learning model inference work is predicted by using the updated polynomial regression model and the long short-term memory network LSTM model respectively;

[0019] The prediction results of the updated polynomial regression model and the prediction results of the long short-term memory network LSTM model are fused to obtain the initial load prediction data sequence.

[0020] Preferably,

[0021] The prediction results of the updated polynomial regression model and the prediction results of the long short-term memory network LSTM model are integrated to obtain the initial load prediction data sequence, which includes:

[0022] Get the first mean square error of the most recent historical forecast results of the updated polynomial regression model, and get the second mean square error of the forecast results of the long short-term memory network (LSTM) model for the same time period;

[0023] Obtaining the output weight of the polynomial regression model and the output weight of the long short-term memory network LSTM model respectively according to the first mean square error and the second mean square error;

[0024] The prediction results of the updated polynomial regression model and the long short-term memory network LSTM model are weightedly fused according to their respective weights to obtain the initial load prediction data sequence.

[0025] Preferably,

[0026] The acquiring of the workload fluctuation data sequence according to the initial load prediction data sequence and the corresponding actual load data sequence comprises:

[0027] The workload fluctuation data sequence is obtained by subtracting the initial load prediction data sequence from the corresponding load actual data sequence at the same time point.

[0028] Preferably,

[0029] The acquiring of the error compensation value based on the workload fluctuation data sequence comprises:

[0030] The 95th percentile value of the workload fluctuation data sequence is obtained, and the 95th percentile value of the workload fluctuation data sequence is used as the error compensation value.

[0031] Preferably,

[0032] Adjusting the number of currently running instances to match the total number of instances required in the future includes:

[0033] When it is predicted that the final load forecast data is on an upward trend and the difference from the preset load peak is less than a preset threshold, increase the number of currently running instances or adjust the batch size;

[0034] When it is predicted that the final load forecast data is on a downward trend and is always smaller than the preset load peak, the number of currently running instances is reduced.

[0035] According to a second aspect of an embodiment of the present invention, there is provided an apparatus for reducing response latency of a deep learning model, the apparatus comprising:

[0036] Historical load collection module: used to obtain historical load data sequences for large-scale deep learning model inference work through a sliding window sampling method;

[0037] Load forecasting module: used to obtain the initial load forecast data sequence for large-scale deep learning model inference work based on the historical load data sequence through a pre-built polynomial regression model and a long short-term memory network LSTM model;

[0038] Fluctuation sequence acquisition module: used to obtain the load actual data sequence corresponding to the initial load prediction data sequence, and obtain the workload fluctuation data sequence based on the initial load prediction data sequence and the corresponding load actual data sequence;

[0039] A compensation value acquisition module is configured to acquire an error compensation value based on the workload fluctuation data sequence;

[0040] A load prediction compensation module is configured to perform error compensation on the initial load prediction data sequence using the error compensation value to obtain a final load prediction data sequence;

[0041] A future instance acquisition module is configured to determine the computing throughput required in a future period based on the final load prediction data sequence, and obtain the total number of instances required in the future period according to the throughput of the unit instance and the computing throughput required in the future period;

[0042] Adjustment module: Used to reduce the response latency of large-scale deep learning model inference by adjusting the number of currently running instances to match the total number of instances required in the future.

[0043] According to a third aspect of an embodiment of the present invention, a storage medium is provided, wherein the storage medium stores a computer program, and when the computer program is executed by a host controller, each step in the above method is implemented.

[0044] The technical solutions provided by the embodiments of the present invention may have the following beneficial effects:

[0045] This application obtains historical load through sliding window sampling, and obtains the initial load prediction data sequence based on the historical load and the pre-built polynomial regression model and long short-term memory network LSTM model; obtains workload fluctuations through the corresponding load actual data sequence, obtains error compensation values ​​based on the workload fluctuations, and compensates the initial load prediction data through the error compensation values ​​to obtain final load prediction data; determines the computing throughput required in the future period based on the final load prediction data sequence, and obtains the total number of instances required in the future period according to the throughput of the unit instance and the computing throughput required in the future period; reduces the response delay of large-scale deep learning model inference work by adjusting the number of currently running instances to match the total number of instances required in the future period; this application can effectively reduce the response delay of model inference, improve the achievement rate of service level goals, and reduce service costs through accurate workload prediction and resource scheduling.

[0046] It is to be understood that the foregoing general description and the following detailed description are exemplary and explanatory only and are not restrictive of the invention. BRIEF DESCRIPTION OF THE DRAWINGS

[0047] The accompanying drawings, which are incorporated in and constitute a part of this specification, illustrate embodiments consistent with the invention and, together with the description, serve to explain the principles of the invention.

[0048] Figure 1 This is a flow chart illustrating a method for reducing the response delay of a deep learning model according to an exemplary embodiment;

[0049] Figure 2 is a system schematic diagram of an apparatus for reducing response latency of a deep learning model according to another exemplary embodiment;

[0050] In the attached figure: 1-historical load acquisition module, 2-load prediction module, 3-fluctuation sequence acquisition module, 4-compensation value acquisition module, 5-predicted load compensation module, 6-future instance acquisition module, 7-adjustment module. DETAILED DESCRIPTION

[0051] Exemplary embodiments will be described in detail herein, examples of which are illustrated in the accompanying drawings. In the following description, when referring to the drawings, like numbers in different figures represent like or similar elements unless otherwise indicated. The embodiments described in the following exemplary embodiments are not intended to represent all possible embodiments consistent with the present invention. Rather, they are merely examples of apparatus and methods consistent with certain aspects of the present invention, as detailed in the appended claims.

[0052] Example 1

[0053] Figure 1 FIG. 1 is a flow chart of a method for reducing the response delay of a deep learning model according to an exemplary embodiment. Figure 1 As shown, the method includes:

[0054] S1, obtains historical load data sequences for large-scale deep learning model inference work through a sliding window sampling method;

[0055] S2, based on the historical load data sequence, obtains an initial load forecast data sequence for large-scale deep learning model inference work through a pre-built polynomial regression model and a long short-term memory network (LSTM) model;

[0056] S3, obtaining a load actual data sequence corresponding to the initial load prediction data sequence, and obtaining a workload fluctuation data sequence based on the initial load prediction data sequence and the corresponding load actual data sequence;

[0057] S4, obtaining an error compensation value based on the workload fluctuation data sequence;

[0058] S5, performing error compensation on the initial load prediction data sequence using the error compensation value to obtain a final load prediction data sequence;

[0059] S6, determining the required computing throughput for a future period of time based on the final load prediction data sequence, and obtaining the total number of instances required for the future period of time according to the throughput of the unit instance and the required computing throughput for the future period of time;

[0060] S7 reduces the response latency of large-scale deep learning model inference by adjusting the number of currently running instances to match the total number of instances required in the future.

[0061] It is understandable that this embodiment uses a sliding window method to sample and obtain the current large-scale deep learning model inference workload historical data. Next, the dynamic joint prediction (DJP) mechanism is used to perform long-term trend prediction on the workload data. The workload prediction interval is set, and the current workload is calculated and recorded using the sliding window method on a regular basis. Specifically, the following steps are included:

[0062] Establish a polynomial regression model and initialize the polynomial function; build a long short-term memory (LSTM) model and pre-train the model using the initial data; the polynomial regression model can be a polynomial of any degree, and the long short-term memory (LSTM) model can be an LSTM variant, including BiLSTM and Linear LSTM.

[0063] The historical workload is fitted into a polynomial function using the least squares method, and the polynomial regression model is updated. If the amount of historical workload data collected meets the requirements for one training of the LSTM model, the LSTM model is trained and updated using this data.

[0064] The workload is predicted using a polynomial regression model and an LSTM model. The sampling weights for the final workload prediction value are determined based on the recent errors of the two models. Specifically, the following are involved:

[0065] Define the workload prediction time length as λ, and define the workload prediction value of the polynomial regression model at time t+λ as The workload prediction value of the LSTM model at time t+λ is ;

[0066] Define the historical prediction mean square error MSE of the polynomial regression model and LSTM model as follows: and , calculate the weights for workload prediction of the polynomial regression model:

[0067]

[0068] In the initial stage of the model inference service, since there is no data for LSTM model training, the workload prediction weight is set to 1.0, that is, the prediction results of the polynomial regression model are fully adopted;

[0069] The initial workload prediction value at time t+λ is obtained based on the sampling weights of the two models:

[0070]

[0071] As the inference service progresses, the mean squared error (MSE) of the workload prediction values ​​of the two models is recalculated periodically;

[0072] The above steps can overcome the large local errors and slow startup issues of existing systems using a single polynomial regression model or LSTM model by leveraging the historical prediction errors of the two large-scale model inference service workload prediction models. This also achieves a dynamic weighted combination of deep learning prediction methods and polynomial fitting prediction methods, thereby ensuring basically accurate workload prediction values ​​at any stage of the model inference service.

[0073] A historical fluctuation compensation method is used to analyze the fluctuation intensity of historical workload data and provide compensation factors to optimize resource scheduling decisions, including:

[0074] Get the actual sequence corresponding to the above predicted sequence ; and compare it with the predicted initial workload forecast sequence Subtract to get the workload fluctuation series , the formula is:

[0075]

[0076] Calculate the 95th percentile of the fluctuation series as the error compensation for the workload forecast at time t+λ , thereby avoiding insufficient resource allocation due to small fluctuations in workload, which leads to increased model inference latency;

[0077] According to the above steps, the accurate prediction value of the workload at time t+λ is obtained :

[0078]

[0079] The workload error compensation quantile can be adjusted in the range of 50 to 100 according to the service level objective (SLO) requirements;

[0080] The above steps can overcome the problem that existing workload prediction cannot cope with instantaneous workload fluctuations, thereby achieving accurate and precise workload prediction values. This provides a resource scheduling foundation for large-scale deep learning model inference, ensuring that the inference request response latency meets the service level objective (SLO).

[0081] To facilitate understanding of the above scheme, this embodiment briefly describes the above scheme, taking the initial current moment as "4", the historical load sequences as "1", "2", and "3" respectively, and the future load prediction sequences "5", "6", and "7" are obtained based on the dynamic joint prediction mechanism DJP through the historical load sequences as "1", "2", and "3". It should be noted that this process is carried out dynamically over time, that is, as time goes by, the actual loads "{5}", "{6}", and "{7}" corresponding to the initial load prediction sequences "5", "6", and "7" can be obtained. At this time, the dynamic joint prediction mechanism DJP has already predicted the initial load prediction sequences "8", "9", and "10". The initial load prediction sequences "5", "6", and "7" are subtracted from the corresponding actual loads "{5}", "{6}", and "{7}" to obtain a fluctuation sequence. An error compensation value is obtained based on the fluctuation sequence. The initial load prediction sequences "8", "9", and "10" are compensated for by the compensation value to obtain accurate workload prediction values ​​at moments "8", "9", and "10". ;

[0082] Dynamically adjust resource allocation strategies based on accurate long-term trend forecasts, including:

[0083] Determine the required computing throughput for the future based on accurate workload forecasts, and derive the total number of instances required based on the throughput per instance (an instance can be understood as a server; the more instances, the more load it can handle). Adjust the number of running service instances based on the current system's resource usage and total resource forecasts to match the predicted workload requirements. When workload peaks are predicted, increase system processing capacity by adding service instances or adjusting batch sizes to avoid excessive latency. When workload decreases are predicted, proactively reduce the number of service instances to conserve resources and reduce costs.

[0084] This embodiment will also adjust the model configuration parameters (including batch size and number of parallel model replicas) based on the performance indicators of the service instance (including CPU utilization, GPU utilization, and memory usage) and the performance requirements of model inference to optimize inference efficiency and response latency.

[0085] It is worth emphasizing that in order to solve the model cold start problem and request response timeout problem in large-scale deep learning model inference, thereby reducing the response latency of the model real-time inference service, this embodiment also provides a method for dynamic resource and request management, specifically including:

[0086] (1) Model Cold Start Optimizer AoT Tuner module, which pre-analyzes the optimal CUDA operators corresponding to each layer of the deep learning model to achieve fast model loading and initialization at runtime and minimize cold start time. Specifically, it includes:

[0087] Offline analysis of each deep learning model is performed to determine the optimal CUDA operator for each layer of the model. This is done using NVIDIA's TensorRT framework, an optimizer for high-performance deep learning inference that performs precision calibration, layer fusion, and operator optimization. Each layer of a large-scale deep learning model is traversed, trying different combinations of CUDA operators and recording the execution time of each combination. The inference execution time for each operator combination is compared, and the CUDA operator with the shortest execution time is selected as the optimal operator, with its configuration parameters recorded. The optimal CUDA operator and its configuration parameters for each layer are saved in the model's optimization configuration file, which is stored in random access memory (RAM) for fast access.

[0088] Use the TensorFlow framework to cache the optimized model for fast loading during inference: Convert the TensorRT-optimized model to the SavedModel format supported by TensorFlow; Save the optimized model and the configuration file of the optimal CUDA operator in the TensorFlow model repository, which can be deployed on a solid-state drive with high bandwidth and low latency; Optionally, maintain a version control mechanism for each model in the model repository to ensure that the latest or specified version of the model and configuration file can be quickly located when an inference request arrives; To improve caching efficiency, use model compression techniques such as weight quantization and pruning to reduce the model size while maintaining inference performance;

[0089] When an inference request arrives, the loaded model and configuration file are directly passed to the GPU, which uses NVIDIA's CUDA technology for fast initialization, including memory allocation and operator binding, to reduce model loading and initialization time;

[0090] Through the above step (1), the AoT Tuner mechanism can ensure that the model is quickly loaded and initialized when an inference request arrives, reducing the delay caused by model loading and initialization, and improving the response speed and overall performance of large-scale deep learning model inference services.

[0091] (2) Dynamic Request Management Adaptive Batch Drop method monitors the system resource status and request latency in real time during the inference process. Based on the current resource status and request latency requirements, it dynamically decides whether to execute the requested batch processing, including:

[0092] Define the system resource state R(t) as the total amount of computing resources available in the system at time t, including the number of GPU cores and memory capacity; define the request latency Di as the time required for the i-th request to be submitted and completed; continuously track and record the values ​​of R(t) and Di to make dynamic scheduling decisions.

[0093] Dynamically decide whether to execute the requested batch based on the current resource status and the request latency requirements: For each request batch to be executed, calculate its estimated completion time :

[0094]

[0095] in, n is the number of requests in the batch, W j For the j The weight of the request, T j For the j The estimated execution time of each request, R(t) The current system resource status.

[0096] When system resources are insufficient and cannot meet the delay requirements of all requests, some requests are selectively discarded according to the preset discarding rules: the maximum allowable delay Dmax is defined as the delay limit set by the system according to the service level objective SLO requirements; for each request i ,if D i > D max , then the request is considered to have a timeout risk; according to the priority of the request P i and delay sensitivity S i Calculate the drop probability for each request P drop ( i ):

[0097]

[0098] in, g ( P i )and h ( Si ) is a predefined function used to adjust the drop probability based on the priority and delay sensitivity of the request.

[0099] By adjusting the batch size and execution order, the overall request processing process is optimized: the batch size is dynamically adjusted according to the current resource status R(t) and the priority information in the request queue to maximize resource utilization and meet the delay requirements; a priority queue is defined in which each request is sorted according to its priority and delay requirements; when the system resources are insufficient to process all requests, the request is discarded according to the probability of the request being discarded. P drop ( i ) and priority queues to selectively discard some requests while ensuring that key requests are given priority; for discarded requests, you can choose to re-queue or directly return an error response. The specific strategy depends on the business needs and SLO requirements of the system.

[0100] Through the above-mentioned Adaptive Batch Drop method, this embodiment can effectively manage resources and reduce timeout requests caused by queuing while ensuring the service quality of key requests, thereby improving the overall performance of large-scale deep learning model inference services and the response latency of inference services.

[0101] Example 2:

[0102] Figure 2 2 is a system diagram illustrating an apparatus for reducing response latency of a deep learning model according to another exemplary embodiment, the apparatus comprising:

[0103] Historical load collection module 1: used to obtain historical load data sequences for large-scale deep learning model inference work through a sliding window sampling method;

[0104] Load forecasting module 2: used to obtain the initial load forecast data sequence for large-scale deep learning model inference work based on the historical load data sequence through a pre-built polynomial regression model and a long short-term memory network LSTM model;

[0105] Fluctuation sequence acquisition module 3: used to obtain the load actual data sequence corresponding to the initial load prediction data sequence, and obtain the workload fluctuation data sequence based on the initial load prediction data sequence and the corresponding load actual data sequence;

[0106] Compensation value acquisition module 4: used to acquire an error compensation value based on the workload fluctuation data sequence;

[0107] Predicted load compensation module 5: used to perform error compensation on the initial load prediction data sequence using the error compensation value to obtain a final load prediction data sequence;

[0108] Future instance acquisition module 6: used to determine the computing throughput required in the future period based on the final load prediction data sequence, and obtain the total number of instances required in the future period according to the throughput of the unit instance and the computing throughput required in the future period;

[0109] Adjustment Module 7: Used to reduce the response latency of large-scale deep learning model inference work by adjusting the number of currently running instances to match the total number of instances required in the future.

[0110] Example 3:

[0111] This embodiment provides a storage medium, wherein the storage medium stores a computer program, and when the computer program is executed by a host controller, each step in the above method is implemented;

[0112] It is understandable that the storage medium mentioned above can be a read-only memory, a magnetic disk or an optical disk, etc.

[0113] It can be understood that the same or similar parts of the above embodiments can be referenced to each other, and the contents not described in detail in some embodiments can refer to the same or similar contents in other embodiments.

[0114] It should be noted that, in the description of the present invention, the terms "first," "second," etc. are used for descriptive purposes only and should not be understood as indicating or implying relative importance. In addition, in the description of the present invention, unless otherwise specified, "a small number of sparsely distributed" means at least two.

[0115] Any process or method description in a flowchart or otherwise described herein may be understood to represent a module, segment or portion of code comprising one or less sparsely distributed executable instructions for implementing the steps of a specific logical function or process, and the scope of the preferred embodiments of the present invention includes alternative implementations in which functions may be performed out of the order shown or discussed, including performing functions in a substantially simultaneous manner or in the reverse order depending on the functions involved, which should be understood by those skilled in the art to which the embodiments of the present invention pertain.

[0116] It should be understood that various components of the present invention can be implemented using hardware, software, firmware, or a combination thereof. In the above-described embodiment, a small number of sparsely distributed steps or methods can be implemented using software or firmware stored in a memory and executed by a suitable instruction execution system. For example, if implemented using hardware, as in another embodiment, any one of the following technologies known in the art or a combination thereof can be used: a discrete logic circuit having logic gate circuits for implementing logical functions on data signals, an application-specific integrated circuit having suitable combinational logic gate circuits, a programmable gate array (PGA), a field-programmable gate array (FPGA), etc.

[0117] Those skilled in the art will understand that all or part of the steps in the method of the above embodiment can be completed by instructing related hardware through a program, and the program can be stored in a computer-readable storage medium. When the program is executed, it includes one or a combination of the steps of the method embodiment.

[0118] In addition, the functional units in the various embodiments of the present invention may be integrated into a single processing module, or each unit may exist physically separately, or two or more units may be integrated into a single module. The aforementioned integrated modules may be implemented in the form of hardware or in the form of software functional modules. If the integrated modules are implemented in the form of software functional modules and sold or used as independent products, they may also be stored in a computer-readable storage medium.

[0119] The storage medium mentioned above can be a read-only memory, a magnetic disk or an optical disk, etc.

[0120] Throughout this specification, references to terms such as "one embodiment," "some embodiments," "examples," "specific examples," or "some examples" indicate that the specific features, structures, materials, or characteristics described in conjunction with that embodiment or example are included in at least one embodiment or example of the present invention. In this specification, schematic representations of the above terms do not necessarily refer to the same embodiment or example. Furthermore, the specific features, structures, materials, or characteristics described may be combined in any suitable manner in any one or a small number of sparsely distributed embodiments or examples.

[0121] Although the embodiments of the present invention have been shown and described above, it will be understood that the above embodiments are illustrative and are not to be construed as limitations on the present invention. A person skilled in the art may change, modify, replace and modify the above embodiments within the scope of the present invention.

Claims

1. A method for reducing the response delay of a deep learning model, characterized in that: The method comprises: Acquire historical load data sequences for large-scale deep learning model inference tasks using a sliding window sampling method; Based on the historical load data sequence, an initial load forecast data sequence for large-scale deep learning model inference work is obtained through a pre-built polynomial regression model and a long short-term memory network LSTM model; Obtaining a load actual data sequence corresponding to an initial load prediction data sequence, and obtaining a workload fluctuation data sequence based on the initial load prediction data sequence and the corresponding load actual data sequence; obtaining an error compensation value based on the workload fluctuation data sequence; Performing error compensation on the initial load prediction data sequence using the error compensation value to obtain a final load prediction data sequence; Determine the required computing throughput for a future period of time based on the final load prediction data sequence, and obtain the total number of instances required for the future period of time based on the throughput of the unit instance and the required computing throughput for the future period of time; By adjusting the number of currently running instances to match the total number of instances required in the future, the response latency of large-scale deep learning model inference workloads can be reduced.

2. The method according to claim 1, characterized in that The initial load forecast data sequence for large-scale deep learning model inference work obtained by using the pre-built polynomial regression model and long short-term memory network LSTM model includes: Establish a polynomial regression model and initialize the polynomial function; build a long short-term memory network LSTM model; Fitting the historical load data sequence into a polynomial function using the least squares method, and updating the polynomial regression model using the fitted polynomial function; Training and updating the long short-term memory network (LSTM) model using the historical load data sequence; The workload of the large-scale deep learning model inference work is predicted by using the updated polynomial regression model and the long short-term memory network LSTM model respectively; The prediction results of the updated polynomial regression model and the prediction results of the long short-term memory network LSTM model are fused to obtain the initial load prediction data sequence.

3. The method according to claim 2, characterized in that The prediction results of the updated polynomial regression model and the prediction results of the long short-term memory network LSTM model are integrated to obtain the initial load prediction data sequence, which includes: Get the first mean square error of the most recent historical forecast results of the updated polynomial regression model, and get the second mean square error of the forecast results of the long short-term memory network (LSTM) model for the same time period. Obtaining the output weight of the polynomial regression model and the output weight of the long short-term memory network LSTM model respectively according to the first mean square error and the second mean square error; The prediction results of the updated polynomial regression model and the long short-term memory network LSTM model are weightedly fused according to their respective weights to obtain the initial load prediction data sequence.

4. The method according to claim 3, characterized in that The acquiring of the workload fluctuation data sequence according to the initial load prediction data sequence and the corresponding actual load data sequence comprises: The workload fluctuation data sequence is obtained by subtracting the initial load prediction data sequence from the corresponding load actual data sequence at the same time point.

5. The method according to claim 4, characterized in that The acquiring of the error compensation value based on the workload fluctuation data sequence comprises: The 95th percentile value of the workload fluctuation data sequence is obtained, and the 95th percentile value of the workload fluctuation data sequence is used as the error compensation value.

6. The method according to claim 5, characterized in that Adjusting the number of currently running instances to match the total number of instances required in the future includes: When it is predicted that the final load forecast data is on an upward trend and the difference from the preset load peak is less than a preset threshold, increase the number of currently running instances or adjust the batch size; When it is predicted that the final load forecast data is on a downward trend and is always smaller than the preset load peak, the number of currently running instances is reduced.

7. A device for reducing the response delay of a deep learning model, characterized in that: The device comprises: Historical load collection module: used to obtain historical load data sequences for large-scale deep learning model inference work through a sliding window sampling method; Load forecasting module: used to obtain the initial load forecast data sequence for large-scale deep learning model inference work based on the historical load data sequence through a pre-built polynomial regression model and a long short-term memory network LSTM model; Fluctuation sequence acquisition module: used to obtain the load actual data sequence corresponding to the initial load prediction data sequence, and obtain the workload fluctuation data sequence based on the initial load prediction data sequence and the corresponding load actual data sequence; A compensation value acquisition module is configured to acquire an error compensation value based on the workload fluctuation data sequence; A load prediction compensation module is configured to perform error compensation on the initial load prediction data sequence using the error compensation value to obtain a final load prediction data sequence; A future instance acquisition module is configured to determine the computing throughput required in a future period based on the final load prediction data sequence, and obtain the total number of instances required in the future period according to the throughput of the unit instance and the computing throughput required in the future period; Adjustment module: Used to reduce the response latency of large-scale deep learning model inference by adjusting the number of currently running instances to match the total number of instances required in the future.

8. A storage medium, characterized in that: The storage medium stores a computer program, which, when executed by the main controller, implements the various steps in the method for reducing the response delay of a deep learning model as described in any one of claims 1 to 6.

Citation Information

Patent Citations

  • Heading machine performance prediction method and device based on deep learning

    CN118428409A

  • Dynamic resource scheduling method and device for multi-GPU cluster reasoning service

    CN119311396A