Large-model high-performance reasoning acceleration method and system based on Triton-Inference-Server
By adopting a large-scale high-performance inference acceleration method based on Triton-Inference-Server in a multi-model inference service system, the system's shortcomings in management complexity, performance optimization and robustness are solved, and high stability, low latency and high throughput inference services are achieved.
Patent Information
- Application Number
- CN202411904871.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2024-12-23
- Publication Date
- 2025-05-16
- Estimated Expiration
- Not applicable · inactive patent
AI Technical Summary
The existing multi-model inference service system has shortcomings in management complexity, performance optimization, robustness and high concurrent request processing, resulting in unstable model services, prolonged inference time and low throughput.
The large-scale high-performance inference acceleration method based on Triton-Inference-Server is adopted. By receiving user model configuration information, detecting and downloading models, supporting model upload and optimization, quantifying processing, building a containerized operating environment, using dynamic batch processing and resource management strategies, combining log analysis and performance monitoring tools.
It significantly reduces the complexity and error rate of manual operation, improves the stability of model services and real-time inference capabilities in high concurrency scenarios, reduces inference delay and resource occupation, and enhances the robustness and maintainability of the system.
Smart Images

Figure CN120012915A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of artificial intelligence and machine learning technology, and in particular to a large-model high-performance reasoning acceleration method and system based on Triton-Inference-Server. Background Art
[0002] In the current field of artificial intelligence and machine learning, multi-model reasoning service systems are widely used to meet reasoning needs in different scenarios. Existing multi-model reasoning service systems usually need to manage multiple base models at the same time. These models involve multiple frameworks and formats, such as TensorFlow, PyTorch, ONNX, etc. Although these systems can support basic reasoning functions, they still have significant deficiencies in management complexity, performance optimization, and robustness.
[0003] First, in terms of multi-model management, existing technologies usually require manual deployment, updating and maintenance of multiple base models, which is not only time-consuming and labor-intensive, but also prone to human errors, resulting in instability of model services. In addition, the coexistence of multiple frameworks and formats increases the difficulty of management and operational complexity. Secondly, the high latency and low throughput of large model reasoning are also the main bottlenecks of existing technologies. Since large models require a large amount of computing resources, such as video memory and processing power, during operation, it is difficult for existing reasoning service systems to effectively allocate and optimize resources, resulting in slow response speed and low throughput for reasoning requests, which cannot meet the needs of high concurrent requests in actual production scenarios. Finally, existing systems are also deficient in the robustness of model services and request backtracking capabilities. When anomalies or errors occur in the reasoning service, there is a lack of efficient error location and recovery mechanisms, which affects the availability of the service. At the same time, the system is weak in the ability to handle the history and backtracking of requests, making it difficult to meet the needs of comprehensive monitoring and auditing of the service process. Summary of the invention
[0004] The purpose of the present invention is to provide a large-model high-performance inference acceleration method and system based on Triton-Inference-Server to solve the problems of limitations of the prior art multi-model inference service system in terms of management complexity, high latency, low throughput and service robustness.
[0005] To achieve the above object, the present invention provides the following technical solution: a large model high-performance inference acceleration method based on Triton-Inference-Server, characterized in that the method comprises:
[0006] S1. Receive the model configuration information provided by the user, detect whether there is a local model copy according to the model configuration, and if not, download the model from the remote storage and generate a configuration file that meets the requirements of the inference server;
[0007] S2. Support users to upload customized models or fine-tune models, and load, optimize and fuse parameters of the models;
[0008] S3. Quantify the model according to user needs and use a variety of quantization methods to optimize reasoning performance;
[0009] S4. Build an operating environment based on the inference server, including deploying containerized inference service instances;
[0010] S5. Start the inference service in the specified hardware environment and use dynamic batch processing technology and resource management strategies to improve inference efficiency.
[0011] S6. Collect operation data through log analysis and performance monitoring tools to optimize and enhance the robustness of reasoning services.
[0012] Preferably, the model configuration information includes model name, model path, inference accuracy, parallel mode, batch size and maximum input sequence length.
[0013] Preferably, the step of downloading the model includes calling a remote storage interface, including a ModelScope API or a Hugging Face Transformers Hub API, to obtain model weights and configuration files.
[0014] Preferably, the optimization and parameter fusion steps optimize the model by using the TensorRT engine building tool and adjust the reasoning path for specific task types to improve efficiency.
[0015] Preferably, the quantization processing includes FP16 half-precision compression, INT8 quantization and INT4 quantization, and a calibration table is generated using a calibration data set to optimize quantization accuracy.
[0016] Preferably, the calibration data set is generated from historical data, specifically including sparse matrix or time series data.
[0017] Preferably, the operating environment includes deploying an inference service instance using a Docker image, the image supports multiple CUDA versions, and integrates a log monitoring plug-in and a load scheduling module.
[0018] Preferably, the dynamic batch processing technology includes merging requests by time window through the dynamic batch processing API of the inference server, and setting an upper limit on the batch size and a delay threshold.
[0019] Preferably, the performance monitoring tools include Prometheus and Grafana, which are used to monitor GPU utilization, inference latency and throughput in real time.
[0020] A large-model high-performance inference acceleration system based on Triton-Inference-Server is used to implement the steps of the large-model high-performance inference acceleration method based on Triton-Inference-Server, and the system includes:
[0021] The model management module is used to receive the model configuration information provided by the user, detect whether there is a local model copy according to the model configuration, and if not, download the model from the remote storage and generate a configuration file that meets the requirements of the inference server;
[0022] The model loading and optimization module is connected to the model management module to support users to upload customized models or fine-tune models, and load, optimize and fuse parameters of the models;
[0023] The model quantization module is connected to the model loading and optimization module, and is used to quantize the model according to user needs and use a variety of quantization methods to optimize the inference performance;
[0024] The operating environment construction module is connected to the model quantization module and is used to build an operating environment based on the inference server, including deploying containerized inference service instances;
[0025] The inference service management module is connected to the operating environment building module to start the inference service in the specified hardware environment and improve the inference efficiency through dynamic batch processing technology and resource management strategies;
[0026] The log analysis and performance monitoring module is connected to the inference service management module to collect the operation data of the inference service and realize the optimization and robustness enhancement of the service.
[0027] It can be seen from the above technical solution that the present invention has the following beneficial effects:
[0028] This large-model high-performance inference acceleration method based on Triton-Inference-Server receives model configuration information provided by the user, detects whether there is a local model copy based on the model configuration, and if not, downloads the model from remote storage and generates a configuration file that meets the requirements of the inference server. It supports users to upload customized models or fine-tune models, and loads, optimizes and fuses parameters of the models. It quantifies the models according to user needs, uses a variety of quantization methods to optimize inference performance, and builds an operating environment based on the inference server, including deploying containerized inference service instances, starting inference services in a specified hardware environment, using dynamic batch processing technology and resource management strategies to improve inference efficiency, and collecting operating data through log analysis and performance monitoring tools to optimize and enhance the robustness of inference services, greatly reduce the complexity and error rate of manual operations, improve the stability of model services, and meet real-time inference requirements in high-concurrency scenarios. For example, through model optimization, the end-to-end reasoning rate can be greatly improved, the operating cost of the reasoning service can be reduced, and the utilization efficiency of hardware resources can be improved. Service anomalies can be quickly located and repaired, the reliability and continuous availability of model services can be improved, and a comprehensive audit of the reasoning process can be achieved to provide a basis for service optimization and problem troubleshooting. It can flexibly adapt to different hardware environments and production requirements, and is suitable for a variety of practical application scenarios. It greatly reduces the user's technical threshold and enables developers to focus on model design and business logic, thereby improving development efficiency and solving the limitations of existing multi-model reasoning service systems in terms of management complexity, high latency, low throughput and service robustness. BRIEF DESCRIPTION OF THE DRAWINGS
[0029] Figure 1 is a flow chart of the method of the present invention;
[0030] Figure 2 This is a connection diagram of the system modules of the present invention. DETAILED DESCRIPTION
[0031] The following will be combined with the drawings in the embodiments of the present invention to clearly and completely describe the technical solutions in the embodiments of the present invention. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of the present invention.
[0032] like Figure 1 As shown, the present invention provides a technical solution: a large model high-performance reasoning acceleration method based on Triton-Inference-Server, characterized in that the method includes:
[0033] S1. Receive the model configuration information provided by the user, detect whether there is a local model copy according to the model configuration, and if not, download the model from the remote storage and generate a configuration file that meets the requirements of the inference server;
[0034] S2. Support users to upload customized models or fine-tune models, and load, optimize and fuse parameters of the models;
[0035] S3. Quantify the model according to user needs and use a variety of quantization methods to optimize reasoning performance;
[0036] S4. Build an operating environment based on the inference server, including deploying containerized inference service instances;
[0037] S5. Start the inference service in the specified hardware environment and use dynamic batch processing technology and resource management strategies to improve inference efficiency.
[0038] S6. Collect operation data through log analysis and performance monitoring tools to optimize and enhance the robustness of reasoning services.
[0039] The above method uses the powerful reasoning capabilities and flexible deployment features of Triton-Inference-Server, combined with the optimized configuration of hardware resources, to achieve high-performance large-model reasoning services. By detecting local model copies and automatically downloading and configuring from remote storage, the complexity of user operations is reduced; it supports uploading user-defined models to meet diverse reasoning needs, and combines with tools such as TensorRT for optimization to improve reasoning efficiency and model performance. The model is processed using quantization technology, such as FP16, INT8, and INT4 quantization, which significantly reduces the reasoning latency and resource usage by reducing the demand for model calculation accuracy. The containerized design of the operating environment not only improves the convenience of deployment, but also optimizes resource allocation between multiple reasoning requests through dynamic batching technology, significantly improving throughput and computing efficiency. Performance monitoring tools and log analysis tools provide detailed data support for system tuning by collecting indicators such as GPU utilization and reasoning latency, while enhancing the robustness and maintainability of the system. Through model optimization and quantization processing technology, the inference latency is significantly reduced while the throughput is improved; dynamic batch processing and resource management strategies further optimize performance, support users to upload custom models or fine-tune models to meet the needs of different application scenarios, and automated download and configuration file generation, as well as containerized deployment, greatly reduce the usage threshold. Through performance monitoring tools and log analysis tools, the status of the inference service is monitored in real time, providing a basis for system optimization and improving system reliability and maintainability.
[0040] Model configuration information includes model name, model path, inference accuracy, parallel mode, batch size, and maximum input sequence length. Model configuration information is used to guide the operation and resource allocation of Triton-Inference-Server. Model name and path are used to identify and locate model files, and inference accuracy determines the range of numerical accuracy required for calculation, such as FP32, FP16, or INT8. Parallel mode improves multi-threaded inference performance by specifying the task allocation method of inference service instances. Batch size defines the amount of input data for each inference, which helps to improve computing resource utilization. The maximum input sequence length limits the size of input data, ensuring that inference tasks run within a controllable range while avoiding exceeding memory or hardware limitations. These configuration information are passed to Triton-Inference-Server in the form of configuration files, which are loaded and applied when the service is started. For each configuration parameter, it can be flexibly adjusted according to user needs to achieve a balance between performance and resource utilization. For example, adjusting the batch size can directly affect throughput and latency, while reducing inference accuracy can significantly save computing resources. It allows users to customize the model's inference accuracy, batch size and other parameters according to actual needs, adapt to various inference scenarios, and reasonably set the batch size and parallel mode, which helps to improve system throughput and reduce inference latency. By limiting the maximum input sequence length, it avoids resource waste and optimizes the performance of the inference server under high load. The clearly defined model configuration information structure allows users to quickly configure inference tasks without complex settings.
[0041] The steps to download the model include calling the remote storage interface, including ModelScope API or HuggingFace Transformers Hub API, to obtain the model weights and configuration files. The remote model download function is implemented through standardized storage interface calls, supporting access to multiple mainstream model repositories such as ModelScope API and Hugging Face Transformers Hub API. These interfaces provide a simple calling method, allowing users to quickly obtain the weights and related configuration files of the specified model. During the implementation process, the system will check the availability of the model in the remote storage through API calls based on the model name and path information entered by the user. If the model exists, the system will download its weights and configuration files and store them in the local directory. To ensure the stability of the download process, the system can implement the breakpoint resume function and ensure the integrity and consistency of the file through hash verification. After the download is complete, the system will automatically parse the model's configuration file and regenerate the adapted configuration file format according to the requirements of the inference server to ensure that the model can be correctly loaded and run by Triton-Inference-Server. It supports multiple model repository interfaces to meet users' needs for obtaining models from different platforms. Through standardized API calls and automated download processes, it reduces manual intervention and improves efficiency. The breakpoint resume function and file verification mechanism improve the reliability of the download process. After the download is complete, an adaptive configuration file is automatically generated to reduce the complexity of user model configuration.
[0042] The steps of optimization and parameter fusion are to optimize the model by using the TensorRT engine building tool, and adjust the inference path for specific task types to improve efficiency. As a high-performance inference optimization tool, the TensorRT engine uses a variety of technologies to improve model inference efficiency: layer fusion: merge continuous operations in the calculation graph into one operation to reduce computational redundancy, operator optimization: select the optimal operator implementation supported by the hardware and use GPU-specific instructions for acceleration, memory optimization: dynamically allocate and reuse memory buffers to reduce video memory usage and improve throughput, and specific task adjustments: optimize the weight calculation method for attention mechanism models such as NLP; for computer vision tasks, optimize the convolution kernel structure. During the optimization process, the system analyzes the user's task type and adjusts the inference path according to the task requirements. For example, for multi-input tasks, redundant paths are prioritized; and for time series data processing, streaming inference capabilities are increased to ensure real-time response. Redundant calculations are reduced through layer fusion and operator optimization, dynamic memory management significantly reduces video memory usage, and inference paths are optimized for different tasks to meet a wide range of application needs. The optimized model performs more stably under high-load environments.
[0043] Quantization processing includes FP16 half-precision compression, INT8 quantization, and INT4 quantization, and calibration data sets are used to generate calibration tables to optimize quantization accuracy. Quantization technology reduces computing resource requirements by converting high-precision parameters (such as FP32) into low-precision formats. The system supports the following quantization types: FP16 half-precision compression: maintains a high value range while reducing computing requirements, INT8 quantization: further reduces model accuracy requirements and uses integer format to represent weights and activation values, INT4 quantization: used in extremely low-resource scenarios, suitable for edge devices, and during the quantization process, the system uses calibration data sets to generate calibration tables. The calibration table samples and analyzes the distribution of the model on specific input data to ensure that the quantized model still has high inference accuracy at low precision. In addition, a mixed precision strategy can be used to retain high precision at key layers of the model to further reduce precision loss. Quantization significantly accelerates inference calculations, especially on GPUs and TPUs, reduces video memory usage and power consumption, supports multiple quantization schemes, covers scenario requirements from the cloud to the edge, and reduces precision loss caused by quantization through calibration tables.
[0044] The calibration data set is generated from historical data, specifically including sparse matrices or time series data. The generation of the calibration data set is based on historical data. The following are two typical scenarios: Sparse matrices: suitable for data scenarios with high sparsity, such as the user-item matrix in the recommendation system. The calibration process uses the characteristics of sparse matrices to sample only non-zero elements, reduce the amount of calculation and improve calibration efficiency. Time series data: suitable for prediction tasks, such as financial transactions or equipment monitoring. During calibration, the key eigenvalues of the data (such as extreme values, means, etc.) are extracted to ensure the adaptability of the model to dynamic changes. During the generation of the calibration data set, data preprocessing techniques (such as normalization and standardization) can be combined to improve the calibration accuracy, and support real-time data stream sampling and dynamic update of the calibration table. Supports multiple data types, covering application scenarios such as sparse and time series, reduces computing requirements based on sparsity or feature extraction, and improves system flexibility through real-time sampling and updating.
[0045] The operating environment includes deploying inference service instances using Docker images. The images support multiple CUDA versions and integrate log monitoring plug-ins and load scheduling modules. Containerized deployment provides a consistent operating environment through Docker images. The image is pre-installed with multiple versions of CUDA to support the operating requirements of different hardware. The following modules are integrated in the image: Log monitoring plug-in: real-time collection of inference service operation logs, including exception information and performance data, load scheduling module: dynamic allocation of hardware resources to ensure service stability under high load conditions, the image building process can be customized according to user needs, such as adding specific dependency libraries or scripts to further improve flexibility. Containerization simplifies the environment configuration process, supports different GPU drivers and CUDA versions, integrates load scheduling modules, and optimizes resource utilization.
[0046] Dynamic batching technology includes merging requests by time window through the dynamic batching API of the inference server, and setting the batch size upper limit and latency threshold. Dynamic batching technology aims to optimize the resource utilization and response efficiency of the inference server when processing multiple requests: Time window merging: Through the dynamic batching API of the inference server, the system collects requests in fixed time windows (such as 10ms), merges the collected requests into a batch at the end of the time window, and submits it to the inference engine for execution. Batch size upper limit: Set the maximum number of requests merged in each batch (such as 64 requests) to prevent insufficient memory resources or high inference latency due to too large a batch. Latency threshold: If the waiting time reaches the threshold (such as 50ms) but the number of requests does not reach the batch size upper limit, the system will immediately execute the current batch to reduce the response delay perceived by the user. Dynamic tuning: According to the real-time load situation, adjust the time window length, batch size and latency threshold to find the optimal balance between throughput and response latency. For example, for short text classification tasks, the system can optimize response time through shorter time windows (such as 5ms) and smaller batch sizes (such as 16 requests); while for large-scale model reasoning tasks, it tends to use longer time windows (such as 20ms) and larger batch sizes (such as 128 requests) to improve throughput. Merging small batch requests reduces the overhead of GPU core startup and data transmission, significantly improving reasoning efficiency. The latency threshold mechanism maximizes system resource utilization without sacrificing user experience, dynamically adjusts batch parameters, and adapts to a variety of scenarios from real-time interaction to offline processing, avoiding resource competition or insufficient memory caused by too large batches.
[0047] Performance monitoring tools include Prometheus and Grafana, which are used to monitor GPU utilization, inference latency, and throughput in real time. Performance monitoring tools use a distributed architecture to monitor key performance indicators of inference services in real time: Prometheus actively pulls performance indicators of inference services (such as GPU utilization, video memory usage, inference latency, and throughput). The inference server provides standardized indicator interfaces (such as HTTP / REST API or gRPC) to facilitate Prometheus to collect data regularly. The indicator collection frequency (such as every 1 second) and storage duration (such as 7 days) can be configured to meet different monitoring needs. Prometheus stores the collected data as a time series database and supports a variety of aggregation functions (such as average, maximum, quantile, etc.) to analyze performance change trends. Grafana provides rich visualization functions and supports users to create custom dashboards to display key indicators (such as real-time GPU utilization line charts, inference latency bar charts, etc.). Users can use Grafana's alarm function to set thresholds (such as GPU utilization > 90% or latency > 50ms) to trigger email or SMS notifications when the thresholds are exceeded. Integrated log analysis tools (such as ELK Stack) associate log information with performance indicators to facilitate problem location, and support exporting historical data for offline analysis, such as predicting future load trends through machine learning models. Grafana provides an intuitive performance monitoring interface to facilitate users to quickly understand the system status. Combined with Prometheus's data analysis capabilities, it can quickly identify performance bottlenecks (such as high GPU usage or increased latency), analyze performance trends based on time series data, guide resource allocation and parameter adjustment, and promptly discover and respond to abnormal conditions through the alarm mechanism.
[0048] A large-model high-performance inference acceleration system based on Triton-Inference-Server is also provided, which is used to implement the steps of the large-model high-performance inference acceleration method based on Triton-Inference-Server, and the system includes:
[0049] The model management module is used to receive the model configuration information provided by the user, detect whether there is a local model copy according to the model configuration, and if not, download the model from the remote storage and generate a configuration file that meets the requirements of the inference server;
[0050] The model loading and optimization module is connected to the model management module to support users to upload customized models or fine-tune models, and load, optimize and fuse parameters of the models;
[0051] The model quantization module is connected to the model loading and optimization module, and is used to quantize the model according to user needs and use a variety of quantization methods to optimize the inference performance;
[0052] The operating environment construction module is connected to the model quantization module and is used to build an operating environment based on the inference server, including deploying containerized inference service instances;
[0053] The inference service management module is connected to the operating environment building module to start the inference service in the specified hardware environment and improve the inference efficiency through dynamic batch processing technology and resource management strategies;
[0054] The log analysis and performance monitoring module is connected to the reasoning service management module to collect the operation data of the reasoning service to achieve service optimization and robustness enhancement.
[0055] The system achieves high-performance inference acceleration through modular design. The model management module is responsible for receiving the model configuration information input by the user, detecting whether there is a model copy locally, downloading the required model from the remote storage, and automatically generating the configuration file. The model loading and optimization module supports users to upload fine-tuned models, and combines the TensorRT optimization tool to perform layer fusion, operator optimization, and inference path adjustment operations to improve model performance. The model quantization module generates a calibration table through the calibration data set, supports multiple quantization methods such as FP16, INT8, and INT4, and can dynamically adjust the quantization accuracy to adapt to different load requirements. The operating environment construction module uses Docker containerization to deploy service instances and integrates multiple plug-ins to improve the monitoring capability and scalability of the service. The inference service management module combines dynamic batch processing technology to merge requests by time window to optimize hardware resource utilization. The log analysis and performance monitoring module collects performance data and operation logs through Prometheus and ELK Stack, and combines Grafana to achieve real-time visualization and abnormal alarms, thereby providing data-driven performance optimization methods. Through the collaboration between modules, the system ensures the efficiency of the entire process from model management to inference service deployment and operation. The system has high flexibility due to its modular design and supports full-process automation from model management to inference service management. Through a variety of optimization technologies (such as dynamic batching and model quantization), the inference efficiency is significantly improved and resource usage is reduced. The containerized deployment method makes the system easy to install and maintain, while improving compatibility with different hardware environments. Log analysis and performance monitoring functions enhance the robustness and maintainability of the system, helping users quickly locate performance bottlenecks and optimize services. Overall, the system has high performance, high reliability, and wide applicability, and can meet the needs of various inference scenarios from the cloud to edge devices.
[0056] Although embodiments of the present invention have been shown and described, it will be appreciated by those skilled in the art that various changes, modifications, substitutions and variations may be made to the embodiments without departing from the principles and spirit of the present invention, and that the scope of the present invention is defined by the appended claims and their equivalents.
Claims
1. A large-model high-performance inference acceleration method based on Triton-Inference-Server, characterized in that: The method comprises: S1. Receive the model configuration information provided by the user, detect whether there is a local model copy according to the model configuration, and if not, download the model from the remote storage and generate a configuration file that meets the requirements of the inference server; S2. Support users to upload customized models or fine-tune models, and load, optimize and fuse parameters of the models; S3. Quantify the model according to user needs and use a variety of quantization methods to optimize reasoning performance; S4. Build an operating environment based on the inference server, including deploying containerized inference service instances; S5. Start the inference service in the specified hardware environment and use dynamic batch processing technology and resource management strategies to improve inference efficiency. S6. Collect operation data through log analysis and performance monitoring tools to optimize and enhance the robustness of reasoning services.
2. According to claim 1, the large-model high-performance inference acceleration method based on Triton-Inference-Server is characterized in that: The model configuration information includes model name, model path, inference accuracy, parallel mode, batch size and maximum input sequence length.
3. The large-model high-performance inference acceleration method based on Triton-Inference-Server according to claim 1 is characterized in that: The step of downloading the model includes calling a remote storage interface, including a ModelScope API or a Hugging Face Transformers Hub API, to obtain model weights and configuration files.
4. The large-model high-performance inference acceleration method based on Triton-Inference-Server according to claim 1 is characterized in that: The optimization and parameter fusion steps optimize the model by using the TensorRT engine building tool and adjust the reasoning path for specific task types to improve efficiency.
5. The large-model high-performance inference acceleration method based on Triton-Inference-Server according to claim 1 is characterized in that: The quantization process includes FP16 half-precision compression, INT8 quantization and INT4 quantization, and a calibration table is generated using a calibration data set to optimize quantization accuracy.
6. The large-model high-performance inference acceleration method based on Triton-Inference-Server according to claim 1 is characterized in that: The calibration data set is generated from historical data, and specifically includes sparse matrix or time series data.
7. The large-model high-performance inference acceleration method based on Triton-Inference-Server according to claim 1 is characterized in that: The operating environment includes deploying an inference service instance using a Docker image, which supports multiple CUDA versions and integrates a log monitoring plug-in and a load scheduling module.
8. The large-model high-performance inference acceleration method based on Triton-Inference-Server according to claim 1 is characterized in that: The dynamic batch processing technology includes merging requests according to time windows through the dynamic batch processing API of the inference server and setting the batch size upper limit and delay threshold.
9. The large-model high-performance inference acceleration method based on Triton-Inference-Server according to claim 1 is characterized in that: The performance monitoring tools include Prometheus and Grafana, which are used to monitor GPU utilization, inference latency, and throughput in real time.
10. A large-model high-performance reasoning acceleration system based on Triton-Inference-Server, used to implement the steps of the large-model high-performance reasoning acceleration method based on Triton-Inference-Server according to any one of claims 1 to 9, characterized in that: The system comprises: The model management module is used to receive the model configuration information provided by the user, detect whether there is a local model copy according to the model configuration, and if not, download the model from the remote storage and generate a configuration file that meets the requirements of the inference server; The model loading and optimization module is connected to the model management module to support users to upload customized models or fine-tune models, and load, optimize and fuse parameters of the models; The model quantization module is connected to the model loading and optimization module, and is used to quantize the model according to user needs and use a variety of quantization methods to optimize the inference performance; The operating environment construction module is connected to the model quantization module and is used to build an operating environment based on the inference server, including deploying containerized inference service instances; The inference service management module is connected to the operating environment building module to start the inference service in the specified hardware environment and improve the inference efficiency through dynamic batch processing technology and resource management strategies; The log analysis and performance monitoring module is connected to the reasoning service management module to collect the operation data of the reasoning service to achieve service optimization and robustness enhancement.
Citation Information
Patent Citations
Deep learning model reasoning batch processing optimization method and system
CN113902116A
Model reasoning acceleration method and system based on GPU equipment
CN114092313A
Deep learning model dynamic batch processing scheduling method and system based on resource adjustment
CN114217966A
Optimized deployment method and device for computer vision deep learning model
CN116048542A
Automatic deployment method based on TensorRT-LLM model reasoning acceleration service
CN117992078A