Heterogeneous hardware environment large model inference engine device
By adopting standardized interfaces and distributed design in the inference engine, the problems of adaptation difficulties and insufficient performance of the inference engine in the existing technology in different GPU environments are solved, and efficient and flexible inference service deployment and operation are achieved.
Patent Information
- Application Number
- CN202411960193.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2024-12-30
- Publication Date
- 2025-05-30
AI Technical Summary
Existing inference engines lack flexibility and scalability, making it difficult to deploy inference services on GPUs of different brands or models, increasing adaptation costs and maintenance complexity, affecting inference performance and deployment scope.
A heterogeneous hardware environment large-model inference engine device is designed, adopting multiple standardized inference interfaces and distributed designs, providing a unified inference interface, supporting multiple GPU hardware environments, dynamically identifying and loading GPUs of different models, automatically adapting performance parameters, and processing large models in parallel on multiple GPUs.
It significantly simplifies the adaptation process of different GPU models, reduces development and maintenance costs, improves the operation efficiency and response time of large models, adapts to multiple GPU hardware environments, meets the needs of different enterprises, and provides convenience for hardware expansion and upgrading.
Smart Images

Figure CN120069057A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of large model processing, and particularly to a large model inference engine device in a heterogeneous hardware environment. Background Art
[0002] With the vigorous development of deep learning technology and the popular application of large pre-trained models (such as GPT, BERT, etc.), while enterprises enjoy the convenience brought by these advanced technologies, they also face unprecedented challenges. Due to differences in business requirements, budget constraints, and technical preferences among different enterprises, the GPU brands, models, and configurations purchased show significant diversity. The complexity of this hardware environment is particularly prominent during the model inference process, not only increasing the cost of adapting to different hardware but also making the construction and maintenance of the inference system more cumbersome.
[0003] Most existing inference engines focus on specific hardware platforms and lack sufficient flexibility and scalability. This means that when enterprises need to deploy inference services on GPUs of different brands or models, they often need to spend a lot of time and effort on adaptation and optimization. This limitation not only restricts the deployment scope of the inference service but also may affect the improvement of inference performance, making it difficult to meet the urgent needs of enterprises for efficient and reliable inference services. Summary of the Invention
[0004] In view of this, the present invention provides a large model inference engine device in a heterogeneous hardware environment, which can meet the needs of enterprises in different hardware environments, reduce the adaptation cost, and improve the inference efficiency.
[0005] An embodiment of the present invention provides a large model inference engine device in a heterogeneous hardware environment, and the device includes:
[0006] Multiple standardized inference interfaces for providing multiple standardized functions, loading a large model into the GPU hardware driver, and performing model inference calculations;
[0007] A distributed design for running large models with different parameter quantities in different GPUs, where the different GPUs are heterogeneous hardware.
[0008] Further, the standardized inference interface includes:
[0009] Multiple standardized interfaces for providing multiple standardized functions, where the multiple standardized functions include model inference, model fine-tuning, Embedding generation, and Reranker functions;
[0010] Multiple hardware driver interfaces for dynamically identifying and loading different models of GPUs, automatically adapting their performance parameters, and loading a large model into the GPU hardware driver.
[0011] Further, the design of each standardized interface includes determining the request path, request parameters, and response parameters.
[0012] Further, the interface path of the standardized interface corresponding to the model inference is: / reasoning, the request method is the newly added POST, the request parameters include the model unique identifier and input data, and the response parameters include the inference result and the first status code, where the first status code is used to indicate whether the inference is successful.
[0013] Further, the standardized interface path corresponding to the model fine-tuning is: / finetuning, the request method is the newly added POST, the request parameters include the unique identifier of the model to be fine-tuned, fine-tuning parameters, and training data, and the response parameters include the unique identifier of the fine-tuned model and the second status code, where the second status code is used to indicate whether the fine-tuning is successful.
[0014] Further, the standardized interface path corresponding to the Embedding generation is: / embedding, the request method is the newly added POST, the request parameters include the input data and the dimension of the Embedding vector, and the response parameters include the generated Embedding vector and the third status code, where the third status code is used to indicate whether the Embedding generation is successful.
[0015] Further, the standardized interface path corresponding to the Reranker function is: / reranking, the request method is the newly added POST, the request parameters include the initial ranking result and Reranker parameters, and the response parameters include the generated re-ranked result and the fourth status code, where the fourth status code is used to indicate whether the Reranker function is successful.
[0016] Further, the design of multiple hardware driver interfaces includes the following modules:
[0017] HAL creation module, used to create the Hardware Abstraction Layer (HAL), where the HAL includes the access and control functions of GPU registers, as well as the management and scheduling functions of GPU resources;
[0018] Dynamic recognition module, used to detect the installed GPU model in the system at startup;
[0019] Performance parameter adaptation module, used to automatically adjust the performance parameters of the driver according to the recognized GPU model
[0020] Large model loading module, used to run large models with different numbers of parameters on different GPUs, where the different GPUs are heterogeneous hardware.
[0021] Further, the distributed design interface supports a task scheduling and load balancing mechanism, a model sharding mechanism, and a fault tolerance mechanism.
[0022] Further, the task scheduling and load balancing mechanism uses an intelligent algorithm to dynamically schedule inference tasks and dynamically adjusts the amount of resources allocated to each GPU according to the real-time status of the GPU and the task requirements;
[0023] The model sharding mechanism allocates each layer or some layers of the large model to different GPUs, or divides the parameters of the large model into multiple segments according to preset rules and allocates them to different GPUs;
[0024] The fault tolerance mechanism automatically reallocates tasks to other GPUs in case of GPU failure or network anomaly.
[0025] The technical solution provided by the present invention designs the large model inference engine as multiple standardized interfaces and a distributed design. Among them, the multiple standardized inference interfaces are used to provide multiple standardized functions, load the large model into the GPU hardware driver, and execute model inference calculations. The distributed design is used to deploy multiple large models with different numbers of parameters to multiple GPU graphics card hardwares simultaneously. Thus, the inference engine in the embodiments of the present invention provides a unified inference interface, greatly simplifies the adaptation process for different GPU models, significantly reduces the development and maintenance costs. Thanks to its distributed design, the engine supports parallel inference, significantly improves the running efficiency of the large model, greatly shortens the response time. The engine can adapt to a variety of GPU hardware environments, flexibly meet the specific needs of different enterprises, provides great convenience for subsequent hardware expansion and upgrade. The standardized interface design reduces the learning threshold for developers, accelerates the model deployment and application process, makes the development process smoother, and can be widely applied to multiple fields such as natural language processing, computer vision, and multi-modal, showing strong market potential and application value. BRIEF DESCRIPTION OF THE DRAWINGS
[0026] Figure 1 is a schematic structural diagram of a large model inference engine device in a heterogeneous hardware environment provided by an embodiment of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0027] In order to make the objectives, technical solutions, and advantages of the present invention clearer, the present invention will be further described in detail below with reference to the accompanying drawings. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of the present invention.
[0028] Embodiment 1
[0029] See Figure 1 , Figure 1 which is a schematic structural diagram of the large model inference engine device in the heterogeneous hardware environment provided by the embodiment of the present invention.
[0030] Among them, the heterogeneous hardware environment large model inference engine device includes a plurality of standardized inference interfaces 11 and a distributed design interface 12. Among them, the plurality of standardized inference interfaces 11 are used to provide a plurality of standardized functions and load the large model into the GPU hardware driver, and perform model inference calculations. The distributed design interface 12 is used to run large models with different parameter quantities in different GPUs, and the different GPUs are heterogeneous hardware.
[0031] The heterogeneous hardware environment large model inference engine device provided by the embodiment of the present invention includes two key components: a plurality of standardized inference interfaces 11 and a distributed design interface 12.
[0032] Among them, the plurality of standardized inference interfaces provide a series of standardized functions. These functions include but are not limited to model inference (i.e., using the loaded model to predict or analyze data), fine-tuning (adjusting the parameters of the pre-trained model on a specific dataset to improve performance), Embedding generation (converting input data into a low-dimensional vector representation, usually used as the input of a machine learning model), and Reranker function (sorting or re-evaluating the inference results according to specific metrics). At the same time, the standardized inference interface is also responsible for loading the large model into the GPU hardware driver. This means that the interface can interact with the GPU hardware to ensure that the model can efficiently use the computing resources of the GPU to perform inference calculations.
[0033] The distributed design interface is designed to support the simultaneous processing of multiple large models with different sizes and parameters. In a heterogeneous hardware environment, this means that the interface needs to be able to manage inference tasks across multiple GPUs, or even across multiple servers or clusters.
[0034] In some embodiments of the present invention, the standardized inference interface 11 includes a plurality of standardized interfaces 111 and a plurality of hardware driver interfaces 112. Among them, the plurality of standardized interfaces 111 are used to provide a plurality of standardized functions, and the plurality of standardized functions include model inference, model fine-tuning, Embedding generation, and Reranker function. The plurality of hardware driver interfaces 112 are used to dynamically identify and load different models of GPUs, automatically adapt their performance parameters, and load the large model into a plurality of different GPU hardware drivers, where the plurality of GPU hardware drivers are heterogeneous hardware drivers.
[0035] The standardized interface 111 provides a series of standardized functions to ensure that the inference engine can execute different inference tasks in a unified and efficient manner, including model inference: using the loaded model to predict or analyze data; model fine-tuning: adjusting the parameters of the pre-trained model on a specific dataset to make it more suitable for a specific task or dataset; Embedding generation: converting input data (such as text, images, etc.) into low-dimensional vector representations, which are usually used as inputs for machine learning models; Reranker function: sorting or re-evaluating the inference results, usually used to improve the quality of search results or recommendation lists.
[0036] The hardware driver interface 112 is responsible for dynamically identifying and loading different models of GPUs and ensuring that the inference engine can automatically adapt to the performance parameters of the GPUs. The hardware driver interface can identify the GPU models installed in the system and load the corresponding driver programs as needed. That is to say, the inference engine can flexibly run in a hardware environment with different GPU configurations. Once the GPU is identified and loaded, the hardware driver interface will automatically adapt to its performance parameters, such as memory size, computing speed, etc. This ensures that the inference engine can make full use of the computing resources of the GPU while avoiding resource waste or performance bottlenecks. Most importantly, the hardware driver interface is responsible for loading the large model into the GPU hardware driver. Therefore, the model data will be transferred to the memory of the GPU and be ready for inference calculation. Since the GPU is good at parallel computing, this loading method can significantly improve the inference speed.
[0037] In some embodiments, the design of each standardized interface includes determining the request path, request parameters, and response parameters.
[0038] The interface path of the standardized interface corresponding to the model inference is: / reasoning, the request method is the newly added POST, the request parameters include the model unique identifier and the input data, and the response parameters include the inference result and the first status code, where the first status code is used to indicate whether the inference is successful.
[0039] Among them, the interface path / reasoning: the URL path of the standardized interface corresponding to the model inference function. When model inference needs to be executed, the client will send a request to this path.
[0040] The newly added POST request method: indicates that the client needs to use the POST request method to send data to the / reasoning interface. The POST request is usually used to submit data to the server for processing, such as submitting form data or uploading files.
[0041] Model unique identifier in the request parameters: This is a unique identifier used to identify a specific model. When sending an inference request, the client needs to specify which model to use for inference, and this identifier is used to achieve this. Input data: This is the input data that needs to be inferred. The format and content of the input data depend on the model and task being used. Inference result: This is the result returned by the server after successfully performing the inference.
[0042] The specific format and content of the inference result in the response parameters depend on the model and task being used. First status code: This is a status code used to indicate whether the inference was successful. The status code is usually a number or string that indicates the result of processing the request. For example, a status code of 200 may indicate successful inference, while status codes in the 4xx or 5xx range may indicate an error.
[0043] In some embodiments of the present invention, the standardized interface path corresponding to the model fine-tuning is: / finetuning, the request method is the new POST, the request parameters include the unique identifier of the model to be fine-tuned, the fine-tuning parameters, and the training data, and the response parameters include the unique identifier of the fine-tuned model and the second status code, and the second status code is used to indicate whether the fine-tuning was successful.
[0044] Standardized interface path / finetuning: The URL path of the standardized interface corresponding to the model fine-tuning function. When the model needs to be fine-tuned, the client sends a request to this path.
[0045] Request method new POST: This indicates that the client needs to use the POST request method to send data to the / finetuning interface. The POST request is used here to submit the parameters and data required for fine-tuning.
[0046] Unique identifier of the model to be fine-tuned in the request parameters: This is a unique identifier used to identify a specific model, indicating which model needs to be fine-tuned. Fine-tuning parameters: These parameters are used to specify various settings in the fine-tuning process, such as the learning rate, number of iterations, regularization terms, etc. The specific values of these parameters depend on the requirements of the fine-tuning task and the model being used. Training data: This is the training data used to fine-tune the model. This data is usually a labeled dataset used to guide the model on how to adjust its parameters during fine-tuning.
[0047] Unique identifier of the fine-tuned model in the response parameters: This is a new unique identifier used to identify the fine-tuned model. This identifier is usually generated by the server after successful fine-tuning. Second status code: This is a status code used to indicate whether the fine-tuning was successful. The status code is usually a number or string indicating the result of the request processing. For example, a status code of 200 may indicate successful fine-tuning, while status codes in the 4xx or 5xx range may indicate an error.
[0048] In some embodiments of the present invention, the standardized interface path corresponding to the Embedding generation is: / embedding, the request method is the new POST, the request parameters include the input data and the dimension of the Embedding vector, and the response parameters include the generated Embedding vector and the third status code, where the third status code is used to indicate whether the Embedding generation was successful.
[0049] Standardized interface path / embedding: The URL path of the standardized interface corresponding to the Embedding generation function. When an Embedding vector of data needs to be generated, the client sends a request to this path.
[0050] Request method new POST: This means that the client needs to use the POST request method to send data to the / embedding interface. The POST request is used here to submit the input data for which the Embedding needs to be generated and the dimension of the Embedding vector.
[0051] Input data in the request parameters: This is the original data for which the Embedding vector needs to be generated. The format and content of the input data depend on the model and task used, such as text, images, or other types of data. Dimension of the Embedding vector: This is a parameter specifying how many dimensions the generated Embedding vector should have. The size of the dimension usually affects the expressive power and computational complexity of the Embedding vector.
[0052] Generated Embedding vector in the response parameters: This is the vector returned by the server after successful execution of the Embedding generation. This vector is a low-dimensional numerical representation, usually used as input or feature representation for machine learning models. Third status code: This is a status code used to indicate whether the Embedding generation was successful. The status code is usually a number or string indicating the result of the request processing. For example, a status code of 200 may indicate successful Embedding generation, while status codes in the 4xx or 5xx range may indicate an error.
[0053] The standardized interface path corresponding to the Reranker function is: / reranking. The request method is the newly added POST. The request parameters include the initial ranking results and the Reranker parameters. The response parameters include the generated re-ranked results and the fourth status code. The fourth status code is used to indicate whether the Reranker function is successful.
[0054] Interface path / reranking: The URL path of the standardized interface corresponding to the Reranker function. When a set of initial ranking results needs to be re-ranked, the client sends a request to this path.
[0055] Request method newly added POST: The client uses the POST request method to send data to the / reranking interface. The POST request is used here to submit the initial ranking results that need to be re-ranked and the parameters required by the Reranker.
[0056] Initial ranking results in the request parameters: This is a list or array that contains the results preliminarily ranked according to a certain standard. These results may be a list of documents returned by a search engine, a list of products generated by a recommendation system, or any other list of entities that need to be ranked. Reranker parameters: These parameters are used to specify various settings and algorithms used by the Reranker during the re-ranking process. These parameters may include model weights, feature selection, ranking criteria, etc., depending on the implementation of the Reranker and the type of task being processed.
[0057] Generated re-ranked results in the response parameters: This is the new ranked list returned by the server after successfully executing the Reranker function. The order of the elements in this list is obtained by re-ranking the initial ranking results according to the algorithm and parameters of the Reranker. Fourth status code: This is a status code used to indicate whether the Reranker function is successful. The status code is a number or string used to indicate the processing result of the request. For example, a status code of 200 may indicate that the Reranker function was successfully executed, while a status code in the 4xx or 5xx range may indicate an error.
[0058] In some embodiments of the present invention, when designing multiple hardware driver interfaces, it includes a HAL creation module, a dynamic recognition module, a performance parameter adaptation module, and a large model loading module, where,
[0059] The HAL creation module is used to create the Hardware Abstraction Layer (HAL), which includes functions for accessing and controlling GPU registers, as well as functions for managing and scheduling GPU resources. The dynamic identification module is used to detect the GPU model installed in the system at startup. The performance parameter adaptation module is used to automatically adjust the performance parameters of the driver according to the identified GPU model. The large model loading module is used to load the large model into the GPU memory.
[0060] HAL is an intermediate layer between the operating system and the hardware. It provides a unified interface to the hardware devices, enabling the operating system and upper-layer applications to not have to concern themselves with the specific implementation details of the underlying hardware. The HAL creation module will implement a set of functions for accessing and controlling GPU registers, as well as functions for managing and scheduling GPU resources (such as video memory, computing units, etc.). These functions allow upper-layer applications to efficiently utilize GPU resources through the HAL. The dynamic identification module can determine the GPU model in the current system by reading the identifiers on the GPU hardware (such as PCI device ID, manufacturer ID, etc.). This information is crucial for subsequent driver loading and performance parameter adaptation. Different GPU models may have different performance characteristics and power consumption requirements. The performance parameter adaptation module will adjust parameters such as the clock frequency, voltage settings, and power consumption limits of the driver based on this information to optimize the performance and power consumption performance of the GPU. In deep learning and other high-performance computing applications, models are usually very large and require a large amount of memory and computing resources. The large model loading module is responsible for transferring the model data from the hard disk or other storage devices to the GPU memory and preparing the relevant computing resources for subsequent model inference or training operations.
[0061] In some embodiments of the present invention, the distributed design interface supports a task scheduling and load balancing mechanism, a model sharding mechanism, and a fault tolerance mechanism. Among them, the task scheduling and load balancing mechanism uses intelligent algorithms to dynamically schedule inference tasks and dynamically adjusts the amount of resources allocated to each GPU according to the real-time status of the GPU and the task requirements. The model sharding mechanism allocates each layer or some layers of the large model to different GPUs, or divides the parameters of the large model into multiple segments according to preset rules and allocates them to different GPUs. The fault tolerance mechanism automatically reallocates tasks to other GPUs in the event of GPU failures or network anomalies.
[0062] The task scheduling and load balancing mechanism uses intelligent algorithms to dynamically schedule inference tasks, and dynamically adjusts the amount of resources allocated to each GPU according to the real-time status of the GPU and task requirements. This can be achieved in the following ways: It may include using intelligent algorithms to evaluate factors such as task queues, GPU loads, network latencies, etc., and making scheduling decisions based on this. The intelligent algorithm may include heuristic algorithms, machine learning algorithms, or rule-based algorithms. Then, according to the real-time status of the GPU (such as temperature, power consumption, memory usage, etc.) and task requirements (such as computational complexity, real-time requirements, etc.), the task allocation on the GPU is dynamically adjusted to optimize the overall performance and resource utilization. Finally, ensure that the loads among the GPUs are relatively balanced, avoiding the situation where some GPUs are overloaded while others are idle, thereby improving the overall throughput and response speed of the system.
[0063] When implementing the model sharding mechanism, it can be divided into hierarchical sharding and parameter sharding. Among them, hierarchical sharding: For deep learning models, different layers can be allocated to different GPUs according to the hierarchical structure of the model to achieve parallel computing. Parameter sharding: For models with a large number of parameters, the parameters can be divided into multiple segments and allocated to different GPUs to reduce the memory occupancy of a single GPU and improve computational efficiency. During the sharding process, it is necessary to ensure synchronization and communication among different GPUs to ensure the consistency and correctness of the model.
[0064] Fault tolerance mechanism is a crucial part of distributed systems, especially in task processing environments involving GPU acceleration. When a GPU fails or the network experiences anomalies, this mechanism can automatically reassign the tasks originally allocated to the faulty GPU to other healthy GPUs, thus ensuring the continuity of the system and the smooth completion of tasks. The implementation methods may include the following: (1) Fault monitoring, using hardware monitoring tools to detect the status of GPUs such as temperature, power consumption, error codes, etc. Or monitor the stability of network connections and bandwidth usage, and then set thresholds and rules to trigger fault alarms when anomalies are detected; (2) Fault reallocation, maintaining a task queue and a GPU resource pool, recording the tasks currently allocated to each GPU. When a GPU fault is detected, mark the tasks on this GPU in the task queue as "to be reallocated", and according to the available resources in the GPU resource pool, use a certain scheduling algorithm (such as round-robin, weighted round-robin, least connections, etc.) to allocate the to-be-reallocated tasks to other GPUs. If network failures prevent communication between GPUs, it may be necessary to consider migrating tasks to other GPUs within the same network partition; (3) Data consistency, ensuring the consistency and integrity of data during task reallocation. If the tasks involve distributed computing or data parallelism, data migration and synchronization issues need to be addressed, and persistent storage or distributed caches are used to save intermediate results and status information; (4) Logging and monitoring, recording the logs of fault detection, task reallocation, and any related operations, and providing a monitoring interface or API so that administrators can view the status and performance of the system in real time; (5) Recovery strategies, for temporary faults (such as network jitters), fast recovery strategies can be implemented, such as retrying connections or retrying after waiting for a period of time. For permanent faults (such as GPU hardware damage), more complex recovery processes are required, which may include replacing faulty hardware, reconfiguring the system, etc. By implementing these fault tolerance mechanisms, distributed systems can maintain a high degree of reliability and stability in the face of challenges such as GPU faults or network anomalies, ensuring the smooth completion of tasks and the continuous operation of the system.
[0065] The technical solution provided by the present invention designs the large model inference engine as multiple standardized interfaces and a distributed design. Among them, the multiple standardized inference interfaces are used to provide multiple standardized functions, load the large model into the GPU hardware driver, and perform model inference calculations. The distributed design is used to deploy multiple large models with different parameter quantities to multiple GPU graphics card hardwares simultaneously. Thus, the inference engine in the embodiments of the present invention provides a unified inference interface, greatly simplifies the adaptation process for different GPU models, significantly reduces the development and maintenance costs. Thanks to its distributed design, the engine supports parallel inference, significantly improving the operation efficiency of the large model and greatly shortening the response time. The engine can adapt to a variety of GPU hardware environments, flexibly meet the specific needs of different enterprises, and provides great convenience for subsequent hardware expansion and upgrade. The standardized interface design reduces the learning threshold for developers, accelerates the model deployment and application process, makes the development process smoother, and can be widely applied to multiple fields such as natural language processing, computer vision, and multi-modal, showing strong market potential and application value.
[0066] The above specific implementation manners do not constitute a limitation on the protection scope of the present invention. Those skilled in the art should understand that various modifications, combinations, sub-combinations, and substitutions can be made according to design requirements and other factors. Any modifications, equivalent substitutions, and improvements made within the spirit and principle of the present invention shall be included within the protection scope of the present invention.
Claims
1. A large model inference engine device in a heterogeneous hardware environment, characterized in that: The device comprises: Multiple standardized reasoning interfaces to provide multiple standardized functions and load large models into GPU hardware drivers and perform model reasoning calculations; The distributed design is used to run large models with different parameter quantities on different GPUs, where the different GPUs are heterogeneous hardware.
2. The device according to claim 1, characterized in that The standardized reasoning interface includes: A plurality of standardized interfaces for providing a plurality of standardized functions, wherein the plurality of standardized functions include model reasoning, model fine-tuning, embedding generation and reranker functions; Multiple hardware driver interfaces are used to dynamically identify and load different types of GPUs, automatically adapt their performance parameters, and load large models into the GPU hardware driver.
3. The device according to claim 2, characterized in that The design of each standardized interface includes determining the request path, request parameters, and response parameters.
4. The device according to claim 3, characterized in that The interface path of the standardized interface corresponding to the model reasoning is: / reasoning, the request method is the newly added POST, the request parameters include the model unique identifier and input data, and the response parameters include the reasoning result and the first status code, and the first status code is used to indicate whether the reasoning is successful.
5. The device according to claim 3, characterized in that The standardized interface path corresponding to the model fine-tuning is: / finetuning, the request method is the newly added POST, the request parameters include the unique identifier of the model to be fine-tuned, the fine-tuning parameters, and the training data, and the response parameters include the unique identifier of the fine-tuned model and the second status code, and the second status code is used to indicate whether the fine-tuning is successful.
6. The device according to claim 3, characterized in that The standardized interface path corresponding to the Embedding generation is: / embedding, the request method is the newly added POST, the request parameters include the input data and the dimensions of the Embedding vector, the response parameters include the generated Embedding vector and the third status code, and the third status code is used to indicate whether the Embedding generation is successful.
7. The device according to claim 3, characterized in that The standardized interface path corresponding to the Reranker function is: / reranking, the request method is the newly added POST, the request parameters include the initial ranking results and the Reranker parameters, and the response parameters include the generated re-ranked results and the fourth status code. The fourth status code is used to indicate whether the Reranker function is successful.
8. The device according to claim 3, characterized in that Multiple hardware driver interface designs include the following modules: A HAL creation module, used to create a hardware abstraction layer (HAL), wherein the HAL includes functions for accessing and controlling GPU registers, and functions for managing and scheduling GPU resources; Dynamic identification module, used to detect the GPU model installed in the system at startup; Performance parameter adaptation module, used to automatically adjust the performance parameters of the driver according to the identified GPU model The large model loading module is used to run large models with different parameter quantities in different GPUs, where the different GPUs are heterogeneous hardware.
9. The device according to claim 1, characterized in that The distributed design interface supports task scheduling and load balancing mechanisms, model sharding mechanisms and fault tolerance mechanisms.
10. The device according to claim 9, characterized in that The task scheduling and load balancing mechanism uses intelligent algorithms to dynamically schedule inference tasks and dynamically adjust the amount of resources allocated to each GPU based on the real-time status of the GPU and task requirements; The model sharding mechanism allocates each layer or some layers of the large model to different GPUs, or divides the parameters of the large model into multiple fragments according to preset rules and allocates them to different GPUs; The fault-tolerance mechanism automatically reallocates tasks to other GPUs in the event of a GPU failure or network anomaly.