Inference service resource configuration method, electronic equipment and readable storage medium
By parsing the model file to determine the minimum storage requirements of the inference service, and combining the inference cluster node status, it automatically configures resource specifications and policies, solving the problems of low automation and poor utilization of model inference service resource configuration, and achieving efficient storage resource management.
Patent Information
- Application Number
- CN202511232793.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-08-29
- Publication Date
- 2025-10-03
- Estimated Expiration
- 2045-08-29
AI Technical Summary
The resource configuration automation level of model inference services in existing technologies is low, and storage resource utilization is poor, resulting in a high risk of manual operation errors and waste or insufficient resources.
By parsing the model file to determine the minimum storage resource requirements, and based on the storage resource status of the inference cluster nodes, automatically configure resource specifications and policies that meet the requirements to achieve automated storage resource configuration.
It improves the utilization of storage resources, reduces the risk of manual operation errors, avoids resource waste or shortage, and optimizes resource allocation efficiency.
Smart Images

Figure CN120743558A_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of computer technology, and in particular to a method for configuring inference service resources, an electronic device, and a readable storage medium. Background Art
[0002] With the rapid development of computer technology, pre-trained models are increasingly being used in fields such as natural language processing and computer vision. Model inference services require a significant amount of storage resources (such as graphics memory) to store data and ensure proper execution. Conventional technologies typically require manual evaluation and configuration of storage resources required for inference services before model deployment and inference service execution. This results in a low level of automation and poor storage resource utilization. Summary of the Invention
[0003] The present application provides a method for configuring inference service resources, an electronic device, and a readable storage medium to at least solve the problems of low automation and poor utilization of storage resources when configuring resources for inference services in related technologies.
[0004] This application provides a method for configuring inference service resources, including: Obtaining a model file for an inference service, and parsing the model file to determine a minimum storage resource requirement for running the inference service; Determining a resource specification that meets the minimum storage resource requirement, where the resource specification is used to characterize the nodes required to run the inference service and the storage resource configuration corresponding to the nodes; According to the storage resource status of the nodes in the inference cluster, a resource configuration policy that matches the resource specification is determined, and storage resources are configured in the inference cluster according to the resource configuration policy to run the inference service.
[0005] The present application also provides a computer program product, comprising: A first processing module is configured to obtain a model file of an inference service and parse the model file to determine a minimum storage resource requirement for running the inference service; A second processing module is used to determine a resource specification that meets the minimum storage resource requirement, where the resource specification is used to represent the nodes required to run the inference service and the storage resource configuration corresponding to the nodes; The third processing module is used to determine a resource configuration strategy that matches the resource specifications according to the storage resource status of the nodes in the inference cluster, and configure storage resources in the inference cluster according to the resource configuration strategy to run the inference service.
[0006] The present application also provides an electronic device, comprising: a memory for storing a computer program; and a processor for implementing the steps of any of the above-mentioned inference service resource configuration methods when executing the computer program.
[0007] The present application also provides a non-volatile computer-readable storage medium, in which a computer program is stored. When the computer program is executed by a processor, the steps of any of the above-mentioned reasoning service resource configuration methods are implemented.
[0008] The present application also provides another computer program product, including a computer program, which implements the steps of any of the above-mentioned reasoning service resource configuration methods when executed by a processor.
[0009] Through the inference service resource configuration method provided in this application, the model file of the inference service is obtained, and the model file is parsed to determine the minimum storage resource requirement required to run the inference service, providing a data basis for subsequent automated resource configuration. After determining the minimum storage resource requirement, the resource specifications that meet the minimum storage resource requirement are determined. According to the storage resource status resource configuration strategy of the nodes in the inference cluster, the resource configuration strategy is used to indicate how to configure storage resources in the inference cluster to match the resource specifications. Storage resources are configured in the inference cluster according to the resource allocation strategy to run the inference service. Therefore, it can solve the technical problems of low automation level and poor storage resource utilization when configuring resources for the inference service, realize the automatic determination of storage resources required for the inference service and the automatic execution of storage resource configuration, reduce the error risk and cost of manual operation, avoid resource waste or shortage due to fragmentation, and improve the utilization of storage resources. BRIEF DESCRIPTION OF THE DRAWINGS
[0010] In order to more clearly illustrate the embodiments of the present application, the following is a brief introduction to the drawings required for use in the embodiments. Obviously, the drawings described below are only some embodiments of the present application. For ordinary technicians in this field, other drawings can be obtained based on these drawings without any creative work.
[0011] Figure 1 This is a flowchart of a method for configuring inference service resources provided in some embodiments of the present application; Figure 2 This is a second flowchart of a method for configuring inference service resources provided in some embodiments of the present application; Figure 3 This is a flowchart of a method for configuring inference service resources according to some embodiments of the present application; Figure 4 This is a fourth flowchart of a method for configuring inference service resources provided in some embodiments of the present application; Figure 5 Schematic diagram of a computer program product provided for some embodiments of the present application.
[0012] Description of reference numerals: 500: computer program product; 501: first processing module; 502: second processing module; 503: The third processing module. DETAILED DESCRIPTION
[0013] The following will be combined with the accompanying drawings in the embodiments of this application to clearly and completely describe the technical solutions in the embodiments of this application. Obviously, the embodiments described are only part of the embodiments of this application, not all of them. Based on the embodiments in this application, all other embodiments obtained by ordinary technicians in this field without making creative efforts are within the scope of protection of this application.
[0014] It should be noted that, in the description of this application, the terms "comprises," "includes," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or device comprising a series of elements includes not only those elements, but also other elements not explicitly listed, or elements inherent to such process, method, article, or device. The terms "first," "second," etc., in this application are used to distinguish similar objects, and are not used to describe a particular order or sequence.
[0015] In order to enable those skilled in the art to better understand the present application, the present application is further described in detail below with reference to the accompanying drawings and specific implementation methods.
[0016] The embodiments of the present application provide a method for configuring inference service resources, and the method is described in detail in conjunction with the execution process of the method for configuring inference service resources.
[0017] Figure 1 This is one of the flow charts of the method for configuring inference service resources provided in some embodiments of this application. Figure 1 As shown, the reasoning service resource configuration method includes: step 110, step 120 and step 130.
[0018] Step 110: Obtain a model file of the inference service, and parse the model file to determine the minimum storage resource requirement for running the inference service.
[0019] Understandably, inference services involve deploying trained machine learning or deep learning models into production environments to perform real-time predictions and inferences on new data. The inference process typically involves processing input data and calculating outputs through model calculations. Model files contain the parameters and structure of the model used to execute inference services. These files are used to load and execute the model during inference. Model files can be considered to contain the information necessary to perform inference tasks.
[0020] Storage resources refer to the hardware and software resources used to store data, including but not limited to graphics processing unit (GPU) memory (also known as video memory), hard disk storage, solid-state storage, and in-memory storage. The performance and capacity of storage resources directly impact data read and write speeds and storage capacity. Minimum storage resource requirements refer to the minimum storage resources required for the normal operation of inference services. Understandably, complex deep learning models typically require more storage space to store model parameters and intermediate calculation results, while high-frequency real-time inference services require faster storage access speeds.
[0021] Parse the model file to determine the minimum storage resource requirements for running the inference service. For example, the minimum storage space required for the service operation can be calculated by analyzing the model size, data set size, or expected load of the model (such as the amount of model data, etc.) contained in the model file. In some embodiments, an automated model analysis tool can be used to automatically evaluate the storage resources required for running the inference service. If there is historical operation data of the target inference service, the resource requirements of the target inference service can also be analyzed based on the historical operation data.
[0022] In some embodiments, obtaining a model file for an inference service and parsing the model file to determine the minimum storage resource requirement for running the inference service includes: Obtain a model file for the inference service, and parse the model file to obtain model parameters and model quantization accuracy corresponding to the inference service; Based on the model parameter quantity and the model quantization accuracy, a minimum storage resource requirement required to run the inference service is determined.
[0023] Model parameterization refers to the total number of weights and bias parameters in a model. It is often used to measure the complexity and scale of a model and directly affects the model's storage and computing resource requirements. Model quantization accuracy refers to the numerical precision used in model parameters and model inference.
[0024] For example, this can be done by directly extracting model parameter quantities and quantization accuracy information by accessing the model file, or by parsing the model file using automated scripts or specialized tools. Determining the minimum storage resource requirements for running an inference service based on model parameter quantities and quantization accuracy can involve calculating the resources required for storage and computation based on the model parameter quantities and quantization accuracy, or using simulation or machine learning methods to predict resource requirements for different parameter quantities and accuracies.
[0025] By parsing the model file, the model parameter quantity and quantization accuracy are determined, and then the minimum storage resource requirements required to run the inference service are determined, providing a data basis for subsequent automated resource allocation. This helps to improve the efficiency and accuracy of resource allocation while ensuring that the inference service can operate normally.
[0026] Here is an example: Parse the model file to obtain model parameters , in billions. For example, you can get it by loading the model file and calculating the number of tensor arrays. Parse the model file to obtain model quantization accuracy For example, you can obtain the model quantization accuracy by reading the parameter quantization_config.quant_method (a parameter in the config.json file) in the config.json (a common configuration file format) file, or reading the parameter torch_dtype (a parameter in the config.json file) in the file. Commonly used model quantization accuracy includes 8-bit floating point (FP8), 16-bit floating point (FP16), 16-bit brain floating point (FP16), and 32-bit floating point (FP32). Using storage resources as video memory For example, the following method can be used to implement the model parameter quantity And the model quantization accuracy , determine the minimum storage resource requirements required to run the inference service: The model quantization precision is 8-bit floating point ( ) ; The model quantization precision is 16-bit floating point ( ) ; The model quantization precision is 32-bit floating point ( ) .
[0027] Step 120: Determine a resource specification that meets the minimum storage resource requirement, where the resource specification is used to characterize the nodes required to run the inference service and the storage resource configuration corresponding to the nodes.
[0028] It's understood that a node refers to an independent unit within a cluster. For example, a node can be any device equipped with one or more graphics processing units (GPUs). Nodes work together to provide computing and storage capabilities. An inference cluster is a group of nodes working together to provide scalable inference capabilities and support large-scale data processing.
[0029] Resource specifications guide how to configure storage resources within an inference cluster to meet the storage requirements of the inference service (at least the minimum storage resource requirements). Specifically, a resource specification can be considered a set of requirements and configurations for nodes and their corresponding storage resources for running the inference service. In some embodiments, resource specifications may also include requirements for compute resources, storage resources, memory, network bandwidth, and more.
[0030] For the inference service, there can be multiple resource specifications that meet the minimum storage resource requirements.
[0031] Table 1 Resource specification example
[0032] Taking the resource specification example table in Table 1 as an example, for an inference service with a minimum storage resource requirement of approximately 809 GB, its resource specification can be that running the inference service requires a node including 6 A-type GPUs with a memory of 141 GB, it can be a node including 9 B-type GPUs with a memory of 96 GB, or it can be a node including 21 C-type GPUs with a memory of 40 GB. If the inference service adopts a multi-node distributed deployment, the resource specification can also be that running the inference service requires two nodes including 5 B-type GPUs with a memory of 96 GB.
[0033] The above example illustrates that multiple resource specifications can exist for an inference service that meet the minimum storage resource requirements. In practice, resource specifications can be used to represent more information, such as the number of nodes, the number of inference service instances, and the parallelization strategy of the inference service.
[0034] Figure 2This is a second flow chart of the method for configuring inference service resources provided in some embodiments of this application. Figure 2 As shown, in some embodiments, the reasoning service resource configuration method further includes: Receive inference service execution requests; obtain and parse the inference service model file; Determine whether model information corresponding to the inference service exists; If it exists, directly calculate the graphics memory required to run the inference service (corresponding to the minimum storage resource requirement mentioned above); If it does not exist, you need to calculate the model parameters and obtain the model quantization accuracy before calculating the graphics memory required to run the inference service; After calculating the graphics memory required to run the inference service, obtain the specifications of the existing GPUs in the cluster and determine the number of GPUs of different specifications required to run the inference service. Save the calculated model information (such as model parameter quantity, model quantization accuracy, etc.).
[0035] The above provides an example of determining the number of GPUs required to meet the minimum storage resource requirement when the storage resource is a video memory resource.
[0036] Step 130: Determine a resource configuration strategy that matches the resource specifications based on the storage resource status of the nodes in the inference cluster, and configure storage resources in the inference cluster according to the resource configuration strategy to run the inference service.
[0037] In actual applications, inference clusters typically have storage resources of varying specifications (for example, GPUs vary in memory size). Therefore, the storage resources of nodes in an inference cluster can be managed as a unified resource pool. Storage resource status can be considered to refer to the detailed information about the storage resources provided by each node in the cluster, such as the total amount of available storage resources on each node and the model parameters of each storage device on each node. It can also refer to the usage status of each node's storage resources, such as whether they are idle or have sufficient storage capacity.
[0038] The resource allocation strategy can be thought of as the storage resource selection or configuration strategy determined within an inference cluster based on the resource specifications of the inference service and the storage resource status of each node in the cluster. For example, this could be determining whether the inference service should be deployed on a single node or distributed across multiple nodes. It could be prioritizing the allocation of lower-load nodes (matching the resource specifications) and their corresponding storage resources based on the node's storage resource load when configuring storage resources to balance the overall cluster load. It could also be prioritizing nodes with the same storage resource specifications when deploying the inference service on multiple nodes to balance the performance of the inference service running on different nodes.
[0039] Since the storage resources of the nodes in the inference cluster may have multiple resource specifications corresponding to the inference service at the same time, in some embodiments, different resource specifications need to be evaluated when determining the resource configuration strategy, such as the impact on the performance of the inference service, the feasibility on each node in the cluster, and the scalability when the demand for inference services increases.
[0040] After determining the resource allocation strategy, configure storage resources in the inference cluster according to the determined resource allocation strategy to run the inference service. Correctly configuring storage resources and ensuring that the inference service can obtain the storage resources it needs to run is the foundation for the inference service to operate.
[0041] The inference service resource configuration method provided in the embodiment of the present application obtains a model file of the inference service and parses the model file to determine the minimum storage resource requirement required to run the inference service, thereby evaluating the minimum storage resource required to run the inference service, providing a data basis for subsequent automated resource configuration, helping to improve the accuracy of resource configuration, and avoiding waste or insufficient storage resources. After determining the minimum storage resource requirement, the resource specification that meets the minimum storage resource requirement is determined. The resource specification characterizes the nodes required to run the inference service and the storage resource configuration corresponding to the nodes, providing guidance for the subsequent determination of resource configuration strategies and resource configuration, ensuring that resources can meet the operation requirements of the inference service. The storage resource status provides comprehensive information about the storage resources of each node for the resource allocation strategy, enabling the resource configuration strategy to be configured based on the storage resource status of the nodes in the inference cluster. The resource configuration strategy is used to indicate how to configure storage resources in the inference cluster to match the resource specifications. Storage resources are configured in the inference cluster according to the resource allocation strategy to run the inference service, thereby realizing automated determination of the storage resources required for the inference service and automated execution of storage resource configuration, reducing the risk and cost of errors in manual operations, avoiding waste or insufficient resources due to fragmentation, and improving storage resource utilization.
[0042] In some embodiments of the present application, determining a resource configuration strategy that matches the resource specification based on the storage resource status of the nodes in the inference cluster includes: Determining whether to deploy the inference service in a single node or in a multi-node distributed manner according to whether the available storage resources of a single node in the inference cluster match the resource specifications; Wherein, when performing single-node deployment, a target single node that meets the minimum storage resource requirement is automatically selected in the inference cluster, so as to configure storage resources corresponding to the resource specifications on the target single node; When performing multi-node distributed deployment, a plurality of target nodes that meet the distributed deployment conditions are automatically selected in the inference cluster to allocate storage resources corresponding to the resource specifications among the plurality of target nodes.
[0043] The available storage resources of a single node can be considered as the details of the storage resources that a single node can provide, such as the total amount of storage resources available on the node, the model parameters of each storage device on each node, etc. It can also refer to the usage status of the node's storage resources, such as whether it is idle or whether it has sufficient storage capacity.
[0044] It is understood that single-node deployment refers to the complete deployment of the inference service on a single node. It is suitable for scenarios where the resource requirements of the inference service can be met by a single node and distributed processing is not required. Multi-node distributed deployment refers to the distribution of inference service components or inference service instances across multiple nodes. It is used in scenarios where the resource requirements of the inference service cannot be met by a single node or distributed processing is required.
[0045] In an embodiment of the present application, whether the inference service is deployed as a single node or as a multi-node distributed deployment is determined based on whether the available storage resources of a single node in the inference cluster match the resource specifications. For example, the storage resources of each individual node in the inference cluster may be checked one by one to see if they match the resource specifications. If the available storage resources of a single node do match the resource specifications, a determination may be made to deploy the inference service as a single node. This may be achieved by evaluating the available storage resources of the node through a cluster management tool or by manually comparing the available storage resources.
[0046] If no single node in the inference cluster has storage resources that match the resource specifications, multiple target nodes that meet the distributed deployment requirements are selected and the storage resources corresponding to the resource specifications (i.e., the requirements for nodes and their corresponding storage resources as defined by the resource specifications) are allocated among these nodes. Node selection and resource classification can be accomplished, for example, through the scheduling function of a cluster management tool or through enumeration and traversal of each node.
[0047] Generally speaking, if the available storage resources of a single node match the resource specifications and a multi-node distributed deployment is not required, single-node deployment should be preferred. Specifically, single-node deployment simplifies the deployment of inference services and subsequent management processes. It eliminates the need for cross-node communication and data transfer, simplifies the architecture of inference services and storage resources, and offers lower inference latency than multi-node distributed deployment.
[0048] However, when the available storage resources of a single node do not match the resource specifications, multi-node distributed deployment can effectively disperse the node load, avoid single-node overload, and enhance the robustness of the inference cluster. When the inference service is deployed in a multi-node distributed manner, the scalability of the inference service can be effectively improved compared to a single-node deployment. When the load of the inference service, such as the number of visits or input data, increases, the storage resources configured for the inference service can be expanded across multiple nodes. In addition, when deploying inference services in a multi-node distributed manner, giving priority to nodes with the same storage resource specifications helps to balance the performance of the inference service running across different nodes.
[0049] The inference service resource configuration method provided in the embodiment of the present application determines whether the inference service will be deployed on a single node or in a multi-node distributed manner based on whether the available storage resources of a single node in the inference cluster match the resource specifications. When the available storage resources of a single node match the resource specifications, single-node deployment is preferred to simplify the inference service architecture and reduce latency. If the available storage resources of a single node do not meet the resource specifications, multi-node distributed deployment can effectively disperse the load and improve the scalability of the inference service. In single-node deployment, a node that meets the minimum storage requirements is automatically selected, and the corresponding storage resources are configured at the node. In multi-node distributed deployment, multiple nodes that meet the conditions are selected, and the required storage resources are allocated among these nodes. The inference solution is attempted to be deployed through both single-node deployment and multi-node distributed deployment to ensure that the inference service can be successfully deployed as much as possible, thereby optimizing the use of storage resources for the inference cluster.
[0050] In some embodiments of the present application, configuring storage resources in the inference cluster according to the resource configuration policy to run the inference service includes: On the target node or multiple target nodes, respectively configure storage resources corresponding to the resource specifications; After completing the storage resource configuration, starting the inference service instance corresponding to the configured storage resource on the configured node to run the inference service; In the case of multi-node distributed deployment, communication links are established between different inference service instances to collaboratively run the inference services.
[0051] It is understandable that a target single node refers to a single node selected in an inference cluster for running an inference service; multiple target nodes refer to multiple nodes selected in an inference cluster for distributed running of an inference service.
[0052] Configure storage resources corresponding to the resource specifications on the target node or nodes. Specifically, configure the storage resources for the target node or nodes based on the storage resource configuration required to run the inference service as indicated by the resource specifications. (If the inference service is deployed on a single node, configure the storage resources for the target node; if the inference service is deployed on multiple nodes, configure the storage resources for multiple target nodes.) Correctly configuring storage resources is essential for running the inference service instance and enabling the execution of some inference services.
[0053] An inference service instance (also known as a Worker instance) typically refers to an independent execution unit used to process a portion of the inference service. An inference service instance can be considered a component of the inference service. Specifically, it can be a process, a thread, or a container instance, depending on the architecture and implementation of the inference service (model).
[0054] Multiple inference service instances can work in parallel, improving the processing power and throughput of the inference service. Furthermore, by distributing tasks across multiple inference service instances, load balancing can be achieved, preventing overloading of a single inference service instance. In some embodiments, an inference task is split into multiple subtasks, each of which is processed by a separate inference service instance. By increasing or decreasing the number of inference service instances, the processing power of the inference service can be expanded to accommodate varying load requirements.
[0055] In the case of multi-node distributed deployment, communication links are established between different inference service instances, allowing the inference service instances in the distributed deployment to share data and status. For example, the communication between the inference service instances can also be simplified by using a service grid.
[0056] The inference service resource configuration method provided in the embodiment of the present application configures the storage resources of the target single node when the inference service is deployed on a single node, and configures the storage resources of multiple target nodes when the inference service is deployed on a multi-node distributed basis. By accurately configuring the storage resources, the inference service instance can obtain the required storage resources, thereby improving the utilization efficiency of the storage resources. After the storage resource configuration is completed, the inference service instance corresponding to the configured storage resources is started on the configured node to enable the normal operation of the inference service. When the inference service is deployed on a multi-node distributed basis, a communication link is established so that the inference service instances in the distributed deployment can work collaboratively. Through parallel processing between different inference service instances, the inference service can process more requests, which helps to improve the overall processing power and throughput of the inference service.
[0057] In some embodiments of the present application, determining a resource configuration strategy that matches the resource specification based on the storage resource status of the nodes in the inference cluster further includes: In a case where the pre-population phase and the decoding phase of the inference service are deployed separately, determining, according to the resource specifications, a first target node set for running the pre-population phase and a second target node set for running the decoding phase; An inference service instance of a pre-filling stage is deployed in the first target node set, and an inference service instance of a decoding stage is deployed in the second target node set, so as to deploy the pre-filling stage separately from the decoding stage.
[0058] It can be understood that PD separation (Prefill-Decoding Disaggregation) deployment divides the inference process into two main stages - prefill and decoding, and deploys these two stages on different hardware resources to improve resource utilization efficiency and inference performance.
[0059] The prefill phase is mainly used to process the input tokens of the input sequence and store the intermediate calculation results for reuse in the subsequent decoding phase. The decoding phase generates output tokens one by one based on the intermediate calculation results constructed in the prefill phase. The computational load in this phase is relatively light.
[0060] The first target node set refers to a set of nodes selected for running the pre-filling phase. The second target node set refers to a set of nodes selected for running the decoding phase.
[0061] In some embodiments, the number of inference service instances in the pre-filling phase and the number of inference service instances in the decoding phase of the inference service are both one, and the corresponding first target node set and second target node set each include only one node.
[0062] It is understood that if the first target node set includes only one node, the pre-population phase of the inference service adopts a single-node deployment; if the first target node set includes more than one node, the pre-population phase of the inference service adopts a multi-node distributed deployment; Similarly, if the second target node set includes only one node, the decoding phase of the inference service adopts a single-node deployment; if the second target node set includes more than one node, the decoding phase of the inference service adopts a multi-node distributed deployment; Among them, the relevant contents about single-node deployment and multi-node distributed deployment can be referred to the description in the aforementioned embodiments.
[0063] Based on resource specifications, a first target node set for running the pre-population phase and a second target node set for running the decoding phase are determined. Specifically, the pre-population phase of the inference service is compute-intensive, while the decoding phase of the inference service is memory-intensive. Therefore, high-specification storage resources (such as GPUs) are typically prioritized for deploying inference service instances in the pre-population phase, while low-specification storage resources (such as GPUs) are prioritized for deploying inference service instances in the decoding phase. For example, automated tools or scripts can be used to scan and obtain detailed parameter information about the storage resources of each node before determining the first and second target node sets. By assigning tasks to the most appropriate nodes, the pre-population and decoding phases of the inference service can be run on nodes that best meet their resource requirements.
[0064] After determining the first target node set and the second target node set, the inference service instance of the pre-filling stage is deployed on the first target node set, and the inference service instance of the decoding stage is deployed on the second target node set. Exemplarily, the automatic deployment and management of the inference service instance can be achieved through an orchestration tool, realizing the separate deployment of the pre-filling and decoding stages, so that each stage can run efficiently on appropriate resources.
[0065] Figure 3 This is a flowchart of the method for configuring inference service resources provided in some embodiments of the present application. Figure 3 As shown, in some embodiments, the reasoning service resource configuration method further includes: After starting the inference service deployment, determine whether the inference service requires PD separation deployment.
[0066] If the inference service is deployed with PD separation, determine whether there is a deployment template (also called a historical operation template) corresponding to the inference service. If a corresponding deployment template exists, use it to deploy the service. If not, set the number of inference service instances in the pre-population phase and the number of inference service instances in the decoding phase to one. Determine whether the (inference) cluster resources are sufficient. If the cluster resources are sufficient, the inference service can be run. If the cluster resources are insufficient, wait for idle resources.
[0067] If the inference service does not adopt PD separation deployment, obtain the resource specification list corresponding to the model inference service and obtain the cluster idle resource information; Determine whether the inference service is running on a single node. If so, the inference service can be run if the cluster resources are sufficient. If not, the inference service can be run on multiple nodes in a distributed manner if the cluster resources are sufficient.
[0068] The above process provides an example of the deployment process of the inference service, including the judgment of PD separation deployment, the evaluation of cluster resources, and the final deployment decision.
[0069] The inference service resource configuration method provided by the embodiment of the present application, when the pre-filling stage and the decoding stage of the inference service are deployed separately, determines the first target node set for running the pre-filling stage and the second target node set for running the decoding stage according to the resource specifications, prioritizes using high-specification storage resources for deploying inference service instances in the pre-filling stage, and uses low-specification storage resources for deploying inference service instances in the decoding stage, thereby matching resource specifications with storage devices corresponding to each node, allowing both stages of the inference service to run in a suitable environment, deploying the inference service instance of the pre-filling stage in the first target node set, and deploying the inference service instance of the decoding stage in the second target node set, thereby achieving separate deployment of the pre-filling stage and the decoding stage, optimizing the utilization of storage resources, and helping to improve the operating efficiency and reliability of the inference service.
[0070] In some embodiments of the present application, the method further comprises: In a case where there is a historical running template of the inference service, determining a first sub-resource specification and a second sub-resource specification based on the historical running template; Determining, based on the historical operation template, the number of inference service instances of the inference service in the pre-filling phase and the number of inference service instances of the inference service in the decoding phase; configuring storage resources in the pre-population phase based on the first sub-resource specification to deploy an inference service instance of the inference service in the pre-population phase; The storage resources of the decoding stage are configured based on the second sub-resource specification to deploy the inference service instance of the inference service in the decoding stage.
[0071] The historical operation template can be considered as a record of the performance and resource usage of the inference service during past deployment and operation of the inference cluster. In the embodiment of the present application, it can be used to guide the current resource configuration, which includes the first sub-resource specification, the second sub-resource specification, the number of inference service instances in the pre-population phase of the inference service, and the number of inference service instances in the decoding phase of the inference service.
[0072] It is understood that the first sub-resource specification is used to represent the nodes required to run the pre-population phase of the inference service and the storage resource configuration corresponding to the nodes. Specifically, the first sub-resource specification is used to guide how to configure storage resources in the inference cluster to meet the storage requirements of the pre-population phase of the inference service.
[0073] The second sub-resource specification is used to represent the nodes required to run the decoding phase of the inference service and the storage resource configuration corresponding to the nodes. Specifically, the second sub-resource specification is used to guide how to configure storage resources in the inference cluster to meet the storage requirements of the decoding phase of the inference service.
[0074] In some embodiments, the first sub-resource specification or the second sub-resource specification may also include requirements for computing resources, storage resources, memory, network bandwidth, etc.
[0075] For an inference service, there may be multiple first sub-resource specifications or second sub-resource specifications. In some embodiments, the first sub-resource specification is equal to the second sub-resource specification and each needs to meet the minimum storage resource requirements required to run a complete inference service.
[0076] In some embodiments, the inference service instance in the pre-filling phase or the inference service instance in the decoding phase may be dynamically adjusted to respond to real-time load changes.
[0077] In some embodiments, the method further comprises: Obtaining an identifier of the target reasoning service; generating a separate operation template for the target inference service based on the number of task execution services of the inference service in the pre-population phase, the number of task execution services of the inference service in the decoding phase, the first resource specification, and the second resource specification; A mapping relationship is established between the identifier and the separate operation template.
[0078] The inference service resource configuration method provided in the embodiment of the present application determines the first sub-resource specification and the second sub-resource specification based on the historical operation template when there is a historical operation template of the inference service. The historical data provides an accurate reference for the actual runtime resource usage, which helps to more accurately determine the resource requirements of the different stages of the inference service, thereby performing more reasonable resource configuration; based on the historical operation template, the number of inference service instances of the inference service in the pre-filling stage and the number of inference service instances of the inference service in the decoding stage are determined. The number of instances directly affects the processing capacity and resource consumption of the inference service. The number of inference service instances is determined based on historical data, which can ensure that the service avoids resource waste while meeting performance requirements. The historical operation template is used to realize automatic configuration of storage resources and deployment instances, reduce manual intervention in the resource configuration process, and improve resource configuration efficiency and accuracy.
[0079] In some embodiments of the present application, the method further comprises: When the inference cluster contains heterogeneous nodes, based on the node performance parameters of each node, a set of candidate nodes that meet the resource specifications are screened; Based on the operation requirements of the reasoning service, determining a target node that can host the reasoning service from the set of candidate nodes; According to the resource configuration policy, storage resources are configured on the target node, and a corresponding inference service instance is started.
[0080] It is understandable that in an inference cluster, nodes may have different hardware configurations, such as different models of GPUs or different capacities of memory, that is, heterogeneous nodes. The heterogeneity of nodes needs to be considered when allocating resources to ensure that resource configuration matches node capabilities.
[0081] Node performance parameters refer to parameters that describe the performance characteristics of a node, such as computing power, memory capacity, and storage bandwidth. They can be used to evaluate whether a node meets the resource requirements of a specific inference service.
[0082] The candidate node set can be considered as a set of nodes that meet the resource specification requirements. These nodes are potential candidates for running the reasoning service. In an embodiment of the present application, the most suitable node is selected from the candidate node set based on the running requirements of the reasoning service to deploy the reasoning service. For example, on the premise that the candidate node is determined to be able to carry the reasoning service, it can be preferred to select a node with a higher storage resource specification as the target node, or it can be preferred to select a node with the same storage device specification as the target node. This application does not specifically limit how to determine the target node from the candidate node set.
[0083] After determining the target node, configure storage resources on the target node according to the resource configuration policy to run the inference service. Properly configuring storage resources and ensuring that the inference service can obtain the storage resources it needs to run is the basis for the inference service to run.
[0084] The inference service resource configuration method provided in the embodiment of the present application, when the inference cluster contains heterogeneous nodes, screens a set of candidate nodes that meet the resource specifications based on the node performance parameters of each node, ensures that the nodes in the candidate node set can meet the resource requirements of the inference service, and determines the target node that can host the inference service from the candidate node set based on the operation requirements of the inference service. The suitable node can provide the required computing and storage capabilities. The most suitable node is selected for deployment, which can optimize the utilization of storage resources. Storage resources are configured on the target node and the inference service instance is started, ensuring that the service can obtain the required resources and run smoothly, thereby improving the reliability and performance of the service.
[0085] In some embodiments of the present application, the method further comprises: During the operation of the inference service, collecting inference performance parameters of the inference service; Based on the comparison result of the inference performance parameter and the preset threshold, the expansion operation of the storage resource is automatically triggered to optimize the inference performance of the inference service.
[0086] It is understood that inference performance parameters refer to key indicators used to measure inference performance during the operation of the inference service, such as latency, throughput, and accuracy. Preset thresholds refer to target values or upper and lower thresholds set for inference performance parameters, which are used for comparison with actual inference performance parameters to determine whether resource adjustments are needed.
[0087] In an embodiment of the present application, during the operation of the inference service, the inference performance parameters of the inference service are collected. Exemplarily, the inference performance parameters of the inference service can be monitored and recorded in real time through a monitoring tool. After the inference performance parameters are collected, the collected performance parameters are compared with a preset threshold, and based on the comparison result, it is determined whether to trigger the expansion operation of the storage resource. Through comparison, insufficient or excessive performance can be identified, thereby determining whether expansion is needed. For example, when the inference delay exceeds the preset threshold, or when the service throughput exceeds the preset threshold, it can be determined that the expansion operation of the storage resource is triggered. Automatic expansion can quickly respond to load changes and avoid delays and inaccuracies caused by manual adjustments.
[0088] The inference service resource configuration method provided in the embodiment of the present application collects the inference performance parameters of the inference service during the operation of the inference service, and timely understands the operation status of the inference service by real-time monitoring of the inference performance parameters, providing a basis for resource adjustment. Based on the comparison result of the inference performance parameters with the preset threshold, by comparing with the preset threshold, it is accurately judged whether resource adjustment is needed, and the expansion operation of storage resources is automatically triggered, which reduces manual intervention in the expansion process and can make resource allocation more flexible and dynamic, thereby flexibly optimizing the inference performance of the inference service and enabling the inference service to adapt to changing load requirements.
[0089] In some embodiments of the present application, the inference performance parameter includes at least the first word time and the per-word time; and automatically triggering the storage resource expansion operation based on the comparison result of the inference performance parameter with a preset threshold value includes: In a case where the first word meta-time exceeds the first word meta-time standard value, increasing the number of inference service instances deployed by the inference service in the pre-population phase to expand the capacity of the inference service; In a case where the per-word time exceeds a standard per-word time value, the number of inference service instances deployed by the inference service in the decoding phase is increased to expand the capacity of the inference service.
[0090] It's important to note that the Time to First Token (TTFT), also known as the maximum latency for generating the first token, refers to the time it takes for the inference service to return the first token response. This metric measures the responsiveness of the model inference service. A shorter TTFT indicates that the inference service can respond quickly to requests.
[0091] Time Per Output Token (TPOT), or the average time it takes to generate each output token. This metric measures the speed at which the model generates output, representing the average time interval between the generation of one token and the generation of the next. A lower TPOT value indicates that the model generates output quickly.
[0092] It is understandable that TTFT primarily measures the initial response speed of the inference service to requests, while TPOT measures the continuous efficiency of the inference service in the process of generating output. The combination of the two can comprehensively evaluate the inference performance of the model.
[0093] It is understandable that by adding more inference service instances, more inference requests can be processed in parallel and throughput, that is, the number of inference tasks completed per unit time, can be increased. As the number of instances increases, the inference service can respond to new inference requests more quickly, reducing user waiting time. The pre-filling phase is mainly used to process input data, perform preliminary calculations, and prepare for generating the output sequence. The time to first token (TTFT) is a key performance indicator of the pre-filling phase, which is used to measure the speed of generating the first token. The decoding phase is mainly used to continue generating the rest of the sequence based on the output of the pre-filling phase. The time per token (TPOT) is a key performance indicator of the decoding phase, which is used to measure the average speed of generating each subsequent token.
[0094] Here is a specific example: Monitor the real-time load of the inference service instance during the pre-population phase and the decoding phase to obtain the current first-word time. and the current per-word time , Respectively The time standard value of the first word To compare, and the standard value of time per word Compare with the standard value and expand according to the following rules: exist In the case of ,increase the number of inference service instances in the pre-population phase of the inference service; exist In the case of ,inference service, the number of inference service instances in the decoding phase is increased.
[0095] The inference service resource configuration method provided by the embodiment of the present application mainly includes the inference performance parameters of the first word element time and the per-word element time. In the first word element time, in the pre-filling stage of the inference service, the input data is mainly processed and preliminary calculations are performed. By increasing the number of inference service instances in the pre-filling stage, the speed of generating the first word element can be increased, thereby shortening the response time; in the decoding stage of the inference service, the remaining part of the sequence is mainly generated based on the output of the pre-filling stage. By increasing the number of inference service instances in the decoding stage, the average speed of generating each subsequent word element can be increased, thereby improving the overall inference efficiency of the inference service.
[0096] In some embodiments of the present application, when there are multiple reasoning services, the method further includes: Obtaining model files of the multiple reasoning services, and parsing the model files respectively to determine the minimum storage resource requirements required to run the multiple reasoning services respectively; Based on the storage resource status of the nodes in the inference cluster, determine one or more candidate resource specifications corresponding to any inference service; Combining and generating a plurality of global resource configuration strategies based on the candidate resource specifications corresponding to the plurality of inference services; Selecting a target global resource configuration strategy that can run the most inference services from the multiple global resource configuration strategies; Storage resources are configured according to the target global resource configuration policy to run the multiple inference services.
[0097] In an embodiment of the present application, multiple different inference services need to be run in the same inference cluster.
[0098] The global resource allocation policy is a strategy for selecting or allocating storage resources within an inference cluster, based on the resource specifications of multiple inference services and the storage resource status of each node in the cluster. It can be thought of as a resource allocation plan generated based on the resource requirements of multiple inference services to optimize overall resource utilization and enable the operation of as many inference services as possible.
[0099] In some embodiments, obtaining the model files of the multiple reasoning services and parsing the model files respectively to determine the minimum storage resource requirements required to run the multiple reasoning services respectively include: For each of the multiple inference services, a model file of the inference service is obtained, and the model file is parsed to determine a minimum storage resource requirement required to run the inference service.
[0100] In some embodiments, determining one or more candidate resource specifications corresponding to any inference service based on the storage resource status of nodes in the inference cluster includes: For each of the multiple inference services, a resource specification that satisfies the minimum storage resource requirement is determined, where the resource specification is used to characterize nodes required to run the inference service and storage resource configurations corresponding to the nodes.
[0101] It is understood that for each of the multiple inference services, a minimum storage resource requirement is determined, and one or more corresponding candidate resource specifications are determined. Candidate resource specifications refer to one or more possible resource configuration options determined for each inference service, which need to be combined with the candidate resource specifications corresponding to other inference services to generate a global resource configuration strategy.
[0102] Ideally, the global resource allocation policy should include the candidate resource specifications for each inference service, meaning that each inference service can be run based on the global resource allocation policy. However, due to factors such as the limited storage resources of nodes in the inference cluster, the global resource allocation policy typically cannot include the candidate resource specifications for each inference service. Therefore, a target global resource allocation policy that can run the most inference services is used to run as many inference services as possible.
[0103] It is understandable that for each reasoning service included in the global resource configuration strategy and the candidate resource specifications corresponding to each reasoning service, the reasoning service resource configuration can be independently performed based on the resource specifications.
[0104] In some embodiments, configuring storage resources according to the target global resource configuration policy to run the multiple inference services includes: For each inference service with corresponding resource specifications in the global resource configuration policy, a resource configuration policy that matches the resource specifications is determined based on the storage resource status of the nodes in the inference cluster, and storage resources are configured in the inference cluster according to the resource configuration policy to run the inference service.
[0105] Figure 4 This is a flowchart of the method for configuring inference service resources provided in some embodiments of the present application. Figure 4 As shown, in some embodiments, the reasoning service resource configuration method further includes: When there are multiple inference services to be scheduled, obtain the cluster's idle GPU topology; Based on the creation time of each inference service, sort the inference services in ascending order to obtain a list of inference services to be scheduled; Traverse each inference service in the list of inference services to be scheduled to determine whether all inference services have been traversed. If completed, record the current inference service deployment plan (corresponding to the global resource allocation strategy mentioned above). If not completed, try to deploy the current inference service; Get the resource specification list of the current inference service (consisting of one or more resource specifications corresponding to the inference service), traverse each resource specification in the resource specification list and apply for cluster resources (cluster resource application can be considered as the process of determining whether cluster resources meet the resource specifications); If the cluster resource application corresponding to the current resource specification is successful, the cluster resource application is performed based on the next resource specification of the current resource specification, and the next inference service of the current inference service is used as the current inference service. The system returns to determine whether all inference services have been traversed. If completed, the current inference service deployment plan is recorded (corresponding to the global resource configuration strategy mentioned above). If not completed, the system attempts to deploy the current inference service and continue execution. If the cluster resource application corresponding to the current resource specification is unsuccessful, determine whether each resource specification in the resource specification list has been traversed. If so, if there is no successful cluster resource application corresponding to the resource specification, terminate the current inference service deployment plan and record it. If the traversal is not completed, apply for cluster resources based on the next resource specification of the current resource specification. Sort the inference service deployment plans and obtain the plan with the largest number of deployed inference services.
[0106] This flowchart provides an example of how to determine the inference service deployment plan (corresponding to the global resource allocation strategy described above) to request resources for multiple inference services and ensure that as many services as possible can run. It includes steps such as service sorting, resource application, deployment attempts, and the final resource allocation decision.
[0107] The inference service resource configuration method provided in the embodiment of the present application obtains and parses model files of multiple inference services to determine the minimum storage resource requirements required for the operation of each service, and determines one or more candidate resource specifications for each inference service based on the storage resource status of each node in the inference cluster. By analyzing the resource requirements of each inference service, the pertinence and accuracy of resource configuration are improved, which helps to improve resource utilization efficiency. Based on the candidate resource specifications corresponding to each of the multiple inference services, a variety of global resource configuration strategies are combined to generate, and the target global resource configuration strategy that can run the most inference services is further screened out from the multiple global resource configuration strategies. The purpose of the screening process is to ensure the global optimality of resource configuration, give priority to the configuration scheme that can support the most services, thereby improving the service capabilities of the cluster and ensuring that as many inference services as possible can run smoothly under resource constraints.
[0108] In some embodiments of the present application, the generation of multiple global resource configuration strategies based on the candidate resource specifications corresponding to the multiple reasoning services includes: For any candidate resource specification of the inference service, determining whether the available storage resources of the nodes in the inference cluster meet the current candidate resource specification; If there is a node that meets the current candidate resource specifications, add the current candidate resource specifications to the current global resource configuration policy; The next inference service of the current inference service is taken as the new current inference service, and the candidate resource specifications for any of the inference services are returned to. The step of determining whether the available storage resources of the nodes in the inference cluster meet the current candidate resource specifications is continued until the current global resource configuration strategy includes a candidate resource specification corresponding to each of the multiple inference services. For each inference service included in the strategy, a matching resource configuration strategy is determined based on the storage resource status of the nodes in the cluster, and storage resources are configured in the cluster to run the service.
[0109] In one embodiment of the present application, determining whether the available storage resources of each node in the cluster meet the candidate resource specifications for the currently targeted inference service can, for example, be performed by using an automated tool to scan the cluster resources and automatically match them with the candidate resource specifications. If the available storage resources of the node meet the candidate resource specifications, the candidate resource specifications are added to the current global resource configuration policy, and the next inference service is selected as the current service, repeating the resource matching and configuration steps.
[0110] Here is a specific example: There are currently M inference services. , and set up each inference service There are N corresponding resource specifications (In actual applications, the resource specifications corresponding to each inference service need to be confirmed based on the actual scenario); For any inference service Candidate resource specifications , to determine whether the available storage resources of the nodes in the inference cluster meet ,in, Representation Reasoning Service The corresponding i-th resource specification; If satisfied, the candidate resource specifications Join the current global resource configuration strategy , and returns to the candidate resource specifications for any of the inference services, and continues to execute the step of determining whether the available storage resources of the nodes in the inference cluster meet the current candidate resource specifications until the current global resource configuration strategy Include A candidate resource specification corresponding to each inference service.
[0111] That is, inference service Candidate resource specifications Join the current global resource configuration strategy After that, for the inference service Candidate resource specifications Determine whether the available storage resources of the nodes in the inference cluster meet the requirements. , if the candidate resource specifications are met Join the current global resource configuration strategy , and repeat the above steps until the current global resource configuration policy Contains a candidate resource specification corresponding to each inference service.
[0112] The inference service resource configuration method provided in the embodiment of the present application determines, for each candidate resource specification of an inference service, whether the available storage resources of the nodes in the cluster meet these candidate resource specifications. If there is a node that meets the current candidate resource specification, the current candidate resource specification will be added to the current global resource configuration strategy. Subsequently, the next inference service will be processed, and this process will be repeated until the candidate resource specifications of all inference services are considered and added to the global resource configuration strategy. By determining the most appropriate resource specification for each inference service, as many inference services as possible can be run, which helps to improve the utilization of storage resources.
[0113] In some embodiments of the present application, the method further comprises: If there is no node that meets the current candidate resource specifications, the current global resource configuration policy is used as a global resource configuration policy, and a new global resource configuration policy is generated based on the current global resource configuration policy and used as the current global resource configuration policy; The next inference service of the current inference service is used as the new current inference service, and the step of returning to the candidate resource specifications for any of the inference services and determining whether the available storage resources of the nodes in the inference cluster meet the current candidate resource specifications is continued.
[0114] In an embodiment of the present application, it is determined whether the available storage resources of each node in the cluster meet the candidate resource specifications of the currently targeted inference service. For example, an automated tool can be used to scan the cluster resources and automatically match them with the candidate resource specifications. If the available storage resources of the node do not meet the candidate resource specifications, the generation of the current global resource configuration policy is stopped (that is, the current global resource configuration policy has been used as a global resource configuration policy for subsequent screening). A new global resource configuration policy is generated based on the current global resource configuration policy and used as the current global resource configuration policy. Specifically, a new global resource configuration policy can be generated by copying the current global resource configuration policy. The next inference service is used as the current service, and the resource matching and configuration steps are repeated.
[0115] Here is a specific example: Still assume that there are M inference services in total , and set up each inference service There are N corresponding resource specifications (In actual applications, the resource specifications corresponding to each inference service need to be confirmed based on the actual scenario); For any inference service Candidate resource specifications , to determine whether the available storage resources of the nodes in the inference cluster meet ,in, Representation Reasoning Service The corresponding i-th resource specification; If not satisfied, stop generating the current global resource configuration strategy (That is, at this time No longer generated), based on the current global resource allocation strategy Generate a new global resource configuration policy And serve as the current global resource allocation strategy; The current inference service Next inference service As the new current inference service, and returning to the candidate resource specifications for any of the inference services, the step of determining whether the available storage resources of the nodes in the inference cluster meet the current candidate resource specifications continues to be executed.
[0116] That is to say, Determine whether the available storage resources of the nodes in the inference cluster meet the requirements. , and determine the subsequent steps based on whether they are met.
[0117] The inference service resource configuration method provided in the embodiment of the present application determines, for each candidate resource specification of an inference service, whether the available storage resources of the nodes in the cluster meet these candidate resource specifications. If there is no node that meets the current candidate resource specifications, the generation of the current global resource configuration policy is stopped, and a new global resource configuration policy is generated based on the current global resource configuration policy, which is continued to be executed as the current policy, and the next inference service is used as the current service, and the resource matching and configuration steps are repeated. This helps to achieve the goal of running as many inference services as possible even when resources are limited, thereby improving resource utilization efficiency and the service capabilities of the system.
[0118] In some embodiments of the present application, the method further comprises: If there is a node that meets the current candidate resource specifications, a new global resource configuration policy is generated based on the current global resource configuration policy and used as the current global resource configuration policy; Determine the next candidate resource specification of the current candidate resource specification; Determining whether the available storage resources of the nodes in the inference cluster meet the next candidate resource specifications; If there is a node that meets the next candidate resource specification, add the next candidate resource specification to the current global resource configuration strategy; If there is no node that meets the next candidate resource specification, stop generating the current global resource configuration strategy; Generate a new global resource configuration policy based on the current global resource configuration policy, and use the next inference service of the current inference service as the new current inference service, and return to the candidate resource specifications for any of the inference services, and continue to execute the step of determining whether the available storage resources of the nodes in the inference cluster meet the current candidate resource specifications.
[0119] It can be understood that when there are nodes that meet the current candidate resource specifications, a new global resource configuration policy is generated based on the current global resource configuration policy and used as the current global resource configuration policy, and the next candidate resource specification of the current candidate resource specification is judged to determine whether there are nodes that meet the next candidate resource specification. If there are nodes that meet the next candidate resource specification, the next candidate resource specification is added to the new current global resource configuration policy. If there are no nodes that meet the next candidate resource specification, the generation of the new current global resource configuration policy is stopped (that is, it has been used as a global resource configuration policy for subsequent screening).
[0120] Here is a specific example: Still assume that there are M inference services in total , and set up each inference service There are N corresponding resource specifications (In actual applications, the resource specifications corresponding to each inference service need to be confirmed based on the actual scenario); For any inference service Candidate resource specifications , to determine whether the available storage resources of the nodes in the inference cluster meet ,in, Representation Reasoning Service The corresponding i-th resource specification; If satisfied, based on the current global resource allocation strategy Generate a new global resource configuration policy And serve as the current global resource allocation strategy; Determine the current candidate resource specifications Next candidate resource specification ; Determine whether the available storage resources of the nodes in the inference cluster meet the next candidate resource specifications ; If there is a resource that meets the next candidate specification The node will be the next candidate resource specification Join the current global resource configuration strategy ; If there is no next candidate resource that meets the requirements Stop generating the current global resource configuration policy for the node (That is, at this time No longer generated); Based on the current global resource allocation strategy Generate a new global resource configuration policy , and the current inference service Next inference service As the new current inference service, and returning to the candidate resource specifications for any of the inference services, the step of determining whether the available storage resources of the nodes in the inference cluster meet the current candidate resource specifications continues to be executed.
[0121] That is to say, Determine whether the available storage resources of the nodes in the inference cluster meet the requirements. , and determine the subsequent steps based on whether they are met.
[0122] The inference service resource configuration method provided by the embodiment of the present application generates a new global resource configuration policy based on the current global resource configuration policy when a node that meets the current candidate resource specification exists, and uses it as the current policy. In addition, the next candidate resource specification of the current candidate resource specification is determined, and it is judged whether there is a node that meets the next candidate resource specification in the cluster. If there is a node that meets the next candidate resource specification, the candidate resource specification will be added to the current global resource configuration policy; if not, the generation of the current global resource configuration policy is stopped, and a new global resource configuration policy is generated based on the current policy, and the processing of the next inference service continues. Through this step-by-step iterative method, the judgment of each candidate resource specification of each inference service can be realized, which helps to realize the operation of as many inference services as possible, thereby improving resource utilization efficiency.
[0123] In some embodiments of the present application, the method further comprises: If there is no node that meets the current candidate resource specifications, stop generating the current global resource configuration policy; Generate a new global resource configuration policy based on the current global resource configuration policy and use it as the current global resource configuration policy; The next inference service of the current inference service is used as the new current inference service, and the step of returning to the candidate resource specifications for any of the inference services and determining whether the available storage resources of the nodes in the inference cluster meet the current candidate resource specifications is continued.
[0124] Here is a specific example: Still assume that there are M inference services in total , and set up each inference service There are N corresponding resource specifications (In actual applications, the resource specifications corresponding to each inference service need to be confirmed based on the actual scenario); For any inference service Candidate resource specifications , to determine whether the available storage resources of the nodes in the inference cluster meet ; If not satisfied, stop generating the current global resource configuration strategy (That is, at this time No longer generated); Based on the current global resource allocation strategy Generate a new global resource configuration policy And serve as the current global resource allocation strategy; The current inference service Next inference service As the new current inference service, and returning to the candidate resource specifications for any of the inference services, the step of determining whether the available storage resources of the nodes in the inference cluster meet the current candidate resource specifications continues to be executed.
[0125] That is to say, Determine whether the available storage resources of the nodes in the inference cluster meet the requirements. , and determine the subsequent steps based on whether they are met.
[0126] The reasoning service resource configuration method provided in the embodiment of the present application stops generating the current global resource configuration policy when it is determined that there is no node that meets the current candidate resource specifications. Subsequently, a new global resource configuration policy is generated based on the existing global resource configuration policy and set as the current policy. Next, the next reasoning service is used as the new current reasoning service, and the process returns to the resource specification judgment step to continue the resource matching process. By iteratively processing each reasoning service, the most appropriate resource configuration is determined for each reasoning service, which helps to achieve the maximum possible number of running reasoning services, thereby improving resource utilization efficiency.
[0127] Through the description of the above implementation methods, those skilled in the art can clearly understand that the method according to the above embodiment can be implemented by means of software plus the necessary general hardware platform, and of course it can also be implemented by hardware, but in many cases the former is a better implementation method.
[0128] Figure 5 Schematic diagram of a computer program product provided in some embodiments of the present application. Figure 5 As shown, the present application further provides a computer program product 500 including: The first processing module 501 is configured to obtain a model file of an inference service and parse the model file to determine the minimum storage resource requirement for running the inference service. A second processing module 502 is configured to determine a resource specification that satisfies the minimum storage resource requirement, where the resource specification is used to represent the nodes required to run the inference service and the storage resource configuration corresponding to the nodes; The third processing module 503 is used to determine a resource configuration strategy that matches the resource specifications according to the storage resource status of the nodes in the inference cluster, and configure storage resources in the inference cluster according to the resource configuration strategy to run the inference service.
[0129] In some embodiments, the third processing module 503 is configured to: Determining whether to deploy the inference service in a single node or in a multi-node distributed manner according to whether the available storage resources of a single node in the inference cluster match the resource specifications; Wherein, when performing single-node deployment, a target single node that meets the minimum storage resource requirement is automatically selected in the inference cluster, so as to configure storage resources corresponding to the resource specifications on the target single node; When performing multi-node distributed deployment, a plurality of target nodes that meet the distributed deployment conditions are automatically selected in the inference cluster to allocate storage resources corresponding to the resource specifications among the plurality of target nodes.
[0130] In some embodiments, the third processing module 503 is configured to: On the target node or multiple target nodes, respectively configure storage resources corresponding to the resource specifications; After completing the storage resource configuration, starting the inference service instance corresponding to the configured storage resource on the configured node to run the inference service; In the case of multi-node distributed deployment, communication links are established between different inference service instances to collaboratively run the inference services.
[0131] In some embodiments, the third processing module 503 is further configured to: In a case where the pre-population phase and the decoding phase of the inference service are deployed separately, determining, according to the resource specifications, a first target node set for running the pre-population phase and a second target node set for running the decoding phase; An inference service instance of a pre-filling stage is deployed in the first target node set, and an inference service instance of a decoding stage is deployed in the second target node set, so as to deploy the pre-filling stage separately from the decoding stage.
[0132] In some embodiments, the computer program product 500 further includes a fourth processing module, configured to: In a case where there is a historical running template of the inference service, determining a first sub-resource specification and a second sub-resource specification based on the historical running template; Determining, based on the historical operation template, the number of inference service instances of the inference service in the pre-filling phase and the number of inference service instances of the inference service in the decoding phase; configuring storage resources in the pre-population phase based on the first sub-resource specification to deploy an inference service instance of the inference service in the pre-population phase; The storage resources of the decoding stage are configured based on the second sub-resource specification to deploy the inference service instance of the inference service in the decoding stage.
[0133] In some embodiments, the computer program product 500 further includes a fifth processing module, configured to: When the inference cluster contains heterogeneous nodes, based on the node performance parameters of each node, a set of candidate nodes that meet the resource specifications are screened; Based on the operation requirements of the reasoning service, determining a target node that can host the reasoning service from the set of candidate nodes; According to the resource configuration policy, storage resources are configured on the target node, and a corresponding inference service instance is started.
[0134] In some embodiments, the computer program product 500 further includes a sixth processing module, configured to: During the operation of the inference service, collecting inference performance parameters of the inference service; Based on the comparison result of the inference performance parameter and the preset threshold, the expansion operation of the storage resource is automatically triggered to optimize the inference performance of the inference service.
[0135] In some embodiments, the inference performance parameter includes at least the first word time and the per-word time; and automatically triggering the storage resource expansion operation based on the comparison result of the inference performance parameter with a preset threshold value includes: In a case where the first word meta-time exceeds the first word meta-time standard value, increasing the number of inference service instances deployed by the inference service in the pre-population phase to expand the capacity of the inference service; In a case where the per-word time exceeds a standard per-word time value, the number of inference service instances deployed by the inference service in the decoding phase is increased to expand the capacity of the inference service.
[0136] In some embodiments, when there are multiple inference services, the computer program product 500 further includes a seventh processing module, the seventh processing module being configured to: Reading model files of the multiple reasoning services and parsing the model files respectively to determine the minimum storage resource requirements required to run the multiple reasoning services respectively; Based on the storage resource status of the nodes in the inference cluster, determine one or more candidate resource specifications corresponding to any inference service; Combining and generating a plurality of global resource configuration strategies based on the candidate resource specifications corresponding to the plurality of inference services; Selecting a target global resource configuration strategy that can run the most inference services from the multiple global resource configuration strategies; Storage resources are configured according to the target global resource configuration policy to run the multiple inference services.
[0137] In some embodiments, the generating of multiple global resource configuration strategies based on the candidate resource specifications corresponding to the multiple reasoning services includes: For any candidate resource specification of the inference service, determining whether the available storage resources of the nodes in the inference cluster meet the current candidate resource specification; If there is a node that meets the current candidate resource specifications, add the current candidate resource specifications to the current global resource configuration policy; The next inference service of the current inference service is used as the new current inference service, and the candidate resource specifications for any of the inference services are returned to, and the step of determining whether the available storage resources of the nodes in the inference cluster meet the current candidate resource specifications is continued until the current global resource configuration strategy includes a candidate resource specification corresponding to each of the multiple inference services.
[0138] In some embodiments, the computer program product 500 further includes an eighth processing module, the eighth processing module being configured to: If there is a node that meets the current candidate resource specifications, a new global resource configuration policy is generated based on the current global resource configuration policy and used as the current global resource configuration policy; Determine the next candidate resource specification of the current candidate resource specification; Determining whether the available storage resources of the nodes in the inference cluster meet the next candidate resource specifications; If there is a node that meets the next candidate resource specification, add the next candidate resource specification to the current global resource configuration strategy; If there is no node that meets the next candidate resource specification, stop generating the current global resource configuration strategy; Generate a new global resource configuration policy based on the current global resource configuration policy, and use the next inference service of the current inference service as the new current inference service, and return to the candidate resource specifications for any of the inference services, and continue to execute the step of determining whether the available storage resources of the nodes in the inference cluster meet the current candidate resource specifications.
[0139] In some embodiments, the eighth processing module is further configured to: If there is no node that meets the current candidate resource specifications, stop generating the current global resource configuration policy; Generate a new global resource configuration policy based on the current global resource configuration policy and use it as the current global resource configuration policy; The next inference service of the current inference service is used as the new current inference service, and the step of returning to the candidate resource specifications for any of the inference services and determining whether the available storage resources of the nodes in the inference cluster meet the current candidate resource specifications is continued.
[0140] In some embodiments, the eighth processing module is further configured to: If there is no node that meets the current candidate resource specifications, stop generating the current global resource configuration policy, generate a new global resource configuration policy based on the current global resource configuration policy and use it as the current global resource configuration policy; The next inference service of the current inference service is used as the new current inference service, and the step of returning to the candidate resource specifications for any of the inference services and determining whether the available storage resources of the nodes in the inference cluster meet the current candidate resource specifications is continued.
[0141] For the description of the features in the embodiment corresponding to the computer program product, please refer to the relevant description of the embodiment corresponding to the reasoning service resource configuration method, which will not be repeated here.
[0142] An embodiment of the present application further provides an electronic device, comprising a memory and a processor, wherein the memory stores a computer program, and the processor is configured to run the computer program to execute the steps in any of the above-mentioned inference service resource configuration method embodiments.
[0143] An embodiment of the present application further provides a non-volatile computer-readable storage medium, in which a computer program is stored, wherein the computer program is configured to execute the steps of any of the above-mentioned inference service resource configuration method embodiments when running.
[0144] In an exemplary embodiment, the non-volatile computer-readable storage medium may include, but is not limited to, various media that can store computer programs, such as a USB flash drive, a read-only memory (ROM), a random access memory (RAM), a mobile hard disk, a magnetic disk, or an optical disk.
[0145] An embodiment of the present application further provides another computer program product, which includes a computer program. When the computer program is executed by a processor, the steps of any of the above-mentioned inference service resource configuration method embodiments are implemented.
[0146] Professionals may further appreciate that the units and algorithm steps of each example described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, computer software, or a combination of the two. In order to clearly illustrate the interchangeability of hardware and software, the above description has generally described the components and steps of each example according to their functions. Whether these functions are performed in hardware or software depends on the specific application and design constraints of the technical solution. Professionals and technicians may use different methods to implement the described functions for each specific application, but such implementation should not be considered beyond the scope of this application.
[0147] The above is a detailed introduction to the reasoning service resource configuration method, electronic device and readable storage medium provided by this application. This article uses specific examples to illustrate the principles and implementation methods of this application. The description of the above embodiments is only applicable to help understand the method and core ideas of this application. It should be pointed out that for ordinary technicians in this technical field, without departing from the principles of this application, several improvements and modifications can be made to this application, and these improvements and modifications also fall within the scope of protection of the claims of this application.
Claims
1. A method for configuring inference service resources, characterized in that: include: Obtaining a model file for an inference service, and parsing the model file to determine a minimum storage resource requirement for running the inference service; Determining a resource specification that meets the minimum storage resource requirement, where the resource specification is used to characterize the nodes required to run the inference service and the storage resource configuration corresponding to the nodes; According to the storage resource status of the nodes in the inference cluster, a resource configuration policy that matches the resource specification is determined, and storage resources are configured in the inference cluster according to the resource configuration policy to run the inference service.
2. The method for configuring inference service resources according to claim 1, wherein: The determining of a resource configuration strategy that matches the resource specifications based on the storage resource status of the nodes in the inference cluster includes: Determining whether to deploy the inference service in a single node or in a multi-node distributed manner according to whether the available storage resources of a single node in the inference cluster match the resource specifications; Wherein, when performing single-node deployment, a target single node that meets the minimum storage resource requirement is automatically selected in the inference cluster, so as to configure storage resources corresponding to the resource specifications on the target single node; When performing multi-node distributed deployment, a plurality of target nodes that meet the distributed deployment conditions are automatically selected in the inference cluster to allocate storage resources corresponding to the resource specifications among the plurality of target nodes.
3. The method for configuring inference service resources according to claim 2, wherein: Configuring storage resources in the inference cluster according to the resource configuration policy to run the inference service includes: On the target single node or multiple target nodes, respectively configure storage resources corresponding to the resource specifications; After completing the storage resource configuration, starting the inference service instance corresponding to the configured storage resource on the configured node to run the inference service; In the case of multi-node distributed deployment, communication links are established between different inference service instances to collaboratively run the inference services.
4. The method for configuring inference service resources according to claim 1, wherein: The determining of a resource configuration strategy that matches the resource specifications based on the storage resource status of the nodes in the inference cluster further includes: In a case where the pre-population phase and the decoding phase of the inference service are deployed separately, determining, according to the resource specifications, a first target node set for running the pre-population phase and a second target node set for running the decoding phase; An inference service instance of a pre-filling stage is deployed in the first target node set, and an inference service instance of a decoding stage is deployed in the second target node set, so as to deploy the pre-filling stage separately from the decoding stage.
5. The method for configuring inference service resources according to claim 4, wherein: The method further comprises: In a case where there is a historical running template of the inference service, determining a first sub-resource specification and a second sub-resource specification based on the historical running template; Determining, based on the historical operation template, the number of inference service instances of the inference service in the pre-filling phase and the number of inference service instances of the inference service in the decoding phase; configuring storage resources in the pre-population phase based on the first sub-resource specification to deploy an inference service instance of the inference service in the pre-population phase; The storage resources of the decoding stage are configured based on the second sub-resource specification to deploy the inference service instance of the inference service in the decoding stage.
6. The method for configuring inference service resources according to claim 1, wherein: The method further comprises: When the inference cluster contains heterogeneous nodes, based on the node performance parameters of each node, a set of candidate nodes that meet the resource specifications are screened; Based on the operation requirements of the reasoning service, determining a target node that can host the reasoning service from the set of candidate nodes; According to the resource configuration policy, storage resources are configured on the target node, and a corresponding inference service instance is started.
7. The method for configuring inference service resources according to any one of claims 1 to 6, characterized in that: The method further comprises: During the operation of the inference service, collecting inference performance parameters of the inference service; Based on the comparison result of the inference performance parameter and the preset threshold, the expansion operation of the storage resource is automatically triggered to optimize the inference performance of the inference service.
8. The method for configuring inference service resources according to claim 7, wherein: The inference performance parameters include at least first word time and per word time; The automatically triggering the expansion operation of the storage resource based on the comparison result of the inference performance parameter and the preset threshold value includes: In a case where the first word meta-time exceeds the first word meta-time standard value, increasing the number of inference service instances deployed by the inference service in the pre-population phase to expand the capacity of the inference service; In a case where the per-word time exceeds a standard per-word time value, the number of inference service instances deployed by the inference service in the decoding phase is increased to expand the capacity of the inference service.
9. The method for configuring inference service resources according to claim 1, wherein: In the case where there are multiple inference services, the method further includes: Obtaining model files of the multiple reasoning services, and parsing the model files respectively to determine the minimum storage resource requirements required to run the multiple reasoning services respectively; Based on the storage resource status of the nodes in the inference cluster, determine one or more candidate resource specifications corresponding to any inference service; Combining and generating a plurality of global resource configuration strategies based on the candidate resource specifications corresponding to the plurality of inference services; Selecting a target global resource configuration strategy that can run the most inference services from the multiple global resource configuration strategies; Storage resources are configured according to the target global resource configuration policy to run the multiple inference services.
10. The method for configuring inference service resources according to claim 9, wherein: The multiple global resource configuration strategies are generated based on the candidate resource specifications corresponding to the multiple reasoning services, including: For any candidate resource specification of the inference service, determining whether the available storage resources of the nodes in the inference cluster meet the current candidate resource specification; If there is a node that meets the current candidate resource specifications, add the current candidate resource specifications to the current global resource configuration policy; The next inference service of the current inference service is used as the new current inference service, and the candidate resource specifications for any of the inference services are returned to, and the step of determining whether the available storage resources of the nodes in the inference cluster meet the current candidate resource specifications is continued until the current global resource configuration strategy includes a candidate resource specification corresponding to each of the multiple inference services.
11. The method for configuring inference service resources according to claim 10, wherein: The method further comprises: If there is a node that meets the current candidate resource specifications, a new global resource configuration policy is generated based on the current global resource configuration policy and used as the current global resource configuration policy; Determine the next candidate resource specification of the current candidate resource specification; Determining whether the available storage resources of the nodes in the inference cluster meet the next candidate resource specifications; If there is a node that meets the next candidate resource specification, add the next candidate resource specification to the current global resource configuration strategy; If there is no node that meets the next candidate resource specification, stop generating the current global resource configuration strategy; Generate a new global resource configuration policy based on the current global resource configuration policy, and use the next inference service of the current inference service as the new current inference service, and return to the candidate resource specifications for any of the inference services, and continue to execute the step of determining whether the available storage resources of the nodes in the inference cluster meet the current candidate resource specifications.
12. The method for configuring inference service resources according to claim 11, wherein: The method further comprises: If there is no node that meets the current candidate resource specifications, stop generating the current global resource configuration policy; Generate a new global resource configuration policy based on the current global resource configuration policy and use it as the current global resource configuration policy; The next inference service of the current inference service is used as the new current inference service, and the step of returning to the candidate resource specifications for any of the inference services and determining whether the available storage resources of the nodes in the inference cluster meet the current candidate resource specifications is continued.
13. The method for configuring inference service resources according to claim 10, wherein: The method further comprises: If there is no node that meets the current candidate resource specifications, stop generating the current global resource configuration policy, generate a new global resource configuration policy based on the current global resource configuration policy and use it as the current global resource configuration policy; The next inference service of the current inference service is used as the new current inference service, and the step of returning to the candidate resource specifications for any of the inference services and determining whether the available storage resources of the nodes in the inference cluster meet the current candidate resource specifications is continued.
14. An electronic device, characterized in that: include: memory for storing computer programs; A processor, configured to implement the steps of the method for configuring inference service resources according to any one of claims 1 to 13 when executing the computer program.
15. A non-volatile computer-readable storage medium, characterized in that: The non-volatile computer-readable storage medium stores a computer program, wherein when the computer program is executed by a processor, the steps of the method for configuring inference service resources according to any one of claims 1 to 13 are implemented.
Citation Information
Patent Citations
Load scheduling method and system for heterogeneous large model
CN118519777A
Inference service management method, equipment, medium and computer program product
CN119415273A
Model virtualization deployment method and device, storage medium and computer equipment
CN119597394A
Inference service method, processing device, equipment, storage medium and program product
CN119862958A
Inference service deployment method and apparatus, device, and storage medium
EP4280051A1
Cited By
Computing power resource allocation method and device, storage medium and program product
CN121008933A
Storage system management method and device, electronic equipment, medium and product
CN121050659A
Distribution method, distribution device, electronic equipment and storage medium
CN121070753A