A computing power resource scheduling method, related equipment, storage medium and program product
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- TENCENT TECHNOLOGY (SHENZHEN) CO LTD
- Filing Date
- 2025-02-05
- Publication Date
- 2026-08-07
AI Technical Summary
由于传统的技术方案对算力资源的调度具有较大的局限性,使得模型处理业务难以得到灵活且高效的业务处理,业务扩展性也受到相应限制
[0051]在本申请实施例中,将模型处理业务部署至异构资源集群,并针对模型处理业务的不同业务场景配置不同的运行检测指标与资源调度策略,然后基于运行检测指标在业务运行过程中呈现出的指标值,来选取对应的资源调度策略进行算力资源的调度,可以实现在异构算力资源下以指标为粒度的差异化调度,从而有效提升模型处理业务的扩展性与的灵活性。此外,通过将模型处理业务部署在异构资源集群,还可以使得同一模型处理业务的业务处理能够联合多种不同计算特性的算力资源来实现,从而提升业务处理的效率。
Smart Images

Figure CN122526720A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of computer technology, and in particular to a computing resource scheduling method, related equipment, storage medium and program product. Background Technology
[0002] Model processing applications (such as intelligent question answering and speech recognition) typically involve complex data processing and computation. To ensure efficient execution, service providers usually need to allocate substantial computing resources for these processes. However, traditional technologies have significant limitations in allocating computing resources, hindering flexible and efficient model processing and restricting scalability. Therefore, how to allocate computing resources to achieve flexible and efficient model processing has become a key research topic. Summary of the Invention
[0003] This application provides a computing resource scheduling method, related equipment, storage medium, and program products, which can realize flexible and efficient processing of model processing services.
[0004] On the one hand, embodiments of this application provide a method for scheduling computing resources, including:
[0005] On the resource scheduling platform, corresponding operation detection indicators and resource scheduling strategies under the operation detection indicators are configured for different business scenarios of model processing business. The model processing business refers to data processing business based on artificial intelligence model.
[0006] If the model processing service is detected to be running, the target computing power container running the model processing service is determined from the heterogeneous resource cluster connected to the resource scheduling platform.
[0007] Obtain the target business scenario associated with the target computing power container, and detect the measured index value of the target computing power container under the target operation detection index. The target operation detection index refers to the operation detection index configured by the model processing business under the target business scenario.
[0008] Based on the measured index values and the target resource scheduling strategy of the target computing power container under the target operation detection index, resource scheduling is performed on the computing power resources contained in the target computing power container.
[0009] Furthermore, embodiments of this application provide a computing resource scheduling device, including:
[0010] The configuration unit is used to configure corresponding operation detection indicators and resource scheduling strategies under the operation detection indicators for different business scenarios of model processing business on the resource scheduling platform. The model processing business refers to data processing business based on artificial intelligence model.
[0011] The container determination unit is used to determine the target computing power container running the model processing service from the heterogeneous resource cluster connected to the resource scheduling platform if the model processing service is detected to be running.
[0012] The indicator detection unit is used to obtain the target business scenario associated with the target computing power container and detect the measured indicator value of the target computing power container under the target operation detection indicator. The target operation detection indicator refers to the operation detection indicator configured by the model processing business under the target business scenario.
[0013] The resource scheduling unit is used to schedule the computing resources contained in the target computing container based on the measured index value and the target resource scheduling strategy of the target computing container under the target operation detection index.
[0014] In one implementation, after scheduling the computing resources contained in the target computing container based on the measured index values and the target resource scheduling strategy, the index detection unit can also be used to perform:
[0015] After a preset time interval, the current index value of the target computing power container under the target operation detection index is obtained, and the expected index value range of the target operation detection index is obtained.
[0016] If the current indicator value is outside the expected indicator value range, then the target resource scheduling strategy is optimized based on the current indicator value, the measured indicator value, and the expected indicator value range.
[0017] In another implementation, the target resource scheduling strategy includes parameter information for at least two resource scheduling parameters; when the indicator detection unit optimizes the target resource scheduling strategy based on the current indicator value, the measured indicator value, and the expected indicator value range, it may specifically perform the following:
[0018] Select a reference indicator value from the expected indicator value range, and obtain a first difference between the current indicator value and the reference indicator value, and a second difference between the measured indicator value and the reference indicator value;
[0019] From the resource scheduling parameters included in the target resource scheduling strategy, select the resource scheduling parameters to be adjusted, and based on the first difference and the second difference, determine the parameter adjustment direction and parameter adjustment magnitude of the resource scheduling parameters to be adjusted.
[0020] According to the adjustment direction and adjustment magnitude of the parameters, update the parameter information of the resource scheduling parameters to be updated in the target resource scheduling strategy.
[0021] In another implementation, when the configuration unit configures corresponding operation detection indicators and resource scheduling strategies under the operation detection indicators for different business scenarios of model processing services on the resource scheduling platform, it can specifically execute:
[0022] Receive the service access request of the model processing service, and determine from the service access request the service scenario associated with the model processing service, as well as the operation detection indicators to be configured in the service scenario;
[0023] From the resource scheduling platform, query the resource scheduling strategy that matches the operation detection index;
[0024] Obtain the business processing requirements of the model in the business scenario, and predict the matching degree between the business processing requirements and the queried resource scheduling strategy;
[0025] If the matching degree is greater than the preset matching degree, then the queried resource scheduling strategy is configured under the running detection index.
[0026] In another implementation, before scheduling the computing resources contained in the target computing power container based on the measured index value and the target resource scheduling strategy, the configuration unit may also be used to perform:
[0027] Obtain the expected value range of the operational detection indicators;
[0028] Based on the operational monitoring indicators and the expected indicator value range, an indicator anomaly event is created for the operational monitoring indicators. The indicator anomaly event is triggered when the indicator value of the operational monitoring indicators is outside the expected indicator value range.
[0029] Based on the target resource scheduling strategy and the abnormal indicator event, a resource scheduling component is generated and deployed in the target computing power container. The resource scheduling component is used to execute the target resource scheduling strategy when the abnormal indicator event is detected.
[0030] In another implementation, the business scenario includes a model training scenario, and the operation detection metrics under the model training scenario include at least one of resource utilization, the queue length of the training task queue, and the container status.
[0031] The resource scheduling strategy under the resource utilization rate is used to indicate that when the resource utilization rate is less than the preset utilization rate, the number of computing power containers associated with the model training scenario should be reduced.
[0032] The resource scheduling strategy under the queue length of the training task queue is used to indicate that when the queue length is greater than the preset length, based on the model training requirements of the model processing business in the model training scenario, the task priority is configured for each training task contained in the training task queue in the computing power container associated with the model training scenario.
[0033] The resource scheduling strategy under the container state is used to instruct that when the container state is in an abnormal state, breakpoint resume training should be performed based on the training record data generated by the target computing power container.
[0034] In another implementation, if the target business scenario is the model training scenario, and the target operation detection metric includes the container status, the resource scheduling unit can also be used to execute:
[0035] During the operation of the target computing power container, the training record data generated by the target computing power container is stored in a preset storage space according to a preset storage period. The preset storage space is located outside the target computing power container.
[0036] When scheduling computing resources contained in the target computing power container based on the measured index values and the target resource scheduling strategy, the following can be specifically executed:
[0037] If the measured index value of the container state is used to indicate that the target computing power container is in an abnormal state, then the historical index value of the target computing power container is obtained, and the historical index value includes the index value of the container state in the previous storage cycle.
[0038] When the historical indicator value is used to indicate that the target computing power container is in a normal state, the training record data generated by the target computing power container in the previous storage cycle is obtained from the preset storage space, and a new computing power container is created.
[0039] The acquired training record data is loaded into the new computing power container, and the model processing service under the model training scenario is run in the new computing power container based on the loaded training record data.
[0040] In another implementation, the business scenario includes a model inference scenario, and the operational monitoring metrics under the model inference scenario include at least one of the execution success rate of the inference task and the container load.
[0041] The resource scheduling strategy under the execution success rate is used to indicate that when the execution success rate is less than the preset success rate, at least one computing power container associated with the model inference scenario is rebuilt.
[0042] The resource scheduling strategy under container load is used to indicate that when a container is overloaded, the number of computing power containers associated with the model inference scenario should be increased.
[0043] In another implementation, if the target business scenario is the model inference scenario, and the target operation detection metric includes the container load, the resource scheduling unit, when scheduling the computing resources contained in the target computing container based on the measured metric value and the target resource scheduling strategy, may specifically perform the following:
[0044] If the measured index value of the container load is used to indicate container overload, then at least one new computing power container is created.
[0045] The model is deployed in each new computing power container to process the service, so as to obtain at least one new target computing power container.
[0046] In another aspect, embodiments of this application provide a computer device, including:
[0047] A memory, wherein a computer program is stored;
[0048] A processor for loading the computer program to implement the method as described in the first aspect.
[0049] In another aspect, embodiments of this application also provide a computer-readable storage medium storing a computer program adapted to be loaded by a processor and executed as described in the first aspect.
[0050] In another aspect, embodiments of this application also provide a computer program product, the computer program product including computer instructions, wherein a processor of a computer device reads the computer instructions and executes the method as described in the first aspect.
[0051] In this embodiment, model processing services are deployed to a heterogeneous resource cluster. Different operational monitoring metrics and resource scheduling strategies are configured for different business scenarios within the model processing services. Then, based on the metric values exhibited during service operation, the corresponding resource scheduling strategy is selected to allocate computing resources. This enables differentiated scheduling at the metric level under heterogeneous computing resources, effectively improving the scalability and flexibility of the model processing services. Furthermore, by deploying model processing services on a heterogeneous resource cluster, the same model processing service can be implemented by combining computing resources with different computational characteristics, thereby improving the efficiency of service processing. Attached Figure Description
[0052] To more clearly illustrate the technical solutions of the embodiments of this application, the drawings used in the description of the embodiments will be briefly introduced below. Obviously, the drawings described below are some embodiments of this application. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0053] Figure 1 This is a schematic diagram of the structure of a heterogeneous resource cluster provided in an embodiment of this application;
[0054] Figure 2 This is a schematic diagram of the structure of a computing resource scheduling system provided in an embodiment of this application;
[0055] Figure 3 This is a schematic diagram illustrating the strategy configuration principle provided in the embodiments of this application;
[0056] Figure 4 This is a flowchart illustrating a computing resource scheduling method provided in an embodiment of this application;
[0057] Figure 5 This is a schematic diagram of an indicator change information provided in an embodiment of this application;
[0058] Figure 6 This is a schematic diagram of a resource scheduling platform architecture provided in an embodiment of this application;
[0059] Figure 7 This is a schematic diagram illustrating the policy configuration principle in a resource scheduling platform provided in an embodiment of this application;
[0060] Figure 8 This is a schematic diagram of a strategy management method provided in an embodiment of this application;
[0061] Figure 9 This is a schematic diagram of a strategy structure provided in an embodiment of this application;
[0062] Figure 10This is a schematic diagram illustrating the application principle of a resource management method provided in an embodiment of this application;
[0063] Figure 11 This is a schematic diagram of the structure of a computing resource scheduling device provided in an embodiment of this application;
[0064] Figure 12 This is a schematic diagram of the structure of a computer device provided in an embodiment of this application. Detailed Implementation
[0065] It should be noted in advance that, in order to enable those skilled in the art to better understand the technical solutions proposed in the embodiments of this application, the embodiments of this application will be described clearly and completely in conjunction with one or more accompanying drawings. Furthermore, the accompanying drawings shown in the embodiments of this application are merely illustrative examples; for instance, the execution order of each step in the drawings can be adaptively adjusted according to the actual application scenario.
[0066] Furthermore, in the embodiments of this application, the block diagrams, modules, and units shown in the accompanying drawings are merely functional entities and do not necessarily correspond to physically independent entities. Each module or unit can be part of a larger module or unit that includes the functionality of that module or unit. That is, the terms "module" or "unit" mentioned in the embodiments of this application refer to a computer program or part of a computer program with a predetermined function, which can work together with other related parts to achieve a predetermined goal. It can be implemented wholly or partially using software, hardware (such as processing circuitry or memory), or a combination thereof, or implemented in different network and / or processor devices and / or microcontroller devices. Similarly, a processor (or multiple processors or memory) can be used to implement one or more modules or units.
[0067] Specifically, this application provides a computing resource scheduling scheme for model processing services of artificial intelligence models. This scheme deploys the model processing service within computing containers of a heterogeneous computing cluster and configures corresponding operational monitoring indicators and resource scheduling strategies for different business scenarios of the model processing service. Each resource scheduling strategy can be associated with at least one operational monitoring indicator. While the model processing service is running, the target computing container running the model processing service and the target business scenario associated with that target computing container under the model processing service are first determined from the heterogeneous computing cluster. Then, the indicator values of the operational monitoring indicators configured under the target business scenario are detected in the target computing container to obtain the actual measured indicator value corresponding to the operational monitoring indicator in the target computing container. Finally, based on the actual measured indicator value and the resource scheduling strategy associated with the operational monitoring indicator, the computing resources contained in the target computing container are scheduled.
[0068] Here, model processing services refer to data processing services implemented based on artificial intelligence models. Optionally, model processing services may include, but are not limited to, speech processing services implemented based on speech processing models (such as speech recognition, speech conversion, etc.), text processing tasks implemented based on natural language processing models (such as semantic understanding, intelligent question answering, language translation, etc.), and image processing services implemented based on image processing models (such as image recognition, image generation, etc.), etc., without limitation. Heterogeneous resource clusters are used to provide heterogeneous computing resources for model processing services. Heterogeneous computing resources refer to composite computing resources obtained by integrating at least two types of computing resources. Different types of computing resources usually have different computing characteristics (such as efficiency, applicable data formats, etc.).
[0069] In other words, the heterogeneous resource cluster in this application can be used to provide at least two types of computing resources. In this case, deploying the model processing service to the heterogeneous resource cluster allows the service to utilize computing resources with different computing characteristics, thereby effectively improving the execution efficiency and flexibility of the model processing service. Furthermore, configuring different operational monitoring metrics and resource scheduling strategies for different business scenarios of the model processing service enables differentiated strategy configuration at the granular level of metrics. Providing differentiated resource scheduling methods for different business scenarios not only improves the scalability of the model processing service but also further enhances its flexibility. Therefore, this application, by combining heterogeneous resource clusters with differentiated strategy configuration at the granular level of operational monitoring metrics, ensures that the model processing service can be processed flexibly and efficiently.
[0070] In one feasible implementation, the structure of a heterogeneous resource cluster can be exemplarily described as follows: Figure 1 .like Figure 1 As shown, the heterogeneous resource cluster can include a container cluster built based on computing power resources provided by at least two heterogeneous devices (or heterogeneous cards). The computing containers in this container cluster can be used to deploy model processing services. Here, heterogeneous devices refer to dedicated hardware devices used to accelerate specific types of computing tasks; different types of heterogeneous devices have different computing characteristics. Optionally, the heterogeneous devices used in this application may include, but are not limited to, those mentioned above. Figure 1 The graphics processing unit (GPU), tensor processing unit (TPU), and field-programmable gate array (FPGA) shown in the figure.
[0071] In this application, a heterogeneous device is typically used to provide a computing power resource. A computing power container can be constructed based on a computing power resource, and the number of each type of computing power container can be at least one (e.g., ...). Figure 1 The container (within the same dashed box) allows for at least two different types of computing containers with varying computing characteristics to exist within a heterogeneous resource cluster. Building computing containers based on a single computing resource simplifies the utilization logic of heterogeneous computing resources, thereby improving container construction efficiency and reducing the difficulty of container management and resource scheduling complexity. Furthermore, containerizing model processing services allows for dynamic adjustment of the number of containers to respond promptly to sudden changes in traffic demand, thus maintaining the functional stability of the model processing services.
[0072] based on Figure 1 The heterogeneous resource cluster shown in this application also provides a resource scheduling system, the structure of which can be exemplarily described in [reference needed]. Figure 2 .like Figure 2 A computing resource scheduling system can include a resource scheduling platform, a heterogeneous resource cluster controlled or managed by the resource scheduling platform, and at least one model processing service that has established a service connection with the resource scheduling platform. As an example, a service connection between the model processing service and the resource scheduling platform can be established by connecting the model processing service to the heterogeneous resource service provided by the resource scheduling platform. Here, heterogeneous resource service refers to data processing services (such as data computation and data storage) provided based on the heterogeneous computing resources contained in the heterogeneous resource cluster. In specific implementation scenarios, heterogeneous resource service can be provided through computing power containers, allowing access to heterogeneous resource service to be achieved by deploying the model processing service to a computing power container within the heterogeneous resource cluster.
[0073] In one exemplary implementation, the resource scheduling platform can jointly deploy model processing services into different types of computing power containers (such as...). Figure 2 (Image recognition service in the context of computing power) can leverage the computational characteristics of various computing resources to collaboratively execute related business processes for the model processing service, thereby improving the processing efficiency of the model processing service. Optionally, the resource scheduling platform can also deploy the model processing service to different computing power containers, with each container independently executing the relevant business processes for the model processing service, thereby achieving load balancing and improved robustness of the model processing service. These different computing power containers can be different types of computing power containers (such as...). Figure 2 The computing power containers deployed in the middle can be either computing power containers that provide image recognition services or different computing power containers of the same type (such as...). Figure 2 (The computing power container deployed in the middle has voice processing services), no restrictions are imposed here.
[0074] During (or after) the deployment of model processing services, the resource scheduling platform can also configure operational monitoring metrics (hereinafter referred to as metrics) and resource scheduling policies applicable to the model processing services. The configured operational monitoring metrics and resource scheduling policies can be stored on the resource scheduling platform or a storage device connected to the resource scheduling platform; there are no restrictions on this. For example, the principle of policy configuration can be found in [link to relevant documentation]. Figure 3 .
[0075] like Figure 3 This application allows for the configuration of different metrics (e.g., ...) for different business scenarios of model processing. Figure 3 For image recognition services, resource scheduling strategies can differ under different metrics, while the same resource scheduling strategy can be used for the same metric (e.g., ...). Figure 3 Indicator 3) or different (e.g. Figure 3 (1) In other words, the same resource scheduling strategy in this application can be reused under different indicators. In practical applications, to ensure business execution efficiency, only resource scheduling strategies with higher quality can be reused. The quality of the strategy can be evaluated based on the changes in indicators after the strategy is implemented. By reusing resource scheduling strategies, the creation of resource scheduling strategies can be avoided during strategy configuration in some cases, thereby effectively reducing the workload during strategy configuration and improving strategy configuration efficiency to a certain extent.
[0076] It should be noted here that, in one optional implementation, the runtime monitoring metrics and resource scheduling strategies can be provided by the model processing business based on its own business processing needs, or they can be predicted by the resource scheduling platform based on the business processing characteristics of the model processing business. For example, business processing characteristics can be obtained by collecting business data generated by the model processing business and then performing data analysis on that data, or by analyzing the business description information of the model processing business (such as by comparing or similaring it to other businesses), without any limitations.
[0077] Based on the aforementioned computing resource scheduling scheme and related systems, this application proposes a specific computing resource scheduling method. This method can be executed by a resource scheduling platform, and a schematic flowchart of this method can be found here. Figure 4 .like Figure 4 As shown, the method may include steps S401-S404:
[0078] S401. On the resource scheduling platform, configure corresponding operation detection indicators and resource scheduling strategies under the operation detection indicators for different business scenarios of model processing business. Model processing business refers to data processing business based on artificial intelligence models.
[0079] In a specific embodiment, the resource scheduling platform can provide heterogeneous resource services to model processing services. These services primarily execute relevant data processing within the model processing services by scheduling the computing resources contained in corresponding computing containers within the heterogeneous resource cluster. Optionally, the resource scheduling platform can connect multiple different model processing services, each of which can be configured with corresponding operational monitoring metrics and resource scheduling strategies within the platform. Operational monitoring metrics refer to metrics whose values need to be monitored during service operation. These can include business-level metrics (such as user churn rate, service exposure, etc.) or device-level metrics (such as resource utilization rate, processing latency, etc.). The resource scheduling strategy is at least used to indicate the scheduling method of computing resources when executing relevant computational tasks; optionally, it can also be used to indicate the strategy triggering conditions.
[0080] In one implementation, differentiated business deployment and / or strategy configuration can be performed for different business scenarios processed by a single model. Differentiated business deployment refers to selecting computing power containers adapted to the computing needs of different business scenarios from the computing power containers contained in the heterogeneous resource cluster for deploying the model processing business. Differentiated strategy configuration refers to configuring different operation monitoring indicators and / or resource scheduling strategies for different business scenarios of the same model processing business, thereby achieving customized resource scheduling for the model processing business in each business scenario.
[0081] Here, "business scenario" refers to the application scenario in which the model processes business. Different business scenarios typically have different business processing requirements. Generally, the business scenario for model processing can include at least one of the following: model training scenario and model inference scenario. By implementing differentiated business deployment or strategy configuration based on different business scenarios, the resource scheduling platform can flexibly adjust the allocation of computing resources to model processing based on the actual processing requirements of the model, thereby effectively avoiding resource waste.
[0082] In one exemplary implementation, the principle of differentiated business deployment can be illustrated by the following example: Assume that model processing service 1's computational requirements in the model inference scenario include high concurrency, and the GPUs in heterogeneous devices are well-suited for task processing scenarios requiring extensive parallel computing. Therefore, when a computing power container a, built based on GPU computing resources, exists in the heterogeneous resource cluster, model processing service 1 can be preferably deployed in computing power container a. Similarly, assuming that model processing service 2's computational requirements in the model training scenario include fast data access, and the TPUs in heterogeneous devices are well-suited for data stream processing in neural network computations, then when a computing power container b, built based on TPUs, exists in the heterogeneous resource cluster, model processing service 2 can be preferably deployed in computing power container b.
[0083] Optionally, after determining the computing power container suitable for deploying model processing services in the business scenario, a relationship between the computing power container and the business scenario can be further established. This allows the resource scheduling platform to quickly select the appropriate computing power container through the relationship when it receives the computing task of the model processing service in the business scenario, thereby efficiently processing the relevant tasks and effectively improving the adaptability of computing power resources to the business scenario.
[0084] Similarly, in one exemplary implementation, the differentiated strategy deployment principle may include any one of the following (1)-(3):
[0085] (1) Configure different operation detection indicators for different business scenarios, and configure the same resource scheduling strategy under different operation indicators.
[0086] (2) Configure different operation detection indicators for different business scenarios, and configure different resource scheduling strategies under different operation detection indicators.
[0087] (3) Configure the same operation detection indicators for different business scenarios, and configure different resource scheduling strategies under the same operation detection indicators.
[0088] It's worth mentioning that different resource scheduling strategies can be based on different resource scheduling methods or different triggering conditions. The triggering conditions can be configured based on the value range (or threshold) of the runtime monitoring metrics. For example, when configuring different resource scheduling strategies for the same runtime monitoring metric, resource scheduling strategy a can be configured based on metric value range A, and resource scheduling strategy b can be configured based on metric value range B. The resource scheduling methods indicated by resource scheduling strategies a and b can be the same (e.g., both are container scaling), but resource scheduling strategy a can be triggered when the metric value 1 is within metric value range A, and resource scheduling strategy b can be triggered when the metric value 1 is within metric value range B. Of course, resource scheduling strategies can also be triggered when the metric value is outside the corresponding metric value range, depending on specific business needs; there are no restrictions here.
[0089] It should be understood that, in the practical application of this application, resource scheduling strategies can also be deployed in the corresponding computing power containers to achieve automatic implementation of resource scheduling strategies.
[0090] Specifically, taking the resource scheduling strategy triggering condition determined based on the value range of the operational monitoring indicators as an example, the automatic implementation of the resource scheduling strategy can be achieved by deploying an automatically executable resource scheduling component in the computing power container. For example, the resource scheduling platform can first obtain the expected value range of the operational monitoring indicators, and then create an indicator anomaly event based on the operational monitoring indicators and their expected value range. Here, the expected value range refers to the range of indicator values that the operational monitoring indicator is expected to be in during the operation of the model processing business, and the indicator anomaly event refers to an event where the value of the operational monitoring indicator is outside the indicator value range. Then, based on the resource scheduling strategy under the operational monitoring indicators and the indicator anomaly event, a resource scheduling component can be generated. This resource scheduling component is used to automatically trigger the execution of the resource scheduling strategy when an indicator anomaly event is detected. By deploying or embedding this resource scheduling component in the computing power container where the model processing business is deployed, the resource scheduling strategy can be automatically triggered and executed within the computing power container when an indicator anomaly event is detected, without the need for additional control from third-party devices, which can effectively improve the container response speed.
[0091] In one feasible implementation, the operational monitoring metrics and / or resource scheduling strategies configured for model processing tasks in various business scenarios can be automatically determined by the resource scheduling platform based on the business requirements of the model processing business. In this case, the resource scheduling platform can pre-store one or more resource scheduling strategies in the database, and the model processing business can initiate a business access request to the resource scheduling platform.
[0092] For example, after receiving a business access request from a model processing service, the resource scheduling platform can determine the business scenario associated with the model processing service and the operational monitoring indicators to be configured under that business scenario from the access request. Then, it queries a pre-stored resource scheduling strategy that matches the operational monitoring indicator. Optionally, the resource scheduling strategy matching the operational monitoring indicator can be a resource scheduling strategy used by other model processing services under that operational monitoring indicator, or it can be a pre-set resource scheduling strategy for that operational monitoring indicator. After finding a resource scheduling strategy, the platform can further obtain the business processing requirements of the model processing service under that business scenario, and then predict the matching degree between the business processing requirements and the queried resource scheduling strategy. If the matching degree is greater than a preset matching degree, the queried resource scheduling strategy is used as the resource scheduling strategy to be configured under that operational monitoring indicator, and the corresponding configuration is executed. When the matching degree is less than or equal to the preset matching degree, the platform can optionally reselect or generate a resource scheduling strategy, prompt the model processing service to provide a corresponding resource scheduling strategy, or optimize an existing resource scheduling strategy to obtain a final resource scheduling strategy adapted to the operational monitoring indicator.
[0093] In another feasible implementation, the operational monitoring metrics and / or resource scheduling strategies configured for the model processing task in various business scenarios can also be directly provided by the model processing business. When the business scenario includes a model training scenario, the operational monitoring metrics provided by the model processing task in the model training scenario can include at least one of the following: resource utilization, queue length of the training task queue, and container status.
[0094] For example, the resource scheduling strategy provided for resource utilization can be used to instruct the number of computing power containers associated with the model training scenario to be reduced when the resource utilization is less than a preset utilization, thereby releasing idle resources and improving resource utilization. The resource scheduling strategy provided for the queue length of the training task queue can be used to instruct the task priority of each training task in the training task queue within the computing power containers associated with the model training scenario, based on the model training requirements of the model processing business in the model training scenario, when the queue length is greater than a preset length. The resource scheduling strategy provided for container status can be used to instruct the execution of breakpoint resume training based on the training record data generated by the computing power container in the abnormal state when the container status is abnormal. Breakpoint resume training refers to the process of continuing training from the point of interruption after an abnormal interruption. The implementation of breakpoint resume training typically requires saving a snapshot of the training state during the training process (called a "breakpoint" or "checkpoint"). In this embodiment of the application, during the operation of the computing power container, the training record data (equivalent to a snapshot) of the computing power container within the preset storage period can be stored in the preset storage space according to the preset storage period. The training record data may include training samples and parameter change data during the training process.
[0095] When a business scenario includes a model inference scenario, the operational monitoring metrics provided by the model processing task in the model inference scenario can include at least one of the inference task's execution success rate and container load. The resource scheduling strategy provided for the execution success rate can be used to instruct the reconstruction of at least one computing power container associated with the model inference scenario when the execution success rate is lower than a preset success rate; the resource scheduling strategy provided for container load can be used to instruct the increase of the number of computing power containers associated with the model inference scenario when containers are overloaded. In practical applications, container load can be specifically measured through more granular metrics such as resource utilization, waiting time of a single inference task, and queue length of inference tasks.
[0096] S402. If the model processing service is detected to be running, the target computing power container running the model processing service is determined from the heterogeneous resource cluster connected to the resource scheduling platform.
[0097] In a specific embodiment, when any computing power container in a heterogeneous resource cluster that deploys a certain model processing service is in a running state, it can be considered that the model processing service is in a running state, and this computing power container can be referred to as a target computing power container. It is understood that there is at least one target computing power container, and the business scenarios associated with different target computing power containers can be different. The business scenario associated with a target computing power container can refer to the business scenario associated with the service deployment, or it can refer to the business scenario currently being applied by the target computing power container. Furthermore, exemplarily, the currently applied business scenario can be determined based on the type of business data processed by the target computing power container (such as training type, inference type, etc.), or the type of the business data sending interface (such as training interface, inference interface, etc.).
[0098] S403. Obtain the target business scenario associated with the target computing power container, and detect the measured index value of the target computing power container under the target operation detection index. The target operation detection index refers to the operation detection index configured for the model processing business under the target business scenario.
[0099] In a specific embodiment, the resource scheduling platform can regard the business scenario of the model processing business in the target computing power container as the target business scenario, take the operation detection index configured for the model processing business in the target business scenario as the target operation detection index, and then obtain the index value of the target operation detection index in the target computing power container to obtain the measured index value of the target operation detection index.
[0100] S404. Based on the measured index values and the target resource scheduling strategy of the target computing power container under the target operation detection index, perform resource scheduling on the computing power resources contained in the target computing power container.
[0101] As mentioned in the relevant embodiments of step S401 above, in a specific implementation, the business scenario of model processing can include a model training scenario. In the model training scenario, the runtime detection metric can include the container status. The resource scheduling strategy under this metric can be used to indicate that when the container status is abnormal, breakpoint resumption training should be performed based on the training record data generated by the computing power container in that abnormal state. The training record data can be recorded periodically, and the recording period can also be called the storage period. It is understood that when the runtime detection metric is a container status, the metric value is the status indicator value of the container status (such as abnormal state, normal state, idle state, etc.).
[0102] In other words, if the target operational monitoring metric includes container status, then the measured value of the target operational monitoring metric can be used to indicate the current operational status of the target computing power container. In this case, for example, the resource scheduling platform can perform resource scheduling on the computing power resources contained in the target computing power container to achieve breakpoint resume training in the following way:
[0103] First, when the measured metric value indicates that the target computing power container is in an abnormal state, the historical state (also known as historical metric value) of the target computing power container is obtained. The historical state can be the container state of the target computing power container in the Xth storage period (X is a positive integer, such as 1) prior to the current storage period. If the historical state is normal, the training record data generated by the target computing power container in the Xth storage period (such as the previous storage period) is obtained from the preset storage space. The preset storage space can be located outside the target computing power container to avoid data inaccessibility when the container is abnormal. Then, a new computing power container is created, and the obtained training record data is loaded into the new computing power container. Then, the model processing business in the model training scenario is run in the new computing power container based on the loaded training record data. In this way, the new computing power container takes over the training task of the target computing power container, realizing breakpoint resume training. Optionally, after successfully executing breakpoint resume training, the computing power resources of the target computing power container can be released, such as by destroying the target computing power container, so that the computing power resources can be reused.
[0104] Similarly, as mentioned in the relevant embodiments of step S401 above, in another specific implementation, the business scenario of model processing can include a model inference scenario, and in the model inference scenario, the operation detection metric can include container load. The resource scheduling strategy under this metric can be used to indicate that when a container is overloaded, the number of computing power containers associated with the model inference scenario should be increased (equivalent to container expansion). Likewise, it is easy to understand that when the operation detection metric is container load, the metric value can be a status indicator value of the load state (such as overload, light load, etc.).
[0105] In other words, if the target operational monitoring metric includes container load, then the measured value of the target operational monitoring metric can be used to indicate the current load status of the target computing power container. In this case, for example, the resource scheduling platform can perform resource scheduling on the computing power resources contained in the target computing power container to achieve container expansion in the following way:
[0106] First, when the measured load metric indicates container overload, at least one new computing power container is created. Then, the model processing service corresponding to the target computing power container is deployed in each new computing power container to obtain at least one new target computing power container. After obtaining the new target computing power container, the model processing service deployed within the container can be run, allowing the new target computing power container to share the request traffic of the model processing service, thereby reducing the load on a single target computing power container. Of course, in other embodiments, if there are also idle target computing power containers in the heterogeneous resource cluster, the resource scheduling platform can also directly start the target computing power container to share the request traffic. This application does not restrict the method of container expansion.
[0107] It is worth mentioning that the resource scheduling strategy in this embodiment can be iteratively optimized. Specifically, it can be determined based on the changes in indicators after the strategy is implemented (i.e., resource scheduling of the target computing power container based on the resource scheduling strategy). For example, after a preset period of time has elapsed since the strategy was implemented, the resource scheduling platform can obtain the current indicator value of the target computing power container under the target operation detection indicator, as well as the expected indicator value range of the target operation detection indicator. If the current indicator value is outside the expected indicator value range, the target resource scheduling strategy is optimized based on the current indicator value, the measured indicator value, and the expected indicator value range.
[0108] In one feasible implementation, the target resource scheduling strategy can include parameter information for at least two resource scheduling parameters. These resource scheduling parameters refer to the parameters used when scheduling computing resources, such as concurrency, storage location, and data transmission frequency. The parameter information can specifically include parameter identifiers (e.g., parameter names) and parameter values.
[0109] In this scenario, when optimizing the target resource scheduling strategy, the resource scheduling platform can first select a reference indicator value from the expected indicator value range. The reference indicator value can be any value randomly selected within this range, or a pre-specified value at a preset position, such as an upper bound, lower bound, or median value; there are no restrictions here. Then, the platform can obtain the first difference between the current indicator value and the reference indicator value, and the second difference between the measured indicator value and the reference indicator value. The first difference can be used to measure the magnitude and distance between the current indicator value and the reference indicator value, and the second difference can be used to measure the magnitude and distance between the measured indicator value and the reference indicator value. Based on the first and second differences, changes in the target operation detection indicators (such as...) can be observed. Figure 5The curve shown can be used to accurately determine whether the implementation effect of the resource adjustment strategy has achieved the expected result. Optionally, the change information can include the direction of change (e.g., increasing, decreasing) and the magnitude of change (e.g., 20%, 0.9, etc.). If the expected result is not achieved, the resource scheduling parameter to be adjusted can be selected from the resource scheduling parameters included in the target resource scheduling strategy. Then, based on the first difference and the second difference, the parameter adjustment direction and the parameter adjustment magnitude of the resource scheduling parameter are determined. Then, the parameter information of the resource scheduling parameter in the target resource scheduling strategy is updated according to the parameter adjustment direction and the parameter adjustment magnitude.
[0110] When updating resource scheduling parameters, the principle of "fast reduction, slow increase" can be followed. "Fast reduction" means that when the parameter adjustment direction is to decrease, a larger parameter adjustment range can be set; "slow increase" means that when the parameter adjustment direction is to increase, a smaller parameter adjustment range can be set. Furthermore, the number of resource scheduling parameters to be adjusted can be one or more, and each resource scheduling parameter can be randomly selected or selected sequentially according to a permutation and combination method; there are no restrictions here.
[0111] In this embodiment, model processing services are deployed to a heterogeneous resource cluster. Different operational monitoring metrics and resource scheduling strategies are configured for different business scenarios within the model processing services. Then, based on the metric values exhibited during service operation, the corresponding resource scheduling strategy is selected to allocate computing resources. This enables differentiated scheduling at the metric level under heterogeneous computing resources, effectively improving the scalability and flexibility of the model processing services. Furthermore, by deploying model processing services on a heterogeneous resource cluster, the same model processing service can be implemented by combining computing resources with different computational characteristics, thereby improving the efficiency of service processing.
[0112] exist Figure 4 In the computing resource scheduling method shown, the architecture of the resource scheduling platform can be exemplarily as follows: Figure 6 As shown. Figure 6 The resource scheduling platform can include a business module, a policy configuration module, and a computing resource module. The business module is used to establish service connections with one or more model processing services, which can be configured as follows: Figure 6 This includes mixed-model processing, game model processing, and visual model processing, among others. The strategy configuration for each model processing business can be completed in the strategy configuration module. Specifically, the strategy configuration module is used to configure the runtime detection metrics that need to be monitored during the operation of the model processing business. Figure 6 The indicators (abbreviated as indicators in Chinese) and the resource scheduling strategies used to configure each operational monitoring indicator. Figure 6 (Abbreviation strategy). For example... Figure 6As can be seen, the configuration of metrics and strategies is related to the business scenario of model processing, and the metrics and their corresponding strategies can be stored in the database. The strategies in the database can be implemented in the computing power containers included in the computing power resource module, and model processing services can be deployed in the computing power containers.
[0113] Optionally, the database may contain one or more preset resource scheduling policies for each model to choose from when processing business. For example, the default resource scheduling policies may be as shown in Table 1.
[0114] Table 1
[0115]
[0116]
[0117] Furthermore, the policy configurations for model training and model inference scenarios can be exemplified in Table 2.
[0118] Table 2
[0119]
[0120] In a specific application scenario, when configuring policies in a resource scheduling platform, you can refer to, for example... Figure 7 It is implemented according to the principles shown. For example... Figure 7 Different operational monitoring metrics can be configured for different business scenarios within the same service, and each operational monitoring metric can be associated with one or more resource scheduling policies. It's worth noting that the initial configuration of metrics and policies can be termed "registration," which can occur when the model processing service first connects to the resource scheduling platform. Registered policies can be iteratively optimized. The overall policy management approach can be as follows: Figure 8 It should be noted that, Figure 8 The threshold in the text is merely an example of a triggering condition for the strategy, and can be replaced with other information in practical applications.
[0121] Optionally, the structure of the resource scheduling strategy can be as follows: Figure 9 .based on Figure 9As can be seen, a complete resource scheduling strategy can include a strategy execution entity and strategy execution configuration items. Strategy execution configuration items can include the strategy's triggering conditions, the resource scheduling parameters used during resource scheduling and their information, and the display method of resource scheduling information (e.g., whether to output, and in what form). The strategy execution entity refers to the object used to execute the strategy, specifically including one or more strategy execution programs and strategy execution environments. For example, the strategy execution program can be the resource scheduling component mentioned in step S401, and the strategy execution environment can be one or more computing power containers with certain attributes (e.g., computing characteristics, environment, resource type, deployed services, etc.) in a heterogeneous resource cluster.
[0122] To facilitate a quick understanding and application of the various implementation methods provided in this application by those skilled in the art, the embodiments of this application also propose an application principle for a computing resource management method, which can be found in the following details. Figure 10 .like Figure 10 This principle can be exemplarily represented by processes S1-S14:
[0123] First, the personnel responsible for model processing can connect their business to the resource scheduling platform (S1) and provide business metrics to the platform upon connection. Then, the resource scheduling platform will register the business metrics for that model processing business (S2). It should be noted that business metrics refer to operational monitoring metrics. Figure 10 For simplicity, "business metrics" or "metrics" are used as a euphemism. After registering a business metric, it checks whether there is an associated operation strategy (S3). Similar to business metrics, the operation strategy here is a euphemism for resource scheduling strategy. If no associated operation strategy is detected in the resource scheduling platform, the business personnel can be notified to provide an operation strategy for that business metric (S4). Here, "no associated operation strategy" means that the operation strategy configured under that business metric does not exist in the resource scheduling platform's database, and the business personnel have not provided an operation strategy for that business metric. If an associated operation strategy is detected, it is determined whether that operation strategy exists only in the database (S5).
[0124] If the data exists only in the database, the corresponding operation strategy is selected according to the business requirements of the model (i.e., S6). These business requirements can include one or both of the business scenario and the specific data processing requirements within that scenario. If the data does not exist only in the database (e.g., business personnel provide corresponding operation strategies), it is determined whether the operation strategy currently provided by the business personnel meets the business requirements (i.e., S7). If it does not meet the business requirements, a new operation strategy can be constructed (i.e., S8). If it meets the business requirements, the operation strategy can be registered under that business metric (i.e., S9). The new operation strategy can be obtained by updating the existing operation strategy, or it can be rewritten or reconstructed.
[0125] After registering the operational strategies under the business metrics, when the model is processing the business in the running state (i.e., S10), it can detect whether there are any anomalies in the business metrics corresponding to the business being processed by the model (i.e., S11). If the metrics are normal, the model continues to process the business to perform the relevant calculations in that business (i.e., return to S10). If the metrics are abnormal, the operational strategy under that business metric is executed (i.e., S12). Then, it is observed whether the business metric is improved after the operational strategy is executed (i.e., S13), such as whether the metric value changes in the direction of reverting to the normal value. If the metric value is improved, the model continues to process the business to perform the relevant calculations in that business (i.e., return to S10); if the metric value is not improved, the operational strategy can be optimized (i.e., S14).
[0126] This application embodiment performs relevant operations on computing power containers based on business indicator values, enabling the rational allocation and efficient utilization of heterogeneous computing power. In this embodiment, business indicators, operation strategies, and triggering conditions can all be flexibly configured when the business is accessed, and the operation strategies and triggering conditions can be updated during business operation. By iteratively updating the operation strategies or triggering conditions, the optimal solution of the operation strategy can be achieved as much as possible, which is conducive to improving the quality of heterogeneous resource services and the efficient utilization of heterogeneous computing power. This improves service flexibility, reduces service deployment and operation costs, and can also alleviate the operational pressure on relevant technical personnel to a certain extent.
[0127] Based on the aforementioned method embodiments, this application also provides a computing resource scheduling device. Specifically, please refer to... Figure 11 , Figure 11 The structure of the device is shown; this device can be mounted on a computer device and used to achieve the above. Figure 4 Some or all of the functions described in the method embodiments. For example... Figure 11 The device may include a configuration unit 1101, a container determination unit 1102, an indicator detection unit 1103, and a resource scheduling unit 1104, wherein:
[0128] Configuration unit 1101 is used to configure corresponding operation detection indicators and resource scheduling strategies under the operation detection indicators for different business scenarios of model processing business on the resource scheduling platform. The model processing business refers to data processing business based on artificial intelligence model.
[0129] The container determination unit 1102 is used to determine the target computing power container running the model processing service from the heterogeneous resource cluster connected to the resource scheduling platform if the model processing service is detected to be running.
[0130] The indicator detection unit 1103 is used to obtain the target business scenario associated with the target computing power container and detect the measured indicator value of the target computing power container under the target operation detection indicator. The target operation detection indicator refers to the operation detection indicator configured by the model processing business under the target business scenario.
[0131] The resource scheduling unit 1104 is used to perform resource scheduling on the computing resources contained in the target computing container based on the measured index value and the target resource scheduling strategy of the target computing container under the target operation detection index.
[0132] In one implementation, after scheduling the computing resources contained in the target computing power container based on the measured index values and the target resource scheduling strategy, the index detection unit 1103 can also be used to perform:
[0133] After a preset time interval, the current index value of the target computing power container under the target operation detection index is obtained, and the expected index value range of the target operation detection index is obtained.
[0134] If the current indicator value is outside the expected indicator value range, then the target resource scheduling strategy is optimized based on the current indicator value, the measured indicator value, and the expected indicator value range.
[0135] In another implementation, the target resource scheduling strategy includes parameter information for at least two resource scheduling parameters; when the indicator detection unit 1103 optimizes the target resource scheduling strategy based on the current indicator value, the measured indicator value, and the expected indicator value range, it may specifically perform the following:
[0136] Select a reference indicator value from the expected indicator value range, and obtain a first difference between the current indicator value and the reference indicator value, and a second difference between the measured indicator value and the reference indicator value;
[0137] From the resource scheduling parameters included in the target resource scheduling strategy, select the resource scheduling parameters to be adjusted, and based on the first difference and the second difference, determine the parameter adjustment direction and parameter adjustment magnitude of the resource scheduling parameters to be adjusted.
[0138] According to the adjustment direction and adjustment magnitude of the parameters, update the parameter information of the resource scheduling parameters to be updated in the target resource scheduling strategy.
[0139] In another implementation, when the configuration unit 1101 configures corresponding operation detection indicators and resource scheduling strategies under the operation detection indicators for different business scenarios of model processing services on the resource scheduling platform, it can specifically execute:
[0140] Receive the service access request of the model processing service, and determine from the service access request the service scenario associated with the model processing service, as well as the operation detection indicators to be configured in the service scenario;
[0141] From the resource scheduling platform, query the resource scheduling strategy that matches the operation detection index;
[0142] Obtain the business processing requirements of the model in the business scenario, and predict the matching degree between the business processing requirements and the queried resource scheduling strategy;
[0143] If the matching degree is greater than the preset matching degree, then the queried resource scheduling strategy is configured under the running detection index.
[0144] In another embodiment, before scheduling the computing resources contained in the target computing power container based on the measured index value and the target resource scheduling strategy, the configuration unit 1101 may also be used to perform:
[0145] Obtain the expected value range of the operational detection indicators;
[0146] Based on the operational monitoring indicators and the expected indicator value range, an indicator anomaly event is created for the operational monitoring indicators. The indicator anomaly event is triggered when the indicator value of the operational monitoring indicators is outside the expected indicator value range.
[0147] Based on the target resource scheduling strategy and the abnormal indicator event, a resource scheduling component is generated and deployed in the target computing power container. The resource scheduling component is used to execute the target resource scheduling strategy when the abnormal indicator event is detected.
[0148] In another implementation, the business scenario includes a model training scenario, and the operation detection metrics under the model training scenario include at least one of resource utilization, the queue length of the training task queue, and the container status.
[0149] The resource scheduling strategy under the resource utilization rate is used to indicate that when the resource utilization rate is less than the preset utilization rate, the number of computing power containers associated with the model training scenario should be reduced.
[0150] The resource scheduling strategy under the queue length of the training task queue is used to indicate that when the queue length is greater than the preset length, based on the model training requirements of the model processing business in the model training scenario, the task priority is configured for each training task contained in the training task queue in the computing power container associated with the model training scenario.
[0151] The resource scheduling strategy under the container state is used to instruct that when the container state is in an abnormal state, breakpoint resume training should be performed based on the training record data generated by the target computing power container.
[0152] In another implementation, if the target business scenario is the model training scenario, and the target operation detection metric includes the container status, the resource scheduling unit 1104 can also be used to execute:
[0153] During the operation of the target computing power container, the training record data generated by the target computing power container is stored in a preset storage space according to a preset storage period. The preset storage space is located outside the target computing power container.
[0154] When scheduling computing resources contained in the target computing power container based on the measured index values and the target resource scheduling strategy, the following can be specifically executed:
[0155] If the measured index value of the container state is used to indicate that the target computing power container is in an abnormal state, then the historical index value of the target computing power container is obtained, and the historical index value includes the index value of the container state in the previous storage cycle.
[0156] When the historical indicator value is used to indicate that the target computing power container is in a normal state, the training record data generated by the target computing power container in the previous storage cycle is obtained from the preset storage space, and a new computing power container is created.
[0157] The acquired training record data is loaded into the new computing power container, and the model processing service under the model training scenario is run in the new computing power container based on the loaded training record data.
[0158] In another implementation, the business scenario includes a model inference scenario, and the operational monitoring metrics under the model inference scenario include at least one of the execution success rate of the inference task and the container load.
[0159] The resource scheduling strategy under the execution success rate is used to indicate that when the execution success rate is less than the preset success rate, at least one computing power container associated with the model inference scenario is rebuilt.
[0160] The resource scheduling strategy under container load is used to indicate that when a container is overloaded, the number of computing power containers associated with the model inference scenario should be increased.
[0161] In another implementation, if the target business scenario is the model inference scenario, and the target operation detection metric includes the container load, the resource scheduling unit 1104, when scheduling the computing resources contained in the target computing container based on the measured metric value and the target resource scheduling strategy, may specifically execute:
[0162] If the measured index value of the container load is used to indicate container overload, then at least one new computing power container is created.
[0163] The model is deployed in each new computing power container to process the service, so as to obtain at least one new target computing power container.
[0164] In one embodiment, Figure 11 Each unit in the illustrated device can be individually or entirely combined into one or more other units, or some of the units can be further divided into multiple functionally smaller units. This achieves the same operation without affecting the technical effects of the embodiments of this application. In other words, the above units are based on logical functional division. In practical applications, the function of one unit can be implemented by multiple units, or the function of multiple units can be implemented by one unit. In other embodiments of this application, Figure 11 The device shown may also include other units, and in practical applications, these functions may also be implemented with the assistance of other units, and may be implemented by multiple units working together.
[0165] According to another embodiment of this application, the following can be executed by running on a computing device including processing elements and storage elements such as a central processing unit (CPU), random access memory (RAM), and read-only memory (ROM). Figure 4 The computer program (including program code) for each step involved in the corresponding method shown, to construct such... Figure 11 The apparatus shown. A computer program may be recorded on, for example, a computer-readable storage medium, loaded onto the apparatus via the computer-readable storage medium, and executed therein.
[0166] Based on the descriptions of the above method and device embodiments, this application also provides a computer device in which a resource scheduling platform can run. Specifically, please refer to... Figure 12 , Figure 12 This application provides a schematic diagram of the structure of a computer device, as shown in the embodiment of the present application. Figure 12 As shown, the computer device may include a processor 1201, a memory 1202, and a communication interface 1203, and the processor 1201, the memory 1202, and the communication interface 1203 may be connected by a bus or other means.
[0167] The processor 1201 (or Central Processing Unit, CPU) is the computing and control core of a computer device. It can parse various instructions within the computer device and process various data. For example, the CPU can be used to parse service access requests and control the computer device to execute corresponding policy configuration tasks; the CPU can also transmit program instructions and other information between internal structures of the computer device, and so on.
[0168] Memory 1202 is a memory device in a computer device used to store programs and data. It is understood that memory 1202 here can include both the computer device's internal memory and any extended memory supported by the computer device.
[0169] The communication interface 1203 may optionally include a standard wired interface or a wireless interface (such as Wi-Fi, mobile communication interface, etc.), and can be used to send and receive data under the control of the processor 1201; the communication interface 1203 can also be used for the transmission and interaction of data within the computer device.
[0170] In a specific embodiment, the processor 1201 may load and execute one or more computer programs stored in the memory 1202 to implement the steps in the method described in the above embodiments.
[0171] This application embodiment also provides a computer-readable storage medium (Memory), which is a memory device in a computer device for storing programs and data. It is understood that the computer-readable storage medium here can include both the built-in storage medium in the computer device and extended storage media supported by the computer device. The computer-readable storage medium provides storage space that stores the processing system of the computer device. Furthermore, the storage space also stores one or more instructions suitable for loading and execution by the processor 1201, which may be one or more computer programs (including program code). It should be noted that the computer-readable storage medium here may be high-speed RAM memory or non-volatile memory, such as at least one disk storage device; optionally, it may also be at least one readable storage medium located remotely from the aforementioned processor.
[0172] This application also provides a computer program product, which includes computer instructions, and the processor of a computer device can execute the above-described method embodiments by loading the computer instructions.
[0173] Based on the same inventive concept, the principles and beneficial effects of the devices, computer equipment, computer-readable storage media and computer program products provided in the embodiments of this application are similar to those described in the foregoing corresponding method embodiments. Therefore, the corresponding method implementation principles and beneficial effects can be referred to, which will not be repeated here for the sake of brevity.
[0174] It should be further noted that the steps in the methods of this application embodiment can be adjusted, merged, and deleted according to actual needs, and the modules in the device of this application embodiment can be merged, divided, and deleted according to actual needs. Furthermore, those skilled in the art will understand that all or part of the steps in the various methods of the above embodiments can be implemented by a program instructing related hardware. It should also be particularly emphasized that when the above embodiments of this application are applied to specific products or technologies, the data acquisition involved in each specific implementation of this application requires the permission or consent of the relevant parties, and the collection, use, and processing of related data must comply with the relevant laws, regulations, and standards of the relevant countries and regions.
[0175] The above-disclosed embodiments are merely preferred embodiments of this application and should not be construed as limiting the scope of this application. Those skilled in the art will understand that all or part of the processes for implementing the above embodiments and equivalent variations made in accordance with the claims of this application are still within the scope of this application.
Claims
1. A method for scheduling computing resources, characterized in that, The method includes: On the resource scheduling platform, corresponding operation detection indicators and resource scheduling strategies under the operation detection indicators are configured for different business scenarios of model processing business. The model processing business refers to data processing business based on artificial intelligence model. If the model processing service is detected to be running, the target computing power container running the model processing service is determined from the heterogeneous resource cluster connected to the resource scheduling platform. Obtain the target business scenario associated with the target computing power container, and detect the measured index value of the target computing power container under the target operation detection index. The target operation detection index refers to the operation detection index configured by the model processing business under the target business scenario. Based on the measured index values and the target resource scheduling strategy of the target computing power container under the target operation detection index, resource scheduling is performed on the computing power resources contained in the target computing power container.
2. The method according to claim 1, characterized in that, After scheduling the computing resources contained in the target computing power container based on the measured index values and the target resource scheduling strategy, the method further includes: After a preset time interval, the current index value of the target computing power container under the target operation detection index is obtained, and the expected index value range of the target operation detection index is obtained. If the current indicator value is outside the expected indicator value range, then the target resource scheduling strategy is optimized based on the current indicator value, the measured indicator value, and the expected indicator value range.
3. The method according to claim 2, characterized in that, The target resource scheduling strategy includes parameter information for at least two resource scheduling parameters; the strategy optimization based on the current indicator value, the measured indicator value, and the expected indicator value range includes: Select a reference indicator value from the expected indicator value range, and obtain a first difference between the current indicator value and the reference indicator value, and a second difference between the measured indicator value and the reference indicator value; From the resource scheduling parameters included in the target resource scheduling strategy, select the resource scheduling parameters to be adjusted, and based on the first difference and the second difference, determine the parameter adjustment direction and parameter adjustment magnitude of the resource scheduling parameters to be adjusted. According to the adjustment direction and adjustment magnitude of the parameters, update the parameter information of the resource scheduling parameters to be updated in the target resource scheduling strategy.
4. The method according to claim 1, characterized in that, The configuration of corresponding operational monitoring indicators and resource scheduling strategies under these indicators on the resource scheduling platform for different business scenarios of model processing includes: Receive the service access request of the model processing service, and determine from the service access request the service scenario associated with the model processing service, as well as the operation detection indicators to be configured in the service scenario; From the resource scheduling platform, query the resource scheduling strategy that matches the operation detection index; Obtain the business processing requirements of the model in the business scenario, and predict the matching degree between the business processing requirements and the queried resource scheduling strategy; If the matching degree is greater than the preset matching degree, then the queried resource scheduling strategy is configured under the running detection index.
5. The method according to claim 1, characterized in that, Before scheduling the computing resources contained in the target computing power container based on the measured index values and the target resource scheduling strategy, the method further includes: Obtain the expected value range of the operational detection indicators; Based on the operational monitoring indicators and the expected indicator value range, an indicator anomaly event is created for the operational monitoring indicators. The indicator anomaly event is triggered when the indicator value of the operational monitoring indicators is outside the expected indicator value range. Based on the target resource scheduling strategy and the abnormal indicator event, a resource scheduling component is generated and deployed in the target computing power container. The resource scheduling component is used to execute the target resource scheduling strategy when the abnormal indicator event is detected.
6. The method according to claim 1, characterized in that, The business scenario includes a model training scenario, and the operation detection metrics under the model training scenario include at least one of resource utilization, training task queue length, and container status. The resource scheduling strategy under the resource utilization rate is used to indicate that when the resource utilization rate is less than the preset utilization rate, the number of computing power containers associated with the model training scenario should be reduced. The resource scheduling strategy under the queue length of the training task queue is used to indicate that when the queue length is greater than the preset length, based on the model training requirements of the model processing business in the model training scenario, the task priority is configured for each training task contained in the training task queue in the computing power container associated with the model training scenario. The resource scheduling strategy under the container state is used to instruct that when the container state is in an abnormal state, breakpoint resume training should be performed based on the training record data generated by the target computing power container.
7. The method according to claim 6, characterized in that, If the target business scenario is the model training scenario, and the target runtime detection metric includes the container state, the method further includes: During the operation of the target computing power container, the training record data generated by the target computing power container is stored in a preset storage space according to a preset storage period. The preset storage space is located outside the target computing power container. The step of scheduling computing resources contained in the target computing power container based on the measured index value and the target resource scheduling strategy includes: If the measured index value of the container state is used to indicate that the target computing power container is in an abnormal state, then the historical index value of the target computing power container is obtained, and the historical index value includes the index value of the container state in the previous storage cycle. When the historical indicator value is used to indicate that the target computing power container is in a normal state, the training record data generated by the target computing power container in the previous storage cycle is obtained from the preset storage space, and a new computing power container is created. The acquired training record data is loaded into the new computing power container, and the model processing service under the model training scenario is run in the new computing power container based on the loaded training record data.
8. The method according to claim 1, characterized in that, The business scenario includes a model inference scenario, and the operational monitoring metrics under the model inference scenario include at least one of the execution success rate of the inference task and the container load. The resource scheduling strategy under the execution success rate is used to indicate that when the execution success rate is less than the preset success rate, at least one computing power container associated with the model inference scenario is rebuilt. The resource scheduling strategy under container load is used to indicate that when a container is overloaded, the number of computing power containers associated with the model inference scenario should be increased.
9. The method according to claim 8, characterized in that, If the target business scenario is the model inference scenario, and the target operation detection metric includes the container load, the step of scheduling the computing resources contained in the target computing container based on the measured metric value and the target resource scheduling strategy includes: If the measured index value of the container load is used to indicate container overload, then at least one new computing power container is created. The model is deployed in each new computing power container to process the service, so as to obtain at least one new target computing power container.
10. A computing resource scheduling device, characterized in that, include: The configuration unit is used to configure corresponding operation detection indicators and resource scheduling strategies under the operation detection indicators for different business scenarios of model processing business on the resource scheduling platform. The model processing business refers to data processing business based on artificial intelligence model. The container determination unit is used to determine the target computing power container running the model processing service from the heterogeneous resource cluster connected to the resource scheduling platform if the model processing service is detected to be running. The indicator detection unit is used to obtain the target business scenario associated with the target computing power container and detect the measured indicator value of the target computing power container under the target operation detection indicator. The target operation detection indicator refers to the operation detection indicator configured by the model processing business under the target business scenario. The resource scheduling unit is used to schedule the computing resources contained in the target computing container based on the measured index value and the target resource scheduling strategy of the target computing container under the target operation detection index.
11. A computer device, characterized in that, include: A memory, wherein a computer program is stored; A processor for loading the computer program to implement the method as described in any one of claims 1-9.
12. A computer-readable storage medium, characterized in that, The computer storage medium stores a computer program adapted to be loaded by a processor and executed as described in any one of claims 1-9.
13. A computer program product, characterized in that, The computer program product includes computer instructions, and the processor of the computer device reads the computer instructions and executes the method as described in any one of claims 1-9.