Resource creation method and device for model service, equipment and medium

By converting the deployment requirements of model service into standardized service templates and generating resource scheduling instructions, the problem of model service resource creation relying on manual operations is solved, automated resource creation and unified operation and maintenance are realized, and deployment efficiency and resource management capabilities are improved.

CN120447913APending Publication Date: 2025-08-08JINAN INSPUR DATA TECH CO LTD
View PDF 0 Cites 3 Cited by

Patent Information

Application Number
CN202510564647.5
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-30
Publication Date
2025-08-08

AI Technical Summary

Technical Problem

In the prior art, model service resource creation is heavily dependent on manual operations, resulting in inefficient deployment across resource types, difficulty in quickly responding to enterprise business needs, and lack of unified resource management and monitoring mechanisms.

Method used

Convert the model service deployment requirements entered by users into standardized service templates, analyze the template to generate resource scheduling instructions, and use the target driver to create cloud resources in the target resource pool to realize automated resource creation and unified operation and maintenance monitoring.

Benefits of technology

Significantly shorten the deployment cycle, reduce labor costs, improve delivery efficiency, and enable enterprises to flexibly and efficiently deploy AI model services on private cloud platforms to adapt to rapidly changing business needs.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120447913A_ABST
    Figure CN120447913A_ABST
Patent Text Reader

Abstract

The invention discloses a resource creation method, device and equipment for model service and a medium, and relates to the technical field of computers, and the method comprises the steps: converting a model service deployment demand input by a user into a standardized service template; analyzing the standardized service template to obtain a model service resource scheduling instruction, the model service resource scheduling instruction being used for indicating a target resource pool; and determining a target driver corresponding to a target resource pool indicated by the model service resource scheduling instruction, so as to create cloud resources of the model service in the target resource pool by using the target driver, thereby realizing automatic resource creation of the model service, greatly shortening the deployment period through full-process automatic operation, reducing the labor cost, and improving the deployment efficiency. The heterogeneous resource management process is effectively integrated, the delivery efficiency is remarkably improved, and an enterprise can flexibly and efficiently deploy various AI model services on a private cloud platform and deal with quickly changing business requirements leisurely.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of computer technology, and in particular to a resource creation method, apparatus, device, and medium for a model service. Background Art

[0002] With the development of artificial intelligence (AI) technology, the demand for deploying a large number of AI model services in enterprise private cloud environments is growing. These model services have widely varying computing resource requirements and may require adaptation to a variety of heterogeneous resource carriers, including virtual machine clusters, Kubernetes container environments, and bare metal servers.

[0003] Currently, the resource creation method for model services in related technologies mainly relies on manual operations, that is, operation and maintenance personnel need to plan resources, build the operating environment, and configure parameters for each model service one by one. In addition, different types of resources and model services need to be managed through independent monitoring tools. The whole process is cumbersome and the process is fragmented.

[0004] It can be seen that the resource creation of model services in related technologies relies heavily on manual operations. When it comes to cross-resource type deployment, it often takes a lot of time and manpower costs, the delivery efficiency is extremely low, and it is difficult to meet the rapidly changing business needs of enterprises. Summary of the Invention

[0005] The present application provides a resource creation method, apparatus, equipment and medium for model services, in order to at least solve the problems in related technologies such as heavy reliance on manual labor for model service resource creation, inefficient cross-resource deployment, and delayed delivery. The method converts the model service deployment requirements input by the user into a standardized service template; parses the standardized service template to obtain a model service resource scheduling instruction, which is used to indicate a target resource pool; determines the target driver corresponding to the target resource pool indicated by the model service resource scheduling instruction, and uses the target driver to create cloud resources for the model service in the target resource pool, thereby realizing automated resource creation for the model service. The full-process automated operation greatly shortens the deployment cycle, reduces labor costs, effectively integrates heterogeneous resource management processes, and significantly improves delivery efficiency, enabling enterprises to flexibly and efficiently deploy various AI model services on private cloud platforms and calmly respond to rapidly changing business needs.

[0006] This application provides a resource creation method for a model service, including:

[0007] Convert the model service deployment requirements entered by the user into a standardized service template;

[0008] Parse the standardized service template to obtain the model service resource scheduling instruction, which is used to indicate the target resource pool;

[0009] Determine the target driver corresponding to the target resource pool, and use the target driver to create cloud resources for the model service in the target resource pool.

[0010] This application also provides a resource creation device for a model service, including:

[0011] The conversion unit is used to convert the model service deployment requirements input by the user into a standardized service template;

[0012] A parsing unit, configured to parse the standardized service template to obtain a model service resource scheduling instruction, wherein the model service resource scheduling instruction is used to indicate a target resource pool;

[0013] The deployment unit is used to determine the target driver corresponding to the target resource pool, so as to use the target driver to create cloud resources for the model service in the target resource pool.

[0014] The present application also provides an electronic device, comprising: a memory for storing a computer program; and a processor for implementing the steps of the resource creation method of any of the above-mentioned model services when executing the computer program.

[0015] The present application also provides a computer-readable storage medium, in which a computer program is stored, wherein when the computer program is executed by a processor, the steps of the resource creation method of any of the above-mentioned model services are implemented.

[0016] The present application also provides a computer program product, including a computer program, which implements the steps of the resource creation method of any of the above-mentioned model services when the computer program is executed by a processor.

[0017] Through this application, the model service deployment requirements input by the user are converted into a standardized service template; the standardized service template is parsed to obtain the model service resource scheduling instruction, which is used to indicate the target resource pool; the target driver corresponding to the target resource pool indicated by the model service resource scheduling instruction is determined, and the target driver is used to create cloud resources for the model service in the target resource pool, thereby realizing automated resource creation for the model service. The full-process automated operation greatly shortens the deployment cycle, reduces labor costs, effectively integrates heterogeneous resource management processes, and significantly improves delivery efficiency, enabling enterprises to flexibly and efficiently deploy various AI model services on private cloud platforms and calmly respond to rapidly changing business needs. BRIEF DESCRIPTION OF THE DRAWINGS

[0018] In order to more clearly illustrate the embodiments of the present application, the following is a brief introduction to the drawings required for use in the embodiments. Obviously, the drawings described below are only some embodiments of the present application. For ordinary technicians in this field, other drawings can be obtained based on these drawings without any creative work.

[0019] Figure 1 A flowchart of a resource creation method for a model service provided in an embodiment of the present application;

[0020] Figure 2 A flowchart of a resource scenario method for another model service provided in an embodiment of the present application;

[0021] Figure 3 A schematic diagram of creating cloud resources in a target resource pool through a heterogeneous resource plug-in module provided in an embodiment of the present application;

[0022] Figure 4 A schematic diagram of a resource creation architecture for a specific model service provided in an embodiment of the present application;

[0023] Figure 5 A schematic diagram of the structure of a resource creation device for a model service provided in an embodiment of the present application. DETAILED DESCRIPTION

[0024] The following will be combined with the accompanying drawings in the embodiments of this application to clearly and completely describe the technical solutions in the embodiments of this application. Obviously, the embodiments described are only part of the embodiments of this application, not all of them. Based on the embodiments in this application, all other embodiments obtained by ordinary technicians in this field without making creative efforts are within the scope of protection of this application.

[0025] It should be noted that, in the description of this application, the terms "comprises," "includes," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or device comprising a series of elements includes not only those elements, but also other elements not explicitly listed, or elements inherent to such process, method, article, or device. The terms "first," "second," etc., in this application are used to distinguish similar objects, and are not used to describe a particular order or sequence.

[0026] In enterprise private cloud environments, with the development of artificial intelligence technology, the demand for deploying a large number of AI model services (such as TensorFlow and PyTorch models) is increasing. Since these model services have very different requirements for computing resources, they may need to be deployed on different resource carriers such as virtual machine clusters, Kubernetes container environments, or bare metal servers. Traditional deployment methods mainly rely on manual configuration. When adapting heterogeneous resources such as virtual machines, containers, and bare metal to different model services, separate deployment processes need to be designed, and there is a lack of a unified mechanism for resource scheduling. In specific operations, operation and maintenance personnel need to manually plan resources, install the environment, and configure parameters for each model service, and manage different types of resources and model services separately through their own independent monitoring tools. The entire process is cumbersome and the process is fragmented.

[0027] Currently, the model service deployment methods used in related technologies have significant flaws. First, resource adaptation is complex. Due to the lack of unified standards, the deployment process is fragmented when different model services adapt to heterogeneous resources, resulting in difficult deployment and high error rates. Second, elasticity is insufficient. The lack of a unified resource orchestration mechanism makes it impossible to dynamically adjust resources according to changes in model service load, which can easily lead to resource waste or insufficient service performance. Third, delivery efficiency is low, with heavy reliance on manual configuration. Deployment across resource types is time-consuming and difficult to quickly respond to enterprise business needs. Fourth, operation and maintenance monitoring is fragmented. The monitoring systems for virtual machines, containers, bare metal machines, and model resources are independent of each other and lack a unified view. This makes it difficult for operation and maintenance personnel to fully grasp the system's operating status, increasing the difficulty of troubleshooting and performance optimization.

[0028] In order to solve the technical problems existing in the relevant technologies, the embodiments of the present application provide a resource creation method for a model service, which converts the model service deployment requirements input by the user into a standardized service template, and then parses the template to generate resource scheduling instructions and determines the target resource pool and the corresponding driver to realize the automatic creation of cloud resources. This method adopts a standardized and low-coupling design, which effectively solves the drawbacks of the resource deployment method in the relevant technologies. Through standardized service templates, the deployment specifications are unified and the resource adaptation process is simplified; based on the instruction-driven resource scheduling and target-driven creation mechanism, the intelligent orchestration and dynamic expansion and contraction of resources are realized; the full-process automated operation greatly improves the delivery efficiency and reduces manual intervention; at the same time, it lays the foundation for the subsequent establishment of a unified operation and maintenance monitoring system, and ultimately realizes the efficient and flexible deployment, resource management and monitoring operation and maintenance of model services in the private cloud platform, thereby improving the overall efficiency and competitiveness of enterprise AI applications.

[0029] In order to enable those skilled in the art to better understand the present application, the present application is further described in detail below with reference to the accompanying drawings and specific implementation methods.

[0030] Figure 1A flowchart of a resource creation method for a model service provided in an embodiment of the present application.

[0031] like Figure 1 As shown, the method comprises the following steps:

[0032] Step 101: Convert the model service deployment requirements input by the user into a standardized service template.

[0033] In an embodiment of the present application, in a scenario where an enterprise uses a cloud management platform to deploy AI model services, users have the need to deploy model services. These needs may include the type of model (such as TensorFlow, PyTorch model, etc.), specific requirements for computing resources (such as how much CPU and memory are required, whether to use a GPU, and what specifications of GPU to use, etc.), and expected service quality (SLA level, such as network bandwidth, latency requirements, etc.).

[0034] The cloud management platform's model service interaction layer is responsible for receiving user input for model service deployment requirements. The model service standardization engine predefines atomic capabilities for model service lifecycle management. These capabilities cover the fundamental operations of each key step in the model service lifecycle, from initial deployment, possible scaling during operation, version updates, and final destruction.

[0035] When the model service standardization engine receives user model service deployment requirements from the cloud management platform's model service interaction layer, it converts these requirements into standardized service templates according to specific rules and formats. This standardized service template is a unified and standardized expression that presents complex and diverse user requirements in a form that the system can clearly understand and process, facilitating subsequent operations and processing.

[0036] Step 102: Parse the standardized service template to obtain a model service resource scheduling instruction, where the model service resource scheduling instruction is used to indicate a target resource pool.

[0037] In the embodiment of the present application, after obtaining the standardized service template in step 101, the template needs to be analyzed and interpreted. The intelligent orchestration and scheduling module is responsible for parsing the standardized service template.

[0038] During the parsing process, the intelligent orchestration and scheduling module analyzes the various model service configuration parameters in the standardized service template, such as the model's resource requirements (whether virtual machines, containers, or bare metal resources are required, and what resource specifications are required), and service quality requirements. Based on these model service configuration parameters, it applies specific algorithms and rules to generate resource scheduling instructions for the model service.

[0039] This model service resource scheduling directive specifies the target resource pool to which the model service should be assigned. This target resource pool can be a collection of different types of resources, such as a virtual machine cluster, a Kubernetes container environment, or a bare metal server. In other words, the model service resource scheduling directive determines the specific resource environment in which the model service will run, providing guidance for subsequent resource creation.

[0040] Step 103: Determine the target driver corresponding to the target resource pool, and use the target driver to create cloud resources for the model service in the target resource pool.

[0041] In an embodiment of the present application, after obtaining the model service resource scheduling instruction and clarifying the target resource pool, the present application can actually create the cloud resources required for the model service in the target resource pool according to the heterogeneous resource driver plug-in module. The heterogeneous resource plug-in layer can provide a variety of drivers, such as virtual machine API driver, K8sCRD driver, bare metal IPMI driver, etc. These drivers correspond to different types of resources (virtual machines, containers, bare metal).

[0042] Based on the type of target resource pool (e.g., a virtual machine resource pool or a container resource pool), the heterogeneous resource plug-in layer determines the corresponding target driver. This target driver is then used to perform a series of operations in the target resource pool to create the cloud resources required for the model service. For example, if the target resource pool is a virtual machine resource pool, the virtual machine API driver is used to create a virtual machine instance with a specific configuration as the carrier for the model service.

[0043] In addition, the present application can also obtain the operating indicators of at least one resource pool through a unified operation and maintenance monitoring module; determine whether the operating indicators in at least one resource pool meet the preset indicator range; if the first operating indicator in the first resource pool is not within the first preset indicator range, determine the preset operation corresponding to the first operating indicator to execute the preset operation corresponding to the first operating indicator.

[0044] The unified operation and maintenance monitoring module obtains real-time operational indicators of at least one resource pool (which can be a virtual machine resource pool, a container resource pool, a bare metal resource pool, etc.). These operational indicators include but are not limited to CPU utilization, memory utilization, GPU utilization, network bandwidth, and latency.

[0045] The unified operation and maintenance monitoring module then compares these acquired operating indicators with the preset indicator range to determine whether the operating indicators of each resource pool meet the requirements. For example, the preset normal range of CPU utilization is between 20% and 80%. If the CPU utilization (i.e., the first operating indicator) of a resource pool (here referred to as the first resource pool) is outside this range, for example, exceeding 80% (i.e., CPU overload), then it indicates that the operating status of the resource pool has become abnormal.

[0046] Once the first operating metric in the first resource pool is detected to be outside the first preset range, the unified operation and maintenance monitoring module determines the default action for this abnormal operating metric based on pre-defined rules. For example, if CPU utilization is excessively high, the default action might be to expand resources (i.e., immediately triggering the horizontal autoscaling (HPA) mechanism to add resources and logging the decision-making process for subsequent review and analysis, such as increasing the number of CPU cores in a virtual machine or increasing the number of containers). If network latency is excessively high, the default action might be to optimize the network configuration. If a bare metal node failure is detected, the default action might be to migrate the service to a backup resource with equivalent conditions and issue a timely alert to operations personnel. Finally, the unified operation and maintenance monitoring module executes the default action corresponding to this first operating metric to adjust the resource pool's operating status back to the normal operating range, ensuring the stable operation of the model service. Furthermore, the unified operation and maintenance monitoring module aggregates and integrates utilization data from all resources and feeds it back to the intelligent orchestration and scheduling module, providing a powerful basis for optimizing subsequent resource scheduling strategies and improving overall resource utilization efficiency.

[0047] In summary, according to the resource creation method of the model service of the present application, the model service deployment requirements input by the user are converted into a standardized service template; the standardized service template is parsed to obtain the model service resource scheduling instruction, and the model service resource scheduling instruction is used to indicate the target resource pool; the target driver corresponding to the target resource pool indicated by the model service resource scheduling instruction is determined, so as to use the target driver to create cloud resources for the model service in the target resource pool, thereby realizing automated resource creation for the model service, greatly shortening the deployment cycle with full-process automated operation, reducing labor costs, effectively integrating heterogeneous resource management processes, and significantly improving delivery efficiency, so that enterprises can flexibly and efficiently deploy various AI model services on private cloud platforms and calmly respond to rapidly changing business needs.

[0048] based on Figure 1 The embodiment shown, Figure 2 The following further illustrates a flowchart of another resource creation method for a model service proposed in this application. Figure 2 The following steps may be included.

[0049] Step 201 : Based on the predefined lifecycle management atomic capabilities, the model service deployment requirements are converted into a standard service template according to a preset template.

[0050] In this application, the standard service template includes the model service configuration parameters corresponding to the model service, and the lifecycle management atomic capability is used to represent the set of basic operation units in the entire lifecycle of the model service.

[0051] Specifically, lifecycle management atomic capabilities are a collection of predefined basic operational units that cover every key step in the model service lifecycle, from deployment to destruction. These include deployment and rollout, dynamic resource scaling (scaling) based on business needs, service version updates, and service termination (destruction). These atomic capabilities are the foundation of model service management and provide a clear operational framework for subsequent template conversion and service deployment.

[0052] A preset template is a pre-set format specification that specifies what content a standard service template should contain and how to organize it. With the help of preset templates, it can be ensured that the converted standard service template has a unified structure and format, which is convenient for the system to perform subsequent processing and parsing. The preset template in this application can specifically be a YAML / JSON template. The YAML / JSON template can define the resource requirements, dependencies, and SLA levels of the model service, support parameterized predefined configurations (model service type, resource type used, GPU model, storage type, etc.), and support on-demand parameter expansion.

[0053] A standard service template includes the model service configuration parameters corresponding to the model service. These configuration parameters describe in detail the various conditions required for the model service to run, such as resource requirements, dependencies, and service-level agreements (SLAs). Standard service templates serve as an important basis for subsequent resource scheduling and service deployment.

[0054] In an embodiment of the present application, before performing template conversion, the present application also needs to perform a compliance check on the model service deployment requirements, specifically including: based on preset rules, performing a compliance check on the model service deployment requirements to determine whether the model service deployment requirements are compliant; if the model service deployment requirements are compliant, the model service deployment requirements are converted into a standardized service template.

[0055] Specifically, this application requires a strict compliance check before converting the model service deployment requirements into a standard service template. This step comprehensively reviews the model service deployment requirements submitted by users based on preset rules to ensure the rationality and legality of the requirements. For example, it checks whether there are any security group conflicts and whether the resource requirements exceed the system's carrying capacity. Only when the model service deployment requirements pass the compliance check and are determined to be compliant will the next step of the conversion process be entered.

[0056] Once model service deployment requirements are determined to be compliant, the model service standardization engine converts these requirements into standard service templates based on pre-set templates. During this conversion process, the service definition standardization feature is leveraged to provide pre-set templates (e.g., YAML / JSON templates) to define in detail the model service's resource requirements, dependencies, SLA levels, and other information. Furthermore, parameterized predefined configurations are supported, allowing users to pre-define parameters such as the model service type, resource types used, GPU model, and storage type based on actual conditions. Parameter expansion is also supported on demand to meet individual user needs. Finally, validation and conversion are performed on the user-submitted configurations, further checking and converting them into standardized service templates (e.g., Intermediate Representation (IR), a unified, standardized format that masks the diversity and complexity of the original configuration, allowing demand information from different sources and formats to be understood and processed in a consistent manner by subsequent modules), facilitating subsequent processing and scheduling.

[0057] Step 202 : parsing the standardized service template, and screening all resource pools according to the model service configuration parameters in the standardized service template to obtain at least one candidate resource pool that matches the model service configuration parameters.

[0058] In some embodiments, the present application can parse the standardized service template through the intelligent orchestration module to obtain the model service configuration parameters in the standard service template; based on the model service configuration parameters, at least one candidate resource pool that matches the model service configuration parameters is screened from all resource pools.

[0059] Model service configuration parameters can include various conditions required for model service operation, such as resource type (whether containers, virtual machines, or bare metal resources are required), whether a GPU is required, GPU memory requirements, network bandwidth requirements, latency requirements, and scaling strategies. Taking the YAML-formatted model service definition file given above as an example, fields such as "resource_type" (specifying containerized deployment, virtual machines, or bare metal deployment), "gpu" (whether GPU nodes must be used), and "sla" (model scheduling SLA requirements, such as GPU memory, network bandwidth, and network latency) are all important configuration parameters.

[0060] Based on the extracted model service configuration parameters, the intelligent orchestration and scheduling module further screens all available resource pools. It checks each resource pool individually to see if it meets the model service configuration requirements. For example, if the model service configuration parameters explicitly require the use of container resources, virtual machine and bare metal resource pools are filtered out. If GPU nodes are required, resource pools without GPUs are excluded. In this way, at least one candidate resource pool that matches the model service configuration parameters is selected from all resource pools. These candidate resource pools serve as the basis for further determination of the target resource pool, and all have the potential to become the resource environment that hosts the model service.

[0061] Step 203: Score at least one candidate resource pool, and determine a target resource pool based on the evaluation result of at least one candidate resource pool to generate a model service resource scheduling instruction.

[0062] In some embodiments, the present application can parse the model service standardization scheduling template issued by the model service standardization engine through the intelligent orchestration scheduling module, calculate the target resource pool according to the scheduling algorithm, convert it into instructions that can be recognized by the heterogeneous resource driver plug-in module, and issue it to the heterogeneous resource driver plug-in module, and the heterogeneous resource driver plug-in module completes the creation of the corresponding cloud resources. Specifically, it includes: scoring at least one candidate resource pool according to the resource utilization data corresponding to at least one candidate resource pool to obtain at least one scoring result; selecting a target scoring result from at least one scoring result, and determining the resource pool corresponding to the target scoring result as the target resource pool; generating a model service resource scheduling instruction according to the target resource pool.

[0063] Specifically, the intelligent orchestration and scheduling module may score each candidate resource pool according to resource utilization data corresponding to at least one candidate resource pool.

[0064] The specific scoring formula is:

[0065] CpuUsed represents the CPU utilization of each candidate resource pool, MemUsed represents the memory utilization of each candidate resource pool, GpuUsed represents the GPU utilization of each candidate resource pool, and IsIronic indicates whether it is a physical node. α, β, and γ are the corresponding weights.

[0066] The weights α, β, and γ in this application can be dynamically adjusted according to actual conditions to reflect the importance of different resource utilizations in the scoring, and are not limited in the embodiments of this application.

[0067] δIsIronic means that when the candidate resource pool being calculated is a physical machine resource pool, a weight score is added to increase the score of the physical machine resource pool.

[0068] This application can use this scoring formula to calculate each candidate resource pool and obtain a scoring result for each candidate resource pool. From the at least one scoring result obtained, a target scoring result is selected, and the resource pool corresponding to the target scoring result is determined as the target resource pool. Generally speaking, the resource pool with the highest scoring result is selected as the target resource pool because it best meets the requirements of the model service in terms of resource utilization and other aspects, and can provide a better operating environment for the model service.

[0069] After determining the target resource pool, the intelligent orchestration and scheduling module generates a model service resource scheduling instruction based on the target resource pool. This instruction includes detailed information about the target resource pool and instructions for the heterogeneous resource driver plug-in module to create cloud resources for the model service in the target resource pool.

[0070] In the process of scoring and determining the target resource pool, this application will further filter based on other requirements in the scheduling template. For example, when there is a GPU demand, GPU bare metal nodes will be allocated first; according to the SLA requirements in the template (such as GPU memory ≥ xGB, network bandwidth ≥ yGbps, network latency ≥ z ms), the target resource pool will be screened here to ensure that the target resource pool finally selected can meet the high-quality operation requirements of the model service.

[0071] In addition, the present application also includes elastic strategies, specifically including: parsing the standardized service template to obtain the model service configuration parameters in the standard service template; determining whether the model service configuration parameters meet the preset conditions, and if the model service configuration parameters meet the preset conditions, determining the preset operations corresponding to the preset conditions to generate corresponding model service resource scheduling instructions based on the preset operations.

[0072] The specific elasticity strategies are shown in Table 1:

[0073] Table 1

[0074]

[0075] Referring to Table 1, if CPU utilization exceeds 80% for 5 minutes (a trigger for vertical scaling), the VM specification is upgraded from 4vCPUs to 8vCPUs (an example of a vertical scaling action). If the service QPS exceeds the threshold and latency exceeds 200ms (a trigger for horizontal scaling), three Pod replicas are added via Kubernetes HPA (an example of a horizontal scaling action). If a bare metal node fails (a trigger for cross-resource migration), the service is migrated to a VM cluster with equivalent GPUs (an example of a cross-resource migration action). These elastic policies enable the model service to automatically adjust resource allocation based on actual conditions during operation, improving service stability and performance.

[0076] Step 204 : Determine a target driver corresponding to the target resource pool according to the target resource pool indicated by the model service resource scheduling.

[0077] Step 205: call the image file corresponding to the target driver.

[0078] Step 206: Create cloud resources for the model service in the target resource pool according to the image file and the target driver.

[0079] In an embodiment of the present application, the heterogeneous resource plug-in module of the present application can provide virtual machine API driver, K8sCRD driver, and bare metal IPMI driver, shielding the differences between different resources in the private cloud and providing a unified and abstract resource scheduling atomic capability for the upper-level intelligent orchestration and scheduling module.

[0080] The heterogeneous resource driver plug-in module determines the corresponding target driver according to the type of the target resource pool. This is because different types of resource pools require different drivers for operation. For example, a virtual machine resource pool requires a virtual machine API driver for management and operation, a Kubernetes container resource pool requires a K8s CRD driver, and a bare metal resource pool requires a bare metal IPMI driver. Figure 3 As shown, the present application provides a specific schematic diagram of creating cloud resources in a target resource pool through a heterogeneous resource plug-in module. Figure 3 , heterogeneous resource plug-in modules (i.e. Figure 3 The heterogeneous resource plug-in layer adapter in the model service receives the model service resource scheduling instructions issued by the intelligent orchestration and scheduling module, including resource creation / modification / startup / shutdown / deletion instructions, and selects the appropriate target resource pool (virtual machine, container, bare metal, etc.) according to the configuration in the model service resource scheduling instructions (including resource type, resource specification and other parameters) to create the corresponding model service carrier cloud resource instance.

[0081] After determining the target driver, the Heterogeneous Resource Driver plug-in module needs to obtain the image file used to create the model service cloud resource. This can be obtained from the hybrid image repository. The hybrid image repository is used to store various image files, including virtual machine images, container images, bare metal system images, and model weight files. Based on the target driver type, the Heterogeneous Resource Driver plug-in module calls the corresponding image file from the hybrid image repository.

[0082] For example, if the target driver is a Kubernetes Custom Resource Definition (CRD) driver for container deployment, the pre-configured model service Docker image is pulled from the repository. If the target driver is a bare metal driver, an ISO image containing a customized kernel is transferred. For a virtual machine driver, a pre-configured model service VM template image is used. These images contain the operating system, software environment, and model-related files required to run the model service, and serve as the foundation for creating cloud resources.

[0083] After the heterogeneous resource driver plug-in module determines the target driver and the corresponding image file, various driver plug-ins in the heterogeneous resource plug-in module (such as virtual machine API driver, K8s CRD driver, bare metal IPMI driver) will perform specific operations in the target resource pool to create cloud resources based on the information of the target driver and image file.

[0084] For example, for container scenarios: Use the K8s CRD driver to generate a Deployment YAML file (a configuration file used to describe container deployment in Kubernetes), then use the Kubectl tool (Kubernetes command-line tool) to deploy the container and deploy the model service into the container environment.

[0085] For bare metal scenarios: Through the bare metal IPMI driver, the Redfish interface (an interface standard for managing server hardware) is used to trigger PXE (Preboot Execution Environment) to install the specified ISO image, thereby installing the operating system and the environment required for the model service on the bare metal server, creating bare metal cloud resources capable of running the model service.

[0086] For virtual machine scenarios: Call the virtual machine API driver, use the virtual machine API to create a virtual machine instance with GPU resources, deploy the model service to the virtual machine, and enable the model service to run in the virtual machine environment.

[0087] In summary, the present invention discloses an automated orchestration process for model services on a cloud management platform, which can effectively improve the efficiency of model service deployment in complex private cloud scenarios: the model service delivery time is shortened from hours to minutes, effectively improving resource utilization: and reducing idle resources by more than 30% through elastic scaling.

[0088] based on Figure 1 and Figure 2 The embodiment shown, as Figure 4 As shown, this application provides a schematic diagram of a resource creation architecture for a specific model service.

[0089] Reference Figure 4 The resource creation architecture of the model service of this application includes a cloud management platform model service interaction layer, a model service standardization engine, an intelligent orchestration and scheduling module, a heterogeneous resource driver plug-in module, and a hybrid image warehouse.

[0090] In an embodiment of the present application, the model service interaction layer of the cloud management platform can obtain the model service deployment requirements input by the user. The model service standardization engine converts the model service deployment requirements into a standardized service template through the predefined lifecycle management atomic capabilities and preset templates. Afterwards, the intelligent orchestration and scheduling module parses the standardized service template, selects the target resource pool from multiple resource pools, and generates a model service resource scheduling instruction that can be recognized by the heterogeneous resource driver plug-in module. The heterogeneous resource driver plug-in module identifies the model service resource scheduling instruction, determines the target driver corresponding to the target resource pool in the cloud basic resources; and calls the image file corresponding to the target driver from the hybrid image repository; based on the image file and the target driver, creates the cloud resources of the model service in the target resource pool. For example, for virtual machine drivers, you can call virtual machine API instructions and use the virtual machine template image of the preset model service as the basis. Whether it is an OpenStack resource pool or a VMware resource pool, you can create a virtual machine instance equipped with GPU resources to meet scenarios with special requirements for computing resources; for container drivers, you can extend the K8S API function through Kubernetes CRD, combine the Redfish interface to trigger the PXE installation process, and deploy the Docker image of the preset model service to the K8S resource pool to achieve automated orchestration and management of containerized applications; for bare metal drivers, you can use the Redfish / IPMI protocol to operate directly in the physical machine resource pool, create cloud resource instances, and provide a deployment environment for applications that require exclusive physical resources, effectively leveraging the high performance and low latency of bare metal.

[0091] Among them, reference Figure 4The resource creation architecture of the model service of the present application also includes the same operation and maintenance monitoring module, which can obtain the operating indicators of at least one resource pool in real time; determine whether the operating indicators in at least one resource pool meet the preset indicator range; if the first operating indicator in the first resource pool is not within the first preset indicator range, determine the preset operation corresponding to the first operating indicator to execute the preset operation corresponding to the first operating indicator.

[0092] It is understandable that the specific functions of the cloud management platform model service interaction layer, model service standardization engine, intelligent orchestration and scheduling module, heterogeneous resource driver plug-in module and hybrid image warehouse in the resource creation architecture of the model service in this application can refer to the above Figure 1 and Figure 2 The embodiments shown are not described in detail here.

[0093] Through the description of the above implementation methods, those skilled in the art can clearly understand that the method according to the above embodiment can be implemented by means of software plus the necessary general hardware platform, and of course it can also be implemented by hardware, but in many cases the former is a better implementation method.

[0094] The embodiment of the present application also provides a resource creation device 500 for a model service. Figure 5 A schematic diagram of a resource creation device for a model service provided in an embodiment of the present application is shown in FIG. Figure 5 As shown, including:

[0095] A conversion unit 501 is used to convert the model service deployment requirements input by the user into a standardized service template;

[0096] A parsing unit 502 is configured to parse the standardized service template to obtain a model service resource scheduling instruction, where the model service resource scheduling instruction is used to indicate a target resource pool;

[0097] The deployment unit 503 is configured to determine a target driver corresponding to the target resource pool, so as to create cloud resources of the model service in the target resource pool by using the target driver.

[0098] Furthermore, in a possible implementation of the embodiment of the present application, the conversion unit 501 is used to: based on the predefined lifecycle management atomic capabilities, convert the model service deployment requirements into a standard service template according to a preset template, the standard service template includes the model service configuration parameters corresponding to the model service, and the lifecycle management atomic capabilities are used to represent the basic operation unit set in the entire life cycle of the model service.

[0099] Furthermore, in a possible implementation of the embodiment of the present application, the conversion unit 501 is used to: perform a compliance check on the model service deployment requirements based on preset rules to determine whether the model service deployment requirements are compliant; if the model service deployment requirements are compliant, convert the model service deployment requirements into a standardized service template.

[0100] Furthermore, in a possible implementation of the embodiment of the present application, the parsing unit 502 is used to: parse the standardized service template to obtain the model service configuration parameters in the standard service template; based on the model service configuration parameters, screen at least one candidate resource pool that matches the model service configuration parameters from all resource pools; score at least one candidate resource pool based on the resource utilization data corresponding to the at least one candidate resource pool to obtain at least one scoring result; select a target scoring result from the at least one scoring result, and determine the resource pool corresponding to the target scoring result as the target resource pool; and generate a model service resource scheduling instruction based on the target resource pool.

[0101] Furthermore, in a possible implementation method of the embodiment of the present application, the parsing unit 502 is used to: parse the standardized service template to obtain the model service configuration parameters in the standard service template; determine whether the model service configuration parameters meet the preset conditions, and if the model service configuration parameters meet the preset conditions, determine the preset operation corresponding to the preset conditions to generate the corresponding model service resource scheduling instructions according to the preset operation.

[0102] Furthermore, in a possible implementation of an embodiment of the present application, the deployment unit 503 is used to: determine the target driver corresponding to the target resource pool according to the target resource pool indicated by the model service resource scheduling; call the image file corresponding to the target driver; and create cloud resources for the model service in the target resource pool according to the image file and the target driver.

[0103] Furthermore, in a possible implementation of an embodiment of the present application, the resource creation device 500 of the model service also includes a determination unit, which is used to: obtain the operating indicators of at least one resource pool; determine whether the operating indicators in at least one resource pool meet the preset indicator range; if the first operating indicator in the first resource pool is not within the first preset indicator range, determine the preset operation corresponding to the first operating indicator to execute the preset operation corresponding to the first operating indicator.

[0104] For the description of the features in the embodiment corresponding to the resource creation device of the model service, please refer to the relevant description of the embodiment corresponding to the resource creation method of the model service, and no further details will be given here.

[0105] An embodiment of the present application further provides an electronic device comprising a memory and a processor, wherein the memory stores a computer program, and the processor is configured to run the computer program to execute the steps in any of the above-mentioned resource creation method embodiments of the model service.

[0106] An embodiment of the present application further provides a computer-readable storage medium storing a computer program, wherein the computer program is configured to execute the steps of any of the above-mentioned resource creation method embodiments of the model service when running.

[0107] In an exemplary embodiment, the computer-readable storage medium may include, but is not limited to, various media that can store computer programs, such as a USB flash drive, a read-only memory (ROM), a random access memory (RAM), a mobile hard disk, a magnetic disk, or an optical disk.

[0108] An embodiment of the present application further provides a computer program product, which includes a computer program. When the computer program is executed by a processor, the steps in any of the above-mentioned resource creation method embodiments of the model service are implemented.

[0109] An embodiment of the present application also provides another computer program product, including a non-volatile computer-readable storage medium, which stores a computer program. When the computer program is executed by a processor, it implements the steps in the resource creation method embodiment of any of the above-mentioned model services.

[0110] Professionals may further appreciate that the units and algorithm steps of each example described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, computer software, or a combination of the two. In order to clearly illustrate the interchangeability of hardware and software, the above description has generally described the components and steps of each example according to their functions. Whether these functions are performed in hardware or software depends on the specific application and design constraints of the technical solution. Professionals and technicians may use different methods to implement the described functions for each specific application, but such implementation should not be considered beyond the scope of this application.

[0111] The above is a detailed introduction to the resource creation method, device, equipment and medium of a model service provided by this application. This article uses specific examples to illustrate the principles and implementation methods of this application. The description of the above embodiments is only used to help understand the method of this application and its core idea. It should be pointed out that for ordinary technicians in this technical field, without departing from the principles of this application, several improvements and modifications can be made to this application, and these improvements and modifications also fall within the scope of protection of the claims of this application.

Claims

1. A resource creation method for a model service, characterized in that: The method comprises: Convert the model service deployment requirements entered by the user into a standardized service template; Parsing the standardized service template to obtain a model service resource scheduling instruction, wherein the model service resource scheduling instruction is used to indicate a target resource pool; A target driver corresponding to the target resource pool is determined, so as to create cloud resources of the model service in the target resource pool by using the target driver.

2. The method according to claim 1, characterized in that Converting the model service deployment requirements input by the user into a standardized service template includes: Based on the predefined lifecycle management atomic capabilities, the model service deployment requirements are converted into the standard service template according to the preset template. The standard service template includes the model service configuration parameters corresponding to the model service. The lifecycle management atomic capabilities are used to represent the basic operation unit set in the entire life cycle of the model service.

3. The method according to claim 1, characterized in that Converting the model service deployment requirements input by the user into a standardized service template includes: Based on preset rules, a compliance check is performed on the model service deployment requirements to determine whether the model service deployment requirements are compliant; If the model service deployment requirement is compliant, the model service deployment requirement is converted into the standardized service template.

4. The method according to claim 1, wherein The step of parsing the standardized service template to obtain the model service resource scheduling instruction includes: Parsing the standardized service template to obtain model service configuration parameters in the standard service template; According to the model service configuration parameters, screening from all resource pools to obtain at least one candidate resource pool that matches the model service configuration parameters; Scoring the at least one candidate resource pool according to the resource utilization data corresponding to the at least one candidate resource pool to obtain at least one scoring result; Selecting a target scoring result from the at least one scoring result, and determining a resource pool corresponding to the target scoring result as the target resource pool; A model service resource scheduling instruction is generated according to the target resource pool.

5. The method according to claim 1, wherein The step of parsing the standardized service template to obtain the model service resource scheduling instruction includes: Parsing the standardized service template to obtain model service configuration parameters in the standard service template; Determine whether the model service configuration parameters meet the preset conditions. If the model service configuration parameters meet the preset conditions, determine the preset operation corresponding to the preset conditions to generate the corresponding model service resource scheduling instructions according to the preset operation.

6. The method according to claim 1, characterized in that The determining of a target driver corresponding to the target resource pool to create cloud resources for a model service in the target resource pool by using the target driver includes: Determining a target driver corresponding to the target resource pool according to the target resource pool indicated by the model service resource scheduling; Calling the image file corresponding to the target driver; According to the image file and the target driver, cloud resources of the model service are created in the target resource pool.

7. The method according to any one of claims 1 to 5, characterized in that The method comprises: Get the operating indicators of at least one resource pool; Determining whether an operating indicator in the at least one resource pool meets a preset indicator range; If the first operating indicator in the first resource pool is not within a first preset indicator range, a preset operation corresponding to the first operating indicator is determined to execute the preset operation corresponding to the first operating indicator.

8. A resource creation device for a model service, characterized in that: The device comprises: The conversion unit is used to convert the model service deployment requirements input by the user into a standardized service template; a parsing unit, configured to parse the standardized service template to obtain a model service resource scheduling instruction, wherein the model service resource scheduling instruction is used to indicate a target resource pool; The deployment unit is configured to determine a target driver corresponding to the target resource pool, so as to create cloud resources of the model service in the target resource pool by using the target driver.

9. An electronic device, characterized in that: include: a processor and a memory for storing a computer program capable of being executed on the processor, Wherein, when the processor is used to run the computer program, it executes the resource creation method of the model service according to any one of claims 1-7.

10. A non-transitory computer-readable storage medium storing computer instructions, characterized in that: The computer instructions are used to enable the computer to execute the resource creation method for the model service according to any one of claims 1 to 7.

Citation Information

Cited By

  • Application resource processing method and device, cloud platform, medium and product

    CN120856710A

  • Heterogeneous resource scheduling method and device of cloud data center, medium and product

    CN121433917A

  • A method, device, medium, and product for scheduling heterogeneous resources in a cloud data center.

    CN121433917B