Cluster computing resource allocation method
By adopting virtual machines, container cloud and bare metal resource services in artificial intelligence simulation experiments, combining high-availability design and dynamic resource scheduling, the problem of multi-user/multi-task/multi-sample online collaborative parallel deployment is solved, and efficient experimental resource allocation and automatic deployment is achieved, ensuring the high performance and high reliability of the system.
Patent Information
- Application Number
- CN202510670308.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-23
- Publication Date
- 2025-08-12
AI Technical Summary
Traditional computing cluster resource allocation methods are difficult to meet the needs of multi-user/multi-task/multi-sample online collaborative parallel deployment, especially in artificial intelligence simulation experiments, it is difficult to achieve new needs such as automatic deployment strategy generation of test resources, efficient distribution of large-scale test resource images and one-click deployment, and release of test resources.
Three resource service forms: virtual machine, container cloud and bare metal, combined with virtual cluster management and high availability design, realize automatic deployment strategy generation of experimental resources and flexible scheduling of computing resources, supports rapid deployment and load balancing of virtual machines, container cloud and bare metal, generate container images through Dockerfile files, and use Kubernetes technology to dynamically generate and deploy containers, combine Prometheus+Node Exporter architecture for real-time data acquisition and preprocessing, and dynamically adjust resource weights for elastic expansion and expansion.
It realizes high-performance and highly reliable cluster computing resource allocation, supports multi-user/multi-task/multi-sample online collaborative parallel deployment, realizes automatic deployment and efficient distribution of experimental resources, and ensures the high availability and ease of use of the system.
Smart Images

Figure CN120469769A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the fields of computer network technology, artificial intelligence and cloud computing, and specifically designs a cluster computing resource allocation method. Background Art
[0002] Artificial intelligence simulation experiments have typical characteristics such as "state space combination explosion", "high computing power requirements", "large use case scale", and "self-learning and self-evolution". Traditional computing cluster resource allocation methods are difficult to meet the needs of multi-user / multi-task / multi-sample online collaborative parallel deployment, and are difficult to meet new needs such as automatic deployment strategy generation of test resources, efficient distribution and one-click deployment of large-scale test resource images, and release of test resources.
[0003] Therefore, it is urgent to propose a cluster computing resource allocation method that meets the needs of large-sample, high-concurrency artificial intelligence simulation experiments for multiple users. Summary of the Invention
[0004] In response to the problems existing in the prior art, the purpose of the present invention is to provide a cluster computing resource allocation method that is oriented towards different types of test objects, different types of test tasks, and different test modes, flexibly adopts a variety of resource service forms, supports the rapid deployment of typical virtual machine instances, supports the orchestration, scheduling, deployment and operation of containerized applications, elastic expansion and contraction, load balancing and high availability capabilities, and supports the stable and effective operation of database-specific functional system software.
[0005] To achieve the above-mentioned objectives, the present invention provides a cluster computing resource allocation method, which supports three resource service forms: virtual machines, container clouds, and bare metal machines. Among them, virtual machine resources are used for the deployment of system software with UI interfaces and C / S architectures; container cloud resources are used for the deployment of system software with B / S architectures and the operation of large experimental samples; bare metal resources are used for the deployment of specific functional system software, Kafka middleware and / or databases; and deployment service computing resources are elastically scheduled and controlled to realize cluster computing resource allocation.
[0006] Furthermore, 86% of the computing nodes are allocated to the container cloud, and several nodes are used as control nodes, where the control nodes are distributed on different computing boards and switch boards.
[0007] Furthermore, 2% of the computing nodes are allocated to virtual machines, and several nodes are used as control nodes, where the control nodes are distributed on different computing boards and switch boards.
[0008] Furthermore, 2% of the computing nodes are allocated for use as bare metal resources.
[0009] Furthermore, 10% of computing nodes are reserved as backup, and will be dynamically added to the container cloud and / or virtual machine cluster based on the platform resource usage.
[0010] Furthermore, in the container cloud resources, the steps for containerized deployment are as follows: S1.1 Container image generation: Generate functional system or application container images, using Dockerfile-based methods to quickly create container images; S1.2 Container-based image distribution and deployment; use container and container cluster technology to complete dynamic generation and deployment of containers; load and schedule container images from the image library, and deploy agents in the virtual server to start the operation of container images.
[0011] Further, the virtual machine deployment steps are as follows: S2.1 Virtual Machine Image Generation: Generate functional system or application virtual machine images based on cloud environment virtualization technology, and create virtual machine images based on operation support tools and cloud environment management tools; S2.2 Image distribution and deployment based on virtual machines.
[0012] Furthermore, step S2.1 further includes: S2.1.1: Start running a blank image instance; S2.1.2: Install the operation support tools required for operation; S2.1.3: Save the virtual machine instance as a virtual machine image; S2.1.4: Package and upload the image and register it in the repository.
[0013] Furthermore, the automatic deployment process of a functional system or application based on a virtual machine is divided into two steps. The specific process is as follows: S2.2.1: Dynamically create virtual machines and associate the mapping relationship between the functional system or application to be deployed and the virtual machine; S2.2.2: The virtual machine automatically downloads the functional system or application to be deployed from the database, loads the configuration file, and runs the functional system or application.
[0014] Furthermore, in step S1.1, the process of generating the functional system or application container image based on the Dockerfile file is as follows: S1.1.1: Package the functional system or application resources; package the functional system or application and its running dependencies into an image, load it into a container, compile, run, and test it; S1.1.2: Generate a Dockerfile; load the base container image that the functional system or application depends on, set the environment variables that the container image depends on, declare the port that the service in the image listens on, and specify the default entry command for the container image, as well as the username and ID for running the container; S1.1.3: Create a container image; load the Dockerfile file, generate the container image by calling the Build command, and upload the generated container image to the image library. Once the container image is successfully uploaded, it can be queried and reused through the image service.
[0015] The beneficial effects of the present invention are as follows: The present invention flexibly adopts a variety of resource service forms. By applying virtual cluster management and high-availability design, it can realize the automatic deployment strategy generation of experimental resources and the elastic scheduling and control of computing resources, providing users with high-performance, highly reliable, service-oriented and easy-to-use cluster computing resource allocation and deployment services. BRIEF DESCRIPTION OF THE DRAWINGS
[0016] Figure 1 A schematic diagram of the test resource container image generation process of the present invention is shown; Figure 2 A schematic diagram of the distribution and deployment process of the experimental resource container image of the present invention is shown; Figure 3 A schematic diagram of the test resource virtual machine image generation process of the present invention is shown; Figure 4 A schematic diagram of the test resource virtual machine image distribution and deployment process of the present invention is shown. DETAILED DESCRIPTION
[0017] The following will clearly and completely describe the technical solution of the present invention in conjunction with the accompanying drawings. Obviously, the embodiments described are only some embodiments of the present invention, not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without making creative efforts are within the scope of protection of the present invention.
[0018] In the description of the present invention, it should be noted that the terms "center," "upper," "lower," "left," "right," "vertical," "horizontal," "inner," and "outer," etc., indicating orientations or positional relationships, are based on the orientations or positional relationships shown in the accompanying drawings and are intended solely to facilitate and simplify the description of the present invention. They are not intended to indicate or imply that the devices or components referred to must have, be constructed, or operate in a specific orientation, and therefore should not be construed as limitations on the present invention. Furthermore, the terms "first," "second," and "third" are used for descriptive purposes only and should not be construed as indicating or implying relative importance.
[0019] In the description of the present invention, it should be noted that, unless otherwise expressly specified or limited, the terms "mounted," "connected," and "connected" should be understood in a broad sense. For example, they may refer to fixed, detachable, or integral connections; mechanical or electrical connections; direct or indirect connections through an intermediate medium; and internal communication between two components. Those skilled in the art will understand the specific meanings of the above terms in the present invention based on specific circumstances.
[0020] The following combination Figure 1-Figure 4 The specific embodiments of the present invention are described in detail. It should be understood that the specific embodiments described herein are only used to illustrate and explain the present invention and are not used to limit the present invention.
[0021] This invention provides a cluster computing resource allocation method that supports three resource service forms: virtual machines, container clouds, and bare metal. Virtual machine resources are primarily used for deploying system software related to a client-server architecture with a UI interface; container cloud resources are primarily used for deploying system software with a client-server architecture and running large experimental samples; and bare metal resources are used for deploying specific functional system software, Kafka middleware, databases, and other applications. By applying virtual cluster management and high-availability design, this method provides flexible and reliable service guarantees for system intelligent applications. This invention flexibly utilizes multiple resource service forms. By applying virtual cluster management and high-availability design, it can achieve automatic deployment strategy generation for experimental resources and flexible scheduling and control of computing resources, providing users with a high-performance, highly reliable, service-oriented, and easy-to-use cluster computing resource allocation and deployment service.
[0022] 1. Container Cloud In the simulation test environment, container cloud resources are primarily planned for B / S architecture system software deployment and large-scale test sample operation. Considering the significant demand for container cloud resources from various systems and intelligent simulation tests, approximately 86% of compute nodes are allocated to the container cloud, with several nodes used as control nodes. These control nodes are distributed across different compute and switch boards to facilitate offline maintenance in the event of hardware failures at the control nodes. The steps for containerized deployment are as follows:
[0023] S1.1 Container Image Generation Functional system / application container image generation is planned to use the Dockerfile file method to quickly create container images. Figure 1 The process of generating a functional system / application container image based on the Dockerfile file is given as follows: S1.1.1: After the user submits a request to encapsulate a functional system or application resource container, the functional system / application resource is encapsulated. The functional system / application and its running dependency packages are encapsulated into an image, loaded into the container for compilation, operation, and testing. The implementation language is determined. For functional systems / applications developed in compiled languages, including C, C++, and Go, the corresponding code compiler is dynamically called to compile the code and generate an executable code package. For functional systems / applications developed in interpreted languages, including Java and Python, the pre-set code packaging environment is called to package the code and generate an executable program or code package.
[0024] S1.1.2: Generate a Dockerfile. Load the base container image that the functional system / application depends on, set the environment variables and parameters that the container image depends on, copy the code package to the base image, set the service listening port, declare the port that the service in the image listens on, and specify the default entry point for the container image, as well as the username and ID for running the container.
[0025] S1.1.3: Create a container image. Load the Dockerfile, generate the container image by calling the Build command, and upload the generated container image to the image repository. Once the container image is successfully uploaded, it can be retrieved and reused through the Image Service.
[0026] S1.2 Container-based image distribution and deployment To speed up container image deployment, we use container (Docker) and container cluster (Kubernetes) technologies to dynamically generate and deploy containers. We load and schedule container images from the image library, and deploy agents in virtual servers to start running container images. Figure 2 The experimental resource container image distribution and deployment steps based on cloud computing server clusters, container image libraries, deployment queues, virtual networks, and cloud computing are given as follows: S1.2.1: Generate a container deployment file, including the version number, Pod name, container image name, container startup command parameters, container operating mode, port number on the virtual host where the container is located, port protocol, and the CPU and memory resources required by the container image.
[0027] S1.2.2: Container Image Loading and Deployment. The deployment agent sends a request to the master node API server and creates a Controller object, Deployment. The controller Deployment loads the corresponding image file from the image repository and requests the API server to create a certain number of Pod objects. Simultaneously, the scheduler on the master node assigns the container to the selected worker node and starts it.
[0028] S1.2.3: To provide a fixed access point for Pod objects, a request is made to the API Server in the Master node to create a Service object, which dynamically allocates a dedicated Cluster IP address and port, allowing functional systems / applications to access the service through its service name and Cluster IP.
[0029] 2. Virtual Machine In the intelligent simulation test environment, virtual machine resources are primarily planned for the deployment of system software with a UI client / server architecture. Based on computing resource requirements and considering the relatively low virtual machine resource requirements of various functional systems and test applications, approximately 2% of computing nodes are allocated to virtual machines, with several nodes used as control nodes. These control nodes are distributed across different computing boards and switch boards to facilitate offline maintenance in the event of hardware failures at the control nodes. The steps for virtual machine deployment are as follows: S2.1 Virtual Machine Image Generation Generate functional system / application virtual machine images based on cloud environment virtualization technology. You need to prepare dependent operation support tools in advance and use cloud environment management tools to create virtual machine images. Figure 3 The specific implementation process is given and described as follows: S2.1.1: Start running a blank image instance: Based on the QEMU / KVM virtualization framework, call the Libvirt API to dynamically create a blank virtual machine instance; Load a lightweight Linux kernel (customized and tailored to retain only necessary drivers, and the kernel size is compressed from 70MB to 28MB); Initialize disk partitions using Btrfs+Zstd compression to improve the efficiency of subsequent incremental layer writing; It can automatically match the optimal kernel version based on the target hardware characteristics (CPU architecture / GPU model), bypass traditional BIOS boot through memory mapping technology, reduce boot time from 15s to 3.2s, and achieve zero-copy boot.
[0030] S2.1.2: Install the runtime support tools required for the runtime: Dependency graph construction: Use the directed acyclic graph (DAG) model to analyze the application's library dependencies; Minimal installation strategy: Select required dependencies based on a greedy algorithm, filtering out non-essential components (such as documentation / debugging symbols); Parallel installation: Use Ansible Playbook to implement concurrent execution of multiple tasks and task distribution algorithm: - name: Batch installation dependencies hosts: localhost strategy: free # Full concurrency mode tasks: - apt: name={{ item}} state=present loop: "{{ minimal_packages}}".
[0031] S2.1.3: Save the virtual machine instance as a virtual machine image; S2.1.4: Package and upload the image and register it in the repository; Upload process: Block upload: Cut the image into 10MB blocks and upload them to Swift object storage in parallel; Global index registration: records image metadata in MongoDB; CDN preheating: Pre-distribute images to edge nodes through the Anycast network to shorten the initial deployment delay.
[0032] Once the image is successfully created and uploaded, it can be queried and reused through the image service. This invention optimizes the entire process of virtual machine image generation and provides highly available, lightweight, and intelligent infrastructure support for subsequent flexible scheduling.
[0033] S2.2 Image distribution and deployment based on virtual machines Figure 4 The automatic deployment process of functional systems / applications based on virtual machines is divided into two steps. The specific implementation process is as follows: S2.2.1: Dynamically create virtual machines and associate the mapping relationship between the functional systems / applications to be deployed and the virtual machines; S2.2.2: The virtual machine automatically downloads the functional system / application to be deployed from the database, loads the configuration file, and runs the functional system / application.
[0034] 3. Bare Metal To meet the high stability requirements of related software tools and databases for hard disk storage resources, 2% of computing nodes can be allocated as bare metal resources for the deployment of specific functional system software, Kafka middleware, databases, etc.
[0035] 4. Backup nodes Taking into account the multi-user and large-sample parallel testing requirements of artificial intelligence simulation experiments, 10% of computing nodes can be reserved as backup, and then dynamically added to the container cloud / virtual machine cluster based on the platform resource usage.
[0036] In order to achieve flexible scheduling of computing resources, the present invention adopts the following steps: 1. Step S1: Real-time data acquisition and preprocessing 1. Multi-dimensional parameter collection Deploy the Prometheus+Node Exporter architecture and collect the following metrics every 10 seconds: CPU utilization:
[0037] Memory pressure index: ×100% Network bandwidth saturation: ×100% Task queue depth:
[0038] Resource volatility: (Rolling calculation of the standard deviation for the last 5 minutes); CPU utilization: Used to determine whether a node is overloaded (e.g., >90%) or idle (e.g., <20%), and identify bottlenecks in compute-intensive tasks.
[0039] Memory pressure index: Used to detect the degree of memory fragmentation and memory exhaustion risk, such as when >95% triggers automatic recycling.
[0040] Network bandwidth saturation: Used to identify network congestion (such as >85%) to avoid cross-node communication delays.
[0041] Task queue depth: Used to monitor task accumulation (such as >80%), triggering automatic capacity expansion.
[0042] Resource volatility: Used to identify sudden traffic or abnormal fluctuations (such as It provides a basis for predictive scheduling.
[0043] in, Meaning: total time (the total working time of the CPU in the statistical period); Meaning: idle time (the cumulative time the CPU is in idle state); Meaning: Used memory (the amount of memory actually occupied by the application); Meaning: buffer memory (memory cached by the kernel buffer, used to temporarily store I / O data); Meaning: total memory (total physical memory of the node); Meaning: number of bytes sent (the total amount of data sent by the network interface during the statistical period); Meaning: maximum bandwidth (the theoretical maximum transmission rate of the network interface, unit: bps); Meaning: number of pending tasks (number of tasks waiting to be executed); Meaning: maximum queue length (the maximum task capacity allowed by the queue); Meaning: resource values at each moment (such as CPU utilization and memory usage at a certain second); Meaning: Mean (arithmetic mean of resource indicators in the last 5 minutes); Meaning: number of data points (the number of data points collected in the last 5 minutes, for example: 5 minutes × 60 seconds / 10-second interval = 30 data points), Refers to a time value.
[0044] 2. Data normalization Perform extreme value normalization on each parameter:
[0045] Represents the original parameter value; max / min: historical extreme values within a 72-hour rolling window; Output range: 0 (lowest load) ~ 100 (highest load).
[0046] Step S2: Dynamic Load Assessment 1. Dynamic weight allocation Adjust the weights based on resource type (container cloud / virtual machine / bare metal) and real-time scenario:
[0047] Meaning is the memory weight coefficient; indicates the memory pressure index The weight of memory resources in the comprehensive resource score is used to control the impact of memory resources on the overall load.
[0048] Meaning is the task queue weight coefficient; indicates the depth of the task queue The weight ratio in the comprehensive resource score reflects the impact of task backlog on system stability.
[0049] Adjustment rules: when near Significant improvement in time (to prevent task abandonment); in real-time computing scenarios (such as stream processing), a higher default value is assigned. .
[0050] It means the response time decay factor; it is used to control the decay rate of historical response time data in weight calculations to prevent old data from excessively influencing current decisions.
[0051] Value range: 0 < ≤ 1 (e.g. 0.9 means the weight of historical data decays by 10% every period).
[0052] Adjustment rules: Reduce in scenarios with severe fluctuations (such as burst traffic) (fast response to changes); improves in steady-state load Maintaining strategy stability).
[0053] This parameter represents the response time weight coefficient, which indicates the weight of the system's average response time in the overall score and directly affects the trigger sensitivity of automatic scaling.
[0054] Adjustment rules: When SLA (Service Level Agreement) requirements are strict, it will significantly increase (Prioritize response speed); can be reduced in batch tasks (allowing for moderate delays).
[0055] In the present invention, the comprehensive load scoring formula is:
[0056] when > the threshold, triggering scaling.
[0057] in, Indicates the average system response time. Service Level Agreement Threshold refers to the critical value of key performance indicators (such as response time and availability) agreed upon in the Service Level Agreement (SLA).
[0058] 2. Comprehensive load index calculation Calculate the resource type R (R∈{container cloud, virtual machine, bare metal}):
[0059] Variable Description: : The dynamic weight of parameter i in resource type R; : normalized value of parameter i; : Indicates parameters The normalized value of Is a parameter of a resource indicator The purpose of normalizing the load scores (such as CPU utilization and memory pressure index) is to eliminate the dimensional differences of different resource types (container cloud / virtual machine / bare metal) so that they can participate in the comprehensive load score calculation fairly.
[0060] : Resource volatility (amplifies the impact of unstable scenarios).
[0061] Step S3: Elastic Decision Generation 1. Prioritization: Arrange resource types in descending order of load index to generate a capacity expansion priority queue: PriorityQueue = sort( , , , reverse=True).
[0062] 2. Dynamic allocation strategy: 1) Emergency capacity expansion (load index ≥ 150) Trigger condition: When the comprehensive load index of a certain type of resource (container cloud / virtual machine / bare metal) When the value is ≥150, it is considered an emergency capacity expansion scenario.
[0063] Allocation rules: Number of nodes required =
[0064] Formula Description: : Every time the threshold is exceeded by 50 points, one baseline expansion unit is triggered; ×2: Double the allocation of resources in emergency scenarios to quickly suppress the avalanche effect; Example: If the container cloud load index =162, then: Required number of nodes node.
[0065] 2) Conventional expansion (120 ≤ load index < 150) Trigger condition: When 120≤LR<150, it is determined to be a regular expansion scenario.
[0066] Allocation rules: Number of nodes required =
[0067] Formula Description: : For every 30 points exceeding the threshold, one baseline expansion unit is allocated; Example: If the virtual machine load index =135, then: .
[0068] 3) Reduced capacity (load index < 80) Trigger condition: When When <80, it is determined to be a resource-inefficient scenario and redundant nodes need to be released.
[0069] Scaling rules:
[0070] Formula Description: Release 20% (rounded down) of the number of nodes currently allocated to this resource type. Example: If the bare metal currently has 10 nodes allocated and =75, then: Number of released nodes node.
[0071] 4) Priority execution logic Resource processing order: Generate a priority queue in descending order of load index (the higher the load, the higher the priority); Process the resource types in the queue one by one, and the execution order is: Emergency expansion → Regular expansion → Reduction
[0072] Allocation termination conditions: When the reserved nodes are exhausted (number of remaining nodes = 0), the allocation is terminated immediately; The scaling-down operation is not limited by the number of reserved nodes, and the released nodes are directly returned to the reserved pool.
[0073] 5) Dynamic weight correction mechanism If a resource type triggers the following scenarios, the weight must be corrected before calculating the load index: High volatility scenarios (resource volatility >15%): ← +0.1; Task backlog scenario (task queue depth >20%): ← +0.05× ; SLA breach scenarios (response time deviation >30%): ←0.3 (forced to overwrite the original weight).
[0074] Policy execution example: Input conditions: 1) Current node allocation: 86 container cloud nodes, 2 virtual machine nodes, and 2 bare metal nodes; 2) Number of reserved nodes: 10 nodes; 3) Load index: Container cloud 162, virtual machine 105, bare metal 88
[0075] Execution process: 1) Priority queue generation: container cloud (162) → virtual machine (105) → bare metal (88); 2) Emergency capacity expansion: Required number of container cloud nodes = 2, remaining reserved nodes = 10-2 = 8; 3) Conventional expansion: This function is not triggered for VMs (105<120); 4) Scaling: Bare machines (88 ≥ 80) are not triggered; 5) Output: Container Cloud allocation +2 nodes (88 nodes in total); Remaining reserved nodes = 8.
[0076] Through this strategy, the system can complete decision calculations within 5 milliseconds and achieve resource expansion response delay of less than 1 second for loads > 150, ensuring the real-time and accuracy of resource scheduling.
[0077] 4. Step S4: Resource Allocation Execution 1. Node migration strategy Calculate the migration cost and only allow low-cost migrations: ; .
[0078] in, Meaning: Used memory, which indicates the physical memory capacity of the node currently actually occupied by applications (excluding cache and buffer). Meaning: Available bandwidth, which indicates the remaining network transmission capacity between nodes or between a node and storage, that is, the currently unoccupied bandwidth. Meaning: Service downtime, which indicates the cumulative time that business services are unavailable during the node migration process (from the start of migration to the full restoration of services).
[0079] 2. Dynamic replenishment of reserved resources Allocation threshold: When a resource is >110, triggers reserved resource allocation; Supplementary formula: ; Recycling mechanism: If the resource load <60 for 10 minutes, releasing 50% of the allocated reserved nodes.
[0080] in, Meaning: the number of nodes that need to be supplemented, indicating the number of nodes that need to be dynamically added based on the current resource gap and load forecast to ensure the redundant capacity of the resource pool.
[0081] Meaning: the number of remaining allocatable nodes, which indicates the number of idle nodes in the resource pool that are currently unoccupied and can be scheduled immediately.
[0082] : Round up function, which rounds the decimal result towards positive infinity to ensure that the calculated result is the smallest integer not less than the original value.
[0083] 5. Step S5: Dynamic Adjustment and Rebalancing 1. Periodic rebalancing (every 5 minutes) 2. Exception handling mechanism Hardware failure: If a physical node fails to respond for three consecutive times, it is automatically marked as a failed node, triggering the migration task = ; Network partition: Re-elect the master node based on the RAFT algorithm to ensure high availability of the scheduler.
[0084] Technical solution advantages: 1. Dynamic weight mechanism: Adapts to different scenarios (such as sudden traffic and task backlogs) by adjusting weights in real time.
[0085] 2. Nonlinear expansion strategy: During emergency expansion, the number of allocated nodes increases exponentially (×2 coefficient in the formula), quickly suppressing the avalanche effect.
[0086] 3. Cost-aware migration: Avoid high-cost migration and improve scheduling economy through the Cmigrate model.
[0087] 4. Predictive reservation allocation: Predict future load based on historical trends and allocate reserved resources in advance (requires integration of LSTM prediction module).
[0088] Beneficial technical effects of the present invention: The present invention proposes a cluster computing resource allocation method to solve the problems of online collaborative parallel deployment of multiple users / multi-tasks / multi-samples in intelligent experiments, automatic deployment strategy generation of experimental resources, efficient distribution and one-click deployment of large-scale experimental resource images, and release of experimental resources, thus effectively supporting the implementation of artificial intelligence joint simulation experiments and test evaluation.
[0089] Any process or method described in the flowchart of the present invention or in other ways herein can be understood as representing a module, segment or portion of code including one or more executable instructions for implementing specific logical functions or process steps, which can be implemented in any computer-readable medium for use by an instruction execution system, device or apparatus. The computer-readable medium can be any medium that stores, communicates, propagates or transmits a program for use by an execution system, device or apparatus, including read-only memory, magnetic disk or optical disk, etc.
[0090] Throughout this specification, reference to terms such as "embodiment" and "example" indicates that a specific feature, structure, material, or characteristic described in conjunction with that embodiment or example is included in at least one embodiment or example of the present invention. In this specification, the schematic representations of these terms do not necessarily refer to the same embodiment or example. Furthermore, those skilled in the art may combine or integrate different embodiments or examples described in this specification, as well as features therein, without creating any inconsistency.
[0091] Although the above content has shown and described the embodiments of the present invention, it can be understood that the above embodiments are exemplary and cannot be understood as limitations of the present invention. Ordinary technicians in this field can perform update operations such as changes, modifications, replacements and variations on the above embodiments within the scope of the present invention.
Claims
1. A cluster computing resource allocation method, characterized in that: The method supports three types of resource services: virtual machines, container clouds, and bare metal. Virtual machine resources are used for the deployment of system software with a UI interface and a client-server architecture; container cloud resources are used for the deployment of system software with a client-server architecture and for running large experimental samples; and bare metal resources are used for the deployment of specific functional system software, Kafka middleware, and / or databases. Flexible scheduling and control of deployment service computing resources are implemented to achieve cluster computing resource allocation.
2. The cluster computing resource allocation method according to claim 1, characterized in that: 86% of the computing nodes are allocated to the container cloud, and several nodes are used as control nodes, where the control nodes are distributed on different computing boards and switch boards.
3. The cluster computing resource allocation method according to claim 1, wherein: Allocate 2% of the computing nodes to virtual machines, and use several nodes as control nodes, where the control nodes are distributed on different computing boards and switch boards.
4. The cluster computing resource allocation method according to claim 1, wherein: Allocate 2% of the compute nodes for bare metal resources.
5. The cluster computing resource allocation method according to claim 1, characterized in that: 10% of computing nodes are reserved as backup, and will be dynamically added to the container cloud and / or virtual machine cluster based on platform resource usage.
6. The cluster computing resource allocation method according to claim 2, characterized in that: In container cloud resources, the steps for containerized deployment are as follows: S1.1 Container image generation: Generate functional system or application container images, using Dockerfile-based methods to quickly create container images; S1.2 Container-based image distribution and deployment; use container and container cluster technology to complete dynamic generation and deployment of containers; load and schedule container images from the image library, and deploy agents in the virtual server to start the operation of container images.
7. The cluster computing resource allocation method according to claim 3, characterized in that: The steps for deploying a virtual machine are as follows: S2.1 Virtual Machine Image Generation: Generate functional system or application virtual machine images based on cloud environment virtualization technology, and create virtual machine images based on operation support tools and cloud environment management tools; S2.2 Image distribution and deployment based on virtual machines.
8. The cluster computing resource allocation method according to claim 7, characterized in that: Step S2.1 further comprises: S2.1.1: Start running a blank image instance; S2.1.2: Install the operation support tools required for operation; S2.1.3: Save the virtual machine instance as a virtual machine image; S2.1.4: Package and upload the image and register it in the repository.
9. The cluster computing resource allocation method according to claim 7, characterized in that: The automatic deployment process of a functional system or application based on a virtual machine is divided into two steps. The specific process is as follows: S2.2.1: Dynamically create virtual machines and associate the mapping relationship between the functional system or application to be deployed and the virtual machine; S2.2.2: The virtual machine automatically downloads the functional system or application to be deployed from the database, loads the configuration file, and runs the functional system or application.
10. The cluster computing resource allocation method according to claim 6, characterized in that: In step S1.1, the process of generating a functional system or application container image based on the Dockerfile file is as follows: S1.1.1: Package the functional system or application resources; package the functional system or application and its running dependencies into an image, load it into a container, compile, run, and test it; S1.1.2: Generate Dockerfile file; Load the basic container image that the functional system or application depends on, set the environment variables that the container image depends on, declare the port that the service in the image listens on, and specify the default entry instruction of the container image, the user name and ID when running the container; S1.1.3: Create a container image; load the Dockerfile file, generate the container image by calling the Build command, and upload the generated container image to the image library. Once the container image is successfully uploaded, it can be queried and reused through the image service.
Citation Information
Cited By
Edge scene-oriented file-level container mirror image storage service implementation method and system
CN121116931A