Cloud computing-based computing resource management method and system

Through dynamic monitoring and load feature analysis of heterogeneous cloud resource pools, resource supply attributes are constructed and resource sharding rules with elastic boundaries are generated, resource fragmentation and load imbalance problems are solved, and the utilization rate of computing resources and system performance are improved.

CN120276852AInactive Publication Date: 2025-07-08HAINAN BIDEN TECHNOLOGY CO LTD
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
CN202510385781.9
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-29
Publication Date
2025-07-08
Estimated Expiration
Not applicable · inactive patent

AI Technical Summary

Technical Problem

The existing cloud-based computing resource management methods lack sufficient dynamic monitoring and efficient scheduling of heterogeneous computing resource pools, and cannot obtain the operating data of physical nodes and virtualized instances in real time, resulting in fragmentation and load imbalance in resource allocation, reducing resource utilization and system performance.

Method used

By collecting dynamic operation data of physical nodes and virtualized instances from heterogeneous computing resource pools, performing load characteristic parameter analysis, constructing resource supply attributes across availability zones, dynamically assessing fragmentation index and load balancing in resource scheduling process, generating resource sharding rules with elastic boundaries, and dynamically correcting the resource allocation structure based on real-time monitoring data.

Benefits of technology

Accurate analysis of computing resources is achieved, resource utilization is improved, resource waste is reduced, task execution efficiency and system stability are improved.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120276852A_ABST
    Figure CN120276852A_ABST
Patent Text Reader

Abstract

The invention provides a computing resource management method and system based on cloud computing, and relates to the technical field of computing resource management, under a hybrid cloud architecture, resource demand analysis is performed on load characteristic parameters to obtain resource classification identifiers corresponding to computing nodes to be classified in computing resources, and the resource classification identifiers are classified according to the load characteristic parameters. According to the resource classification identifier, constructing a cross-available-area resource supply attribute; determining efficiency association characteristics of a resource fragmentation index and a load balance degree in computing resource scheduling, fragmenting a task allocation strategy of computing resources according to the efficiency association characteristics and resource supply attributes, and generating a resource fragmentation rule with an elastic boundary in the computing resource scheduling; and splitting the calculation task into resource adjustment information matched with the resource fragmentation rule, performing resource mapping on the resource adjustment information, and dynamically correcting a resource allocation structure according to the real-time monitoring data. According to the method and the device, the computing resource demand can be accurately analyzed under the hybrid cloud architecture, so that the utilization rate of the computing resource is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the technical field of computing resource management. More specifically, this application relates to a computing resource management method and system based on cloud computing. Background Art

[0002] Computing resource management refers to the process of efficiently allocating, scheduling, monitoring, and optimizing computing resources in a cloud computing environment to ensure the rational use of resources and the maximization of system performance. With the wide application of cloud computing, the dynamic, heterogeneous, and distributed nature of computing resources makes resource management more complex. Traditional static resource configuration methods can no longer meet the large-scale and high-concurrency computing requirements in modern cloud platforms. Therefore, computing resource management methods based on cloud computing need to possess characteristics such as dynamic scheduling, load balancing, fault tolerance mechanisms, and elastic scaling.

[0003] However, existing computing resource management methods based on cloud computing usually lack sufficient dynamic monitoring and efficient scheduling of heterogeneous computing resource pools, cannot obtain the running data of physical nodes and virtualization instances in real time and accurately, nor can they effectively analyze load characteristics and resource requirements, resulting in fragmentation phenomena and load imbalance problems during the resource allocation process, further reducing the resource utilization rate and the optimization space of system performance, causing significant resource waste in the scheduling and task allocation of computing resources, and being unable to meet the increasingly complex application requirements and efficient resource management. Therefore, how to accurately analyze the computing resource requirements in a hybrid cloud architecture to improve the utilization rate of computing resources is an issue faced by the industry. Summary of the Invention

[0004] This application provides a computing resource management method and system based on cloud computing, which can accurately analyze the computing resource requirements in a hybrid cloud architecture to improve the utilization rate of computing resources.

[0005] In a first aspect, this application provides a computing resource management method based on cloud computing. The computing resource management method includes the following steps: Collect the dynamic running data of physical nodes and virtualization instances from a heterogeneous computing resource pool, and synchronously obtain the load characteristic parameters of user computing tasks; Under a hybrid cloud architecture, perform resource requirement analysis on the load characteristic parameters to obtain the resource classification identifiers corresponding to the to-be-classified computing nodes in the computing resources, and construct a cross-availability zone resource supply attribute according to the resource classification identifiers; Conduct a dynamic efficiency evaluation on the computing resource scheduling process to obtain the efficiency correlation characteristics of the resource fragmentation index and the load balance degree in the computing resource scheduling. Fragment the task allocation strategy of the computing resources according to the efficiency correlation characteristics and the resource supply attribute, and generate a resource fragmentation rule with an elastic boundary in the computing resource scheduling; Split the computing task into resource adjustment information that matches the resource sharding rule, perform resource mapping on the resource adjustment information, and dynamically correct the resource allocation structure according to real-time monitoring data.

[0006] In this embodiment, under the hybrid cloud architecture, performing resource requirement analysis on the load characteristic parameters to obtain the resource classification identifier corresponding to the computing node to be classified in the computing resources specifically includes: Eliminate the differences in the load characteristic parameters to obtain a standardized load vector; Perform multi-dimensional feature decoupling on the standardized load vector to obtain a computing throughput fluctuation coefficient; Determine the discrete probability distribution corresponding to the computing node to be classified in the computing resources according to the computing throughput fluctuation coefficient; Calculate the resource classification identifier corresponding to the computing node to be classified in the computing resources from the discrete probability distribution.

[0007] In this embodiment, constructing the resource supply attribute across availability zones according to the resource classification identifier specifically includes: Determine the topological attribute set according to the computing resources of the nodes across availability zones; Determine the resource supply capacity characteristics of the availability zone during computing according to the topological attribute set; Determine the resource supply attribute of the availability zone from the resource classification identifier and the resource supply capacity characteristics.

[0008] In this embodiment, the resource supply attribute represents the characteristics of the computing resources that the availability zone can provide under specific conditions.

[0009] In this embodiment, performing dynamic efficiency evaluation on the computing resource scheduling process to obtain the efficiency correlation characteristics between the resource fragmentation index and the load balance degree in the computing resource scheduling specifically includes: Determine the resource fragmentation index and the load balance degree in the computing resource scheduling process; Adopt a sliding window mechanism to capture the load migration trajectory associated with the resource fragmentation index and the load balance degree; Determine the efficiency correlation characteristics between the resource fragmentation index and the load balance degree in the computing resource scheduling according to the load migration trajectory.

[0010] In this embodiment, sharding the task allocation strategy of the computing resources according to the efficiency correlation characteristics and the resource supply attribute to generate a resource sharding rule with an elastic boundary in the computing resource scheduling specifically includes: Construct a resource sharding constraint space based on the elastic decay factor in the efficiency correlation characteristics and the failure boundary parameter in the resource supply attribute; Perform multi-dimensional decoupling analysis on the resource sharding constraint space and inject the elastic attenuation factor to form an initial sharding rule set; Perform elastic boundary calculation on the initial sharding rule set to obtain dynamic sharding parameters in resource scheduling; Determine the resource sharding rule with an elastic boundary in computing resource scheduling according to the dynamic sharding parameters.

[0011] In this embodiment, the resource sharding rule represents the rule for determining that each computing task should be assigned to the corresponding resource shard during the resource scheduling process.

[0012] In this embodiment, splitting the computing task into resource adjustment information matching the resource sharding rule specifically includes: Determine the initial task unit set during computing resource adjustment based on the resource sharding rule; Determine the fault isolation degree index during computing task allocation according to the initial task unit set; Perform resource mapping adaptation on the fault isolation degree index to obtain resource adjustment information.

[0013] In this embodiment, performing execution resource mapping on the resource adjustment information and dynamically correcting the resource allocation structure according to real-time monitoring data specifically includes: Use the dynamic weight allocation algorithm to map the atomic task units in the resource adjustment information to the target computing nodes to generate the initial resource allocation characteristics; Calculate the load volatility of the computing nodes through lightweight runtime monitoring; Generate the corrected resource allocation structure based on the initial resource allocation characteristics and the load volatility.

[0014] In a second aspect, the present application provides a computing resource management system based on cloud computing for executing a computing resource management method based on cloud computing. The computing resource management system includes: A data acquisition module for collecting the dynamic operation data of physical nodes and virtualization instances from a heterogeneous computing resource pool and synchronously obtaining the load characteristic parameters of user computing tasks; A resource classification module for performing resource demand analysis on the load characteristic parameters in a hybrid cloud architecture to obtain the resource classification identifiers corresponding to the computing nodes to be classified in the computing resources, and constructing the resource supply attributes across availability zones according to the resource classification identifiers; A task sharding module for performing dynamic efficiency evaluation on the computing resource scheduling process to obtain the efficiency correlation characteristics of the resource fragmentation index and the load balance degree in computing resource scheduling, and sharding the task allocation strategy of the computing resources according to the efficiency correlation characteristics and the resource supply attributes to generate a resource sharding rule with an elastic boundary in computing resource scheduling; A resource allocation module, configured to split a computing task into resource adjustment information matching the resource sharding rule, perform execution resource mapping on the resource adjustment information, and dynamically correct the resource allocation structure according to real-time monitoring data.

[0015] The technical solutions provided by the embodiments disclosed in this application have the following beneficial effects: Collect dynamic operation data of physical nodes and virtualization instances from a heterogeneous computing resource pool, and synchronously obtain load characteristic parameters of user computing tasks; under a hybrid cloud architecture, perform resource demand analysis on the load characteristic parameters to obtain resource classification identifiers corresponding to the to-be-classified computing nodes in the computing resources, and construct cross-availability zone resource supply attributes according to the resource classification identifiers; perform dynamic efficiency evaluation on the computing resource scheduling process to obtain efficiency correlation characteristics of resource fragmentation index and load balance degree in the computing resource scheduling, and slice the task allocation strategy of the computing resources according to the efficiency correlation characteristics and the resource supply attributes to generate a resource sharding rule with elastic boundaries in the computing resource scheduling; split the computing task into resource adjustment information matching the resource sharding rule, perform execution resource mapping on the resource adjustment information, and dynamically correct the resource allocation structure according to real-time monitoring data.

[0016] It can be seen that in this application, the efficiency of computing resources can be dynamically evaluated; among them, by collecting dynamic operation data of physical nodes and virtualization instances in real time, as well as the load characteristics of user computing tasks, the resource usage situation and task requirements can be comprehensively understood, providing an accurate basis for resource demand analysis and optimization decision-making, thereby improving the utilization efficiency of computing resources; by performing resource demand analysis on the load characteristic parameters, the resource demand characteristics and performance differences of different computing nodes can be accurately identified, and then cross-availability zone resource supply attributes can be constructed to ensure the reasonable allocation and scheduling of resources; through dynamic efficiency evaluation, the resource fragmentation and load balance degree problems in the computing resource scheduling can be captured in real time, generating flexible resource sharding rules, further improving the elasticity and adaptability of resource scheduling; by dynamically correcting the resource allocation structure, the resource mapping can be adjusted according to real-time monitoring data to make the computing task more matching the resource sharding rule, improving the system's adaptability to load changes, reducing resource waste, and enhancing the task execution efficiency and overall stability of the system.

[0017] In summary, the technical solution adopted in this application can accurately analyze the computing resource requirements under a hybrid cloud architecture to improve the utilization rate of computing resources. BRIEF DESCRIPTION OF THE DRAWINGS

[0018] To more clearly illustrate the technical solutions in the embodiments of the present application or the prior art, the following will briefly introduce the accompanying drawings required for the description of the embodiments or the prior art. Obviously, the accompanying drawings in the following description are only the embodiments of the present application. For those of ordinary skill in the art, without creative efforts, other drawings can also be obtained based on these drawings.

[0019] Figure 1 is a flowchart of a computing resource management method based on cloud computing provided by the present application; Figure 2 is a schematic flowchart of determining a resource classification identifier provided by the present application; Figure 3 is a schematic flowchart of determining a resource sharding rule provided by the present application; Figure 4 is a module structure diagram of a computing resource management system based on cloud computing provided by the present application. Detailed implementation manners

[0020] The following will clearly and completely describe the technical solutions in the embodiments of the present application with reference to the accompanying drawings in the embodiments of the present application. Obviously, the described embodiments are only a part of the embodiments of the present application, rather than all the embodiments. Based on the embodiments of the present application, all other embodiments obtained by those of ordinary skill in the art without creative efforts belong to the scope of protection of the present application.

[0021] The embodiments of the present application provide a computing resource management method and system based on cloud computing. The core is to collect the dynamic operation data of physical nodes and virtualization instances from a heterogeneous computing resource pool, and synchronously obtain the load characteristic parameters of user computing tasks; in a hybrid cloud architecture, perform resource demand analysis on the load characteristic parameters to obtain the resource classification identifiers corresponding to the computing nodes to be classified in the computing resources, and construct the resource supply attributes across availability zones according to the resource classification identifiers; perform dynamic efficiency evaluation on the computing resource scheduling process to obtain the efficiency correlation characteristics of the resource fragmentation index and the load balance degree in the computing resource scheduling, and shard the task allocation strategy of the computing resources according to the efficiency correlation characteristics and the resource supply attributes to generate a resource sharding rule with elastic boundaries in the computing resource scheduling; split the computing tasks into resource adjustment information matching the resource sharding rule, perform execution resource mapping on the resource adjustment information, and dynamically correct the resource allocation structure according to the real-time monitoring data.

[0022] Embodiment 1. To better understand the above technical solutions, the following will describe the above technical solutions in detail with reference to the accompanying drawings of the specification and specific implementation manners. Refer to Figure 1As shown, the figure is an exemplary flowchart of a cloud computing-based computing resource management method according to this embodiment of the present application. The computing resource management method includes the following steps: In step S1, collect the dynamic operation data of physical nodes and virtualization instances from the heterogeneous computing resource pool, and synchronously obtain the load characteristic parameters of user computing tasks.

[0023] When specifically implemented, collecting the dynamic operation data of physical nodes and virtualization instances from the heterogeneous computing resource pool can be achieved in the following ways: First, install lightweight data collection agents, such as Telegraf or NodeExporter, on physical nodes to collect key metrics such as CPU utilization, memory occupancy, disk I / O, and network throughput, and send the data to time series databases such as Prometheus. Second, in the virtualization environment, obtain the resource usage of each virtual instance through the virtualization management layer, and store the data using InfluxDB or VictoriaMetrics for efficient querying and analysis. Finally, use the data bus to converge data from different sources, and use a stream processing framework for data cleaning, anomaly detection, and format standardization, and use the processed data as dynamic operation data.

[0024] It should be noted that in the present application, the heterogeneous computing resource pool represents a collection of computing resources composed of computing devices with different architectures and types; physical nodes represent physical servers or computing devices that directly provide computing, storage, and network resources, support the operation of operating systems and virtualization instances, and can be used as the basic units of the computing resource pool; virtualization instances represent virtual machines or containers running on physical nodes, with independent computing, storage, and network resources, and support isolation and dynamic scheduling; dynamic operation data represents the key metrics that continuously change during the operation of computing resources.

[0025] In addition, when specifically implemented, synchronously obtaining the load characteristic parameters of user computing tasks can be achieved in the following ways: First, in the task submission phase, integrate the job scheduling system, record the resource request parameters of the task, including the number of CPU cores, memory size, GPU requirements, etc., and store them in the database. Second, during the task execution process, use monitoring tools such as cAdvisor or Prometheus to collect real-time runtime metrics such as CPU utilization, memory occupancy, and I / O load of the task, and perform data streaming transmission through the Kafka message queue. Then, use the stream processing framework to preprocess the data, calculate the load change trend, peak, and fluctuation of the task, and analyze the resource consumption pattern of the task based on the time series prediction algorithm. Finally, store the task load characteristic parameters in the data warehouse, and obtain the load characteristic parameters by reading the data warehouse.

[0026] It should be noted that in this application, the load characteristic parameters represent the usage patterns and their changing trends of the CPU, memory, I / O, and network resources during the execution of the computing task.

[0027] In step S2, under the hybrid cloud architecture, perform resource requirement analysis on the load characteristic parameters to obtain the resource classification identifier corresponding to the computing node to be classified in the computing resources, and construct the resource supply attribute across availability zones according to the resource classification identifier.

[0028] Preferably, in this embodiment, under the hybrid cloud architecture, perform resource requirement analysis on the load characteristic parameters to obtain the resource classification identifier corresponding to the computing node to be classified in the computing resources, and refer to Figure 2 As shown in the figure, which is a schematic flowchart of determining the resource classification identifier in some embodiments of this application. The determination of the resource classification identifier in this embodiment can be implemented by the following steps: In step S21, eliminate the differences in the load characteristic parameters to obtain a standardized load vector; In step S22, perform multi-dimensional feature decoupling on the standardized load vector to obtain the computing throughput fluctuation coefficient; In step S23, determine the discrete probability distribution corresponding to the computing node to be classified in the computing resources according to the computing throughput fluctuation coefficient; In step S24, calculate the resource classification identifier corresponding to the computing node to be classified in the computing resources from the discrete probability distribution.

[0029] In specific implementation, first, collect the load characteristic parameters of computing nodes from different availability zones, such as CPU utilization rate, memory occupancy rate, disk I / O throughput, network bandwidth utilization rate, etc. Since the computing nodes in different availability zones may have different hardware configurations, and different hardware configurations include: CPU main frequency, storage speed; which leads to inconsistent measurement standards, so standardization processing is required. The Z-score standardization method is adopted to convert the load data of all computing nodes into standardized data with zero mean and unit variance, and this standardized data is used as the standardized load vector. Then, the load characteristics of the computing nodes are multi-dimensional, including CPU, memory, I / O, etc., so methods such as principal component analysis or independent component analysis need to be used for dimensionality reduction and feature decoupling. The PCA method is adopted to calculate the covariance matrix and perform eigenvalue decomposition, extract the first few main eigenvectors, so that the data has the largest variance in these dimensions, calculate the time series correlation of the load characteristics, use wavelet transform or fast Fourier transform to analyze the frequency components of each load characteristic, strip the noise, and calculate the computing throughput fluctuation coefficient of the computing node through the method based on information entropy, that is: computing throughput fluctuation coefficient of the computing node = standard deviation of computing throughput / mean of computing throughput. Then, use kernel density estimation to calculate the probability density function of the throughput fluctuation, and then calculate the cumulative distribution function to obtain the probability of the computing node at different throughput levels. Classify according to the quantiles of the computing node in the probability distribution, where the quantiles can be 25%, 50%, 75%, which are not limited here, that is, obtain the discrete probability distribution corresponding to the computing nodes to be classified in the computing resources. Finally, according to the discrete probability distribution of the computing nodes, divide the resource categories and generate resource classification identifiers; adopt the K-Means clustering or Gaussian mixture model method to classify the computing nodes according to the throughput fluctuation coefficient, and the generated resource classification identifiers include: compute-intensive, suitable for HPC computing; memory-intensive, suitable for large-scale data processing; storage-optimized, suitable for storage-intensive tasks; low-power, suitable for ARM architecture.

[0030] It should be noted that in this application, the computing nodes to be classified in the computing resources refer to the computing nodes that have not been classified in the computing resource pool. These computing nodes need to be analyzed according to their load characteristics to determine the appropriate resource classification and scheduling strategies; the standardized load vector refers to the standardized data for cross-regional analysis of computing resource allocation; the computing throughput fluctuation coefficient represents the stability of the throughput capacity of the computing node, and the larger the value of the computing throughput fluctuation coefficient, the greater the computing load fluctuation; the discrete probability distribution represents the probability distribution of the computing node at different load levels; the resource classification identifier is the information that identifies the resource characteristics of the computing node to optimize the scheduling and matching of computing tasks.

[0031] In this embodiment, constructing the resource supply attributes across availability zones according to the resource classification identifier can be implemented by the following steps: Determine the topological attribute set according to the computing resources of the cross-availability zone nodes; Determine the resource supply capacity characteristics of the availability zones during computing according to the topological attribute set; Determine the resource supply attributes of the availability zones from the resource classification identifier and the resource supply capacity characteristics.

[0032] Specifically, in implementation, first, use the API of the cloud platform to obtain the detailed information of the availability zones, including the physical location, network latency, bandwidth, storage type, etc. of each node; on this basis, collect the computing node configurations of different availability zones, and perform topological modeling of the resources between nodes through the cloud resource management tool; according to the network connection, latency and bandwidth, use the graph theory method to perform topological structure modeling on the nodes of each availability zone to form a topological attribute set including computing nodes and network connectivity relationships. Then, on the basis of obtaining the topological attribute set, further obtain the resource supply capacity characteristics of each availability zone through the computing resource monitoring tool, mainly including CPU utilization rate, memory usage, network bandwidth, storage capacity, etc.; use the load balancing algorithm and resource allocation strategy to analyze the resource supply capacity of different availability zones, and calculate the resource supply capacity of each availability zone under a given load; for example, use the weighted resource pool method to evaluate the capacity according to the load of the nodes, available resources and the topological relationship of the availability zones; through multi-dimensional analysis, combine the time series prediction model to estimate the future change trend of resource requirements, that is, obtain the resource supply capacity characteristics. Finally, according to the resource classification identifier and the resource supply capacity characteristics of the cross-availability zone nodes, use the decision tree algorithm or the weighted average model to comprehensively evaluate the resource supply attributes of each availability zone; the resource supply attributes can include resource availability, load balancing degree, latency, etc.; among them, the decision tree model: classify the resources of each availability zone by training the decision tree to judge the type of tasks it is suitable to carry; the weighted average method: after weighting the resource supply capacity characteristics and the classification identifier, calculate the supply capacity of each availability zone, obtain the comprehensive attribute, and use this comprehensive attribute as the resource supply attribute of the availability zone.

[0033] It should be noted that in this application, the availability zone represents an independent physical data center area in the cloud computing platform, and each availability zone has an independent power supply, network and cooling system to ensure high availability and fault tolerance; the topological attribute set represents a set of information describing the connections, network latency, bandwidth, etc. between the cross-availability zone computing resources; the resource supply capacity characteristics represent the capabilities of computing, storage and network resources that each availability zone can provide within a specific time period; the resource supply attributes represent the characteristics of the computing resources that the availability zone can provide under specific conditions.

[0034] In step S3, a dynamic performance evaluation is performed on the computing resource scheduling process to obtain the performance correlation characteristics of the resource fragmentation index and the load balance degree in the computing resource scheduling. According to the performance correlation characteristics and the resource supply attributes, the task allocation strategy of the computing resources is sliced to generate a resource slicing rule with an elastic boundary in the computing resource scheduling.

[0035] In this embodiment, the dynamic performance evaluation of the computing resource scheduling process to obtain the performance correlation characteristics of the resource fragmentation index and the load balance degree in the computing resource scheduling can be implemented by the following steps: Determine the resource fragmentation index and the load balance degree in the computing resource scheduling process; Adopt a sliding window mechanism to capture the load migration trajectory associated with the resource fragmentation index and the load balance degree; Determine the performance correlation characteristics of the resource fragmentation index and the load balance degree in the computing resource scheduling according to the load migration trajectory.

[0036] In specific implementation, first, the resource fragmentation index in the computing resource scheduling process can be determined based on parameters such as memory fragmentation, CPU idle cores, and disk space waste, that is: resource fragmentation index = (U / T) - (S / T), where U represents the allocated resources, S represents the actual used resources, and T represents the total resources. The higher the resource fragmentation index, the lower the utilization efficiency of the resources; the load balance degree can be represented by the standard deviation or the imbalance coefficient. The lower the load balance degree, the more uniform the load, which will not be elaborated here. Then, define a time window with a fixed size for real-time monitoring of the load and resource fragmentation situation during the scheduling process; within each time window, calculate the resource fragmentation index and the load balance degree within the current window and record their change trajectories. By setting the window size, such as 5 minutes or 10 minutes, the short-term fluctuations in the resource scheduling process can be captured; continuous calculation and data update are performed using the sliding window to ensure the accuracy of real-time monitoring. Each time the sliding window moves forward, the data of the previous time period is discarded and new data is added for calculation, and each calculation result is used as the load migration trajectory. Finally, by analyzing the resource fragmentation index and the load balance degree captured in multiple time windows, a correlation analysis method, such as the Pearson correlation coefficient or the Spearman rank correlation coefficient, is used to evaluate the relationship between the two, that is, the performance correlation characteristics of the resource fragmentation index and the load balance degree. If the correlation is strong (close to 1 or -1), it indicates that there is a close connection between the resource fragmentation index and the load balance degree.

[0037] It should be noted that in this application, the computing resource scheduling process refers to the process of dynamically allocating computing resources to different nodes or availability zones according to the requirements of computing tasks, the availability of resources, and the load conditions, ensuring efficient resource utilization and task execution; the sliding window mechanism is to set a window with a fixed time interval to monitor and update the changes in the resource fragmentation index and the load balance degree in real time to capture the trend of time-series data; the resource fragmentation index refers to the degree of inefficient utilization of computing resources, reflecting the resource waste in the computing resource scheduling process; the load balance degree refers to the degree of uniformity of load distribution among multiple nodes. A low load balance degree indicates a more uniform load distribution; the load migration trajectory refers to the path and historical record of the migration of tasks or loads between different computing nodes or availability zones during the computing resource scheduling process, reflecting the dynamic change process of the load; the efficacy correlation feature refers to the relationship feature between the resource fragmentation index and the load balance degree.

[0038] Preferably, in this embodiment, the task allocation strategy of computing resources is sliced according to the efficacy correlation feature and the resource supply attribute to generate a resource slicing rule with an elastic boundary in computing resource scheduling. Refer to Figure 3 As shown, this figure is a schematic flow chart of determining the resource slicing rule in some embodiments of this application. The determination of the resource slicing rule in this embodiment can be implemented by the following steps: In step S31, based on the elastic decay factor in the efficacy correlation feature and the failure boundary parameter in the resource supply attribute, a resource slicing constraint space is constructed; In step S32, multi-dimensional decoupling analysis is performed on the resource slicing constraint space, and the elastic decay factor is injected to form an initial slicing rule set; In step S33, elastic boundary calculation is performed on the initial slicing rule set to obtain dynamic slicing parameters in resource scheduling; In step S34, according to the dynamic slicing parameters, a resource slicing rule with an elastic boundary in computing resource scheduling is determined.

[0039] In specific implementation, first, when constructing the resource sharding constraint space, a constraint optimization algorithm can be used to simultaneously consider the elastic decay factor and the failure boundary parameter, define the allocation constraints of resources in the time, space, and load dimensions, and form the resource sharding constraint space. Next, multi-dimensional decoupling analysis is to perform multi-dimensional decoupling analysis on the resource sharding constraint space, which involves different dimensions of resources. Through methods such as principal component analysis or independent component analysis, the constraint space is decomposed into different dimensions so that each dimension can be analyzed independently and the elastic decay factor can be injected; the injection of the elastic decay factor is to introduce the elastic decay factor into the resource sharding constraint space and adjust the weight of each dimension. The weighted average method or the adaptive adjustment algorithm can be used to adjust the resource allocation strategy in real time, enabling the system to dynamically adjust the elasticity of resources when the load fluctuates or resources are scarce, and taking the final result as the initial sharding rule set. Then, elastic boundary calculation is based on the elastic decay factor of resources and the resource supply capacity, and calculates the maximum load range of each resource shard through a dynamic boundary algorithm. This process can adopt dynamic programming or simulated annealing algorithms to calculate appropriate boundary values according to the real-time resource status and load changes; the determination of dynamic sharding parameters is through elastic boundary calculation to obtain the dynamic sharding parameters of each resource shard, such as the allocated computing resource amount, the boundary value of task migration, the maximum threshold of resource supply, etc. These dynamic sharding parameters will be used to adjust the task allocation strategy to make resource allocation more flexible and efficient. Finally, the rule engine or the constraint optimization algorithm is used to formulate the final resource sharding rules, which ensure that in the actual scheduling process, each task can allocate resources according to the calculated elastic boundary to maintain the efficient use of resources; the resource sharding rules can be dynamically adjusted based on factors such as the priority of tasks, the actual demand for resources, and the load change trend. For example, the priority scheduling strategy or the task migration strategy can be adopted to adjust the allocation ratio of tasks in different resource shards according to the dynamic sharding parameters.

[0040] It should be noted that in this application, the task allocation strategy refers to the way of determining the allocation of tasks among different computing nodes or resource shards according to the requirements of tasks and the availability of resources during the computing resource scheduling process; the elastic boundary refers to the flexible range in resource allocation, allowing adjustment according to load fluctuations or resource changes; the elastic decay factor represents the degree to which the effectiveness of resources gradually decays during the resource scheduling process as the resource usage time prolongs or the load fluctuates; the failure boundary parameter represents the maximum load threshold of resources or the boundary of resource supply, and when the resource usage reaches this boundary, the resource may fail or be unable to meet the task requirements, which can be dynamically obtained through the load monitoring and resource health check mechanisms; the resource shard constraint space represents the constraint range of resource allocation under specific conditions; the initial shard rule set represents the set of rules guiding task allocation in the resource scheduling process; the dynamic shard parameter represents the metrics guiding resource sharding and task allocation in the resource scheduling process; the resource shard rule represents the rule for determining the corresponding resource shard to which each computing task should be allocated during the resource scheduling process.

[0041] In step S4, the computing task is split into resource adjustment information that matches the resource shard rule, the resource adjustment information is subjected to execution resource mapping, and the resource allocation structure is dynamically corrected according to the real-time monitoring data.

[0042] In this embodiment, splitting the computing task into resource adjustment information that matches the resource shard rule can be implemented by the following steps: Determine the initial task unit set during computing resource adjustment based on the resource shard rule; Determine the fault isolation degree index during computing task allocation according to the initial task unit set; Perform resource mapping adaptation on the fault isolation degree index to obtain resource adjustment information.

[0043] In specific implementation, first, a task splitting algorithm, such as a divide-and-conquer algorithm or a workflow scheduling method, is used to split large-scale computing tasks into independent and parallelizable small task units. Each task unit is initially allocated according to the constraints of resource sharding rules, such as resource type, resource capacity, and load balancing requirements, that is, an initial task unit set during computing resource adjustment is obtained. A load balancing algorithm can be used to dynamically allocate the initial task units to ensure that each unit after the computing task is split can match the resource sharding rules. Then, a fault tolerance algorithm and a redundancy allocation mechanism are used to evaluate the task unit set to ensure that each task unit has a certain fault isolation ability during resource allocation. At this time, the fault isolation degree can be quantified by calculating the resource dependency relationship between each task unit and other units. Using the isolation degree calculation formula, the redundancy and dependency of resources are defined according to the resource sharing degree between task units, and a fault isolation degree index is constructed. For example, the lower the resource dependency, the higher the isolation degree index, indicating that the task unit can execute independently in case of a failure. Finally, resource mapping adaptation dynamically maps the resources of the computing task units through the fault isolation degree index and the resource sharding rules. In the mapping process, resources that meet the isolation degree conditions are selected for allocation according to the fault isolation degree requirements of the task units. The resource adaptation algorithm is used to allocate resources to ensure that each task unit meets its fault isolation degree requirements during resource mapping. In this way, the resource allocation process can be optimized, and it is ensured that the computing task has a certain fault tolerance ability during execution. Through the fault tolerance scheduling mechanism, it is ensured that when the computing task is allocated to a certain node, the node can execute independently of other nodes, thus avoiding single point of failure, and resource adjustment information is obtained.

[0044] It should be noted that in this application, the initial task unit set refers to the set of basic execution units obtained after splitting the computing task; the fault isolation degree index represents the ability to independently handle faults of the computing task unit during resource allocation; the resource adjustment information represents the relevant data for dynamically adjusting the resource allocation scheme.

[0045] In this embodiment, the execution resource mapping of the resource adjustment information and the dynamic correction of the resource allocation structure according to the real-time monitoring data can be implemented by the following steps: The atomic task units in the resource adjustment information are mapped to the target computing nodes by using a dynamic weight allocation algorithm to generate an initial resource allocation feature; The load volatility of the computing nodes is monitored through lightweight runtime; A corrected resource allocation structure is generated based on the initial resource allocation feature and the load volatility.

[0046] In the specific implementation, first, a dynamic weight allocation algorithm is used to map each atomic task unit to the target computing node according to its resource requirements. These weights are usually adjusted dynamically based on the priority of the task, the computing resource requirements and the availability of the resource node. The mapped result is used as the initial resource allocation feature. During the mapping process, the load balancing algorithm is used to ensure the balanced allocation of task units and avoid overloading a node, taking into account the node load, network bandwidth and the computing power of the node. Then, the load fluctuation rate is calculated by monitoring the CPU usage, memory occupancy, disk I / O and network bandwidth of the node. The load fluctuation rate reflects whether the load of the computing node is stable. Severe fluctuations indicate large load changes. Lightweight runtime monitoring: Use lightweight monitoring tools to collect the load data of the computing node in real time. The monitoring module collects the resource usage of the computing node in the background with low overhead and calculates the load fluctuation rate. The load fluctuation of the node is modeled by the time series data analysis method to calculate the load change trend and volatility of each node. Finally, the initial resource allocation feature generated by the dynamic weight allocation algorithm describes how each task unit is allocated to the target computing node. The initial resource allocation feature usually contains information such as the resource requirements of the task unit and the node resource capacity. The impact of load fluctuation rate: According to the real-time load fluctuation rate, the resource allocation structure can be adjusted. For example, when the load fluctuation of a node is large, it may be necessary to adjust the task allocation of the node to reduce the load pressure of the node and avoid overload. The initial resource allocation structure is corrected according to the load fluctuation rate through the resource adjustment algorithm.

[0047] It should be noted that, in the present application, the atomic task unit means decomposing the computing task into the smallest independent execution units, each unit can independently complete the calculation and coordinate with other units, and each atomic task unit contains the computing requirements, resource requirements, fault tolerance and other information of the task; the initial resource allocation characteristics represent the data describing the resource requirements and allocation relationship between the task unit and the computing node; the load fluctuation rate represents the rate of change of the load of the computing node within a certain period of time; the resource allocation structure represents the resource allocation plan between the task unit and the computing node during the computing resource scheduling process.

[0048] It can be seen that in this application, the effectiveness of computing resources can be dynamically evaluated. Among them, by collecting the dynamic operation data of physical nodes and virtualized instances in real time, as well as the load characteristics of user computing tasks, it is possible to comprehensively understand the resource usage and task requirements, provide an accurate basis for resource requirement analysis and optimization decision-making, thereby improving the utilization efficiency of computing resources. By performing resource requirement analysis on the load characteristic parameters, it is possible to accurately identify the resource requirement characteristics and performance differences of different computing nodes, and then construct the resource supply attributes across availability zones to ensure the reasonable allocation and scheduling of resources. Through dynamic effectiveness evaluation, it is possible to capture the resource fragmentation and load balance problems in computing resource scheduling in real time, generate flexible resource sharding rules, and further improve the elasticity and adaptability of resource scheduling. By dynamically modifying the resource allocation structure, it is possible to adjust the resource mapping according to the real-time monitoring data, make the computing tasks more matched with the resource sharding rules, improve the system's adaptability to load changes, reduce resource waste, and enhance the task execution efficiency and the overall stability of the system.

[0049] In summary, the technical solution adopted in this application can accurately analyze the computing resource requirements under the hybrid cloud architecture to improve the utilization rate of computing resources.

[0050] Embodiment 2. This application provides a computing resource management system based on cloud computing. Refer to Figure 4 As shown in the figure, which is a module structure diagram of the computing resource management system based on cloud computing according to this embodiment of this application, the computing resource management system includes: A data acquisition module 100, configured to collect the dynamic operation data of physical nodes and virtualized instances from a heterogeneous computing resource pool, and synchronously obtain the load characteristic parameters of user computing tasks; A resource classification module 200, configured to perform resource requirement analysis on the load characteristic parameters under the hybrid cloud architecture to obtain the resource classification identifiers corresponding to the computing nodes to be classified in the computing resources, and construct the resource supply attributes across availability zones according to the resource classification identifiers; A task sharding module 300, configured to perform dynamic effectiveness evaluation on the computing resource scheduling process to obtain the effectiveness correlation characteristics of resource fragmentation index and load balance degree in computing resource scheduling, and shard the task allocation strategy of computing resources according to the effectiveness correlation characteristics and the resource supply attributes to generate resource sharding rules with elastic boundaries in computing resource scheduling; A resource allocation module 400, configured to split the computing tasks into resource adjustment information matching the resource sharding rules, perform execution resource mapping on the resource adjustment information, and dynamically modify the resource allocation structure according to the real-time monitoring data.

[0051] This application is described with reference to the flowcharts and / or block diagrams of methods, apparatus (systems), and computer program products according to embodiments of the present application. It should be understood that each flow and / or block in the flowcharts and / or block diagrams can be implemented by computer program instructions, and the combination of flows and / or blocks in the flowcharts and / or block diagrams can also be implemented. These computer program instructions can be provided to the processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing devices to generate a machine, so that the instructions executed by the processor of the computer or other programmable data processing devices generate a means for implementing the functions specified in one Figure 1 flow or multiple flows and / or blocks Figure 1 block or multiple blocks.

[0052] Those of ordinary skill in the art can understand that all or part of the steps in the various methods of the above embodiments can be completed by instructing relevant hardware through a program. This program can be stored in a computer-readable storage medium, which includes read-only memory (ROM), random access memory (RAM), programmable read-only memory (PROM), erasable programmable read-only memory (EPROM), one-time programmable read-only memory (OTPROM), electrically-erasable programmable read-only memory (EEPROM), compact disc read-only memory (CD-ROM), or other optical disc memories, magnetic disc memories, tape memories, or any other medium that can be used to carry or store data and is computer-readable.

[0053] It should also be noted that the term "comprising" or any other variant thereof is intended to cover non-exclusive inclusion, so that a process, method, commodity, or device comprising a series of elements not only includes those elements, but also includes other elements not expressly listed, or further includes elements inherent to such process, method, commodity, or device. Without further limitation, an element defined by the statement "comprising one..." does not exclude the existence of another identical element in the process, method, commodity, or device comprising the element.

Claims

1. A computing resource management method based on cloud computing, characterized in that, The described computing resource management method includes the following steps: Collect the dynamic operation data of physical nodes and virtualization instances from the heterogeneous computing resource pool, and synchronously obtain the load characteristic parameters of user computing tasks; Under the hybrid cloud architecture, conduct resource demand analysis on the load characteristic parameters to obtain the resource classification identifiers corresponding to the computing nodes to be classified in the computing resources, and construct the resource supply attributes across availability zones according to the resource classification identifiers; Conduct dynamic efficiency evaluation on the computing resource scheduling process to obtain the efficiency correlation characteristics of the resource fragmentation index and load balance degree in the computing resource scheduling, and slice the task allocation strategy of the computing resources according to the efficiency correlation characteristics and the resource supply attributes to generate a resource slicing rule with elastic boundaries in the computing resource scheduling; Split the computing tasks into resource adjustment information matching the resource slicing rule, perform execution resource mapping on the resource adjustment information, and dynamically correct the resource allocation structure according to the real-time monitoring data.

2. The method for managing computing resources based on cloud computing according to claim 1, wherein, Under the hybrid cloud architecture, the specific steps for conducting resource demand analysis on the load characteristic parameters to obtain the resource classification identifiers corresponding to the computing nodes to be classified in the computing resources include: Eliminate the differences in the load characteristic parameters to obtain a standardized load vector; Perform multi-dimensional feature decoupling on the standardized load vector to obtain the computing throughput fluctuation coefficient; Determine the discrete probability distribution corresponding to the computing nodes to be classified in the computing resources according to the computing throughput fluctuation coefficient; Calculate the resource classification identifiers corresponding to the computing nodes to be classified in the computing resources from the discrete probability distribution.

3. The method for managing computing resources based on cloud computing according to claim 1, wherein The specific steps for constructing the resource supply attributes across availability zones according to the resource classification identifiers include: Determine the topological attribute set according to the computing resources of the nodes across availability zones; Determine the resource supply capacity characteristics of the availability zones during computing according to the topological attribute set; Determine the resource supply attributes of the availability zones from the resource classification identifiers and the resource supply capacity characteristics.

4. The method for managing computing resources based on cloud computing according to claim 1, wherein The resource supply attribute represents the characteristics of the computing resources that the availability zone can provide under specific conditions.

5. The method for managing computing resources based on cloud computing according to claim 1, characterized in that, The specific steps for conducting dynamic efficiency evaluation on the computing resource scheduling process to obtain the efficiency correlation characteristics of the resource fragmentation index and load balance degree in the computing resource scheduling include: Determine the resource fragmentation index and load balance degree in the computing resource scheduling process; Adopt a sliding window mechanism to capture the load migration trajectory associated with the resource fragmentation index and the load balance degree; Determine the efficiency correlation characteristics of the resource fragmentation index and load balance degree in the computing resource scheduling according to the load migration trajectory.

6. The computational resource management method based on cloud computing according to claim 1, characterized in that The specific steps for slicing the task allocation strategy of the computing resources according to the efficiency correlation characteristics and the resource supply attributes to generate a resource slicing rule with elastic boundaries in the computing resource scheduling include: Based on the elastic decay factor in the efficiency correlation characteristics and the failure boundary parameter in the resource supply attributes, construct a resource slicing constraint space; Conduct multi-dimensional decoupling analysis on the resource slicing constraint space and inject the elastic decay factor to form an initial slicing rule set; Perform elastic boundary calculation on the initial slicing rule set to obtain the dynamic slicing parameters in the resource scheduling; Determine a resource fragmentation rule with an elastic boundary in computing resource scheduling according to the dynamic fragmentation parameter.

7. The method for managing computing resources based on cloud computing according to claim 1, wherein The resource fragmentation rule represents the rule for determining the corresponding resource fragment to which each computing task should be allocated during the resource scheduling process.

8. The method for managing computing resources based on cloud computing according to claim 1, wherein Splitting the computing task into resource adjustment information matching the resource fragmentation rule specifically includes: Determining an initial task unit set during computing resource adjustment based on the resource fragmentation rule; Determining a fault isolation degree index during computing task allocation according to the initial task unit set; Performing resource mapping adaptation on the fault isolation degree index to obtain resource adjustment information.

9. The method for managing computing resources based on cloud computing according to claim 1, characterized in that Performing execution resource mapping on the resource adjustment information and dynamically correcting the resource allocation structure according to real-time monitoring data specifically includes: Using a dynamic weight allocation algorithm to map the atomic task units in the resource adjustment information to target computing nodes to generate an initial resource allocation feature; Monitoring the load volatility of computing nodes through lightweight runtime monitoring; Generating a corrected resource allocation structure based on the initial resource allocation feature and the load volatility.

10. A computing resource management system based on cloud computing, which is used to execute a computing resource management method based on cloud computing as described in any one of claims 1 to 9, characterized in that, The computing resource management system includes: A data acquisition module, configured to collect dynamic operation data of physical nodes and virtualization instances from a heterogeneous computing resource pool, and synchronously obtain load characteristic parameters of user computing tasks; A resource classification module, configured to perform resource requirement analysis on the load characteristic parameters in a hybrid cloud architecture to obtain resource classification identifiers corresponding to computing nodes to be classified in computing resources, and construct a resource supply attribute across availability zones according to the resource classification identifiers; A task fragmentation module, configured to perform dynamic efficiency evaluation on the computing resource scheduling process to obtain an efficiency correlation feature of resource fragmentation index and load balance degree in computing resource scheduling, and fragment the task allocation strategy of computing resources according to the efficiency correlation feature and the resource supply attribute to generate a resource fragmentation rule with an elastic boundary in computing resource scheduling; A resource allocation module, configured to split a computing task into resource adjustment information matching the resource fragmentation rule, perform execution resource mapping on the resource adjustment information, and dynamically correct the resource allocation structure according to real-time monitoring data.

Citation Information

Cited By

  • Resource allocation method and device, equipment and medium

    CN121116651A