Computing power sharing system and method in multi-tenant environment

Through real-time monitoring and dynamic configuration of the multi-tenant computing power sharing system, the problems of waste of resources and insufficient fault diagnosis are solved, efficient and flexible resource management and fault repair are achieved, and system stability and business continuity are ensured.

CN120276861AInactive Publication Date: 2025-07-08YUNJU DATA TECH (SHANGHAI) CO LTD
View PDF 0 Cites 6 Cited by

Patent Information

Application Number
CN202510425585.X
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-07
Publication Date
2025-07-08
Estimated Expiration
Not applicable · inactive patent

AI Technical Summary

Technical Problem

The existing technology is difficult to optimize resource utilization in a multi-tenant environment, resulting in waste of resources, increased energy consumption and high operation and maintenance costs, and lack of effective failure prevention and rapid diagnosis mechanisms, affecting business continuity.

Method used

Through the computing power monitoring module, load analysis and prediction module, task priority determination module, resource optimization configuration module and fault management module, real-time resource monitoring and dynamic configuration are realized, task priority and fault diagnosis are optimized, and resource efficiency utilization and system stability are ensured.

Benefits of technology

It realizes the accuracy and flexibility of resource allocation, avoids waste, improves response capabilities, ensures that computing resources are fully utilized during peak demand periods, reduces costs caused by fault response delays, and improves system stability and reliability.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120276861A_ABST
    Figure CN120276861A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of computing power sharing, in particular to a computing power sharing system and method in a multi-tenant environment, and the system comprises a computing power monitoring module, a load analysis and prediction module, a task priority judgment module, a resource optimization configuration module and a fault management module. According to the invention, by implementing real-time resource monitoring and data analysis and accurately capturing the speed and mode of resource consumption, the resource allocation is more accurate, the waste of computing power resources is avoided, the response capability is improved, the load prediction is generated by using historical and real-time data, the computing power resources can be ensured to be fully utilized in the peak period of demand, overload is avoided, and the service life of the system is prolonged. Through intelligent analysis of the resource utilization rate and the predicted completion time and optimization of priority ranking of the tasks, the emergency tasks can be rapidly responded, implementation of a real-time computing power distribution strategy is achieved, high-efficiency computing power resource use is guaranteed, the stability and reliability of the system are improved, and the cost problem caused by fault response delay is reduced.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of computing power sharing, and particularly to a computing power sharing system and method in a multi-tenant environment. Background Art

[0002] Computing power sharing technology is an important cloud computing application, which mainly focuses on how to efficiently allocate and manage computing resources among different users or organizations. This technology allows multiple users to execute complex computing tasks by sharing infrastructure, thereby optimizing resource utilization and reducing costs. Computing power sharing can be implemented based on different architectures, including centralized and distributed architectures. The centralized architecture is usually managed by a single organization, and all computing resources are concentrated in one location. The distributed architecture disperses computing resources at multiple locations and supports a wider geographical distribution and flexible resource scheduling through network connections.

[0003] Among them, the computing power sharing system in a multi-tenant environment is a system for multiple tenants (i.e., different user or customer organizations) to jointly use and manage computing resources. In this environment, each tenant can independently use the allocated computing power resource part according to their own needs without affecting the operations and data security of other tenants. Its main purpose is to improve the utilization efficiency of computing power resources, reduce costs, and provide greater flexibility and scalability. Computing power sharing is common in cloud service providers, where resources such as servers, storage, and network devices are shared by multiple tenants while maintaining their respective isolation and security.

[0004] The prior art adopts a static resource allocation mode, which is difficult to achieve the optimization of resource utilization. Although the centralized architecture is easy to manage, it is difficult to accurately allocate resources when meeting the complex computing needs of multi-tenants, easily causing resource surplus at some nodes and shortage at other nodes. Although the distributed architecture supports a wide geographical distribution, it is insufficient in real-time monitoring and dynamic resource allocation, lacking flexibility and unable to quickly adapt to the temporarily increased resource requirements of tenants, resulting in low processing efficiency. In addition, due to the inflexible resource allocation strategy, it often leads to resource waste, increasing energy consumption and operation and maintenance costs. The prior art lacks an effective fault prevention and rapid diagnosis mechanism. Once a resource configuration error or system failure occurs, the complex recovery process seriously affects business continuity, especially in business scenarios with a high dependence on computing resources, and it is difficult to effectively support efficient and secure resource management. Summary of the Invention

[0005] To address the technical problems in the prior art, such as the difficulty in achieving optimal resource utilization with the static resource allocation mode, the centralized architecture being easy to manage but difficult to accurately allocate resources when meeting the complex computing requirements of multi-tenants, resulting in resource surplus at some nodes and shortage at others, the distributed architecture supporting wide geographical distribution but lacking in real-time monitoring and dynamic resource allocation, lacking flexibility and being unable to quickly adapt to the temporarily increased resource requirements of tenants, leading to low processing efficiency. In addition, due to the inflexible resource allocation strategy, it often causes resource waste, increases energy consumption and operation and maintenance costs, and the prior art lacks effective fault prevention and rapid diagnosis mechanisms. Once resource configuration errors or system failures occur, the complex recovery process seriously affects business continuity, especially in business scenarios with high dependence on computing resources, and it is difficult to effectively support efficient and secure resource management, the embodiments of the present invention provide a computing power sharing system and method in a multi-tenant environment. The technical solutions are as follows: On the one hand, a computing power sharing system in a multi-tenant environment is provided. The system includes: The computing power monitoring module monitors the utilization rates of CPU cores, memory blocks, and storage partitions based on the virtual computing unit usage data of multi-tenants, integrates the data according to resource types and usage frequencies, analyzes the resource usage patterns of multi-tenants, and statistically analyzes the resource occupancy and change trends to obtain an overview of the dynamic resource usage; The load analysis and prediction module collects the time series data of CPU cores and memory blocks based on the overview of the dynamic resource usage, analyzes the task scheduling behaviors of multi-tenants, predicts the resource load trends in future time periods, and determines the interval of peak computing power resource time periods to obtain peak period prediction indicators; The task priority determination module evaluates the resource consumption and completion time of tasks according to the task scheduling requirements of multi-tenants based on the peak period prediction indicators, classifies the tasks, calculates the priority of each task, and obtains a task priority mapping table; The resource optimization and configuration module analyzes the task priorities and computing power resource usage data based on the task priority mapping table, reconfigures the existing computing power resources, and reallocates virtual computing units to high-priority tasks to optimize the computing power resource configuration efficiency and obtain an overview of the real-time computing power allocation; The fault management module monitors the abnormal states of multi-tenant resources based on the overview of the real-time computing power allocation, analyzes the load shift nodes and task migration situations in the abnormal states, conducts fault diagnosis, and performs fault repair operations to restore the computing power resources to the normal state and obtain a log of the recovery operation effects.

[0006] On the other hand, the dynamic overview of resource usage includes utilization statistics, consumption speed metrics, and space occupancy data. The peak period prediction metrics specifically include CPU usage peaks, memory occupancy peaks, and resource demand prediction results. The task priority mapping table includes resource demand classification results, completion time limit assessment results, and priority ranking results. The real-time computing power allocation overview includes resource configuration status and virtual unit allocation data. The recovery operation effect log specifically includes anomaly detection status, fault diagnosis results, and recovery status update results.

[0007] On the other hand, the computing power monitoring module includes a resource tracking sub-module, a data integration sub-module, and a trend analysis sub-module; The resource tracking sub-module monitors the real-time utilization of CPU cores, memory blocks, and storage partitions based on the usage data of multi-tenant virtual computing units, and real-time tracks the resource consumption speed and space occupancy to obtain a resource consumption snapshot; The data integration sub-module classifies and integrates the data according to resource types and usage frequencies based on the resource consumption snapshot, and performs quality inspection and cleaning on the integrated data to obtain an integrated data view; The trend analysis sub-module analyzes the multi-tenant resource usage patterns based on the integrated data view, identifies the resource occupancy and change trends, determines the future change amount of computing power resource usage, and obtains the dynamic overview of resource usage.

[0008] On the other hand, the load analysis and prediction module includes a data collection sub-module, a behavior analysis sub-module, and a load prediction sub-module; The data collection sub-module collects the time series data of CPU cores and memory blocks based on the dynamic overview of resource usage, records the usage peaks and minimum values of resources, and statistically analyzes the change trends in continuous time periods to obtain a time series data record; The behavior analysis sub-module analyzes the task scheduling behaviors of multi-tenants based on the time series data record, statistically analyzes the average occupancy ratio of tasks to resources, and identifies the time distribution law of task execution to obtain the task scheduling behavior analysis result; The load prediction sub-module compares the resource occupancy peaks and resource idle periods of multiple tasks based on the task scheduling behavior analysis result, predicts the load trend of computing power resources in future periods, and determines the peak period interval of computing power resources to obtain the peak period prediction metrics.

[0009] On the other hand, the task priority determination module includes a demand assessment sub-module, a task classification sub-module, and a priority calculation sub-module; The demand assessment sub-module evaluates the resource consumption value and completion time range corresponding to tasks item by item based on the peak period prediction metrics according to the task scheduling demand data of multi-tenants, and combines the resource occupancy of tasks to perform data aggregation to obtain the task resource demand details; The task classification sub-module classifies tasks based on the details of the task resource requirements by comparing the task resource utilization rate and the completion time range, and calculates the total amount of resources occupied by multiple groups of tasks to obtain a summary of task classification. The priority calculation sub-module calculates the priority of each task based on the task classification summary according to the resource occupancy and completion time requirements in the task group, and determines the task execution order through the score ranking of the priorities to obtain a task priority mapping table.

[0010] On the other hand, according to the resource occupancy and completion time requirements in the task group, the formula is adopted: ; Calculate the priority of each task one by one, determine the task execution order through the score ranking of the priorities, and obtain a task priority mapping table, where represents the priority score of task ; represents the completion time requirement of task ; represents the resource occupancy of task ; represents the resource occupancy of the th task represents the total number of tasks.

[0011] On the other hand, the resource optimization configuration module includes a resource analysis sub-module, a configuration adjustment sub-module, and a resource allocation sub-module; The resource analysis sub-module analyzes the relationship between task priorities and computing power resource allocation based on the task priority mapping table, and counts the resource redundancy and deficiency situations for the resource requirements and current allocation status of high-priority tasks to obtain resource configuration analysis information; The configuration adjustment sub-module calculates the demand difference for computing power resources of high-priority tasks based on the resource configuration analysis information, makes resource ratio allocation adjustments, reduces the resource allocation quota for low-priority tasks, and allocates more computing power resources to high-priority tasks to obtain resource reconfiguration data; The resource allocation sub-module reallocates virtual computing unit resources to high-priority tasks according to task priorities based on the resource reconfiguration data, and optimizes the computing power resource configuration efficiency to obtain a real-time computing power allocation overview.

[0012] On the other hand, when calculating the demand difference for computing power resources of high-priority tasks and making resource ratio allocation adjustments to reduce the resource allocation quota for low-priority tasks, the formula is adopted: ; and ; Allocate more computing power resources to high-priority tasks to obtain resource reconfiguration data, where represents the demand difference between high-priority tasks and the current computing power resource allocation, represents the total amount of adjusted computing power resources, represents the th amount of computing power resources currently allocated to a high-priority task, represents the th expected computing power resource demand of a high-priority task, represents the total number of high-priority tasks, represents the number of high-priority tasks, represents the original total amount of computing power resources, represents the total amount of computing power resources allocated to low-priority tasks, represents the total amount of computing power resources allocated to all tasks.

[0013] On the other hand, the fault management module includes a status monitoring sub-module, a fault analysis sub-module, and a repair execution sub-module; The status monitoring sub-module, based on the real-time computing power allocation profile, monitors the abnormal status of multi-tenant resources in real time, records the time and location of the abnormality, and continuously tracks the development trend of the abnormality to obtain abnormal monitoring details; The fault analysis sub-module, based on the abnormal monitoring details, analyzes the load shift nodes and task migration situations, compares the data differences between normal load and abnormal status, and identifies the patterns and paths causing the fault to locate the fault source and obtain fault source analysis information; The repair execution sub-module, based on the fault source analysis information, implements repair operations, adjusts system-related settings or replaces faulty hardware, monitors the repair process and evaluates the repair effect, and restores the computing power resources to the normal state to obtain a log of the recovery operation effect.

[0014] On the other hand, a computing power sharing method in a multi-tenant environment is provided. This method is applied to a computing power sharing system in a multi-tenant environment and includes the following steps: S1: Based on the virtual computing unit usage data of multi-tenants, monitor the utilization rates of CPU cores, memory blocks, and storage partitions, analyze the multi-tenant resource usage patterns, and count the resource occupancy and change trends to obtain an overview of resource usage dynamics; S2: Based on the overview of resource usage dynamics, collect the time series data of CPU cores and memory blocks, predict the future resource load trends, and determine the peak period interval of computing power resources to obtain peak period prediction indicators; S3: Based on the peak period prediction indicators, according to the task scheduling requirements of multi-tenants, evaluate the resource consumption and completion time of tasks, calculate the priority of each task, and obtain a task priority mapping table; S4: Based on the task priority mapping table, analyze the task priorities and computing power resource usage data, reconfigure the existing computing power resources, and reallocate virtual computing units to high-priority tasks to obtain a real-time computing power allocation overview; S5: Based on the real-time computing power allocation overview, monitor the abnormal states of multi-tenant resources, analyze the load shift nodes and task migration situations in the abnormal states, perform fault repair operations, and restore the computing power resources to the normal state to obtain a log of the recovery operation effects.

[0015] The beneficial effects brought by the technical solution provided by the embodiments of the present invention at least include: By implementing real-time resource monitoring and data analysis, accurately capture the speed and pattern of resource consumption, making resource allocation more precise, avoiding waste of computing power resources and enhancing the response ability, generating load predictions using historical and real-time data to ensure that computing power resources can be fully utilized during peak demand periods and avoid overload, optimizing the priority ranking of tasks through intelligent analysis of resource utilization rates and estimated completion times, enabling urgent tasks to be quickly responded to. The implementation of the real-time computing power allocation strategy not only ensures high-efficiency use of computing power resources but also improves the stability and reliability of the system, reducing the cost issues caused by delayed fault responses. BRIEF DESCRIPTION OF THE DRAWINGS

[0016] In order to more clearly illustrate the technical solutions in the embodiments of the present invention, the following will briefly introduce the drawings required for the description of the embodiments. Obviously, the drawings in the following description are only some embodiments of the present invention. For those of ordinary skill in the art, without creative efforts, other drawings can be obtained based on these drawings.

[0017] Figure 1 It is a schematic diagram of the system of the present invention; Figure 2 It is a schematic diagram of the system framework of the present invention; Figure 3 It is a flowchart of the computing power monitoring module of the present invention; Figure 4 It is a flowchart of the load analysis and prediction module of the present invention; Figure 5 It is a flowchart of the task priority determination module of the present invention; Figure 6 It is a flowchart of the resource optimization and configuration module of the present invention; Figure 7 It is a flowchart of the fault management module of the present invention; Figure 8 It is a flowchart of the method steps of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0018] The technical solutions in the present invention will be described below in conjunction with the accompanying drawings.

[0019] In the embodiments of the present invention, words such as "exemplarily" and "for example" are used to indicate examples, illustrations or explanations. Any embodiment or design solution described as an "example" in the present invention should not be construed as being more preferred or having more advantages than other embodiments or design solutions. Rather, the use of the word "example" is intended to present concepts in a specific manner. In addition, in the embodiments of the present invention, the meaning expressed by "and / or" can be both, or either one of the two.

[0020] In the embodiments of the present invention, "image" and "picture" can sometimes be used interchangeably. It should be noted that when the difference is not emphasized, the meanings they express are the same. "(of)", "corresponding", and "corresponding" can sometimes be used interchangeably. It should be noted that when the difference is not emphasized, the meanings they express are the same.

[0021] In the embodiments of the present invention, sometimes subscripts such as W1 may be written in a non-subscript form such as W1. When the difference is not emphasized, the meanings they express are the same.

[0022] To make the technical problems to be solved, technical solutions and advantages of the present invention clearer, the following will be described in detail in conjunction with the accompanying drawings and specific embodiments.

[0023] The embodiments of the present invention provide a computing power sharing system in a multi-tenant environment, as Figure 1 shown, the system includes: The computing power monitoring module monitors the utilization rates of CPU cores, memory blocks, and storage partitions based on the virtual computing unit usage data of multi-tenants, integrates the data according to resource types and usage frequencies, analyzes the multi-tenant resource usage patterns, and statistically analyzes the resource occupancy and change trends to obtain an overview of the dynamic resource usage; The load analysis and prediction module collects the time series data of CPU cores and memory blocks based on the overview of the dynamic resource usage, analyzes the task scheduling behaviors of multi-tenants, predicts the resource load trends in future time periods, and determines the peak time period interval of computing power resources to obtain peak time period prediction indicators; The task priority determination module evaluates the resource consumption and completion time of tasks based on the peak time period prediction indicators, classifies tasks according to the task scheduling requirements of multi-tenants, calculates the priority of each task, and obtains a task priority mapping table; The resource optimization and configuration module analyzes the task priorities and computing power resource usage data based on the task priority mapping table, reconfigures the existing computing power resources, and reallocates virtual computing units to high-priority tasks to optimize the computing power resource configuration efficiency and obtain an overview of the real-time computing power allocation; Based on the real-time computing power allocation overview, the fault management module monitors the abnormal states of multi-tenant resources, analyzes the load shift nodes and task migration situations in the abnormal states, conducts fault diagnosis, and performs fault repair operations to restore the computing power resources to the normal state, obtaining the log of the restoration operation effect.

[0024] The dynamic overview of resource usage includes utilization statistics information, consumption speed metrics, space occupancy data. The peak period prediction metrics are specifically the CPU usage peak, memory occupancy peak, and resource demand prediction results. The task priority mapping table includes resource demand classification results, completion time limit assessment results, and priority sorting results. The real-time computing power allocation overview includes resource configuration status and virtual unit allocation data. The log of the restoration operation effect is specifically the abnormal detection status, fault diagnosis results, and restoration status update results.

[0025] As Figure 2 and Figure 3 shown, the computing power monitoring module includes a resource tracking sub-module, a data integration sub-module, and a trend analysis sub-module; Based on the virtual computing unit usage data of multi-tenants, the resource tracking sub-module monitors the real-time utilization rates of CPU cores, memory blocks, and storage partitions, and real-time tracks the resource consumption speed and space occupancy, obtaining a snapshot of resource consumption; Collect the real-time usage data of CPU cores, memory blocks, and storage partitions by multi-tenants in the virtualization environment, set monitoring tasks for each virtual computing unit, with a collection period of 1 second. The monitored data includes core metrics such as CPU usage rate, memory usage, and remaining space of the storage partition. Standardize the data through the defined collection protocol and API interface, call the real-time data stream processing framework to clean and parse the data stream collected each time, filter out the collected error or missing data information, and then store the cleaned data in the time series database for archiving, recording the resource consumption value and timestamp of each collection point. Calculate the snapshot value of real-time resource consumption by calculating the stored data.

[0026] Based on the snapshot of resource consumption, the data integration sub-module classifies and integrates the data according to resource type and usage frequency, and conducts quality inspection and cleaning on the integrated data, obtaining an integrated data view; First, classify the snapshot data by resource type (such as CPU, memory, storage), call the data partitioning algorithm to group the data with tenant ID and resource ID as key-value pairs, and sort the time series in each group in chronological order. Define filtering rules to eliminate invalid snapshot points caused by noise or abnormal collection. Subsequently, sample the classified data at a specified frequency, aggregate the high-frequency sampled data into low-frequency statistical data such as average value, maximum value, minimum value, etc. Process the redundant or duplicate records in the sampled data through a data cleaning algorithm to generate a data view that has been cleaned and integrated.

[0027] Based on the integrated data view, the trend analysis sub-module analyzes the resource usage patterns of multiple tenants, identifies resource occupancy and change trends, determines the future change amount of computing power resource usage, and obtains an overview of the dynamic resource usage. First, calculate the average usage rate and peak occupancy rate of different tenants under the same resource type through statistical analysis methods. Use the linear regression analysis algorithm to fit the resource occupancy data in a continuous time period and obtain the resource change rate. Analyze the slope and change amplitude of the resource usage curve. Divide the resource usage situation into three categories: growth type, stable type, and decline type by defining classification rules for tenant resource utilization patterns. At the same time, adopt short-term and long-term trend analysis models to evaluate the resource usage changes within the next 7 days and 30 days respectively. Finally, generate a dynamic overview diagram including resource usage patterns and future change amounts.

[0028] As Figure 2 and Figure 4 shown, the load analysis and prediction module includes a data collection sub-module, a behavior analysis sub-module, and a load prediction sub-module. Based on the overview of dynamic resource usage, the data collection sub-module collects time series data of CPU cores and memory blocks, records the peak and minimum usage of resources, and counts the change trends in continuous time periods to obtain time series data records. Set up monitoring tasks to collect the usage of CPU cores and memory blocks. By defining the collection period and time window, generate time series data points for the resource usage status per second, record the CPU core occupancy rate and the usage amount of memory blocks at each time point. Store the peak, minimum, and resource utilization change curves of resource occupancy in each time period as key indicators. Filter and denoise the original data to eliminate outliers or missing values during the collection process. Segment the time series data through a defined window sliding mechanism, count the average resource occupancy rate and change trends in continuous time periods, and store the statistical results in the resource usage database to form time series data records with timestamps.

[0029] The behavior analysis sub-module analyzes the task scheduling behaviors of multi-tenants based on time series data records, calculates the average resource occupancy ratio of tasks, identifies the time distribution pattern of task execution, and obtains the analysis results of task scheduling behaviors. Group and analyze the task scheduling behaviors of each tenant, extract the start time and end time of task execution, match the resource occupancy data with task timestamps, calculate the average occupancy ratio of each task for CPU cores and memory blocks, set task classification rules to divide tasks into high, medium, and low occupancy ratio categories, determine the time distribution of different tasks throughout the day by calculating the frequency distribution of task occupancy periods, count the execution density and time period distribution characteristics of tasks, and generate the analysis results of task scheduling behaviors.

[0030] Based on the analysis results of task scheduling behaviors, the load prediction sub-module compares the resource occupancy peaks of multiple tasks and resource idle periods, predicts the future trend of computing resource load, and determines the peak period interval of computing resources to obtain peak period prediction indicators. Extract the resource occupancy peaks of tasks and the resource idle periods between tasks, split the time series data, model the peak period and off-peak period separately, calculate the changing trend of future resource occupancy through linear interpolation, set the calculation rules for the resource peak period interval, predict the resource demand for the next period based on the resource usage change rate in the previous period, define the time period with a resource usage rate exceeding the preset threshold as the peak period, and finally obtain the peak period prediction indicators for computing resources.

[0031] As Figure 2 and Figure 5 shown, the task priority determination module includes a demand assessment sub-module, a task classification sub-module, and a priority calculation sub-module. Based on the peak period prediction indicators, the demand assessment sub-module evaluates the corresponding resource consumption values and completion time ranges of tasks item by item according to the task scheduling demand data of multi-tenants, combines the resource occupancy of tasks, and performs data aggregation to obtain the detailed task resource requirements. First, extract the basic parameters of tasks from the multi-tenant task scheduling demand data, including the start time, end time, required number of computing cores, size of memory blocks, and execution frequency of tasks. Group the task demand data through the time axis, match the task execution period with the peak period prediction indicators, count the resource usage amount and occupied time length required for each task during execution, calculate the resource occupancy quantity of tasks during the peak period and the resource requirements during the off-peak period item by item according to the set calculation rules, combine the time span of tasks and the scheduling cycle, obtain the completion time range corresponding to each task, and finally generate the detailed task resource requirements through organizing the calculation results.

[0032] The task classification sub-module classifies tasks based on the detailed task resource requirements, by comparing the task resource utilization rate and the completion time range, and counts the total amount of resources occupied by multiple groups of tasks to obtain the task classification summary; First, sort the tasks according to the resource usage of the tasks, and classify and store high-occupancy tasks, medium-occupancy tasks, and low-occupancy tasks respectively according to the classification threshold set by the resource occupancy ratio. Subsequently, count the completion time range of each task, and divide the tasks into two categories: long-time tasks and short-time tasks according to the grouping rules set by the time length. Through the cross-matching of the resource usage and the time range, the tasks are further refined into two types: resource-intensive tasks and time-sensitive tasks. Count the total distribution of different classified tasks in terms of resource types to generate the task classification summary.

[0033] The priority calculation sub-module calculates the priority of each task one by one based on the task classification summary, according to the resource occupancy and the completion time requirements in the task group, and determines the task execution order through the score sorting of the priorities to obtain the task priority mapping table; Extract the resource usage and the completion time requirements of each task in the task group, allocate weight parameters according to the priority calculation rules, take the resource usage and the time requirements as the core indicators of the priority, and assign different weight values respectively. By evaluating the resource occupancy ratio and the urgency of the completion time of the tasks, calculate the priority score of each task one by one, sort the tasks in the task group according to the scores, arrange the priorities from high to low, organize the sorting results to generate the priority list of the tasks, and finally output the corresponding relationship between the tasks and the priorities to form the task priority mapping table.

[0034] According to the resource occupancy and the completion time requirements in the task group, use the formula: ; Calculate the priority of each task one by one, determine the task execution order through the score sorting of the priorities to obtain the task priority mapping table, where represents the priority score of task , represents the completion time requirement of task , represents the resource occupancy of task , represents the resource occupancy of the th task, represents the total number of tasks; The following task data is obtained through data collection and monitoring: The total number of tasks is ; The completion time requirements of each task are respectively: ; The resource usage of each task is: ; The total amount of task resources is calculated by the cumulative formula: ; Calculate the priority score of Task 1: ; Calculate the priority score of Task 2: ; Calculate the priority score of Task 3: ; Calculate the priority score of Task 4: ; According to the above calculation results, the priority scores of each task are: ; ; ; ; By sorting by priority scores, the execution order of tasks is: Task 4>Task 2>Task 1>Task 3. This result shows that tasks with higher priority scores have shorter completion time requirements and larger resource usage, while tasks with lower priority scores have longer completion time requirements or smaller resource usage, thereby optimizing resource allocation efficiency and time management.

[0035] like Figure 2 and Figure 6 As shown, the resource optimization configuration module includes a resource analysis submodule, a configuration adjustment submodule, and a resource allocation submodule; The resource analysis submodule analyzes the relationship between task priority and computing resource allocation based on the task priority mapping table. It counts resource redundancy and shortage for high-priority tasks and current allocation status to obtain resource configuration analysis information. Extract the priority score, resource allocation status and resource amount required for each task. According to the resource allocation requirements of high-priority tasks, count the current allocation value and actual demand value of each type of resource in high-priority tasks in turn. Mark the part where the demand is greater than the allocation as the resource-deficient area, and mark the part where the allocation is greater than the demand as the resource-redundant area. Summarize all resource-redundant and insufficient areas, re-sort the resource status according to the task priority, and add up the number of resources in each area. Finally, form resource configuration analysis information and store it in a structured data format.

[0036] Based on the resource configuration analysis information, the configuration adjustment sub-module calculates the demand difference of high-priority tasks for computing power resources, adjusts the resource ratio allocation, reduces the resource allocation quota of low-priority tasks, and allocates more computing power resources to high-priority tasks to obtain resource reconfiguration data; According to the resource requirements of high-priority tasks, calculate the difference between the resources required by the tasks and the currently allocated resources, extract the resource allocation quota of low-priority tasks and calculate its total amount, reduce the resource allocation quota of low-priority tasks according to a preset ratio, divide the released resources into various resource pools, and gradually adjust the resource allocation by matching the demand difference of high-priority tasks with the released resource amount. Finally, reallocate each type of resource from low-priority tasks to high-priority tasks, generate the reconfiguration data of high-priority tasks and save it.

[0037] Calculate the demand difference of high-priority tasks for computing power resources, adjust the resource ratio allocation, reduce the resource allocation quota of low-priority tasks, using the formula: ; and ; Allocate more computing power resources to high-priority tasks to obtain resource reconfiguration data, where represents the demand difference between high-priority tasks and the current computing power resource configuration, represents the total amount of adjusted computing power resources, represents the th amount of computing power resources currently allocated to the high-priority task, represents the th expected computing power resource requirement of the high-priority task, represents the total number of high-priority tasks, represents the number of high-priority tasks, represents the original total amount of computing power resources, represents the total amount of computing power resources allocated to low-priority tasks, the total amount of computing power resources allocated to all tasks; The actually monitored data is as follows: (Current resource allocation of high-priority tasks) is [10, 20, 15] GFLOPS; (Ideal resource requirements of high-priority tasks) is [12, 18, 17] GFLOPS; (Total number of high-priority tasks) is 3 tasks; (Original total resource amount) is 100 GFLOPS; The total amount of low-priority task resource allocation is 30 GFLOPS; The total amount of resource allocation is 100 GFLOPS; Calculation : ; ; Calculation : ; The calculation results show that the adjusted total amount of resources is 99.86 GFLOPS, showing the resource allocation after adjusting 0.14 GFLOPS computing power resources from low-priority tasks to high-priority tasks. The purpose of this resource reallocation is to ensure that high-priority tasks obtain an allocation closer to their ideal resource requirements, thereby improving the efficiency and response ability of the entire system.

[0038] Based on the resource reconfiguration data, the resource allocation sub-module reallocates virtual computing unit resources to high-priority tasks according to task priorities, and optimizes the efficiency of computing power resource allocation to obtain a real-time computing power allocation profile; Execute the reallocation of virtual computing unit resources one by one according to the allocation requirements of high-priority tasks, update the resource status of each virtual computing unit in real time, record the allocated resource status data into the real-time resource database in chronological order, and at the same time conduct an overall assessment of the computing power resource allocation status of all current task groups, adjust and optimize the resource allocation order according to task priorities, and finally generate a real-time computing power allocation profile for high-priority tasks.

[0039] Such as Figure 2 and Figure 7 shown, the fault management module includes a status monitoring sub-module, a fault analysis sub-module, and a repair execution sub-module; Based on the real-time computing power allocation profile, the status monitoring sub-module monitors the abnormal status of multi-tenant resources in real time, records the time and location of the abnormality, and continuously tracks the development trend of the abnormality to obtain the details of abnormal monitoring; Extract the current allocation status of multi-tenant resources and the usage of computing power resources, set the monitoring period as a second-level unit, collect and record the resource utilization rate of each virtual computing unit in real time, define the determination conditions for abnormal status, including resource occupancy exceeding the threshold, uneven resource allocation, or resource unavailability, etc., compare and analyze the monitoring data for each time period, mark the status determined to be abnormal according to time and resource location, store the duration of the abnormality and the node information where it occurs, conduct a trend analysis on the abnormal changes in consecutive time periods, record the change direction and fluctuation range of the abnormal status, and finally form the details of abnormal monitoring.

[0040] Based on the abnormal monitoring details, the fault analysis sub-module analyzes the load offset nodes and task migration situations, compares the data differences between normal load and abnormal states, identifies the patterns and paths causing the faults, locates the fault sources, and obtains fault source analysis information; Compare the resource usage in the normal load state and the abnormal state. Calculate the differences between the resource usage and allocation amounts involved in the abnormal state and the average values in the normal state respectively. Identify the change trends and diffusion paths of abnormal loads. Track each node involved in task migration one by one, extract the start time, end time of the migrated tasks, and the resource allocation situation during migration. Compare the resource state changes before and after migration. Define the identification rules for abnormal nodes. Determine the location of the abnormal source according to the resource usage deviation and load trend matching rules. Finally, locate the patterns and paths causing the faults to form fault source analysis information.

[0041] Based on the fault source analysis information, the repair execution sub-module implements repair operations, adjusts the system-related settings or replaces faulty hardware, monitors the repair process and evaluates the repair effect, restores the computing power resources to the normal state, and obtains the recovery operation effect log; Adjust the system configuration parameters or trigger resource reallocation operations. If the fault involves hardware problems, confirm the hardware status and replace the faulty components. Monitor the resource state changes during the repair process, record the resource usage after repair, compare the resource utilization rate and allocation status changes before and after repair, evaluate whether the repair meets the normal standards, and store the operation records involved in the repair process as log data.

[0042] As Figure 8 shown, a computing power sharing method in a multi-tenant environment includes the following steps: S1: Based on the virtual computing unit usage data of multi-tenants, monitor the utilization rates of CPU cores, memory blocks, and storage partitions, analyze the multi-tenant resource usage patterns, and count the resource occupancy and change trends to obtain an overview of resource usage dynamics; S2: Based on the overview of resource usage dynamics, collect the time series data of CPU cores and memory blocks, predict the resource load trends in future periods, and determine the peak period interval of computing power resources to obtain peak period prediction indicators; S3: Based on the peak period prediction indicators, evaluate the resource consumption and completion time of tasks according to the task scheduling requirements of multi-tenants, calculate the priority of each task, and obtain a task priority mapping table; S4: Based on the task priority mapping table, analyze the task priorities and computing power resource usage data, reconfigure the existing computing power resources, and reallocate virtual computing units to high-priority tasks to obtain an overview of real-time computing power allocation; S5: Based on the real-time computing power allocation profile, monitor the abnormal status of multi-tenant resources, analyze the load offset nodes and task migration in the abnormal status, perform fault repair operations, restore the computing power resources to the normal state, and obtain the log of the restoration operation effect.

[0043] It should be understood that the term "and / or" in this article is merely a description of the association relationship of associated objects, indicating that there can be three relationships. For example, A and / or B can represent: A exists alone, A and B exist simultaneously, and B exists alone. Here, A and B can be singular or plural. In addition, the character " / " in this article generally represents an "or" relationship between the front and back associated objects, but it may also represent an "and / or" relationship, which can be specifically understood by referring to the context before and after.

[0044] In the present invention, "at least one" means one or more, and "multiple" means two or more. "At least one of the following" or its similar expressions refer to any combination of these items, including any combination of single item (s) or plural item (s). For example, at least one of a, b, or c can represent: a, b, c, a - b, a - c, b - c, or a - b - c, where a, b, and c can be single or multiple.

[0045] It should be understood that in various embodiments of the present invention, the magnitudes of the sequence numbers of the above processes do not mean the order of execution. The order of execution of each process should be determined by its function and internal logic, and should not constitute any limitation to the implementation process of the embodiments of the present invention.

[0046] Those of ordinary skill in the art can realize that the units and algorithm steps of each example described in combination with the embodiments disclosed in this article can be implemented by electronic hardware, or a combination of computer software and electronic hardware. Whether these functions are executed in a hardware or software manner depends on the specific application and design constraints of the technical solution. Professional technicians can use different systems for each specific application to implement the described functions, but such implementation should not be considered to exceed the scope of the present invention.

[0047] Those skilled in the art can clearly understand that for the convenience and simplicity of description, the specific working processes of the devices, apparatuses, and units described above can refer to the corresponding processes in the foregoing system embodiments, and will not be elaborated herein.

[0048] In several embodiments provided by the present invention, it should be understood that the disclosed devices, apparatuses, and systems can be implemented in other ways. For example, the device embodiments described above are merely illustrative. For example, the division of the units is only a logical function division. In actual implementation, there may be other division methods. For example, multiple units or components can be combined or integrated into another device, or some features can be ignored or not executed. Another point is that the couplings or direct couplings or communication connections shown or discussed with each other can be through some interfaces. The indirect couplings or communication connections of the devices or units can be in electrical, mechanical, or other forms.

[0049] The units described as separate components may or may not be physically separated. The components shown as units may or may not be physical units, that is, they can be located in one place or distributed to multiple network units. Some or all of the units can be selected according to actual needs to achieve the purpose of the solution of this embodiment.

[0050] In addition, in each embodiment of the present invention, the functional units can be integrated into one processing unit, or each unit can exist physically alone, or two or more units can be integrated into one unit.

[0051] If the function is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present invention, in essence, or the part that contributes to the prior art, or a part of this technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions for causing a computer device (which can be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the systems described in each embodiment of the present invention. The aforementioned storage medium includes: various media such as USB flash drives, mobile hard disks, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical discs that can store program codes.

[0052] As described above, the above are only specific implementation manners of the present invention, but the protection scope of the present invention is not limited thereto. Any person skilled in the art within the technical scope disclosed by the present invention can easily think of changes or substitutions, which should all be covered by the protection scope of the present invention. Therefore, the protection scope of the present invention should be subject to the protection scope of the claims.

Claims

1. A computing power sharing system in a multi-tenant environment, characterized in that The system includes: The computing power monitoring module monitors the utilization rates of CPU cores, memory blocks, and storage partitions based on the usage data of multi-tenant virtual computing units, integrates the data according to resource types and usage frequencies, analyzes the resource usage patterns of multi-tenants, and statistically analyzes the resource occupancy and change trends to obtain a dynamic overview of resource usage; The load analysis and prediction module collects time series data of CPU cores and memory blocks based on the dynamic overview of resource usage, analyzes the task scheduling behaviors of multi-tenants, predicts the resource load trends in future time periods, and determines the peak time period intervals of computing power resources to obtain peak time period prediction indicators; The task priority determination module evaluates the resource consumption and completion time of tasks according to the task scheduling requirements of multi-tenants based on the peak time period prediction indicators, classifies tasks, calculates the priority of each task, and obtains a task priority mapping table; The resource optimization and configuration module analyzes the task priorities and computing power resource usage data based on the task priority mapping table, reconfigures the existing computing power resources, and reallocates virtual computing units to high-priority tasks to optimize the computing power resource configuration efficiency and obtain a real-time overview of computing power allocation; The fault management module monitors the abnormal states of multi-tenant resources based on the real-time overview of computing power allocation, analyzes the load shift nodes and task migration situations in the abnormal states, performs fault diagnosis, and executes fault repair operations to restore the computing power resources to the normal state and obtain a log of the recovery operation effects; 2. The computing power sharing system in a multi-tenant environment according to claim 1, wherein The dynamic overview of resource usage includes utilization rate statistical information, consumption speed indicators, and space occupancy data. The peak time period prediction indicators are specifically the CPU usage peak, memory occupancy peak, and resource demand prediction results. The task priority mapping table includes resource demand classification results, completion time limit evaluation results, and priority ranking results. The real-time overview of computing power allocation includes resource configuration status and virtual unit allocation data. The log of the recovery operation effects is specifically the abnormal detection status, fault diagnosis results, and recovery status update results.

3. The computing power sharing system in a multi-tenant environment according to claim 1, characterized in that, The computing power monitoring module includes: The resource tracking sub-module monitors the real-time utilization rates of CPU cores, memory blocks, and storage partitions based on the usage data of multi-tenant virtual computing units, and real-time tracks the resource consumption speed and space occupancy to obtain a resource consumption snapshot; The data integration sub-module classifies and integrates the data according to resource types and usage frequencies based on the resource consumption snapshot, and performs quality inspection and cleaning on the integrated data to obtain an integrated data view; The trend analysis sub-module analyzes the resource usage patterns of multi-tenants based on the integrated data view, identifies the resource occupancy and change trends, and determines the future change amount of computing power resource usage to obtain a dynamic overview of resource usage.

4. The computing power sharing system in a multi-tenant environment according to claim 1, wherein The load analysis and prediction module includes: The data collection sub-module collects time series data of CPU cores and memory blocks based on the dynamic overview of resource usage, records the usage peaks and minimum values of resources, and statistically analyzes the change trends in continuous time periods to obtain a record of time series data; The behavior analysis sub-module analyzes the task scheduling behavior of multi-tenants based on the time series data record, calculates the average resource occupancy ratio of tasks, identifies the time distribution law of task execution, and obtains the task scheduling behavior analysis result; The load prediction sub-module compares the resource occupancy peaks and resource idle periods of multiple tasks based on the task scheduling behavior analysis result, predicts the computing power resource load trend in the future period, and determines the high computing power resource peak period interval to obtain the peak period prediction index.

5. The computing power sharing system in a multi-tenant environment according to claim 1, wherein The task priority determination module includes: The demand assessment sub-module evaluates the resource consumption value and completion time range corresponding to tasks item by item based on the peak period prediction index and the task scheduling demand data of multi-tenants, and combines the resource occupancy of tasks to perform data summarization to obtain the task resource demand details; The task classification sub-module classifies tasks by comparing the task resource usage rate and completion time range based on the task resource demand details, and calculates the total amount of resources occupied by multiple groups of tasks to obtain the task classification summary; The priority calculation sub-module calculates the priority of tasks one by one according to the resource occupancy and completion time requirements in the task group based on the task classification summary, and determines the task execution order through the score sorting of priorities to obtain the task priority mapping table.

6. The computing power sharing system in a multi-tenant environment according to claim 5, wherein According to the resource occupancy and completion time requirements in the task group, the formula is used: ; Calculate the priority of each task one by one, sort by the score of the priority, determine the task execution order, and obtain the task priority mapping table, where represents the priority score of the task . represents the completion time requirement of the task . represents the resource occupancy of the task . represents the th resource occupancy of the task represents the total number of tasks.

7. The computing power sharing system in a multi-tenant environment according to claim 1, wherein The resource optimal configuration module includes: The resource analysis sub-module analyzes the relationship between task priorities and computing power resource allocation based on the task priority mapping table, and counts the resource redundancy and deficiency for the resource requirements and current allocation status of high-priority tasks to obtain the resource configuration analysis information; The configuration adjustment sub-module calculates the demand difference of high-priority tasks for computing power resources based on the resource configuration analysis information, performs resource ratio allocation adjustment, reduces the resource allocation quota of low-priority tasks, and allocates more computing power resources to high-priority tasks to obtain the resource reconfiguration data; The resource allocation sub-module reallocates virtual computing unit resources to high-priority tasks according to task priorities based on the resource reconfiguration data, and optimizes the computing power resource configuration efficiency to obtain the real-time computing power allocation overview.

8. The computing power sharing system in a multi-tenant environment according to claim 7, wherein Calculating the demand difference of high-priority tasks for computing power resources, performing resource ratio allocation adjustment, and reducing the resource allocation quota of low-priority tasks, the formula is used: ; and ; Allocate more computing power resources to the highest-priority tasks to obtain resource reconfiguration data, where represents the demand difference between the high-priority tasks and the current computing power resource allocation, represents the total adjusted computing power resources, represents the number of computing power resources currently allocated to the th high-priority task, represents the expected computing power resource demand for the th high-priority task, represents the total number of high-priority tasks, represents the original total computing power resources, represents the total computing power resource allocation for low-priority tasks, represents the total computing power resource allocation for all tasks.

9. The computing power sharing system in a multi-tenant environment according to claim 1, wherein The fault management module includes: The status monitoring sub-module monitors the abnormal status of multi-tenant resources in real time based on the real-time computing power allocation overview, records the time and location of the abnormality, and continuously tracks the abnormal development trend to obtain the abnormal monitoring details; The fault analysis sub-module analyzes the load offset nodes and task migration situations based on the abnormal monitoring details, compares the data differences between normal load and abnormal status, and identifies the patterns and paths causing the faults to locate the fault source and obtain the fault source analysis information; The repair execution sub-module implements repair operations based on the fault source analysis information, adjusts the system-related settings or replaces faulty hardware, monitors the repair process and evaluates the repair effect, and restores the computing power resources to the normal state to obtain the recovery operation effect log.

10. A computing power sharing method in a multi-tenant environment, the computing power sharing method in the multi-tenant environment is used to implement the computing power sharing system in the multi-tenant environment according to any one of claims 1-9, characterized in that, including the following steps: S1: Based on the data of virtual computing units used by multi-tenants, monitor the utilization rates of CPU cores, memory blocks, and storage partitions, analyze the resource usage patterns of multi-tenants, and count the resource occupancy and change trends to obtain an overview of the dynamic resource usage. S2: Based on the overview of the dynamic resource usage, collect the time series data of CPU cores and memory blocks, predict the resource load trend in the future period, and determine the interval of peak computing power resource periods to obtain peak period prediction indicators. S3: Based on the peak period prediction indicators, evaluate the resource consumption and completion time of tasks according to the task scheduling requirements of multi-tenants, calculate the priority of each task, and obtain a task priority mapping table. S4: Based on the task priority mapping table, analyze the task priorities and computing power resource usage data, reconfigure the existing computing power resources, and reallocate virtual computing units to high-priority tasks to obtain an overview of real-time computing power allocation. S5: Based on the overview of the real-time computing power allocation, monitor the abnormal status of multi-tenant resources, analyze the load shift nodes and task migration situations in the abnormal status, perform fault repair operations, restore the computing power resources to the normal state, and obtain a log of the recovery operation effects.

Citation Information

Cited By

  • Computing power dynamic segmentation method and system, computer equipment and storage medium

    CN120469816A

  • Data platform resource consumption monitoring system

    CN120469904A

  • Distributed hierarchical scheduling method and device based on complex scene

    CN120540861A

  • Data hotspot identification and computing power dynamic reservation scheduling method

    CN120750797A

  • Data hotspot identification and computing power dynamic reservation scheduling method

    CN120750797B