Cloud computing big data all-in-one machine

By designing a cloud computing big data all-in-one machine with multiple collaborative working modules, the limitations of traditional systems in resource management and task scheduling are solved, efficient resource utilization and task scheduling are achieved, and the overall performance and reliability of the system are improved.

CN119938308APending Publication Date: 2025-05-06CHINA SOUTHERN POWER GRID COMPANY
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
CN202411787775.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2024-12-05
Publication Date
2025-05-06

AI Technical Summary

Technical Problem

Traditional cloud computing big data systems have limitations in resource management and task scheduling, and are difficult to meet complex and changeable data processing needs, including inefficient resource utilization, inefficient task execution efficiency, extended response time, and potential risks in security and reliability.

Method used

A cloud computing big data all-in-one machine is designed, including task resource extraction module, task load detection module, resource utilization analysis module, adaptive adjustment module, abnormal detection and early warning module, automatic load balancing module, task priority setting module and human-computer interaction module. Through the coordinated work of these modules, dynamic resource management and intelligent task scheduling are realized.

Benefits of technology

By accurately evaluating task requirements and resource utilization, efficient utilization of resources and efficient task scheduling can be achieved, task execution efficiency is significantly improved, computing costs are reduced, and system security and reliability are enhanced.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119938308A_ABST
    Figure CN119938308A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of cloud computing and big data, and particularly discloses a cloud computing and big data all-in-one machine which comprises a task resource extraction module, a task load detection module, a resource utilization rate analysis module, a self-adaptive adjustment module, an anomaly detection and early warning module, an automatic load balancing module, a task priority setting module and a man-machine interaction module. The task load change coefficient and the resource utilization rate change coefficient of each calculation node are calculated, the change trend of the task load is analyzed according to the task load change coefficient, the task load report is generated, the change trend of the resource utilization rate is analyzed according to the resource utilization rate change coefficient, and the resource utilization rate report is generated. Through judgment and analysis of the abnormity detection early warning module and the preset threshold value, management personnel can be reminded in time to adjust abnormal data, reasonable utilization of resources is ensured through introduction of dynamic adjustment and load balancing technologies, and resource waste is avoided.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of cloud computing and big data, and in particular to a cloud computing and big data integrated machine. Background Art

[0002] The rise of cloud computing and big data systems is an inevitable product of the development of information technology, aimed at meeting the growing demand for data processing. With the popularization of the Internet and the development of Internet of Things technology, enterprises and organizations need to process and analyze massive amounts of data every day. These data are not only large in scale, but also diverse in type and complex in structure, which places extremely high demands on computing resources and storage resources. In order to meet this challenge, cloud computing and big data technologies are combined to build a flexible and scalable data processing platform to provide users with efficient and convenient data services.

[0003] Traditional cloud computing big data systems are usually based on virtualization technology, which manages computing resources, storage resources and network resources in a pooled manner. Users can dynamically apply for and use these resources according to actual needs. These systems usually include multiple modules such as data acquisition, data storage, data processing and data analysis, and can support parallel processing and real-time analysis of large-scale data. However, traditional cloud computing big data systems still have some limitations in resource management and task scheduling, and it is difficult to fully meet the complex and changing data processing needs.

[0004] Traditional cloud computing big data systems have obvious deficiencies in resource allocation and task scheduling. On the one hand, these systems often adopt static resource allocation strategies and cannot be dynamically adjusted according to the complexity of the task and the real-time load, resulting in low resource utilization and serious waste. On the other hand, traditional task scheduling algorithms often lack intelligence and adaptability and cannot accurately assess the resource requirements and priorities of tasks, resulting in low task execution efficiency and extended response time. In addition, traditional cloud computing big data systems cannot achieve real-time reporting of resource utilization, automatic load balancing, adaptive adjustment, anomaly detection, and task priority setting. There are also some potential risks and challenges in terms of security and reliability, which need to be further improved and optimized. Summary of the invention

[0005] In order to overcome the above-mentioned defects of the prior art, an embodiment of the present invention provides a cloud computing big data integrated machine to solve the problems raised in the above-mentioned background technology.

[0006] To achieve the above-mentioned objectives, the present invention provides the following technical solutions: a cloud computing big data integrated machine, including a task resource extraction module, a task load detection module, a resource utilization analysis module, an adaptive adjustment module, an anomaly detection and warning module, an automatic load balancing module, a task priority setting module and a human-computer interaction module.

[0007] Task resource extraction module: used to receive and parse task requests, extract task requirement information, evaluate the complexity of task resources based on the task requirement information, and generate a task resource requirement report;

[0008] Task load detection module: used to calculate the task load variation coefficient of each computing node, analyze the change trend of the task load according to the task load variation coefficient, generate a task load report, and transmit the task load variation coefficient to the anomaly detection and warning module;

[0009] Resource utilization analysis module: used to calculate the resource utilization change coefficient of each computing node, analyze the change trend of resource utilization according to the resource utilization change coefficient, generate a resource utilization report, and transmit the resource utilization change coefficient to the anomaly detection and early warning module;

[0010] Adaptive adjustment module: used to adaptively adjust the allocation of computing resources, storage resources, and network resources using elastic scaling adjustment strategies according to changes in task resource requirements, task load, and resource utilization;

[0011] Anomaly detection and early warning module: used to compare the task load change coefficient and resource utilization change coefficient with the preset thresholds to determine whether anomalies occur. Based on the abnormal data, a machine learning algorithm is used to perform anomaly detection and generate early warning information to remind the administrator to handle it in time.

[0012] Automatic load balancing module: automatically adjusts the distribution of tasks on each computing node according to the task load change coefficient and resource utilization change coefficient in the anomaly detection and early warning module and the preset threshold judgment result, and uses genetic algorithm to optimize the task scheduling strategy;

[0013] Task priority setting module: used by managers to set task priorities and prioritize task resources according to the urgency and importance of tasks;

[0014] Human-computer interaction module: used to receive the warning information and task priority division information from the abnormal detection and warning module, so that managers can manage and make decisions on abnormal data.

[0015] Preferably, the task requirement information includes computing resources, storage resources and network resource requirement information required for the task, the computing resources include task load value, CPU utilization and processing power hardware facilities, such as central processing unit CPU, graphics processing unit GPU, field programmable gate array FPGA and accelerator; the storage resources cover equipment and technologies for data storage, including hard disk drives HDD, solid state drives SSD, memory RAM and distributed storage systems; the network resource requirement information includes bandwidth, delay, throughput, network topology and quality of service QoS parameters.

[0016] Preferably, the calculation steps of the task load variation coefficient of each computing node are specifically as follows:

[0017] Step S1: Calculate the average load AL. The calculation model is as follows:

[0018] Where, L = { L1,L2,…,L m } It is represented as a set of task load values ​​collected within two hours, m is represented as the number of sampling points; j is represented as a sampling point of the task load value, j = 1, 2, 3, ... m;

[0019] Step S2: Calculate the task load variation coefficient of each computing node. The calculation model is as follows:

[0020] Among them, LVC represents the task load variation coefficient of each computing node, L j It is represented as the load value of the jth sampling point, and AL is represented as the average load.

[0021] Preferably, the steps for calculating the resource utilization change coefficient of each computing node are specifically as follows:

[0022] Step S1: Calculate the average utilization AU. The calculation model is as follows:

[0023] Among them, U= { U1,U2,…,U n } It represents the CPU utilization set collected within two hours, n represents the number of sampling points; i represents the CPU utilization sampling point, i = 1, 2, 3, ... n;

[0024] Step S2: Calculate the standard deviation of utilization σ u, the calculation model is as follows:

[0025] Among them, U iIt is expressed as the CPU utilization of the i-th sampling point;

[0026] Step S3: Calculate the resource utilization change coefficient of each computing node. The calculation model is as follows:

[0027] Among them, RVC represents the resource utilization change coefficient of each computing node, AU represents the average utilization, σ u is denoted as the standard deviation of utilization.

[0028] Preferably, the specific content of the abnormal detection and early warning module comparing the task load change coefficient with a preset threshold value is:

[0029] Extract the task load variation coefficient LVC of each computing node and obtain the task load variation deviation formula: If the task load change coefficient is less than or equal to the preset load change threshold LVC 预 , then the judgment result is that the task load is normal. If the task load change coefficient is greater than the preset load change threshold LVC 预 , the result is judged as abnormal task load, the abnormal result is sent to the management personnel, and the automatic load balancing module is started.

[0030] Preferably, the specific content of the abnormal detection and early warning module comparing the resource utilization rate change coefficient with a preset threshold value is:

[0031] Extract the resource utilization change coefficient RVC of each computing node and obtain the resource utilization change deviation formula: If the resource utilization change coefficient is less than or equal to the preset resource utilization change threshold RVC 预 , then the judgment result is that the resource utilization is normal. If the resource utilization change coefficient is greater than the preset resource utilization change threshold RVC 预 , the result is that the resource utilization is abnormal, the abnormal result is sent to the management personnel, and the automatic load balancing module is started.

[0032] Technical effects and advantages of the present invention:

[0033] 1. The present invention realizes accurate assessment and prediction of task requirements and resource utilization through task resource complexity demand analysis and task load real-time detection technology, and realizes efficient utilization of resources and efficient scheduling of tasks through dynamic reporting and adaptive adjustment technology;

[0034] 2. The present invention can report resource utilization in real time, automatically perform load balancing, adaptive adjustment, anomaly detection, and task priority setting, and realize dynamic matching of task requirements and resource utilization. Through the application of this technology, the task execution efficiency can be significantly improved, the computing cost can be reduced, and the further development of cloud computing and big data technology can be promoted;

[0035] 3. The present invention calculates the task load variation coefficient and the resource utilization variation coefficient, and analyzes the variation trend of the task load according to the task load variation coefficient to generate a task load report. It analyzes the variation trend of the resource utilization according to the resource utilization variation coefficient to generate a resource utilization report. Through the abnormal detection and early warning module and the preset threshold judgment analysis, the management personnel can be reminded to adjust the abnormal data in time. By introducing dynamic adjustment and load balancing technology, the rational utilization of resources is ensured and the waste of resources is avoided. BRIEF DESCRIPTION OF THE DRAWINGS

[0036] The present invention is further described using the accompanying drawings, but the embodiments in the accompanying drawings do not constitute any limitation to the present invention. A person skilled in the art can obtain other drawings based on the following drawings without creative work.

[0037] Figure 1 It is a schematic diagram of the overall structure of the present invention. DETAILED DESCRIPTION

[0038] The following will be combined with the drawings in the embodiments of the present invention to clearly and completely describe the technical solutions in the embodiments of the present invention. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of the present invention.

[0039] See also Figure 1 As shown, the present invention provides a cloud computing big data integrated machine, including a task resource extraction module, a task load detection module, a resource utilization analysis module, an adaptive adjustment module, an anomaly detection and warning module, an automatic load balancing module, a task priority setting module and a human-computer interaction module.

[0040] The output end of the task resource extraction module is respectively connected to the input end of the task load detection module and the input end of the resource utilization analysis module, the output end of the task load detection module is linked to the input end of the adaptive adjustment module, the output end of the resource utilization analysis module is linked to the input end of the adaptive adjustment module, the output end of the adaptive adjustment module is linked to the input end of the anomaly detection and early warning module, the output end of the anomaly detection and early warning module is respectively linked to the input end of the automatic load balancing module and the input end of the task priority setting module, and the output end of the task priority setting module is linked to the input end of the human-computer interaction module.

[0041] The task resource extraction module is used to receive and parse the task request, extract the task requirement information, evaluate the complexity of the task resources according to the task requirement information, and generate a task resource requirement report;

[0042] In this embodiment, it should be specifically explained that the task requirement information includes the computing resources, storage resources and network resource requirement information required for the task, and the computing resources include the task load value, CPU utilization and processing power hardware facilities, such as central processing unit (CPU), graphics processing unit (GPU), field programmable gate array (FPGA) and other special accelerators. These hardware are responsible for executing computing tasks and are the core of data processing and application operation. CPU is mainly used for general computing tasks, while GPU and FPGA are good at parallel processing and acceleration of specific types of workloads, such as machine learning, graphics rendering, etc.; in addition, computing resources also include related cooling systems, power supply and other supporting facilities to ensure that the computing equipment can run continuously and stably.

[0043] The storage resources cover various devices and technologies used for data storage, including hard disk drives (HDDs), solid-state drives (SSDs), memory (RAM), and distributed storage systems. HDDs and SSDs are persistent storage solutions, with HDDs being cheaper but slower, and SSDs being faster but more expensive. Memory is volatile storage, used to temporarily store data being processed to increase data access speed. Distributed storage systems connect multiple storage devices through a network to form a logically large-capacity storage pool to improve data availability and fault tolerance. In addition, storage resources also include related backup and recovery mechanisms to ensure data security.

[0044] The network resource demand information includes bandwidth, delay, throughput, network topology and quality of service (QoS) parameters. Bandwidth refers to the maximum data transmission rate of the network connection, which determines the upper limit of the data transmission speed. Delay refers to the time required for a data packet to travel from the sender to the receiver, which is particularly important for real-time applications. Throughput reflects the total amount of data that can be transmitted by the network in a certain period of time. The network topology describes the connection method between each node in the network, which affects the efficiency and reliability of data transmission. The quality of service parameters include indicators such as packet loss rate and jitter, which are used to ensure the stability and reliability of network transmission, especially to ensure the data transmission quality of key applications under congestion conditions.

[0045] The task load detection module is used to calculate the task load variation coefficient of each computing node, analyze the change trend of the task load according to the task load variation coefficient, generate a task load report, and transmit the task load variation coefficient to the anomaly detection and early warning module;

[0046] In this embodiment, it should be specifically explained that the calculation steps of the task load variation coefficient of each computing node are as follows:

[0047] Step S1: Calculate the average load AL. The calculation model is as follows:

[0048] Where, L = { L1,L2,...,L m } It is represented as a set of task load values ​​collected within two hours, m is represented as the number of sampling points; j is represented as a sampling point of the task load value, j = 1, 2, 3, ... m;

[0049] Step S2: Calculate the task load variation coefficient of each computing node. The calculation model is as follows:

[0050] Among them, LVC represents the task load variation coefficient of each computing node, L j It is represented as the load value of the jth sampling point, and AL is represented as the average load.

[0051] In this embodiment, it should be specifically explained that the LVC task load variation coefficient is used to measure the magnitude of load fluctuation; if the task load variation coefficient value is large, it means that the load changes drastically; conversely, if the task load variation coefficient value is small, it indicates that the load is relatively stable.

[0052] The resource utilization analysis module is used to calculate the resource utilization change coefficient of each computing node, analyze the change trend of resource utilization according to the resource utilization change coefficient, generate a resource utilization report, and transmit the resource utilization change coefficient to the anomaly detection and early warning module;

[0053] In this embodiment, it should be specifically explained that the steps for calculating the resource utilization rate variation coefficient of each computing node are as follows:

[0054] Step S1: Calculate the average utilization AU. The calculation model is as follows:

[0055] Among them, U= { U1,U2,...,U n } It represents the CPU utilization set collected within two hours, n represents the number of sampling points; i represents the CPU utilization sampling point, i = 1, 2, 3, ... n;

[0056] Step S2: Calculate the standard deviation of utilization σ u, the calculation model is as follows:

[0057] Among them, U i It is expressed as the CPU utilization of the i-th sampling point;

[0058] Step S3: Calculate the resource utilization change coefficient of each computing node. The calculation model is as follows:

[0059] Among them, RVC represents the resource utilization change coefficient of each computing node, AU represents the average utilization, σ u is denoted as the standard deviation of utilization.

[0060] In this embodiment, it should be specifically explained that the RVC resource utilization variation coefficient is used to measure the magnitude of resource utilization fluctuations; if the resource utilization variation coefficient value is large, it means that the resource utilization changes dramatically; conversely, if the resource utilization variation coefficient value is small, it indicates that the resource utilization is relatively stable.

[0061] The adaptive adjustment module is used to adaptively adjust the allocation of computing resources, storage resources and network resources using elastic scaling adjustment strategies according to changes in task resource requirements, task loads and resource utilization;

[0062] The anomaly detection and early warning module is used to compare the task load change coefficient and the resource utilization change coefficient with the preset thresholds to determine whether an anomaly occurs, use a machine learning algorithm to perform anomaly detection based on the abnormal data, and generate early warning information to remind the administrator to handle it in time;

[0063] In this embodiment, it should be specifically explained that the specific content of the abnormal detection and early warning module comparing the task load change coefficient with the preset threshold value is:

[0064] Extract the task load variation coefficient LVC of each computing node and obtain the task load variation deviation formula: If the task load change coefficient is less than or equal to the preset load change threshold LVC 预 , then the judgment result is that the task load is normal. If the task load change coefficient is greater than the preset load change threshold LVC 预 , the result is judged as abnormal task load, the abnormal result is sent to the management personnel, and the automatic load balancing module is started.

[0065] In this embodiment, it should be specifically explained that the specific content of the abnormal detection and early warning module comparing the resource utilization rate change coefficient with the preset threshold value is:

[0066] Extract the resource utilization change coefficient RVC of each computing node and obtain the resource utilization change deviation formula: If the resource utilization change coefficient is less than or equal to the preset resource utilization change threshold RVC 预 , then the judgment result is that the resource utilization is normal. If the resource utilization change coefficient is greater than the preset resource utilization change threshold RVC 预 , the result is that the resource utilization is abnormal, the abnormal result is sent to the management personnel, and the automatic load balancing module is started.

[0067] The automatic load balancing module automatically adjusts the distribution of tasks on each computing node according to the task load change coefficient and resource utilization change coefficient in the abnormal detection and early warning module and the preset threshold judgment result, and uses a genetic algorithm to optimize the task scheduling strategy;

[0068] The task priority setting module is used for managers to set task priorities and prioritize task resources according to the urgency and importance of tasks;

[0069] The human-computer interaction module is used to receive the warning information and task priority division information from the abnormal detection and warning module, so that management personnel can manage and make decisions on abnormal data.

[0070] Finally: The above description is only a preferred embodiment of the present invention and is not intended to limit the present invention. Any modifications, equivalent substitutions, improvements, etc. made within the spirit and principles of the present invention should be included in the protection scope of the present invention.

[0071] The above is only a specific implementation of the present application, but the protection scope of the present application is not limited thereto. Any person skilled in the art who is familiar with the present technical field can easily think of changes or substitutions within the technical scope disclosed in the present application, which should be included in the protection scope of the present application. Therefore, the protection scope of the present application should be based on the protection scope of the claims.

Claims

1. A cloud computing and big data integrated machine, characterized in that: include: Task resource extraction module: used to receive and parse task requests, extract task requirement information, evaluate the complexity of task resources based on the task requirement information, and generate a task resource requirement report; Task load detection module: used to calculate the task load variation coefficient of each computing node, analyze the change trend of the task load according to the task load variation coefficient, generate a task load report, and transmit the task load variation coefficient to the anomaly detection and warning module; Resource utilization analysis module: used to calculate the resource utilization change coefficient of each computing node, analyze the change trend of resource utilization according to the resource utilization change coefficient, generate a resource utilization report, and transmit the resource utilization change coefficient to the anomaly detection and early warning module; Adaptive adjustment module: used to adaptively adjust the allocation of computing resources, storage resources, and network resources using elastic scaling adjustment strategies according to changes in task resource requirements, task load, and resource utilization; Anomaly detection and early warning module: used to compare the task load change coefficient and resource utilization change coefficient with the preset thresholds to determine whether anomalies occur. Based on the abnormal data, a machine learning algorithm is used to perform anomaly detection and generate early warning information to remind the administrator to handle it in time. Automatic load balancing module: automatically adjusts the distribution of tasks on each computing node according to the task load change coefficient and resource utilization change coefficient in the anomaly detection and early warning module and the preset threshold judgment result, and uses genetic algorithm to optimize the task scheduling strategy; Task priority setting module: used by managers to set task priorities and prioritize task resources according to the urgency and importance of tasks; Human-computer interaction module: used to receive the warning information and task priority division information from the abnormal detection and warning module, so that managers can manage and make decisions on abnormal data.

2. The cloud computing and big data integrated machine according to claim 1, characterized in that: The task requirement information includes the computing resources, storage resources and network resource requirement information required for the task. The computing resources include the task load value, CPU utilization and processing power hardware facilities: central processing unit CPU, graphics processing unit GPU, field programmable gate array FPGA and accelerator; the storage resources cover the equipment and technologies used for data storage, including hard disk drives HDD, solid state drives SSD, memory RAM and distributed storage systems; the network resource requirement information includes bandwidth, latency, throughput, network topology and service quality parameters.

3. The cloud computing and big data integrated machine according to claim 1, characterized in that: The calculation steps of the task load variation coefficient of each computing node are as follows: Step S1: Calculate the average load AL. The calculation model is as follows: Where, L = { L1,L2,…,L m } It is represented as a set of task load values ​​collected within two hours, m is represented as the number of sampling points; j is represented as a sampling point of the task load value, j = 1, 2, 3, ... m; Step S2: Calculate the task load variation coefficient of each computing node. The calculation model is as follows: Among them, LVC represents the task load variation coefficient of each computing node, L j It is represented as the load value of the jth sampling point, and AL is represented as the average load.

4. The cloud computing and big data integrated machine according to claim 1, characterized in that: The steps for calculating the resource utilization change coefficient of each computing node are specifically as follows: Step S1: Calculate the average utilization AU. The calculation model is as follows: Among them, U= { U1,U2,…,U n } It represents the CPU utilization set collected within two hours, n represents the number of sampling points; i represents the CPU utilization sampling point, i = 1, 2, 3, ... n; Step S2: Calculate the standard deviation of utilization σ u, the calculation model is as follows: Among them, U i It is expressed as the CPU utilization of the i-th sampling point; Step S3: Calculate the resource utilization change coefficient of each computing node. The calculation model is as follows: Among them, RVC represents the resource utilization change coefficient of each computing node, AU represents the average utilization, σ u is denoted as the standard deviation of utilization.

5. The cloud computing and big data integrated machine according to claim 1, characterized in that: The specific content of the abnormal detection and early warning module comparing the task load change coefficient with the preset threshold value is: Extract the task load variation coefficient LVC of each computing node and obtain the task load variation deviation formula: If the task load change coefficient is less than or equal to the preset load change threshold LVC 预 , then the judgment result is that the task load is normal. If the task load change coefficient is greater than the preset load change threshold LVC 预 , the result is judged as abnormal task load, the abnormal result is sent to the management personnel, and the automatic load balancing module is started.

6. The cloud computing and big data integrated machine according to claim 1, characterized in that: The specific content of the abnormal detection and early warning module comparing the resource utilization rate change coefficient with the preset threshold value is: Extract the resource utilization change coefficient RVC of each computing node and obtain the resource utilization change deviation formula: If the resource utilization change coefficient is less than or equal to the preset resource utilization change threshold RVC 预 , then the judgment result is that the resource utilization is normal. If the resource utilization change coefficient is greater than the preset resource utilization change threshold RVC 预 , the result is that the resource utilization is abnormal, the abnormal result is sent to the management personnel, and the automatic load balancing module is started.

Citation Information

Cited By

  • Enterprise-level cloud computing resource dynamic allocation and management system and implementation method thereof

    CN120935183A