A Multi-Source Computing Power Data Integration and Intelligent Scheduling System and Method
By establishing a unified resource abstraction layer, dynamic data integration module and adaptive intelligent scheduling engine, the problem of low efficiency of existing computing resource scheduling methods is solved, efficient, flexible and precise scheduling of multi-source heterogeneous computing resources is achieved, and the adaptability and data processing capabilities of the system are enhanced.
Patent Information
- Application Number
- CN202411200939.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-08-29
- Publication Date
- 2025-07-29
- Estimated Expiration
- 2044-08-29
AI Technical Summary
The existing computing resource scheduling methods are not intelligent and flexible enough, resulting in low execution efficiency, serious idle resources and waste, and insufficient adaptability to network fluctuations and load changes.
Establish a unified resource abstraction layer for virtualization management, build a dynamic data integration module and an adaptive intelligent scheduling engine, design an edge-cloud-terminal collaborative computing framework, integrate secure multi-party computing and differential privacy technologies, determine elastic scaling mechanisms, and realize intelligent scheduling of multi-source heterogeneous computing resources.
It improves the utilization rate of multi-source heterogeneous computing resources, can match task requirements more intelligently, more flexible and more accurately, reduce resource idleness and waste, enhance the system's ability to adapt to network fluctuations and load changes, and improve data processing capabilities and data quality control.
Smart Images

Figure CN118916147B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of intelligent control, and particularly to a multi-source computing power data integration and intelligent scheduling system and method. Background Art
[0002] In today's rapidly developing information age, computing power resources have become a key driving force for social progress. With the rapid development of cutting-edge technologies such as cloud computing, big data, and artificial intelligence, the scheduling and management of computing power resources face many new challenges.
[0003] How to achieve the efficient allocation and utilization of computing power resources has become the focus of attention in the industry. Computing power scheduling refers to the process of, after receiving a computing power task request, allocating the task to appropriate computing power resources for processing through an optimization algorithm according to the task characteristics and computing power requirements. Existing computing power task scheduling methods are not intelligent and flexible enough, and the execution efficiency is relatively low. Summary of the Invention
[0004] Based on the above problems, the present invention proposes a multi-source computing power data integration and intelligent scheduling system and method. Through the solution of the present invention, the utilization rate of multi-source heterogeneous computing resources can be greatly improved, and the task requirements and available resources can be matched more intelligently, flexibly, and accurately, reducing resource idleness and waste; the adaptability of the system to network fluctuations and load changes can be improved, and the data processing ability and data quality control of the system can be enhanced.
[0005] In view of this, one aspect of the present invention proposes a multi-source computing power data integration and intelligent scheduling system, including: a central control server, a cloud platform, edge devices, and local terminals;
[0006] The central control server is configured to:
[0007] Establish a unified resource abstraction layer to standardize the description of multi-source heterogeneous computing resources and perform virtualization management;
[0008] Construct a dynamic data integration module for realizing the real-time collection, cleaning, and standardization processing of multi-source heterogeneous data;
[0009] Design an adaptive intelligent scheduling engine, integrate machine learning algorithms, and perform task allocation and resource scheduling according to historical task execution data and real-time system state data;
[0010] Construct an edge-cloud-terminal collaborative computing framework that can support the dynamic migration and load balancing of computing tasks;
[0011] Establish a global performance optimization model, and continuously optimize the overall performance and resource utilization rate of the system through reinforcement learning technology;
[0012] Integrate secure multi-party computation and differential privacy technologies to ensure data privacy protection and security during integration and processing;
[0013] Determine an elastic scaling mechanism to support dynamic adjustment of the system scale and seamless access to new types of computing resources.
[0014] Another aspect of the present invention provides a multi-source computing power data integration and intelligent scheduling method, including:
[0015] Establish a unified resource abstraction layer to standardize the description of multi-source heterogeneous computing resources and perform virtualization management;
[0016] Construct a dynamic data integration module for realizing real-time acquisition, cleaning, and standardization processing of multi-source heterogeneous data;
[0017] Design an adaptive intelligent scheduling engine, integrate machine learning algorithms, and perform task allocation and resource scheduling based on historical task execution data and real-time system status data;
[0018] Construct an edge-cloud-terminal collaborative computing framework that can support dynamic migration and load balancing of computing tasks;
[0019] Establish a global performance optimization model and continuously optimize the overall system performance and resource utilization through reinforcement learning technology;
[0020] Integrate secure multi-party computation and differential privacy technologies to ensure data privacy protection and security during integration and processing;
[0021] Determine an elastic scaling mechanism to support dynamic adjustment of the system scale and seamless access to new types of computing resources.
[0022] Optionally, the step of establishing a unified resource abstraction layer to standardize the description of multi-source heterogeneous computing resources and perform virtualization management includes:
[0023] Identify all types of computing resources in the system;
[0024] Classify the computing resources and establish a resource type hierarchy;
[0025] Design a general computing resource description model;
[0026] Define the key attributes of each type of computing resource;
[0027] Develop an automated tool to collect detailed information about each specific computing resource;
[0028] Map the collected information into the computing resource description model;
[0029] Design a virtualization management interface to abstract the underlying hardware details;
[0030] Implement a resource pooling mechanism to convert physical resources into manageable virtual resource units;
[0031] Design a unified API for resource allocation, monitoring, and recycling;
[0032] Generate a cross-platform resource management protocol;
[0033] Deploy distributed monitoring agents to collect resource usage in real time;
[0034] Design a central monitoring panel to display the overall resource status;
[0035] Build a dynamic resource discovery and registration mechanism;
[0036] Build a role-based access control system;
[0037] Design an encrypted communication channel to protect resource information transmission;
[0038] Build a caching mechanism to accelerate the query of frequently accessed resource information;
[0039] Optimize the storage and retrieval algorithms for computing resource descriptions;
[0040] Design a modular architecture to support the access of new computing resource types;
[0041] Build a distributed resource management framework to support large-scale clusters.
[0042] Optionally, the steps of building a dynamic data integration module for realizing real-time collection, cleaning, and standardization processing of multi-source heterogeneous data include:
[0043] Identify all data source types, classify the data sources, and establish a data source directory;
[0044] Define a common data structure and design a metadata model to describe the attributes, sources, and timestamps of the data;
[0045] Develop dedicated data collectors for each data source type and build a data access layer that supports multiple protocols;
[0046] Build a stream processing framework and design and implement a streaming data processing pipeline;
[0047] Design a data format conversion function, develop a data verification and error handling mechanism, and design a data cleaning algorithm;
[0048] Perform data field mapping and conversion, develop a unit conversion and standardization function library, and design a data semantic conversion mechanism;
[0049] Build a real-time data quality detection algorithm, design a data quality scoring system, and develop an abnormal data processing mechanism;
[0050] Design a data fusion algorithm to handle the conflicts and complementarities of multi-source data, and develop a data correlation analysis function and a data consistency check mechanism;
[0051] Build a data caching mechanism and design a parallel processing architecture to optimize data storage and indexing strategies;
[0052] Adopt a microservices architecture, determine dynamic configuration management, and design a distributed processing framework;
[0053] Develop a data flow monitoring dashboard to achieve visual display of the data processing process, and design an alarm and notification mechanism.
[0054] Optionally, the steps of designing an adaptive intelligent scheduling engine, integrating machine learning algorithms, and performing task allocation and resource scheduling according to historical task execution data and real-time system status data include:
[0055] Design a data collection module according to the data collector, and use the data collection module to obtain historical task execution data and real-time system status data;
[0056] Perform a data cleaning and feature extraction process on the historical task execution data and real-time system status data, and build a data storage system to support fast query and analysis;
[0057] Define a task feature set according to the features extracted from the historical task execution data, develop a task classification algorithm, identify the type of tasks, and build a task requirement prediction model;
[0058] Design a resource status feature set in combination with the features extracted from the real-time system status data, and develop a resource performance prediction model and a resource availability evaluation algorithm;
[0059] Select a machine learning algorithm, and design a model training process and a model evaluation / verification mechanism;
[0060] According to the task requirement prediction model, the resource performance prediction model, and the resource availability evaluation algorithm, develop a resource scheduling decision model based on a machine learning model, determine a multi-objective optimization mechanism, and design a scheduling strategy scoring system;
[0061] Build a real-time inference module to generate scheduling decisions;
[0062] Determine the decision execution and feedback mechanism;
[0063] Develop an anomaly detection and handling process.
[0064] Optionally, the steps of constructing an edge-cloud-terminal collaborative computing framework that can support dynamic migration and load balancing of computing tasks include:
[0065] Develop lightweight edge devices for preliminary data processing and analysis at the location of the data source;
[0066] Design a communication protocol between edge devices to support distributed collaboration;
[0067] Develop dynamic resource management and task scheduling algorithms to achieve load balancing between edge devices;
[0068] Establish an elastic and scalable cloud platform for processing large-scale data and complex computing tasks;
[0069] Develop cloud platform resource management and task scheduling algorithms based on machine learning to achieve dynamic migration and load balancing across edge devices;
[0070] Provide API interfaces to support two-way communication and collaboration between edge devices and the cloud platform;
[0071] Formulate a collaboration mechanism between edge devices and the cloud platform, including task allocation strategies, data transmission protocols, and security authentication mechanisms;
[0072] Develop cross-layer resource monitoring and scheduling algorithms to achieve dynamic load balancing of edge-cloud-terminals;
[0073] Design a fault tolerance mechanism to ensure the continuous and reliable operation of computing tasks.
[0074] Optionally, the steps of establishing a global performance optimization model to continuously optimize the overall system performance and resource utilization through reinforcement learning technology include:
[0075] Determine the key metrics describing the overall system performance;
[0076] Establish a global performance model covering edge devices, communication networks, local terminals, and cloud platforms to describe the relationships and influencing factors among the metrics at each layer;
[0077] Define the system state space;
[0078] Design a reinforcement learning algorithm suitable for this problem;
[0079] Determine the reward function with the overall system performance metric as the goal;
[0080] Establish a real-time monitoring mechanism for the edge-cloud to collect system operation status data;
[0081] Design data preprocessing and feature extraction methods to provide input data for reinforcement learning;
[0082] Utilize the pre - processed and feature - extracted operation status data to train a reinforcement learning model and learn the optimal resource scheduling and task migration strategies;
[0083] Deploy the trained model into the system to perform real - time performance optimization decisions;
[0084] Continuously monitor the system operation status and iteratively optimize the reinforcement learning model;
[0085] Integrate the reinforcement learning optimization module into the edge - cloud - terminal collaboration framework to ensure the coordinated operation of each component.
[0086] Optionally, the steps of integrating secure multi - party computation and differential privacy technologies to ensure privacy protection and security during data integration and processing include:
[0087] Clarify the privacy requirements and compliance requirements of all parties in the system;
[0088] Design a comprehensive privacy protection framework based on secure multi - party computation and differential privacy, covering all links of data collection, transmission, storage, and processing;
[0089] Select the corresponding secure multi - party computation algorithm;
[0090] Design a secure communication and computation protocol among multiple parties to ensure that each party can only obtain the final result and cannot peek into the intermediate process;
[0091] Implement the secure multi - party computation module and integrate it with other components of the system;
[0092] Identify the privacy - sensitive data in the system and determine the corresponding differential privacy protection objectives;
[0093] Design a differential privacy algorithm based on noise injection to ensure data availability while minimizing privacy leakage;
[0094] Integrate the differential privacy module into the data processing flow to ensure privacy protection of data in each link;
[0095] Establish a privacy risk assessment mechanism to regularly evaluate the privacy protection effect of the system;
[0096] Design a real - time monitoring module to detect possible privacy leakage events and trigger response measures;
[0097] According to the evaluation results and monitoring feedback, continuously optimize the privacy protection plan;
[0098] Design targeted privacy protection test scenarios covering different attack models and data processing scenarios;
[0099] Verify the effectiveness and reliability of secure multi - party computation and differential privacy technologies through testing;
[0100] Optimize and improve according to the test results.
[0101] Optionally, the step of determining the elastic scaling mechanism to support the dynamic adjustment of the system scale and the seamless access of new computing resources includes:
[0102] Analyze the business requirements and service characteristics of the system, and determine the goals and key indicators of elastic scaling;
[0103] Design the architecture and mechanism of elastic scaling according to the business requirements and technical characteristics;
[0104] Establish a system resource monitoring module to collect CPU, memory, and network metric data in real time;
[0105] Based on the monitoring data and historical load patterns, use machine learning methods to predict the future trend of resource demand changes;
[0106] According to the prediction results of the resource demand change trend, formulate a dynamic scaling decision-making strategy;
[0107] Implement an automated scaling execution mechanism;
[0108] Design an access mechanism for new computing resources to support the seamless docking of heterogeneous hardware devices and cloud services;
[0109] Implement resource discovery, scheduling, and orchestration functions to ensure that new resources can be effectively utilized by the system;
[0110] Establish an optimization algorithm based on resource utilization and service metrics to allocate and schedule computing resources;
[0111] Adopt an appropriate resource scheduling strategy according to the resource characteristics and application requirements;
[0112] Design redundant backup and self-healing mechanisms to ensure that the system can still operate smoothly in case of resource failures or fluctuations;
[0113] Design functions for fault detection, isolation, and automatic recovery to improve the reliability of the system.
[0114] Optionally, the step of using machine learning methods to predict the future trend of resource demand changes based on the monitoring data and historical load patterns includes:
[0115] Collect monitoring data from cloud platforms, edge devices, and local terminals in real time;
[0116] Perform operations such as cleaning, missing value filling, and outlier detection on the collected monitoring data to obtain the first monitoring data;
[0117] Extract appropriate feature variables from the first monitoring data according to the business characteristics and monitoring metrics;
[0118] Normalize and standardize the feature variables to ensure the same dimension between features;
[0119] Select a machine learning algorithm suitable for time series prediction as the model to be trained;
[0120] Obtain the historical monitoring data and historical load patterns of the cloud platform, edge devices, and terminals;
[0121] Divide the historical monitoring data into a training set and a validation set, train and cross-validate the model to be trained, and obtain the first model;
[0122] Evaluate the performance metrics of the first model and iteratively optimize the first model;
[0123] Using the trained first model, input the historical load pattern and feature variables, and predict the future resource demand change trends for different resource types within a certain period of time.
[0124] Adopting the technical solution of the present invention, the multi-source computing power data integration and intelligent scheduling method includes establishing a unified resource abstraction layer to standardize and virtualize the description of multi-source heterogeneous computing resources; constructing a dynamic data integration module for realizing real-time collection, cleaning, and standardization processing of multi-source heterogeneous data; designing an adaptive intelligent scheduling engine, integrating machine learning algorithms, and performing task allocation and resource scheduling according to historical execution task data and real-time system status data; constructing an edge-cloud-terminal collaborative computing framework that can support the dynamic migration and load balancing of computing tasks; establishing a global performance optimization model to continuously optimize the overall system performance and resource utilization through reinforcement learning technology; integrating secure multi-party computing and differential privacy technologies to ensure data privacy protection and security during the integration and processing process; determining an elastic scaling mechanism to support the dynamic adjustment of the system scale and the seamless access of new computing resources. Through this solution, the utilization rate of multi-source heterogeneous computing resources can be greatly improved, and it can more intelligently, flexibly, and accurately match task requirements and available resources, reducing resource idleness and waste; it can improve the system's adaptability to network fluctuations and load changes and enhance the system's data processing ability and data quality control. BRIEF DESCRIPTION OF THE DRAWINGS
[0125] Figure 1 is a schematic block diagram of a multi-source computing power data integration and intelligent scheduling system provided by an embodiment of the present invention;
[0126] Figure 2 is a flowchart of a multi-source computing power data integration and intelligent scheduling method provided by an embodiment of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0127] In order to more clearly understand the above-mentioned objects, features and advantages of the present invention, the present invention is further described in detail below in conjunction with the accompanying drawings and specific embodiments. It should be noted that, in the absence of conflict, the embodiments of the present application and the features therein can be combined with each other.
[0128] In the following description, many specific details are set forth to facilitate a full understanding of the present invention. However, the present invention may also be implemented in other ways different from those described herein. Therefore, the scope of protection of the present invention is not limited to the specific embodiments disclosed below.
[0129] The terms "first", "second", etc. in the specification and claims of this application and the above-mentioned drawings are used to distinguish different objects, rather than to describe a specific order. In addition, the terms "including" and "having" and any variations thereof are intended to cover non-exclusive inclusions. For example, a process, method, system, product or device that includes a series of steps or units is not limited to the listed steps or units, but optionally includes steps or units that are not listed, or optionally includes other steps or units inherent to these processes, methods, products or devices.
[0130] Reference to "embodiments" herein means that a particular feature, structure, or characteristic described in conjunction with the embodiments may be included in at least one embodiment of the present application. The appearance of the phrase in various locations in the specification does not necessarily refer to the same embodiment, nor is it an independent or alternative embodiment that is mutually exclusive with other embodiments. It is explicitly and implicitly understood by those skilled in the art that the embodiments described herein may be combined with other embodiments.
[0131] Refer to the following Figures 1 to 2 To describe a multi-source computing power data integration and intelligent scheduling system and method provided according to some embodiments of the present invention.
[0132] like Figure 1 As shown, one embodiment of the present invention provides a multi-source computing power data integration and intelligent scheduling system, including: a central control server, a cloud platform, an edge device and a local terminal;
[0133] The central control server is configured to:
[0134] Establish a unified resource abstraction layer to standardize the description of multi-source heterogeneous computing resources and perform virtualization management;
[0135] Build a dynamic data integration module for real-time collection, cleaning and standardization of multi-source heterogeneous data;
[0136] Design an adaptive intelligent scheduling engine, integrate machine learning algorithms, and perform task allocation and resource scheduling based on historical task execution data and real-time system status data;
[0137] Build an edge-cloud-terminal collaborative computing framework that can support the dynamic migration and load balancing of computing tasks;
[0138] Establish a global performance optimization model and continuously optimize the overall system performance and resource utilization through reinforcement learning technology;
[0139] Integrate secure multi-party computing and differential privacy technologies to ensure data privacy protection and security during the integration and processing;
[0140] Determine an elastic scaling mechanism to support the dynamic adjustment of the system scale and the seamless access of new computing resources.
[0141] It should be known that Figure 1 The block diagram of the multi-source computing power data integration and intelligent scheduling system shown is only for illustration, and the number of each module shown does not limit the protection scope of the present invention. The multi-source computing power data integration and intelligent scheduling system provided in this embodiment can be used to execute the embodiments of the corresponding multi-source computing power data integration and intelligent scheduling method. For the specific implementation process, please refer to the descriptions of the following method embodiments, which will not be elaborated here.
[0142] Please refer to Figure 2 , another embodiment of the present invention provides a multi-source computing power data integration and intelligent scheduling method, including:
[0143] Establish a unified resource abstraction layer, standardize the description of multi-source heterogeneous computing resources and perform virtualization management;
[0144] In this step, the unified resource abstraction layer includes: a resource description model for standardizing the description of different types of computing resources; a virtualization management interface for uniformly managing physical resources and virtual resources; a resource monitoring module for real-time monitoring of the status and performance indicators of various resources.
[0145] Build a dynamic data integration module for realizing the real-time collection, cleaning, and standardization processing of multi-source heterogeneous data;
[0146] In this step, the dynamic data integration module includes: a distributed data collection subsystem that supports parallel collection of multi-source data; a real-time data stream processing engine for immediate cleaning and conversion of data; a data quality evaluation and control mechanism to ensure the consistency and reliability of the integrated data.
[0147] Design an adaptive intelligent scheduling engine, integrate machine learning algorithms, and perform task allocation and resource scheduling based on historical task execution data and real-time system status data;
[0148] In this step, the adaptive intelligent scheduling engine includes: a multi-objective optimization algorithm that comprehensively considers factors such as task completion time, resource utilization rate, and energy consumption; a prediction model that predicts task execution time and resource requirements based on historical data; and a dynamic scheduling policy generator that dynamically adjusts the scheduling policy according to the system state and prediction results.
[0149] Build an edge-cloud-terminal collaborative computing framework that can support the dynamic migration and load balancing of computing tasks;
[0150] In this step, the edge-cloud-terminal collaborative computing framework includes: a task decomposition module that decomposes complex tasks into subtasks that can be executed at different levels; a dynamic task migration mechanism that dynamically adjusts the task execution location according to network conditions and computing load; and a distributed state synchronization mechanism that ensures data consistency between different levels.
[0151] Establish a global performance optimization model to continuously optimize the overall system performance and resource utilization rate through reinforcement learning technology;
[0152] In this step, the global performance optimization model includes: a performance metric collector that collects real-time performance data from all aspects of the system; a reinforcement learning model that optimizes resource allocation and scheduling decisions through continuous learning; and a simulation environment that is used to simulate and evaluate the effects of different scheduling policies.
[0153] Integrate secure multi-party computing and differential privacy technologies to ensure privacy protection and security during data integration and processing;
[0154] In this step, the integration of secure multi-party computing and differential privacy technologies includes: a data encryption module that ensures the security of data during transmission and storage; a secure multi-party computing protocol that supports cross-domain computing while protecting privacy; and a differential privacy processor that adds noise to data analysis and output results to protect individual privacy.
[0155] Determine an elastic scaling mechanism to support the dynamic adjustment of the system scale and the seamless access of new computing resources.
[0156] In this step, the elastic scaling mechanism includes: an automatic resource discovery module that can automatically identify and access newly added computing resources; a dynamic load balancer that dynamically adjusts resource allocation according to the system load; and horizontal and vertical scaling support that allows the system to adjust its scale without interrupting services.
[0157] In an embodiment of the present invention, a "system" refers to a computing system or infrastructure environment that requires resource scheduling and management, and can be one of the following situations: a cloud computing environment (such as a public cloud, private cloud, or hybrid cloud platform, including various virtual computing resources), a data center (including various hardware resources such as physical servers, storage devices, network devices, etc.), an edge computing environment (computing resources deployed at devices close to the data source and at the network edge), a high-performance computing cluster (a parallel computing system composed of multiple computing nodes, including heterogeneous resources such as CPUs and GPUs), an embedded system (such as industrial control devices, Internet of Things devices, etc., which may include special hardware such as FPGAs and dedicated accelerators), and so on. Whether it is a cloud, data center, edge, or embedded system, it is necessary to identify and manage various heterogeneous computing resources in the system to provide efficient computing services for upper-layer applications; this is also the basis for implementing intelligent scheduling strategies, and it is necessary to accurately master the types and performance indicators of available computing resources in the system. In short, a "system" here refers to a computing environment that requires resource management and scheduling, and can be different forms of computing infrastructure such as a cloud platform, data center, edge device, etc., or an overall composed of these facilities. Identifying various heterogeneous computing resources in the system is a prerequisite for implementing intelligent scheduling.
[0158] Adopting the technical solution of this embodiment, through the unified resource abstraction layer and the adaptive intelligent scheduling engine, the utilization rate of multi-source heterogeneous computing resources can be significantly improved; it can more accurately match task requirements and available resources, reducing resource idleness and waste; the edge-cloud-terminal collaborative computing framework enables the system to flexibly adjust the execution location of computing tasks according to real-time situations, which can not only optimize performance but also improve the system's adaptability to network fluctuations and load changes; the dynamic data integration module supports large-scale and real-time heterogeneous data processing, significantly enhancing the system's data processing ability and data quality control; the global performance optimization model can continuously improve the overall performance of the system, including key indicators such as task completion time and throughput, through continuous learning and optimization; the integrated secure multi-party computing and differential privacy technologies ensure that while improving computing efficiency, sensitive data and user privacy can also be protected, making the system applicable to a wider range of application scenarios; the elastic scaling mechanism enables the system to easily handle scale changes, seamlessly integrate new computing resources, and improve the system's scalability and adaptability; the unified management interface and automated scheduling mechanism can significantly reduce the management complexity of large-scale heterogeneous systems and reduce the need for human intervention; through machine learning algorithms, the system can more accurately predict task execution time and resource requirements, and thus make better scheduling decisions; intelligent scheduling and global optimization can improve energy utilization efficiency, reduce unnecessary energy consumption, and contribute to achieving green computing; by optimizing resource allocation and task scheduling, the system can provide more stable and responsive services for end users. Generally speaking, the solution of this embodiment achieves remarkable technical effects in multiple aspects such as resource utilization efficiency, system performance, data processing ability, security, scalability, and management convenience.
[0159] In some possible embodiments of the present invention, the step of establishing a unified resource abstraction layer to standardize the description of multi-source heterogeneous computing resources and perform virtualized management includes:
[0160] Identifying all types of computing resources in the system (such as CPUs, GPUs, FPGAs, dedicated accelerators, etc.);
[0161] In this step, the system's hardware topology can be automatically scanned through the system management tool or API to obtain information on various computing resources, including CPU models, GPU models, FPGA models, dedicated accelerator models, etc.; collect performance monitoring metrics for various computing resources, such as CPU utilization, GPU load, FPGA resource occupancy, etc.; the monitoring tools built into the operating system or third-party monitoring platforms can be used to extract characteristic metrics for various computing resources based on the collected monitoring data, such as CPU instruction set support, GPU memory bandwidth, FPGA logic unit utilization, etc.; an identification model for computing resource types can be established by combining resource characteristics; machine learning algorithms can be used to identify and classify unknown resources. In this way, it is possible to comprehensively and automatically identify various computing resources involved in the system and obtain their detailed performance characteristics. Through the solution in this step, accurately identifying various computing resources helps to optimize resource allocation and scheduling and improve utilization; for heterogeneous computing resources, dynamic scheduling can be performed according to the real-time load to improve the scalability of the system; the complete resource topology and performance data help to build a comprehensive resource visualization and monitoring system; the rich resource characteristic data can be used as the input of the machine learning model to achieve more accurate resource demand prediction. In short, this computing resource identification technology can provide key support for the dynamic resource management of heterogeneous systems and improve the overall utilization efficiency of computing resources.
[0162] Classify the computing resources and establish a hierarchical structure of resource types;
[0163] In this step, through system scanning and monitoring, basic information of various computing resources is collected, including model, manufacturer, architecture, etc. These metadata form the basis for resource classification. According to the collected resource metadata, attributes such as functional characteristics and performance indicators of various resources are analyzed, such as the instruction set of the CPU, the graphics processing ability of the GPU, and the parallel computing characteristics of the FPGA. Combining the types of computing resources identified in the previous step, different types of computing resources are hierarchically classified according to similar attribute characteristics. A tree structure or a network structure can be used to organize the relationships between different types of resources. Standardized tags are defined for each type of resource to describe the type, performance, usage, etc. of the resource. The tag system is beneficial for subsequent resource retrieval and scheduling. The classified resources are organized according to the system topology structure to form a logical topology map of the resources. The topology model reflects the hierarchical and interconnection relationships between resources. Through the above steps, a complete hierarchical structure of computing resource types can be established, realizing a clear resource classification system and tags, enhancing the visualization and manageability of resources. Based on the resource topology model, more intelligent resource discovery, selection, and scheduling can be achieved. The hierarchical resource classification is applicable to various heterogeneous computing environments, enhancing generality. Standardized resource types and tags are beneficial for cross-system resource sharing and collaboration. The resource classification system provides important input features for machine learning models, supporting automated resource management. In short, establishing a hierarchical structure of computing resource types is the key foundation for realizing dynamic and intelligent resource management and can bring significant technical effects.
[0164] Design a general computing resource description model (such as based on XML or JSON);
[0165] In this step, according to the aforementioned resource classification system / hierarchical structure of resource types, the basic elements and attributes required to describe computing resources are determined, such as resource type, performance indicators, functional characteristics, deployment environment, etc. The resource elements and attributes are organized into a data model in XML or JSON format. The model should have a good hierarchical structure and scalability. Syntax rules for the resource description language are formulated, such as element names, attribute definitions, value types, etc. The semantic definitions of various resource attributes are clarified to ensure cross-system interoperability. Software tools for parsing, querying, and generating resource description documents are developed to support the verification, conversion, and management of resource description documents. Such a general resource description language can achieve: a standardized description method, which helps resources migrate between different systems; a unified description language, which facilitates resource discovery, orchestration, and scheduling; cross-system resource description, which provides a basis for resource sharing and collaboration; structured resource information, which provides rich input for machine learning models; and enhanced information interaction between different systems based on the standard description language.
[0166] Define the key attributes of each type of computing resource (such as processing power, memory size, energy consumption characteristics, etc.);
[0167] In this step, various main types of computing resources, such as CPUs, GPUs, FPGAs, storage devices, etc., can be identified according to the aforementioned resource classification system / resource type hierarchy; for each resource type, systematically analyze its key functional characteristics and performance metrics; for example, the processing power, memory size, instruction set support, etc. of the CPU, the floating-point computing performance, video memory capacity, memory bandwidth, etc. of the GPU, and the capacity, IOPS, latency, etc. of the storage device; for each resource type, determine a set of parameter metrics that can comprehensively describe its characteristics; the parameters should cover all aspects such as the processing power, storage capacity, energy consumption characteristics, and deployment environment of the resource. Taking CPU resources as an example, the following key attribute parameters can be defined: processor architecture (x86, ARM, RISC-V, etc.), number of cores (physical cores of the CPU), number of threads (logical threads of the CPU), instruction set (instruction set support such as SSE, AVX, NEON, etc.), clock frequency (CPU main frequency), cache (L1 / L2 / L3 cache size), memory support (memory type, capacity, bandwidth), power consumption (rated power consumption or thermal design power consumption), performance (single-core performance metrics, multi-core performance metrics), deployment environment (operating system, virtualization support); organize the key attributes of various resources into a standardized data model for subsequent resource description and management, and XML, JSON or other structured formats can be used to define the attribute model; collect and maintain the attribute data of various computing resources to build a comprehensive resource attribute database; the database should support functions such as attribute query, comparison, and analysis. Defining the key attributes of each computing resource in this way can achieve a comprehensive description of resource attributes, enhance the visualization and manageability of resources; based on the intelligent matching of the attribute model, more accurate resource selection and scheduling can be achieved; the standardized attribute definition is applicable to computing resources of different types and different manufacturers; the sharing of attribute data is conducive to cross-system resource collaboration and business cooperation; the structured attribute information provides important input features for machine learning models.
[0168] Develop an automated tool to collect detailed information about each specific computing resource;
[0169] Map the collected information to the computing resource description model;
[0170] Design a virtualization management interface to abstract the underlying hardware details;
[0171] Implement a resource pooling mechanism to convert physical resources into manageable virtual resource units;
[0172] Design a unified API for resource allocation, monitoring, and recycling;
[0173] Generate a cross-platform resource management protocol;
[0174] Deploy distributed monitoring agents to collect resource usage in real time;
[0175] Design a central monitoring panel to display the overall resource status;
[0176] Build a dynamic resource discovery and registration mechanism;
[0177] Build a role-based access control (RBAC) system;
[0178] Design an encrypted communication channel to protect the transmission of resource information;
[0179] Build a caching mechanism to accelerate the query of frequently accessed resource information;
[0180] Optimize the storage and retrieval algorithms for computing resource descriptions;
[0181] Design a modular architecture to support the access of new computing resource types;
[0182] In this step, a set of standardized resource interfaces can be designed to describe and manage different types of computing resources; the interfaces should include key functions such as resource discovery, resource description, and resource allocation; the interface definition should follow the principles of openness and extensibility to facilitate the access of new resource types in the future; based on the standard interfaces, a unified resource management layer is constructed to be responsible for coordinating and scheduling various types of computing resources; the resource management layer should possess core functions such as resource registration, resource monitoring, and resource allocation; the management layer can adopt a microservices or component-based design to improve flexibility and extensibility; for each type of computing resource, a corresponding resource adapter component is developed; the adapter is responsible for converting the attributes and functions of specific resources into the format defined by the standard interfaces; the adapter can be an independent module or integrated into the resource management layer; the resource management layer should provide a convenient extension mechanism to support the dynamic access of new resource types; new resources only need to implement the adapters of the standard interfaces to seamlessly integrate into the resource management system; above the resource management layer, a resource orchestration layer is constructed to be responsible for the coordinated scheduling of cross-heterogeneous resources; the orchestration layer can flexibly combine different types of computing resources according to business requirements; the orchestration layer should provide a declarative resource orchestration language to facilitate users to define resource topologies and constraints. Through the above design, a modular computing resource management architecture can be realized. Through the standardized interface design, the access of new types of resources becomes simple and efficient; the modular architecture supports dynamic expansion and can support the development of new technologies without a full-scale reconstruction; the unified resource management layer can coordinate and schedule different types of computing resources; the resource adapter mechanism shields the underlying resource details and provides a unified abstract view; the resource orchestration layer supports the declarative definition of resource topologies and realizes the coordinated scheduling of cross-heterogeneous resources; the orchestration layer can allocate and optimize resources according to business requirements to improve resource utilization efficiency; the unified resource management interface is conducive to the monitoring and visual presentation of resource status; the standardized resource attribute definition provides a basis for realizing automated resource management. In short, through the modular architecture design, the access of new computing resource types can be easily supported, and the openness, flexibility, and automation level of the system can be greatly enhanced.
[0183] Build a distributed resource management framework to support large-scale clusters.
[0184] Through the solution of this embodiment, the unified management of heterogeneous resources is realized, simplifying the system complexity; improving the resource management efficiency and reducing the management overhead; through the global view, better resource allocation and load balancing can be achieved; resource fragmentation is reduced and the overall utilization efficiency is improved; dynamic resource scheduling is supported, and the system can quickly respond to changes in computing requirements; it is convenient to introduce new types of computing resources and improve the scalability of the system; detailed and standardized resource information is provided for intelligent scheduling algorithms; more accurate task-resource matching is achieved and the computing efficiency is improved; a unified resource access interface is provided to reduce the complexity of application development; cross-platform deployment and migration of applications are supported; through real-time monitoring, resource failures can be quickly discovered and processed; dynamic replacement and failover of resources are supported; through better resource allocation strategies, energy use is optimized; intelligent scheduling based on energy consumption characteristics is supported; the unified access control mechanism enhances the system security, and the fine-grained resource management supports more accurate security policies; through better resource planning and utilization, the hardware investment and operation costs are reduced; cost-based intelligent resource allocation is supported; the standardized resource description helps to more accurately estimate the task performance and supports performance modeling and optimization based on historical data. Through this unified resource abstraction layer, the system can more effectively manage and utilize multi-source heterogeneous computing resources, thus significantly improving the overall computing efficiency and system performance.
[0185] In some possible implementation manners of the present invention, the steps of constructing a dynamic data integration module for realizing real-time collection, cleaning and standardization processing of multi-source heterogeneous data include:
[0186] Identify all data source types (such as databases, file systems, APIs, sensors, etc.), classify the data sources, and establish a data source directory;
[0187] In this step, by comprehensively sorting out all types of data sources used inside and outside the organization, including databases, file systems, API interfaces, sensors, etc.; collecting basic information about each data source, such as name, type, location, connection method, etc.; classifying the data sources according to their characteristics, such as storage type, access method, data format, etc.; a hierarchical classification system can be adopted, for example, classified as structured, semi-structured, and unstructured according to the storage type; classified as batch, real-time, etc. according to the access method; organize the data source information collected through research into a comprehensive data source directory; the directory should contain detailed metadata for each data source, such as name, type, location, connection method, data format, owner, etc.; adopt a unified data model and metadata standard to ensure the consistency and readability of the directory information; establish a management and maintenance mechanism for the data source directory to ensure the timely update and accuracy of the directory information; formulate standardized processes for the addition, change, and withdrawal of data sources and implement automated management; consider using metadata management tools or data catalog services to improve management efficiency; based on the directory, establish a registration and discovery mechanism for data sources; allow data users to search for and access the required data sources through the directory; functions such as search, browsing, and subscription can be provided to facilitate users to quickly locate the required data; through this step, a comprehensive and structured data source directory can be established, which centrally displays the detailed information of all data sources within the organization, providing a comprehensive perspective for data management; helps to sort out data assets, discover data silos, and improve data sharing and reuse; the standardized data source registration and discovery mechanism enables data users to conveniently search for the required data sources; improves the accessibility and availability of data assets, reduces the cost of data use; the unified metadata standard and management mechanism ensure the consistency and credibility of the data source directory information; provide basic support for data governance, data quality control, etc.; based on the data source directory, cross-heterogeneous data source orchestration and integration can be achieved; provide standardized data access capabilities for data analysis, business applications, etc.; the standardized processes and automated tools for directory management improve the efficiency and accuracy of data source information maintenance; enhance the agility and sustainability of data asset management. In short, establishing a comprehensive data source directory can bring a panoramic view of data assets to the organization, improve data discoverability, support data governance, and lay a foundation for data integration and automated management.
[0188] Define a common data structure (such as based on JSON or Avro), and design a metadata model to describe the attributes, sources, and timestamps of the data;
[0189] Develop dedicated data collectors for each data source type and build a data access layer that supports multiple protocols (such as MQTT, HTTP, WebSocket, etc.);
[0190] Build stream processing frameworks (such as Apache Flink, Kafka Streams), and design and implement streaming data processing pipelines;
[0191] Design data format conversion functions, develop data validation and error handling mechanisms, and design data cleaning algorithms (such as deduplication, complementation, etc.);
[0192] Perform data field mapping and conversion, develop unit conversion and standardization function libraries, and design data semantic conversion mechanisms;
[0193] Build real-time data quality detection algorithms, design data quality scoring systems, and develop abnormal data handling mechanisms;
[0194] Design data fusion algorithms, handle conflicts and complementarities of multi-source data, and develop data correlation analysis functions and data consistency checking mechanisms;
[0195] Build data caching mechanisms (to improve the processing speed of frequently accessed data) and design parallel processing architectures (to improve data processing throughput), and optimize data storage and indexing strategies;
[0196] Adopt a microservices architecture (to support independent expansion of modules), determine dynamic configuration management (to support immediate access to new data sources), and design a distributed processing framework (to support large-scale data processing);
[0197] Develop a data flow monitoring dashboard to achieve visual display of the data processing process, and design alarm and notification mechanisms.
[0198] The solution of this embodiment can achieve real-time data processing, reduce data latency, and improve data throughput through parallel processing; improve data consistency and reliability through automated cleaning and standardization, and reduce the impact of incorrect data through real-time data quality control; be able to process heterogeneous data in different formats and structures, and realize the correlation analysis and integration of multi-source data; support the rapid access and configuration of new data sources, and adapt to different data processing requirements and scenarios; the standardized data is easier to analyze and utilize, provides a unified data access interface, and simplifies application development; supports real-time data analysis and decision-making support, and can quickly respond to data changes and anomalies; reduces storage requirements through data cleaning and compression, and has intelligent data hierarchical storage to optimize costs and performance; supports large-scale data processing and adapts to the growth of data volume; the modular design is convenient for function expansion and upgrade; the real-time monitoring and alarm mechanism improves system stability, and the data backup and recovery mechanism enhances data security; provides high-quality and standardized data for machine learning and AI analysis, and supports complex data mining and predictive analysis; supports data traceability, facilitates auditing and compliance management, and realizes data privacy protection and sensitive information processing; automated data processing reduces manual intervention; improves data utilization efficiency and reduces data management costs. Through this dynamic data integration module, the system can efficiently process multi-source heterogeneous data, provide high-quality and real-time data support, and lay a solid data foundation for subsequent intelligent scheduling and analysis and decision-making.
[0199] In some possible embodiments of the present invention, the step of designing an adaptive intelligent scheduling engine, integrating machine learning algorithms, and performing task allocation and resource scheduling according to historical task execution data and real-time system state data includes:
[0200] Design a data collection module according to the data collector, and use the data collection module to obtain historical task execution data and real-time system state data;
[0201] Perform a data cleaning and feature extraction process on the historical task execution data and real-time system state data, and construct a data storage system to support fast query and analysis;
[0202] Define a task feature set (such as computational complexity, memory requirements, data dependencies, etc.) according to the features extracted from the historical task execution data, develop a task classification algorithm, identify the type of the task, and construct a task requirement prediction model;
[0203] In this step, analyze the historical task execution data to identify the key features affecting task execution, such as computational complexity, memory requirements, data dependencies, etc.; based on these features, establish a task feature set to provide a basis for subsequent task classification and prediction. The feature set should include quantitative metrics and qualitative descriptions to comprehensively reflect the characteristics of the tasks; based on the defined task feature set, develop suitable task classification algorithms, such as decision trees, K-Means clustering, etc.; use the historical task data for algorithm training and validation to ensure classification accuracy; the classification algorithm should be able to automatically identify and classify newly submitted tasks; collect the execution data of historical tasks, including resource consumption, execution time, quality metrics, etc.; combined with the task feature set, establish a task requirement prediction model, such as a regression model, neural network, etc.; the prediction model can estimate the required resources, execution time, quality metrics, etc. according to the task features; deploy the trained task classification algorithm and requirement prediction model into the actual task management system; when a new task is submitted, automatically perform type identification and requirement prediction to provide a decision-making basis for task scheduling, resource allocation, etc.; as new task data continues to accumulate, regularly update the model parameters to improve prediction accuracy. Through this step, the key features affecting task execution can be comprehensively extracted, laying a foundation for subsequent classification and prediction; the definition of the feature set helps to deeply understand the inherent characteristics of the tasks; the task classification algorithm based on machine learning can accurately identify the types of new tasks; the automated task classification improves the efficiency and accuracy of task management; the prediction model established using historical data can estimate the required resources, time, etc. for new tasks; accurate requirement prediction helps to reasonably schedule tasks and allocate resources, improving the overall operation efficiency; deployed in the actual task management system, it can continuously collect new data and update the model parameters; it has the ability of self-learning and optimization, and continuously improves the prediction accuracy as the business develops; it provides quantitative basis and intelligent suggestions for key decisions such as task scheduling and resource allocation; it helps to improve the scientificity and rationality of task management and reduce the operation cost. In short, through task feature definition, classification algorithm development and requirement prediction model construction, automated task type identification and requirement prediction can be achieved, providing strong support for the enterprise's task management and resource optimization.
[0204] Design a resource status feature set (such as CPU usage, memory occupancy, network bandwidth, etc.) by combining the features extracted from the real-time system status data, and develop a resource performance prediction model and a resource availability assessment algorithm;
[0205] In this step, analyze the real-time system monitoring data to identify key features reflecting the resource status, such as CPU usage, memory occupancy, network bandwidth, etc.; based on these features, establish a comprehensive resource status feature set to provide a basis for subsequent performance prediction and availability assessment; the feature set should include quantitative indicators and dynamic change trends to comprehensively reflect the real-time status of resources; based on the defined resource status feature set, develop suitable performance prediction models, such as time series analysis, neural networks, etc.; use historical monitoring data for model training and verification to ensure prediction accuracy; the prediction model can predict the performance change trend in the future period according to the current resource status; combine the resource status feature set and the performance prediction model to develop a resource availability assessment algorithm; this algorithm can comprehensively consider the current status and future prediction of resources to evaluate whether it can meet business requirements; the availability assessment indicators can include resource usage thresholds, performance indicator warnings, etc.; deploy the trained performance prediction model and availability assessment algorithm into the actual resource management system; when new resource status data is collected, automatically perform performance prediction and availability assessment to provide decision-making basis for resource scheduling, expansion, etc.; as new monitoring data accumulates, regularly update the model parameters to improve prediction accuracy and assessment effect. Through this step, key features reflecting the real-time status of resources can be comprehensively extracted, laying a foundation for subsequent prediction and assessment; the definition of the feature set helps to deeply understand the operation status of resources; the prediction model based on machine learning can accurately predict the performance change trend of resources in the future period; accurate performance prediction provides a basis for resource scheduling and expansion, improving resource utilization efficiency; use the performance prediction results to develop an algorithm for comprehensively evaluating resource availability; the availability assessment results help to timely discover resource bottlenecks and take corresponding countermeasures; deployed in the actual resource management system, it can continuously collect new monitoring data and update model parameters; it has the ability of self-learning and optimization, and continuously improves prediction accuracy and assessment effect with the development of business; it provides quantitative basis and intelligent suggestions for key decisions such as resource scheduling and expansion; it helps to improve the scientificity and rationality of resource management and reduce operation costs. In short, through the definition of resource status features, the development of performance prediction models and the construction of availability assessment algorithms, the perception and prediction of the real-time status of resources can be realized, providing strong support for the resource management and business guarantee of enterprises.
[0206] Select machine learning algorithms (such as random forest, deep neural network, reinforcement learning, etc.), design the model training process (including feature selection, model parameter tuning, etc.) and model evaluation / verification mechanism;
[0207] According to the task requirements prediction model, resource performance prediction model and resource availability assessment algorithm, develop a resource scheduling decision model based on machine learning models, determine the multi-objective optimization mechanism (to balance factors such as performance, energy consumption, cost, etc.), and design a scheduling strategy scoring system;
[0208] In this step, historical monitoring data can be utilized, and methods such as time series analysis and deep learning can be adopted to predict the quantity, type of tasks and their resource requirements within a certain period in the future; the prediction results provide a basis for subsequent resource scheduling; according to business objectives, multiple metrics to be optimized are determined, such as performance, energy consumption, cost, etc.; a multi-objective optimization model is established to balance the conflicts and trade-offs between different objectives and seek the optimal balance point; the optimization objectives can be to maximize system throughput, minimize energy consumption cost, etc.; using the results of the task demand prediction model, resource performance prediction model and resource availability evaluation algorithm, a resource scheduling decision model is constructed; machine learning methods such as reinforcement learning and genetic algorithms are adopted to automatically generate a scheduling strategy that meets multi-objective optimization; in the process of generating the scheduling strategy, factors such as task priority, resource utilization rate, and energy consumption cost can be considered; a scoring mechanism for the scheduling strategy is established, comprehensively considering multiple dimensions such as performance, energy consumption, and cost; according to the predicted task requirements and resource availability, the execution effects of various scheduling schemes are simulated; the scheduling strategy with the highest score is selected as the final execution plan; the generated scheduling strategy is deployed to the actual resource management system; new task requirements and resource status data are continuously collected, and the prediction model and scheduling strategy are updated regularly; according to the actual execution effect, the multi-objective optimization weights are adjusted, and the scheduling strategy is continuously optimized. Through this step, based on the machine learning-based task demand prediction model, the future workload can be predicted more accurately; it provides a reliable basis for resource scheduling and improves the pertinence of the scheduling scheme; comprehensively considering multiple business objectives such as performance, energy consumption, and cost, it seeks the optimal balance; it meets the specific requirements in different scenarios and improves the flexibility of resource management; using the machine learning model to automatically generate a scheduling strategy that meets multi-objective optimization; the process of generating the scheduling strategy is more intelligent and automated, reducing the manual intervention cost; a comprehensive scheduling strategy scoring mechanism is established, comprehensively considering multiple key metrics; by simulating the execution effect, the optimal scheduling strategy is selected to improve the accuracy of decision-making; the scheduling system is deployed to the actual environment, continuously collecting new monitoring data and updating the model; it has the ability of self-learning and optimization, and continuously improves the scheduling effect as the business develops. In short, the resource scheduling decision model based on the machine learning model can achieve more intelligent and automated resource scheduling, provide an excellent performance, low energy consumption and reasonable cost resource management solution for the enterprise, and significantly improve the operation efficiency of the entire system.
[0209] Build a real-time inference module to generate scheduling decisions;
[0210] Determine the decision execution and feedback mechanism;
[0211] Develop an anomaly detection and handling process.
[0212] In this embodiment, it may further include: designing an online learning algorithm to continuously optimize the model; implementing a model update and version control mechanism; developing an A / B test framework to evaluate the performance of new and old models; developing an interface with the underlying resource management system to implement the execution logic of task allocation and resource allocation; designing a rollback and fault recovery mechanism; building a comprehensive performance metric monitoring system, developing a scheduling effect analysis tool, implementing a visualization dashboard to display key performance metrics; designing a modular architecture to support the loose coupling integration of each component; implementing a distributed computing framework to improve large-scale scheduling capabilities; optimizing the system response time and throughput; implementing a data encryption and access control mechanism; designing a privacy protection algorithm to protect sensitive information; developing an audit log system to track the scheduling decision-making process.
[0213] The solution of this embodiment uses machine learning algorithms to achieve more accurate task-resource matching, reduce human intervention, and improve scheduling speed and accuracy; based on historical data and real-time status, it realizes the dynamic allocation of resources, reduces resource waste, and improves overall utilization; the adaptive learning mechanism enables the system to cope with changing workloads and environments and supports the rapid integration of new tasks and resources; through machine learning models, it improves the prediction accuracy of task execution time and resource requirements; supports better load balancing and capacity planning; balances multiple objectives such as performance, energy consumption, and cost to achieve global optimality; supports dynamically adjusting optimization objectives according to different scenarios; improves system stability through an anomaly detection and handling mechanism; real-time monitoring and automatic adjustment reduce system failures; intelligent scheduling considers energy consumption factors to optimize overall energy use; supports green computing and sustainable development goals; provides interpretability of scheduling decisions to increase system credibility; supports the auditing and backtracking of the decision-making process; the distributed architecture supports the efficient scheduling of large-scale clusters, improving the scalability and throughput of the system; through better resource allocation, it reduces task waiting time; improves system response speed and service quality; reduces hardware and power costs by optimizing resource utilization and energy efficiency; reduces human intervention and operation and maintenance costs; provides rich data and insights for system performance analysis and optimization, supporting predictive maintenance and proactive optimization; ensures the security of sensitive data through a privacy protection mechanism. Through this adaptive intelligent scheduling engine, the system can achieve more efficient, flexible, and intelligent task allocation and resource scheduling, significantly improving the overall system performance and resource utilization efficiency, while providing better service quality for users.
[0214] In some possible embodiments of the present invention, the steps of constructing an edge-cloud-terminal collaborative computing framework that can support the dynamic migration and load balancing of computing tasks include:
[0215] Develop lightweight edge devices (such as embedded devices, IoT devices, etc.) for performing preliminary data processing and analysis at the location of the data source.
[0216] Design a communication protocol between edge devices to support distributed collaboration;
[0217] Develop dynamic resource management and task scheduling algorithms to achieve load balancing among edge devices;
[0218] Build a cloud platform with elastic scalability for processing large-scale data and complex computing tasks;
[0219] Develop cloud platform resource management and task scheduling algorithms based on machine learning to achieve dynamic migration and load balancing across edge devices;
[0220] Provide API interfaces to support two-way communication and collaboration between edge devices and the cloud platform;
[0221] Formulate a collaboration mechanism between edge devices and the cloud platform, including task allocation strategies, data transmission protocols, and security authentication mechanisms;
[0222] Develop cross-layer resource monitoring and scheduling algorithms to achieve dynamic load balancing among edge-cloud-terminals;
[0223] Design a fault tolerance mechanism to ensure the continuous and reliable operation of computing tasks.
[0224] Through the solution of this embodiment, a collaborative computing framework among edge devices-cloud platform-local terminals that supports dynamic migration and load balancing can be constructed. Edge devices can quickly process data and reduce latency; the cloud can handle complex computing tasks and improve the overall computing power; dynamic resource management and task scheduling can be carried out to cope with load fluctuations; the fault tolerance mechanism ensures the continuity of critical tasks; cross-layer load balancing can be achieved to improve resource utilization; edge devices and cloud resources can be flexibly expanded according to requirements and support diverse application scenarios. Generally speaking, this edge-cloud-terminal collaborative computing framework can make full use of edge device and cloud resources, achieve dynamic migration and load balancing of computing tasks, and thus improve the performance, reliability, and flexibility of the overall system.
[0225] In this embodiment, it also includes: implementing an end-to-end encryption mechanism to protect data transmission and storage security; developing a distributed identity authentication system to ensure cross-layer access control; designing a privacy-preserving computing framework to support the secure processing of sensitive data; implementing a task checkpoint mechanism to support the rapid recovery of tasks; developing a fault detection and isolation system to improve system reliability; designing a task rescheduling strategy to handle node failures; implementing a preloading and predictive execution mechanism to reduce task startup time; developing a resource reservation strategy to optimize the execution efficiency of critical tasks; designing an adaptive computing offloading strategy to balance local and remote execution; designing a unified task submission and management interface and developing visualization monitoring and debugging tools; providing an SDK and documentation to facilitate the integration of third-party applications; implementing the integration of each module and end-to-end testing; developing a performance benchmark test suite to evaluate the overall system performance; conducting large-scale simulation tests to verify the system's performance in complex scenarios.
[0226] The solution of this embodiment makes full use of the computing resources of each layer of the edge, cloud, and terminal through dynamic task allocation; reduces resource idleness and improves overall computing efficiency; selects the optimal execution location according to task characteristics and resource status; reduces task waiting time and improves response speed through load balancing; supports dynamic resource expansion and contraction to adapt to load changes; realizes the self-regulation and optimization of the system through task migration; reduces the computing burden on end devices and extends battery life; reduces network latency and improves interaction performance through proximity computing; reduces unnecessary data transmission through intelligent data transmission strategies; uses edge nodes for data preprocessing to reduce the network pressure on the cloud; sensitive data can be processed locally or at edge nodes to reduce the risk of data exposure; the distributed security mechanism improves the overall security of the system; edge nodes can provide basic services during network interruptions to achieve local caching and asynchronous synchronization of data; improves the robustness of the system through a multi-level fault tolerance mechanism; the fault isolation and rapid recovery capabilities enhance the usability of the system; the unified task description and scheduling mechanism supports the collaborative work of different hardware platforms; makes full use of the computing power of dedicated hardware (such as GPUs, NPUs); reduces unnecessary data center computing and energy consumption through intelligent task allocation; supports green computing and optimizes overall energy utilization; the unified development framework simplifies the application development and deployment process and supports rapid prototype verification and iterative optimization; supports distributed data processing and analysis to improve big data processing capabilities; edge intelligence provides support for real-time decision-making; reduces hardware investment and operating costs through optimized resource allocation; the flexible computing model supports pay-per-use to optimize the cost structure. Through this edge-cloud-terminal collaborative computing framework, the system can achieve the efficient utilization of computing resources, the intelligent allocation and dynamic migration of tasks, and the effective balance of global load; this not only improves the overall system performance and resource utilization rate, but also provides users with better quality of service and experience, and at the same time provides strong support for the development of new applications and services.
[0227] In some possible embodiments of the present invention, the step of establishing a global performance optimization model and continuously optimizing the overall system performance and resource utilization through reinforcement learning technology includes:
[0228] Determine key metrics that describe the overall system performance (such as response latency, throughput, energy consumption, etc.);
[0229] Establish a global performance model covering edge devices, communication networks, local terminals, and cloud platforms to describe the relationships and influencing factors among the metrics at each level;
[0230] Define the system state space (including key factors such as resource allocation and task load);
[0231] Design a reinforcement learning algorithm suitable for this problem (such as Q-learning, policy gradient, etc.);
[0232] Determine the reward function with the overall system performance metric as the goal;
[0233] Establish a real-time monitoring mechanism between the edge and the cloud to collect system operation status data;
[0234] Design data preprocessing and feature extraction methods to provide input data for reinforcement learning;
[0235] Use the preprocessed and feature-extracted operation status data to train the reinforcement learning model and learn the optimal resource scheduling and task migration strategies;
[0236] In this step, through system monitoring, various operating status metrics such as CPU / GPU utilization, memory usage, and network load are collected; the collected raw data is preprocessed, such as cleaning and normalization, to ensure data quality; combined with business knowledge, effective features related to resource scheduling are extracted to prepare for subsequent reinforcement learning modeling; the resource scheduling and task migration problems are modeled as a Markov decision process (MDP); the state space (system operating status), action space (scheduling strategy and task migration plan), reward function (optimization goal), etc. are defined; deep reinforcement learning algorithms, such as DQN, PPO, etc., are used to train the agent to learn the optimal scheduling and migration strategies; the preprocessed training data is used to iteratively train the reinforcement learning model; the reward function, neural network structure, hyperparameters, etc. are adjusted to optimize the model performance; an offline training and online fine-tuning combination method is adopted to improve the model generalization ability; the trained reinforcement learning model is deployed to the actual resource management system; the scheduling and migration effects of the model are tested in the production environment, and feedback data is collected for further optimization; a monitoring metric system is established to regularly evaluate the operating status and optimization potential of the model. Through this step, the reinforcement learning model can automatically learn the optimal resource scheduling strategy according to the system operating status; the scheduling strategy can dynamically adapt to changes in business requirements and system states, improving resource utilization efficiency; the reinforcement learning model can learn when and where to migrate tasks to maximize system performance and minimize migration overhead; the task migration strategy can make intelligent decisions based on multiple factors such as computing load, energy consumption, and cost; by designing a reasonable reward function, the reinforcement learning model can balance multiple optimization goals such as performance, energy consumption, and cost; find the best balance point between various goals, improving the overall efficiency of resource management; after the model is deployed to the actual system, it can continuously collect new monitoring data and continuously optimize its decision-making strategy; showing the ability of self-learning and self-adaptation, continuously improving as business requirements change. In short, using reinforcement learning technology to achieve intelligent resource scheduling and task migration can greatly improve the resource utilization efficiency and overall performance of the computing system, providing more flexible and reliable computing services for enterprises.
[0237] Deploy the trained model to the system and execute performance optimization decisions in real time;
[0238] Continuously monitor the system operating status and iteratively optimize the reinforcement learning model;
[0239] Integrate the reinforcement learning optimization module into the edge-cloud-terminal collaboration framework to ensure the coordinated operation of each component.
[0240] Through the solution of this embodiment, a global performance optimization model based on reinforcement learning can be established, and continuous optimization can be achieved. The resource allocation and task deployment can be dynamically adjusted to optimize the key performance indicators; it can adaptively cope with complex operating environment changes and maintain high performance; it can learn the optimal resource scheduling strategy and make full use of edge device and cloud resources; it can achieve cross-level dynamic load balancing and eliminate resource bottlenecks; the reinforcement learning model can be continuously optimized without manual intervention; it can be flexibly configured and extended for different application scenarios; it can automatically execute performance optimization decisions and reduce manual operation and maintenance costs; it can improve system reliability, reduce failure rate and maintenance costs. Generally speaking, the global performance optimization model based on reinforcement learning can enable the edge-cloud collaborative computing system to achieve autonomous, intelligent and efficient resource management, thereby greatly improving the overall system performance and resource utilization efficiency and providing excellent computing services for various application scenarios.
[0241] Through the solution of this embodiment, the system can continuously learn and adapt to new workloads and environmental changes; through experience accumulation, the decision-making quality is gradually improved; considering the overall system situation, it can avoid global sub-optimality caused by local optimization; it can balance short-term and long-term goals and achieve better resource allocation; the system can automatically adjust strategies to adapt to different workloads and hardware configurations; it can reduce manual intervention and operation and maintenance complexity; through the prediction of future states, it can achieve forward-looking resource scheduling; it can anticipate possible performance bottlenecks and resource competition in advance; at the same time, considering multiple performance indicators, it can find the best balance point; it supports dynamic adjustment of optimization goals to adapt to changes in business requirements; it can learn strategies for handling various anomalies and extreme situations and enhance the robustness and fault tolerance of the system; it can achieve more refined and dynamic resource allocation and reduce resource waste; it can make full use of the system potential and improve the overall efficiency; it can learn energy-saving strategies to reduce energy consumption while ensuring performance; it can support the realization of green computing goals; according to the characteristics of different users or applications, it can provide customized performance optimization; it can improve user satisfaction and service quality; the system can quickly learn and utilize the characteristics of newly added hardware or software components to accelerate the implementation and revenue realization of innovative technologies; by analyzing the decision-making process of the model, it can provide interpretable optimization suggestions to assist managers in understanding system behavior and performance bottlenecks; it can reduce the need for manual configuration and tuning and achieve automated operation and maintenance; it can free up the time of operation and maintenance personnel to focus on higher-level optimization tasks; the model can adapt to the dynamic changes in system scale; it can support the rapid integration and optimization of new services and applications; through more efficient resource utilization, it can reduce hardware investment and operating costs; it can optimize cloud resource usage and achieve better cost control. Through this global performance optimization model based on reinforcement learning, the system can achieve continuous self-optimization and evolution, continuously improve performance and resource utilization rate, and at the same time have strong adaptability and scalability; this can not only improve the overall efficiency of the system, but also bring significant economic benefits and competitive advantages to enterprises.
[0242] In some possible embodiments of the present invention, the steps of integrating secure multi-party computation and differential privacy technologies to ensure privacy protection and security during data integration and processing include:
[0243] Clarify the privacy requirements and compliance requirements of all parties in the system (such as users, enterprises, third parties);
[0244] Design a comprehensive privacy protection framework based on secure multi-party computation and differential privacy, covering all aspects of data collection, transmission, storage, and processing;
[0245] Select corresponding secure multi-party computation algorithms (such as Yao's Garbled Circuit, Secret Sharing, etc.);
[0246] Design secure communication and computation protocols among multiple parties to ensure that each party can only obtain the final result and cannot peek into the intermediate process;
[0247] Implement the secure multi-party computation module and integrate it with other components of the system;
[0248] Identify privacy-sensitive data in the system and determine the corresponding differential privacy protection objectives;
[0249] Design a differential privacy algorithm based on noise injection to ensure data availability while minimizing privacy leakage;
[0250] Integrate the differential privacy module into the data processing flow to ensure privacy protection of data in all aspects;
[0251] Establish a privacy risk assessment mechanism to regularly evaluate the privacy protection effect of the system;
[0252] Design a real-time monitoring module to detect possible privacy leakage events and trigger response measures;
[0253] Continuously optimize the privacy protection solution based on the evaluation results and monitoring feedback;
[0254] Design targeted privacy protection test scenarios covering different attack models and data processing scenarios;
[0255] Verify the effectiveness and reliability of secure multi-party computation and differential privacy technologies through testing;
[0256] Optimize and improve according to the test results.
[0257] Through the solution of this embodiment, secure multi-party computing and differential privacy technology can be integrated into the system to achieve comprehensive data privacy protection. Secure multi-party computing ensures that all parties can only obtain the final result and cannot spy on the intermediate process; differential privacy technology injects noise interference to weaken the attacker's ability to infer sensitive data; differential privacy algorithms minimize privacy leakage while ensuring the availability of data analysis and applications; the system can securely integrate and process data without reducing the value of the data; the privacy protection solution complies with the requirements of relevant laws, regulations and industry standards; the real-time monitoring mechanism can quickly discover and respond to privacy leakage incidents and improve compliance; the privacy protection technology is decoupled from other modules of the system and has good scalability; the privacy protection strategy can be flexibly configured and optimized according to different application scenarios and privacy requirements. In general, by integrating secure multi-party computing and differential privacy technology, the privacy security of data in the system can be effectively protected, and the availability and value of data can be maximized while meeting compliance requirements.
[0258] The solution of this embodiment protects individual privacy during the data processing process, prevents the leakage of identity and sensitive information, and can ensure that individual privacy cannot be easily inferred even in the case of data leakage; supports computing on encrypted data without decrypting the original data; allows multiple parties to perform joint computing while protecting their respective data privacy; reduces the risk of single-point attacks through distributed storage and computing. Even if some participating parties are compromised, it will not lead to the leakage of global data; can perform statistical analysis and machine learning while protecting privacy; provides query results with differential privacy protection to prevent the inference of individual information through multiple queries; meets the requirements of strict data protection regulations such as GDPR and CCPA, provides technical guarantees for enterprise privacy compliance, and reduces legal risks; supports cross-organizational data sharing and analysis on the premise of protecting privacy; promotes data-driven innovation while protecting the privacy rights and interests of individuals and organizations; enhances users' trust in the system through a transparent privacy protection mechanism; gives users more data control rights and improves user participation; restricts the scope of data use through technical means to prevent data from being used for unauthorized purposes; realizes "burn after reading" of data to reduce the risks brought by long-term data storage; trains machine learning models while protecting the privacy of the original data to prevent privacy leakage caused by model reverse engineering; encourages users to provide more real and complete data through privacy protection technology; reduces data loss or distortion caused by privacy concerns; provides complete audit logs for easy supervision and internal control; supports the quantitative evaluation and continuous improvement of privacy protection effects; maintains system performance while protecting privacy through optimized algorithms and hardware acceleration; supports dynamically adjusting the privacy protection level according to the scenario to balance usability and privacy; defends against advanced privacy attack means such as inference attacks and reconstruction attacks; provides long-term privacy protection and remains effective even in the face of future improvements in computing power; realizes complex analysis tasks with privacy protection on large-scale data sets and provides privacy protection solutions for data-intensive applications. By integrating these technologies, the system can ensure the privacy and security of data while providing efficient data processing and analysis capabilities; this not only meets the increasingly strict regulatory requirements but also provides a reliable foundation for data-driven innovation and promotes the healthy development of the data economy.
[0259] In some possible implementation manners of the present invention, the step of determining the elastic scaling mechanism to support the dynamic adjustment of the system scale and the seamless access of new computing resources includes:
[0260] Analyze the business requirements and service characteristics of the system to determine the goals and key metrics of elastic scaling (such as response time, throughput, etc.);
[0261] Design the architecture and mechanism of elastic scaling according to the business requirements and technical characteristics (such as how to monitor resource usage, when to trigger expansion / contraction, etc.);
[0262] Establish a system resource monitoring module to collect CPU, memory, and network metric data in real time;
[0263] Based on the monitoring data and historical load patterns, use machine learning methods to predict the future trend of resource demand changes;
[0264] According to the prediction results of the resource demand change trend, formulate a dynamic scaling decision-making strategy (how to determine when to increase / decrease resources);
[0265] In this step, evaluate the current resource supply situation and conduct a comparative analysis with the predicted demand changes; when the predicted demand exceeds the current supply capacity, automatically trigger the scaling-up operation, otherwise trigger the scaling-down operation; automatically convert the formulated scaling decision into specific resource configuration change operations; continuously monitor the scaling process to ensure that the target service quality indicators are within an acceptable range. In this embodiment, this dynamic scaling strategy based on the prediction of resource demand change trends can perform precise scaling according to the actual demand change prediction results, avoiding the situation of resource surplus or shortage; maximizing the utilization of existing resources and reducing the overall resource cost; perceiving the resource demand change trend in advance and adjusting the resource supply in a timely manner to ensure the stability of business service indicators; automating the decision-making and execution processes of resource scaling, reducing the manual operation and maintenance cost and response time; quickly adapting to the changes in resource demand and improving the overall elasticity and reliability of the IT system. In short, this dynamic scaling strategy can effectively improve the utilization efficiency of cloud computing resources and the quality of business services.
[0266] Implement an automated scaling execution mechanism (able to quickly add or release computing resources);
[0267] Design an access mechanism for new computing resources to support seamless docking of heterogeneous hardware devices and cloud services;
[0268] Implement resource discovery, scheduling, and orchestration functions to ensure that new resources can be effectively utilized by the system;
[0269] Establish an optimization algorithm based on resource utilization rate and service indicators to allocate and schedule computing resources;
[0270] According to the resource characteristics and application requirements, adopt appropriate resource scheduling strategies (such as load balancing, resource affinity, etc.);
[0271] Design a redundant backup and self-healing mechanism to ensure that the system can still operate smoothly in case of resource failures or fluctuations;
[0272] Design functions for fault detection, isolation, and automatic recovery to improve the reliability of the system.
[0273] Through the solution of this embodiment, a complete elastic scaling mechanism can be established to achieve dynamic adjustment of the system scale and seamless access to new computing resources, automatically allocate computing resources according to actual needs, and avoid resource surplus or shortage; new resources can be quickly accessed to improve the overall computing power and resource utilization efficiency; the resource configuration can be adjusted in real time according to business requirements to ensure that service metrics such as response time and throughput meet the standards; an automated fault detection and self-healing mechanism can improve the reliability and sustainability of the service; the dynamic scaling of resources and the access of new resources can be completed without manual intervention; an automated monitoring, decision-making, and execution mechanism can reduce the workload of operation and maintenance personnel; it supports the rapid access of new computing resources such as heterogeneous hardware devices and cloud services; the elastic scaling ability is decoupled from other system modules to improve the flexibility of the overall architecture. Generally speaking, by determining the elastic scaling mechanism, the system can dynamically adjust the resource scale according to actual needs and quickly access new computing devices, thereby improving resource utilization efficiency, ensuring service quality, reducing operation and maintenance costs, and enhancing the flexibility of the overall architecture.
[0274] Through the solution of this embodiment, the system can automatically adjust its scale according to the load change, achieving a seamless transition from small scale to large scale; support rapid response to burst traffic to ensure service quality; dynamically adjust resource allocation to avoid resource waste; minimize resource costs while ensuring performance; automatically expand capacity to handle high loads and prevent system crashes; quickly isolate and replace faulty nodes to improve system fault tolerance; maintain stable response time and throughput even when the load changes; reduce human intervention and lower the risk of operation and maintenance errors; allocate resources on demand to avoid over-configuration; support cost-based intelligent decision-making to optimize operating expenses; easily integrate new types of computing resources, such as dedicated hardware accelerators; support rapid deployment and testing of new services to accelerate the innovation cycle; achieve unified resource management in multi-cloud and hybrid-cloud environments to improve system flexibility; support dynamic resource scheduling across regions to optimize the global user experience; achieve disaster recovery and business continuity guarantee; automatically adjust system configuration according to different workload characteristics; support dynamic switching between different environments such as development, testing, and production; automated scaling operations reduce the need for human intervention and simplify the management complexity of large-scale systems; support independent scaling in a multi-tenant environment to ensure secure isolation during the expansion process; support fine-grained service-level scaling to optimize specific functional modules; achieve application-aware resource allocation to improve application performance; optimize energy consumption by dynamically adjusting computing resources; support green computing goals and reduce the carbon footprint; ensure data consistency and integrity during dynamic scaling; support smooth expansion of stateless and stateful services; reduce service latency and interruption through dynamic resource adjustment; support personalized performance guarantees to meet the needs of different user groups. Through this elastic scaling mechanism, the system can achieve efficient utilization of resources, dynamic optimization of performance, and effective control of costs; not only improve the reliability and scalability of the system, but also provide an enterprise with a technical foundation to cope with complex and changing market environments, supporting the rapid development and innovation of the business.
[0275] In some possible implementation manners of the present invention, the step of predicting the future resource demand change trend using a machine learning method based on monitoring data and historical load patterns includes:
[0276] Real-time collect monitoring data (such as CPU utilization, memory usage, network bandwidth, etc.) from cloud platforms, edge devices, and local terminals;
[0277] Perform operations such as cleaning, missing value filling, and outlier detection on the collected monitoring data to obtain the first monitoring data;
[0278] Extract appropriate feature variables (such as the average value, peak value, etc. statistically calculated by time window) from the first monitoring data according to business characteristics and monitoring metrics;
[0279] Normalize and standardize the feature variables to ensure the consistency of the dimensions between features;
[0280] Select machine learning algorithms suitable for time series prediction (such as time series analysis, neural networks, random forests, etc.) as the model to be trained;
[0281] Obtain the historical monitoring data and historical load patterns of the cloud platform, edge devices, and terminals;
[0282] Divide the historical monitoring data into a training set and a validation set, train and cross-validate the model to be trained, and obtain the first model;
[0283] Evaluate the performance metrics of the first model (such as MSE, RMSE, etc.) and iteratively optimize the first model;
[0284] Use the trained first model, input the historical load pattern and feature variables, and predict the changing trends of resource requirements within a certain period in the future for different resource types (such as CPU, memory, network bandwidth, etc.) respectively.
[0285] In this embodiment, the historical load pattern refers to the regular pattern of resource usage analyzed over a past period of time, specifically including the following aspects: periodic change pattern (analyze the periodic change rules of resource usage on different time scales such as a day, a week, a month, etc. For example, some business systems may have obvious peak and valley differences between weekdays and weekends), seasonal change pattern (analyze the change rules of resource usage between different seasons in a year. Some industries may show obvious seasonal changes due to factors such as peak product sales seasons), event-triggered pattern (analyze the impact of some special events on the resource usage pattern, such as new product launches, marketing activities, etc. that will trigger a sudden increase in short-term resource requirements), abnormal fluctuation pattern (analyze the occasional abnormal fluctuations in resource usage and their causes, such as sudden business peaks or resource exhaustion caused by system failures), and so on. By analyzing these historical load patterns, the internal rules of resource usage can be better understood, providing a basis for subsequent demand prediction modeling; taking these historical load characteristics as input features and combining with real-time monitoring data, and using a machine learning model for prediction can more accurately predict the changing trends of future resource requirements, which is of great significance for realizing dynamic scaling and optimizing resource utilization.
[0286] In this embodiment, the machine learning method can deeply mine complex patterns in monitoring data, improving the accuracy and reliability of predictions; automate the entire process of resource demand prediction and scaling decisions, significantly reducing the manual operation and maintenance costs; promptly sense and respond to changes in resource demands, enhancing the overall elasticity and reliability of the IT system; perform refined resource scheduling and scaling based on accurate demand prediction results, improving the utilization efficiency of resources; predict resource demand changes in advance and allocate resources in a timely manner to ensure the stability of business service quality indicators. In short, this machine learning-based dynamic resource demand prediction method is one of the key technologies for realizing intelligent and adaptive resource management, which can significantly enhance the agility and operation efficiency of the IT system.
[0287] It should be noted that for the foregoing method embodiments, for the sake of simple description, they are all expressed as a series of action combinations. However, those skilled in the art should know that this application is not limited by the described action sequence, because according to this application, certain steps can be performed in other sequences or simultaneously. Secondly, those skilled in the art should also know that the embodiments described in the specification are all preferred embodiments, and the actions and modules involved are not necessarily essential to this application.
[0288] In the above embodiments, the descriptions of the various embodiments have their own emphases. For parts not detailed in a certain embodiment, reference can be made to the relevant descriptions of other embodiments.
[0289] In the several embodiments provided by this application, it should be understood that the disclosed device can be implemented in other ways. For example, the device embodiments described above are merely illustrative. For example, the above division of units is only a logical function division. In actual implementation, there can be other division methods. For example, multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the displayed or discussed coupling or direct coupling or communication connection to each other can be through some interfaces. The indirect coupling or communication connection of the device or unit can be in an electrical or other form.
[0290] The units described as separate components above may or may not be physically separated. The components shown as units may or may not be physical units, that is, they can be located in one place, or distributed to multiple network units. Some or all of the units can be selected according to actual needs to achieve the purpose of the solution of this embodiment.
[0291] In addition, in each embodiment of the present application, the functional units can be integrated into one processing unit, or each unit can exist physically alone, or two or more units can be integrated into one unit. The above-mentioned integrated unit can be implemented in the form of hardware or in the form of a software functional unit.
[0292] If the above-mentioned integrated unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable memory. Based on this understanding, the technical solution of the present application, in essence, or the part that contributes to the prior art, or all or part of this technical solution, can be embodied in the form of a software product. This computer software product is stored in a memory and includes several instructions to enable a computer device (which can be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the above-mentioned methods in each embodiment of the present application. The aforementioned memory includes various media that can store program codes, such as USB flash drives, read-only memories (ROM, Read-Only Memory), random access memories (RAM, Random Access Memory), mobile hard disks, magnetic disks, or optical discs.
[0293] Those of ordinary skill in the art can understand that all or part of the steps in the various methods of the above embodiments can be completed by instructing relevant hardware through a program. This program can be stored in a computer-readable memory. The memory can include: flash drives, read-only memories (abbreviation: ROM), random access memories (abbreviation: RAM), magnetic disks, or optical discs, etc.
[0294] The above has introduced the embodiments of the present application in detail. Specific examples are used in this article to elaborate on the principle and implementation manner of the present application. The description of the above embodiments is only used to help understand the method and its core idea of the present application; at the same time, for those of ordinary skill in the art, according to the idea of the present application, there will be changes in the specific implementation manner and application scope. In summary, the content of this specification should not be construed as a limitation to the present application.
[0295] Although the present invention is disclosed as above, the present invention is not limited thereto. Any person skilled in the art can easily think of changes or substitutions without departing from the spirit and scope of the present invention. All kinds of modifications and changes can be made, including the combination of the above different functions and implementation steps, including the implementation manners of software and hardware, and all are within the protection scope of the present invention.
Claims
1. A multi-source computing power data integration and intelligent scheduling system, characterized in that, Including: A central control server, a cloud platform, edge devices, and local terminals; The central control server is configured to: Establish a unified resource abstraction layer to standardize the description of multi-source heterogeneous computing resources and perform virtualization management, including: identifying all types of computing resources in the system; classifying the computing resources to establish a resource type hierarchy; designing a general computing resource description model; defining the key attributes of each type of computing resource; developing automated tools to collect detailed information about each specific computing resource; mapping the collected information into the computing resource description model; designing a virtualization management interface to abstract the underlying hardware details; implementing a resource pooling mechanism to convert physical resources into manageable virtual resource units; designing a unified API for resource allocation, monitoring, and recycling; generating a cross-platform resource management protocol; deploying distributed monitoring agents to collect resource usage in real time; designing a central monitoring panel to display the overall resource status; building a dynamic resource discovery and registration mechanism; building a role-based access control system; designing an encrypted communication channel to protect the transmission of resource information; building a caching mechanism to accelerate the query of frequently accessed resource information; optimizing the storage and retrieval algorithms for computing resource descriptions; designing a modular architecture to support the access of new computing resource types; building a distributed resource management framework to support large-scale clusters; Build a dynamic data integration module for real-time collection, cleaning, and standardization of multi-source heterogeneous data; Design an adaptive intelligent scheduling engine, integrating machine learning algorithms, for task allocation and resource scheduling based on historical task execution data and real-time system status data; Build an edge-cloud-terminal collaborative computing framework that can support the dynamic migration and load balancing of computing tasks; Establish a global performance optimization model to continuously optimize the overall system performance and resource utilization through reinforcement learning techniques, specifically including: determining the key metrics describing the overall system performance; establishing a global performance model covering edge devices, communication networks, local terminals, and cloud platforms to describe the relationships and influencing factors among the metrics at each level; defining the system state space; designing a reinforcement learning algorithm suitable for the global performance optimization problem; determining the reward function with the overall system performance metric as the goal; establishing a real-time monitoring mechanism for the edge-cloud to collect system operation state data; designing data preprocessing and feature extraction methods to provide input data for reinforcement learning; using the preprocessed and feature-extracted operation state data to train the reinforcement learning model to learn the optimal resource scheduling and task migration strategies; deploying the trained model into the system to execute performance optimization decisions in real time; continuously monitoring the system operation state and iteratively optimizing the reinforcement learning model; integrating the reinforcement learning optimization module into the edge-cloud-terminal collaborative framework to ensure the coordinated operation of each component; Integrate secure multi-party computing and differential privacy technologies to ensure privacy protection and security of data during integration and processing; Determine an elastic scaling mechanism to support the dynamic adjustment of the system scale and the seamless access of new computing resources.
2. A multi-source computing power data integration and intelligent scheduling method, characterized in that, Including: Establish a unified resource abstraction layer to standardize the description of multi-source heterogeneous computing resources and perform virtualization management; Build a dynamic data integration module for real-time collection, cleaning, and standardization processing of multi-source heterogeneous data; Design an adaptive intelligent scheduling engine, integrate machine learning algorithms, and perform task allocation and resource scheduling based on historical task execution data and real-time system status data; Build an edge-cloud-terminal collaborative computing framework that supports dynamic migration and load balancing of computing tasks; Establish a global performance optimization model, and continuously optimize the overall system performance and resource utilization through reinforcement learning technology. Specifically, it includes: determining key metrics that describe the overall system performance; establishing a global performance model covering edge devices, communication networks, local terminals, and cloud platforms to describe the relationships and influencing factors among the metrics at each level; defining the system state space; designing a reinforcement learning algorithm suitable for the global performance optimization problem; determining the reward function with the overall system performance metric as the goal; establishing a real-time monitoring mechanism for the edge-cloud to collect system operation status data; designing data preprocessing and feature extraction methods to provide input data for reinforcement learning; using the preprocessed and feature-extracted operation status data to train the reinforcement learning model and learn the optimal resource scheduling and task migration strategies; deploying the trained model into the system to perform real-time performance optimization decisions; continuously monitoring the system operation status and iteratively optimizing the reinforcement learning model; integrating the reinforcement learning optimization module into the edge-cloud-terminal collaborative framework to ensure the coordinated operation of each component; Integrate secure multi-party computing and differential privacy technologies to ensure privacy protection and security of data during integration and processing; Determine an elastic scaling mechanism to support dynamic adjustment of the system scale and seamless access to new computing resources; Among them, the steps of establishing a unified resource abstraction layer to standardize the description of multi-source heterogeneous computing resources and perform virtualization management include: Identify all types of computing resources in the system; Classify the computing resources and establish a resource type hierarchy; Design a general computing resource description model; Define the key attributes of each type of computing resource; Develop an automated tool to collect detailed information about each specific computing resource; Map the collected information into the computing resource description model; Design a virtualization management interface to abstract the underlying hardware details; Implement a resource pooling mechanism to convert physical resources into manageable virtual resource units; Design a unified API for resource allocation, monitoring, and recycling; Generate a cross-platform resource management protocol; Deploy a distributed monitoring agent to collect resource usage in real time; Design a central monitoring panel to display the overall resource status; Build a dynamic resource discovery and registration mechanism; Build a role-based access control system; Design an encrypted communication channel to protect the transmission of resource information; Build a caching mechanism to accelerate the query of frequently accessed resource information; Optimize the storage and retrieval algorithms for computing resource descriptions; Design a modular architecture to support the access of new computing resource types; Build a distributed resource management framework to support large-scale clusters.
3. The multi-source computing power data integration and intelligent scheduling method according to claim 2, wherein The steps of building a dynamic data integration module for real-time collection, cleaning, and standardization processing of multi-source heterogeneous data include: Identify all data source types, classify the data sources, and establish a data source directory; Define a common data structure and design a metadata model to describe the attributes, sources, and timestamps of the data; Develop dedicated data collectors for each data source type and build a data access layer that supports multiple protocols; Build a stream processing framework and design and implement a streaming data processing pipeline; Design a data format conversion function, develop a data validation and error handling mechanism, and design a data cleaning algorithm; Perform data field mapping and conversion, develop a unit conversion and standardization function library, and design a data semantic conversion mechanism; Build a real-time data quality detection algorithm, design a data quality scoring system, and develop an abnormal data handling mechanism; Design a data fusion algorithm to handle the conflicts and complementarities of multi-source data, and develop a data correlation analysis function and a data consistency check mechanism; Build a data caching mechanism and design a parallel processing architecture, and optimize the data storage and indexing strategies; Adopt a microservices architecture, determine dynamic configuration management, and design a distributed processing framework; Develop a data flow monitoring dashboard to achieve visual display of the data processing process, and design an alarm and notification mechanism.
4. The multi-source computing power data integration and intelligent scheduling method according to claim 3, wherein The steps of designing an adaptive intelligent scheduling engine, integrating machine learning algorithms, and performing task allocation and resource scheduling based on historical task execution data and real-time system status data include: Design a data collection module according to the data collector, and use the data collection module to obtain historical task execution data and real-time system status data; Perform a data cleaning and feature extraction process on the historical task execution data and real-time system status data, and build a data storage system to support fast query and analysis; Define a task feature set according to the features extracted from the historical task execution data, develop a task classification algorithm, identify the type of the task, and build a task requirement prediction model; Design a resource status feature set in combination with the features extracted from the real-time system status data, and develop a resource performance prediction model and a resource availability evaluation algorithm; Select a machine learning algorithm, and design a model training process and a model evaluation / verification mechanism; According to the task requirement prediction model, the resource performance prediction model, and the resource availability evaluation algorithm, develop a resource scheduling decision model based on the machine learning model, determine a multi-objective optimization mechanism, and design a scheduling strategy scoring system; Build a real-time inference module to generate scheduling decisions; Determine the decision execution and feedback mechanism; Develop an anomaly detection and handling process.
5. The multi-source computing power data integration and intelligent scheduling method according to claim 4, wherein, The steps of building an edge-cloud-terminal collaborative computing framework that can support the dynamic migration and load balancing of computing tasks include: Develop lightweight edge devices for performing preliminary data processing and analysis at the location of the data source; Design a communication protocol between edge devices to support distributed collaboration; Develop dynamic resource management and task scheduling algorithms to achieve load balancing between edge devices; Establish an elastic scalable cloud platform for processing large-scale data and complex computing tasks; Develop cloud platform resource management and task scheduling algorithms based on machine learning to achieve dynamic migration and load balancing across edge devices; Provide API interfaces to support two-way communication and collaboration between edge devices and the cloud platform; Formulate a collaboration mechanism between edge devices and cloud platforms, including task allocation strategies, data transmission protocols, and security authentication mechanisms; Develop cross-layer resource monitoring and scheduling algorithms to achieve dynamic load balancing across the edge-cloud-terminal; Design a fault tolerance mechanism to ensure the continuous and reliable operation of computing tasks.
6. The multi-source computing power data integration and intelligent scheduling method according to claim 5, wherein The steps of integrating secure multi-party computation and differential privacy technologies to ensure privacy protection and security during data integration and processing include: Clarify the privacy requirements and compliance requirements of all parties in the system; Design a comprehensive privacy protection framework based on secure multi-party computation and differential privacy, covering all links of data collection, transmission, storage, and processing; Select the corresponding secure multi-party computation algorithm; Design secure communication and computing protocols among multiple parties to ensure that each party can only obtain the final result and cannot peek into the intermediate process; Implement the secure multi-party computation module and integrate it with other components of the system; Identify privacy-sensitive data in the system and determine the corresponding differential privacy protection goals; Design a differential privacy algorithm based on noise injection to ensure data availability while minimizing privacy leakage; Integrate the differential privacy module into the data processing flow to ensure privacy protection of data in all links; Establish a privacy risk assessment mechanism to regularly evaluate the privacy protection effect of the system; Design a real-time monitoring module to detect possible privacy leakage events and trigger response measures; Continuously optimize the privacy protection plan based on the evaluation results and monitoring feedback; Design targeted privacy protection test scenarios covering different attack models and data processing scenarios; Verify the effectiveness and reliability of secure multi-party computation and differential privacy technologies through testing; Optimize and improve according to the test results.
7. The multi-source computing power data integration and intelligent scheduling method according to claim 6, characterized in that The steps of determining an elastic scaling mechanism to support dynamic adjustment of system scale and seamless access to new computing resources include: Analyze the business requirements and service characteristics of the system to determine the goals and key metrics of elastic scaling; Design the architecture and mechanism of elastic scaling according to business requirements and technical characteristics; Establish a system resource monitoring module to collect CPU, memory, and network metric data in real time; Based on the monitoring data and historical load patterns, use machine learning methods to predict the future trend of resource demand changes; According to the prediction results of the resource demand change trend, formulate a dynamic scaling decision-making strategy; Implement an automated scaling execution mechanism; Design an access mechanism for new computing resources to support seamless docking of heterogeneous hardware devices and cloud services; Implement resource discovery, scheduling, and orchestration functions to ensure that new resources can be effectively utilized by the system; Establish an optimization algorithm based on resource utilization and service metrics to allocate and schedule computing resources; Adopt appropriate resource scheduling strategies according to resource characteristics and application requirements; Design redundant backup and self-healing mechanisms to ensure the stable operation of the system in case of resource failures or fluctuations; Design functions for fault detection, isolation, and automatic recovery to improve the reliability of the system.
8. The multi-source computing power data integration and intelligent scheduling method according to claim 7, wherein The steps of using machine learning methods to predict the future trend of resource demand changes based on monitoring data and historical load patterns include: Collect monitoring data from the cloud platform, edge devices, and local terminals in real time; Perform operations such as cleaning, missing value filling, and outlier detection on the collected monitoring data to obtain the first monitoring data; Extract appropriate feature variables from the first monitoring data according to business characteristics and monitoring indicators; Perform normalization and standardization processing on the feature variables to ensure the same dimension between features; Select a machine learning algorithm suitable for time series prediction as the model to be trained; Obtain the historical monitoring data and historical load patterns of the cloud platform, edge devices, and terminals; Divide the historical monitoring data into a training set and a validation set, train and cross-validate the model to be trained to obtain the first model; Evaluate the performance indicators of the first model and perform iterative optimization on the first model; Use the trained first model, input the historical load pattern and feature variables, and predict the resource demand change trend in a certain future period for different resource types respectively.
Citation Information
Patent Citations
Data computing resource configuration method considering privacy protection under edge cloud collaboration framework
CN116614502A
Heterogeneous computing resource integration method and device, electronic equipment, storage medium and computer program
CN118519774A
Method and system for improving computing power efficiency
CN118550711A