An edge-cloud collaborative real-time data processing and computing scheduling method and system
Patent Information
- Application Number
- CN202610969238.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2026-07-01
- Publication Date
- 2026-09-29
AI Technical Summary
[0003]针对上述所显示出来的问题,本发明提供了一种边缘-云端协同的实时数据处理与计算调度方法及系统用以解决背景技术中提到的静态、固定的分配策略无法根据边缘端与云端实时变化的资源负载进行动态调整导致任务响应延迟或者加剧任务排队,延长整体处理时间,降低了整体的实用性和稳定性的问题
Smart Images

Figure CN122845587A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of data processing technology, and in particular to a real-time data processing and computing scheduling method and system for edge-cloud collaboration. Background Technology
[0002] Currently, with the deepening of Industry 4.0 and intelligent manufacturing, a large number of sensors, controllers, and intelligent devices have been deployed in dam industrial sites, leading to an explosive growth in data volume. This data contains immense value for driving production optimization, predictive equipment maintenance, and quality control, but it also poses unprecedented challenges to data processing technologies. Currently, dam industrial data processing mainly faces the following core contradictions: limited computing power at the edge makes it difficult to handle large-scale data computation; while the cloud has sufficient computing power, high data transmission latency fails to meet real-time requirements. To address these contradictions, an architecture combining edge computing and cloud computing has emerged. Existing technologies have seen some simple edge-cloud task allocation schemes, such as statically dividing tasks based on their preset types: fixing tasks with high real-time requirements to be executed at the edge, and fixing computationally intensive tasks to be executed in the cloud. However, this static and fixed allocation strategy has significant drawbacks: it lacks flexibility and dynamic adaptability. It cannot dynamically adjust according to the real-time changes in resource load at the edge and cloud (such as the instantaneous load of CPU and memory at the edge, and the backlog of task queues in the cloud). When the computing power at the edge suddenly becomes overloaded, the newly arrived real-time tasks will still be forcibly allocated to the edge, which may lead to task execution failure or response delay. Conversely, when the cloud load is too high, large tasks will still be distributed to it, which will exacerbate task queuing, prolong the overall processing time, and reduce the overall usability and stability. Summary of the Invention
[0003] To address the aforementioned problems, this invention provides a real-time data processing and computation scheduling method and system for edge-cloud collaboration. This solves the problem mentioned in the background art that static and fixed allocation strategies cannot be dynamically adjusted according to the real-time changes in resource load at the edge and cloud, leading to task response delays or increased task queuing, extending overall processing time, and reducing overall practicality and stability.
[0004] A real-time data processing and computing scheduling method with edge-cloud collaboration includes the following steps: Real-time acquisition of data processing tasks generated in industrial scenarios, and extraction of metadata for each data processing task; Dynamically monitor the real-time computing load, available memory, and network status of edge computing nodes, as well as the resource utilization and task queue status of the cloud distributed computing cluster, and determine the real-time resource status of edge computing nodes and cloud distributed computing cluster respectively. Based on the metadata of each data processing task, the real-time resource status of edge computing nodes and cloud distributed computing clusters, a preset intelligent collaborative scheduling algorithm dynamically allocates suitable computing execution nodes to each data processing task. The edge computing nodes and the cloud-based distributed computing cluster each execute their assigned tasks, and data and model parameters are exchanged between the two. In this process, the edge computing nodes will process and generate key summary data or abnormal event data and synchronize them to the cloud distributed computing cluster. The cloud distributed computing cluster will then distribute the decision logic generated based on global data analysis to the edge computing nodes.
[0005] Preferably, the real-time acquisition of data processing tasks generated in the industrial scenario and the extraction of metadata for each data processing task includes: Through a data interface layer deployed in the industrial field, task trigger signals from at least one data source are monitored and received in real time, including production equipment controllers, manufacturing execution systems, or enterprise resource planning systems. The task trigger signal is parsed to obtain the raw byte stream. The raw byte stream is unpacked, decoded and structured according to the preset data pattern to obtain the start and end boundary description parameters of each data processing task. The matching metadata description block task type is determined based on the task start and end boundary description parameters. The task types include: real-time equipment status monitoring, dynamic adjustment of production parameters, quality anomaly detection, raw data cleaning and transformation, and local emergency command generation. The metadata to be collected metrics are determined based on the metadata description block task type, and the metadata for each data processing task is extracted based on the metadata to be collected metrics.
[0006] Preferably, the step of parsing the task trigger signal to obtain the raw byte stream, unpacking, decoding, and structuring the raw byte stream according to a preset data pattern, and obtaining the start and end boundary description parameters for each data processing task includes: The task trigger signal is parsed to determine the identifier of the target industrial data source and to obtain the raw byte stream. Match and call the corresponding data pattern configuration file from the protocol template library based on the identifier of the target industrial data source; Load the data pattern corresponding to the target data source according to the data pattern configuration file, and determine the communication protocol specification, data packet structure and task semantic rule definition parameters according to the data pattern; According to the communication protocol specifications, data packet structure, and task semantic rules, the raw byte stream is sequentially unpacked and decoded to obtain structured task information data. Based on the boundary rules contained in the structured task information data and data patterns, the descriptive parameters of the start and end boundaries of each data processing task are identified and extracted.
[0007] Preferably, the dynamic monitoring of the real-time computing load, available memory, network status of the edge computing nodes, and the resource utilization and task queue status of the cloud-based distributed computing cluster, and the determination of the real-time resource status of the edge computing nodes and the cloud-based distributed computing cluster respectively, includes: Deploy lightweight monitoring components on edge computing nodes and periodically collect the first type of resource status data of the edge computing nodes. The first type of resource status data includes: CPU utilization, available memory capacity, local disk I / O throughput, and network latency and bandwidth fluctuation data for communication with the cloud. Deploy a resource monitoring service component on the management node of the cloud-based distributed computing cluster and periodically collect the second type of resource status data of the cloud-based distributed computing cluster. The second type of resource status data includes: the overall CPU and memory utilization of the cluster, the number of pending tasks and the average waiting time of each computing node's task queue, and the available capacity and I / O performance of the cluster's storage system. The collected first-class and second-class resource status data are standardized and normalized respectively to obtain edge resource status vectors and cloud resource status vectors with unified measurement standards. Based on the resource status vectors at the edge and in the cloud, the real-time resource status levels of the edge computing nodes and the distributed computing cluster in the cloud are determined respectively.
[0008] Preferably, the step of dynamically allocating suitable computing execution nodes to each data processing task based on the metadata of each data processing task, the real-time resource status of edge computing nodes and cloud distributed computing clusters, and through a preset intelligent collaborative scheduling algorithm includes: The metadata of each data processing task, the real-time resource status of edge computing nodes and cloud distributed computing clusters are input into the preset intelligent collaborative scheduling algorithm model. The algorithm calculates the first fit score between each data processing task and the edge computing node, as well as the second fit score with the cloud-based distributed computing cluster, using a pre-set intelligent collaborative scheduling algorithm model. Compare the first fitness score with the second fitness score, assign each data processing task to the computing node with the higher fitness score, and output a scheduling decision based on the assignment result; Based on the scheduling decision, corresponding task scheduling instructions are generated and each data processing task is sent to the assigned computing execution node.
[0009] Preferably, the calculation factors for calculating the first fit score between each data processing task and the edge computing node include: the matching degree between the real-time requirements of each data processing task and the ability of the edge computing node to process real-time tasks, the matching degree between the data volume of each data processing task and the current available memory of the edge computing node, and the current network status of the edge computing node. The factors used to calculate the second fit score between each data processing task and the cloud-based distributed computing cluster include: the matching degree between the data volume of each data processing task and the throughput capacity of the cloud-based distributed computing cluster, and the current waiting time of the task queue of the cloud-based distributed computing cluster. The intelligent collaborative scheduling algorithm model is predefined by the following rules: If the real-time requirement of each data processing task is extremely high and the resource status of the edge computing node is not overloaded, then the decision is to allocate it to the edge computing node. If the amount of data for each data processing task exceeds a preset threshold, or if each data processing task is a non-real-time processing task, the decision is to allocate it to a cloud-based distributed computing cluster. If the resource status level of the edge computing node is overloaded while the cloud is normal or idle, then one or more tasks that should have been assigned to the edge computing node but have a non-extremely high real-time requirement will be assigned to the cloud distributed computing cluster instead.
[0010] Preferably, the step of executing assigned tasks through edge computing nodes and cloud-based distributed computing clusters, and interacting with each other for data and model parameters, includes: The edge computing node receives and executes the assigned first type of task, which is a task with high real-time requirements. The execution process includes: performing real-time calculations based on local data and local decision models, and generating first type of result data, which includes equipment status warnings, production parameter adjustment instructions, or standardized data that has been cleaned and transformed. The second type of task is received and executed by the cloud-based distributed computing cluster. The second type of task is a data-intensive or computationally intensive non-real-time task. The execution process includes: aggregating historical and real-time data from multiple edge computing nodes, running a global analysis model, and generating a second type of result data. The second type of result data includes equipment health assessment, production trend prediction, or optimized global decision-making plan. Extract key data subsets from the first type of result data and feature data generated during the execution process that are valuable for global cloud analysis, and upload them synchronously to the cloud distributed computing cluster. Based on the optimized global decision-making plan in the second type of result data, key model parameters are extracted and distributed to edge computing nodes to update or replace their local decision-making models.
[0011] Preferably, the step of extracting key model parameters based on the optimized global decision-making plan in the second type of result data, and distributing the key model parameters to edge computing nodes to update or replace their local decision-making models, includes: Key model parameters are extracted from the optimized global decision-making plan in the second type of result data. The key model parameters include: model weights, bias terms, decision tree structure, or some layer parameters in the neural network. An optimized global decision model is generated based on key model parameters. The optimized global decision model is then compared with the local decision model of the edge computing node. The set of difference parameters between the old and new models is determined based on the comparison results. The set of differences in parameters is encapsulated into an update package and sent to the edge computing nodes; The update package is merged with the local decision model through edge computing nodes to complete the model update.
[0012] Preferably, the method further includes: Set up a data caching queue on the edge computing nodes to temporarily store data to be uploaded to the distributed computing cluster when the network is interrupted or of poor quality. Determine the priority and timeliness of each data processing task, and sort and manage the data to be transmitted in the data cache queue according to the priority and timeliness, prioritizing the transmission of high-priority data; When the network connection is detected to be restored, the breakpoint resume mechanism and data verification mechanism are automatically triggered to continue transmitting data from the breakpoint, and data integrity verification is performed after the transmission is completed.
[0013] An edge-cloud collaborative real-time data processing and computing scheduling system, the system comprising: The extraction module is used to acquire data processing tasks generated in industrial scenarios in real time and extract the metadata of each data processing task. The determination module is used to dynamically monitor the real-time computing load, available memory, and network status of edge computing nodes, as well as the resource utilization and task queue status of the cloud distributed computing cluster, and to determine the real-time resource status of the edge computing nodes and the cloud distributed computing cluster respectively. The allocation module is used to dynamically allocate suitable computing execution nodes to each data processing task based on the metadata of each data processing task, the real-time resource status of edge computing nodes and cloud distributed computing clusters, and through a preset intelligent collaborative scheduling algorithm. The interaction module is used to execute the assigned tasks by the edge computing nodes and the cloud distributed computing cluster respectively, and to exchange data and model parameters between the two. In this process, the edge computing nodes will process and generate key summary data or abnormal event data and synchronize them to the cloud distributed computing cluster. The cloud distributed computing cluster will then distribute the decision logic generated based on global data analysis to the edge computing nodes.
[0014] Other features and advantages of the invention will be set forth in the following description, and will be apparent in part from the description, or may be learned by practicing the invention. The objects and other advantages of the invention may be realized and obtained by means of the structures particularly pointed out in the written description and the accompanying drawings.
[0015] The technical solution of the present invention will be further described in detail below with reference to the accompanying drawings and embodiments. Attached Figure Description
[0016] The accompanying drawings are provided to further illustrate the invention and form part of the specification. They are used together with the embodiments of the invention to explain the invention and do not constitute a limitation thereof.
[0017] Figure 1 The flowchart illustrates the real-time data processing and computation scheduling method for edge-cloud collaboration provided by this invention. Figure 2 Another flowchart of the real-time data processing and computing scheduling method for edge-cloud collaboration provided by the present invention; Figure 3 This is another flowchart of a real-time data processing and computing scheduling method for edge-cloud collaboration provided by the present invention; Figure 4 This is a schematic diagram of the structure of a real-time data processing and computing scheduling system with edge-cloud collaboration provided by the present invention. Detailed Implementation
[0018] Exemplary embodiments will now be described in detail, examples of which are illustrated in the accompanying drawings. When the following description relates to the drawings, unless otherwise indicated, the same numerals in different drawings denote the same or similar elements. The embodiments described in the following exemplary embodiments do not represent all embodiments consistent with this disclosure. Rather, they are merely examples of apparatuses and methods consistent with some aspects of this disclosure as detailed in the appended claims.
[0019] A real-time data processing and computation scheduling method for edge-cloud collaboration, such as... Figure 1 As shown, it includes the following steps: Step S101: Acquire data processing tasks generated in the industrial scenario in real time and extract metadata for each data processing task; Step S102: Dynamically monitor the real-time computing load, available memory, network status of edge computing nodes, as well as the resource utilization and task queue status of the cloud distributed computing cluster, and determine the real-time resource status of edge computing nodes and cloud distributed computing cluster respectively. Step S103: Based on the metadata of each data processing task, the real-time resource status of edge computing nodes and cloud distributed computing clusters, dynamically allocate suitable computing execution nodes for each data processing task through a preset intelligent collaborative scheduling algorithm. Step S104: The edge computing nodes and the cloud distributed computing cluster execute their respective assigned tasks, and data and model parameters are exchanged between the two. In this process, the edge computing nodes will process and generate key summary data or abnormal event data and synchronize them to the cloud distributed computing cluster. The cloud distributed computing cluster will then distribute the decision logic generated based on global data analysis to the edge computing nodes.
[0020] In this embodiment, metadata refers to data that describes the attributes of a data processing task, including but not limited to task type, data source identifier, data format, data size, timestamp and frequency of data generation, maximum end-to-end latency required by the task, and target endpoint of task output. In this embodiment, the preset data pattern is represented as a predefined description specification of communication protocols and data structures for a specific industrial data source. It includes communication protocol specifications, data packet structures (such as frame headers, data payloads, checksums, frame trailers, etc.), and task semantic rules (such as data meaning, units, value ranges, etc.). The data pattern is stored in a protocol template library in the form of a configuration file and is dynamically loaded according to the data source identifier.
[0021] The working principle of the above technical solution is as follows: real-time acquisition of data processing tasks generated in industrial scenarios, and extraction of metadata for each data processing task; dynamic monitoring of the real-time computing load, available memory, network status of edge computing nodes, and resource utilization and task queue status of the cloud distributed computing cluster, and determination of the real-time resource status of the edge computing nodes and the cloud distributed computing cluster respectively; based on the metadata of each data processing task and the real-time resource status of the edge computing nodes and the cloud distributed computing cluster, dynamic allocation of suitable computing execution nodes for each data processing task through a preset intelligent collaborative scheduling algorithm; execution of the assigned tasks by the edge computing nodes and the cloud distributed computing cluster respectively, and interaction of data and model parameters between the two.
[0022] The beneficial effects of the above technical solution are as follows: global optimization of computing resources is achieved through intelligent scheduling, which effectively solves the contradiction between limited computing power at the edge and excessive latency in the cloud in industrial scenarios. It achieves a balance between real-time and comprehensive data processing, which can meet the high real-time requirements of equipment early warning and parameter adjustment at the edge, and can also complete the in-depth mining and trend prediction of large-scale data in the cloud. It solves the problem mentioned in the existing technology that the static and fixed allocation strategy cannot be dynamically adjusted according to the real-time changes in resource load at the edge and the cloud, which leads to task response delay or exacerbates task queuing, prolongs the overall processing time, and reduces the overall practicality and stability.
[0023] In one embodiment, such as Figure 2 As shown, the real-time acquisition of data processing tasks generated in industrial scenarios and the extraction of metadata for each data processing task include: Step S201: Through the data interface layer deployed in the industrial site, listen in real time and receive task trigger signals from at least one data source, including production equipment controller, manufacturing execution system or enterprise resource planning system; Step S202: Parse the task trigger signal to obtain the raw byte stream, unpack, decode and structure the raw byte stream according to the preset data mode, and obtain the start and end boundary description parameters of each data processing task; Step S203: Determine the matching metadata description block task type based on the task start and end boundary description parameters. The task type includes: real-time equipment status monitoring, dynamic adjustment of production parameters, quality anomaly detection, raw data cleaning and transformation, and local emergency command generation. Step S204: Determine the metadata indicators to be collected based on the metadata description block task type, and extract the metadata for each data processing task based on the metadata indicators to be collected.
[0024] In this embodiment, the metadata description block is represented as a data packet header containing attributes such as task priority, data size, and latency sensitivity.
[0025] The beneficial effects of the above technical solution are as follows: it realizes the automated and standardized perception and description of multi-source heterogeneous data tasks in industrial sites, providing accurate decision-making basis for subsequent intelligent scheduling. Furthermore, by parsing the original byte stream through preset data patterns, it can adapt to various industrial private or standard protocols, enhancing universality and scalability.
[0026] In one embodiment, parsing the task trigger signal to obtain the raw byte stream, unpacking, decoding, and structuring the raw byte stream according to a preset data pattern, and obtaining the start and end boundary description parameters for each data processing task include: The task trigger signal is parsed to determine the identifier of the target industrial data source and to obtain the raw byte stream. Match and call the corresponding data pattern configuration file from the protocol template library based on the identifier of the target industrial data source; Load the data pattern corresponding to the target data source according to the data pattern configuration file, and determine the communication protocol specification, data packet structure and task semantic rule definition parameters according to the data pattern; According to the communication protocol specifications, data packet structure, and task semantic rules, the raw byte stream is sequentially unpacked and decoded to obtain structured task information data. Based on the boundary rules contained in the structured task information data and data patterns, the descriptive parameters of the start and end boundaries of each data processing task are identified and extracted.
[0027] In this embodiment, the unpacking operation includes: identifying the frame header, data payload, checksum, and frame trailer from the original byte stream according to the data packet structure defined in the data mode; The start of the task data packet is determined based on the frame header, and the end of the task data packet is determined based on the frame tail or the data packet length field. In this embodiment, the decoding operation includes: parsing the unpacked data payload according to the encoding rules defined in the data mode. The encoding rules include, but are not limited to: big-endian / little-endian byte order conversion, data type parsing, and conversion of the custom encoding mapping table; In this embodiment, the protocol template library is represented as a library that stores various industrial communication protocols and data pattern configuration files. Each configuration file corresponds to one or a type of data source and contains all the parameters and rules required for packet decryption and decoding. By matching the data source identifier, the system can call the corresponding configuration file to parse the data stream. In this embodiment, boundary rules refer to rules used to identify the start and end of data processing task data packets. These include delimitation rules based on specific start or end character or byte sequences, dynamic delimitation rules based on the data packet length field, and time rules based on silent time intervals.
[0028] The beneficial effects of the above technical solution are as follows: through a configurable protocol template library, it is possible to quickly adapt to the communication protocols of different data sources, reducing the complexity and cost of data integration. Furthermore, it can accurately identify the start and end boundaries of task data, ensuring the integrity and correctness of data packets, and providing a reliable data foundation for subsequent processing.
[0029] In one embodiment, the dynamic monitoring of the real-time computing load, available memory, and network status of edge computing nodes, as well as the resource utilization and task queue status of the cloud-based distributed computing cluster, and the determination of the real-time resource status of the edge computing nodes and the cloud-based distributed computing cluster respectively, includes: Deploy lightweight monitoring components on edge computing nodes and periodically collect the first type of resource status data of the edge computing nodes. The first type of resource status data includes: CPU utilization, available memory capacity, local disk I / O throughput, and network latency and bandwidth fluctuation data for communication with the cloud. Deploy a resource monitoring service component on the management node of the cloud-based distributed computing cluster and periodically collect the second type of resource status data of the cloud-based distributed computing cluster. The second type of resource status data includes: the overall CPU and memory utilization of the cluster, the number of pending tasks and the average waiting time of each computing node's task queue, and the available capacity and I / O performance of the cluster's storage system. The collected first-class and second-class resource status data are standardized and normalized respectively to obtain edge resource status vectors and cloud resource status vectors with unified measurement standards. Based on the resource status vectors at the edge and in the cloud, the real-time resource status levels of the edge computing nodes and the distributed computing cluster in the cloud are determined respectively.
[0030] In this embodiment, the lightweight monitoring component refers to resource monitoring software or modules deployed on edge computing nodes. It collects data such as CPU utilization, memory capacity, and disk I / O by calling the underlying interface of the operating system, and obtains network latency and bandwidth data through network probing tools. In this embodiment, the resource status vector is represented as a multi-dimensional vector, where each dimension represents a resource status indicator (such as CPU utilization, available memory, etc.). Through standardization (converting the data to a mean of 0 and a variance of 1) and normalization (scaling the data to the [0,1] interval), the original data with different dimensions and ranges are converted into comparable values.
[0031] In this embodiment, the standardization and normalization process includes mapping the collected raw data to a numerical range of 0 to 1; In this embodiment, the real-time resource status level is determined by inputting the resource status vector into a preset resource status evaluation model. The values of each dimension of the resource status vector are compared with the corresponding reference thresholds to comprehensively evaluate the discrete status level, which represents idle, normal, busy, or overloaded. The resource status evaluation model adopts a weighted scoring algorithm, which assigns different weights to different dimensions of the resource status vector, calculates the comprehensive load score, and determines the final resource status level based on the load score range.
[0032] The beneficial effects of the above technical solution are as follows: it realizes unified and quantitative monitoring of two types of heterogeneous computing resources, namely edge and cloud, so that scheduling can make decisions based on accurate and comparable status information. Furthermore, it refines the original monitoring data into intuitive resource status levels, which greatly simplifies the decision-making logic of the scheduling algorithm and improves scheduling efficiency.
[0033] In one embodiment, such as Figure 3 As shown, the method of dynamically allocating suitable computing execution nodes for each data processing task based on the metadata of each data processing task, the real-time resource status of edge computing nodes and cloud distributed computing clusters, and through a preset intelligent collaborative scheduling algorithm includes: Step S301: Input the metadata of each data processing task, the real-time resource status of the edge computing node and the cloud distributed computing cluster into the preset intelligent collaborative scheduling algorithm model; Step S302: Calculate the first fit score between each data processing task and the edge computing node, and the second fit score between each task and the cloud distributed computing cluster using a preset intelligent collaborative scheduling algorithm model. Step S303: Compare the first fitness score with the second fitness score, assign each data processing task to the computing node with the higher fitness score, and output the scheduling decision based on the allocation result; Step S304: Generate corresponding task scheduling instructions based on the scheduling decision and send each data processing task to the allocated computing execution node.
[0034] In this embodiment, the intelligent collaborative scheduling algorithm model is represented as a decision system based on rules and / or machine learning models. It receives task metadata and resource state vectors as input, calculates the adaptability score between the task and the edge and cloud through predefined rules (such as real-time priority rules and load balancing rules) and / or machine learning models (such as classifiers and regression models), and finally outputs a scheduling decision. In this embodiment, the fit score is represented as a numerical value, indicating the degree of matching between the task and the computing node. Computational factors include the matching degree between the task's real-time requirements and the node's ability to process real-time tasks (e.g., the matching degree between real-time requirements and the node's current latency), the matching degree between data volume and the node's currently available memory, and the node's network status. The total score is obtained by weighted summation or other combinations of these factors.
[0035] The beneficial effects of the above technical solution are as follows: by making scheduling decisions through quantitative adaptation scores, the task allocation process becomes more objective and accurate, avoiding the problem of uneven resource utilization that may be caused by single rule scheduling. It can flexibly guide tasks to the most suitable computing nodes according to the real-time changes in task requirements and resource status, thereby improving the adaptability and practicality of task allocation.
[0036] In one embodiment, the calculation factors for calculating the first fit score between each data processing task and the edge computing node include: the matching degree between the real-time requirements of each data processing task and the ability of the edge computing node to process real-time tasks, the matching degree between the data volume of each data processing task and the current available memory of the edge computing node, and the current network status of the edge computing node. The factors used to calculate the second fit score between each data processing task and the cloud-based distributed computing cluster include: the matching degree between the data volume of each data processing task and the throughput capacity of the cloud-based distributed computing cluster, and the current waiting time of the task queue of the cloud-based distributed computing cluster. The intelligent collaborative scheduling algorithm model is predefined by the following rules: If the real-time requirement of each data processing task is extremely high and the resource status of the edge computing node is not overloaded, then the decision is to allocate it to the edge computing node. If the amount of data for each data processing task exceeds a preset threshold, or if each data processing task is a non-real-time processing task, the decision is to allocate it to a cloud-based distributed computing cluster. If the resource status level of the edge computing node is overloaded while the cloud is normal or idle, then one or more tasks that should have been assigned to the edge computing node but have a non-extremely high real-time requirement will be assigned to the cloud distributed computing cluster instead.
[0037] The beneficial effects of the above technical solution are as follows: by assigning the highest weight to real-time requirements, it ensures that high real-time tasks are always prioritized, meeting the stringent requirements of core industrial applications for immediate response. Furthermore, by combining preset scheduling rules, it not only ensures scheduling determinism in key scenarios, but also handles complex scheduling scenarios that require trade-offs through a scoring model, thereby enhancing the robustness and intelligence of the system.
[0038] In one embodiment, the step of executing assigned tasks through edge computing nodes and a cloud-based distributed computing cluster, and interacting with each other for data and model parameters, includes: The edge computing node receives and executes the assigned first type of task, which is a task with high real-time requirements. The execution process includes: performing real-time calculations based on local data and local decision models, and generating first type of result data, which includes equipment status warnings, production parameter adjustment instructions, or standardized data that has been cleaned and transformed. The second type of task is received and executed by the cloud-based distributed computing cluster. The second type of task is a data-intensive or computationally intensive non-real-time task. The execution process includes: aggregating historical and real-time data from multiple edge computing nodes, running a global analysis model, and generating a second type of result data. The second type of result data includes equipment health assessment, production trend prediction, or optimized global decision-making plan. Extract key data subsets from the first type of result data and feature data that are valuable for global cloud analysis generated during the execution process, and upload them synchronously to the cloud distributed computing cluster. Based on the optimized global decision-making plan in the second type of result data, key model parameters are extracted and distributed to edge computing nodes to update or replace their local decision-making models.
[0039] In this embodiment, the key data subset refers to the data that is of great value to the global analysis in the cloud and is selected from the processing results at the edge. The selection criteria include the importance level of the event represented by the data, the sample data of a specified type in the cloud, or the data that triggers the cloud deep analysis process. The key data subset is identified by preset rules or machine learning models. In this embodiment, the global decision-making plan is represented as an optimized decision-making scheme generated by the cloud based on global data analysis, including model parameters (such as neural network weights and decision tree structure) and business rules (such as threshold adjustment and strategy update).
[0040] The beneficial effects of the above technical solution are as follows: it clarifies the division of responsibilities and collaboration between the edge and the cloud, forms an efficient collaborative paradigm of real-time edge response and macro-optimization in the cloud, and establishes a two-way intelligent interaction closed loop between the edge and the cloud, enabling the edge to make more intelligent decisions locally, while the cloud can continuously optimize the entire data processing process based on global information.
[0041] In one embodiment, the step of extracting key model parameters based on the optimized global decision-making plan from the second type of result data, and distributing the key model parameters to edge computing nodes to update or replace their local decision-making models, includes: Key model parameters are extracted from the optimized global decision-making plan in the second type of result data. The key model parameters include: model weights, bias terms, decision tree structure, or some layer parameters in the neural network. An optimized global decision model is generated based on key model parameters. The optimized global decision model is then compared with the local decision model of the edge computing node. The set of difference parameters between the old and new models is determined based on the comparison results. The set of differences in parameters is encapsulated into an update package and sent to the edge computing nodes; The update package is merged with the local decision model through edge computing nodes to complete the model update.
[0042] The beneficial effects of the above technical solution are: by using the differential update mechanism, only the model difference parameters are transmitted, rather than the entire model file, which greatly reduces the network transmission load and improves the data transmission efficiency.
[0043] In one embodiment, the method further includes: Set up a data caching queue on the edge computing nodes to temporarily store data to be uploaded to the distributed computing cluster when the network is interrupted or of poor quality. Determine the priority and timeliness of each data processing task, and sort and manage the data to be transmitted in the data cache queue according to the priority and timeliness, prioritizing the transmission of high-priority data; When the network connection is detected to be restored, the breakpoint resume mechanism and data verification mechanism are automatically triggered to continue transmitting data from the breakpoint, and data integrity verification is performed after the transmission is completed.
[0044] The beneficial effects of the above technical solution are as follows: through breakpoint resume, data verification and priority management mechanisms, the robustness and reliability of data transmission in unreliable industrial network environments are significantly improved, ensuring that critical data can be transmitted in a priority and complete manner after network interruption, preventing the loss of important production data and ensuring the continuity of system services.
[0045] In this embodiment, a mechanism for edge computing nodes to make degradation decisions in the event of a communication interruption is also included, specifically: A lightweight knowledge graph is pre-stored locally on the edge computing node. The lightweight knowledge graph is generated by the cloud-based distributed computing cluster based on global data and periodically distributed. It contains entities, relationships and attributes related to the jurisdiction of the edge computing node. When a communication interruption is detected between the edge computing node and the cloud-based distributed computing cluster, the edge computing node executes the following fallback decision process: Receive real-time data streams from the dam industrial site, parse the real-time data streams, and extract multiple data feature units; Determine the data source identifier and feature type for each data feature unit; Using the data source identifier and feature type as the joint query key, entity matching queries are performed in a pre-stored lightweight knowledge graph; When the query is successful, the data feature unit is bound to the corresponding target entity in the lightweight knowledge graph; Based on the relationships of the target entity in the lightweight knowledge graph, path traversal and reasoning are performed to generate one or more downgrade decision candidates. Starting with the target entity, a multi-step path traversal is performed along predefined semantic relationships in the lightweight knowledge graph to discover all reachable, decision-related terminal entities, including fault mode entities or coping strategy entities. For each valid path from the target entity to the terminal entity, extract the sequence of entities and relationships traversed along the path to form a candidate inference chain; Based on a predefined rule base and confidence model, an initial weight score is calculated for each candidate inference chain; All candidate inference chains are sorted and filtered based on the initial weight scores, and the operations contained in the terminal entities or paths corresponding to the top N inference chains are used as multiple downgrade candidate decisions. Obtain the initial confidence score for each downgrade candidate decision and input the initial confidence score into the preset decision filtering output engine; The preset decision filtering output engine filters and prioritizes downgraded candidate decisions based on multi-dimensional filtering rules, and outputs a sorted list of downgraded candidate decisions that meet the requirements. Select the highest-ranked target downgrade candidate decision from the downgrade candidate decision ranking list as the final decision, execute the final decision, and record the decision process and results locally.
[0046] In this embodiment, the lightweight knowledge graph is represented as an optimized industrial knowledge subgraph, whose data size is specifically pruned to meet the following characteristics: the number of entities is limited to less than 5,000, it supports completing a 3-hop path traversal starting from any entity within 100ms, and it only includes entities directly related to the devices managed by the edge nodes and their associated entities. In this embodiment, the data feature unit is a standardized data semantic block, which includes at least the following fields: unique identifier of the data source, feature type code (e.g., temperature, pressure, vibration frequency), feature value and its unit, data quality identifier, etc. In this embodiment, the predefined semantic relationships include at least the following three types: causal relationship: representing a causal logical chain between entities; component: representing the inclusion relationship between the whole and its parts; symptom manifestation: representing the correspondence between faults and symptoms; handling strategy: representing the mapping relationship between problems and countermeasures, etc. In this embodiment, the multi-dimensional screening rules include: security dimension: excluding decisions that may cause equipment damage or safety accidents; resource dimension: the computing resources and execution time required for the decision must be within the range of currently available resources on the edge node; timeliness dimension: the estimated execution time of the decision must be less than the time threshold for problem deterioration; priority dimension: weighted adjustment based on the preset priority of different decisions. The beneficial effects of the above technical solution are as follows: By using a local lightweight knowledge graph and inference engine, even when the dam loses contact with the cloud, the edge nodes can still make intelligent decisions based on local knowledge, continuously monitor and regulate the dam's operating status, avoid business shutdowns due to communication interruptions, enable the edge nodes to have autonomous decision-making capabilities, and proactively take a series of degraded but optimized measures to maintain the dam's critical functions, greatly improving the system's resilience and survivability. Furthermore, through multi-dimensional screening rules and confidence quantification assessment, the safest, most feasible, and most effective final decision in the current environment can be quickly and automatically selected from multiple possible decisions, avoiding blindness and arbitrariness in decision-making and ensuring the implementation of decisions.
[0047] In one embodiment, this embodiment also discloses an edge-cloud collaborative real-time data processing and computing scheduling system, such as... Figure 4 As shown, the system includes: The extraction module 401 is used to acquire data processing tasks generated in industrial scenarios in real time and extract metadata for each data processing task. The determination module 402 is used to dynamically monitor the real-time computing load, available memory, network status of edge computing nodes, as well as the resource utilization and task queue status of the cloud distributed computing cluster, and to determine the real-time resource status of the edge computing nodes and the cloud distributed computing cluster respectively. The allocation module 403 is used to dynamically allocate suitable computing execution nodes to each data processing task based on the metadata of each data processing task, the real-time resource status of edge computing nodes and cloud distributed computing clusters, and through a preset intelligent collaborative scheduling algorithm. The interaction module 404 is used to execute the assigned tasks by the edge computing nodes and the cloud distributed computing cluster respectively, and to exchange data and model parameters between the two. In this process, the edge computing nodes will process and generate key summary data or abnormal event data and synchronize them to the cloud distributed computing cluster. The cloud distributed computing cluster will then distribute the decision logic generated based on global data analysis to the edge computing nodes.
[0048] The working principle and beneficial effects of the above technical solution have been explained in the method embodiments, and will not be repeated here.
[0049] Other embodiments of this disclosure will readily occur to those skilled in the art upon consideration of the specification and practice of the disclosure herein. This application is intended to cover any variations, uses, or adaptations of this disclosure that follow the general principles of this disclosure and include common knowledge or customary techniques in the art not disclosed herein. The specification and examples are to be considered exemplary only, and the true scope and spirit of this disclosure are indicated by the following claims.
[0050] It should be understood that this disclosure is not limited to the precise structures described above and shown in the accompanying drawings, and various modifications and changes can be made without departing from its scope. The scope of this disclosure is limited only by the appended claims.
Claims
1. A real-time data processing and computation scheduling method for edge-cloud collaboration, characterized in that, Includes the following steps: Real-time acquisition of data processing tasks generated in industrial scenarios, and extraction of metadata for each data processing task; Dynamically monitor the real-time computing load, available memory, and network status of edge computing nodes, as well as the resource utilization and task queue status of the cloud distributed computing cluster, and determine the real-time resource status of edge computing nodes and cloud distributed computing cluster respectively. Based on the metadata of each data processing task, the real-time resource status of edge computing nodes and cloud distributed computing clusters, a preset intelligent collaborative scheduling algorithm dynamically allocates suitable computing execution nodes to each data processing task. The edge computing nodes and the cloud-based distributed computing cluster each execute their assigned tasks, and data and model parameters are exchanged between the two. In this process, the edge computing nodes will process and generate key summary data or abnormal event data and synchronize them to the cloud distributed computing cluster. The cloud distributed computing cluster will then distribute the decision logic generated based on global data analysis to the edge computing nodes.
2. The edge-cloud collaborative real-time data processing and computation scheduling method according to claim 1, characterized in that, The real-time acquisition of data processing tasks generated in industrial scenarios and the extraction of metadata for each data processing task include: Through a data interface layer deployed in the industrial field, task trigger signals from at least one data source are monitored and received in real time, including production equipment controllers, manufacturing execution systems, or enterprise resource planning systems. The task trigger signal is parsed to obtain the raw byte stream. The raw byte stream is unpacked, decoded and structured according to the preset data pattern to obtain the start and end boundary description parameters of each data processing task. The matching metadata description block task type is determined based on the task start and end boundary description parameters. The task types include: real-time equipment status monitoring, dynamic adjustment of production parameters, quality anomaly detection, raw data cleaning and transformation, and local emergency command generation. The metadata to be collected metrics are determined based on the metadata description block task type, and the metadata for each data processing task is extracted based on the metadata to be collected metrics.
3. The edge-cloud collaborative real-time data processing and computation scheduling method according to claim 2, characterized in that, The process involves parsing the task trigger signal to obtain the raw byte stream, unpacking, decoding, and structuring the raw byte stream according to a preset data pattern, and obtaining start and end boundary description parameters for each data processing task, including: The task trigger signal is parsed to determine the identifier of the target industrial data source and to obtain the raw byte stream. Match and call the corresponding data pattern configuration file from the protocol template library based on the identifier of the target industrial data source; Load the data pattern corresponding to the target data source according to the data pattern configuration file, and determine the communication protocol specification, data packet structure and task semantic rule definition parameters according to the data pattern; According to the communication protocol specifications, data packet structure, and task semantic rules, the raw byte stream is sequentially unpacked and decoded to obtain structured task information data. Based on the boundary rules contained in the structured task information data and data patterns, the descriptive parameters of the start and end boundaries of each data processing task are identified and extracted.
4. The edge-cloud collaborative real-time data processing and computation scheduling method according to claim 1, characterized in that, The dynamic monitoring of real-time computing load, available memory, network status of edge computing nodes, and resource utilization and task queue status of the cloud distributed computing cluster, and the determination of real-time resource status of edge computing nodes and cloud distributed computing cluster respectively, includes: Deploy lightweight monitoring components on edge computing nodes and periodically collect the first type of resource status data of the edge computing nodes. The first type of resource status data includes: CPU utilization, available memory capacity, local disk I / O throughput, and network latency and bandwidth fluctuation data for communication with the cloud. Deploy a resource monitoring service component on the management node of the cloud-based distributed computing cluster and periodically collect the second type of resource status data of the cloud-based distributed computing cluster. The second type of resource status data includes: the overall CPU and memory utilization of the cluster, the number of pending tasks and the average waiting time of each computing node's task queue, and the available capacity and I / O performance of the cluster's storage system. The collected first-class and second-class resource status data are standardized and normalized respectively to obtain edge resource status vectors and cloud resource status vectors with unified measurement standards. Based on the resource status vectors at the edge and in the cloud, the real-time resource status levels of the edge computing nodes and the distributed computing cluster in the cloud are determined respectively.
5. The edge-cloud collaborative real-time data processing and computation scheduling method according to claim 1, characterized in that, The method, based on the metadata of each data processing task, the real-time resource status of edge computing nodes and cloud-based distributed computing clusters, dynamically allocates suitable computing execution nodes to each data processing task through a preset intelligent collaborative scheduling algorithm, including: The metadata of each data processing task, the real-time resource status of edge computing nodes and cloud distributed computing clusters are input into the preset intelligent collaborative scheduling algorithm model. The algorithm calculates the first fit score between each data processing task and the edge computing node, as well as the second fit score with the cloud-based distributed computing cluster, using a pre-set intelligent collaborative scheduling algorithm model. Compare the first fitness score with the second fitness score, assign each data processing task to the computing node with the higher fitness score, and output a scheduling decision based on the assignment result; Based on the scheduling decision, corresponding task scheduling instructions are generated and each data processing task is sent to the assigned computing execution node.
6. The edge-cloud collaborative real-time data processing and computation scheduling method according to claim 5, characterized in that, The factors used to calculate the first fit score between each data processing task and the edge computing node include: the matching degree between the real-time requirements of each data processing task and the ability of the edge computing node to process real-time tasks, the matching degree between the data volume of each data processing task and the current available memory of the edge computing node, and the current network status of the edge computing node. The factors used to calculate the second fit score between each data processing task and the cloud-based distributed computing cluster include: the matching degree between the data volume of each data processing task and the throughput capacity of the cloud-based distributed computing cluster, and the current waiting time of the task queue of the cloud-based distributed computing cluster. The intelligent collaborative scheduling algorithm model is predefined by the following rules: If the real-time requirement of each data processing task is extremely high and the resource status of the edge computing node is not overloaded, then the decision is to allocate it to the edge computing node. If the amount of data for each data processing task exceeds a preset threshold, or if each data processing task is a non-real-time processing task, the decision is to allocate it to a cloud-based distributed computing cluster. If the resource status level of the edge computing node is overloaded while the cloud is normal or idle, then one or more tasks that should have been assigned to the edge computing node but have a non-extremely high real-time requirement will be assigned to the cloud distributed computing cluster instead.
7. The edge-cloud collaborative real-time data processing and computing scheduling method according to claim 1, characterized in that, The process of executing assigned tasks through edge computing nodes and cloud-based distributed computing clusters, and interacting with each other for data and model parameters, includes: The edge computing node receives and executes the assigned first type of task, which is a task with high real-time requirements. The execution process includes: performing real-time calculations based on local data and local decision models, and generating first type of result data, which includes equipment status warnings, production parameter adjustment instructions, or standardized data that has been cleaned and transformed. The second type of task is received and executed by the cloud-based distributed computing cluster. The second type of task is a data-intensive or computationally intensive non-real-time task. The execution process includes: aggregating historical and real-time data from multiple edge computing nodes, running a global analysis model, and generating a second type of result data. The second type of result data includes equipment health assessment, production trend prediction, or optimized global decision-making plan. Extract key data subsets from the first type of result data and feature data that are valuable for global cloud analysis generated during the execution process, and upload them synchronously to the cloud distributed computing cluster. Based on the optimized global decision-making plan in the second type of result data, key model parameters are extracted and distributed to edge computing nodes to update or replace their local decision-making models.
8. The edge-cloud collaborative real-time data processing and computation scheduling method according to claim 7, characterized in that, The step of extracting key model parameters based on the optimized global decision-making plan from the second type of result data, and distributing the key model parameters to edge computing nodes to update or replace their local decision-making models includes: Key model parameters are extracted from the optimized global decision-making plan in the second type of result data. The key model parameters include: model weights, bias terms, decision tree structure, or some layer parameters in the neural network. An optimized global decision model is generated based on key model parameters. The optimized global decision model is then compared with the local decision model of the edge computing node. The set of difference parameters between the old and new models is determined based on the comparison results. The set of differences in parameters is encapsulated into an update package and sent to the edge computing nodes; The update package is merged with the local decision model through edge computing nodes to complete the model update.
9. The edge-cloud collaborative real-time data processing and computing scheduling method according to claim 1, characterized in that, The method further includes: Set up a data caching queue on the edge computing nodes to temporarily store data to be uploaded to the distributed computing cluster when the network is interrupted or of poor quality. Determine the priority and timeliness of each data processing task, and sort and manage the data to be transmitted in the data cache queue according to the priority and timeliness, prioritizing the transmission of high-priority data; When the network connection is detected to be restored, the breakpoint resume mechanism and data verification mechanism are automatically triggered to continue transmitting data from the breakpoint, and data integrity verification is performed after the transmission is completed.
10. A real-time data processing and computing scheduling system with edge-cloud collaboration, characterized in that, The system includes: The extraction module is used to acquire data processing tasks generated in industrial scenarios in real time and extract the metadata of each data processing task. The determination module is used to dynamically monitor the real-time computing load, available memory, and network status of edge computing nodes, as well as the resource utilization and task queue status of the cloud distributed computing cluster, and to determine the real-time resource status of the edge computing nodes and the cloud distributed computing cluster respectively. The allocation module is used to dynamically allocate suitable computing execution nodes to each data processing task based on the metadata of each data processing task, the real-time resource status of edge computing nodes and cloud distributed computing clusters, and through a preset intelligent collaborative scheduling algorithm. The interaction module is used to execute the assigned tasks by the edge computing nodes and the cloud distributed computing cluster respectively, and to exchange data and model parameters between the two. In this process, the edge computing nodes will process and generate key summary data or abnormal event data and synchronize them to the cloud distributed computing cluster. The cloud distributed computing cluster will then distribute the decision logic generated based on global data analysis to the edge computing nodes.