Data processing and analysis method and system for edge computing gateway
Through the edge computing gateway, the task is analyzed and cached state recognition is optimized, and the task data processing and unloading strategies are solved, which is the problem of low data dependency processing efficiency and incomplete consistency checks in complex task scenarios, and the data processing efficiency and consistency guarantee capabilities are improved.
Patent Information
- Application Number
- CN202510379394.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-28
- Publication Date
- 2025-06-06
- Estimated Expiration
- 2045-03-28
AI Technical Summary
When existing edge computing gateways deal with complex task scenarios, insufficient task correlation analysis, low data dependency processing efficiency, inflexible distributed offload decisions, and incomplete data consistency inspection of result data.
Through the edge computing gateway, a correlation analysis of the computing tasks of multiple user devices is generated, and a task association graph is used to identify the cache status and data dependencies of the task in combination with the pre-stored cache matrix. Based on this information, task data is preprocessed, task unloading strategies are dynamically determined, and task execution status is tracked for consistency checks and in-depth analysis.
It improves the data processing efficiency and consistency guarantee capabilities of edge computing gateways, optimizes task scheduling and resource utilization, and ensures data processing quality in complex task scenarios.
Smart Images

Figure CN119883661B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of edge gateway technology, and in particular to a data processing and analysis method and system for an edge computing gateway. Background Art
[0002] With the rapid development of edge computing, edge computing gateways, as important nodes connecting user devices and cloud servers, undertake key functions such as computing task allocation, data processing and transmission. Its efficient data processing and analysis capabilities have played an important role in smart cities, industrial Internet of Things, and intelligent transportation. However, with the continuous increase in the number of user devices and the increase in the complexity of computing tasks, edge computing gateways face challenges such as a sharp increase in data volume, limited computing resources, and complex data dependencies between tasks. More efficient data processing and analysis methods are urgently needed.
[0003] Existing edge computing gateways mostly use static task allocation strategies or simple data preprocessing methods. This method can meet basic needs when the number of tasks is small or the data flow relationship is simple, but in complex scenarios, it lacks in-depth analysis of task relevance and data dependency, which can easily lead to inefficient task scheduling, data flow bottlenecks, and consistency problems in processing results. In addition, the task offloading strategy in distributed processing is often based on fixed rules, which is difficult to dynamically adjust according to real-time task requirements and network conditions, further limiting the system's computing efficiency and resource utilization. Summary of the invention
[0004] The main purpose of the present invention is to solve the technical problems of insufficient task correlation analysis, low data dependency processing efficiency, inflexible distributed offloading decision-making and imperfect result data consistency check when the existing edge computing gateway processes complex task scenarios;
[0005] A first aspect of the present invention provides a data processing and analysis method for an edge computing gateway, the data processing and analysis method for the edge computing gateway comprising:
[0006] The edge computing gateway performs correlation analysis on computing tasks of multiple user devices to obtain a task correlation graph, and identifies the cache status and data dependency of each computing task based on the task correlation graph and a pre-stored cache matrix to obtain a task data processing requirement list and an inter-task data flow graph;
[0007] According to the task data processing requirement list, preprocess the task data of the computing task to obtain preprocessed task data, and according to the preset offloading decision and the inter-task data flow diagram, transmit the preprocessed task data to the local processing unit of the edge computing gateway and the cooperating edge device for processing to obtain a distributed task data set;
[0008] Tracking the execution status of the distributed task data set to obtain a task status table, collecting and transmitting the processing results of the distributed task data set according to the task status table to obtain intermediate result data;
[0009] A consistency check is performed on the intermediate result data, and after the consistency check passes, a predefined analysis model is used to perform an in-depth analysis on the result data after the consistency check passes to obtain a comprehensive analysis result, and the comprehensive analysis result is distributed to relevant terminal devices.
[0010] Optionally, in a first implementation of the first aspect of the present invention, performing correlation analysis on computing tasks of multiple user devices by the edge computing gateway to obtain a task correlation graph includes:
[0011] Scanning the task data and expected output type of each computing task through the edge computing gateway to obtain task data features;
[0012] According to the task data features, a task similarity matrix between computing tasks is constructed, and a clustering algorithm is applied to the task similarity matrix to obtain task groups;
[0013] According to the task groups and task data features, a directed acyclic graph is constructed to obtain an initial task association graph, and the initial task association graph is topologically sorted and redundant edges are eliminated to obtain a task association graph.
[0014] Optionally, in a second implementation of the first aspect of the present invention, identifying the cache status and data dependency of each computing task according to the task association graph and the pre-stored cache matrix, and obtaining the task data processing requirement list and the inter-task data flow graph includes:
[0015] Traversing each node in the task association graph, searching for a corresponding cache item in a pre-stored cache matrix for the task ID of each node, and obtaining a task cache status list;
[0016] Marking each edge in the task association graph according to the task cache status list to obtain a task association graph with cache marks;
[0017] Traversing the task association graph with cache marks, extracting the predecessor task information and the follow-up task information of each computing task, and obtaining a task dependency table;
[0018] Calculate the data processing requirements of each computing task according to the task dependency table and the task cache status list to obtain a task data processing requirement list;
[0019] Based on the task association graph with cache marks and the task data processing requirement list, a data flow graph is constructed to obtain an inter-task data flow graph.
[0020] Optionally, in a third implementation of the first aspect of the present invention, according to the preset offloading decision and the inter-task data flow diagram, the pre-processed task data is transmitted to the local processing unit of the edge computing gateway and the cooperating edge device for processing, and the distributed task data set is obtained, including:
[0021] According to the preset offloading decision, the computing tasks are classified into local processing tasks and remote processing tasks, and a task classification list is obtained;
[0022] According to the task classification list and the data flow diagram between tasks, a data transmission instruction is generated for each computing task to obtain a task data transmission instruction set;
[0023] According to the task data transmission instruction set, the pre-processing task data of the local processing task is loaded into the local processing unit of the edge computing gateway to obtain a local task data subset;
[0024] According to the task data transmission instruction set, the pre-processed task data of the remote processing task is transmitted to the cooperating edge device through a secure channel to obtain a remote task data subset;
[0025] The local task data subset and the remote task data subset are merged, and task metadata is added to obtain a distributed task data set with labels.
[0026] Optionally, in a fourth implementation of the first aspect of the present invention, the generating of a data transmission instruction for each computing task by comparing the task classification list and the inter-task data flow diagram to obtain a task data transmission instruction set includes:
[0027] Traversing each computing task in the task classification list, extracting classification information and task identifiers of each computing task, and obtaining a task basic information table;
[0028] According to the task basic information table, locate the node of each computing task in the inter-task data flow graph, identify the predecessor task and the successor task of each computing task based on the node, and obtain the task dependency table;
[0029] Based on the task dependency table, calculating the data transmission priority of each computing task to obtain a task transmission priority list;
[0030] According to the task transmission priority list and task classification information, a target execution unit is assigned to each computing task to obtain a task distribution mapping table;
[0031] In combination with the task distribution mapping table and the inter-task data flow diagram, a data transmission instruction including a source address, a target address, a data size and a transmission time window is generated to form a task data transmission instruction set.
[0032] Optionally, in a fifth implementation of the first aspect of the present invention, tracking the execution status of the distributed task data set to obtain a task status table, collecting and transmitting the processing results of the distributed task data set according to the task status table, and obtaining the intermediate result data includes:
[0033] Starting a timer for each computing task in the distributed task data set, and periodically checking the task execution progress to obtain real-time task progress information;
[0034] According to the real-time task progress information, the task status table is updated to record the current status, start time and expected completion time of each computing task;
[0035] Monitor the task status table, identify completed tasks, and determine subsequent tasks that can be started based on the data flow diagram between tasks to obtain a task execution queue;
[0036] According to the task execution queue, processing results of completed tasks are collected from local processing units and cooperating edge devices to obtain an original result data set;
[0037] The original result data set is subjected to format unification and data normalization processing to obtain intermediate result data.
[0038] Optionally, in a sixth implementation of the first aspect of the present invention, the performing a consistency check on the intermediate result data, and after the consistency check passes, using a predefined analysis model to perform an in-depth analysis on the result data after the consistency check passes to obtain a comprehensive analysis result, and distributing the comprehensive analysis result to relevant terminal devices includes:
[0039] Applying a data integrity check algorithm to the intermediate result data to generate a check code, and comparing it with the expected check code to obtain a data consistency check result;
[0040] According to the data consistency check result, a data subset that passes the check is screened out to form verified result data;
[0041] Input the verified result data into a predefined analysis model, perform feature extraction, pattern recognition and anomaly detection, and obtain preliminary analysis results;
[0042] The preliminary analysis results are subjected to multi-dimensional cross-validation and result fusion to obtain comprehensive analysis results. According to a preset data distribution strategy, the comprehensive analysis results are encrypted and divided into multiple data packets, which are transmitted to relevant terminal devices through a secure channel.
[0043] A second aspect of the present invention provides a data processing and analysis system for an edge computing gateway, the data processing and analysis system for the edge computing gateway comprising:
[0044] A correlation analysis module is used to perform correlation analysis on computing tasks of multiple user devices through the edge computing gateway to obtain a task correlation graph, and identify the cache status and data dependency of each computing task based on the task correlation graph and a pre-stored cache matrix to obtain a task data processing requirement list and a data flow graph between tasks;
[0045] A task allocation module is used to pre-process the task data of the computing task according to the task data processing requirement list to obtain pre-processed task data, and transmit the pre-processed task data to the local processing unit of the edge computing gateway and the cooperating edge device for processing according to the preset offloading decision and the inter-task data flow diagram to obtain a distributed task data set;
[0046] A tracking module, used to track the execution status of the distributed task data set, obtain a task status table, collect and transmit the processing results of the distributed task data set according to the task status table, and obtain intermediate result data;
[0047] The inspection and distribution module is used to perform consistency check on the intermediate result data, and after the consistency check passes, use a predefined analysis model to perform in-depth analysis on the result data after the consistency check passes, obtain a comprehensive analysis result, and distribute the comprehensive analysis result to relevant terminal devices.
[0048] The data processing and analysis method and system of the above-mentioned edge computing gateway, by performing correlation analysis on the computing tasks of multiple user devices, generates a task correlation graph, combines the cache status and data dependency of the task with the cache matrix to identify, and generates a task data processing requirement list and an inter-task data flow graph. The task data is preprocessed according to the task data processing requirement list, and based on the preset offloading decision and the inter-task data flow graph, the preprocessed task data is transmitted to the local processing unit and the collaborative edge device to generate a distributed task data set. The execution status of the distributed task data set is tracked, the processing result data is collected and consistency checked, the result data is deeply analyzed using the analysis model, and a comprehensive analysis result is generated and distributed to the relevant terminal devices. The present invention improves the data processing efficiency and consistency assurance capability of the edge computing gateway through dynamic task scheduling and distributed collaborative processing.
[0049] Other features and advantages of the present invention will be described in the following description, and partly become apparent from the description, or understood by practicing the present invention. The purpose and other advantages of the present invention are realized and obtained by the structures particularly pointed out in the description, claims and drawings.
[0050] In order to make the above-mentioned objects, features and advantages of the present invention more obvious and easy to understand, preferred embodiments are given below and described in detail with reference to the accompanying drawings. BRIEF DESCRIPTION OF THE DRAWINGS
[0051] Figure 1 This is a schematic diagram of a first embodiment of a data processing and analysis method for an edge computing gateway in an embodiment of the present invention;
[0052] Figure 2 The present invention is a schematic diagram of an embodiment of a data processing and analysis system for an edge computing gateway. DETAILED DESCRIPTION
[0053] In order to make the purpose, technical solution and advantages of the embodiments of the present invention clearer, the technical solution of the present invention will be clearly and completely described below in conjunction with the accompanying drawings. Obviously, the described embodiments are part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of the present invention.
[0054] The terms "including" and "having" and any variations thereof mentioned in the embodiments of the present invention are intended to cover non-exclusive inclusions. For example, a process, method, system, product or device end including a series of steps or units is not limited to the listed steps or units, but may optionally include other steps or units that are not listed, or may optionally include other steps or units that are inherent to these processes, methods, products or device ends.
[0055] To facilitate understanding of this embodiment, a data processing and analysis method for an edge computing gateway disclosed in an embodiment of the present invention is first described in detail. Figure 1 As shown, the method comprises the following steps:
[0056] 101. Perform correlation analysis on computing tasks of multiple user devices through the edge computing gateway to obtain a task correlation graph, and identify the cache status and data dependency of each computing task based on the task correlation graph and the pre-stored cache matrix to obtain a task data processing requirement list and a data flow graph between tasks;
[0057] In one embodiment of the present invention, the edge computing gateway performs correlation analysis on computing tasks of multiple user devices to obtain a task association graph, including: scanning the task data and expected output type of each computing task by the edge computing gateway to obtain task data characteristics; constructing a task similarity matrix between computing tasks based on the task data characteristics, and applying a clustering algorithm to the task similarity matrix to obtain task groups; constructing a directed acyclic graph based on the task groups and task data characteristics to obtain an initial task association graph, and topologically sorting and eliminating redundant edges on the initial task association graph to obtain a task association graph.
[0058] Specifically, when the edge computing gateway performs correlation analysis on computing tasks of multiple user devices, it first scans the task data and expected output type of each computing task. This step involves using data analysis algorithms to extract the features of the task. Specifically, the edge computing gateway analyzes the structure, size, and type (such as text, image, audio, etc.) of the input data, as well as the expected output format. For task data, the gateway calculates statistical features (such as mean, variance, distribution), structural features (such as data dimension, sparsity), and semantic features (such as keywords, topics). The analysis of the expected output type includes determining whether the task is of classification, regression, or generation type. This information is integrated into a multidimensional vector, where each dimension represents a specific attribute of the task, such as data complexity, computational intensity, memory requirements, expected execution time, etc. This multidimensional vector is the task data feature, which comprehensively describes the characteristics of each computing task.
[0059] Specifically, the edge computing gateway constructs a task similarity matrix between computing tasks based on the characteristics of the task data. This process first selects an appropriate similarity measurement method, commonly used are Euclidean distance, cosine similarity, Jaccard coefficient, etc. The selection criteria depend on the nature of the task characteristics. For example, for high-dimensional feature vectors, cosine similarity is usually more suitable because it focuses on the direction of the vector rather than the absolute size. During calculation, the gateway calculates the similarity score for each pair of tasks to form an n×n symmetric matrix, where n is the total number of tasks and the elements on the diagonal are 1. After completing the similarity matrix, the gateway applies a clustering algorithm to group the tasks. Commonly used clustering algorithms include K-means, hierarchical clustering, DBSCAN, etc., and factors such as data distribution characteristics, expected number of clusters, and algorithm complexity are considered when selecting. The clustering process groups similar tasks into a group to form a task group. The tasks in each group have similar characteristics, and this grouping reflects the inherent relationship between tasks.
[0060] Specifically, the edge computing gateway first represents each task as a node in the graph, which contains information such as task ID, group, and main features. Then, directed edges are added based on the data dependencies between tasks. If the output of task A is the input of task B, add a directed edge between A and B. This process requires a detailed analysis of the data flow of each task to ensure that all necessary dependencies are captured. In the process of adding edges, the gateway needs to be careful to avoid forming loops, as loops may cause deadlocks or infinite loops. If a potential loop is detected, the gateway will resolve it by splitting or redesigning the task. The constructed initial task association graph contains all tasks and the relationships between them, but there may be unnecessary complexity.
[0061] Specifically, the edge computing gateway optimizes the initial task association graph. The first is topological sorting, which ensures that the tasks in the graph are arranged in the order of dependency. Topological sorting uses depth-first search or Kahn algorithm, starting from nodes without incoming edges and gradually processing all nodes until all nodes are visited. The sorting result reflects the execution order of the tasks. Next, redundant edge elimination is performed, which removes edges that can be implied by other paths. For example, if there is a path A→B→C, the edge directly connecting A→C may be redundant. Redundant edges are removed by comparing each edge with the existing path. If there is a path with a length greater than 1 connecting the same two nodes, then the edge directly connecting the two nodes is considered redundant. Through these optimization steps, a concise and effective task association graph is finally obtained, which accurately reflects the essential association between tasks.
[0062] Furthermore, the method of identifying the cache status and data dependency of each computing task according to the task association graph and the pre-stored cache matrix to obtain a task data processing requirement list and a data flow graph between tasks includes: traversing each node in the task association graph, searching for the corresponding cache item in the pre-stored cache matrix for the task ID of each node, and obtaining a task cache status list; marking each edge in the task association graph according to the task cache status list, and obtaining a task association graph with cache marks; traversing the task association graph with cache marks, extracting the predecessor task information and subsequent task information of each computing task, and obtaining a task dependency table; calculating the data processing requirements of each computing task according to the task dependency table and the task cache status list, and obtaining a task data processing requirement list; constructing a data flow graph based on the task association graph with cache marks and the task data processing requirement list, and obtaining a data flow graph between tasks.
[0063] Specifically, when traversing each node in the task association graph, the edge computing gateway first creates an empty task cache status list. For each node, the gateway extracts its task ID and then searches for the corresponding cache item in the pre-stored cache matrix. The cache matrix is a two-dimensional array in which rows represent tasks and columns represent different execution devices (such as local processing units or collaborative edge devices). Each element in the matrix represents the cache status of a specific task on a specific device, usually represented by a value between 0 and 1, where 0 represents uncached, 1 represents fully cached, and an intermediate value represents partially cached. The search process uses the task ID as a row index and traverses all columns to obtain the cache status of the task on all devices. This information is integrated into a cache status object, which contains the task ID and the cache ratio on each device. The gateway adds this object to the task cache status list. This process is repeated until all nodes in the task association graph are processed. The resulting task cache status list contains the complete cache information of each task, providing an important basis for subsequent data processing and scheduling decisions.
[0064] Specifically, according to the task cache status list, the edge computing gateway then marks each edge in the task association graph. This process starts from the starting node of the task association graph and traverses the entire graph along the directed edges. For each edge, the gateway checks the cache information of its starting node and the ending node in the task cache status list. If the task of the starting node has been fully or partially cached on the device where the ending node is located, the edge is marked as a "cached edge". The marking method can be to add a Boolean value or a numerical value to the attributes of the edge to indicate the cache status and ratio. If the two tasks connected by the edge are not cached or are cached on different devices, the edge is marked as a "non-cached edge". In addition, the gateway also calculates the data transfer volume of each edge, which is based on the output size and cache ratio of the starting task. For example, if the output data size of a task is 100MB and there is 50% cache on the target device, the actual data transfer volume of this edge is 50MB. This information is added to the attributes of the edge. Through this process, the original task association graph is converted into a task association graph with cache marking, which contains rich cache and data transfer information, providing a detailed basis for subsequent task scheduling and data flow analysis.
[0065] Specifically, traversing the task association graph with cache marks and extracting the predecessor task information and successor task information of each computing task is a key step in building the task dependency table. The edge computing gateway uses the depth-first search (DFS) or breadth-first search (BFS) algorithm to traverse each node in the graph. For each node, the gateway checks all its incoming and outgoing edges. The starting node of the incoming edge is identified as the predecessor task of the current task, and the ending node of the outgoing edge is identified as the successor task. The gateway creates a dependency object for each task, which contains the task ID, the predecessor task list, and the successor task list. In this process, the gateway also needs to consider the cache mark on the edge. For the case marked as "cache edge", the gateway adds a flag to the dependency to indicate that the dependency involves cache data. This is very important for subsequent task scheduling because it affects data transmission and processing strategies. For example, if all predecessor tasks of a task are connected through cache edges, then the task may be able to use cached data directly without waiting for the predecessor tasks to complete. The gateway also records the amount of data transmission for each dependency, which comes from the attributes of the edge. Finally, all this information is integrated into the task dependency table, where each entry contains the task ID, the list of predecessor tasks (with cache marks), the list of successor tasks (with cache marks), and the corresponding data transfer amount.
[0066] Specifically, according to the task dependency table and the task cache status list, the edge computing gateway calculates the data processing requirements of each computing task and generates a task data processing requirements list. This process first considers the input requirements of each task. The gateway checks all predecessor tasks of the task and calculates the total amount of data that needs to be transmitted. For cached data, only the uncached part needs to be considered. For example, if the output of a predecessor task is 100MB, but 50% is already cached, then only 50MB needs to be transmitted. The gateway also needs to consider the computational complexity of the task itself, which can be obtained from the attributes of the task or historical execution data. Computational complexity is usually expressed as the number of CPU cycles required or the estimated execution time. The gateway combines the input data volume and computational complexity to estimate the overall resource requirements of the task, including computing resources, memory requirements, and storage space. In addition, the gateway also needs to consider the output data size of the task, which affects the data transmission requirements of subsequent tasks. All this information is integrated into a task data processing requirements object, which contains the task ID, input data volume, computational complexity, resource requirements, and output data volume. The gateway creates such an object for each task, and finally forms a complete task data processing requirements list. This list provides detailed guidance for task scheduling and resource allocation, enabling the edge computing gateway to make more optimized decisions.
[0067] Specifically, based on the task association graph with cache marks and the task data processing requirements list, the edge computing gateway constructs a data flow graph to obtain the data flow graph between tasks. This process starts from the task association graph, but focuses on the flow of data rather than the dependencies between tasks. The gateway first creates a new graph structure in which nodes represent tasks and edges represent data flows. For each edge in the task association graph, the gateway checks its cache mark and the relevant information in the task data processing requirements list. If the edge is marked as a non-cache edge, the gateway adds a corresponding edge in the data flow graph and marks the expected data transfer volume. For the cache edge, the gateway calculates the actual amount of data that needs to be transferred, which is equal to the total amount of data minus the cached part. The gateway also considers the output data volume of each task, adds edges from the task to all its subsequent tasks in the data flow graph, and marks the expected output data volume. In addition, the gateway marks the data processing capacity of each task node in the graph, which comes from the task data processing requirements list. This information helps identify potential data processing bottlenecks. Finally, the gateway may need to add virtual source nodes and sink nodes to represent the initial input and final output of the data to complete the entire data flow. The data flow graph constructed in this way not only shows the data dependencies between tasks, but also contains detailed data volume information and processing capabilities, providing a comprehensive view for subsequent task scheduling and resource allocation.
[0068] 102. Preprocess the task data of the computing task according to the task data processing requirement list to obtain preprocessed task data, and transmit the preprocessed task data to the local processing unit of the edge computing gateway and the cooperating edge device for processing according to the preset offloading decision and the data flow diagram between tasks to obtain a distributed task data set;
[0069] In one embodiment of the present invention, according to the preset offloading decision and the inter-task data flow diagram, the pre-processed task data is transmitted to the local processing unit of the edge computing gateway and the cooperating edge device for processing respectively to obtain a distributed task data set, including: according to the preset offloading decision, the computing tasks are classified into local processing tasks and remote processing tasks to obtain a task classification list; according to the task classification list and the inter-task data flow diagram, a data transmission instruction is generated for each computing task to obtain a task data transmission instruction set; according to the task data transmission instruction set, the pre-processed task data of the local processing task is loaded into the local processing unit of the edge computing gateway to obtain a local task data subset; according to the task data transmission instruction set, the pre-processed task data of the remote processing task is transmitted to the cooperating edge device through a secure channel to obtain a remote task data subset; the local task data subset and the remote task data subset are merged, and task metadata is added to obtain a labeled distributed task data set.
[0070] Specifically, according to the task data processing requirements list, the task data of the computing task is preprocessed, and the process of obtaining the preprocessed task data involves multiple specific steps. First, the edge computing gateway traverses each task in the task data processing requirements list. For each task, the gateway determines the specific operations of preprocessing based on the information in its requirements list. These operations may include data format conversion, data compression, data cleaning, feature extraction, etc. Data format conversion ensures that the data meets the input requirements of the processing unit, such as converting unstructured data into structured data. Data compression is used to reduce transmission and storage requirements, especially for tasks that require remote processing. Data cleaning includes removing noise, processing missing values, correcting erroneous data, etc. to improve the accuracy of subsequent processing. Feature extraction is to extract key information from the original data to reduce the computational burden of the processing unit. The gateway selectively applies these preprocessing steps according to the specific requirements of the task. For example, for image processing tasks, preprocessing may include image scaling, color space conversion, histogram equalization, etc. For text analysis tasks, it may involve word segmentation, stop word removal, word frequency statistics, etc. The preprocessed data will be temporarily stored and associated with the original task ID to form a preprocessing task data set. This preprocessing process not only optimizes the data, but also prepares it for subsequent task distribution and execution.
[0071] Specifically, according to the predetermined offloading decision, the edge computing gateway classifies the computing tasks into local processing tasks and remote processing tasks, and obtains a task classification list. This process first requires the gateway to access the pre-defined offloading decision rules. These rules may be based on a variety of factors, such as the computational complexity of the task, the data size, the network status, the device capabilities, etc. The gateway checks the characteristics of each task one by one and matches it with the offloading decision rules. For example, if the computing requirements of a task exceed the capabilities of the local processing unit, or require specific hardware resources (such as GPU) that are only available on the remote device, then the task will be classified as a remote processing task. On the contrary, if the task has high real-time requirements or involves sensitive data that is not suitable for transmission, it will be classified as a local processing task. The gateway also needs to consider load balancing in this process to ensure that not all tasks are concentrated on one processing unit. The classification results will be recorded in a task classification list, each entry of which contains the task ID and its classification (local or remote). This list not only guides the subsequent data transmission process, but also provides an important basis for resource scheduling.
[0072] Specifically, next, the edge computing gateway generates data transfer instructions for each computing task by comparing the task classification list and the data flow diagram between tasks, and obtains the task data transfer instruction set. This process requires the gateway to comprehensively consider the classification of tasks (local or remote) and the data dependencies between tasks. The gateway first checks the data flow diagram between tasks to determine the source of input data and the destination of output data for each task. For tasks classified as local processing, the gateway generates instructions to load data into the local processing unit. These instructions include information such as the location of the data source, the target memory address, and the size of the data. For remote processing tasks, the gateway generates data transfer instructions that specify the data source, the target device address, the transmission protocol, the priority, etc. The gateway also needs to consider the dependencies between tasks to ensure that the data transfer order conforms to the task execution order. For example, if task B depends on the output of task A, then the output data transfer instruction of task A must be executed before the input data transfer instruction of task B. In addition, the gateway also needs to generate instructions for returning the results, especially for remote tasks. These instruction sets form a complete task data transfer instruction set, which provides a detailed operation guide for subsequent data movement.
[0073] Specifically, according to the task data transfer instruction set, the edge computing gateway loads the pre-processed task data of the local processing task into the local processing unit to obtain a local task data subset. This process first involves the gateway identifying all tasks marked for local processing. For each local task, the gateway looks for its corresponding pre-processed task data. The loading process usually involves transferring data from a storage device (such as a hard disk) to the memory of the processing unit. This process needs to consider memory management to ensure that there is enough space to accommodate all local task data. If there is insufficient memory, the gateway may need to implement data paging or task queue mechanisms. During the loading process, the gateway also needs to verify data integrity to ensure that no data is corrupted or lost. After the loading is completed, the gateway will update the task status and mark the data as ready for processing. All of these locally loaded task data form a local task data subset. This subset contains the input data of all tasks to be executed in the local processing unit, preparing for subsequent local computing.
[0074] Specifically, at the same time, the edge computing gateway transmits the pre-processed task data of the remote processing task to the collaborative edge device through a secure channel according to the task data transmission instruction set, and obtains the remote task data subset. This process first requires the gateway to establish a secure connection with the collaborative edge device. The secure channel is usually implemented through an encryption protocol (such as SSL / TLS) to ensure the confidentiality and integrity of data transmission. The gateway transmits the data of the remote task one by one according to the priority order in the transmission instruction set. During the transmission process, the gateway needs to monitor the network status, such as bandwidth usage, latency, etc., and may adjust the transmission strategy according to the real-time situation. For example, if network congestion is detected, the gateway may reduce the transmission priority of non-critical tasks. For large data sets, the gateway may adopt a block transmission strategy and reassemble it at the receiving end. After the transmission is completed, the gateway will request the receiver to confirm the data integrity, and if a transmission error is found, the retransmission mechanism will be started. All task data successfully transmitted to the remote device constitute the remote task data subset.
[0075] Specifically, at last, the edge computing gateway merges the local task data subset and the remote task data subset, and adds task metadata to obtain a tagged distributed task data set. The merging process is not a simple data splicing, but a logical structure is created to maintain the relationship between the two subsets. The gateway adds metadata tags to each task data, including task ID, processing location (local or remote), data size, expected processing time, etc. These metadata are essential for subsequent task scheduling and monitoring. The merged data set also contains dependency information between tasks, which comes from the task association graph built previously. The gateway also adds a globally unique identifier to track the entire distributed data set. This tagged distributed task data set is a comprehensive data structure that not only contains the input data of all tasks, but also includes all the contextual information required for task execution. It provides a unified view of the entire distributed computing process, enabling the edge computing gateway to effectively manage and coordinate the execution of tasks distributed on different processing units.
[0076] Furthermore, the task classification list and the data flow diagram between tasks are compared to generate data transmission instructions for each computing task, and the task data transmission instruction set is obtained, including: traversing each computing task in the task classification list, extracting the classification information and task identifier of each computing task, and obtaining a task basic information table; according to the task basic information table, locating the node of each computing task in the data flow diagram between tasks, identifying the predecessor task and the successor task of each computing task based on the node, and obtaining a task dependency table; based on the task dependency table, calculating the data transmission priority of each computing task, and obtaining a task transmission priority list; according to the task transmission priority list and the task classification information, assigning a target execution unit to each computing task, and obtaining a task distribution mapping table; combining the task distribution mapping table and the data flow diagram between tasks, generating a data transmission instruction including a source address, a target address, a data size and a transmission time window, and forming a task data transmission instruction set.
[0077] Specifically, the process of traversing each computing task in the task classification list, extracting the classification information and task identifier of each computing task, and obtaining the task basic information table involves systematically processing the pre-generated task classification list. The edge computing gateway first creates an empty data structure for storing the task basic information. Then, the gateway visits each entry in the task classification list in turn. For each entry, the gateway extracts two key information: the classification of the task (local processing or remote processing) and the unique task identifier (usually a number or string ID). This information is organized into a new data object containing fields such as "taskID" and "classificationType". The gateway may also add additional metadata such as the task creation timestamp or task priority (if included in the original classification). This newly created object is added to the task basic information table. The traversal process continues until all tasks are processed. The final generated task basic information table is a structured data set with each entry representing the basic attributes of a computing task. This table formats the original classification information to make it easier to process and query in subsequent steps. The task basic information table becomes an important reference for subsequent task scheduling and resource allocation decisions because it provides basic classification and identification information for each task, which is convenient for fast retrieval and processing.
[0078] Specifically, according to the task basic information table, the node of each computing task is located in the inter-task data flow graph, and the predecessor and successor tasks of each computing task are identified based on the nodes. The process of obtaining the task dependency table involves in-depth analysis of the data flow relationship between tasks. The edge computing gateway first loads the inter-task data flow graph, which is a directed graph structure that describes the data transmission relationship between tasks. The gateway traverses each task in the task basic information table and uses the task identifier to locate the corresponding node in the data flow graph. For each located node, the gateway performs two main operations: forward search and backward search. During the forward search process, the gateway identifies all edges pointing to the current node, and the starting nodes of these edges represent the predecessor tasks of the current task. During the backward search process, the gateway identifies all edges starting from the current node, and the ending nodes of these edges represent the successor tasks of the current task. The gateway records the identifiers of these predecessor and successor tasks to form a dependency list. In addition, the gateway also records the attributes of each dependency, such as the amount of data transmission, whether cached data is involved, etc. This information is organized into a new data structure, namely the task dependency table. Each entry contains fields such as "taskID", "predecessors" (predecessor task list), "successors" (successor task list) and related dependency properties. This process is repeated until all tasks are processed. The final generated task dependency table provides the position and relationship of each task in the entire computing process, which becomes the key input for subsequent task scheduling and data transmission planning.
[0079] Specifically, based on the task dependency table, the data transmission priority of each computing task is calculated, and the process of obtaining the task transmission priority list involves a complex priority evaluation algorithm. The edge computing gateway first defines a set of priority calculation rules, which consider multiple factors, such as the criticality of the task, the length of the dependency chain, the size of the data, etc. The gateway traverses each task in the task dependency table and first evaluates its position in the dependency graph. Tasks at the beginning of the dependency chain usually get higher priority because their execution is a prerequisite for the start of other tasks. The gateway also considers the number of subsequent tasks for each task. Tasks with more subsequent tasks may get higher priority because their delays affect more subsequent operations. The amount of data is also an important factor. The transmission of large amounts of data may need to start earlier to avoid becoming a bottleneck. The gateway combines these factors to calculate a numerical priority score for each task. The calculation process may involve weighted summation or more complex mathematical models. After the score calculation is completed, the gateway arranges the tasks in descending order of priority score to form a task transmission priority list. This list contains the identifier of each task and the corresponding priority score, and may also include a description of the key factors that determine the priority. The task transmission priority list provides clear guidance for subsequent data transmission scheduling, ensuring that data of critical tasks can be transmitted first, thereby optimizing the efficiency of the entire computing process.
[0080] Specifically, according to the task transmission priority list and task classification information, a target execution unit is assigned to each computing task. The process of obtaining the task distribution mapping table involves comprehensive consideration of task priority and execution location. The edge computing gateway first creates a new data structure to store the task distribution information. The gateway processes each task in the task transmission priority list in turn, while referring to the task classification information (local processing or remote processing). For tasks classified as local processing, the gateway assigns them to the local processing unit, which is usually the gateway's own computing resources. For remote processing tasks, the gateway needs to select the most suitable one from the available collaborative edge devices. The selection process considers multiple factors, such as the current load of the device, processing power, and matching degree with task requirements. The gateway may maintain a collaborative device resource status table to record the real-time status and capabilities of each device. Based on this information, the gateway selects the optimal target execution unit for each remote task. After the selection is completed, the gateway creates a mapping entry containing fields such as "taskID", "executionUnitID", "executionType" (local / remote), etc. This process is repeated until all tasks are assigned. The final generated task distribution mapping table records in detail which execution unit each task will run on, providing clear guidance for subsequent actual task distribution and execution. This mapping table takes into account both the priority of tasks, ensuring that high-priority tasks are processed in a timely manner, and the effective use of system resources, thus optimizing task allocation.
[0081] Specifically, combining the task distribution mapping table and the inter-task data flow graph, a data transmission instruction containing the source address, the target address, the data size and the transmission time window is generated. The process of forming a task data transmission instruction set is the last step of the data transmission plan. The edge computing gateway first creates a new data structure to store the transmission instructions. The gateway traverses each entry in the task distribution mapping table and determines the data dependency of each task by comparing it with the inter-task data flow graph. For each data dependency, the gateway generates a transmission instruction. The source address of the instruction is the location where the data is generated, which may be local storage or a collaborative device; the target address is the location where the task is executed, which comes from the task distribution mapping table. The data size information is extracted from the inter-task data flow graph, which determines the time and resources required for the transmission. The determination of the transmission time window needs to consider multiple factors: the priority of the task, the network bandwidth, the current load of the source and target devices, etc. The gateway may use a complex scheduling algorithm to optimize the overall transmission plan to ensure that the data of high-priority tasks can arrive in time when needed while avoiding network congestion. Each transmission instruction also includes other necessary information, such as transmission protocol, encryption requirements, etc. All these instructions collectively form the task data transmission instruction set, which specifies in detail all necessary data movements in the entire distributed computing process. This instruction set serves as a direct guide for data transmission execution, ensuring that the right data is transmitted to the right place at the right time, supporting the smooth execution of the entire distributed task.
[0082] 103. Track the execution status of the distributed task data set to obtain a task status table, collect and transmit the processing results of the distributed task data set according to the task status table, and obtain intermediate result data;
[0083] In one embodiment of the present invention, the tracking of the execution status of the distributed task data set to obtain a task status table, collecting and transmitting the processing results of the distributed task data set according to the task status table, and obtaining intermediate result data includes: starting a timer for each computing task in the distributed task data set, and periodically checking the task execution progress to obtain real-time task progress information; updating the task status table according to the real-time task progress information to record the current status, start time and expected completion time of each computing task; monitoring the task status table to identify completed tasks, and determining subsequent tasks that can be started according to the data flow diagram between tasks to obtain a task execution queue; according to the task execution queue, collecting the processing results of completed tasks from the local processing unit and the collaborating edge devices to obtain the original result data set; and performing format unification and data normalization processing on the original result data set to obtain intermediate result data.
[0084] Specifically, the process of starting a timer for each computing task in the distributed task data set and periodically checking the task execution progress to obtain real-time task progress information involves precise time management and task monitoring mechanisms. The edge computing gateway first creates a timer object for each task, which contains information such as the unique identifier of the task, the start time, and the estimated execution time. After the timer is started, the gateway sends a progress query request to the execution unit at a preset time interval (for example, every 100 milliseconds or every second). For local tasks, this query is directly performed on the local processing unit; for remote tasks, the gateway needs to send query instructions to the cooperating edge device through the network. After receiving the query, the execution unit returns the execution status of the current task, including information such as the percentage of completed workload, the amount of data processed, and the current processing stage. The gateway receives this information and compares it with the overall requirements of the task to calculate the accurate progress percentage. At the same time, the gateway also records the timestamp of each query to analyze the execution speed of the task and predict the remaining time. This information is integrated into a real-time task progress object, which contains fields such as "taskID", "progressPercentage", "currentStage", "estimatedTimeRemaining", etc. The gateway continuously updates these objects to form a dynamic real-time task progress information set. This set not only reflects the current status of each task, but can also be used to identify execution anomalies or performance bottlenecks, providing an important basis for task scheduling and resource allocation.
[0085] Specifically, the process of updating the task status table and recording the current status, start time, and estimated completion time of each computing task based on real-time task progress information involves continuous data updating and predictive analysis. The edge computing gateway maintains a dynamic task status table, which stores the latest status information of all tasks in a formatted manner. Whenever new real-time task progress information is received, the gateway triggers the update process of the status table. The update process first locates the entry of the corresponding task and then replaces the old data with the new progress information. The current status of the task may include categories such as "waiting", "executing", "completed", and "error". The start time is usually recorded when the task first enters the "executing" state. The calculation of the estimated completion time is relatively complex, and the gateway will dynamically predict it based on the current progress, the executed time, and the overall requirements of the task. The prediction algorithm may consider factors such as historical execution data and current system load, and use linear regression or more complex machine learning models to improve accuracy. If the predicted completion time deviates significantly from the initial estimate, the gateway may mark the task as requiring special attention. In addition, the gateway will also calculate and record other relevant indicators, such as the deviation between the actual execution time of the task and the expected time, resource utilization, etc. This updated information is integrated into the task status table, where each entry contains fields such as "taskID", "currentStatus", "startTime", "estimatedCompletionTime", "actualProgress", "resourceUtilization", etc. By continuously updating this table, the gateway can maintain an accurate, real-time global view of task execution, providing key basis for the system's dynamic scheduling and optimization decisions.
[0086] Specifically, the process of monitoring the task status table, identifying completed tasks, and determining the subsequent tasks that can be started based on the inter-task data flow graph to obtain the task execution queue involves complex dependency analysis and scheduling decisions. The edge computing gateway continuously monitors the updated task status table, paying special attention to the tasks whose status has changed to "completed". Whenever a task is detected to be completed, the gateway triggers a series of operations. First, the gateway locates the node of the completed task in the inter-task data flow graph, and then traverses all the edges starting from the node to find all potential subsequent tasks. For each potential subsequent task, the gateway checks the status of all its predecessor tasks. A task is considered to be startable only when all its predecessor tasks have been completed. The gateway also needs to consider resource availability to ensure that there are enough computing and storage resources to execute new tasks. Tasks that meet all conditions are added to a queue to be executed. This queue is not just a simple list of tasks, but a data structure with priorities. The priority is determined based on multiple factors, such as the criticality of the task, the expected execution time, resource requirements, etc. The gateway may use complex algorithms, such as weighted fair queues or multi-level feedback queues, to manage this execution queue. Each entry in the queue contains information such as the task identifier, priority score, resource requirements, etc. The gateway continuously updates this queue, inserting new tasks into the appropriate position when they become executable, and removing them from the queue when they start executing. In this way, the gateway ensures the correct order of task execution while optimizing overall execution efficiency and resource utilization.
[0087] Specifically, according to the task execution queue, the processing results of completed tasks are collected from the local processing unit and the collaborative edge devices, and the process of obtaining the original result data set involves distributed data collection and management. The edge computing gateway first creates a data structure to store the collected results. Then, the gateway traverses the task execution queue and identifies all completed tasks. For each completed task, the gateway needs to determine its execution location (local processing unit or specific collaborative edge device). For tasks executed locally, the gateway directly extracts the result data from the local storage. For tasks executed on collaborative edge devices, the gateway needs to initiate a data transmission request. These requests are sent through a pre-established secure channel, and methods such as HTTPS or proprietary encryption protocols may be used to ensure the security of data transmission. During the data transmission process, the gateway needs to handle possible network delays or interruptions, implement retry mechanisms and data integrity checks. The received data is temporarily stored in the gateway's buffer. For large result data, the gateway may need to implement streaming or block transmission strategies to avoid memory overflow. The result data of each task is tagged with metadata, including information such as task ID, execution time, and data size. These tagged result data sets form the original result data set. The gateway also needs to verify whether the collected data is complete and conforms to the expected format and size. If an anomaly is found, the gateway may need to re-request data or mark the task as needing to be re-executed. Through this process, the gateway integrates the task results scattered across different processing units into a centrally managed data set, preparing for subsequent data processing and analysis.
[0088] Specifically, the process of format unification and data normalization of the original result data set to obtain the intermediate result data involves complex data conversion and standardization operations. The edge computing gateway first analyzes each data item in the original result data set to identify its format, structure, and data type. Since the results may come from different processing units and task types, the data format may be different. The gateway needs to define a unified target format that should be compatible with all types of result data and facilitate subsequent analysis and processing. The format unification process may include data structure conversion (such as unifying different JSON or XML formats), data type standardization (such as ensuring that all numerical types use the same precision), and metadata normalization (such as unified timestamp format). In this process, the gateway may need to use a complex data conversion library or a custom parser. Data normalization involves adjusting the value range, handling outliers, and filling missing values. For example, values of different dimensions may need to be standardized and mapped to a unified range (such as between 0 and 1). Outliers may need to be identified and properly handled by statistical methods to avoid undue impact on subsequent analysis. For missing data, the gateway may use mean filling, regression prediction, or more complex interpolation methods. In addition, the gateway also needs to ensure data consistency and resolve possible data conflicts. For example, if multiple tasks produce different results about the same entity, a merging or selection strategy needs to be implemented. The processed data is organized into a structured intermediate result dataset, and each data item follows a unified format and standard. This intermediate result dataset not only has a unified format, but also has significantly improved data quality and consistency, providing a high-quality data foundation for subsequent in-depth analysis and decision-making.
[0089] 104. Perform a consistency check on the intermediate result data, and after the consistency check passes, use a predefined analysis model to perform an in-depth analysis on the result data after the consistency check passes, obtain a comprehensive analysis result, and distribute the comprehensive analysis result to relevant terminal devices.
[0090] In one embodiment of the present invention, the consistency check is performed on the intermediate result data, and after the consistency check passes, a predefined analysis model is used to perform an in-depth analysis on the result data after the consistency check passes to obtain a comprehensive analysis result, and the comprehensive analysis result is distributed to the relevant terminal device, including: applying a data integrity check algorithm to the intermediate result data, generating a check code, and comparing it with the expected check code to obtain a data consistency check result; based on the data consistency check result, screening out a data subset that passes the check to form verified result data; inputting the verified result data into a predefined analysis model, performing feature extraction, pattern recognition and anomaly detection to obtain a preliminary analysis result; performing multi-dimensional cross-validation and result fusion on the preliminary analysis result to obtain a comprehensive analysis result, and according to a preset data distribution strategy, encrypting and dividing the comprehensive analysis result into multiple data packets, and transmitting them to the relevant terminal device through a secure channel.
[0091] Specifically, the process of applying a data integrity check algorithm to the intermediate result data, generating a check code, and comparing it with the expected check code to obtain the data consistency check result involves multiple levels of data verification. The edge computing gateway first selects an appropriate check algorithm, such as cyclic redundancy check (CRC), MD5, or SHA series hash functions. The selected algorithm needs to balance computational efficiency and security. The gateway applies this algorithm to the entire intermediate result data set to generate a unique check code. At the same time, the gateway also generates a separate check code for each independent data item in the data set. These check codes are compared with the check codes pre-calculated during data transmission or processing. The comparison process not only checks the consistency of the overall data set, but also verifies the integrity of each independent data item. If any mismatch is found, the gateway will mark the corresponding data item as a potential problem. In addition, the gateway also performs a structural integrity check to ensure that all necessary data fields are present and in the correct format. This includes checking whether the data type, range, length, etc. meet the predefined specifications. The gateway also checks the logical relationship between the data to ensure that the association between different data items remains consistent. The results of all these checks are integrated into a detailed data consistency check report, which contains information such as the overall consistency status, the inspection results of each data item, the problems found and their severity. This report not only identifies problems in the data, but also provides reliability assurance for subsequent data processing and analysis.
[0092] The process of filtering out a subset of data that has passed the check based on the results of the data consistency check to form the verified result data involves strict data filtering and reconstruction. The edge computing gateway first sets a series of screening criteria based on the previous consistency check results, including data integrity, format correctness, logical consistency, etc. The gateway evaluates each data item one by one, and only data that fully meets all the criteria will be retained. For data that fails the check, the gateway will adopt different processing strategies depending on the nature of the problem. For minor format problems, the gateway may try to automatically correct them; for cases where the value range is slightly beyond the expected range, moderate adjustments may be made. However, for serious data corruption or inconsistency, the relevant data items will be completely excluded. In this process, the gateway will also evaluate the importance and irreplaceability of the data. For critical but minor data, the gateway may mark them for manual review. The screening process also includes data deduplication and merging to ensure that the final data set does not contain redundant information. The gateway will also reorganize the data structure to optimize the storage and access efficiency of the data. Finally, the gateway generates a new data set, the verified result data. This data set contains not only the original data that passed the check, but also the verification status of each data item, any corrections or adjustments made, and the credibility score of the data. This rigorously validated dataset provides high-quality, reliable input for subsequent in-depth analysis, significantly reducing the risk of analytical errors caused by data issues.
[0093] The process of inputting the verified result data into the predefined analysis model, performing feature extraction, pattern recognition and anomaly detection, and obtaining preliminary analysis results involves complex data analysis techniques and algorithm applications. The edge computing gateway first loads the predefined analysis models, which may include machine learning algorithms, statistical analysis tools, rule engines, etc. In the feature extraction stage, the gateway uses a variety of techniques such as principal component analysis (PCA) and convolutional neural network (CNN) to extract key features from the raw data. These features may include statistical indicators, time series features, frequency domain features, etc., depending on the nature of the data and the analysis objectives. In the pattern recognition stage, the gateway applies classification algorithms such as support vector machines (SVM) and random forests to identify important patterns and categories in the data. This may involve clustering the data to find natural groupings in the data. The anomaly detection stage uses algorithms such as isolation forests and single-class SVM to identify data points that deviate significantly from normal patterns. The gateway also performs time series analysis to predict future trends or identify periodic patterns. Throughout the analysis process, the gateway needs to dynamically adjust model parameters to ensure the accuracy and reliability of the analysis results. The preliminary analysis results contain multiple levels of information, such as the main features identified and their importance, the main patterns in the data and their distribution, the detected anomalies and their detailed descriptions, the predicted trends and possible future states, etc. These results are organized into a structured report, including numerical results, chart visualization, text description and other forms, to provide comprehensive data support for subsequent in-depth analysis and decision-making.
[0094] The preliminary analysis results are cross-validated and fused in multiple dimensions to obtain comprehensive analysis results. According to the preset data distribution strategy, the comprehensive analysis results are encrypted and divided into multiple data packets, and transmitted to the relevant terminal devices through a secure channel. The process involves complex data integration, security processing and transmission mechanisms. The edge computing gateway first cross-validates the preliminary analysis results, which includes repeated analysis of the same data set using different analysis methods and comparing the consistency of the results of different methods. The gateway also applies statistical methods such as cross-validation and bootstrap to evaluate the stability and reliability of the results. In the result fusion stage, the gateway uses ensemble learning techniques such as bagging, boosting or stacking to combine the outputs of different models into more robust and accurate predictions. This process also includes the processing of conflicting results, which may involve expert systems or rule-based decision logic. The fused results form a comprehensive analysis report containing key findings, prediction results, recommended measures, etc. Next, the gateway determines the type and level of information that each terminal device should receive according to the preset data distribution strategy. The gateway then encrypts the comprehensive analysis results and uses strong encryption algorithms such as AES or RSA to ensure data security. The encrypted data is split into multiple smaller packets, each with a checksum and sequence number added. The gateway then establishes secure channels, possibly using SSL / TLS protocols, to encrypt communications with each end device. The packets are transmitted through these secure channels and reassembled and decrypted at the receiving end. Throughout the process, the gateway continuously monitors the transmission status to ensure that all packets are successfully delivered and correctly reassembled. This approach not only ensures the accuracy and reliability of the analysis results, but also protects the security of the data during transmission, allowing critical information to be distributed safely and efficiently to the end users in need.
[0095] In this embodiment, by performing correlation analysis on the computing tasks of multiple user devices, a task correlation graph is generated, and the cache status and data dependency of the task are identified in combination with the cache matrix to generate a task data processing requirement list and an inter-task data flow graph. The task data is preprocessed according to the task data processing requirement list, and based on the preset offloading decision and the inter-task data flow graph, the preprocessed task data is transmitted to the local processing unit and the collaborative edge device to generate a distributed task data set. The execution status of the distributed task data set is tracked, the processing result data is collected and a consistency check is performed, the result data is deeply analyzed using an analysis model, and a comprehensive analysis result is generated and distributed to relevant terminal devices. The present invention improves the data processing efficiency and consistency assurance capabilities of the edge computing gateway through dynamic task scheduling and distributed collaborative processing.
[0096] The above describes the data processing and analysis method of the edge computing gateway in the embodiment of the present invention. The following describes the data processing and analysis system of the edge computing gateway in the embodiment of the present invention. Figure 2 , an embodiment of the data processing and analysis system of the edge computing gateway in the embodiment of the present invention includes:
[0097] The association analysis module 201 is used to perform association analysis on the computing tasks of multiple user devices through the edge computing gateway to obtain a task association graph, and identify the cache status and data dependency of each computing task according to the task association graph and the pre-stored cache matrix to obtain a task data processing requirement list and a data flow graph between tasks;
[0098] The task allocation module 202 is used to pre-process the task data of the computing task according to the task data processing requirement list to obtain pre-processed task data, and transmit the pre-processed task data to the local processing unit of the edge computing gateway and the cooperating edge device for processing according to the preset offloading decision and the inter-task data flow diagram to obtain a distributed task data set;
[0099] A tracking module 203 is used to track the execution status of the distributed task data set, obtain a task status table, collect and transmit the processing results of the distributed task data set according to the task status table, and obtain intermediate result data;
[0100] The inspection and distribution module 204 is used to perform a consistency check on the intermediate result data, and after the consistency check passes, use a predefined analysis model to perform an in-depth analysis on the result data after the consistency check passes, obtain a comprehensive analysis result, and distribute the comprehensive analysis result to relevant terminal devices.
[0101] In an embodiment of the present invention, the data processing and analysis system of the edge computing gateway runs the data processing and analysis method of the edge computing gateway. The data processing and analysis system of the edge computing gateway performs correlation analysis on the computing tasks of multiple user devices, generates a task correlation graph, combines the cache status and data dependency of the task with the cache matrix, and generates a task data processing requirement list and an inter-task data flow graph. The task data is preprocessed according to the task data processing requirement list, and the preprocessed task data is transmitted to the local processing unit and the collaborative edge device based on the preset unloading decision and the inter-task data flow graph to generate a distributed task data set. The execution status of the distributed task data set is tracked, the processing result data is collected and consistency checked, the result data is deeply analyzed using the analysis model, and a comprehensive analysis result is generated and distributed to the relevant terminal devices. The present invention improves the data processing efficiency and consistency guarantee capability of the edge computing gateway through dynamic task scheduling and distributed collaborative processing.
[0102] Those skilled in the art can clearly understand that, for the convenience and brevity of description, the specific working process of the above-described system, device, or unit can refer to the corresponding process in the aforementioned method embodiment and will not be repeated here.
[0103] If the integrated unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present invention is essentially or the part that contributes to the prior art or the whole or part of the technical solution can be embodied in the form of a software product. The computer software product is stored in a storage medium, including several instructions to enable a computer device (which can be a personal computer, server, or network device, etc.) to perform all or part of the steps of the method described in each embodiment of the present invention. The aforementioned storage medium includes: U disk, mobile hard disk, read-only memory (ROM), random access memory (RAM), disk or optical disk and other media that can store program code.
[0104] As described above, the above embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit the same. Although the present invention has been described in detail with reference to the aforementioned embodiments, those skilled in the art should understand that the technical solutions described in the aforementioned embodiments may still be modified, or some of the technical features thereof may be replaced by equivalents. However, these modifications or replacements do not deviate the essence of the corresponding technical solutions from the spirit and scope of the technical solutions of the embodiments of the present invention.
Claims
1. A data processing and analysis method for an edge computing gateway, characterized in that: The data processing and analysis method of the edge computing gateway includes: The edge computing gateway performs correlation analysis on the computing tasks of multiple user devices to obtain a task correlation graph, and traverses each node in the task correlation graph, searches for the corresponding cache item in the pre-stored cache matrix for the task ID of each node, and obtains a task cache status list; according to the task cache status list, each edge in the task correlation graph is marked to obtain a task correlation graph with cache marks; the task correlation graph with cache marks is traversed to extract the predecessor task information and the subsequent task information of each computing task to obtain a task dependency table; according to the task dependency table and the task cache status list, the data processing requirements of each computing task are calculated to obtain a task data processing requirement list; based on the task correlation graph with cache marks and the task data processing requirement list, a data flow diagram is constructed to obtain a data flow diagram between tasks; According to the task data processing requirement list, preprocess the task data of the computing task to obtain preprocessed task data, and according to the preset offloading decision and the inter-task data flow diagram, transmit the preprocessed task data to the local processing unit of the edge computing gateway and the cooperating edge device for processing to obtain a distributed task data set; Tracking the execution status of the distributed task data set to obtain a task status table, collecting and transmitting the processing results of the distributed task data set according to the task status table to obtain intermediate result data; A consistency check is performed on the intermediate result data, and after the consistency check passes, a predefined analysis model is used to perform an in-depth analysis on the result data after the consistency check passes to obtain a comprehensive analysis result, and the comprehensive analysis result is distributed to relevant terminal devices.
2. The data processing and analysis method of the edge computing gateway according to claim 1 is characterized in that: The performing correlation analysis on the computing tasks of multiple user devices by the edge computing gateway to obtain a task correlation graph includes: Scanning the task data and expected output type of each computing task through the edge computing gateway to obtain task data features; According to the task data features, a task similarity matrix between computing tasks is constructed, and a clustering algorithm is applied to the task similarity matrix to obtain task groups; According to the task groups and task data features, a directed acyclic graph is constructed to obtain an initial task association graph, and the initial task association graph is topologically sorted and redundant edges are eliminated to obtain a task association graph.
3. The data processing and analysis method of the edge computing gateway according to claim 1 is characterized in that: According to the preset offloading decision and the inter-task data flow diagram, the pre-processed task data is transmitted to the local processing unit of the edge computing gateway and the cooperating edge device for processing, and the distributed task data set is obtained, including: According to the preset offloading decision, the computing tasks are classified into local processing tasks and remote processing tasks, and a task classification list is obtained; According to the task classification list and the data flow diagram between tasks, a data transmission instruction is generated for each computing task to obtain a task data transmission instruction set; According to the task data transmission instruction set, the pre-processing task data of the local processing task is loaded into the local processing unit of the edge computing gateway to obtain a local task data subset; According to the task data transmission instruction set, the pre-processed task data of the remote processing task is transmitted to the cooperating edge device through a secure channel to obtain a remote task data subset; The local task data subset and the remote task data subset are merged, and task metadata is added to obtain a distributed task data set with labels.
4. The data processing and analysis method of the edge computing gateway according to claim 3 is characterized in that: The step of generating a data transmission instruction for each computing task by comparing the task classification list and the data flow diagram between tasks to obtain a task data transmission instruction set includes: Traversing each computing task in the task classification list, extracting classification information and task identifiers of each computing task, and obtaining a task basic information table; According to the task basic information table, locate the node of each computing task in the inter-task data flow graph, identify the predecessor task and the successor task of each computing task based on the node, and obtain the task dependency table; Based on the task dependency table, calculating the data transmission priority of each computing task to obtain a task transmission priority list; According to the task transmission priority list and task classification information, a target execution unit is assigned to each computing task to obtain a task distribution mapping table; In combination with the task distribution mapping table and the inter-task data flow diagram, a data transmission instruction including a source address, a target address, a data size and a transmission time window is generated to form a task data transmission instruction set.
5. The data processing and analysis method of the edge computing gateway according to claim 1, characterized in that: The tracking of the execution status of the distributed task data set to obtain a task status table, collecting and transmitting the processing results of the distributed task data set according to the task status table, and obtaining the intermediate result data includes: Starting a timer for each computing task in the distributed task data set, and periodically checking the task execution progress to obtain real-time task progress information; According to the real-time task progress information, the task status table is updated to record the current status, start time and expected completion time of each computing task; Monitor the task status table, identify completed tasks, and determine subsequent tasks that can be started based on the data flow diagram between tasks to obtain a task execution queue; According to the task execution queue, processing results of completed tasks are collected from local processing units and cooperating edge devices to obtain an original result data set; The original result data set is subjected to format unification and data normalization processing to obtain intermediate result data.
6. The data processing and analysis method of the edge computing gateway according to claim 1, characterized in that: The performing a consistency check on the intermediate result data, and after the consistency check passes, performing an in-depth analysis on the result data after the consistency check passes using a predefined analysis model to obtain a comprehensive analysis result, and distributing the comprehensive analysis result to relevant terminal devices includes: Applying a data integrity check algorithm to the intermediate result data to generate a check code, and comparing it with the expected check code to obtain a data consistency check result; According to the data consistency check result, a data subset that passes the check is screened out to form verified result data; Input the verified result data into a predefined analysis model, perform feature extraction, pattern recognition and anomaly detection, and obtain preliminary analysis results; The preliminary analysis results are subjected to multi-dimensional cross-validation and result fusion to obtain comprehensive analysis results. According to a preset data distribution strategy, the comprehensive analysis results are encrypted and divided into multiple data packets, which are transmitted to relevant terminal devices through a secure channel.
7. A data processing and analysis system for an edge computing gateway, characterized in that: The data processing and analysis system of the edge computing gateway includes: An association analysis module is used to perform association analysis on the computing tasks of multiple user devices through the edge computing gateway to obtain a task association graph, and traverse each node in the task association graph, search for the corresponding cache item in the pre-stored cache matrix for the task ID of each node, and obtain a task cache status list; mark each edge in the task association graph according to the task cache status list to obtain a task association graph with cache marks; traverse the task association graph with cache marks, extract the predecessor task information and subsequent task information of each computing task, and obtain a task dependency table; calculate the data processing requirements of each computing task according to the task dependency table and the task cache status list, and obtain a task data processing requirement list; based on the task association graph with cache marks and the task data processing requirement list, construct a data flow diagram to obtain a data flow diagram between tasks; A task allocation module is used to pre-process the task data of the computing task according to the task data processing requirement list to obtain pre-processed task data, and transmit the pre-processed task data to the local processing unit of the edge computing gateway and the cooperating edge device for processing according to the preset offloading decision and the inter-task data flow diagram to obtain a distributed task data set; A tracking module, used to track the execution status of the distributed task data set, obtain a task status table, collect and transmit the processing results of the distributed task data set according to the task status table, and obtain intermediate result data; The inspection and distribution module is used to perform consistency check on the intermediate result data, and after the consistency check passes, use a predefined analysis model to perform in-depth analysis on the result data after the consistency check passes, obtain a comprehensive analysis result, and distribute the comprehensive analysis result to relevant terminal devices.
Citation Information
Patent Citations
Task cache-based computing migration method in edge computing
CN112860350A
Task cooperative scheduling method and device, equipment and medium
CN118245184A