Data acquisition method and device for edge gateway of Internet of Things, and medium

By deploying monitoring and acquisition processes in edge gateway clusters, and using dynamic load balancing algorithms and real-time health status detection, the problems of inefficiency and high latency under traditional data acquisition methods are solved, and efficient and reliable data acquisition and task migration are achieved.

CN120034559AActive Publication Date: 2025-05-23INSPUR GENERSOFT CO LTD

Patent Information

Application Number
CN202510226724.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-02-27
Publication Date
2025-05-23
Estimated Expiration
2045-02-27

AI Technical Summary

Technical Problem

In the scenario where edge gateways face large-scale device data acquisition, data acquisition efficiency under traditional data acquisition methods is inefficient, and task migration relies on manual intervention, and the takeover delay is long, which cannot meet the scenario requirements of high real-time.

Method used

By deploying at least one monitoring process and multiple acquisition processes in each server of the edge gateway cluster, a preset dynamic load balancing algorithm is used to allocate acquisition tasks in multiple acquisition processes, and the process health status is monitored in real time. When an abnormality occurs, the takeover target is automatically determined to achieve seamless task connection.

Benefits of technology

It improves the speed and efficiency of data acquisition, reduces the risk of data acquisition and interruption, ensures the stability and reliability of the system, meets the requirements of high real-timeness, and has good adaptability and scalability.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120034559A_ABST
    Figure CN120034559A_ABST
Patent Text Reader

Abstract

The embodiment of the invention discloses a data collection method and device for an edge gateway of the Internet of Things and a medium, and relates to the technical field of data collection, and the method comprises the steps: determining at least one monitoring process and a plurality of corresponding collection processes which are deployed in each server in an edge gateway cluster in advance; through the monitoring process, distributing the collection task in the plurality of collection processes according to a preset dynamic load balancing algorithm so as to determine a plurality of online collection processes corresponding to the collection task; executing an acquisition task through the plurality of online acquisition processes, detecting health states between the monitoring process and the plurality of online acquisition processes, and determining real-time health state data; and when the specified monitoring process and / or the specified online acquisition process included in the real-time health state data are / is abnormal, determining a takeover target in the edge gateway cluster so as to process an acquisition task of the specified monitoring process and / or the specified online acquisition process through the takeover target, thereby realizing data acquisition of the Internet of Things equipment.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of data collection technology, and in particular to a data collection method, device and medium for an edge gateway of the Internet of Things. Background Art

[0002] With the rapid development of IoT technology, more and more smart devices are connected to the network, forming a huge network of connected devices. The amount of data generated by these devices has increased dramatically, which has put forward higher requirements for data collection, processing and analysis. The traditional centralized data processing model has been unable to meet the needs of modern IoT applications due to its inherent limitations, such as bandwidth limitations, latency issues, and single-point abnormality risks of central nodes. Especially in the field of edge computing, edge gateways, as key nodes connecting devices and cloud platforms, undertake the initial processing and forwarding of data.

[0003] However, when the edge gateway is faced with the collection of large-scale device data, the processing capacity of a single node is often difficult to cope with. On the one hand, the processing capacity of a single node is limited, and it is difficult to cope with massive concurrent data, resulting in low data collection efficiency; on the other hand, once a single node fails, the data collection work of the surrounding devices will be interrupted, seriously affecting the stability and reliability of the system. In the traditional data collection method, a distributed architecture is used for data collection, but the monitoring of the process usually relies on a single heartbeat detection mechanism, resulting in misjudgment or missed judgment; and in the event of an abnormal situation, task migration relies on manual intervention, and the takeover delay is as long as minutes, which cannot meet the needs of scenarios with high real-time requirements. Therefore, in the scenario where the edge gateway is faced with the collection of large-scale device data, the data collection efficiency under the traditional data collection method is low, and the task migration relies on manual intervention, and the takeover delay is long, which cannot meet the needs of scenarios with high real-time requirements. Summary of the invention

[0004] One or more embodiments of the present specification provide a data collection method, device and medium for an IoT edge gateway, which are used to solve the following technical problems: In a scenario where the edge gateway is faced with collecting large-scale device data, the data collection efficiency under traditional data collection methods is low, and task migration relies on manual intervention, with long takeover delays, which cannot meet the needs of scenarios with high real-time requirements.

[0005] One or more embodiments of this specification adopt the following technical solutions:

[0006] One or more embodiments of the present specification provide a data collection method for an edge gateway of an Internet of Things, the method comprising: determining at least one monitoring process and corresponding multiple collection processes pre-deployed in each server in an edge gateway cluster; obtaining a pre-configured collection task, and distributing the collection task among the multiple collection processes according to a preset dynamic load balancing algorithm through the monitoring process to determine multiple online collection processes corresponding to the collection task; executing the collection task through the multiple online collection processes, and detecting the health status between the monitoring process and the multiple online collection processes to determine real-time health status data; when the real-time health status data contains an abnormality in a specified monitoring process and / or a specified online collection process, determining a takeover target in the edge gateway cluster to process the collection tasks of the specified monitoring process and / or the specified online collection process through the takeover target to realize data collection of Internet of Things devices.

[0007] One or more embodiments of this specification provide a data acquisition device for an edge gateway of the Internet of Things, including:

[0008] at least one processor; and,

[0009] a memory communicatively connected to the at least one processor; wherein,

[0010] The memory stores instructions that can be executed by the at least one processor, and the instructions are executed by the at least one processor to enable the at least one processor to perform the above method.

[0011] One or more embodiments of the present specification provide a non-volatile computer storage medium storing computer executable instructions, wherein the computer executable instructions are configured to execute the above method.

[0012] At least one of the above-mentioned technical solutions adopted in the embodiments of this specification can achieve the following beneficial effects: through the technical solutions of the embodiments of this specification, the traditional centralized data processing mode is limited by bandwidth and delay, and is unable to cope with large-scale data collection. The embodiments of this specification can reasonably distribute the collection tasks to multiple online collection processes according to the real-time load status of the collection process, such as CPU occupancy, memory occupancy and network bandwidth utilization, through a preset dynamic load balancing algorithm, effectively solving the problem of insufficient processing capacity of a single node, avoiding low collection efficiency caused by excessive data volume, greatly improving the speed of data collection, and being able to quickly process massive concurrent data to meet the requirements of modern Internet of Things for data collection efficiency; in the traditional mode, a single point abnormality of the central node will interrupt data collection, and the traditional distributed architecture relies on a single heartbeat detection and is prone to misjudgment and omission. The embodiments of this specification accurately identify abnormal processes through comprehensive health status detection of the monitoring process and the online collection process. Once a specified monitoring process or a specified online collection process is found to be abnormal, it can be quickly detected in the edge gateway cluster. The takeover target is determined in the process, and the task takeover is automatically completed without human intervention, which greatly reduces the risk of data collection interruption and ensures the stability and reliability of the system; the traditional data collection method relies on manual task migration under abnormal circumstances, and the takeover delay is as high as minutes, which cannot meet the high real-time requirements. The embodiment of this specification can monitor the process status in real time. When the real-time health status data shows an abnormality, it can quickly determine the takeover target and complete the task migration, and achieve seamless connection of tasks with extremely short delays. For example, in scenarios with extremely high real-time requirements such as industrial automation and intelligent transportation, the timeliness of data collection can be ensured; the scale of IoT devices and data traffic are constantly changing dynamically, and traditional models are difficult to adapt. The embodiment of this specification is based on the edge gateway cluster architecture. Whether it is the increase in data volume brought by the access of new devices, or the device offline, process abnormalities, etc., it can be flexibly responded to. The dynamic load balancing algorithm can adjust the task allocation at any time, automatically detect and process abnormal processes, and realize self-optimization and adjustment of the system. It can run stably in a complex and changeable IoT environment and has strong adaptability and scalability. BRIEF DESCRIPTION OF THE DRAWINGS

[0013] In order to more clearly illustrate the technical solutions in the embodiments of this specification or the prior art, the following briefly introduces the drawings required for use in the embodiments or the prior art description. Obviously, the drawings described below are only some embodiments recorded in this specification. For ordinary technicians in this field, other drawings can be obtained based on these drawings without creative labor. In the drawings:

[0014] Figure 1 A flow chart of a data collection method for an edge gateway of the Internet of Things provided in an embodiment of this specification;

[0015] Figure 2 A schematic diagram of a system architecture for an application scenario provided in an embodiment of this specification;

[0016] Figure 3 A flowchart of a monitoring process failure and recovery provided in an embodiment of this specification;

[0017] Figure 4 A flowchart of failure and recovery of a collection process provided by an embodiment of this specification;

[0018] Figure 5 A schematic diagram of the structure of a data acquisition device for an edge gateway of the Internet of Things provided in an embodiment of this specification. DETAILED DESCRIPTION

[0019] In order to enable those skilled in the art to better understand the technical solutions in this specification, the technical solutions in the embodiments of this specification will be clearly and completely described below in conjunction with the drawings in the embodiments of this specification. Obviously, the described embodiments are only part of the embodiments of this specification, not all of the embodiments. Based on the embodiments of this specification, all other embodiments obtained by ordinary technicians in this field without creative work should fall within the scope of protection of this specification.

[0020] The embodiments of this specification provide a data collection method for an edge gateway of the Internet of Things. It should be noted that the execution subject in the embodiments of this specification can be a server or any device with data processing capabilities. Figure 1 A flow chart of a data collection method for an edge gateway of the Internet of Things provided in an embodiment of this specification, such as Figure 1 As shown, it mainly includes the following steps:

[0021] Step S101, determining at least one monitoring process and corresponding multiple collection processes pre-deployed in each server in the edge gateway cluster.

[0022] In one embodiment of the present specification, the edge gateway system in the embodiment of the present specification adopts a layered architecture, which includes two programs: a monitoring process and a collection process. Through division of labor and cooperation, resource allocation is optimized to improve the overall performance of the system. Figure 2 A schematic diagram of a system architecture for an application scenario provided in an embodiment of this specification, such as Figure 2As shown in the figure, the edge gateway cluster includes a monitoring process and corresponding multiple collection processes. The monitoring process is used for synchronization and coordination between processes and emergency handling and transfer in case of failure. Its main functions include monitoring function, failover function and collection task scheduling function. The monitoring function is used to continuously monitor the running status of the collection process to ensure its normal operation. The failover mechanism is used to quickly reallocate tasks to other normally running collection processes when a failure of the collection process is detected; similarly, if the monitoring process on the server fails, the system can intelligently transfer all collection tasks of the server to other healthy servers to ensure the continuity and stability of data collection. The collection task scheduling function is used to intelligently allocate and schedule tasks according to the real-time load of each collection process to achieve load balancing and improve system efficiency. The collection process is used to perform data collection tasks. By working in collaboration with the monitoring process, it shares the pressure of data collection and improves the data processing capability of the system.

[0023] In one embodiment of the present specification, a monitoring process and multiple collection processes are deployed in advance on each server of the edge gateway cluster, making full use of the computing resources of the server while ensuring the high availability and load balancing of the system. The monitoring processes on different servers exchange information through an efficient communication mechanism to achieve data synchronization and state sharing. At the same time, the monitoring process is responsible for scheduling the data collection tasks of the collection process to ensure the orderly progress of data collection work.

[0024] Step S102, obtaining a pre-configured collection task, and distributing the collection task among multiple collection processes according to a preset dynamic load balancing algorithm by monitoring the process, so as to determine multiple online collection processes corresponding to the collection task.

[0025] In the field of IoT data collection, the traditional task allocation method is highly dependent on static device lists. The lists are based on the access status of devices at specific time nodes, data transmission requirements and other information, and are relatively stable within a certain period of time. However, in the actual IoT operating environment, the device scale and data traffic are in a state of dynamic change. For example, in the case of dynamic changes caused by sudden device access, these devices that suddenly accessed are not taken into account when formulating the task allocation strategy, resulting in the inability of the collection process to collect and process the data of these devices in a timely manner, causing data backlog. In addition, in a complex IoT environment, devices may be offline for a variety of reasons, such as hardware failure, unstable network connection, etc. When a device is offline, the task allocation method based on the static device list cannot adjust the collection task in time, and the collection resources originally allocated to the offline device will be idle, while other devices that are operating normally and have a large amount of data cannot obtain these idle resources, making the resource utilization efficiency of the entire system low and unable to achieve reasonable resource allocation. In summary, the traditional task allocation method that relies on the static device list lacks an effective response mechanism when facing the dynamic changes in device scale and data traffic such as sudden device access and offline, and cannot cope with the dynamically changing device scale and data traffic, resulting in uneven resource utilization.

[0026] Through the monitoring process, the collection task is allocated among the multiple collection processes according to a preset dynamic load balancing algorithm to determine the multiple online collection processes corresponding to the collection task, specifically including: obtaining real-time load indicator data of the collection process, wherein the real-time load indicator data includes CPU occupancy, memory occupancy and network bandwidth utilization; determining the comprehensive load coefficient corresponding to the collection process according to the real-time load indicator data, and determining the dynamic weight of each of the collection processes based on the comprehensive load coefficient; performing virtual slot number mapping through the dynamic weight of the collection process to determine the virtual slot number corresponding to each of the collection processes; and performing weighted round-robin allocation among the multiple collection processes according to the virtual slot number corresponding to each of the collection processes to determine the multiple online collection processes corresponding to the collection task.

[0027] In one embodiment of the present specification, the monitoring process collects three-dimensional resource indicators of each collection process, including CPU occupancy rate Ci, memory occupancy rate Mi, and network bandwidth utilization rate Ni, through an embedded probe at a preset time (e.g., 200ms) as a period, and at the same time, calculates the resource fluctuation variance in the past time period. For example, in the past 5 minutes, the resource fluctuation variances corresponding to the CPU occupancy rate, memory occupancy rate, and network bandwidth utilization rate are respectively According to the comprehensive load factor Among them, μ takes the value of 0.1 as the historical fluctuation penalty coefficient, which is used to avoid load oscillation nodes. The setting of the historical fluctuation penalty coefficient can prevent load oscillation nodes from participating in task allocation and prevent the system stability from being affected by frequent and large fluctuations in node load; ε = 1*10-5, which is used to prevent small constants with a denominator of 0. α, β, and γ can be set according to demand, for example, they can be set to 0.4, 0.3, and 0.3, and dynamically adjusted according to the business scenario. The dynamic weight corresponding to the collection process is calculated according to the method of Wi = 1 / Li for the obtained comprehensive load coefficient.

[0028] First remove " Offline ” or " Waiting for restart ” The node in the state, where the node refers to the collection process, is excluded from the current allocation pool. According to the product of the weight ratio of each node and the preset fixed total number of slots, the number of virtual slots corresponding to each collection process is generated. Here, for the convenience of probability mapping, the fixed total number of slots can be set to 100. Taking 100 as an example, assuming that W1 is 0.6, W2 is 0.3, and W3 is 0.1, the corresponding number of slots is 60, 30 and 10 respectively, a virtual slot sequence is constructed, and the slot interval 0-99 is divided into P1: [0-59], P2: [60-89], P3: [90-99] in proportion. Each time a task is assigned, the corresponding collection process is selected according to the slot order. In the above example, the first collection process has the highest weight and occupies 60% of the slots, which means a 60% probability of selection, that is, the node with the lightest load undertakes the most tasks. As tasks are assigned, when the load of the first collection process gradually increases → Its weight may be reduced to 0.4, the slot adjusted to 40, and subsequent task allocations automatically reduced to the frequency of this node.

[0029] Through the above technical solution, three-dimensional resource indicators such as CPU occupancy rate, memory occupancy rate and network bandwidth utilization rate, as well as the resource fluctuation variance in the past time period are comprehensively considered, and the comprehensive load coefficient and dynamic weight are calculated through complex formulas, which can more accurately reflect the actual load capacity of the collection process; the number of virtual slots is divided according to the weight ratio, and the collection process is selected in slot order, so that the node with the lightest load undertakes the most tasks, and the task allocation frequency is automatically adjusted as the node load changes, thereby realizing dynamic and reasonable task allocation; reasonable task allocation enables resources to be more effectively allocated to each collection process, avoiding the situation where some processes are overloaded and some process resources are idle; α, β, and γ can be dynamically adjusted according to business scenarios, which can meet the emphasis on CPU, memory and network resources in different IoT application scenarios. In scenarios with high requirements for network transmission, the value of γ can be appropriately increased, so that task allocation is more inclined to collection processes with high network bandwidth utilization, thereby improving the operating efficiency of the system in specific scenarios.

[0030] Step S103, executing the collection task through multiple online collection processes, and detecting the health status between the monitoring process and the multiple online collection processes to determine real-time health status data.

[0031] The health status between the monitoring process and the multiple online collection processes is detected to determine the real-time health status data, specifically including: through the information interaction between the multiple monitoring processes in the gateway cluster, according to the preset majority decision mechanism, the health status of the monitoring process is detected to determine the first health status data corresponding to the monitoring process; through the monitoring process, the health status of the multiple online collection processes is detected to determine the second health status data corresponding to the online collection process.

[0032] In one embodiment of the present specification, when the monitoring processes of different servers are started, the collection configuration information of the server is broadcast to all monitoring processes. The monitoring process starts the required collection process according to the configuration information to ensure that the collection process is ready. Through the information interaction between multiple monitoring processes in the gateway cluster, according to the preset majority decision mechanism, the health status of the monitoring process is detected to determine the first health status data corresponding to the monitoring process. Figure 3 A flowchart of a monitoring process failure and recovery provided in an embodiment of this specification is shown in FIG. Figure 3 As shown in the figure, the monitoring process records all received collection configurations to ensure the integrity of the information. The monitoring processes perform health checks by sending ping messages to each other, marking the processes that respond normally as " Online ” After multiple retries, it still did not respond normally and was "pending confirmation ” ; When most monitoring processes mark a monitoring process as " To be confirmed ” , the process is finally marked as " Offline ” , and notify other processes.

[0033] Figure 4 A flowchart of a failure and recovery process of a collection process provided by an embodiment of this specification is shown as follows: Figure 4 As shown, after the devices to be collected are assigned to each collection process, a notification to start collection is sent to the collection process to officially start the data collection work. Through the monitoring process, the health status of the multiple online collection processes is detected to determine the second health status data corresponding to the online collection process. Specifically, the monitoring process regularly sends ping messages to the collection process, and the collection process replies with pong messages to confirm its online status. If the collection process does not reply to the pong message within the specified time, the monitoring process will retry to send the ping message. If no reply is received for three consecutive times, the process is marked as" Waiting for restart ” The monitoring process attempts to restart in " Waiting for restart ” If the collection process still cannot be restored after restart, it will be marked as " Offline ” .

[0034] Through information interaction between multiple monitoring processes in the gateway cluster, according to a preset majority decision mechanism, a health status detection is performed on the monitoring process to determine the first health status data corresponding to the monitoring process, specifically including: according to a preset period, periodically collecting the first real-time resource data corresponding to the monitoring process, and broadcasting it to other monitoring processes in the gateway cluster, wherein the first real-time resource data includes real-time CPU data, real-time memory data and the number of real-time threads; through other monitoring processes in the gateway cluster, the first real-time resource data is hashed to determine the real-time resource fingerprint data corresponding to the monitoring process; according to the real-time resource fingerprint data and the majority decision mechanism, a health status detection is performed on the monitoring process to determine the first health status data corresponding to the monitoring process.

[0035] According to the real-time resource fingerprint data and the majority decision mechanism, the health status of the monitoring process is detected to determine the first health status data corresponding to the monitoring process, specifically including: if the first fingerprint deviation of the real-time resource fingerprint data of the monitoring process received by any other monitoring process and the preset local cluster average value is not greater than a preset threshold, the monitoring process is judged to be in a healthy state; if the first fingerprint deviation exceeds the threshold value N times in a row, the resource fingerprint sequence of the monitoring process in the historical time period is collected to predict the predicted resource fingerprint in the next time period; when the second fingerprint deviation between the predicted resource fingerprint and the real-time resource fingerprint data exceeds the preset threshold, the monitoring process is judged to be a monitoring process to be confirmed, a voting request is initiated to the gateway cluster, and the first health status data corresponding to the monitoring process is determined according to the voting result.

[0036] In one embodiment of the present specification, the first real-time resource data such as real-time CPU data, real-time memory data and real-time thread number of the monitoring process are collected at a preset period and broadcast to other monitoring processes in the gateway cluster to realize information sharing. After receiving the broadcast data, other monitoring processes perform hash calculations to convert complex resource data into representative real-time resource fingerprint data for subsequent comparison and analysis. The real-time resource fingerprint data received by each monitoring process is compared with the preset local cluster average value to calculate the first fingerprint deviation. If the deviation is not greater than the preset threshold, the monitoring process is determined to be in a healthy state; if the first fingerprint deviation exceeds the threshold N times in a row, the next step of prediction and analysis is entered. The resource fingerprint sequence within the historical time period of the monitoring process is collected to predict the predicted resource fingerprint for the next time period. If the second fingerprint deviation between the predicted resource fingerprint and the real-time resource fingerprint data exceeds the preset threshold, the monitoring process is determined to be a monitoring process to be confirmed. For the monitoring process to be confirmed, a voting request is initiated to the gateway cluster, and the voting results of other monitoring processes in the cluster are combined to determine the first health status data of the monitoring process and obtain the final health status conclusion.

[0037] For example, each monitoring process periodically (e.g., every 10 seconds) reports its own resource fingerprint (CPU / memory / number of threads) hash value Fingerprint = Hash (CPU% | Memory% | ThreadCount). If the received fingerprint of other processes deviates from the local cluster average by <5%, it is considered a normal response. If the fingerprint deviation exceeds the threshold for three consecutive times, start time series analysis: collect the fingerprint sequence of the process in the past 5 minutes, and predict the fingerprint value of the next cycle through a prediction model, such as the Long Short-Term Memory (LSTM) model. If the difference between the actual value and the predicted value exceeds 20%, it is determined to be an abnormal process. The monitoring process initiates a voting request to the cluster. If more than (N+1) / 2 monitoring processes (rounded up) agree, it is marked as an offline node, triggering task takeover. Here, N is the number of monitoring processes.

[0038] Through the above technical solution, multiple monitoring processes participate and are based on the majority decision-making mechanism, which avoids the one-sidedness of the judgment of a single monitoring process; combined with real-time data and historical data prediction analysis, the health status of the monitoring process is comprehensively evaluated from multiple dimensions, reducing the risk of misjudgment and providing more accurate health status information; preset periodic data collection and dynamic adjustment of judgment strategies can adapt to the dynamic changes in the use of monitoring process resources in the Internet of Things environment, effectively respond to resource fluctuations, and ensure the timeliness and effectiveness of health status detection; through the health status detection mechanism, potential problems can be discovered and handled in a timely manner, reducing data collection interruptions or errors caused by monitoring process failures, and improving the reliability of the entire Internet of Things edge gateway data collection system.

[0039] Through the monitoring process, the health status of the multiple online collection processes is detected to determine the second health status data corresponding to the online collection process, specifically including: sending a first number of ping instructions to the online collection process through the monitoring process; when the pong instruction returned by the online collection process is not received for a first preset number of times in a row, obtaining the real-time resource data of the online collection process; performing a secondary status detection on the online collection process according to the real-time resource data and pre-set historical baseline data; when the real-time resource data and the historical baseline data meet a preset condition, sending a simulation collection instruction to determine the execution result data of the online collection process, so as to determine the second health status data corresponding to the online collection process based on the execution result data.

[0040] In one embodiment of the present specification, the monitoring process sends a first number of ping instructions to the online collection process, and determines whether it is operating normally through the pong instruction returned by the online collection process. If the pong instruction can be received in time, it can be preliminarily considered that the collection process is in a normal working state. When the pong instruction is not received for a first preset number of times, it indicates that there may be a problem with the collection process. At this time, relying solely on the ping instruction detection is not enough to accurately determine its status. Therefore, the monitoring process will obtain the real-time resource data of the online collection process for more in-depth detection.

[0041] Compare and analyze the acquired real-time resource data with the preset historical baseline data. The historical baseline data is obtained based on the long-term accumulation of resource usage of the collection process under normal operation. By comparing the difference between the two, it is judged whether the state of the collection process is abnormal according to the preset conditions. The preset conditions can be the deviation range of resource data, for example, when the real-time CPU usage exceeds a certain percentage of the historical baseline data, it is considered to be abnormal. When the real-time resource data and the historical baseline data meet the preset conditions, that is, there is a situation that may affect the normal operation of the collection process, the monitoring process sends a simulated collection instruction. The simulated collection instruction simulates a real data collection task, and the collection process returns the execution result data after executing the instruction. The monitoring process finally determines whether the online collection process is offline based on this execution result data. If the execution result data shows that the execution verification failed and the execution result data is empty, the online collection process is judged to be offline.

[0042] For example, the monitoring process sends a 1-byte ping command per second, and the collection process immediately replies with a pong. If pong is not received for two consecutive times, it switches to resource probe mode, and the collection process needs to return its current CPU / memory usage. If the resource data deviates too much from the historical baseline (such as a sudden increase in CPU to 90%), send a simulated collection command (such as reading the device 0x00 register) to verify the actual execution capability. Based on this execution result data, it is finally determined whether the online collection process is offline.

[0043] The above technical solution not only relies on a simple ping-pong mechanism, but also combines the comparison of real-time resource data with historical baseline data, as well as the simulation of the collection instruction execution results for comprehensive judgment. This avoids misjudgment caused by accidental factors such as network fluctuations, and can more accurately determine the health status of the online collection process, providing a reliable basis for subsequent decision-making; through multi-level detection methods, it can better adapt to various complex situations, whether it is hardware failure, software anomalies or network fluctuations, it can effectively detect changes in the health status of the collection process, ensuring that the system can run stably even in complex environments.

[0044] Step S104, when the real-time health status data contains an abnormality in the specified monitoring process and / or the specified online collection process, a takeover target is determined in the edge gateway cluster to process the collection tasks of the specified monitoring process and / or the specified online collection process through the takeover target to realize data collection of the Internet of Things devices.

[0045] Determining a takeover target in the edge gateway cluster specifically includes: when an abnormality occurs in the designated online collection process, screening multiple other online collection processes in the server through the monitoring process to determine a takeover collection process; processing the collection task of the designated online collection process through the takeover collection process, and monitoring the health status of the designated online collection process in real time; when the health status of the designated online collection process returns to normal, restoring the collection task to the designated online collection process.

[0046] In one embodiment of the present specification, in the IoT edge gateway data acquisition system, when an abnormality occurs in a designated online acquisition process, the monitoring process plays a key role. The monitoring process filters out the takeover acquisition process from multiple other online acquisition processes on the same server. The screening process here can be achieved through the load balancing algorithm in step S102 combined with the current load conditions of other acquisition processes, with the aim of selecting the process that is most suitable for taking over the abnormal acquisition process task. After determining to take over the acquisition process, the acquisition task of the abnormal collection process is immediately executed to ensure that the data collection work is not interrupted. At the same time, in the process of processing the acquisition task of the designated online collection process by taking over the acquisition process, the health status of the designated online collection process is monitored in real time by the monitoring process. Once" Offline ” The status collection process is restored online, and the monitoring process reallocates the previously transferred tasks back to the process.

[0047] Through the above technical solution, when an abnormality occurs in the collection process, the takeover process is quickly found to continue the collection task, avoiding data collection interruption due to process abnormality, ensuring that the data generated by the equipment can be continuously collected, and ensuring that related business processes are not affected, such as industrial production monitoring, intelligent environment monitoring and other scenarios. Data will not be lost; by real-time monitoring of the health status of the abnormal collection process and restoring the task when it returns to normal, the system resources are fully utilized. Even if some processes have temporary failures, they can maintain stable operation, thereby enhancing the reliability and fault tolerance of the entire edge gateway data collection system and reducing the risk of serious impact on the system due to failure of a single collection process.

[0048] Determining a takeover target in the edge gateway cluster specifically includes: when an abnormality occurs in the designated monitoring process, terminating a plurality of designated collection tasks corresponding to the designated monitoring process, and selecting a target server with the smallest load in the edge gateway cluster; determining the takeover target with the target server, and taking over the collection tasks under the designated monitoring process.

[0049] In one embodiment of the present specification, in the architecture of data collection of the edge gateway of the Internet of Things, when an abnormality is detected in the designated monitoring process, a series of operations will be performed in sequence to ensure the continuity of data collection. First, the multiple collection tasks that the designated monitoring process is responsible for scheduling and managing are immediately terminated. This is to avoid confusion or erroneous execution of the collection tasks when the monitoring process is abnormal, and to prevent errors or incomplete data collection. Screening is performed in the entire edge gateway cluster to find the target server with the smallest load. Here, it is necessary to obtain the load information of each server in the cluster in real time, and determine the server with the lightest load as the target through comparison. Finally, the screened target server is determined as the takeover target to take over the collection tasks that were originally responsible for the abnormal monitoring process. Even if the monitoring process is abnormal, the data collection task can be quickly transferred to other servers for continued execution. Once " Offline ” The monitoring process of the status returns to normal, and the previously transferred tasks will be reallocated back to the collection process under the original monitoring process.

[0050] Through the above technical solution, the collection task is terminated in time when the monitoring process is abnormal, so as to avoid the collection task running under wrong scheduling or management due to the failure of the monitoring process, prevent the collection of wrong, repeated or incomplete data, and ensure the quality of the data; by selecting the target server with the smallest load in the edge gateway cluster to take over the task, it can effectively deal with the sudden situation of abnormal monitoring process. Even if some monitoring processes fail, the whole system can still continue to run, and the data collection work will not be completely stagnant due to the problem of a certain monitoring process; selecting the server with the smallest load as the takeover target can ensure the smooth progress of the data collection task while minimizing the additional consumption of the resources of the entire cluster, avoiding the assignment of tasks to servers with too high load, resulting in server performance degradation or even new failures, realizing the reasonable allocation and efficient utilization of resources, and improving the operation efficiency of the entire edge gateway cluster.

[0051] Through the technical solutions of the embodiments of this specification, the traditional centralized data processing mode is limited by bandwidth and delay, and is unable to cope with large-scale data collection. The embodiments of this specification can reasonably distribute the collection tasks to multiple online collection processes according to the real-time load status of the collection process, such as CPU occupancy, memory occupancy and network bandwidth utilization, through a preset dynamic load balancing algorithm, effectively solving the problem of insufficient processing capacity of a single node, avoiding low collection efficiency caused by excessive data volume, greatly improving the speed of data collection, and being able to quickly process massive concurrent data to meet the requirements of modern Internet of Things for data collection efficiency; in the traditional mode, a single point abnormality of the central node will interrupt data collection, and the traditional distributed architecture relies on a single heartbeat detection that is prone to misjudgment and omission. The embodiments of this specification accurately identify abnormal processes through comprehensive health status detection of the monitoring process and the online collection process. Once a specified monitoring process or a specified online collection process is found to be abnormal, the takeover target can be quickly determined in the edge gateway cluster to automatically complete the task takeover. , no human intervention is required, which greatly reduces the risk of data collection interruption and ensures the stability and reliability of the system; the traditional data collection method relies on manual task migration under abnormal circumstances, and the takeover delay is as high as minutes, which cannot meet the high real-time requirements. The embodiment of this specification can monitor the process status in real time. When the real-time health status data shows an abnormality, it can quickly determine the takeover target and complete the task migration, and achieve seamless connection of tasks with extremely short delays. For example, in scenarios with extremely high real-time requirements such as industrial automation and intelligent transportation, the timeliness of data collection can be ensured; the scale of IoT devices and data traffic are constantly changing dynamically, and traditional models are difficult to adapt. The embodiment of this specification is based on the edge gateway cluster architecture. Whether it is the increase in data volume brought by the access of new devices, or the device offline, process abnormalities, etc., it can respond flexibly. The dynamic load balancing algorithm can adjust task allocation at any time, automatically detect and process abnormal processes, and realize self-optimization and adjustment of the system. It can run stably in a complex and changeable IoT environment and has strong adaptability and scalability.

[0052] The embodiment of this specification also provides a data acquisition device for an edge gateway of the Internet of Things, such as Figure 5 As shown, the device includes: at least one processor; and a memory communicatively connected to the at least one processor; wherein the memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor so that the at least one processor can execute the above method.

[0053] The embodiments of the present specification also provide a non-volatile computer storage medium storing computer executable instructions, wherein the computer executable instructions are configured to execute the above method.

[0054] Each embodiment in this specification is described in a progressive manner, and the same or similar parts between the embodiments can be referred to each other, and each embodiment focuses on the differences from other embodiments. In particular, for the device, equipment, and non-volatile computer storage medium embodiments, since they are basically similar to the method embodiments, the description is relatively simple, and the relevant parts can be referred to the partial description of the method embodiment.

[0055] The above is a description of a specific embodiment of the specification. Other embodiments are within the scope of the appended claims. In some cases, the actions or steps recorded in the claims can be performed in an order different from that in the embodiments and still achieve the desired results. In addition, the processes depicted in the drawings do not necessarily require the specific order or continuous order shown to achieve the desired results. In some embodiments, multitasking and parallel processing are also possible or may be advantageous.

[0056] The devices and media provided in the embodiments of this specification correspond one-to-one to the methods. Therefore, the devices and media also have similar beneficial technical effects as the corresponding methods. Since the beneficial technical effects of the methods have been described in detail above, the beneficial technical effects of the devices and media will not be repeated here.

[0057] Those skilled in the art will appreciate that the embodiments of this specification may be provided as methods, systems, or computer program products. Therefore, this specification may take the form of a complete hardware embodiment, a complete software embodiment, or an embodiment combining software and hardware. Moreover, this specification may take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) that contain computer-usable program code.

[0058] This specification is described with reference to the flowcharts and / or block diagrams of the methods, devices (systems), and computer program products according to the embodiments of this specification. It should be understood that each process and / or box in the flowchart and / or block diagram, as well as the combination of the processes and / or boxes in the flowchart and / or block diagram, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing device to produce a machine, so that the instructions executed by the processor of the computer or other programmable data processing device generate instructions for implementing the processes in the flowchart and / or block diagram. Figure 1 A process or multiple processes and / or boxes Figure 1 A device that provides the functions specified in a block or multiple blocks.

[0059] These computer program instructions can also be stored in a computer-readable memory that can direct a computer or other programmable data processing device to work in a specific manner, such that the instructions stored in the computer-readable memory produce a manufacture including an instruction device that implements the functions specified in one process Figure 1 one process or multiple processes and / or blocks Figure 1 or the functions specified in one block or multiple blocks.

[0060] These computer program instructions can also be loaded onto a computer or other programmable data processing device, such that a series of operational steps are performed on the computer or other programmable device to produce a computer-implemented process, thereby providing the steps for implementing the functions specified in one process Figure 1 one process or multiple processes and / or blocks Figure 1 or the functions specified in one block or multiple blocks on the computer or other programmable device.

[0061] In a typical configuration, a computing device includes one or more processors (CPUs), an input / output interface, a network interface, and memory.

[0062] The memory may include non-permanent memory in the form of computer-readable media, random access memory (RAM) and / or non-volatile memory such as read-only memory (ROM) or flash memory (flash RAM). Memory is an example of computer-readable media.

[0063] Computer-readable media includes permanent and non-permanent, removable and non-removable media that can store information by any method or technology. The information can be computer-readable instructions, data structures, program modules, or other data. Examples of computer storage media include, but are not limited to, phase change memory (PRAM), static random access memory (SRAM), dynamic random access memory (DRAM), other types of random access memory (RAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory or other memory technologies, compact disc read-only memory (CD-ROM), digital versatile disc (DVD) or other optical storage, magnetic cassettes, magnetic tape magnetic disk storage or other magnetic storage devices, or any other non-transitory media that can be used to store information that can be accessed by a computing device. As defined herein, computer-readable media does not include transitory computer-readable media such as modulated data signals and carrier waves.

[0064] It should also be noted that the terms " include ” , " contain ”or any other variation thereof is intended to cover a non-exclusive inclusion such that a process, method, article, or apparatus that includes a list of elements includes not only those elements but also other elements not expressly listed or that are inherent to such process, method, article, or apparatus. In the absence of more limitations, the statement " Includes a ……” The defined elements do not exclude the existence of other identical elements in the process, method, product or equipment including the elements.

[0065] The above description is only one or more embodiments of this specification and is not intended to limit this specification. For those skilled in the art, one or more embodiments of this specification may have various changes and variations. Any modification, equivalent replacement, improvement, etc. made within the spirit and principle of one or more embodiments of this specification shall be included in the scope of the claims of this specification.

Claims

1. A data collection method for an edge gateway of the Internet of Things, characterized in that: The method comprises: Determine at least one monitoring process and corresponding multiple collection processes pre-deployed in each server in the edge gateway cluster; Acquire a pre-configured collection task, and distribute the collection task among the multiple collection processes according to a preset dynamic load balancing algorithm through the monitoring process to determine multiple online collection processes corresponding to the collection task; Executing the collection task through the multiple online collection processes, and detecting the health status between the monitoring process and the multiple online collection processes to determine real-time health status data; When the real-time health status data contains an abnormality in a specified monitoring process and / or a specified online collection process, a takeover target is determined in the edge gateway cluster to process the collection tasks of the specified monitoring process and / or the specified online collection process through the takeover target to realize data collection of the Internet of Things devices.

2. The data collection method of an Internet of Things edge gateway according to claim 1 is characterized in that: The monitoring process allocates the collection task among the multiple collection processes according to a preset dynamic load balancing algorithm to determine multiple online collection processes corresponding to the collection task, specifically including: Acquire real-time load indicator data of the collection process, wherein the real-time load indicator data includes CPU occupancy rate, memory occupancy rate and network bandwidth utilization rate; Determine the comprehensive load coefficient corresponding to the collection process according to the real-time load indicator data, and determine the dynamic weight of each collection process based on the comprehensive load coefficient; Mapping the number of virtual slots is performed through the dynamic weight of the acquisition process to determine the number of virtual slots corresponding to each acquisition process; According to the number of virtual slots corresponding to each of the collection processes, weighted round-robin allocation is performed among the multiple collection processes to determine the multiple online collection processes corresponding to the collection task.

3. The data collection method of an Internet of Things edge gateway according to claim 1 is characterized in that: Detecting the health status between the monitoring process and the multiple online acquisition processes to determine real-time health status data specifically includes: By exchanging information between multiple monitoring processes in the gateway cluster, according to a preset majority decision mechanism, a health status detection is performed on the monitoring process to determine first health status data corresponding to the monitoring process; The health status of the multiple online acquisition processes is detected through the monitoring process to determine the second health status data corresponding to the online acquisition process.

4. The data collection method of an Internet of Things edge gateway according to claim 3 is characterized in that: Through information interaction between multiple monitoring processes in the gateway cluster, according to a preset majority decision mechanism, health status detection is performed on the monitoring process to determine first health status data corresponding to the monitoring process, specifically including: According to a preset period, periodically collect first real-time resource data corresponding to the monitoring process, and broadcast it to other monitoring processes in the gateway cluster, wherein the first real-time resource data includes real-time CPU data, real-time memory data, and real-time thread number; Performing hash calculation on the first real-time resource data through other monitoring processes in the gateway cluster to determine the real-time resource fingerprint data corresponding to the monitoring process; The health status of the monitoring process is detected according to the real-time resource fingerprint data and the majority decision mechanism to determine the first health status data corresponding to the monitoring process.

5. The data collection method of an Internet of Things edge gateway according to claim 3 is characterized in that: Detecting the health status of the multiple online acquisition processes through the monitoring process to determine the second health status data corresponding to the online acquisition process specifically includes: Sending a first number of ping instructions to the online acquisition process through the monitoring process; When the pong instruction returned by the online acquisition process is not received for a first preset number of times, acquiring real-time resource data of the online acquisition process; Performing secondary status detection on the online collection process according to the real-time resource data and the preset historical baseline data; When the real-time resource data and the historical baseline data meet a preset condition, a simulation collection instruction is sent to determine the execution result data of the online collection process, so as to determine the second health status data corresponding to the online collection process based on the execution result data.

6. The data collection method of an Internet of Things edge gateway according to claim 4 is characterized in that: Performing a health status detection on the monitoring process according to the real-time resource fingerprint data and the majority decision mechanism to determine first health status data corresponding to the monitoring process specifically includes: If the deviation between the real-time resource fingerprint data of the monitoring process received by any other monitoring process and the first fingerprint of the preset local cluster average value is not greater than a preset threshold, the monitoring process is judged to be in a healthy state; If the first fingerprint deviation exceeds the threshold value N times continuously, collecting the resource fingerprint sequence of the monitoring process in the historical time period to predict the predicted resource fingerprint in the next time period; When the second fingerprint deviation between the predicted resource fingerprint and the real-time resource fingerprint data exceeds a preset threshold, the monitoring process is determined to be a monitoring process to be confirmed, a voting request is initiated to the gateway cluster, and the first health status data corresponding to the monitoring process is determined according to the voting result.

7. The data collection method of an Internet of Things edge gateway according to claim 1 is characterized in that: Determining a takeover target in the edge gateway cluster specifically includes: When the designated online acquisition process is abnormal, the monitoring process is used to screen multiple other online acquisition processes in the server to determine a takeover acquisition process; Processing the collection tasks of the designated online collection process through the takeover collection process, and monitoring the health status of the designated online collection process in real time; When the health status of the designated online acquisition process returns to normal, the acquisition task is restored to the designated online acquisition process.

8. The data collection method of an Internet of Things edge gateway according to claim 1 is characterized in that: Determining a takeover target in the edge gateway cluster specifically includes: When the designated monitoring process is abnormal, the collection tasks of the designated collection corresponding to the designated monitoring process are terminated, and the target server with the smallest load is selected in the edge gateway cluster; The takeover target is determined by the target server, and the collection task under the designated monitoring process is taken over.

9. A data acquisition device for an edge gateway of the Internet of Things, characterized in that: The device comprises: at least one processor; and, a memory communicatively connected to the at least one processor; wherein, The memory stores instructions that can be executed by the at least one processor, and the instructions are executed by the at least one processor so that the at least one processor can perform the method according to any one of claims 1 to 8.

10. A non-volatile computer storage medium storing computer executable instructions, characterized in that: The computer executable instructions are configured to execute the method according to any one of claims 1 to 8.

Citation Information

Patent Citations

  • Data acquisition method and system thereof

    CN106936618A

  • Health monitoring method and device for server cluster

    CN116723196A

  • Method and system for cluster dynamic balance expansion

    CN117950858A

  • IT centralized monitoring management system based on SAAS

    CN118075091A

  • VPP-based load balancing method, system and device supporting QUIC multipath transmission, and medium

    CN119383128A

Cited By

  • Information data security sharing method based on network security

    CN122119954A