Data acquisition method and system, edge server and storage medium

By introducing multiple edge node groups and dynamic packet management into the data acquisition system, the problems of high failure risk and low acquisition efficiency of existing data acquisition methods are solved, and efficient and reliable data acquisition is achieved.

CN120017677AInactive Publication Date: 2025-05-16SHENZHEN FANHE TECH CO LTD
View PDF 6 Cites 0 Cited by

Patent Information

Application Number
CN202510474685.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-16
Publication Date
2025-05-16
Estimated Expiration
Not applicable · inactive patent

AI Technical Summary

Technical Problem

The existing data acquisition methods have problems such as high failure risk, low acquisition efficiency and long delay, especially in high concurrency and high real-time scenarios.

Method used

By introducing multiple edge node groups into the data acquisition system, each group contains multiple edge nodes. The edge server regularly obtains sensor and edge node information, dynamically determines the target edge node of the sensor, and processes data acquisition tasks and abnormal data in real time.

Benefits of technology

Dynamic grouping management of sensor clusters is realized, the load pressure of a single node is reduced, the risk of acquisition interruption caused by a single point of failure is avoided, and the efficiency and reliability of data acquisition are improved.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120017677A_ABST
    Figure CN120017677A_ABST
Patent Text Reader

Abstract

The invention discloses a data acquisition method and system, an edge server and a storage medium. The data acquisition method is applied to the data acquisition system, the data acquisition system comprises an edge server and at least one edge node group, the edge node group comprises a plurality of edge nodes, and the data acquisition method comprises the following steps: the edge server regularly acquires sensor information and edge node information corresponding to each edge node group; determining an initial edge node of each sensor according to the sensor information and the edge node information; performing statistical analysis on the initial edge node of each sensor, and determining a target edge node of each sensor according to a statistical analysis result; and sending the data acquisition task to the target edge node, and receiving data information sent by the target edge node. The data acquisition efficiency and reliability can be improved, and the risk of acquisition interruption caused by a single-point fault is avoided.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of data processing technology, and in particular to a data acquisition method, system, edge server and storage medium. Background Art

[0002] With the accelerated advancement of smart park construction, the number of Internet of Things (IoT) devices is growing exponentially, and the data collection system needs to cope with the high concurrent access and real-time processing requirements of massive sensors. At present, the industry generally adopts a centralized edge collection method, that is, a single edge server aggregates the data of devices in the park and uploads it to the cloud management layer for business use. However, this model has exposed the following significant defects in actual applications: First, the risk of single-point collection failure is prominent. Traditional centralized edge collection is prone to data collection interruption and data loss risks due to network fluctuations or service interruptions. Second, the transmission efficiency and real-time performance are insufficient. Data is transmitted through a single collection channel, which is limited by bandwidth bottlenecks and protocol conversion delays, resulting in high response delays, making it difficult to meet the millisecond-level response requirements of high-real-time scenarios such as security monitoring and equipment control. Summary of the invention

[0003] Based on this, the main purpose of the present invention is to provide a data collection method, system, edge server and storage medium, aiming to solve the problems of high failure risk, low collection efficiency and long delay in existing data collection methods.

[0004] In a first aspect, the present invention provides a data collection method, which is applied to a data collection system, wherein the data collection system includes an edge server and at least one edge node group, wherein the edge node group includes a plurality of edge nodes, and the data collection method includes the following steps: The edge server periodically obtains sensor information and edge node information corresponding to each edge node group; Determine the initial edge node of each sensor according to the sensor information and the edge node information; Performing statistical analysis on the initial edge nodes of each sensor, and determining the target edge nodes of each sensor according to the statistical analysis results; Send a data collection task to the target edge node, and receive data information sent by the target edge node.

[0005] In one embodiment, the sensor information includes the installation location of the sensor, the data generation frequency and the network status, and the edge node information includes the installation location of the edge node. The step of determining the initial edge node of each sensor according to the sensor information and the edge node information comprises: Calculating the transmission distance between the sensor and the edge node according to the installation position of the sensor and the installation position of the edge node; Calculating the transmission time according to the transmission distance, data generation frequency and network status; An initial edge node of each sensor is determined according to the transmission time.

[0006] In one embodiment, the edge node information includes a load threshold of the edge node, and the step of performing statistical analysis on the initial edge nodes of each sensor and determining the target edge node of each sensor according to the statistical analysis result includes: Counting the initial edge nodes of each sensor to obtain the number of sensors assigned to each edge node; Detecting whether the assigned number of sensors exceeds the load threshold; If so, the initial edge node is adjusted according to the detection result and the transmission time to obtain the target edge node of each sensor.

[0007] In one embodiment, the data collection method further includes: Obtain the operating status information of each edge node in real time; Determine whether each edge node is abnormal according to the operation status information; When it is determined that at least one edge node is abnormal, the edge node with the abnormality is switched.

[0008] In one embodiment, after the step of sending a data collection task to the target edge node and receiving data information sent by the target edge node, the data collection method further includes: Identify the data information using a pre-trained data anomaly identification model to determine whether the data information is abnormal and the type of abnormality; If an exception exists, a target exception handling method is determined according to the exception type, and the data information with the exception is processed by the target exception handling method.

[0009] In one embodiment, the step of determining a target exception handling method according to the exception type and processing the data information with the exception by using the target exception handling method includes: If it is determined that the abnormal type is data missing, the data information with the abnormality is processed by a K-nearest neighbor algorithm; and / or, If it is determined that the abnormal type is data duplication, the data information with the abnormality is processed by a hash algorithm; and / or, If it is determined that the abnormal type is a data abnormality, the abnormal data information is processed by a regression model or a smoothing mechanism algorithm; and / or, If it is determined that the abnormal type is a data unit abnormality, the data information with the abnormality is processed through natural language processing.

[0010] In one embodiment, the data collection system further includes a cloud management layer, and after the step of processing the abnormal data information by the target abnormality processing method, the system further includes: The processed data information is sent to the cloud management layer.

[0011] In a second aspect, the present invention further provides a data collection system, the data collection system comprising an edge server and at least one edge node group, the edge node group comprising a plurality of edge nodes, the edge server being used to: Regularly obtain sensor information and edge node information corresponding to each edge node group; Determine the initial edge node of each sensor according to the sensor information and the edge node information; Performing statistical analysis on the initial edge nodes of each sensor, and determining the target edge nodes of each sensor according to the statistical analysis results; Send a data collection task to the target edge node, and receive data information sent by the target edge node.

[0012] In one embodiment, the sensor information includes the installation location of the sensor, the data generation frequency and the network status, and the edge node information includes the installation location of the edge node. The edge server is specifically used for: Calculating the transmission distance between the sensor and the edge node according to the installation position of the sensor and the installation position of the edge node; Calculating the transmission time according to the transmission distance, data generation frequency and network status; An initial edge node of each sensor is determined according to the transmission time.

[0013] In one embodiment, the edge node information includes a load threshold of the edge node, and the edge server is specifically used to: Counting the initial edge nodes of each sensor to obtain the number of sensors assigned to each edge node; Detecting whether the assigned number of sensors exceeds the load threshold; If so, the initial edge node is adjusted according to the detection result and the transmission time to obtain the target edge node of each sensor.

[0014] In one embodiment, the edge server is further configured to: Obtain the operating status information of each edge node in real time; Determine whether each edge node is abnormal according to the operation status information; When it is determined that at least one edge node is abnormal, the edge node with the abnormality is switched.

[0015] In one embodiment, the edge server is further configured to: Identify the data information using a pre-trained data anomaly identification model to determine whether the data information is abnormal and the type of abnormality; If an exception exists, a target exception handling method is determined according to the exception type, and the data information with the exception is processed by the target exception handling method.

[0016] In one embodiment, the edge server is specifically used for: If it is determined that the abnormal type is data missing, the data information with the abnormality is processed by a K-nearest neighbor algorithm; and / or, If it is determined that the abnormal type is data duplication, the data information with the abnormality is processed by a hash algorithm; and / or, If it is determined that the abnormal type is a data abnormality, the abnormal data information is processed by a regression model or a smoothing mechanism algorithm; and / or, If it is determined that the abnormal type is a data unit abnormality, the data information with the abnormality is processed through natural language processing.

[0017] In a third aspect, the present invention further provides an edge server, comprising: a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the computer program implements the steps of the data collection method in the first aspect when executed by the processor.

[0018] In a fourth aspect, the present invention further provides a computer-readable storage medium having a computer program stored thereon, and when the computer program is executed by a processor, the steps of the data acquisition method in the first aspect are implemented.

[0019] The present invention provides a data acquisition method, system, edge server and storage medium. The data acquisition method is applied to a data acquisition system. The data acquisition system includes an edge server and at least one edge node group. The edge node group includes multiple edge nodes. The edge server periodically obtains sensor information and edge node information corresponding to each edge node group; then, the initial edge node of each sensor is determined according to the sensor information and edge node information; then, the initial edge node of each sensor is statistically analyzed, and the initial edge node of each sensor is reallocated and adjusted according to the statistical analysis result to determine the target edge node of each sensor; finally, a data acquisition task is sent to the target edge node, so that the target edge node only collects the data information of the sensor assigned to it, and then receives the data information sent by the target edge node. The present invention improves the traditional data acquisition system and sets multiple edge nodes. Even if one of the edge nodes fails, the data information of the sensor it is responsible for collecting can be transferred to other edge nodes in the same group for collection, thereby avoiding the risk of collection interruption caused by a single point failure. At the same time, by periodically obtaining sensor information and edge node information, it is used to determine and adjust the target edge node corresponding to each sensor, thereby realizing dynamic grouping management of sensor clusters and reducing the load pressure of a single node. In summary, this application constructs a new type of data acquisition system, which solves the inherent defects of traditional centralized edge acquisition systems through dynamic edge node collaboration and load balancing, and improves the data acquisition efficiency and reliability in high-density sensor scenarios. BRIEF DESCRIPTION OF THE DRAWINGS

[0020] Figure 1 It is a structural diagram of a traditional data acquisition system; Figure 2 A schematic diagram of the structure of a data acquisition system involved in an embodiment of the present invention; Figure 3 A schematic diagram of the structure of an edge server in a hardware operating environment involved in an embodiment of the present invention; Figure 4 Schematic diagram of the flow of the first embodiment of the data collection method of the present invention.

[0021] The realization of the purpose, functional features and advantages of the present invention will be further explained in conjunction with embodiments and with reference to the accompanying drawings. DETAILED DESCRIPTION

[0022] In order to make the purpose, technical solution and advantages of the present invention more clearly understood, the present invention is further described in detail below in conjunction with the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present invention and are not intended to limit the present invention.

[0023] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by technicians in the technical field to which the present invention belongs; the terms used herein are only for the purpose of describing specific embodiments and are not intended to limit the present invention; the terms "including" and "having" in the specification and claims of the present invention and the above-mentioned figure descriptions and any variations thereof are intended to cover non-exclusive inclusions.

[0024] In the description of the embodiments of the present invention, the technical terms "first", "second", etc. are only used to distinguish different objects, and cannot be understood as indicating or implying relative importance or implicitly indicating the number, specific order or primary and secondary relationship of the indicated technical features. In the description of the embodiments of the present invention, the meaning of "multiple" is more than two, unless otherwise clearly and specifically defined.

[0025] Reference to "embodiments" herein means that a particular feature, structure, or characteristic described in conjunction with the embodiments may be included in at least one embodiment of the present invention. The appearance of the phrase in various places in the specification does not necessarily refer to the same embodiment, nor is it an independent or alternative embodiment that is mutually exclusive with other embodiments. It is explicitly and implicitly understood by those skilled in the art that the embodiments described herein may be combined with other embodiments.

[0026] With the accelerated advancement of smart park construction, the number of Internet of Things (IoT) devices is growing exponentially, and data collection systems need to cope with the high concurrent access and real-time processing requirements of massive sensors. At present, the industry generally adopts a centralized edge collection method, that is, a single edge server aggregates the data of devices in the park and uploads it to the cloud management layer for business use. For example, taking the smart park scenario as an example, Figure 1 This is a structural diagram of the traditional data acquisition system. Various equipment in the park will produce massive amounts of data information, such as fire protection system equipment, power system equipment, air conditioning system equipment, etc. The data information generated will be sent directly to the edge server through the sensors of each device, and then uploaded to the cloud management layer for processing. This data acquisition mode has exposed the following significant defects in actual applications: First, the risk of single-point acquisition failure is prominent. Traditional centralized edge acquisition is prone to data collection interruption and data loss risks due to network fluctuations or service interruptions. Second, the transmission efficiency and real-time performance are insufficient. Data is transmitted through a single acquisition channel, which is limited by bandwidth bottlenecks and protocol conversion delays, resulting in high response delays, making it difficult to meet the millisecond-level response requirements of high-real-time scenarios such as security monitoring and equipment control.

[0027] In addition, in the above data collection mode, the collected data information is not processed, and usually needs to rely on manual calibration or fixed threshold filtering. This data processing method is difficult to deal with the problem of abnormal collected data caused by interference from complex factors in the collection link.

[0028] After introducing the background technology of the data collection method provided by the embodiment of the present invention, the implementation environment involved in the data collection method provided by the embodiment of the present invention will be briefly described below. The data collection method provided by the embodiment of the present invention can be applied to Figure 2 In the distributed data acquisition system shown. The distributed data acquisition system includes an edge server and at least one edge node group, each edge node group includes a plurality of edge nodes, each edge node is communicatively connected to the edge server, and can communicate with each other, transmit data, and exchange information, etc. Among them, the edge server refers to a server that places computing resources and services at the location closest to users and devices. It is located at the edge of the network, close to users, and is usually deployed on physical or virtual devices, such as network border routers, switches, load balancers, etc. The number of edge node groups can be set according to parameters such as the number of devices in the park and the amount of data information generated. Similarly, the number of edge nodes in each edge node group can be set according to parameters such as the number of devices corresponding to the edge node group and the amount of data information generated by the devices in the group. For example Figure 2 In the smart park scenario, one edge node group can be set up for every five devices, and two edge nodes can be set in each edge node group.

[0029] The data collection system may also include multiple data generating devices. For example, in a smart park scenario, the data generating devices may include but are not limited to: fire protection system equipment, security system equipment, power system equipment, access control system equipment, elevator system equipment, parking lot system equipment, meter reading system equipment, lighting system equipment, air conditioning system equipment, etc. Each data generating device may communicate, transmit data, and exchange information with the edge nodes in the corresponding edge node group. Specifically, each data generating device may send data information to the edge node through a sensor.

[0030] The data collection system may also include a cloud management layer, which is communicatively connected to the edge server and can receive data information sent by the edge server.

[0031] Those skilled in the art will understand that Figure 2 The structure shown in the figure is only a block diagram of a part of the structure related to the scheme of the present invention, and does not constitute a limitation on the data acquisition system to which the scheme of the present invention is applied. The specific distributed data acquisition system may include more or fewer components than those shown in the figure, or combine certain components, or have a different arrangement of components.

[0032] Further, refer to Figure 3 , Figure 3 The schematic diagram is a structural diagram of an edge server of a hardware operating environment involved in an embodiment of the present invention.

[0033] like Figure 3 As shown, the edge server may include: a processor 1001, such as a CPU, a communication bus 1002, a user interface 1003, a network interface 1004, and a memory 1005. Among them, the communication bus 1002 is used to realize the connection and communication between these components. The user interface 1003 may include a display screen (Display), an input unit such as a keyboard (Keyboard), and the optional user interface 1003 may also include a standard wired interface and a wireless interface. The network interface 1004 may optionally include a standard wired interface and a wireless interface (such as a Wi-Fi interface). The memory 1005 may be a high-speed RAM memory, or a stable memory (non-volatile memory), such as a disk memory. The memory 1005 may also be a storage device independent of the aforementioned processor 1001.

[0034] Those skilled in the art will understand that Figure 3 The structure shown in the figure does not constitute a limitation on the edge server of the present invention, and may include more or less components than those shown in the figure, or combine certain components, or arrange the components differently.

[0035] like Figure 3 As shown, the memory 1005 as a computer storage medium may include an operating system, a network communication module, a user interface module and a computer program.

[0036] exist Figure 3 In the edge server shown, the network interface 1004 is mainly used to connect to the cloud management layer and / or the edge node, and to communicate data with the cloud management layer and / or the edge node; the user interface 1003 is mainly used to connect to the data generating device, and to communicate data with the data generating device; and the processor 1001 can be used to call the computer program stored in the memory 1005 and execute various embodiments of the data acquisition method of the present invention.

[0037] Based on the above, various embodiments of the data collection method of the present invention are proposed.

[0038] Reference Figure 4 , Figure 4 Schematic diagram of the flow of the first embodiment of the data collection method of the present invention.

[0039] In this embodiment, the data collection method is applied to a data collection system, the data collection system includes an edge server and at least one edge node group, the edge node group includes a plurality of edge nodes, and the data collection method includes: Step S10, the edge server periodically obtains sensor information and edge node information corresponding to each edge node group; The data collection method of this embodiment is applied to a data collection system, which includes an edge server and at least one edge node group, wherein the edge node group includes multiple edge nodes, and each edge node is connected to the edge server in communication. The data collection method of this embodiment is implemented by an edge server, which refers to a server that places computing resources and services at the location closest to users and devices. It is located at the edge of the network, close to users, and is usually deployed on physical or virtual devices, such as network border routers, switches, load balancers, etc. The number of edge node groups can be set according to parameters such as the number of devices in the park and the amount of data information generated. Similarly, the number of edge nodes in each edge node group can be set according to parameters such as the number of devices corresponding to the edge node group and the amount of data information generated by the devices in the group. By setting multiple edge node groups and assigning data generating devices of different systems to different edge node groups, resource contention can be avoided, and it is conducive to subsequent dynamic grouping management.

[0040] In this embodiment, the method is applied to Figure 2 The data collection system in the smart park scenario in is used as an example to illustrate. An edge node group is set for every 5 data generating devices, and 2 edge nodes are set in each edge node group. That is, the fire protection system equipment, security system equipment, power system equipment, access control system equipment and elevator system equipment are assigned to the edge nodes of group A, and the data information of these data generating devices can be sent to edge nodes A-1 or A-2 according to the allocation strategy; the parking lot system equipment, meter reading system equipment, lighting system equipment, and air conditioning system equipment are assigned to the edge nodes of group B, and the data information of these data generating devices can be sent to edge nodes B-1 or B-2 according to the allocation strategy.

[0041] In this embodiment, the traditional data acquisition system is improved and multiple edge nodes are set up. Even if one of the edge nodes fails, the data information of the sensors of the data generating device responsible for collecting it can be transferred to other edge nodes in the same group for collection, thereby avoiding the risk of collection interruption due to single point failure.

[0042] In this embodiment, the edge server periodically obtains the sensor information and edge node information corresponding to each edge node group, wherein the sensor information includes the installation location, data generation frequency and network status of the sensor, and the edge node information includes the installation location and load threshold of the edge node. The installation location refers to the physical location, which can be represented by coordinates; the data generation frequency refers to the data acquisition and collection frequency of the data generating device; the network status refers to the transmission network status between the sensor and each edge node, for example, it can be divided into smooth, poor, disconnected, etc.; the load threshold refers to the maximum load capacity that the edge node can bear, for example, the maximum load capacity of an edge node is to receive 50 sensor transmission data at the same time, then the load threshold is defined as 50.

[0043] Step S20, determining the initial edge node of each sensor according to the sensor information and edge node information; After obtaining the sensor information and edge node information, the initial edge node of each sensor is determined according to the sensor information and edge node information. Specifically, the transmission distance between the sensor and the edge node can be calculated according to the installation position of the sensor and the installation position of the edge node; then, the transmission time is calculated according to the transmission distance, data generation frequency and network status; finally, the initial edge node of each sensor is determined according to the transmission time. The specific execution process can refer to the second embodiment described below, which will not be described in detail this time.

[0044] Step S30, performing statistical analysis on the initial edge nodes of each sensor, and determining the target edge nodes of each sensor according to the statistical analysis results; Since the initial edge nodes obtained by preliminary determination may be overloaded, etc., it is necessary to perform statistical analysis on the initial edge nodes of each sensor, determine whether the initial edge nodes need to be adjusted based on the statistical analysis results, and then determine the target edge nodes of each sensor. The specific execution process can refer to the third embodiment described below, which will not be described in detail this time. Through the above analysis and adjustment, dynamic grouping management of the sensor cluster is further realized, which can reduce the load pressure of a single node.

[0045] Step S40: sending a data collection task to the target edge node, and receiving data information sent by the target edge node.

[0046] After determining the target edge node of each sensor, send a data collection task to the target edge node so that the target edge node only collects data information of the assigned sensor and then receives the data information sent by the target edge node. For example, after determining that the target edge node of the fire protection system equipment and security system equipment is A-1, and the target edge node of the power system equipment, access control system equipment and elevator system equipment is A-2, then send data collection tasks to edge nodes A-1 and A-2 respectively, so that edge node A-1 only receives data information of the fire protection system equipment and security system equipment, and edge node A-2 only receives data information of the power system equipment, access control system equipment and elevator system equipment.

[0047] The embodiment of the present invention provides a data collection method, which is applied to a data collection system. The data collection system includes an edge server and at least one edge node group, wherein the edge node group includes multiple edge nodes, and the edge server periodically obtains sensor information and edge node information corresponding to each edge node group; then, the initial edge node of each sensor is determined according to the sensor information and edge node information; then, the initial edge node of each sensor is statistically analyzed, and the initial edge node of each sensor is reallocated and adjusted according to the statistical analysis result to determine the target edge node of each sensor; finally, a data collection task is sent to the target edge node, so that the target edge node only collects the data information of the sensor assigned to it, and then receives the data information sent by the target edge node. The present invention improves the traditional data collection system and sets multiple edge nodes. Even if one of the edge nodes fails, the data information of the sensor it is responsible for collecting can be transferred to other edge nodes in the same group for collection, thereby avoiding the risk of collection interruption caused by a single point failure. At the same time, by periodically obtaining sensor information and edge node information, it is used to determine and adjust the target edge node corresponding to each sensor, thereby realizing dynamic grouping management of the sensor cluster, which can reduce the load pressure of a single node. In summary, this application constructs a new type of data acquisition system, which solves the inherent defects of traditional centralized edge acquisition systems through dynamic edge node collaboration and load balancing, and improves the data acquisition efficiency and reliability in high-density sensor scenarios.

[0048] Furthermore, based on the above-mentioned first embodiment, a second embodiment of the data collection method of the present invention is proposed.

[0049] In this embodiment, the sensor information includes the installation location of the sensor, the data generation frequency and the network status, and the edge node information includes the installation location of the edge node. The above step S20 includes: Step S21, calculating the transmission distance between the sensor and the edge node according to the installation position of the sensor and the installation position of the edge node; Step S22, calculating the transmission time according to the transmission distance, data generation frequency and network status; Step S23: determining the initial edge node of each sensor according to the transmission time.

[0050] In this embodiment, the process of determining the preliminary edge node corresponding to each sensor is as follows: First, the transmission distance between the sensor and the edge node is calculated based on the installation location of the sensor and the installation location of the edge node. The specific calculation method can refer to the existing technology; then, the transmission time is calculated based on the transmission distance, data generation frequency and network status. The specific calculation rules can be set according to actual needs and are not limited this time. Next, the initial edge node of each sensor is determined based on the transmission time. Specifically, the edge node with the shortest transmission time between the sensor and the edge node is determined as the initial edge node of the sensor.

[0051] Further, the edge node information includes a load threshold of the edge node, and the above step S30 includes: Step S31, counting the initial edge nodes of each sensor to obtain the number of sensors assigned to each edge node; Step S32, detecting whether the number of sensors allocated exceeds the load threshold; Step S33: If yes, the initial edge node is adjusted according to the detection result and the transmission time to obtain the target edge node of each sensor.

[0052] In this embodiment, since the initial edge nodes obtained by preliminary determination may have overload and other situations, it is necessary to count the initial edge nodes of each sensor to obtain the number of sensors assigned to each edge node, and then detect whether the number of sensors assigned exceeds the load threshold, which refers to the maximum load capacity that the edge node can bear; if the number of sensors assigned to the edge node exceeds its load threshold, the initial edge node is adjusted according to the excess number and transmission time, so as to obtain the target edge node of each sensor. For example, the load threshold of edge nodes A-1 and A-2 is 3, but it is preliminarily determined that the number of sensors assigned to A-1 and A-2 is 4 and 1 respectively, then one of the sensors corresponding to edge node A-1 is assigned to edge node A-2, and which sensor to adjust specifically can be selected according to the transmission time between these four sensors and edge nodes A-1 and A-2, for example, the one with the longest transmission time with edge node A-1 can be selected, or the one with the shortest transmission time with edge node A-2 can be selected, and the specific adjustment strategy can be flexibly set, which is not limited this time. Furthermore, if the number of sensors assigned to the edge nodes does not exceed the load threshold, then no adjustment is required, and the initial edge node of each sensor is the target edge node of each sensor.

[0053] Furthermore, based on the above first and second embodiments, a third embodiment of the data acquisition method of the present invention is proposed.

[0054] In this embodiment, the data collection method further includes the following steps: Step A, obtaining the operating status information of each edge node in real time; Step B, judging whether each edge node has an abnormality according to the operation status information; Step C: when it is determined that at least one edge node is abnormal, the edge node with the abnormality is switched.

[0055] In this embodiment, the edge server will obtain the operating status information of each edge node in real time. The operating status information is the information about whether the node is running normally, which is used to determine whether each edge node is available. When it is determined that at least one edge node is abnormal, the edge node with the abnormality is automatically switched, for example, the edge node with the abnormality is eliminated, thereby avoiding the sensor data information from being sent to the edge node with the abnormality.

[0056] In this embodiment, the abnormal edge node can be automatically switched to ensure the stability and reliability of the data collection service provided to the sensor of the data generating device.

[0057] Furthermore, based on the above first and second embodiments, a fourth embodiment of the data acquisition method of the present invention is proposed.

[0058] In this embodiment, after the above step S40, the data collection method further includes: Step S50, identifying the data information through a pre-trained data anomaly recognition model to determine whether the data information is abnormal and the type of abnormality; If there is an exception, step S61 is executed to determine a target exception processing method according to the exception type, and process the data information with the exception by using the target exception processing method.

[0059] In this embodiment, the edge server includes an AI error correction and cleaning module, which can be used to correct or remove abnormal data. Specifically, the data information is first identified through a pre-trained data anomaly recognition model to determine whether the received data information has an anomaly and its anomaly type. Among them, the data anomaly recognition model can adopt a deep learning model, such as a deep neural network (Deep Neural Networks, DNN), a long short-term memory network (Long Short-Term Memory, LSTM), and then the model is trained through a training sample set to obtain a pre-trained data anomaly recognition model.

[0060] If there is an exception, the target exception processing method is determined according to the exception type, and the data information with the exception is processed by the target exception processing method. For example, if the exception type is determined to be data missing, the data information with the exception is processed by the K nearest neighbor algorithm; for another example, if the exception type is determined to be data duplication, the data information with the exception is processed by the hash algorithm; for another example, if the exception type is determined to be data anomaly, the data information with the exception is processed by a regression model or a smoothing mechanism algorithm; for another example, if the exception type is determined to be a data unit anomaly, the data information with the exception is processed by natural language processing. The specific processing process can be referred to the fifth embodiment described below, which will not be repeated this time.

[0061] Furthermore, the data collection system further includes a cloud management layer. After the above step S61, the data collection method further includes: The processed data information is sent to the cloud management layer.

[0062] In this embodiment, after the abnormal data is processed by the AI ​​error correction and cleaning module, the processed data information is sent to the cloud management layer for business use.

[0063] Furthermore, after the above step S50, if there is no abnormality, the step of sending the data information to the cloud processor is executed.

[0064] In this embodiment, the collected data is processed through an AI model, including but not limited to error detection, compensation and adaptive calibration, which can reduce human intervention, improve processing efficiency while ensuring the accuracy and integrity of the data, thereby better serving the upper-level applications in the cloud.

[0065] Furthermore, based on the fourth embodiment described above, a fifth embodiment of the data collection method of the present invention is proposed.

[0066] In this embodiment, the above step S61 includes: Step S611, if it is determined that the abnormal type is data missing, the data information with the abnormality is processed by a K nearest neighbor algorithm; and / or, Step S612: if it is determined that the abnormal type is data duplication, the abnormal data information is processed by a hash algorithm; and / or, Step S613: if it is determined that the abnormal type is a data abnormality, the abnormal data information is processed by a regression model or a smoothing mechanism algorithm; and / or, Step S614: If it is determined that the abnormal type is a data unit abnormality, the abnormal data information is processed through natural language processing.

[0067] In this embodiment, for different exception types, the corresponding processing procedures are as follows: If the data anomaly recognition model detects that there is no data at some time points, that is, the anomaly type is determined to be data missing, the K-Nearest Neighbor (KNN) algorithm is used to process the data information with anomalies. Specifically, the mean or weighted value of the neighboring samples can be used to fill in the outliers to improve data integrity.

[0068] If the data anomaly recognition model detects that the same data was collected at a certain time point, that is, the anomaly type is determined to be data duplication, the hash algorithm is used to deduplicate the data information with the anomaly. Specifically, the hash value of each data item is calculated. Since the same data will generate the same hash value, the duplicate data can be filtered out through query judgment.

[0069] If the abnormal type is determined to be a data anomaly, the data information with the anomaly is processed through a regression model or a smoothing mechanism algorithm. Specifically, if it is a significant data anomaly, for example, the normal data is all in the hundreds, and suddenly a data of 10,000 appears. At this time, a smoothing mechanism algorithm is used for correction. Specifically, the local abnormal fluctuation can be corrected by a moving average or exponential smoothing method based on a time series. If it is a non-significant data anomaly, for example, the data deviates from the threshold range, but it is not an obvious abnormal fluctuation, a regression model is used to correct the abnormal data. Specifically, the reasonable value of the abnormal point can be predicted by linear regression, ridge regression, etc., and the abnormal data can be corrected by feature correlation.

[0070] If the abnormality type is determined to be a data unit abnormality, the abnormal data information is processed through natural language processing (NLP). For example, when collecting temperature, inconsistent unit data appears, such as Celsius, such as "25°C", and Fahrenheit, such as "77°F". NLP processing is used to convert units and achieve data unification.

[0071] In this embodiment, by selecting appropriate processing methods for different exception types, the accuracy and completeness of the processed data information can be guaranteed.

[0072] Based on the methods described in all the above embodiments, the present invention further provides a data acquisition system, which includes an edge server and at least one edge node group, wherein the edge node group includes multiple edge nodes, and the edge server is used to: Regularly obtain sensor information and edge node information corresponding to each edge node group; Determine the initial edge node of each sensor according to the sensor information and the edge node information; Performing statistical analysis on the initial edge nodes of each sensor, and determining the target edge nodes of each sensor according to the statistical analysis results; Send a data collection task to the target edge node, and receive data information sent by the target edge node.

[0073] In one embodiment, the sensor information includes the installation location of the sensor, the data generation frequency and the network status, and the edge node information includes the installation location of the edge node. The edge server is specifically used for: Calculating the transmission distance between the sensor and the edge node according to the installation position of the sensor and the installation position of the edge node; Calculating the transmission time according to the transmission distance, data generation frequency and network status; An initial edge node of each sensor is determined according to the transmission time.

[0074] In one embodiment, the edge node information includes a load threshold of the edge node, and the edge server is specifically used to: Counting the initial edge nodes of each sensor to obtain the number of sensors assigned to each edge node; Detecting whether the assigned number of sensors exceeds the load threshold; If so, the initial edge node is adjusted according to the detection result and the transmission time to obtain the target edge node of each sensor.

[0075] In one embodiment, the edge server is further configured to: Obtain the operating status information of each edge node in real time; Determine whether each edge node is abnormal according to the operation status information; When it is determined that at least one edge node is abnormal, the edge node with the abnormality is switched.

[0076] In one embodiment, the data collection system further includes a cloud management layer, and the edge server is further used for: The data information is sent to the cloud management layer.

[0077] In one embodiment, the edge server is further configured to: Identify the data information using a pre-trained data anomaly identification model to determine whether the data information is abnormal and the type of abnormality; If an exception exists, a target exception handling method is determined according to the exception type, and the data information with the exception is processed by the target exception handling method.

[0078] In one embodiment, the edge server is specifically used for: If it is determined that the abnormal type is data missing, the data information with the abnormality is processed by a K-nearest neighbor algorithm; and / or, If it is determined that the abnormal type is data duplication, the data information with the abnormality is processed by a hash algorithm; and / or, If the abnormal type is determined to be a data abnormality, the abnormal data information is processed by a regression model or a smoothing mechanism algorithm; and / or, If it is determined that the abnormal type is a data unit abnormality, the data information with the abnormality is processed through natural language processing.

[0079] In one embodiment, the edge server is further configured to: The processed data information is sent to the cloud management layer.

[0080] Among them, the functional implementation of each module in the above-mentioned data acquisition system corresponds to each step in the above-mentioned data acquisition method embodiment, and its functions and implementation processes are no longer repeated here.

[0081] The present invention also provides a computer-readable storage medium having a computer program stored thereon. When the computer program is executed by a processor, the steps of the data acquisition method described in any one of the above embodiments are implemented.

[0082] The specific embodiments of the computer-readable storage medium of the present invention are substantially the same as the embodiments of the above-mentioned data acquisition method, and are not described in detail herein.

[0083] Those skilled in the art can understand that all or part of the processes in the above-mentioned embodiments can be completed by instructing the relevant hardware through a computer program, and the computer program can be stored in a non-volatile computer-readable storage medium. When the computer program is executed, it can include the processes of the embodiments of the above-mentioned methods. Among them, any reference to the memory, database or other medium used in the embodiments provided by the present invention can include at least one of non-volatile memory and volatile memory. Non-volatile memory can include read-only memory (ROM), magnetic tape, floppy disk, flash memory, optical memory, high-density embedded non-volatile memory, resistive random access memory (ReRAM), magnetic random access memory (MRAM), ferroelectric random access memory (FRAM), phase change memory (PCM), graphene memory, etc. Volatile memory can include random access memory (RAM) or external cache memory, etc. As an illustration and not limitation, RAM can be in various forms, such as static random access memory (SRAM) or dynamic random access memory (DRAM). The database involved in each embodiment provided by the present invention may include at least one of a relational database and a non-relational database. Non-relational databases may include distributed databases based on blockchains, etc., but are not limited to this. The processor involved in each embodiment provided by the present invention may be a general-purpose processor, a central processing unit, a graphics processor, a digital signal processor, a programmable logic unit, a data processing logic unit based on quantum computing, an artificial intelligence (AI) processor, etc., but are not limited to this.

[0084] The technical features of the above embodiments may be arbitrarily combined. To make the description concise, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, they should be considered to be within the scope of the present invention.

[0085] The above embodiments only express several implementation methods of the present invention, and the descriptions thereof are relatively specific and detailed, but they cannot be understood as limiting the scope of the present invention. It should be pointed out that, for those of ordinary skill in the art, several modifications and improvements can be made without departing from the concept of the present invention, and these all belong to the protection scope of the present invention. Therefore, the protection scope of the present invention shall be subject to the attached claims.

Claims

1. A data collection method, characterized in that: Applied to a data acquisition system, the data acquisition system includes an edge server and at least one edge node group, the edge node group includes a plurality of edge nodes, and the data acquisition method includes the following steps: The edge server periodically obtains sensor information and edge node information corresponding to each edge node group; Determine the initial edge node of each sensor according to the sensor information and the edge node information; Performing statistical analysis on the initial edge nodes of each sensor, and determining the target edge nodes of each sensor according to the statistical analysis results; Send a data collection task to the target edge node, and receive data information sent by the target edge node.

2. The data collection method according to claim 1, characterized in that: The sensor information includes the installation location of the sensor, the data generation frequency and the network status, and the edge node information includes the installation location of the edge node. The step of determining the initial edge node of each sensor according to the sensor information and the edge node information comprises: Calculating the transmission distance between the sensor and the edge node according to the installation position of the sensor and the installation position of the edge node; Calculating the transmission time according to the transmission distance, data generation frequency and network status; An initial edge node of each sensor is determined according to the transmission time.

3. The data collection method according to claim 2, characterized in that: The edge node information includes a load threshold of the edge node, The step of performing statistical analysis on the initial edge nodes of each sensor and determining the target edge nodes of each sensor according to the statistical analysis results comprises: Counting the initial edge nodes of each sensor to obtain the number of sensors assigned to each edge node; Detecting whether the assigned number of sensors exceeds the load threshold; If so, the initial edge node is adjusted according to the detection result and the transmission time to obtain the target edge node of each sensor.

4. The data collection method according to any one of claims 1 to 3, characterized in that: The data collection method further comprises: Obtain the operating status information of each edge node in real time; Determine whether each edge node is abnormal according to the operation status information; When it is determined that at least one edge node is abnormal, the edge node with the abnormality is switched.

5. The data collection method according to any one of claims 1 to 3, characterized in that: After the step of sending the data collection task to the target edge node and receiving the data information sent by the target edge node, the data collection method further includes: Identify the data information using a pre-trained data anomaly identification model to determine whether the data information is abnormal and the type of abnormality; If an exception exists, a target exception handling method is determined according to the exception type, and the data information with the exception is processed by the target exception handling method.

6. The data collection method according to claim 5, characterized in that: The step of determining a target exception handling method according to the exception type and processing the data information with the exception by using the target exception handling method comprises: If it is determined that the abnormal type is data missing, the data information with the abnormality is processed by a K-nearest neighbor algorithm; and / or, If it is determined that the abnormal type is data duplication, the data information with the abnormality is processed by a hash algorithm; and / or, If it is determined that the abnormal type is a data abnormality, the abnormal data information is processed by a regression model or a smoothing mechanism algorithm; and / or, If it is determined that the abnormal type is a data unit abnormality, the data information with the abnormality is processed through natural language processing.

7. The data collection method according to claim 5, characterized in that: The data collection system further includes a cloud management layer, and after the step of processing the abnormal data information by the target abnormality processing method, it also includes: The processed data information is sent to the cloud management layer.

8. A data acquisition system, characterized in that: The data collection system includes an edge server and at least one edge node group, wherein the edge node group includes a plurality of edge nodes, and the edge server is used to: Regularly obtain sensor information and edge node information corresponding to each edge node group; Determine the initial edge node of each sensor according to the sensor information and the edge node information; Performing statistical analysis on the initial edge nodes of each sensor, and determining the target edge nodes of each sensor according to the statistical analysis results; Send a data collection task to the target edge node, and receive data information sent by the target edge node.

9. An edge server, characterized in that: The edge server comprises: a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the computer program implements the steps of the data collection method according to any one of claims 1 to 7 when executed by the processor.

10. A computer-readable storage medium, characterized in that: The computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, the steps of the data acquisition method according to any one of claims 1 to 7 are implemented.

Citation Information

Patent Citations

  • Node cluster management method and device, equipment and storage medium

    CN112799789A

  • Task processing method, device and system, computer equipment and storage medium

    CN115729683A

  • Monitoring data acquisition method and device, electronic equipment and storage medium

    CN116055496A

  • Task allocation method and device, equipment and storage medium

    CN116932161A

  • Internet of Things data transmission method and device, computer equipment and storage medium

    CN118264692A