Data acquisition system for big data analysis
Patent Information
- Application Number
- CN202511334973.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-09-18
- Publication Date
- 2025-11-21
AI Technical Summary
现有技术在低温条件下,传感器节点的电量消耗不均衡,导致部分传感器节点提前退出工作,影响大数据分析的数据采集效果。
采用自适应的分簇时间周期计算方法,基于温度和数据转发记录,动态调整传感器节点的簇头和成员角色,优化数据传输路径以平衡电量消耗。
通过自适应的分簇时间周期计算,降低了传感器节点的电量消耗差别,减少了节点提前退出的概率,提高了数据采集系统的稳定性和效率。
Smart Images

Figure CN120994503A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of data acquisition, and more particularly to a data acquisition system for big data analysis. Background Technology
[0002] When analyzing environmental conditions, in order to acquire large-scale data for big data analysis, multiple sensor nodes are usually set up to collect environmental condition-related data.
[0003] In existing technologies, to conserve power in sensor nodes, the operating state of the sensor nodes is typically divided into member nodes and cluster heads through periodic clustering. However, periodic clustering is not effective in balancing power consumption at low temperatures, potentially causing some sensor nodes to prematurely cease operation. This is because the actual battery capacity decreases with decreasing temperature. Summary of the Invention
[0004] The purpose of this invention is to disclose a data acquisition system for big data analysis, thereby solving the technical problems mentioned in the background art.
[0005] To achieve the above objectives, the present invention provides the following technical solution: This invention provides a data acquisition system for big data analysis, including an edge control device and sensor nodes; The edge control device includes a computing module, a clustering module, and a transceiver module; The calculation module is used to adaptively calculate the clustering time period based on the temperature and data forwarding records at the location of the edge control device; The clustering module is used to cluster sensor nodes based on the clustering time period to obtain the set of cluster head node numbers; The transceiver module is used to send the set of cluster head node numbers to the sensor nodes.
[0006] The sensor node is used to determine whether its own number belongs to the cluster head node's number set. If it belongs to the cluster head node's number set, it will act as a cluster head node in the next clustering time period; otherwise, it will act as a member node in the next clustering time period.
[0007] Preferably, member nodes are used to acquire environmental monitoring data and transmit the monitoring data to the cluster head node; The cluster head node is used to forward monitoring data sent by member nodes to the edge control device.
[0008] Preferably, the cluster head node is also used to acquire environmental monitoring data and forward the acquired monitoring data to the edge control device.
[0009] Preferably, forwarding the monitoring data to the edge control device includes: Cluster head nodes scan the surrounding radio signals to determine whether the edge control device is within their communication range. If so, they send the monitoring data to the edge control device; otherwise, they select one of the other cluster head nodes within their communication range as a relay target and send the monitoring data to the relay target.
[0010] Preferably, selecting one of the other cluster head nodes within its communication range as the relay target includes: It periodically sends radio signals containing its own identifier to the surrounding area; Receive wireless charging signals containing numbers from other cluster head nodes; Store the numbers in the received wireless charging signals into the relay target set; When it is necessary to select a relay target, calculate the forwarding status value of the cluster head node corresponding to each number in the relay target set; The cluster head node with the largest forwarding status value in the relay target set is selected as the relay target.
[0011] Preferably, calculating the forwarding status value of the cluster head node corresponding to each number in the relay target set includes: Obtain the forwarding success rate of the cluster head node corresponding to each number in the relay target set during the current clustering time period; The forwarding status value is calculated based on the forwarding success rate and the signal strength of the wireless charging signals containing numbers sent by other cluster head nodes.
[0012] Preferably, the clustering time period is adaptively calculated based on the temperature and data forwarding records at the location of the edge control device, including: Acquire historical temperature data for the location of the edge control device; Calculate the temperature coefficient based on historical temperature data; Calculate the forwarding coefficient based on data forwarding records; Clustering time period is calculated based on temperature coefficient and forwarding coefficient.
[0013] Preferably, the historical temperature data includes the temperatures of the most recent N days, where N is a preset value.
[0014] Preferably, the temperature coefficient is calculated based on historical temperature data, including: Calculate the average temperature and temperature trend value based on the temperature over the most recent N days; The temperature coefficient is calculated based on the average temperature and the temperature trend value.
[0015] Preferably, clustering sensor nodes based on the clustering time period includes: After the previous clustering time period ends, the sensor nodes are re-clustered.
[0016] Beneficial effects: Compared with existing technologies, this invention no longer uses a fixed period to cluster sensor nodes. Instead, it uses an adaptive clustering time period calculation based on temperature and data forwarding records. This allows the clustering time period to match the temperature during the collection of environmental monitoring data for big data analysis, which can better reduce the difference in power consumption rates among various sensor nodes and reduce the probability of sensor nodes prematurely ceasing operation. Attached Figure Description
[0017] To more clearly illustrate the technical solutions of the embodiments of the present invention, the accompanying drawings used in the description of the embodiments will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0018] Figure 1 This is a schematic diagram of a data acquisition system for big data analysis according to the present invention.
[0019] Figure 2 This is a schematic diagram illustrating the process of obtaining transit targets according to the present invention. Detailed Implementation
[0020] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.
[0021] This invention provides a data acquisition system for big data analysis, including an edge control device and sensor nodes.
[0022] The sensor nodes of this invention are used to acquire environmental data and then transmit it to an edge control device. The edge control device is positioned at the center of the monitoring area, thereby reducing the average communication distance between the individual sensor nodes and the edge control device.
[0023] The edge control device then transmits the received data to the server for storage. When big data analysis is needed, the analysis terminal can read the data from the server for analysis.
[0024] In this invention, the edge control device includes a computing module, a clustering module, and a transceiver module.
[0025] The transceiver module is mainly used for communication with the server and for communication with sensor nodes.
[0026] In this invention, the calculation module is used to adaptively calculate the clustering time period based on the temperature and data forwarding records at the location of the edge control device.
[0027] The sensor node of this invention records a temperature every day, thereby providing data support for subsequent calculation of clustering time periods.
[0028] The data forwarding record is obtained by the edge control device through statistical analysis of the received data, representing the number of data collected by each sensor node.
[0029] Preferably, the clustering time period is adaptively calculated based on the temperature and data forwarding records at the location of the edge control device, including: Acquire historical temperature data for the location of the edge control device; Calculate the temperature coefficient based on historical temperature data; Calculate the forwarding coefficient based on data forwarding records; Clustering time period is calculated based on temperature coefficient and forwarding coefficient.
[0030] This invention calculates the clustering time period from two different directions, thereby adapting the clustering time period to changes in data transmission pressure and temperature, which enables the invention to better reduce the differences in power consumption rates among various sensor nodes.
[0031] Preferably, the historical temperature data includes the temperatures of the most recent N days, where N is a preset value.
[0032] For example, N can be 10.
[0033] Preferably, the temperature coefficient is calculated based on historical temperature data, including: Calculate the average temperature and temperature trend based on the temperatures of the most recent N days: The formula for calculating the average temperature is:
[0034] This is the average temperature. This is the temperature on day n after sorting the temperatures by acquisition time.
[0035] For example, temperatures obtained earlier are ranked higher.
[0036] The formula for calculating the temperature trend value is:
[0037] This represents the temperature trend value.
[0038] The temperature coefficient is calculated based on the average temperature and the temperature trend value, including: The formula for calculating the temperature coefficient is:
[0039] This represents the temperature coefficient, and `std` indicates normalization, mapping the values within the parentheses to the interval [0,1]. Calculate the weight for temperature, for example, it could be 0.7.
[0040] The temperature coefficient of this invention is weighted by normalizing the average temperature and the temperature trend value. This results in a larger temperature coefficient when the average temperature is higher and the temperature trend value is larger. This is beneficial for calculating the clustering time period in the subsequent calculation, which results in a larger clustering time period, reduces the clustering frequency, and improves the efficiency of data transmission. Conversely, a smaller temperature coefficient can increase the clustering frequency and more effectively balance the power consumption of the sensor nodes.
[0041] Preferably, calculating the forwarding coefficient based on data forwarding records includes: Get the total number of data collected by each sensor node on day n in the last N days, where n∈[1,N]. The larger n is, the closer day n is to the current time. Calculate the forwarding coefficient using the following formula:
[0042] Forwarding coefficient, Let M be the total number of data collected by the m-th sensor node on day n, where M is the total number of sensor nodes.
[0043] The forwarding coefficient of this invention does not directly calculate the average data collection volume over the most recent N days, because the reference value of data collection volume decreases as time recedes further into the present. Therefore, this invention uses a different approach... The statistical period is divided into two time periods, each with a different weight. This allows newer data to have a greater influence on the calculation of the forwarding coefficient, enabling the forwarding coefficient to more accurately represent the acquisition pressure of the sensor node. This is beneficial for the present invention to more accurately adjust the clustering time period.
[0044] Preferably, the clustering time period is calculated based on the temperature coefficient and the forwarding coefficient, including: The clustering time period is calculated using the following formula:
[0045] For the clustering time period, Calculate the weights (e.g., 0.3) for the clustering time period. This is the maximum value of the set clustering time period.
[0046] This invention, by comprehensively considering the temperature coefficient and the forwarding coefficient, achieves a reduction in clustering frequency when the recent temperature is higher, the trend of increasing temperature is stronger, and the data acquisition pressure is lower, resulting in a longer clustering time period. Conversely, when the recent temperature is lower, the trend of decreasing temperature is stronger, and the data acquisition pressure is higher, resulting in a shorter clustering time period, thereby increasing the clustering frequency and achieving a more balanced power consumption rate among various sensor nodes.
[0047] In this invention, the clustering module is used to cluster sensor nodes based on the clustering time period to obtain the set of cluster head node numbers.
[0048] After the new clustering time period is calculated, the clustering module of the present invention starts to calculate the time length between the moment when the calculation result of the clustering time period is obtained and the current moment. When the time length is greater than the new clustering time period, it indicates that the current clustering period has ended.
[0049] The maximum clustering time period of this invention can be 1 day.
[0050] Preferably, clustering sensor nodes based on the clustering time period includes: After the previous clustering time period ends, the sensor nodes are re-clustered.
[0051] During the re-clustering process, the sensor nodes continue to use the sensor network architecture formed in the previous clustering time period for data transmission.
[0052] Preferably, the clustering process is as follows: The first step is to obtain the maximum x-coordinate of each sensor node. Minimum value of x-axis Maximum value of the ordinate and minimum value of the ordinate ; The second step is to establish region K:
[0053] x and y are the x and y coordinates of a point in the region, respectively; The third step is to divide region K into multiple sub-regions of equal size; The fourth step is to designate the sensor node with the most remaining power in each sub-region as the cluster head node.
[0054] By partitioning the data first and then selecting the cluster head node, we can avoid excessive cluster head nodes, thereby reducing the average backoff waiting time during communication and improving data transmission efficiency.
[0055] Using the sensor node with the most remaining power as the cluster head node can better reduce the power consumption gap between various sensor nodes.
[0056] Preferably, region K is divided into multiple sub-regions of equal area, including: The sub-region is a rectangular region with four sides of equal length; The formula for calculating the length of the rectangular region is:
[0057] s is the length of the rectangular area, and R is the communication radius of the sensor node.
[0058] The length calculated using the above formula can increase the probability that cluster head nodes in adjacent sub-regions can communicate directly.
[0059] In the process of dividing region K into multiple subregions of equal area, the remainder obtained by dividing the length of region K by s is denoted as p1. If p1 is not 0, then region K is expanded in the length direction, and the increased length is... ; If the remainder obtained by dividing the width of region K by s is denoted as p2, and if p2 is not 0, then region K is expanded in the width direction, with the increased width being... ; Then the expanded area is divided into multiple sub-regions of equal size.
[0060] The above process can expand region K when the length and / or width of region K is not an integer multiple of s, so as to meet the need to obtain multiple sub-regions with the same area.
[0061] In this invention, the transceiver module is used to send the set of cluster head node numbers to the sensor node.
[0062] When transmitting the cluster head node set number, the sensor network architecture formed in the previous clustering time period is continued to be used for transmitting the cluster head node set number.
[0063] In this invention, the sensor node is used to determine whether its own number belongs to the number set of the cluster head node. If it belongs to the number set of the cluster head node, it will act as the cluster head node in the next clustering time period; otherwise, it will act as a member node in the next clustering time period.
[0064] In this invention, member nodes are used to acquire environmental monitoring data and transmit the monitoring data to the cluster head node; The cluster head node is used to forward monitoring data sent by member nodes to the edge control device.
[0065] The monitoring data includes temperature, humidity, etc.
[0066] In this invention, the cluster head node is also used to acquire environmental monitoring data and forward the acquired monitoring data to the edge control device.
[0067] Preferably, forwarding the monitoring data to the edge control device includes: Cluster head nodes scan the surrounding radio signals to determine whether the edge control device is within their communication range. If so, they send the monitoring data to the edge control device; otherwise, they select one of the other cluster head nodes within their communication range as a relay target and send the monitoring data to the relay target.
[0068] The sensor nodes of this invention periodically broadcast their own identification number so that other sensor nodes within the communication range can update their neighbor node tables.
[0069] Preferably, such as Figure 2 It selects one of the other cluster head nodes within its communication range as a relay target, including: It periodically sends radio signals containing its own identifier to the surrounding area; Receive wireless charging signals containing numbers from other cluster head nodes; Store the numbers in the received wireless charging signals into the relay target set; When it is necessary to select a relay target, calculate the forwarding status value of the cluster head node corresponding to each number in the relay target set; The cluster head node with the largest forwarding status value in the relay target set is selected as the relay target.
[0070] Preferably, calculating the forwarding status value of the cluster head node corresponding to each number in the relay target set includes: Get the forwarding success rate of the cluster head node corresponding to each number in the relay target set during the current clustering time period: Calculate the forwarding success rate using the following formula:
[0071] Let k be the forwarding success rate of the cluster head node k corresponding to the number in the target set. This represents the total number of times the cluster head node that needs to forward data sends data to cluster head node k within the current clustering time period. This represents the total number of times the cluster head node that needs to forward data resends data to cluster head node k within the current clustering time period. The forwarding status value is calculated based on the forwarding success rate and the signal strength of the wireless charging signals containing numbers sent by other cluster head nodes.
[0072] The forwarding status value of cluster head node k. Calculate a weight (e.g., 0.5) for the forwarding state value. The signal strength of the wireless charging signal containing the number sent by cluster head node k, which is currently being forwarded by the cluster head node that needs to forward data.
[0073] By weighting success rate and signal strength, cluster head nodes with higher forwarding success rates and stronger signal strength can achieve higher forwarding status values, which helps ensure the quality of data transmission.
[0074] The preferred embodiments of the present invention disclosed above are merely illustrative of the invention. These preferred embodiments do not exhaustively describe all details, nor do they limit the invention to the specific implementations described. Clearly, many modifications and variations can be made based on the content of this specification. This specification selects and specifically describes these embodiments to better explain the principles and practical applications of the invention, thereby enabling those skilled in the art to better understand and utilize the invention. The invention is limited only by the claims and their full scope and equivalents.
Claims
1. A data acquisition system for big data analysis, characterized in that, Includes edge control devices and sensor nodes; The edge control device includes a computing module, a clustering module, and a transceiver module; The calculation module is used to adaptively calculate the clustering time period based on the temperature and data forwarding records at the location of the edge control device; The clustering module is used to cluster sensor nodes based on the clustering time period to obtain the set of cluster head node numbers; The transceiver module is used to send the set of cluster head node numbers to the sensor nodes; The sensor node is used to determine whether its own number belongs to the cluster head node's number set. If it belongs to the cluster head node's number set, it will act as a cluster head node in the next clustering time period; otherwise, it will act as a member node in the next clustering time period.
2. The data acquisition system for big data analysis according to claim 1, characterized in that, Member nodes are used to acquire environmental monitoring data and transmit the monitoring data to the cluster head node; The cluster head node is used to forward monitoring data sent by member nodes to the edge control device.
3. The data acquisition system for big data analysis according to claim 1, characterized in that, The cluster head node is also used to acquire environmental monitoring data and forward the acquired monitoring data to the edge control device.
4. A data acquisition system for big data analysis according to claim 2 or 3, characterized in that, Forwarding monitoring data to the edge control device includes: Cluster head nodes scan the surrounding radio signals to determine whether the edge control device is within their communication range. If so, they send the monitoring data to the edge control device; otherwise, they select one of the other cluster head nodes within their communication range as a relay target and send the monitoring data to the relay target.
5. A data acquisition system for big data analysis according to claim 4, characterized in that, Select one of the other cluster head nodes within its communication range as a relay target, including: It periodically sends radio signals containing its own identifier to the surrounding area; Receive wireless charging signals containing numbers from other cluster head nodes; Store the numbers in the received wireless charging signals into the relay target set; When it is necessary to select a relay target, calculate the forwarding status value of the cluster head node corresponding to each number in the relay target set; The cluster head node with the largest forwarding status value in the relay target set is selected as the relay target.
6. A data acquisition system for big data analysis according to claim 5, characterized in that, Calculate the forwarding status value of the cluster head node corresponding to each number in the relay target set, including: Obtain the forwarding success rate of the cluster head node corresponding to each number in the relay target set during the current clustering time period; The forwarding status value is calculated based on the forwarding success rate and the signal strength of the wireless charging signals containing numbers sent by other cluster head nodes.
7. A data acquisition system for big data analysis according to claim 1, characterized in that, Based on the temperature and data forwarding records at the location of the edge control device, the clustering time period is adaptively calculated, including: Acquire historical temperature data for the location of the edge control device; Calculate the temperature coefficient based on historical temperature data; Calculate the forwarding coefficient based on data forwarding records; Clustering time period is calculated based on temperature coefficient and forwarding coefficient.
8. A data acquisition system for big data analysis according to claim 7, characterized in that, Historical temperature data includes the temperatures of the most recent N days, where N is a preset value.
9. A data acquisition system for big data analysis according to claim 8, characterized in that, The temperature coefficient is calculated based on historical temperature data, including: Calculate the average temperature and temperature trend value based on the temperature over the most recent N days; The temperature coefficient is calculated based on the average temperature and the temperature trend value.
10. A data acquisition system for big data analysis according to claim 8, characterized in that, The sensor nodes are clustered based on the clustering time period, including: After the previous clustering time period ends, the sensor nodes are re-clustered.
Citation Information
Patent Citations
Human body data collection method and device
CN113825219A
Mobile sensor network cluster head generation method based on error-random weight evaluation mechanism
CN117651313A
High-altitude photovoltaic module state management system
CN120050744A
Method for reporting and accumulating data in a wireless communication network
US20060187866A1