Internet of Things Data Storage, Processing and Analysis System Based on Massive Data
Through InfluxDB cluster and Spark technology, the problem of insufficient storage and processing capabilities of massive IoT data is solved, efficient data analysis and real-time processing are achieved, and the system's scalability and processing speed are improved.
Patent Information
- Application Number
- CN202111242760.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2021-10-25
- Publication Date
- 2025-05-30
- Estimated Expiration
- 2041-10-25
AI Technical Summary
The existing technology is difficult to effectively process and analyze massive IoT data, the stand-alone storage capacity is limited and the processing capacity is insufficient, and traditional Hadoop is difficult to achieve real-time processing and make full use of memory.
InfluxDB cluster database is used to store time series data, combine Spark processing and analysis data, design a partition fault tolerance consistency model, use a consistent hashing algorithm to ensure the scalability of the cluster, and realize real-time data analysis through Spark Streaming.
It realizes efficient storage and processing of massive IoT data, solves the problem of insufficient single-machine capacity and processing capabilities, improves the speed and efficiency of data analysis, simplifies programming complexity, and reduces the number of hard disk exchanges through memory storage intermediate results.
Smart Images

Figure CN113961562B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of data processing, and specifically to an Internet of Things data storage, processing and analysis system based on massive data. Background Art
[0002] With the continuous development of the Internet of Things, massive amounts of Internet of Things data are collected and stored. However, the storage capacity of a single machine is limited, and the ability to process and analyze data is insufficient. A single machine or a simple distributed solution can no longer meet the requirements.
[0003] Big data technologies represented by the Hadoop ecosystem provide a way for storing data for the continuously growing massive data. HDFS is an important component of the Hadoop ecosystem, which realizes horizontal expansion of storing data in a cluster of ordinary hardware. However, there are great difficulties in using Hadoop to process and analyze data. For example, MapReduce programming is difficult, and different batch processes need to be written for different scenarios. Moreover, Hadoop mainly uses hard disks to store the intermediate results of calculations, and data needs to be continuously swapped in and out of memory, resulting in limited speed and inability to make full use of the large-capacity memory. In addition, HDFS is generally used to process unstructured data, and it is difficult to store Internet of Things data. In addition, traditional Hadoop is difficult to efficiently process data in real time and cannot fully mine the short-term and effective information existing in the data. Summary of the Invention
[0004] To solve the problems existing in the above solutions, the present invention provides an Internet of Things data storage, processing and analysis system based on massive data. The present invention can be easily horizontally expanded, solves the problems of limited storage capacity and insufficient processing capacity of a single machine, and stores, processes and analyzes massive amounts of Internet of Things data by fully utilizing the capabilities of the cluster through using an InfluxDB cluster to store time series data and Spark to process and analyze data.
[0005] The object of the present invention can be achieved by the following technical solutions:
[0006] An Internet of Things data storage, processing and analysis system based on massive data, including a cluster building module, an InfluxDB cluster database, a server, a data processing module and a data analysis module;
[0007] The cluster building module: used for building an InfluxDB cluster. The distributed expansion of the InfluxDB cluster requires designing a partition fault-tolerant consistency model, and different models are designed according to the characteristics of different modules in InfluxDB; and the consistent hashing algorithm is used to ensure the scalability of the number of sub-clusters;
[0008] InfluxDB Cluster Database: used to store data generated by IoT devices; Data Processing Module: used to process IoT data using Spark. First, the client reads IoT data using the InfluxDB cluster, converts it into the data structure required by Spark, and then performs preprocessing;
[0009] Data Analysis Module: used to analyze the data processed by the Data Processing Module using Spark Streaming, and then continuously send the analyzed data into Spark ML for training to continuously iterate the model.
[0010] Furthermore, the partition tolerance consistency model includes three characteristic indicators: consistency, availability, and partition tolerance; among them, consistency is manifested as: different cluster nodes read the same data content or fail; availability is manifested as: the client accesses the cluster and gets the corresponding content but does not guarantee the content is the latest; partition tolerance is manifested as: when messages between nodes are lost, the cluster still works normally.
[0011] Furthermore, the InfluxDB cluster consists of META nodes and DATA nodes. The META nodes are used to save the key information for the system to run, meeting the consistency characteristic indicator; the key information includes database names, table names, and retention policy information; the DATA nodes are used to save specific IoT data, meeting the availability characteristic indicator.
[0012] Furthermore, both the META nodes and the DATA nodes meet the partition tolerance characteristic indicator.
[0013] Furthermore, the specific processing flow of the Data Processing Module is as follows:
[0014] First, the client establishes a connection with the InfluxDB cluster database. At this time, the proxy nodes of the InfluxDB cluster will establish a Session with the client and wait for the client to send commands;
[0015] The client sends a command to query data according to the task. The proxy nodes of the InfluxDB cluster dispatch the query task to different nodes according to the query content, and each node searches for the corresponding data according to the command and returns it;
[0016] The client receives the data returned by the InfluxDB cluster and converts it into the data type required by Spark. Spark saves the data in memory;
[0017] The client preprocesses the data in memory according to the written code. The preprocessing includes: screening data that meets the conditions, converting data types, and regularization.
[0018] Furthermore, the data processing module is also used to store the preprocessed data in HDFS for use by programs or models of other versions; the preprocessed data is unstructured data.
[0019] Furthermore, the specific analysis steps of the data analysis module are as follows:
[0020] First, the data received by the server is stored in the message queue and then in the InfluxDB cluster database; the message queue continuously inputs the data into Spark Streaming;
[0021] Spark continuously receives real-time input data streams, splits them into several batches of data according to a predetermined time interval, then processes these data through the Spark Engine, and finally obtains several batches of processed result data. Then, the several batches of processed result data are continuously sent into Spark ML for training, and the model is continuously iterated.
[0022] Furthermore, the data analysis module is also used to build a machine learning application using ML Pipeline, continuously process the data and send it into Spark ML for training, and continuously iterate the model.
[0023] Compared with the prior art, the beneficial effects of the present invention are:
[0024] 1. The present invention uses the InfluxDB cluster database to store the data generated by IoT devices. Compared with traditional relational databases, it optimizes the scenario of more writes and fewer reads, can improve the read and write efficiency, and can ensure the order of time.
[0025] 2. Designing the InfluxDB cluster can solve the problems of limited single-machine data capacity and insufficient processing power. Through the cluster, disaster recovery backup can also be provided to ensure data security. Through distributed expansion, the storage capacity is increased, and when reading, it can be read from multiple machines simultaneously, reducing the time for reading data during data analysis.
[0026] 3. The present invention uses Spark to process and analyze data, which improves the processing speed, simplifies the programming of batch processing tasks, and has greatly improved usability compared with the traditional Hadoop MapReduce. At the same time, by using in-memory storage of intermediate results, the number of exchanges with the hard disk is greatly reduced, and the time required to complete the task is reduced. Description of the Drawings
[0027] To more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the following will briefly introduce the drawings required for use in the description of the embodiments or the prior art. Obviously, the drawings in the following description are only some embodiments of the present invention. For those of ordinary skill in the art, without creative efforts, other drawings can also be obtained based on these drawings.
[0028] Figure 1 It is the system block diagram of the present invention.
[0029] Figure 2 It is the schematic diagram of building an InfluxDB cluster in the present invention.
[0030] Figure 3 It is the schematic diagram of the process of using Spark to process and analyze Internet of Things data in the present invention.
[0031] Figure 4 It is the schematic diagram of the process of using Spark Streaming to analyze data in real time in the present invention. Detailed implementation manners
[0032] The following will clearly and completely describe the technical solutions of the present invention in combination with the embodiments. Obviously, the described embodiments are only some embodiments of the present invention, rather than all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts belong to the scope of protection of the present invention.
[0033] As Figures 1 to 4 shown, the Internet of Things data storage, processing and analysis system based on massive data includes a cluster building module, an InfluxDB cluster database, a data processing module and a data analysis module;
[0034] The cluster building module is used to build an InfluxDB cluster. The distributed expansion of the InfluxDB cluster requires designing a partition tolerance consistency model, and different models are designed for the characteristics of different modules in InfluxDB; specifically including:
[0035] The design of a distributed system needs to consider three indicators: system consistency, availability and partition tolerance, but the three cannot exist at the same time. Consistency needs to ensure that the data content read by different cluster nodes is the same or fails. Availability needs to ensure that the client can access the cluster and get the corresponding content but does not guarantee that the content is the latest. Partition tolerance ensures that the cluster can still serve when the messages between nodes are lost.
[0036] There is communication among multiple nodes in the cluster, so it is necessary to ensure partition tolerance. Therefore, the design of different modules of InfluxDB needs to make a trade-off between consistency and availability. The distributed system of InfluxDB mainly consists of META nodes and DATA nodes; META nodes store key information for system operation, such as database names, table names, retention policy information, etc. Therefore, data consistency must be ensured, otherwise it will affect the operation of the system, causing some modules to fail to write data due to inconsistent META information.
[0037] For DATA nodes, they mainly store specific IoT data and can tolerate partial inconsistency of data. Therefore, the availability of DATA nodes is selected to be ensured. At the same time, availability reduces the difficulty of horizontal expansion of DATA nodes.
[0038] When the performance of a single cluster reaches a bottleneck, sub-clusters can be used to break through the performance of the single cluster. By using the consistent hashing algorithm, the scalability of the number of sub-clusters can be ensured. The consistent hashing algorithm can reduce the re-mapping of data when the number of clusters increases, reducing the cost of data migration. Compared with traditional proxy and hashing algorithms, it ensures the load balance of each cluster node and reduces the degree of business intrusion.
[0039] The InfluxDB cluster database is used to store data generated by IoT devices; the data processing module is used to process IoT data using Spark. The specific processing process is as follows:
[0040] First, the client establishes a connection with the InfluxDB cluster database. At this time, the proxy node of the InfluxDB cluster will establish a Session with the client and wait for the client to send commands;
[0041] The client sends a command to query data according to the task. The proxy node of the InfluxDB cluster distributes the query task to different nodes according to the query content. Each node searches for the corresponding data according to the command and returns it;
[0042] The client receives the data returned by the InfluxDB cluster and converts it into a data type that can be efficiently utilized by Spark. Spark will keep the data in memory all the time, thereby improving the processing speed of the next step;
[0043] The client processes the data in memory according to the written code. The specific processing content includes steps such as screening data that meets the conditions, conversion of data types, regularization, etc.;
[0044] After the data processing is completed, the data can be analyzed. For example, SparkML can be used to cluster the data to find the relationships between the data, and the results can be visualized. It is also possible to use SparkML to train a model. For example, a model can be trained based on the status information of the power supply equipment collected, so as to evaluate the status of the power supply equipment and predict whether a failure will occur, etc.;
[0045] The data analysis module is used to analyze the data processed by the data processing module using Spark Streaming. The specific steps are as follows:
[0046] First, the data received by the server is stored in the message queue and then in the InfluxDB cluster database;
[0047] The message queue continuously inputs the data into Spark Streaming. Spark can continuously receive real-time input data streams, split them into several batches of data according to a predetermined time interval, and then process these data through the Spark Engine. Finally, several batches of processed result data are obtained, and then the several batches of processed result data are continuously sent into Spark ML for training to continuously iterate the model;
[0048] In addition, the processed data is unstructured data and has a reduced relationship with time. It can be stored in HDFS for use by other versions of programs or models, such as models built with Python and the display of the front-end interface. Compared with directly reading and processing, the speed can be improved;
[0049] In addition, Spark ML Pipline can be used to improve efficiency and reduce the difficulty of coding. Spark ML Pipeline builds a set of High-level APIs based on DataFrame. We can use ML Pipeline to build machine learning applications. It can organize multiple processing processes of a machine learning application and manage the running order of each processing step at the code implementation level, greatly simplifying the difficulty of developing machine learning applications.
[0050] The working principle of the present invention:
[0051] The Internet of Things data storage, processing and analysis system based on massive data, when working, first of all, the cluster building module is used to build an InfluxDB cluster. For the distributed expansion of the InfluxDB cluster, a partition tolerance consistency model needs to be designed, and different models are designed according to the characteristics of different modules in InfluxDB. Then, the data processing module is used to process Internet of Things data by using Spark. First, the client reads data by using the InfluxDB cluster, converts it into the data structure required by Spark, and performs preprocessing, such as data type conversion, missing value filling, regularization, etc. Since all the read data is stored in memory, the data processing speed can be greatly accelerated. The data processing module is used to store the processed data into HDFS for other uses, such as data display, data analysis, etc.
[0052] The data analysis module is used to analyze the data processed by the data processing module by using Spark Streaming, and use Spark ML for data analysis such as clustering and regression. It can also build a machine learning system to continuously process data and send it into model training, and continuously iterate and train the model according to the collected Internet of Things data, etc. This improves the data processing speed, simplifies the programming of batch processing tasks, and has greatly improved the usability compared with the traditional Hadoop MapReduce. At the same time, by using in-memory storage of intermediate results, the number of exchanges with the hard disk is greatly reduced, and the time required to complete the task is reduced.
[0053] In the description of this specification, the descriptions referring to terms such as "one embodiment", "example", "specific example", etc. mean that the specific features, structures, materials or characteristics described in connection with the embodiment or example are included in at least one embodiment or example of the present invention. In this specification, the schematic representations of the above terms do not necessarily refer to the same embodiment or example. Moreover, the specific features, structures, materials or characteristics described can be combined in any one or more embodiments or examples in a suitable manner.
[0054] The above-disclosed preferred embodiments of the present invention are only used to help illustrate the present invention. The preferred embodiments do not describe all the details in detail, nor do they limit the present invention to the specific embodiments described. Obviously, many modifications and variations can be made according to the content of this specification. This specification selects and specifically describes these embodiments in order to better explain the principles and practical applications of the present invention, so that those skilled in the relevant technical fields can understand and utilize the present invention well. The present invention is only limited by the claims and their full scope and equivalents.
Claims
1. An Internet of Things data storage, processing and analysis system based on massive data, characterized in that, it includes a cluster building module, an InfluxDB cluster database, a server, a data processing module and a data analysis module; Cluster building module: used to build an InfluxDB cluster. For the distributed expansion of the InfluxDB cluster, a partition tolerance consistency model needs to be designed, and different models are designed according to the characteristics of different modules in InfluxDB; and the consistent hashing algorithm is used to ensure the scalability of the number of sub-clusters; InfluxDB cluster database: used to store the data generated by Internet of Things devices; Data processing module: used to process Internet of Things data using Spark. First, the client reads the Internet of Things data using the InfluxDB cluster, converts it into the data structure required by Spark, and then performs preprocessing; Data analysis module: used to analyze the data processed by the data processing module using Spark Streaming, and then continuously send the analyzed data into Spark ML for training to continuously iterate the model; The partition tolerance consistency model includes three characteristic indicators: consistency, availability, and partition tolerance; among them, consistency is manifested as: different cluster nodes read the same data content or fail; availability is manifested as: the client accesses the cluster to obtain the corresponding content but does not guarantee the latest content; partition tolerance is manifested as: when the messages of node-to-node communication are lost, the cluster still works normally; The InfluxDB cluster consists of META nodes and DATA nodes. The META nodes are used to save the key information of the system operation and meet the consistency characteristic indicators; the key information includes database names, table names, and retention policy information; the DATA nodes are used to save specific Internet of Things data and meet the availability characteristic indicators; The specific processing process of the data processing module is as follows: First, the client establishes a connection with the InfluxDB cluster database. At this time, the proxy node of the InfluxDB cluster will establish a Session with the client and wait for the client to send commands; The client sends a command to query data according to the task. The proxy node of the InfluxDB cluster distributes the query task to different nodes according to the query content. Each node searches for the corresponding data according to the command and returns it; The client receives the data returned by the InfluxDB cluster and converts it into the data type required by Spark. Spark saves the data in memory; The client preprocesses the data in memory according to the written code. The preprocessing includes: screening data that meets the conditions, conversion of data types, and regularization.
2. The Internet of Things data storage, processing and analysis system based on massive data according to claim 1, characterized in that, both the META nodes and the DATA nodes meet the partition tolerance characteristic indicators.
3. The Internet of Things data storage, processing and analysis system based on massive data according to claim 1, characterized in that, The data processing module is also used to store the preprocessed data in HDFS for use by other versions of programs or models; the preprocessed data is unstructured data.
4. The Internet of Things data storage, processing and analysis system based on massive data according to claim 1, characterized in that the specific analysis steps of the data analysis module are as follows: First, the data received by the server is stored in the message queue and then in the InfluxDB cluster database; the message queue continuously inputs the data into Spark Streaming; Spark continuously receives the real-time input data stream, splits it into several batches of data according to a predetermined time interval, then processes these data through the Spark Engine, and finally obtains several batches of processed result data. Then, the several batches of processed result data are continuously sent into Spark ML for training, and the model is continuously iterated.
5. The Internet of Things data storage, processing and analysis system based on massive data according to claim 4, characterized in that the data analysis module is also used to build a machine learning application using ML Pipeline, continuously process the data and send it into Spark ML for training, and continuously iterate the model.
Citation Information
Patent Citations
Spark Streaming abnormal temperature data warning method based on Spark platform
CN106778033A
Cloud computing data platform construction method based on Kubernetes
CN111327681A