A unified storage and query method and system for multi-sensor data

Through unified storage and query methods, the problems of multi-sensor data storage and query are solved, and flexible processing and efficient storage of sensor data are realized, which is suitable for the real-time writing and query requirements of diversified sensor data.

CN113868217BActive Publication Date: 2025-05-30SHANGHAI YINJIANG SMART INTELLIGENT TECH CO LTD
View PDF 6 Cites 0 Cited by

Patent Information

Application Number
CN202111155813.4
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2021-09-30
Publication Date
2025-05-30
Estimated Expiration
2041-09-30

AI Technical Summary

Technical Problem

The prior art is difficult to effectively store and query multi-sensor data uniformly, resulting in high data accumulation pressure, low query efficiency, and the inability to flexibly handle the diversity of sensors and the need for real-time writes to real-time query.

Method used

A unified storage and query method for multivariate sensor data is adopted. The data is received uniformly by receiving service module, the data analysis service module analyzes and processes data, and the write service module writes data into HDFS, and maintains the index file. The query service module analyzes query requests and realizes flexible data query through the data reading service module.

Benefits of technology

It realizes unified storage and flexible query of multi-sensor data, can access sensor equipment and set sensor packets at any time, without affecting data reception and query, has horizontal scaling capabilities, and makes full use of Hadoop's storage capabilities, improving the performance of data writing and query.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN113868217B_ABST
    Figure CN113868217B_ABST
Patent Text Reader

Abstract

The present invention relates to a method and system for unified storage and query of multi-sensor data. The present invention can access sensor devices at any time without manual intervention; it can also set and modify sensor groups at any time without affecting data reception and query; a unified storage method is adopted for different sensor data, which can easily achieve horizontal expansion; and the powerful storage capacity of Hadoop is fully utilized to achieve flexible data query at the same time.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of information processing, and particularly relates to a unified storage and query method and system for multi-source sensor data. Background Art

[0002] At present, the development of the Internet of Things has matured. Various sensing devices required in the construction of the judicial administrative application system, such as prison management sensing devices and execution correction bracelets, can be classified into on-site types, remote sensing types, etc. from the perspective of observation and measurement standards, and can be divided into fixed sensors and mobile sensors from the mobility of the observation platform. For so many types of sensors, their data types, data compositions, data sending frequencies, and data accuracies are different. The traditional approach is to separately set up a data storage table in a relational database according to different sensor categories, and each piece of data of the sensor is stored in a separate row. As time goes by, these large amounts of accumulated data will surely cause huge pressure on the storage and query of the database. Even if the HBase big data storage technology is used for storage, a great deal of effort is required to design the row keys for these sensor data with different characteristics. With the influx of a large amount of data, problems related to the sustainable use of the system caused by HBase Region splitting will also be faced. At the same time, the method of separately designing tables for multi-source sensors requires adding corresponding data storage services and query services. In addition to spending a great deal of effort in design and development, the parallel operation of multiple services will surely cause contention for system resources, imposing no small resource pressure and operation and maintenance pressure on the server.

[0003] In response to such scenarios with multiple sensors, diverse data formats, and large amounts of data, various storage technologies have emerged. For example, the "Storage Method and Device for Real-Time Streaming Data" with the application number 201710224721.4 receives real-time streaming data; parses the real-time streaming data to obtain a parsing result; determines the number of data entries of the real-time streaming data according to the parsing result; determines whether the number of data entries of the real-time streaming data reaches a preset number of data entries; if so, writes the parsing result of the real-time streaming data into a distributed data query engine; the "Real-Time Storage Method for Unstructured Streaming Data in Rail Transit" with the application number 201910181493.6 also accesses the collected multi-source unstructured streaming data into a data buffer; after the buffer data volume of the queue reaches a threshold, uses the multi-threaded writing method of HBase or the flushCommits() method to write the data into HBase. These methods all use a cache to receive real-time data and batch write it into storage once the cache is full. The advantage is that it can improve the data writing performance, but there are several problems: 1. System failures such as crashes and power outages will inevitably cause the data in the cache to be lost; 2. No solution is proposed for sensor diversity; 3. It cannot meet the application scenario of real-time writing and real-time query, which is unacceptable in the time-critical emergency command system environment; the "Automatic In-warehousing Method for Multisensor Multivariate Heterogeneous Data Streams" with the application publication number CN 104750814 A realizes the data writing method for multiple sensors, and uses a sensor format information table, a monitoring date table, an abnormal character table, and a repair function table to store the corresponding information respectively. This method needs to establish a sensor format information table according to the data format of the sensor before receiving heterogeneous streaming data, and cannot flexibly handle the diversification of multiple sensors. At the same time, using various repair functions to repair data will lose the authenticity of the data, resulting in the inability to implement analysis scenarios such as sensor stability analysis. Usually, these abnormal data should be processed by data users according to actual needs; the "Data Storage Method, Dashcam, Server, and Storage Medium" with the application publication number CN 110737807 A uses the indexing feature of the EXT file system to use the date of data collection as the primary index and the data collection time as the secondary index. This method has a strong dependence on the EXT file system and is not applicable to the HDFS file system. At the same time, multiple file scans are required for data queries within a time range, and its method of storing indexes and data blocks together reduces the efficiency of index scanning; in the "High-Performance File Storage and Management System Based on HDFS" with the application publication number CN 110647497 A, when querying a corresponding file, it is necessary to sequentially scan the index file, and cannot quickly locate the target, and the efficiency is not high, which is not applicable to application scenarios with frequent reads. In addition, for files with frequent reads, the optimization method of copying them to the local has little significance for scenarios with large randomness in reading, and is also not applicable to the application scenario of the present invention. Summary of the Invention

[0004] To overcome the above deficiencies, the present invention aims to provide a unified storage and query method and system for multi-sensor data. The present invention can connect to sensor devices at any time without manual intervention; it can also set and modify sensor groups at any time without affecting data reception and query; a unified storage method is adopted for different sensor data, which can easily achieve horizontal expansion; and the powerful storage capacity of Hadoop is fully utilized to achieve flexible data query at the same time.

[0005] The present invention achieves the above object through the following technical solutions: A unified storage method for multi-sensor data, comprising the following steps:

[0006] (1) The receiving service module uniformly receives data. After the receiving service module writes the data into a log file, it sends the data to the data parsing service module; the data sending method is in the form of a message.

[0007] (2) After receiving the message, the data parsing service module obtains the log file name, and judges whether the file has been opened. If so, it continues to read the file until the end of the file; otherwise, it judges whether the currently opened file has reached the end of the file. If so, it opens a new log file for parsing and deletes the previously parsed log file at the same time. Otherwise, it puts the new log file name into the cache and waits to be read, continues to read the currently opened log file until the end of the file and deletes the file, then extracts the new file name from the cache and opens the file for reading; and sends the name of the parsed log file to the data writing service module; the sending method is in the form of a message.

[0008] (3) After receiving the name of the parsed log file, the data writing service module appends the device registration code and detection data in the parsed log file to the HDFS data file named by the registration code, and maintains the information in the index file.

[0009] Preferably, in the step (1), when the size of the log file expands to the set size limit, the data is written into a new log file, and the receiving service module sends the corresponding log file name to the data parsing service module in the form of a message. The receiving service module uses a multi-threaded method to receive data. In principle, one thread corresponds to one log file.

[0010] Preferably, after the data parsing service module receives the message in step (2), it forks a new process to open and parse the new log file, and at the same time maintains a file list to record the currently open files; the data parsing service module extracts the device's data description information and the detected data from the log file. Specifically, it extracts key information such as device identifier, name, sensor type, sensor coordinates (for fixed sensors), data items, and data types from the data description information, and sends these key information and the monitoring time to the device registration service module. The device registration service module determines whether the device has been registered according to the key information. If it has not been registered, it records the key information, writes the monitoring time, and assigns a registration code, and returns the registration code to the data parsing service module; if it has been registered, after returning the registration code and the last detection time to the data parsing service module, it writes the current monitoring time; the device registration service module sends the registration code, device description information, detected data, monitoring time, and last monitoring time to the writing service module.

[0011] Preferably, after obtaining the name of the parsed log file in step (3), it determines whether the currently open parsed log file is the same as the file name in the message. If the file names are different, it determines whether the currently open file has reached the end of the file. If so, it opens a new log file for parsing and deletes the previously parsed log file at the same time. Otherwise, it puts the name of the newly parsed log file into the cache for waiting to be read, and continues to read the currently open parsed log file; after the data parsing service module receives the message, it forks a new process to open the new log file and at the same time maintains a file list to record the open files; the data writing service module reads the device's registration code and detected data from the parsed log file and appends them to the HDFS data file named after the registration code, and maintains index information. Except for the different suffixes, the file names of the index file group are the same as those of the data file; if the file does not exist, it means that the data of a new device has been received. For a new device, it creates a data file and an index file group named after the registration code, writes the device description information, data item information, data acquisition frequency, etc. into the data header of the data file, and then appends the monitoring data to the data file, and at the same time maintains index information.

[0012] Preferably, the specific way of maintaining the index is as follows: after obtaining information such as the registration code, the position Pdata of the monitoring data in the data file, the monitoring time, and the last detection time, it determines whether the current monitoring time is the first data of the day (new date Dn) according to the monitoring time and the last detection time, and performs different operations according to whether it is a new date:

[0013] If it is a new date, append the position of the current monitoring data in the data file to the end of the data index file, record the current position Pindex of the data index, append the current monitoring time to the time index, record the current position PTime of the time index, and then append PTime to the data index; at the same time, according to whether the current monitoring time matches the last year in the date index, if it is the same year, directly append the data index position Pindex to the day before Dn-1 in the date index, if the day before Dn-1 does not exist, append -1 to the back of Dn-2 to get Dn-1 and then append Pindex to the back of Dn-1, and so on until January 1st; if it is not the same year, append the year and year identifier to the date index and then write Pindex;

[0014] If it is not a new date, compare the current monitoring time with the previous monitoring time to obtain the time difference DIFn (seconds), append the time difference DIFn to the time index, write to the data file in the data file, obtain the data file offset DIFp according to the position before writing and the position after writing, and append the offset DIFp to the data index.

[0015] A unified query method for multi-sensor data includes the following steps:

[0016] 1) After the query service module receives a query request, it sends the query conditions to the query parsing service module;

[0017] 2) The query parsing service module parses information such as device identifier, sensor name, sensor group name, and monitoring time according to the query conditions, and judges whether this information exists in the device registration service module. If it does not exist, it returns an exception message. Otherwise, it sends the obtained registration code and monitoring time to the query service module;

[0018] 3) The query service module sends the device registration code and monitoring time obtained in step 2) to the data reading service module;

[0019] 4) The data reading service module jumps to the specified year and date position in the corresponding date index to obtain the position Pindex in the data index. Pindex points to the position where the data of the first monitoring on the current day is stored in the data file. Immediately following Pindex is the position Ptime of the time of the first monitoring on the current day in the time index; according to Ptime, read the corresponding time Dtime from the time index, and scan the time offset backward. The scanning step size is 32 bits, and the time offset is accumulated to Dtime in turn until Dtime is equal to the specified time, and record the offset times Mcnt; offset Pindex backward Mcnt times, and the offset step size is 64 bits, that is, obtain the data stored at the specified time, read the data and return.

[0020] Preferably, the query service module and the query parsing service module can also be combined together, and the query service and the query parsing service are compatible and integrated with each other.

[0021] A unified storage and query system for multi-sensor data includes a receiving service module, a data parsing service module, a writing service module, a device registration service module, a query service module, a query parsing module, and a data reading service module; the receiving service module is connected to the data parsing service module; the data parsing service module is respectively connected to the writing service module and the device registration service module; the query service module is respectively connected to the query parsing module and the data reading service module; the query parsing module is connected to the data reading service module.

[0022] The beneficial effects of the present invention are as follows: (1) The present invention can access sensor devices at any time without manual intervention; (2) The present invention can set and modify sensor groups at any time without affecting data reception and query; (3) Different sensor data adopt a unified storage method, and horizontal expansion can be easily achieved; (4) The present invention makes full use of the powerful storage capacity of Hadoop and realizes flexible data query at the same time. Brief Description of the Drawings

[0023] Figure 1 It is a schematic diagram of the data file format of the present invention;

[0024] Figure 2 It is a schematic diagram of the data logical structure of the present invention;

[0025] Figure 3 It is a schematic diagram of the data physical structure of the present invention;

[0026] Figure 4 It is a schematic diagram of the system method of the present invention;

[0027] Figure 5 It is a schematic diagram of the working process of the data receiving service module of the present invention;

[0028] Figure 6 It is a schematic diagram of the working process of the data parsing service module of the present invention;

[0029] Figure 7 It is a schematic diagram of the working process of the writing service module of the present invention;

[0030] Figure 8 It is a schematic diagram of the maintenance method process of the index of the present invention. Detailed Embodiments

[0031] The present invention will be further described below in conjunction with specific embodiments, but the protection scope of the present invention is not limited thereto:

[0032] Embodiment: The present invention is based on a big data storage engine of Hadoop 2.0 or higher, and uses data files and corresponding index file groups to describe sensor data. The data file uses a data header + data method to store data in the HDFS file of Hadoop, where the device description data (information) is the information attached when the sensor sends data. The format of the data file is as follows: Figure 1 shown.

[0033] In actual scenarios, the sensor may be outdated, scrapped, under maintenance, or have lost transmission data. The collected detection data may not be strictly evenly distributed in time, and the corresponding sensor data cannot be inferred by time difference. For this reason, we have designed a separate index file group to store the data collection time in the sensor data and the storage location of the collected data in the data file. The index stores the time point and the location of the data monitored by the corresponding sensor in the data file. For ease of understanding, its logical structure is as follows: Figure 2 shown.

[0034] Since ordinary indexes use the form of "key+position", when the data volume is too large, the index volume itself also increases, and the efficiency of index scanning also decreases. This solution uses an improved indexing method, which to a certain extent suppresses the rapid growth of the index file volume, while reducing the sequential scanning operation of the index file and improving the performance of data retrieval. In addition, given the feature that the HDFS file system does not allow file modification, without considering the high memory consumption method of caching, after weighing the pros and cons, the physical structure is as follows: Figure 3 shown.

[0035] A fixed-length array is used in the date index to store the index position of the first monitored data of the day. If the data of the day does not exist, -1 is stored, indicating that there is no data on that day. The year is stored using a 4-byte (32-bit) identifier + a 4-byte value (32-bit); the index position is stored using a long integer of 8 bytes (64 bits).

[0036] Among them, 1. The purpose of using a 4-byte (32-bit) identifier + a 4-byte value for the year is to make up 64 bits, so that 64 bits can be read each time, increasing data reading efficiency. 2. The purpose of adding an identifier in front of the year is to determine whether the read data is a year, which can be used to scan the nearest year from back to front and then jump to the specified year. 3. The purpose of using a fixed length is to be able to calculate the starting position of the next year based on the year, which is used to locate the position of any year.

[0037] The data index stores the position of the first monitored data of the day in the data file. Immediately following it is the position of the first monitoring time in the time index. Then comes the offset of the subsequent monitored data relative to the storage position of the first monitored data, which is saved using a long integer (64 bits). Suppose the storage position of the first monitored data on a certain day is "123456789" and the length of the detected data is fixed at 4. Then the content of the index file is "123456789, 4, 4, 4…". For data types such as pictures, videos, and audios that do not have a fixed size, the content of the index file may be "123456789, 14, 8, 35…". This indexing method can greatly reduce the volume of data indexes using compressed storage.

[0038] The time index stores the time of the first monitoring of the day (64 bits). Immediately following it is the time difference (in seconds) relative to the previous monitoring time of the day, which is saved using an integer (32 bits) value. Similarly, for sensors with a stable acquisition period, there is also a large compression space in the time index file.

[0039] In addition, for hazard sources and key protection objects, generally multiple sensing and control devices are used for joint monitoring. For scenarios where it may be necessary to retrieve the data of multiple sensing and control devices simultaneously, we group these sensing and control devices. The grouping information includes: grouping ID, sensor identifier, hazard source, name of the key protection object, etc. The grouping information is stored in the registration information.

[0040] In addition to storing the grouping information, the registration information also stores information such as sensor identification, sensor name, sensor type, and device registration code.

[0041] The data file is named using the device registration code. In this way, there are as many data files and corresponding index file groups as there are sensors. Such a design is sufficient to handle various large-scale emergency command scenarios even with thousands of sensors in Hadoop.

[0042] The data item information uses an XML structure, and the content includes information such as FieldName and DataType, as shown below:

[0043]

[0044] Such as Figure 4As shown in the figure, a unified storage and query system for multi-sensor data consists of seven service modules, including a receiving service module, a data parsing service module, a writing service module, a device registration service module, a query service module, a query parsing module, and a data reading service module. The receiving service module is connected to the data analysis service module. The data parsing service module is respectively connected to the writing service module and the device registration service module. The query service module is respectively connected to the query parsing module and the data reading service module. The query parsing module is connected to the data reading service module. By decoupling the relationships between service modules, it avoids data loss caused by reasons such as service module exceptions and slow processing, and at the same time plays the role of "peak shaving and valley filling". In the present invention, logs are added between service modules during the process of data storage in the database. To prevent the log file from expanding infinitely, the size limit of the log file can be set according to the actual situation. Using the log method is not a necessary means of the present invention.

[0045] A unified storage method for multi-sensor data includes the following steps:

[0046] Step 1: The process of the data receiving service module is as Figure 5 shown. After the data receiving service module receives the detection data from the sensor, the data is appended to the end of the "***_source_log serial number.log" log file, where the log serial number starts from 1. When the size of the log file expands to the set size limit, the serial number in the new log file name increases. "***" is the process number or thread number of the data receiving service module. After the data receiving service writes the log, it sends the corresponding log file name to the data parsing service in the form of a message. As an optimization solution, the data receiving service module can use a multi-threaded method to receive data. At this time, "***" is the thread number, that is, one thread writes one log file.

[0047] Step 2: After receiving the message, the data parsing service module obtains the name of the log file and determines whether the file has been opened. For the multi-threaded mode, it checks whether the file name is in its own file list "openFileList". If the file names are inconsistent, for example, the currently opened log file is "123_source_1.log" and the received log file name is "123_source_2.log", it determines whether the currently opened file "123_source_1.log" has reached the end of the file. If so, it indicates that the file has been parsed, and then opens the new log file "123_source_2.log" for parsing, and deletes the log file "123_source_1.log" at the same time. Otherwise, it puts the file name "123_source_2.log" into the cache and waits to be read, continues to read and parse the data in "123_source_1.log" until the end of the file, and then takes out the file name "123_source_2.log" from the cache and opens the file for parsing. As an optimization solution, after receiving the message, the data parsing service splits a new process to open and parse the file "123_source_2.log". For the multi-threaded mode, it maintains a file list "openFileList" to record the currently opened files. If there are more than two file names in the file list that are the same except for the file serial numbers, it deletes the file that has reached the end of the file and has a smaller file serial number.

[0048] The data parsing service uses the SensorML model to parse the data in the log file and extracts the device data description information and the detected data. In particular, it extracts the key information from the data description information, such as: the device identifier is "abcdefg", the device name is "Meteorological Sensor No. 12345", the data items are "temperature, pressure, windSpeed, windDirection", the data types are "double, double, double, double", the sensor coordinates are "(x, y)", and the detection time is "2016-08-11T19:58:00.629Z". It sends this key information to the device registration service.

[0049] The device registration service maintains a device list "sensorList" that stores device identifiers, device names, sensor coordinates, data items, and registration codes. Based on this information, the device registration service determines whether the device has been registered. If not, it writes the key information into "sensorList", assigns the registration code "reg_123456789", records the monitoring time "2016-08-11T19:58:00.629Z", and then returns the registration code to the data parsing service. If the device has been registered, it directly returns the registration code and the last monitoring time of the device. Its workflow is as Figure 6 shown.

[0050] The data parsing service module writes the registration code "reg_123456789", device description information, detection time "2016-08-11T19:58:00.629Z", and detected data "19.451170409324014, 1012.6821277249093, 5.697366000667501, 6.327918472858433" into the parsing log. The name of the parsing log is "***_analyze_log sequence number.log", where the log sequence number starts from 1. When the size of the log file expands to the set size limit, the sequence number in the new log file name increases. "***" is the process number or thread number. Finally, the name of the parsing log file is sent to the data writing service in the form of a message.

[0051] Step 3: The writing service module receives the message and, in the same way as in Step 2, determines new files and deletes old files. Its workflow is as Figure 7 shown.

[0052] The writing service module reads the device registration code "reg_123456789" and detection data from the parsing log file and appends them to the HDFS data file "reg_123456789.data" named after the registration code. It writes the writing position of the detection data, such as "6514", and the monitoring time "2016-08-11T19:58:00.629Z" into the index file group. The index file group includes the date index file "reg_123456789.date", data index "reg_123456789.index", and time index file "reg_123456789.time". If the file does not exist, it means that data of a new device has been received. For a new device, it creates a data file named after the registration code and the corresponding index file group, writes the device description information and data item information into the data header of the data file, then appends the monitoring data to the data file, and at the same time writes the monitoring time and the writing position of the monitoring data into the index file group.

[0053] Among them, the process of maintaining the index is as follows Figure 8 as shown below:

[0054] After obtaining information such as the registration code, the position Pdata of the monitoring data in the data file, the monitoring time, and the last detection time, it is judged whether the current monitoring time is the first data of the day (new date Dn) according to the monitoring time and the last monitoring time, and different operations are performed according to whether it is a new date:

[0055] If it is a new date, append the position of the current monitoring data in the data file to the end of the data index file, record the current position Pindex of the data index, append the current monitoring time to the end of the time index, record the current position PTime of the time index, and then append PTime to the end of the data index; at the same time, according to whether the current monitoring time matches the last year in the date index, if it is the same year, directly append the data index position Pindex to the day before Dn-1 in the date index, if the day before Dn-1 does not exist, append -1 to the day after Dn-2 to get Dn-1 and then append Pindex to Dn-1, and so on until January 1st; if it is not the same year, append the year and year identifier to the date index and then write Pindex;

[0056] If it is not a new date, compare the current monitoring time with the last monitoring time to obtain the time difference DIFn (seconds), append the time difference DIFn to the end of the time index, write to the data file in the data file, and obtain the data file offset DIFp according to the position before writing and the position after writing, and append the offset DIFp to the data index.

[0057] A unified query method for multi-sensor data. In this embodiment, the query parsing service is merged into the query service in the query implementation case. Suppose it is required to query the data of the sensor numbered "Meteorological Sensor 12345" with the coordinates "(x, y)" at a certain place on August 11, 2016. The implementation case of the query is as follows:

[0058] Step 1: After receiving the query request, the query service sends the name "Meteorological Sensor 12345" to the device registration service. The device registration service queries whether the device exists in the sensorList according to the device name. If it does not exist, it returns an exception message to the query service. Otherwise, it returns the registration code to the query service. The query service sends the registration code "reg_123456789" and the query time "2016-08-11" to the data reading service.

[0059] Step 2: The data reading service reads the scan of "2016" in "reg_123456789.date" according to the registration code "reg_123456789". If it cannot be found, an exception is returned. Assuming that after finding the position corresponding to 2016, it is calculated that the position PIndex1 of "2016-08-11" is obtained by shifting backward 223 times according to the date. Continue to shift backward until a position PIndex2 that is not -1 is encountered. According to PIndex1 and PIndex2, the corresponding data positions Pdata1 and Pdata2 are read from the data index. These two values are the start position and the end position (the start position of the next day) where the monitoring data of "2016-08-11" is stored in the data file. The data between these two positions is read and converted into corresponding numerical values in sequence according to the data item information in the file header to obtain all the monitoring data of that day. According to PIndex1 and PIndex2, the corresponding data positions Ptime1 and Ptime2 are read from the time index. These two values are the start position and the end position (the start position of the next day) where the monitoring time of "2016-08-11" is stored in the time index. The data at the start position is read to obtain the first monitoring time and scanned backward to the Ptime2 position in sequence. The monitoring time corresponding to the monitoring data of that day is calculated one by one, and then the monitoring data and the monitoring time are returned to the data query service.

[0060] The above are the specific embodiments of the present invention and the technical principles applied. If changes are made according to the concept of the present invention and the functions and effects generated do not exceed the spirit covered by the specification and the drawings, they should still fall within the protection scope of the present invention.

Claims

1. A unified storage method for multi-sensor data, characterized in that, it includes the following steps: (1) The receiving service module uniformly receives data. After writing the data into a log file, the receiving service module sends the data to the data parsing service module; the data sending method is sending in the form of a message. (2) After receiving the message, the data parsing service module obtains the log file name and determines whether the file has been opened. If so, it continues to read the file until the end of the file; otherwise, it determines whether the currently opened file has reached the end of the file. If so, it opens a new log file for parsing and deletes the previously parsed log file at the same time. Otherwise, it puts the new log file name into the cache and waits to be read, continues to read the currently opened log file until the end of the file and deletes the file, then extracts the new file name from the cache and opens the file for reading; and sends the name of the parsed log file to the data writing service module; the sending method is sending in the form of a message. (3) After receiving the name of the parsed log file, the data writing service module appends the device registration code and detection data in the parsed log file to the HDFS data file named by the registration code and maintains the information in the index file.

2. The unified storage method for multi-sensor data according to claim 1, characterized in that: In step (1), when the size of the log file expands to the set size limit, the data is written into a new log file. After the data receiving service module writes the log, it sends the corresponding log file name to the data parsing service module in the form of a message. The receiving service module uses a multi-threaded method to receive data. In principle, one thread corresponds to one log file.

3. The unified storage method for multi-sensor data according to claim 1, characterized in that: After receiving the message in step (2), the data parsing service module forks a new process to open and parse the new log file, and maintains a file list for recording the currently opened files; The data parsing service module extracts the device data description information and the detected data from the log file, and extracts the key information from the data description information. The key information includes device identifier, name, sensor type, sensor coordinates, data item, data type. It sends these key information and the monitoring time to the device registration service module. The device registration service module determines whether the device has been registered according to the key information. If not, it records the key information, writes the monitoring time and assigns a registration code, and returns the registration code to the data parsing service module; if it has been registered, it returns the registration code and the last detection time to the data parsing service module, and then writes the current monitoring time; the device registration service module sends the registration code, device description information, detected data, monitoring time, and last monitoring time to the writing service module.

4. The unified storage method for multi-sensor data according to claim 1, characterized in that: After obtaining the name of the parsed log file in step (3), it is determined whether the currently opened parsed log file is the same as the file name in the message. If the file names are inconsistent, it is determined whether the currently opened file has reached the end of the file. If so, a new log file is opened for parsing, and the previously parsed log file is deleted. Otherwise, the name of the new parsed log file is placed in the cache for reading, and the currently opened parsed log file is continued to be read; after receiving the message, the data parsing service module forks a new process for the new log file to open the new log file and maintains a file list to record the opened files. The data writing service module reads the registration code and detection data of the device from the parsed log file and appends them to the HDFS data file named after the registration code, and maintains the index information. Except for the different suffixes, the file name of the index file group is the same as that of the data file; if the file does not exist, it means that the data of a new device has been received. For the new device, a data file and an index file group named after the registration code are created, and the device description information, data item information, and data acquisition frequency are written into the data header of the data file, and then the monitoring data is appended to the data file, and the index information is maintained at the same time.

5. A unified storage method for multi-sensor data according to claim 4, characterized in that: The specific way of maintaining the index is as follows: after obtaining the registration code, the position Pdata of the monitoring data in the data file, the monitoring time, and the last detection time information, it is judged whether the current monitoring time is the first data of the day according to the monitoring time and the last monitoring time. The first data of the day is the new date Dn, and different operations are performed according to whether it is a new date: If it is a new date, the position of the current monitoring data in the data file is appended to the end of the data index file, the current position Pindex of the data index is recorded, the current monitoring time is appended to the end of the time index, the current position PTime of the time index is recorded, and then PTime is appended to the end of the data index; at the same time, according to whether the current monitoring time matches the last year in the date index, if it is the same year, the data index position Pindex is directly appended to the day before Dn-1 in the date index. If the day before Dn-1 does not exist, -1 is appended to the day after Dn-2 to get Dn-1, and then Pindex is appended to the end of Dn-1, and so on until January 1st; if it is not the same year, the year and year identifier are appended to the date index, and then Pindex is written. If it is not a new date, the current monitoring time is compared with the last monitoring time to obtain the time difference DIFn. The unit of the time difference DIFn is seconds. The time difference DIFn is appended to the end of the time index, the data file is written in the data file, and the data file offset DIFp is obtained according to the position before writing and the position after writing. The offset DIFp is appended to the data index.

6. A unified query method for multi-sensor data, characterized in that, including the following steps: 1) After the query service module receives a query request, it sends the query conditions to the query parsing service module; 2) The query parsing service module parses out the device identifier, sensor name, sensor group name, and monitoring time information based on the query conditions, and determines whether this information exists in the device registration service module. If it does not exist, it returns an exception message. Otherwise, it sends the obtained registration code and monitoring time to the query service module; 3) The query service module sends the device registration code and monitoring time obtained in step 2) to the data reading service module; 4) The data reading service module jumps to the specified year and date positions in the corresponding date index to obtain the position Pindex in the data index. Pindex points to the position where the data of the first monitoring on the current day is stored in the data file. Immediately following Pindex is the position Ptime of the time of the first monitoring on the current day in the time index; according to Ptime, it reads the corresponding time Dtime from the time index, and scans backward by the time offset. The scanning step size is 32 bits, and the time offset is successively added to Dtime until Dtime is equal to the specified time, and the number of offset times Mcnt is recorded; Pindex is offset backward Mcnt times, and the offset step size is 64 bits, that is, the data stored at the specified time is obtained, and the data is read and returned.

7. A unified query method for multi-sensor data according to claim 6, characterized in that: The query service module and the query parsing service module can also be combined together, and the query service and the query parsing service are mutually compatible and integrated.

8. A unified storage and query system for multi-sensor data applying the method according to claim 1, characterized in that, it includes a receiving service module, a data parsing service module, a writing service module, a device registration service module, a query service module, a query parsing module, and a data reading service module; the receiving service module is connected to the data parsing service module; the data parsing service module is respectively connected to the writing service module and the device registration service module; the query service module is respectively connected to the query parsing module and the data reading service module; the query parsing module is connected to the data reading service module.

Citation Information

Patent Citations

  • Multisensor-based multivariate and heterogeneous data steam automatic storage method

    CN104750814A

  • Real-time streaming data storage method and device

    CN108694187A

  • A real-time storage method for unstructured flow data of rail transit

    CN109947896A

  • High-performance file storage and management system based on HDFS

    CN110647497A

  • Data storage method, automobile data recorder, server and storage medium

    CN110737807A