Trusted data management system and method based on multiple acquisition and multiple storage

Through the combination of multiple data acquisition modules, multiple storage and real-time processing, the problems of insufficient comprehensive data and lag in real-time processing in the data management platform are solved, and efficient, accurate and timely processing of data is achieved.

CN120353783APending Publication Date: 2025-07-22陈思恩
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510227128.X
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-02-27
Publication Date
2025-07-22

AI Technical Summary

Technical Problem

The existing data management platform has problems such as insufficient data comprehensiveness, data exceptions and missing in the data processing process. It especially ignores unstructured or semi-structured data sources and is unable to process real-time data in a timely manner, resulting in reading errors and delays in operation processes.

Method used

Multiple data acquisition modules are used to automatically identify and parse data formats from different sources and convert them into a unified format; combined with the multi-data storage module, use MySQL, Redis, HBase and Hive databases, and select the most appropriate storage method according to data characteristics; use the Kafka platform to diversion and processing of emergency data through the data shunt module; and the real-time processing module quickly processes key data.

Benefits of technology

It realizes more comprehensive data acquisition and storage, improves data integrity, accuracy and consistency, reduces data backlog, and ensures rapid processing and feedback of real-time data.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120353783A_ABST
    Figure CN120353783A_ABST
Patent Text Reader

Abstract

The invention discloses a credible data management system and method based on multiple acquisition and multivariate storage. The credible data management system comprises a multiple data acquisition module, a multivariate data storage module, a data distribution module, a data preprocessing and quality assurance module and a message queue and real-time processing module. The various data acquisition modules are respectively connected with the multivariate data storage module and the data distribution module, and are used for automatically identifying and analyzing data formats of different sources and converting the data formats into uniform data formats in the system; the output end of the data distribution module is respectively connected with the data preprocessing and quality assurance module and the message queue and real-time processing module; the output end of the data preprocessing and quality assurance module is connected with the multivariate data storage module; the output end of the message queue and real-time processing module is connected with the multivariate data storage module, and the message queue and real-time processing module is used for processing data needing to be processed in real time.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of data governance, and particularly to a trusted data governance system and method based on multiple collections and diverse storages. Background Art

[0002] With the rapid development of information technology, data has become the core resource for enterprise decision-making and operation. However, there are several significant defects in the existing data management platforms in the data processing process. Specifically, most platforms are limited to a single data collection method, such as only supporting the import of relational databases, ignoring the importance of unstructured or semi-structured data sources such as Internet of Things devices, social media, and log files, resulting in insufficient comprehensiveness of the collected data; in addition, there are often a series of data problems during the data collection process, such as data anomalies, data missing, etc. If this part of the data is directly stored in the database, it is easy to cause data reading errors or inability to obtain correct data information from the read data when reading; moreover, for data that needs to be processed in real time, if the data cannot be processed and stored in time, the subsequent operation process will be delayed.

[0003] In view of this, the inventor of the present invention has specifically designed a trusted data governance system and method based on multiple collections and diverse storages, and this case is thus produced. Summary of the Invention

[0004] The purpose of the present invention is to provide a trusted data governance system and method based on multiple collections and diverse storages. The invention has a more complete data collection mode and stores the data into different types of databases according to the characteristics and requirements of the collected data; in addition, after the data enters the platform, a series of preprocessing will be performed on the data to improve the integrity, accuracy, and consistency of the data content; there is a real-time processing module in the system, which will perform a processing method of producing and consuming data simultaneously on the data entering the processing queue, reducing the backlog of urgently needed data to be processed.

[0005] To achieve the above object, the present invention adopts the following technical solutions:

[0006] A trusted data governance system based on multiple collections and diverse storages, including multiple data collection modules, diverse data storage modules, a data shunting module, a data preprocessing and quality assurance module, and a message queue and real-time processing module; the multiple data collection modules are respectively connected to the diverse data storage modules and the data shunting module, and the multiple data collection modules are used to automatically identify and parse data formats from different sources and convert them into a unified data format within the system; the output end of the data shunting module is respectively connected to the data preprocessing and quality assurance module and the message queue and real-time processing module; the output end of the data preprocessing and quality assurance module is connected to the diverse data storage module; the output end of the message queue and real-time processing module is connected to the diverse data storage module, and the message queue and real-time processing module is used to process data that needs to be processed in real time.

[0007] Further, the multiple data collection modules include HTTP requests, TCP transmissions, file synchronization, identification and parsing, and data conversion; HTTP requests are used to obtain data from a Web server or an API interface; TCP transmissions are used to collect data from real-time data streams or device data; file synchronization is used for the import of batch data; identification and parsing are used to identify and parse the data formats of the collected data; data conversion is used to convert the parsed data into a unified data format.

[0008] Further, the diverse data storage modules include intelligent storage management, MySQL, Redis, HBase, and Hive; intelligent storage management is used to analyze the storage scheme of the content input into the database and deliver it to the specified type of database; MySQL is used to store structured data; Redis is used to store data with fast read requirements; HBase is used to store large-scale unstructured and semi-structured data; Hive is used to store data with high IO read requirements.

[0009] Further, the operations for processing data in the data preprocessing and quality assurance module include data cleaning, thinning, time factor indexing, and spatial indexing; the data preprocessing and quality assurance module delivers the processed data to the diverse data storage modules.

[0010] Further, the operations for processing data in the message queue and real-time processing module include data cleaning, thinning, time factor indexing, and spatial indexing; the message queue and real-time processing module delivers the processed data to the diverse data storage modules.

[0011] Further, the data shunting module shunts the data that needs to be processed after collection to the message queue and real-time processing module or the data preprocessing and quality assurance module through the kafka open-source streaming platform.

[0012] A trusted data governance method based on multiple collections and multiple storages, comprising the following steps:

[0013] S1. Collect data through various collection methods in multiple data collection modules, identify and parse the collected data, convert the data format after parsing into a unified data format within the system, and then deliver the converted data to the data shunt module or the multiple data storage module according to requirements;

[0014] S2. The data delivered to the data shunt module is shunted again through the Kafka open-source streaming platform according to the urgency requirements of the data. The data that needs to be processed urgently flows into the message queue and the real-time processing module, and other data flows into the data preprocessing and quality assurance module;

[0015] S3. Perform data processing such as data cleaning, thinning, time indexing, and space indexing on the data input to the message queue and real-time processing module or the data preprocessing and quality assurance module, and then deliver the processed data to the multiple data storage module;

[0016] S4. The data delivered to the multiple data storage module first selects the optimal storage scheme through intelligent storage management, and then delivers the data to the database of the selected type in the scheme.

[0017] In the above step S3, when there is data that does not meet the requirements, the data preprocessing and quality assurance module will mark or automatically repair this part of the data.

[0018] In the above step S4, the storage scheme selected by the intelligent storage management inside the multiple data storage module is to balance the data reading speed, data writing performance, and the storage space size occupied by the data.

[0019] After adopting the above technical solution, the present invention has the following beneficial effects:

[0020] 1. The system of the present invention supports obtaining data from multiple sources (such as web servers, API interfaces, real-time data streams, device data, batch files, etc.), and can automatically identify and parse data in different formats, and convert them into a unified data format, making the collected data more comprehensive and diverse.

[0021] 2. The present invention realizes the efficient shunting of data by using an open-source streaming platform such as Kafka. According to the urgency and nature of data processing, the data is directed to different processing paths, such as real-time processing or preprocessing and quality assurance, to meet the requirements of different types of data processing.

[0022] 3. The present invention adopts a diversified storage architecture, combines various database technologies such as MySQL, Redis, HBase, and Hive, and selects the most suitable storage method for data with different characteristics. For example, structured data is suitable for being stored in MySQL; data that needs to be read quickly can be put into Redis; large-scale unstructured or semi-structured data is more suitable for HBase; and data with high IO requirements is considered to use Hive. In addition, the intelligent storage management system can analyze and determine the best storage strategy according to the actual situation.

[0023] 4. For data scenarios that require immediate response, the system of the present invention provides a dedicated message queue and real-time processing module, enabling key information to be processed and fed back within an extremely short time, which is very important for application scenarios that rely on real-time data analysis. Brief Description of the Drawings

[0024] Figure 1 It is a schematic structural diagram of the system of the present invention;

[0025] Figure 2 It is a schematic diagram of the working process of the present invention.

[0026] The reference signs in the drawings are represented as:

[0027] 1. Multiple data acquisition modules; 2. Diversified data storage modules; 3. Data shunting modules; 4. Data preprocessing and quality assurance modules; 5. Message queue and real-time processing modules. Detailed Embodiments

[0028] In order to make the objectives, technical solutions and advantages of the present invention clearer and more understandable, the present invention will be further described in detail below with reference to the drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present invention and are not used to limit the present invention.

[0029] Please refer to Figure 1 and Figure 2, A trusted data governance system based on multiple collections and diverse storage, including multiple data collection modules 1, diverse data storage modules 2, data shunt modules 3, data preprocessing and quality assurance modules 4, and message queues and real-time processing modules 5; the multiple data collection modules 1 are respectively connected to the diverse data storage modules 2 and the data shunt modules 3, and the multiple data collection modules 1 are used to automatically identify and parse different source data formats and convert them into a unified data format within the system; the output ends of the data shunt modules 3 are respectively connected to the data preprocessing and quality assurance modules 4 and the message queues and real-time processing modules 5; the output end of the data preprocessing and quality assurance module 4 is connected to the diverse data storage modules 2; the output end of the message queues and real-time processing modules 5 is connected to the diverse data storage modules 2, and the message queues and real-time processing modules 5 are used to process data that needs to be processed in real time.

[0030] As Figure 1 shown, the multiple data collection modules 1 include HTTP requests, TCP transmissions, file synchronization, identification and parsing, and data conversion; HTTP requests are used to obtain data from Web servers or API interfaces; TCP transmissions are used to collect real-time data streams or device data; file synchronization is used for batch data import; identification and parsing are used to identify and parse the data formats of the collected data; data conversion is used to convert the parsed data into a unified data format.

[0031] As Figure 1 shown, the diverse data storage modules 2 include intelligent storage management, MySQL, Redis, HBase, and Hive; intelligent storage management is used to analyze the storage scheme of the content input into the database and deliver it to the specified type of database; MySQL is used to store structured data; Redis is used to store data with fast read requirements; HBase is used to store large-scale unstructured and semi-structured data; Hive is used to store data with high IO read requirements, and Hive is a data warehouse built on top of Apache Hadoop.

[0032] As Figure 1 shown, the operations for processing data in the data preprocessing and quality assurance module 4 include data cleaning, thinning, time indexing, and spatial indexing; the data preprocessing and quality assurance module 4 delivers the processed data to the diverse data storage modules 2, and the data preprocessing and quality assurance module 4 ensures that the data quality meets the requirements by monitoring the integrity, accuracy, and consistency of the data in real time; for data that does not meet the requirements, the system will mark or automatically repair it to avoid negative impacts on subsequent analysis.

[0033] As Figure 1As shown, the operations of processing data in the message queue and real-time processing module 5 include data cleaning, thinning, time indexing and space indexing; the message queue and real-time processing module 5 transmits the processed data to the multivariate data storage module 2; the message queue and real-time processing module 5 implements real-time data production and data digestion. This module is for data that needs to be processed in real time, such as: AIS data, etc.

[0034] Data cleaning in this system includes: 1. Missing value processing: filling or deleting missing values; 2. Outlier detection: identifying and processing outliers, such as through statistical methods or machine learning algorithms; 3. Duplicate data processing: deleting duplicate data records; 4. Format standardization: unifying data formats to ensure data consistency; 5. Data validation: checking whether the data meets the expected business rules and constraints.

[0035] The purpose of data cleaning in this system is to: check and delete completely identical or almost identical records in the database to avoid the analysis results being affected by duplicate data; for data items containing null values, choose appropriate strategies to fill these gaps according to the specific situation, such as using statistical methods such as mean, median or mode to fill, or using predictive models to estimate missing values; identify and correct inaccurate or illogical data points, which may involve format errors, spelling errors, numerical anomalies, etc.

[0036] The rarefaction in this system includes: 1. Sampling: selecting some data samples randomly or according to certain rules; 2. Aggregation: aggregating data by time period, geographic location or other dimensions to generate higher-level statistical data; 3. Filtering: filtering out data that does not meet the requirements based on specific conditions.

[0037] The purpose of rarefaction in this system is to reduce the size of the data set while retaining the key features and information of the data as much as possible. This operation is used to improve data processing efficiency, reduce storage requirements, and speed up data analysis.

[0038] The time index in this system includes: in time series data, a timestamp is assigned to each data point.

[0039] The purpose of time index in this system is: 1. Improve data access speed and query efficiency and quickly locate and analyze data; 2. Calculate the average value, maximum value and other auxiliary statistical methods within a certain period of time; 3. Extract historical data according to timestamp for modeling or simulation prediction.

[0040] The spatial index in this system includes: in the geographic spatial data, a geographic location coordinate is assigned to each data point.

[0041] The purpose of the spatial index in this system is as follows: 1. To improve the access speed and query efficiency of data and quickly locate and analyze data; 2. To find data within a specific area through geographical location coordinates, such as finding the nodes on a certain navigation route; 3. To serve as a reference point for other coordinate geographical locations; 4. To define a certain area for mathematical statistics.

[0042] As Figure 1 shown, the data shunting module 3 shunts the data that needs to be processed after collection to the message queue and the real-time processing module 5 or the data preprocessing and quality assurance module 4 through the Kafka open-source streaming platform.

[0043] As Figure 1 and Figure 2 shown, a trusted data governance method based on multiple collections and multiple storages includes the following steps:

[0044] S1. Collect data through various collection methods in the multiple data collection module 1, identify and parse the collected data, convert the data format after parsing into the unified data format within the system, and then deliver the converted data to the data shunting module 3 or the multiple data storage module 2 according to requirements;

[0045] S2. The data delivered to the data shunting module 3 is shunted again through the Kafka open-source streaming platform according to the urgency requirements of the data. The data that needs to be processed urgently flows into the message queue and the real-time processing module 5, and other data flows into the data preprocessing and quality assurance module 4;

[0046] S3. Perform data processing such as data cleaning, thinning, time indexing, and spatial indexing on the data input to the message queue and the real-time processing module 5 or the data preprocessing and quality assurance module 4, and then deliver the processed data to the multiple data storage module 2;

[0047] S4. The data delivered to the multiple data storage module 2 first selects the optimal storage scheme through intelligent storage management, and then delivers the data to the database of the selected type in the scheme.

[0048] In the above step S3, when there is data that does not meet the requirements, the data preprocessing and quality assurance module will mark or automatically repair this part of the data.

[0049] In the above step S4, the storage scheme selected by the intelligent storage management inside the multiple data storage module is to balance the data reading speed, data writing performance, and the storage space size occupied by the data.

[0050] As described above, it is only a preferred specific embodiment of the present invention, but the protection scope of the present invention is not limited thereto. Any changes or substitutions that can be easily thought of by those skilled in the art within the technical scope disclosed by the present invention should be covered within the protection scope of the present invention. Therefore, the protection scope of the present invention should be subject to the protection scope of the claims.

Claims

1. A trusted data governance system based on multiple collections and diverse storages, characterized in that: It includes multiple data acquisition modules (1), a multi-source data storage module (2), a data shunting module (3), a data preprocessing and quality assurance module (4), and a message queue and real-time processing module (5); the multiple data acquisition modules (1) are respectively connected to the multi-source data storage module (2) and the data shunting module (3), and the multiple data acquisition modules (1) are used to automatically identify and parse data formats from different sources and convert them into a unified data format within the system; the output end of the data shunting module (3) is respectively connected to the data preprocessing and quality assurance module (4) and the message queue and real-time processing module (5); the output end of the data preprocessing and quality assurance module (4) is connected to the multi-source data storage module (2); the output end of the message queue and real-time processing module (5) is connected to the multi-source data storage module (2), and the message queue and real-time processing module (5) is used to process data that needs to be processed in real time.

2. The trusted data governance system based on multiple collections and multiple storages as claimed in claim 1, characterized in that: The multiple data acquisition modules (1) include HTTP requests, TCP transmissions, file synchronization, identification and parsing, and data conversion; The HTTP request is used to obtain data from a Web server or an API interface; the TCP transmission is used to collect data from real-time data streams or device data; the file synchronization is used for the import of batch data; the identification and parsing is used to identify and parse the data formats of the collected data; the data conversion is used to convert the parsed data into a unified data format.

3. The trusted data governance system based on multiple collection and multi - element storage according to claim 1, characterized in that: The multi-source data storage module (2) includes intelligent storage management, MySQL, Redis, HBase, and Hive; the intelligent storage management is used to analyze the storage scheme of the content input into the database and deliver it to the specified type of database; MySQL is used to store structured data; Redis is used to store data with fast read requirements; HBase is used to store large-scale unstructured and semi-structured data; Hive is used to store data with high IO read requirements.

4. The trusted data governance system based on multiple collections and multiple storages according to claim 1, characterized in that: The operations for processing data in the data preprocessing and quality assurance module (4) include data cleaning, thinning, time indexing, and spatial indexing; the data preprocessing and quality assurance module (4) delivers the processed data to the multi-source data storage module (2).

5. The trusted data governance system based on multiple collection and multiple storage according to claim 1, characterized in that: The operations for processing data in the message queue and real-time processing module (5) include data cleaning, thinning, time indexing, and spatial indexing; the message queue and real-time processing module (5) delivers the processed data to the multi-source data storage module (2).

6. The trusted data governance system based on multiple acquisitions and multiple storages according to claim 1, characterized in that: The data shunting module (3) shunts the data that needs to be processed after collection to the message queue and real-time processing module (5) or the data preprocessing and quality assurance module (4) through the kafka open-source streaming platform.

7. A trusted data governance method based on multiple collections and diverse storages, characterized in that, It includes the following steps: S1. Collect data through various collection methods in the multiple data acquisition modules (1), identify and parse the collected data, convert the data format after parsing into a unified data format within the system, and then deliver the converted data to the data shunting module (3) or the multi-source data storage module (2) according to requirements; S2. The data sent to the data shunt module (3) is shunted again through the Kafka open source streaming platform according to the urgency requirements of the data. The data that needs to be processed urgently flows into the message queue and the real-time processing module (5), and other data flows into the data preprocessing and quality assurance module (4). S3. The data input into the data preprocessing and quality assurance module (4) or the message queue and real-time processing module (5) is processed for data cleaning, thinning, time indexing, and space indexing, and then the processed data is sent to the multi-source data storage module (2). S4. The data sent to the multi-source data storage module (2) first selects the optimal storage scheme through intelligent storage management, and then sends the data to the database of the selected type in the scheme.

8. The trusted data governance method based on multiple acquisitions and multiple storages according to claim 7, characterized in that: When there is data that does not meet the requirements, the data preprocessing and quality assurance module (4) will mark or automatically repair this part of the data.

9. The trusted data governance method based on multiple acquisitions and multiple storages according to claim 7, characterized in that: The storage scheme selected by the intelligent storage management inside the multi-source data storage module (2) is to balance the data reading speed, data writing performance, and the storage space occupied by the data.