Data management method and system based on big data

By adopting comprehensive methods of data collection, cleaning, storage, analysis and security modules in big data systems, data quality, security and integration problems in big data are solved, and efficient data governance and full mining of value are achieved.

CN120045553APending Publication Date: 2025-05-27BEIJING HANGYUN SCI & TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510194040.2
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-02-21
Publication Date
2025-05-27

AI Technical Summary

Technical Problem

The existing technology is difficult to effectively solve data quality problems, data security problems, and data integration and sharing problems in big data, making it difficult to ensure the accuracy, security and availability of data.

Method used

A data governance method and system based on big data is adopted, including data collection, data cleaning, data storage, data analysis and mining, data security and other modules. Specific technical means include data acquisition using network crawlers and database connectors, data deduplication and outlier detection of hashing algorithms and statistics, efficient storage of distributed file systems and columnar storage databases, in-depth analysis of machine learning algorithm libraries and distributed computing frameworks, and data security guarantees of AES encryption algorithms and role-based access control models.

Benefits of technology

By automatically detecting and correcting data quality problems, ensuring data security, realizing data standardization and seamless integration, significantly improving data availability and value, and meeting the needs of enterprises for real-time and accuracy.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120045553A_ABST
    Figure CN120045553A_ABST
Patent Text Reader

Abstract

The invention discloses a data governance method and system based on big data. The method comprises the following steps: S1, a data acquisition module; s2, a data cleaning module; s3, a data storage module; s4, a data analysis and mining module; and S5, a data security module. The invention belongs to the technical field of data storage, particularly relates to a data management method and system based on big data, and has the beneficial effects that the data quality is remarkably improved, the data security is guaranteed, the data processing efficiency is greatly improved, and the data value is fully mined.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of data storage, and specifically refers to a data management method and system based on big data. Background Art

[0002] In the digital age, big data has become an important asset for the development of enterprises and organizations. With the widespread application of technologies such as the Internet of Things and mobile Internet, the amount of data has grown exponentially. However, large amounts of data bring not only opportunities, but also severe challenges.

[0003] From the perspective of data quality, data duplication, errors, missing data and other issues are common. For example, in a customer information management system, the same customer's information may be repeatedly entered and the content may be different due to inconsistent data standards collected from different channels. This not only wastes storage space, but also misleads decision-making analysis. From the perspective of data security, data leakage incidents occur frequently. For example, a well-known e-commerce platform once leaked a large amount of user information due to security vulnerabilities, causing huge losses to users and enterprises. In terms of data integration and sharing, the data formats, encoding methods, and data structures of different business systems are different, making it difficult for data to be effectively circulated and integrated within the enterprise, forming "data islands".

[0004] Traditional data governance methods mainly rely on manual operations and simple tools. When faced with massive and highly complex data, they are extremely inefficient and cannot meet the requirements of real-time and accuracy. For example, manual data cleaning is not only time-consuming and laborious, but also prone to omissions and errors. Therefore, it is urgent to develop an advanced, efficient, and intelligent data governance method and system. Summary of the invention

[0005] The core purpose of the present invention is to provide a comprehensive, efficient and intelligent data governance method and system based on big data. The system can automatically detect and correct data quality problems, ensure data security, achieve data standardization and seamless integration, thereby improving the availability and value of data, and providing accurate and reliable data support for enterprise decision-making, business optimization, risk prediction, etc.

[0006] In order to solve the above technical problems, the technical solution adopted by the present invention is as follows:

[0007] A data governance method and system based on big data, comprising the following steps:

[0008] S1, data acquisition module;

[0009] S2, data cleaning module;

[0010] S3, data storage module;

[0011] S4, Data Analysis and Mining Module;

[0012] S5, Data Security Module.

[0013] Preferably, the S1, Data Acquisition Module includes a web crawler and a database connector, and the specific method is as follows:

[0014] S1-1, The web crawler uses the Scrapy framework of Python as the web crawler tool.

[0015] S1-2, For relational databases, the database connector uses MySQL Connector / Python as the connector.

[0016] Preferably, the S2, Data Cleaning Module includes a data deduplication component and an outlier detection module, and the specific method is as follows:

[0017] S2-1, The data deduplication component uses the hash algorithm to implement the data deduplication function.

[0018] S2-2, The outlier detection model constructs an outlier detection model based on the 3σ principle of statistics.

[0019] Preferably, the S3, Data Storage Module includes a distributed file system and a columnar storage database, and the specific method is as follows:

[0020] S3-1, The distributed file system uses Hadoop Distributed File System (HDFS) as the infrastructure for big data storage.

[0021] S3-2, The columnar storage database uses HBase as the columnar storage database, which is built on HDFS and is suitable for storing massive, sparse structured data.

[0022] Preferably, the S4, Data Analysis and Mining Module includes a machine learning algorithm library and a distributed computing framework, and the specific method is as follows.

[0023] S4-1, The machine learning algorithm library (Scikit-learn) uses the integrated Scikit-learn machine learning algorithm library and can provide rich machine learning algorithms and tools.

[0024] S4-2, The distributed computing framework (Spark) uses Apache Spark as the distributed computing framework, which is based on in-memory computing and greatly improves the data processing speed.

[0025] Preferably, the S5, Data Security Module includes an encryption algorithm and an access control component, and the specific method is as follows.

[0026] S5-1. The encryption algorithm uses the Advanced Encryption Standard (AES) to encrypt sensitive data.

[0027] S5-2. The access control component is constructed based on the Role-Based Access Control (RBAC) model.

[0028] The beneficial effects achieved by the present invention with the above structure are as follows:

[0029] 1. The data quality is significantly improved: Through the deduplication component and the outlier detection model in the data cleaning module, it can automatically identify and process duplicate, incorrect, and abnormal situations in the data, improve the accuracy and integrity of the data, and provide a reliable data basis for subsequent data analysis and decision-making.

[0030] 2. The data security is guaranteed: The AES encryption algorithm is used to encrypt sensitive data, combined with the access control component based on the RBAC model, to ensure the security of the data from two aspects of data encryption and permission management, effectively preventing data leakage and illegal access, and protecting the privacy of enterprises and users.

[0031] 3. The data processing efficiency is greatly improved: By using the distributed computing framework Spark and the distributed file system HDFS, it can achieve parallel processing and efficient storage of large-scale data, greatly shortening the data processing time, and meeting the enterprise's demand for real-time data processing.

[0032] 4. The data value is fully mined: With the help of machine learning algorithms in Scikit-learn, it can deeply analyze and mine the data after cleaning and storage, discover potential patterns and rules in the data, provide valuable decision-making basis for the enterprise's market prediction, customer relationship management, product optimization, etc., and enhance the competitiveness of the enterprise. Description of the Drawings

[0033] Figure 1 It is a flowchart of an embodiment of the present invention. Detailed Embodiments

[0034] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described in conjunction with the drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments; based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of the present invention.

[0035] As Figure 1 shown, a data governance method and system based on big data includes the following steps:

[0036] S1. Data collection module; S2. Data cleaning module; S3. Data storage module; S4. Data analysis and mining module; S5. Data security module.

[0037] Preferably, S1, the data collection module includes a web crawler and a database connector, and the specific method is as follows:

[0038] S1-1. Web crawler: Use the Scrapy framework of Python as the web crawler tool. Scrapy is built based on the Twisted asynchronous network framework and has efficient asynchronous I / O operation capabilities, enabling concurrent access to a large number of web pages in a short time. It has extremely strong scalability. By writing different Spider classes and middleware, it can easily adapt to the structures of various websites and data extraction requirements. For example, when crawling data from a news website, XPath or CSS selectors can be used to accurately locate information such as news titles, texts, and release times. At the same time, Scrapy provides a powerful scheduler and downloader, which can intelligently manage the request queue and handle various exceptions during the download process.

[0039] S1-2. Database connector: For relational databases, use MySQL Connector / Python as the connector. It is the Python driver officially provided by MySQL. By following the SQL standard, it can establish a stable and high-speed connection with the MySQL database. Complex SQL query statements can be executed using simple Python code to achieve operations such as data extraction, insertion, and update. For the non-relational database MongoDB, choose the PyMongo library. PyMongo provides rich APIs, supports document-style data operations on MongoDB, and can conveniently obtain collection and document data in the database to meet the needs of different types of data collection.

[0040] Preferably, S2, the data cleaning module includes a data deduplication component and an outlier detection module, and the specific method is as follows:

[0041] S2-1. Data deduplication component: Use the hash algorithm to implement the data deduplication function. This component first maps the key features of the data (such as user ID, order number, etc.) to a hash value of a fixed length. Common hash algorithms such as MD5 and SHA-256 are available. Here, the SHA-256 algorithm is selected because of its higher security and hash value uniqueness. When new data enters, calculate its hash value and compare it with the hash values of the stored data. If the hash values are the same, further compare the detailed content of the data to determine whether it is duplicate data. For example, for a user information table, generate a hash value based on the user ID. If the hash values of two user information are the same and other key information (such as name, contact information) is also the same, it is determined as duplicate data and deleted.

[0042] S2-2, Outlier Detection Model: An outlier detection model is constructed based on the 3σ principle of statistics. For numerical data, first calculate the mean (μ) and standard deviation (σ) of the data. According to the 3σ principle, data values within the interval (μ - 3σ, μ + 3σ) are considered normal, and data points outside this interval are determined as outliers. For example, when analyzing the sales price data of a certain commodity, if a price value is much higher or much lower than the mean plus / minus 3 times the standard deviation, it may be a data entry error or there may be special circumstances, and it can be further reviewed to decide whether to correct or delete it.

[0043] Preferably, S3, the data storage module includes a distributed file system and a columnar storage database, and the specific method is as follows:

[0044] S3-1, Distributed File System: Use Hadoop Distributed File System (HDFS) as the infrastructure for big data storage. HDFS adopts a master-slave architecture, consisting of a NameNode and multiple DataNodes. The NameNode, as the master node, is responsible for managing the namespace and metadata of the file system, recording information such as the mapping relationship between files and data blocks. The DataNode, as the slave node, is responsible for actual data storage, storing data in the form of data blocks on local disks. HDFS has high fault tolerance. When a certain DataNode fails, the data can be obtained from other replica nodes. At the same time, it has good scalability. Just add new DataNode nodes to increase the storage capacity.

[0045] S3-2, Columnar Storage Database: Use HBase as the columnar storage database. It is built on top of HDFS and is suitable for storing massive, sparse structured data. The data model of HBase is organized in the form of tables. Tables consist of rows and column families, and each column family contains multiple columns. When storing data, the data is stored by column, which makes it only necessary to read the required columns when querying data, greatly improving the query efficiency. For example, when storing user behavior data, different types of behavior data (such as browsing records, purchase records, comment records) can be stored in different column families respectively to facilitate quickly querying specific types of user behavior information.

[0046] Preferably, S4, the data analysis and mining module includes a machine learning algorithm library and a distributed computing framework, and the specific method is as follows.

[0047] S4-1. Machine Learning Algorithm Library (Scikit-learn): It integrates the Scikit-learn machine learning algorithm library and can provide a rich variety of machine learning algorithms and tools. Classification algorithms such as decision tree algorithms build a tree structure to make classification decisions on data, continuously splitting nodes according to the characteristics of the data until the termination condition is reached, and can be used in scenarios such as customer classification and risk assessment. Support vector machine algorithms classify data by finding an optimal hyperplane and perform well in small-sample and non-linear classification problems. Clustering algorithms such as the K-Means algorithm divide data into K clusters, making the data within the same cluster highly similar and the data between different clusters less similar, and can be used for user group division, market segmentation, etc.

[0048] S4-2. Distributed Computing Framework (Spark): It uses Apache Spark as the distributed computing framework. Based on in-memory computing, it greatly improves the data processing speed. Spark provides abstract data structures such as RDD (Resilient Distributed Dataset) and DataFrame. RDD is an immutable collection of distributed objects that can process data through a series of transformation operations (such as map, filter, reduce, etc.) and supports parallel computing in the cluster. DataFrame is a distributed dataset with structure information, similar to the table structure in a traditional database, providing richer operation methods and optimization strategies for convenient complex analysis and processing of large-scale data.

[0049] Preferably, S5. The data security module includes an encryption algorithm and an access control component, and the specific method is as follows.

[0050] S5-1. Encryption Algorithm: It uses the Advanced Encryption Standard (AES) to encrypt sensitive data. AES is a symmetric encryption algorithm with the characteristics of fast encryption speed and high security. During data transmission and storage, the original data (such as user ID numbers, bank card numbers, etc.) is encrypted into ciphertext through the AES algorithm. When encrypting, a key is required, and only users with the same key can decrypt the ciphertext back to the original data. For example, when transmitting user login information, the password is encrypted by AES before being transmitted over the network to prevent the password from being stolen.

[0051] S5-2. Access Control Component: Build an access control component based on the Role-Based Access Control (RBAC) model. First, define different user roles, such as administrators, ordinary users, data analysts, etc. Assign corresponding data access permissions to each role. For example, administrators have read and write permissions to all data, ordinary users can only read their own relevant data, and data analysts can read specific data sets for analysis but cannot modify the data. In this way, it is ensured that only authorized users can access specific data resources, effectively preventing illegal access and abuse of data.

[0052] The specific implementation method has the following steps: 1. System deployment; 2. Data collection configuration; 3. Data cleaning settings; 4. Data analysis and mining. 1. System deployment: Install distributed computing frameworks such as Hadoop and Spark and storage systems such as HDFS and HBase on the cluster server. Configure the network communication parameters between nodes to ensure that data can be quickly transmitted between nodes. Install the Python environment and related libraries, such as Scrapy, MySQL Connector / Python, PyMongo, Scikit-learn, etc., to provide software support for the operation of the system; 2. Data collection configuration: According to the data collection requirements, write the Spider class of the Scrapy crawler and define the data extraction rules. Configure the connection parameters of MySQL Connector / Python and PyMongo, including the database address, port, username, password, etc., to ensure that various data sources can be accurately connected and data can be collected; 3. Data cleaning settings: In the data cleaning module, set the hash algorithm parameters of the data deduplication component, such as selecting the hash function type and setting the hash value length. Configure the parameters of the outlier detection model, such as determining the data grouping method, the method of calculating the mean and standard deviation, etc., and optimize according to different data types and business requirements; 4. Data analysis and mining: According to the business problems and data characteristics, select appropriate machine learning algorithms from Scikit-learn. For example, select the decision tree algorithm when classifying customers. Build a data processing process on the Spark platform, load the data into the RDD or DataFrame format, perform data preprocessing, model training and evaluation, and store the analysis results in HBase or other databases for subsequent use.

[0053] Although the embodiments of the present invention have been shown and described, for those of ordinary skill in the art, it can be understood that various changes, modifications, substitutions, and variations can be made to these embodiments without departing from the principles and spirit of the present invention. The scope of the present invention is defined by the appended claims and their equivalents.

[0054] The above describes the present invention and its implementation manners. Such description is not restrictive. What is shown in the drawings is only one of the implementation manners of the present invention, and the actual structure is not limited thereto. In summary, if those of ordinary skill in the art are inspired by it and, without departing from the gist of the present invention, creatively design structural manners and embodiments similar to the technical solution, they shall fall within the protection scope of the present invention.

Claims

1. A data governance method and system based on big data, comprising the following steps: S1, data acquisition module; S2, data cleaning module; S3, data storage module; S4, data analysis and mining module; S5. Data security module.

2. According to the data governance method and system based on big data according to claim 1, it is characterized by: The S1, data acquisition module includes a web crawler and a database connector and the specific method is as follows: S1-1. The web crawler uses Python's Scrapy framework as a web crawler tool. S1-2. The database connector uses MySQL Connector / Python as a connector for a relational database.

3. According to the data governance method and system based on big data according to claim 1, it is characterized by: The S2, data cleaning module includes a data deduplication component and an outlier detection module and the specific method is as follows: S2-1. The data deduplication component uses a hash algorithm to implement data deduplication function. S2-2. The outlier detection model is constructed based on the 3σ principle of statistics.

4. According to a data governance method and system based on big data according to claim 1, it is characterized by: The S3 data storage module includes a distributed file system and a list storage database and the specific method is as follows: S3-1. The distributed file system uses Hadoop Distributed File System (HDFS) as the infrastructure for big data storage. S3-2. The column storage database uses HBase as the column storage database, which is built on HDFS and is suitable for storing massive, sparse structured data.

5. According to the data governance method and system based on big data according to claim 1, it is characterized by: The S4, data analysis and mining module includes a machine learning algorithm library and a distributed computing framework and the specific method is as follows. S4-1. The machine learning algorithm library (Scikit-learn) uses an integrated Scikit-learn machine learning algorithm library and can provide a wealth of machine learning algorithms and tools. S4-2. The distributed computing framework (Spark) uses Apache Spark as the distributed computing framework, which is based on memory computing and greatly improves the data processing speed.

6. The data governance method and system based on big data according to claim 1, characterized in that: The S5, data security module includes an encryption algorithm and an access control component and the specific method is as follows. S5-1. The encryption algorithm uses the Advanced Encryption Standard (AES) to encrypt sensitive data. S5-2. The access control component constructs an access control component based on a role-based access control (RBAC) model.