Big data platform system
By providing diversified data collection capabilities, automated data governance processes, cross-platform data sharing capabilities, unified system management and good scalability and customization capabilities, the problems of single data collection methods, complex governance processes, insufficient sharing capabilities and limited management capabilities in the existing technology are solved, and the diversification, efficiency, automation and unification of the big data collection and governance platform are achieved.
Patent Information
- Application Number
- CN202410931944.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2024-07-12
- Publication Date
- 2025-05-16
AI Technical Summary
In the existing technology, the data collection method is single, the data governance process is complex and lacks automation, the data sharing ability is insufficient, the system management ability is limited, the scalability and customization ability is lacking, and the lack of a unified big data collection and governance platform.
Provides diverse data acquisition capabilities, including the acquisition of structured and unstructured data, and supports access to multiple data sources. Realize the automation and intelligence of data governance, and use functional modules such as data element cleaning, standardization, management and quality management. Provide cross-platform data query, download interface and Kafka message push functions to achieve efficient data sharing. Provides unified operation and maintenance management capabilities, including data resource management, permission security management, log management and other functions. It has good scalability and customization capabilities, and supports user-customized data element collection strategies and customized data element analysis and processing logic.
It has achieved diversified and efficient data collection, automation and intelligence of data governance, cross-platform and efficient data sharing, unified and stable system management, and improved scalability and customized capabilities, solving the problems of single data collection methods, complex governance processes, insufficient sharing capabilities, and limited management capabilities in the existing technology.
Smart Images

Figure FT_1 
Figure FT_2 
Figure FT_3
Abstract
Description
[0001] Technical field: In the field of big data collection and management technology Background Art
[0002] The field of big data collection and governance technology is an emerging technology field. With the advent of the big data era, the technology in this field has developed rapidly. At present, in terms of big data collection and governance, the existing practices mainly include: Use relational databases and NoSQL databases to store and manage data; Use distributed computing frameworks such as Hadoop and Spark for data processing; Use distributed file systems such as HDFS for data storage; Use search engines such as ElasticSearch and HBase to retrieve data; Use ETL tools to clean, transform and load data; Analyze and mine data through data mining and machine learning techniques. Summary of the invention
[0003] Technical problems to be solved: In view of the defects of single data collection methods in the prior art, this patent provides diversified data collection capabilities, including the collection of structured and unstructured data, and supports the access of multiple data sources to meet the diversified needs of data collection in different scenarios. In view of the shortcomings of complex data governance processes and lack of automation, this patent provides functional modules such as data element cleaning, data element standardization, data element control and data element quality management, realizes the automation and intelligence of data governance, and improves the efficiency of data governance. In view of the problem of insufficient data sharing capabilities, this patent realizes cross-platform data query, data download interface and Kafka message push functions, provides an efficient data sharing method, and promotes the circulation and application of data resources. In view of the shortcomings of limited system management capabilities, this patent provides unified operation and maintenance management capabilities, including data resource management, authority security management, log management and other functions, realizes unified management of big data platforms, and improves the stability and security of system operation. In view of the problem of lack of support for the scalability and customization capabilities of big data collection and governance platforms, this patent has good scalability and customization capabilities, supports user-customized data element collection strategies and customized data element analysis and processing logic and processes to meet different business needs and development scenarios. In view of the shortcomings of the existing technology in big data collection and management, which is relatively scattered and lacks a unified big data collection and management platform, this patent provides a unified big data collection and management platform that integrates functional modules such as data collection, data management, data application and system management.
[0004] Technical solution: Step 1: Data collection 1. Collect diverse data, including structured relational databases, text, etc., as well as unstructured files, pictures, etc. 2. Support the input of diverse data source types, such as Oracle, MySQL, SQL Server, etc., as well as FTP, Clickhouse, Hive, etc. 3. Support real-time collection of structured data and achieve real-time data synchronization through Kafka message queues. 4. Provide high-performance file input capabilities, supporting file formats including TXT, CSV and their compressed formats. Step 2: Data governance 1. Data element cleaning, including rich cleaning components, such as deduplication, screening, conversion, etc., support script cleaning, and support output after data cleaning. 2. Data element standardization, establish a unified data standard system, and support batch import and synchronization of data dictionaries. 3. Data element management and control, support multi-type mixed storage mode, data extraction, conversion, cleaning, conversion and mapping functions. 4. Data element quality management, including data quality analysis, basic details analysis, professional algorithm analysis, etc., support data quality rule management. Step 3: Data Application 1. Data search across the entire platform, supporting collision retrieval of structured and unstructured data, and realizing structured data retrieval of the full-text library and unstructured data retrieval of the file library. 2. Data element sharing, supporting data sharing permissions to be controlled at the field level, providing data query and sharing, data download interface, and Kafka message push functions. Step 4: System Management 1. Provide unified operation and maintenance management capabilities, including data resource management, permission security management, and log management. 2. Support monitoring and early warning of its own environment status to ensure the healthy operation and security management of the system. 3. Support a unified management portal for open source big data clusters, covering components such as HDFS, YARN, MapReduce, and Hive. Step 5: Expansion and Customized Services 1. Possess good scalability and customization capabilities to meet different business needs and development scenarios. 2. Support user-customized data element collection strategies, as well as customized data element analysis and processing logic and processes Step 6: Open Services 1. Follow the industry's popular open standards, have cross-platform capabilities, and support mainstream operating systems and hardware platforms. 2. Adopt technical frameworks that are easy to expand horizontally in the future, such as in-memory computing / offline computing, columnar database hybrid mode, streaming computing, and distributed storage technology.
[0005] The innovation and core of this patent: The innovation and core of this patent lies in providing a data element collection and governance platform, which integrates multiple functional modules such as data collection, data governance, data application and system management, forming a complete big data processing and analysis solution. 1. Diversification of data collection: It supports the collection of structured and unstructured data from different types of data sources, as well as the ability of real-time collection and file entry, meeting the needs of data diversity and real-time in the big data era. 2. Comprehensiveness of data governance: It covers multiple aspects such as data cleaning, standardization, control and quality management, and ensures the quality, consistency and availability of data through professional data governance tools and methods. 3. Depth of data application: It not only provides the ability to search and share data, but also realizes cross-platform data query and sharing through technologies such as data bus, improving the application efficiency of data. 4. Unification of system management: It provides a unified operation and maintenance management platform, supports the unified management of big data clusters, and monitors and warns of its own environmental status, ensuring the stability and security of the system. 5. Expansion and customization capabilities: The platform has strong scalability and customization capabilities, and can develop customized data element collection strategies and analysis and processing logic according to the specific needs of users. 6. Open services: Following the popular open standards in the industry and supporting cross-platform operations, this makes the platform highly portable and integrable, and can be easily connected to other systems or platforms.
[0006] The technical effects brought by adopting this patented technical solution: 1. Significantly improved processing speed: Previously, the data processing speed was usually in seconds, but after adopting this patented technical solution, the processing speed is increased to milliseconds, greatly improving the efficiency of data processing.
[0007] 2. Enhanced data processing capability: This patented technical solution can process larger amounts of data, thereby improving data processing capabilities.
[0008] 3. Improved data processing accuracy: By optimizing algorithms and processing procedures, this patented technical solution can process data more accurately, thereby improving the accuracy of data processing.
[0009] 4. Reduced system resource consumption: Compared with previous technologies, this patented technical solution can utilize system resources more efficiently, thereby reducing system resource consumption.
[0010] 5. Simplified data processing flow: By adopting modular design, this patented technical solution simplifies the data processing flow and reduces the difficulty of data processing.
[0011] 6. Improved data processing quality: Through the quality control mechanism, this patented technical solution can ensure the quality of data processing, thereby improving the quality of data processing.
[0012] 7. Enhanced scalability: With modular design, this patented technical solution has good scalability and can be easily expanded and upgraded.
[0013] 8. Wider scope of application: This patented technical solution can be applied to more types of data processing scenarios, thereby broadening its scope of application. In summary, this patented technical solution has significant advantages over previous technologies in terms of data processing speed, capacity, accuracy, resource consumption, process simplification, quality, scalability and scope of application. BRIEF DESCRIPTION OF THE DRAWINGS In order to more clearly illustrate the technical solutions in the embodiments of the present invention, the drawings required for use in the embodiments will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without creative work. Figure 1 A schematic diagram of the big data platform system architecture provided by an embodiment of the present invention; Figure 2 A schematic diagram of four steps for improving the data governance process provided by an embodiment of the present invention; Figure 3 A schematic diagram of a technical architecture diagram provided for an embodiment of the present invention; DETAILED DESCRIPTION
[0014] Key technologies: 1. Multi-source heterogeneous data acquisition technology: An important feature of big data is diversity, which means that the data sources are extremely wide and the data types are extremely complex. This complex data environment brings great challenges to the processing of multi-source data. In order to process big data, it is first necessary to extract and integrate the data from the required data sources, extract relationships and entities from them, and use a uniformly defined structure to store these data after association and aggregation. During data integration and extraction, the data needs to be cleaned to ensure data quality and credibility. The key technologies of data extraction and integration are as follows: engine technology based on materialization or ETL method (Materialization or ETL engine); engine technology based on federation database or middleware method (Federation engine or Mediator); engine technology based on data stream method (Streamengine); search engine method (Search engine) technology.
[0015] 2. Distributed File System: Hadoop Distributed File System can provide high-throughput data access, suitable for large-scale data set applications, and provides storage for massive data. This project uses a distributed file system to store large unstructured files.
[0016] 3. Distributed computing: Distributed computing is a computing technology that provides very large computing power to solve data processing or mining analysis for big data. It divides a large task into multiple parts, and then assigns these parts to multiple computers for processing, and finally combines these calculation results to get the final result. Use Hadoop map / reduce and other distributed computing technologies for log analysis and statistics, intelligent analysis, large-scale indexing, massive data sorting, word frequency statistics and other businesses that do not require high real-time performance. Use memcached, redis and other caching technologies to add high-speed cache between the database and the application for data resources that change less but need to be read frequently, which can effectively reduce database pressure and greatly improve system performance.
[0017] 2 Big data platform technology solutions: 1. Data collection: Data collection refers to the collection of data from different data sources through various technical means and tools, and the integration and processing of data to meet the demand for data and subsequent analysis. Data collection includes the collection of structured and unstructured data, and the management and control of data sources.
[0018] 2. Diversified data collection capabilities: To access data from a relational database to the platform, you need to provide basic information about the database, as well as data tables and table structure information. Details are as follows: (1) Database support type: Supports Oracle, MySQL, SQLServer, and domestic DAMO, GBase8 and other mainstream databases; (2) Basic database information: Chinese name, database name, database type, SID, server IP port number, user name, password; (3) Basic data table information: Chinese name (generated by Chinese comments or English name by default, can be modified, required), English name (cannot be modified, required), Chinese comments (automatically obtained by the system, can be modified, not required), data label; (4) Table structure information: Field name (automatically obtained by the system, cannot be modified, required), field type (automatically obtained by the system, cannot be modified, required), field Chinese (generated by Chinese annotation or English name by default, can be modified, required), field annotation (automatically obtained by the system, cannot be modified, required), data element code (optional, if it is a data element, the data element name should be provided), code item code (optional, if it is a code item, the code value should be provided); (5) Support single and batch entry of data tables, and set full or incremental update of data; (6) Support partitioning of entered data according to time or category. The data acquisition module supports the entry of a variety of data, including structured relational databases, text, etc., as well as unstructured files, pictures, etc., and stores the data.
[0019] 3. High-performance file input capability: The system supports file access. If file access is required, the file needs to be placed in the FTP server and the basic information of the FTP service and the file parsing format are provided. The file formats currently supported by the system are TEXT and CSV formats. The specific information is as follows: (1) Basic information of the FTP server: Chinese name of the FTP server (required), IP address, port number, user name, password, and path; (2) File parsing format: Supported file types include TXT, WORD, CSV, and TXT and CSV files compressed with gz, tar.gz, zip, gzip, and bzip2, and support exporting specified tables and fields in the table from the database to files; (3) When entering files, support specifying the FTP file directory to be entered and support setting the data entry synchronization cycle. Field name (required), field type (required), field Chinese (required), field comment (required), field separator (required), data element code (optional, if it is a data element, the data element name should be provided), code item code (optional, if it is a code item, the code value should be provided). For the aggregation of text data, it supports TXT, CSV and other formats, and supports TXT and CSV files compressed with gz, tar, zip, gzip, bzip2, etc.
[0020] 4. High-performance data retrieval capability: The data collection module supports direct access to data with large volumes and frequent updates, and provides real-time data query and retrieval capabilities. It can complete full and incremental updates of database data, with the finest granularity of the update method supporting an update cycle of 15 minutes. It supports multi-table batch collection and multi-threaded concurrent collection to improve collection efficiency.
[0021] 5. Data governance: The data governance module performs functions such as resource planning, collection, quality improvement, storage, management, statistics, and query for various types of data generated by Mobile Zijin Research Institute in the process of business handling and service provision. The platform performs catalog management, asset registration in accordance with metadata standards, data cleaning, data reconstruction, data fusion, storage of heterogeneous data and massive data, and data cloud management for the collected data. It supports self-service rectification of data quality and tracking of data quality issues. Data governance refers to the process of cleaning, correcting, and converting data. By using data governance capabilities, it is possible to identify and process missing, duplicate, incorrect, and inconsistent problems in data, standardize and unify data to meet the specified standards and requirements. The included functions are data control, data standards, data cleaning, and data quality functions.
[0022] 6. Data element cleaning: Data element cleaning is to "clean" the data by filling in missing values, smoothing noisy data, identifying or removing outliers, and resolving inconsistencies. The main goals are to achieve the following: format standardization, abnormal data removal, error correction, and duplicate data removal.
[0023] 7. Data element standardization: Data element standardization function description: Data element standardization refers to the standardization of raw data of varying quality based on the self-built industry data standard system, combined with data quality and data development module functions, including the establishment of a unified data standard system, data dictionary system, providing standardized table building tools, and supporting batch import and synchronization of data dictionaries, so as to achieve unified standard and standardized data storage and management. Data element standardization implementation steps: (1) View the data dictionary list: pass the query conditions to the background, query the data in the library according to the conditions, and obtain the data return list (2) Modify the data dictionary: pass the code item id to the background, query the code item and code value data in the library according to the id, bind to the form, modify the form data, pass to the background, first delete the original code value associated with the code item, save the code item and the code value passed in by the front end. (3) Create data dictionary (file import): fill in the form information, select the database source, select the imported data dictionary file, pass the file and form information to the background, process the file, and save it to the database (4) Create data dictionary (database import): fill in the form, select the database source, user, table, and code configuration, pass it to the background, and save it to the library (5) Create data dictionary (manual creation): fill in the form, select the data source, pass it to the background, and save it to the library (6) Delete data dictionary: pass the code item id to be deleted to the background, query the data source table field information based on the id, and determine whether the data dictionary is used. If it is not used, delete it from the library, otherwise do not delete it (7) View data dictionary details: pass the id to the background, query the code item and code value data based on the id, and pass it to the front page 8. Data element control: Data element control refers to the use of multi-type hybrid storage modes and flexible data storage management systems to support data extraction, conversion, cleaning, conversion and mapping, verification and validation, real-time data warehouse and offline data warehouse functions, metadata management and data permission control, task tracking and traceability, data distribution, ETL construction and script management, data lineage analysis and data labeling management and other functions to achieve reasonable classification, scenario-based application and comprehensive control of data.
[0024] 9. Data management analysis: (1) Select the database to be analyzed. The backend will analyze the database and enter the table into our platform. (2) After the analysis is completed, an analysis report will be generated for viewing. 10. Data monitoring: Implementation steps: Select the table to be imported. After listing the table fields, you can choose to automatically match according to the field annotations, or you can choose your own optional rules to detect. After the detection, a detection report will be generated. 11. Data element sharing: The data element sharing function realizes the control of data sharing permissions to the field level, and provides cross-platform data query, data download interface and Kafka message push. (1) Data sharing permission control: The data sharing permission control function supports the precise control of data sharing permissions to the field level of the table, ensuring the security and compliance of data access. Users can set access rights for different roles as needed and perform fine-grained control on fields to ensure that only authorized users can access and share data. (2) Data query and sharing: The data query and sharing function achieves the goal of cross-platform and seamless connection with various production system libraries. Users can use the universal REST API to easily access and share data in different data sources. In addition, it supports real-time SQL queries, fully utilizes memory and network resources through resident processes, greatly improves query efficiency, and makes query efficiency exceed that of traditional frameworks.
[0025] Each embodiment in this specification is described in a progressive manner, and the same or similar parts between the embodiments can be referred to each other, and each embodiment focuses on the differences from other embodiments. In particular, for the device embodiment, since it is basically similar to the method embodiment, the description is relatively simple, and the relevant parts can be referred to the partial description of the method embodiment. The above is only a specific implementation method of the present invention, but the protection scope of the present invention is not limited thereto. Any changes or replacements that can be easily thought of by any technician familiar with the technical field within the technical scope disclosed by the present invention should be covered within the protection scope of the present invention. Therefore, the protection scope of the present invention should be based on the protection scope of the claims.
Claims
1. A big data platform system, characterized in that: include: The data collection module is used to collect structured and unstructured data and supports access to multiple data sources, including: Database access submodule, used for connecting with databases and third-party systems and collecting data; File access submodule, used for connecting to the file server and collecting data; A real-time data stream access submodule is used to connect to the real-time data stream and collect data; The data governance module includes functions such as data element cleaning, data element standardization, data element control, and data element quality management, which is used to clean, standardize, control, and manage the quality of the collected data; Data application module, including functions such as full-platform data search and data element sharing, used to provide data analysis and application services; The system management module provides unified operation and maintenance management capabilities for monitoring platform operation status, managing user permissions, configuring system parameters, etc.
2. The big data platform according to claim 1, characterized in that: The data acquisition module also includes: The data conversion submodule is used to convert the collected data, including data format conversion, data type conversion, data mapping, etc. The data cleaning submodule is used to clean the collected data, including removing noise, filling missing values, eliminating outliers, etc. The data standardization submodule is used to standardize the collected data, including data normalization, data discretization, etc. The data integration submodule is used to integrate data from different data sources, including data merging, data association, etc.
3. The big data platform according to claim 1, characterized in that: The data governance module also includes: Data quality management submodule is used to evaluate and manage data quality, including data quality inspection, data quality analysis, data quality report, etc.; The data security management submodule is used to manage data security, including data encryption, access control, security audit, etc. The data standard management submodule is used to manage data standards, including data standard formulation, data standard release, data standard monitoring, etc.
4. The big data platform according to claim 1, characterized in that: The data application module also includes: The data search submodule is used to provide data search functions across the entire platform, supporting users to search for data by keywords, field values, and other methods; The data analysis submodule is used to provide data analysis functions and support users to conduct statistics, analysis, and mining of data; The data visualization submodule is used to provide data visualization functions, supporting users to display data in the form of charts, graphs, etc.; The data sharing submodule is used to provide data sharing functions and support users to share data with other users or systems.
5. The big data platform according to claim 1, characterized in that: The system management module also includes: User management submodule, used to manage user information, user permissions, etc.; The permission management submodule is used to manage system permissions, resource permissions, etc. The log management submodule is used to manage system logs, user operation logs, etc. The monitoring and management submodule is used to monitor the system operation status, performance indicators, etc.
6. The big data platform according to claim 1, characterized in that: The platform adopts a microservice architecture, with each module loosely coupled and able to be independently deployed and expanded.
7. The big data platform according to claim 1, characterized in that: The platform adopts a distributed architecture, can process massive data, and ensure high availability and reliability of the system.