Implementation method of multi-source heterogeneous big data acquisition and conversion service

Through multi-source heterogeneous data acquisition tools and Spark Streaming technology, the problems of big data integration and real-time processing in the existing technology are solved, the timeliness of data assetization and business needs are realized, and the data value is fully utilized.

CN120196667APending Publication Date: 2025-06-24BEIJING WUZI UNIVERSITY
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
CN202510344044.4
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-23
Publication Date
2025-06-24

AI Technical Summary

Technical Problem

It is difficult for the existing technology to effectively integrate multi-source heterogeneous big data to realize data assetization and real-time processing, and it is difficult to meet the timeliness of business needs and the realization of data value.

Method used

It adopts multi-source heterogeneous data acquisition tools to provide customized plug-in framework standards and a variety of data acquisition plug-ins to achieve data integration. Ensure data quality and security through process-based big data cleaning tools and metadata management. Use Spark Streaming technology to carry out real-time PeB-level massive data processing, and provide data sharing and query search services.

Benefits of technology

It realizes the integration and real-time processing of multi-source heterogeneous big data, helping customers to achieve data assetization, meet the timeliness of business needs, and give full play to the value of data.

✦ Generated by Eureka AI based on patent content.
Patent Text Reader

Abstract

The invention relates to a method for realizing multi-source heterogeneous big data acquisition and conversion service, in particular to a method for acquiring data from a data source which acquires multi-source heterogeneous data, provides a user-defined plug-in framework standard, provides various data acquisition plug-ins for different data structures and sources and realizes multi-source heterogeneous data integration. And then, completing millisecond-level processing of the PB-level mass data in real time.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to a method for implementing multi-source heterogeneous big data collection and conversion services, and particularly to a method for obtaining data from a data source that integrates multi-source heterogeneous data by collecting multi-source heterogeneous data, providing a standard for a custom plug-in framework, providing a variety of data collection plug-ins for different data structures and sources, and then completing the millisecond-level processing of real-time PB-level massive data. Technical Background

[0002] The era of big data has arrived. In the fields of business economy, social governance, etc., development and decision-making will increasingly be based on data and analysis, and data will become the fundamental resource in economic operation. How to integrate data from various information systems, conduct data analysis, discover the value of data, and provide services for future business development has become a new challenge and goal for governments at all levels and various enterprises.

[0003] The challenges faced by governments and enterprises in using traditional data analysis are as follows:

[0004] (1) How to achieve data assetization. Data is scattered, huge in volume, complex in structure, and diverse in types, requiring efficient processing and connection means to break information silos. How to quickly locate and extract correct and valuable data from diverse and disorderly big data, and discover the relationships between data to form data assets is one of the challenges faced by big data applications.

[0005] (2) How to use data to solve problems. The timeliness of massive data statistics, analysis, and prediction is poor, making it difficult to meet the growing needs of business departments. Achieving the deep integration of data and business, and accurately establishing an analysis model oriented to problems is the basis for realizing the monetization of data value and also one of the challenges faced by big data applications.

[0006] (3) How to understand the conclusions of big data. The connotation of data is difficult to understand. How to understand the business value contained in the conclusions of big data analysis, apply big data to business development, quickly adapt to the continuously growing and changing business needs, and actually solve business problems is one of the challenges faced by big data applications.

[0007] To solve the above problems, governments and enterprises have introduced big data technologies one after another and started to build big data platforms and service capabilities to break information silos and achieve data assetization. Summary of the Invention

[0008] The technical problem to be solved by the present invention is to realize the assetization of big data collection and management based on a multi-source heterogeneous big data collection and exchange method, establish a scientific and secure data opening and sharing mechanism with the help of a process-based big data cleaning tool, assist customers in achieving data assetization, and lay a solid foundation for giving full play to the value of data.

[0009] A method for data collection and exchange of big data based on multi-source heterogeneity, the method at least includes the following steps:

[0010] Step 1: The collection tool supports the collection of multi-source heterogeneous data, provides a standard for the custom plug-in framework, and provides a variety of data collection plug-ins for different data structures and sources to achieve the integration of multi-source heterogeneous data. The tool has a perfect mechanism for monitoring and alarming abnormal data to ensure data integrity. The HTTPS protocol and the SSL data encryption algorithm are used to ensure data security during the data transfer process. Based on the Spark Streaming technology, the real-time PB-level massive data is processed in milliseconds.

[0011] Step 2: Data processing: The data collection tool realizes the collection of different types of data, and uses technologies such as missing value processing, data transformation, data redundancy, and data reduction to perform data preprocessing and output through various methods. The most direct output method is to connect to the self-service analysis platform and the research and judgment system to complete subsequent data analysis and applications. In addition, it can also be output in various ways such as files, databases, message queues, and APIs to achieve data sharing.

[0012] Step 3: Metadata management: Metadata is the attribute description of unstructured data, including: data type, data unit, data source, data format, compression format, etc. Taking the monitored picture data as an example, the metadata includes the following content: ① Data type: monitored picture data; ② Data format: binary picture; ③ Compression format: uncompressed; ④ Data source: VIO_JDCZP; ⑤ Data fields: photo1, photo2, photo3; ⑥ Primary key source: VIO_SURVE IL_DK, etc.

[0013] Metadata management is to perform operations such as adding, deleting, and editing on metadata.

[0014] Step 4: Data asset management: Provide a function to monitor and track the data status of data resources, and visually display the monitoring information in the form of graphs and tables. Provide a statistical analysis function for the data operation status and data status according to the established conditions, and provide a running anomaly alarm prompt, statistical analysis of abnormal operation status, and processing mechanism function for abnormal operation status.

[0015] Step 5: Shared exchange service: Data sharing and exchange is based on data collection and storage. The data of different application systems is transformed into data that conforms to specific informatization specifications after standardization transformation. Different application systems extract data through this service to achieve information sharing.

[0016] Step 6: Query the search service: Adopt a fully distributed system with a multi-copy mechanism, introduce a heterogeneous distributed storage mode, adapt to the characteristics of large-scale distributed applications, support database applications with different characteristics, and enable the free expansion of database service capabilities. Adapt to the needs of theme and big data storage and analysis services, adapt to high concurrency and high reliability, and provide a basic platform for data mining and user behavior analysis. Specifically, it includes: ① Omnidirectional retrieval means. The data query engine provides a variety of retrieval operators, including various logical combination retrievals of external features and text content (AND, OR, NOT, XOR), position retrievals (in the same paragraph, in the same sentence, with a difference of several words, and related to the front-back order, etc.), secondary retrievals, progressive retrievals, etc.; ② Support various sorts of retrieval results. The retrieval system provides multiple sorting methods such as sorting by relevance, sorting by attributes, and hybrid sorting; ③ Search statistics. Provide the keywords frequently entered by users to the management personnel according to the statistical results, which is convenient for the management personnel to check and fill in the gaps in management methods and problem solutions according to the keyword statistical results.

[0017] The present invention relates to a method for implementing a multi-source heterogeneous big data collection and conversion service, and particularly relates to a method for collecting source heterogeneous data, providing a standard for a custom plug-in framework, and providing a variety of data collection plug-ins for different data structures and sources, so as to obtain data from a data source integrated with multi-source heterogeneous data, and then complete the millisecond-level processing of real-time PB-level massive data. Detailed implementation manners

[0018] 1. Support data collection tools for multi-source heterogeneous data;

[0019] 2. For the collection of different types of data, use techniques such as missing value processing, data transformation, data redundancy, and data reduction to implement data preprocessing and output through various methods.

[0020] 3. Metadata management;

[0021] 4. Provide a function for monitoring and tracking the data status of data resources, and visually display the monitoring information in the form of graphs and tables;

[0022] 5. Data sharing and exchange: Based on data collection and storage, convert the data of different application systems into data that conforms to specific informatization specifications after standardization transformation, and different application systems extract data through this service to achieve information sharing.

[0023] 6. Adopt a fully distributed system with a multi-copy mechanism, introduce a heterogeneous distributed storage mode, adapt to the characteristics of large-scale distributed applications, support database applications with different characteristics, and enable the free expansion of database service capabilities. Adapt to the needs of theme and big data storage and analysis services, adapt to high concurrency and high reliability, and provide a basic platform for data mining and user behavior analysis.

[0024] When the above technical solution is implemented, data is obtained from the data sources for multi-source heterogeneous data integration, and then real-time PB-level massive data is processed in milliseconds.

[0025] Finally, it should be noted that the above embodiments are only used to illustrate rather than limit the technical solutions described in the present invention; therefore, although this specification has described the present invention in detail with reference to the above embodiments, those of ordinary skill in the art should understand that the present invention can still be modified or equivalently replaced; and all technical solutions and their improvements that do not depart from the spirit and scope of the present invention should be covered within the scope of the claims of the present invention.

Claims

1. A method for implementing multi-source heterogeneous big data collection and conversion services, characterized by: The method comprises at least the following steps: Step 1: The collection tool supports multi-source heterogeneous data collection, provides a custom plug-in framework standard, and provides a variety of data collection plug-ins for different data structures and sources to achieve multi-source heterogeneous data integration. The tool has a complete abnormal data monitoring and alarm mechanism to ensure data integrity. The HTTPS protocol and SSL data encryption algorithm are used to ensure data security during data transfer. Based on Spark Streaming technology, real-time PB-level massive data can be processed in milliseconds. Step 2: Data processing: Data collection tools collect different types of data, use technologies such as missing value processing, data transformation, data redundancy, and data reduction to pre-process data, and output it in a variety of ways. The most direct output method is to connect to the self-service analysis platform and the research and judgment system to complete subsequent data analysis and application. In addition, it can also be output in a variety of ways such as files, DB, message queues, and APIs to achieve data sharing. Step 3: Metadata management: Metadata is a description of the attributes of unstructured data, including: data type, data unit, data source, data format, compression format, etc. Taking monitoring image data as an example, the metadata includes the following: ① Data type: monitoring image data; ② Data format: binary image; ③ Compression format: non-compressed; ④ Data source: VIO_JDCZP; ⑤ Data fields: photo1, photo2, photo3; ⑥ Primary key source: VIO_SURVEIL_DK, etc. Metadata management is to perform operations such as adding, deleting, and editing metadata. Step 4: Data asset management: Provides the function of monitoring and tracking the data status of data resources, and visualizes the monitoring information in the form of graphics and tables. Provides the function of statistical analysis of data operation status and data status according to established conditions, and provides abnormal operation alarm prompts, abnormal operation status statistical analysis, and processing mechanism functions for abnormal operation status. Step 5: Shared exchange service: Data sharing and exchange is based on data collection and storage. The data of different application systems are converted into data that conforms to specific information specifications through standardized transformation. Different application systems extract data through this service to realize information sharing. Step 6: Query and search service: adopt a fully distributed, multi-copy system, introduce a heterogeneous distributed storage mode, adapt to large-scale distributed application characteristics, support database applications with different characteristics, and the free expansion of database service capabilities, adapt to the needs of subject and big data storage and analysis services, adapt to high concurrency and high reliability, and provide a basic platform for data mining and user behavior analysis. Specifically include: ① All-round search means, the data query engine provides a variety of search operators, including various logical combination searches of external features and text content (AND, OR, NOT, XOR), position search (same paragraph, same sentence, difference in a few words, and order of front and back), secondary search, progressive search, etc.; ② Support various sorting of search results, the search system provides various sorting methods such as relevance sorting, attribute sorting, and mixed sorting; ③ Search statistics, provide the keywords frequently entered by users to the management personnel according to the statistical results, so that the management personnel can check and fill in the gaps in the management methods and problem solutions according to the keyword statistical results.

Citation Information

Patent Citations

  • Intelligent service application platform and method for multi-source heterogeneous data fusion

    CN107193858A

  • Multi-body data space interoperation method for multi-source heterogeneous data

    CN117931913A

  • Construction method of industrial park big data center and big data center system

    CN119248875A