Multi-source heterogeneous big data aggregation analysis method based on master data modeling

Through master data modeling and graph structure storage mode, combined with machine learning and artificial intelligence technologies, the problem of integrating and analyzing multi-source heterogeneous big data is solved, efficient data management and analysis is achieved, the data processing speed and accuracy are improved, multimodal analysis is supported, and hidden information in the data is revealed.

CN120687639APending Publication Date: 2025-09-23JIANGSU JINLING TECH GRP CORP
View PDF 0 Cites 2 Cited by

Patent Information

Application Number
CN202510505016.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-22
Publication Date
2025-09-23

AI Technical Summary

Technical Problem

Existing technologies make it difficult to effectively integrate and analyze multi-source heterogeneous big data, especially in massive data scenarios. Traditional methods have bottlenecks in data compatibility, processing efficiency, and fusion accuracy, and the connection between structured and unstructured data is weak.

Method used

A master data modeling-based method is adopted to model the master data, define the relationship model between business entities and data entities, use a graph structure storage model, and combine machine learning and artificial intelligence technologies to define the data entity fusion algorithm, access and clean data, build a relational subject library, use lake-warehouse integrated technology for storage and calculation, and perform multimodal analysis.

Benefits of technology

It achieves unified management and efficient analysis of multi-source heterogeneous big data, improves data connection and analysis efficiency, supports graph theory algorithms, reveals hidden information in the data, provides support for intelligent decision-making, and data analysis results can be continuously updated and easily shared.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120687639A_ABST
    Figure CN120687639A_ABST
Patent Text Reader

Abstract

The invention discloses a multi-source heterogeneous big data aggregation analysis method based on master data modeling. The method comprises the following steps: modeling main data, defining a relation mode of a business entity and a data entity, and outputting a graph structure storage mode; classifying the service entities, and defining different data entity fusion algorithms for different data entity types by taking the main attributes of the service entities as the identification basis of the data entities; accessing data, and cleaning the accessed data; service entities and data entities are analyzed from the cleaned data based on a data entity fusion algorithm, and a relation theme library is constructed according to the incidence relation between the service entities and the data entities; storing the cleaned data and a relation theme library thereof; and carrying out modeling analysis on the cleaned data by combining the characteristics of the business entity and the relation theme library. According to the method, unified management of multi-source heterogeneous mass data is carried out from a global perspective, the blood relationship of the data is clear, and the authority is controllable.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of data aggregation and analysis, and in particular to a multi-source heterogeneous big data aggregation and analysis method based on master data modeling. Background Art

[0002] With the rapid development of information technology, especially the recent rise of big data, cloud computing, artificial intelligence, mobile internet, and the Internet of Things, data has experienced explosive growth, and its sources are becoming increasingly diverse. Data primarily encompasses structured, semi-structured, and unstructured data, with unstructured data accounting for over 80% of the total data volume. Master data, as the core business entity data within an enterprise or organization, is typically confined to information systems, failing to realize its immense value. Furthermore, business data is dispersed across various business systems, creating data silos. This multi-source, heterogeneous data environment poses significant challenges for consistent, integrated data analysis. Current applications primarily focus on the classification and processing of structured, semi-structured, and unstructured data, but lack deep understanding of the correlations within heterogeneous data. This is particularly true for applications involving massive amounts of big data, where data is typically simply imported into the system and processed using traditional stream or batch methods. This processing approach faces bottlenecks in data compatibility, processing efficiency, and integration accuracy. In particular, the connections between structured and unstructured data are often weak, making comprehensive and effective analysis difficult using traditional methods. Therefore, developing more efficient and accurate data aggregation and analysis technologies has become an urgent challenge for the industry. This application, through master data modeling technology, provides a unified, integrated view of heterogeneous data at the aggregation level, laying a high-quality data foundation for multi-source heterogeneous big data analysis and enabling collaborative data analysis. This method significantly improves data connectivity and analysis efficiency, and therefore has broad application prospects. Summary of the Invention

[0003] The purpose of the present invention is to provide a multi-source heterogeneous big data aggregation and analysis method based on master data modeling to address the deficiencies in the existing technology.

[0004] To achieve the above objectives, the present invention provides a multi-source heterogeneous big data aggregation and analysis method based on master data modeling, comprising:

[0005] Step 1: Model the master data and analyze the business entities contained in the master data. Define the raw data of each data type as a data entity, then define the relationship model between the business entity and the data entity, and output a graph structure storage model.

[0006] Step 2: Classify the business entities, and use the main attributes of the business entities as the basis for identifying data entities. Define different data entity fusion algorithms for different data entity types.

[0007] Step 3: Access data from different sources, formats, and structures, and clean the accessed data;

[0008] Step 4: Based on the data entity fusion algorithm, business entities and data entities are parsed from the cleaned data, and the association relationship between the business entities and the data entities is written into the graph database through data fusion to construct a relationship theme library;

[0009] Step 5: Store the cleaned data and its related subject database into the database;

[0010] Step 6: Model and analyze the cleaned data based on the characteristics of the business entity and the relationship subject library.

[0011] Furthermore, the business entity is obtained by analyzing in the following manner:

[0012] Based on data research, the asset metadata of master data is divided into standard metadata, connection metadata, and high-value metadata. Different metadata corresponds to different application scenarios.

[0013] Master data modeling is performed based on the type of metadata, combining top-down and bottom-up modeling methods to analyze the business entities contained in the master data in a specific field.

[0014] Furthermore, the business entities include people, events, addresses, objects, organizations, or abstract network accounts, mobile phone numbers, and assets.

[0015] Furthermore, in step 2, machine learning and artificial intelligence large model technology are combined to abstract the data entity fusion algorithm through model training.

[0016] Furthermore, the data accessed in step 3 includes structured data, semi-structured data, and unstructured data, and the data access strategies include one-time import and docking access.

[0017] Furthermore, cleaning of structured data includes correcting data type errors and inconsistencies, cleaning of semi-structured data includes unified format conversion, and cleaning of the unstructured data includes feature alignment to ensure the consistency of the unstructured data in time and space.

[0018] Furthermore, the cleaned data is stored in the warehouse based on the lake-warehouse integration technology, as follows:

[0019] Use data warehouse storage for structured data and use data lake object storage for semi-structured and unstructured data.

[0020] Furthermore, the step 6 specifically includes:

[0021] Poll modeling and analysis tasks to obtain the basic information required for task execution;

[0022] Decompose the modeling and analysis tasks, generate a directed acyclic graph, and determine the dependencies, data flows, and computational order of the subtasks;

[0023] Submit Spark tasks in sequence according to the calculation order, run the subtask modeling operators, and save the running status and results;

[0024] For model accumulation tasks, execution is triggered regularly and modeling analysis results are updated incrementally.

[0025] Furthermore, it also includes:

[0026] Step 7: Generate a visual report of the system operation status, modeling analysis results, theme library analysis results, etc. and output it to the user end.

[0027] Beneficial effects: 1. The present invention provides a global perspective for unified management of multi-source heterogeneous massive data, with clear data lineage and controllable permissions;

[0028] 2. This invention provides a unified fusion platform for heterogeneous data through master data modeling. Benefiting from the rapid development of machine learning and artificial intelligence technologies in recent years, the processing speed and accuracy of unstructured data have greatly benefited;

[0029] 3. Using lake-warehouse integration technology, heterogeneous data storage and computing frameworks are unified, reducing data redundancy, integrating distributed computing engines, and improving batch processing, stream processing, and interactive query performance;

[0030] 4. Structured data, semi-structured data, and unstructured data are integrated using a relational theme library, supporting graph theory algorithms and facilitating multimodal analysis through data modeling. This can reveal hidden information in the data and provide support for intelligent decision-making.

[0031] 5. Through the model solidification capability, data analysis results can be continuously updated and easily shared, which can further maximize business value. BRIEF DESCRIPTION OF THE DRAWINGS

[0032] Figure 1 It is a flow chart of a multi-source heterogeneous big data aggregation and analysis method based on master data modeling according to an embodiment of the present invention. DETAILED DESCRIPTION

[0033] The present invention will be further illustrated below with reference to the accompanying drawings and specific embodiments. These embodiments are implemented based on the technical solutions of the present invention. It should be understood that these embodiments are only used to illustrate the present invention and are not used to limit the scope of the present invention.

[0034] like Figure 1 As shown above, an embodiment of the present invention provides a multi-source heterogeneous big data aggregation and analysis method based on master data modeling, including:

[0035] Step 1: Model the master data and analyze the business entities contained in the master data. Define the raw data of each data type as a data entity, then define the relationship model between the business entity and the data entity, and output the graph structure storage model. In actual applications, master data modeling establishes a relationship mapping between business entities and data entities through research and analysis of data assets in specific fields. In addition to using the general entity relationship modeling design concept, it also combines data research, data standards and business analysis to provide entity classification management, graph model conversion and data traceability capabilities. It should be noted that master data modeling is iterative and should evolve gradually in practice. The main steps include:

[0036] Based on data research, the asset metadata of master data in specific fields is divided into standard metadata, connection metadata, and high-value metadata. Different metadata corresponds to different application scenarios. For example, standard metadata is formulated according to industry data standards, connection metadata refers to metadata extracted from various connection data sources, and high-value metadata is automatically generated by data analysis applications.

[0037] Master data modeling is performed based on the type of metadata. Using a combination of top-down and bottom-up modeling approaches, the modeling process is refined and analyzed to identify the business entities contained in the master data for a specific domain, such as specific people, events, addresses, items, organizations, or abstract network accounts, mobile phone numbers, and assets.

[0038] Define data entities, distinguish between structured, semi-structured, and unstructured data, and each record is an instance of the corresponding data entity;

[0039] Optimize the storage structure and simplify the computational complexity. Business entities only store the primary attribute set, where the primary attribute set is the identifier that distinguishes between entities. The primary attribute of a data entity can be the feature vector or summary of the original data.

[0040] Determine the unique identifier of the business entity as the basis for entity identification. For example, for a person's identity, the unique identifier is usually the ID number. However, if the ID number is not available in the data, the name, date of birth, and place of residence can generally be used to uniquely identify a person.

[0041] Establish associations between different data entities and business entities, and solidify the entity relationship model into the graph database storage model.

[0042] Step 2: Classify the business entities and use the main attributes of the business entities as the basis for identifying data entities. Define different data entity fusion algorithms for different data entity types, thereby building a data entity fusion algorithm library. The main steps include:

[0043] Establish associations between different metadata and entity fusion algorithms. Different feature fusion algorithms are used for structured, semi-structured, and unstructured data. For complex application scenarios, combining machine learning with AI big model technology and model training, data entity fusion algorithms can be abstracted at a higher dimension.

[0044] Map the data metadata to business entities, identify the characteristics of the original data, construct data entities, and generate entity-fused relational subject library subgraphs;

[0045] Characterize the fusion algorithm from the perspectives of performance and accuracy, and incorporate entity feature algorithms, fusion algorithms, and heuristic algorithms into core business assets for unified management.

[0046] Step 3: Access data from different sources, formats, and structures, and clean the accessed data. In practical applications, attention should be paid to the carrying capacity of the big data platform to ensure stable and efficient operation of the platform. The main steps include:

[0047] Create data access tasks to connect multi-source heterogeneous data to the big data platform. Different access strategies can be used, including one-time import or incremental docking import. One-time import means no additional data will be added after import, while docking import means data will be incrementally connected to the system in real time.

[0048] A data cleaning task is created. Once the data is accessed, the cleaning task is scheduled by a background cleaning program to clean the data. In a preferred embodiment of the present invention, different cleaning strategies are adopted for data of different structures. For example, cleaning of structured data includes correcting data type errors and inconsistencies, cleaning of semi-structured data includes unified format conversion, and cleaning of unstructured data includes feature alignment to ensure temporal and spatial consistency of the unstructured data.

[0049] In addition, specialized parsing tools can be developed to uniformly convert semi-structured data into JSON format. Since this application focuses on generating the final aggregated analysis capability, in practical applications, data access and data cleaning should be decoupled. For access scenarios involving massive amounts of data, a big data technology platform should be selected for hosting, such as accessing the Hadoop distributed file system to prevent data from being inaccessible due to disk fullness. Archiving, merging, and compressed storage of small files should also be considered.

[0050] Step 4: Based on the data entity fusion algorithm, business entities and data entities are parsed from the cleaned data, and the association relationship between business entities and data entities is written into the graph database through data fusion to build a relational subject library. In actual applications, business entities and data entities are the basis for the association between data. For massive data application scenarios, the distributed computing capabilities should be fully exploited, and technologies such as Kafka, Spark, and Flink should be used to support business logic. Entity feature fusion tasks are created through the user interface, and the entity feature fusion program in the background will perform unified scheduling and execution. The main steps include:

[0051] Poll entity feature fusion tasks to obtain basic information required for task execution, such as data metadata definition, data storage location, data type, operating status, fusion rules, etc.

[0052] Apply for system resources, submit tasks, and initialize the task execution environment;

[0053] Identify business entities in the data and generate business entity instances;

[0054] Different feature recognition algorithms are used for different data structures. For semi-structured and unstructured data, intelligent algorithms such as text recognition, image recognition, and audio and video recognition are used to extract feature vectors and generate data entity instances.

[0055] Integrate the relationship between business entity instances and data entity instances to generate a relational subject library subgraph instance;

[0056] Combined with the relational theme library, similarity comparison and multimodal data fusion are performed on text, pictures, audio and video, and the results are merged and stored in the relational theme library graph;

[0057] Release resources, update task status, and wait for the next round of task execution.

[0058] Furthermore, related information can be categorized and stored by business entity type, reducing the size of the related table and facilitating subsequent modeling and analysis. Feature vectors and information summaries can serve as a basis for similarity comparisons and multimodal data fusion across text, images, audio, and video. Deep learning and artificial intelligence technologies can also be fully utilized to refine the processing of unstructured data. Algorithms should be categorized and graded based on their strength to ensure that tasks can be completed within a reasonable timeframe.

[0059] Step 5, store the cleaned data and its related subject database into the database. The cleaned data can be stored in the database based on the lake-warehouse integrated technology. Specifically, different storage strategies should be adopted for different data structures. For example, data warehouse storage can be used for structured data, and object storage of data lake can be used for semi-structured and unstructured data. In actual applications, different data storage strategies should be adopted for different data structures, but they should be included in the lake-warehouse integrated architecture for unified storage. In a preferred embodiment of the present invention, the lake-warehouse integrated architecture should support a distributed computing framework, be able to complete high-performance computing and real-time analysis, and adapt to various types of data storage strategies. The underlying data storage of the graph database can use distributed storage engines such as HBase to maintain the unity of the technology stack. In addition, for semi-structured and unstructured object storage, the summary information of the entity feature vector can be used as the primary key ID, and a RESTful interface can be provided to obtain the semi-structured data and unstructured data corresponding to the primary key ID.

[0060] Step 6: Model and analyze the cleaned data based on the characteristics of the business entities and the relationship theme library. In practical applications, modeling and analysis is an important means of data analysis. The present invention can deeply explore the connections between heterogeneous data. System users can create modeling and analysis tasks through visual modeling and analysis, and the background modeling and analysis program will uniformly schedule and execute them. The main steps include:

[0061] Poll modeling and analysis tasks to obtain basic information required for task execution, such as data metadata definition, data storage location, modeling and analysis conditions, and operating status;

[0062] Decompose the modeling and analysis tasks to generate a DAG (Directed Acyclic Graph) to determine the dependencies, data flows, and computation order of the subtasks.

[0063] Submit Spark tasks in sequence according to the calculation order, run the subtask modeling operators, and save the running status and results;

[0064] For model accumulation tasks, execution is triggered regularly, and modeling and analysis results are incrementally updated to form a high-value database.

[0065] Additionally, modeling and analysis can be based on the Spark SQL engine, combined with Milvus vector search and the Spark GraphX ​​distributed graph computing framework to support large-scale heterogeneous data computation. It should also support a variety of operators to complete analysis in different application scenarios, such as filtering, extraction, set intersection and difference, transformation, deduplication, merging, group statistics, and graph theory algorithms. Modeling and analysis programs should be resident in memory as background programs, and analysis should be submitted as tasks. Through reasonable resource scheduling, tasks can be broken down and run, and the status and results of the operations recorded. Intermediate and final results of modeling and analysis should have a lifecycle and be included in the unified management of data assets.

[0066] This embodiment of the present invention further includes: Step 7, generating a visual report based on the system operation status, modeling analysis results, and subject library analysis results, and outputting it to the user. In actual applications, multiple views are provided for different users, such as a data asset view for displaying the retrieval and analysis results of metadata, master data models, raw data, and high-value data, and an operation and maintenance view for displaying the system operation status.

[0067] The present invention has a wide range of application scenarios, including but not limited to finance, telecommunications, e-commerce, education, healthcare, police and other industries. Taking the anti-fraud field as an example, through cross-industry cooperation, after the data in the field is processed by the aggregation method provided by the present invention, the data is effectively organized and managed, business entities are connected through data entities, and data entities are linked to the original data. Through modeling analysis, it is convenient to portray character portraits, and analyze behaviors such as frequent changes of mobile phones, high-frequency communications, and abnormal transactions, thereby locking down criminal suspects and victims, providing strong support for protecting the property safety of the people.

[0068] The above description is merely a preferred embodiment of the present invention. It should be noted that any other aspects not specifically described are considered prior art or common knowledge to those skilled in the art. Improvements and modifications may be made without departing from the principles of the present invention, and such improvements and modifications are also within the scope of protection of the present invention.

Claims

1. A multi-source heterogeneous big data aggregation and analysis method based on master data modeling, characterized by: include: Step 1: Model the master data and analyze the business entities contained in the master data. Define the raw data of each data type as a data entity, then define the relationship model between the business entity and the data entity, and output a graph structure storage model. Step 2: Classify the business entities, and use the main attributes of the business entities as the basis for identifying data entities. Define different data entity fusion algorithms for different data entity types. Step 3: Access data from different sources, formats, and structures, and clean the accessed data; Step 4: Based on the data entity fusion algorithm, business entities and data entities are parsed from the cleaned data, and the association relationship between the business entities and the data entities is written into the graph database through data fusion to construct a relationship theme library; Step 5: Store the cleaned data and its related subject database into the database; Step 6: Model and analyze the cleaned data based on the characteristics of the business entity and the relationship subject library.

2. The multi-source heterogeneous big data aggregation and analysis method based on master data modeling according to claim 1 is characterized in that: The business entity is obtained by analyzing in the following way: Based on data research, the asset metadata of master data is divided into standard metadata, connection metadata, and high-value metadata. Different metadata corresponds to different application scenarios. Master data modeling is performed based on the type of metadata, combining top-down and bottom-up modeling methods to analyze the business entities contained in the master data in a specific field.

3. The multi-source heterogeneous big data aggregation and analysis method based on master data modeling according to claim 2 is characterized in that: The business entities include people, events, addresses, objects, organizations, or abstract network accounts, mobile phone numbers, and assets.

4. The multi-source heterogeneous big data aggregation and analysis method based on master data modeling according to claim 1 is characterized in that: In step 2, machine learning and artificial intelligence large model technology are combined to abstract the data entity fusion algorithm through model training.

5. The multi-source heterogeneous big data aggregation and analysis method based on master data modeling according to claim 1 is characterized in that: The data accessed in step 3 includes structured data, semi-structured data and unstructured data, and the data access strategies include one-time import and docking access.

6. The multi-source heterogeneous big data aggregation and analysis method based on master data modeling according to claim 5 is characterized in that: Cleaning of structured data includes correcting data type errors and inconsistencies, cleaning of semi-structured data includes unified format conversion, and cleaning of unstructured data includes feature alignment to ensure the consistency of unstructured data in time and space.

7. The multi-source heterogeneous big data aggregation and analysis method based on master data modeling according to claim 5 is characterized in that: The cleaned data is stored in the warehouse based on the lake-warehouse integration technology, as follows: Use data warehouse storage for structured data and use data lake object storage for semi-structured and unstructured data.

8. The multi-source heterogeneous big data aggregation and analysis method based on master data modeling according to claim 1 is characterized in that: The step 6 specifically includes: Poll modeling and analysis tasks to obtain the basic information required for task execution; Decompose the modeling and analysis tasks, generate a directed acyclic graph, and determine the dependencies, data flows, and computational order of the subtasks; Submit Spark tasks in sequence according to the calculation order, run the subtask modeling operators, and save the running status and results; For model accumulation tasks, execution is triggered regularly and modeling analysis results are updated incrementally.

9. The multi-source heterogeneous big data aggregation and analysis method based on master data modeling according to claim 1 is characterized in that: Also includes: Step 7: Generate a visual report of the system operation status, modeling analysis results, theme library analysis results, etc. and output it to the user end.

Citation Information

Cited By

  • Digital base system

    CN121689559A

  • Structural domain interaction database construction method based on multi-source data fusion

    CN121979864A