A distributed big data fusion system and method for multi-source heterogeneous data

By using a distributed big data fusion system to collect data pattern metadata and map model alignment from multi-source heterogeneous data, multi-dimensional feature extraction and similarity matrix calculation are achieved. This generates associated data clusters and performs distributed computation, solving the computational bottleneck and consistency issues in multi-source heterogeneous data processing. It also improves the flexibility and accuracy of data fusion and ensures the precision and transparency of globally fused data.

CN122365393APending Publication Date: 2026-07-10SICHUAN VOCATIONAL COLLEGE OF CHEM TECH
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
SICHUAN VOCATIONAL COLLEGE OF CHEM TECH
Filing Date
2026-05-21
Publication Date
2026-07-10

AI Technical Summary

Technical Problem

Traditional methods struggle to handle multi-source heterogeneous data due to issues such as data mismatch, information loss, semantic inconsistency, computational bottlenecks, low processing efficiency, and poor scalability. Furthermore, they lack tracking records of data sources and processing procedures.

Method used

A distributed big data fusion system is adopted to collect data pattern metadata and align the mapping model, perform multi-dimensional feature extraction and similarity matrix calculation, generate associated data clusters, and perform distributed computing node allocation and global weighted aggregation to achieve data standardization and consistent fusion.

Benefits of technology

It improves the flexibility of data access and fusion, quantifies the similarity between data entities, enhances the accuracy and completeness of data association, ensures the accuracy and stability of globally fused data, and improves the transparency and manageability of the data fusion process.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122365393A_ABST
    Figure CN122365393A_ABST
Patent Text Reader

Abstract

This invention relates to the field of distributed big data fusion technology, specifically a distributed big data fusion system and method for multi-source heterogeneous data. The system includes: performing data association matching on a standardized dataset based on a similarity matrix to obtain associated data clusters; generating a distributed computing node allocation scheme based on the associated data clusters; scheduling the associated data clusters to corresponding distributed computing nodes for local fusion based on the allocation scheme to obtain local fusion results; calculating the data consistency coefficient between each distributed computing node based on the local fusion results; and performing global weighted aggregation on the local fusion results based on the data consistency coefficients to obtain globally fused data. This invention effectively improves the transparency and manageability of the data fusion process by recording operation logs and data change information at each stage of the entire process, including data acquisition, feature extraction, similarity calculation, distributed scheduling, and global weighted aggregation.
Need to check novelty before this filing date? Find Prior Art