A heterogeneous data fusion processing method and system based on big data collection

CN122777601APending Publication Date: 2026-09-18HANGZHOU FULL SYNESTHESIA TECHNOLOGY CO LTD +1
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202610789659.2
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2026-06-03
Publication Date
2026-09-18

AI Technical Summary

Technical Problem

[0002]现有针对异构数据源的数据采集环节,普遍采用固定参数与静态配置的采集方式,未针对不同数据源的配置信息开展全面的源特征拆解与深度分析,无法结合数据源的接口类型、数据更新频率、字段结构参数以及历史数据样本制定适配性采集规则,既不能依据接口传输方式合理设定采集时间窗口、根据数据更新频率确定窗口持续时长,也无法通过字段属性与数据分布精准识别关键标识字段与核心数据负载字段,导致采集策略缺乏动态适配与自主调整能力,数据提取过程存在冗余采集与有效信息缺失并存的问题,无法稳定生成标准化、结构化且可直接用于融合处理的待融合数据对象,大幅降低数据采集的精准度与整体效率

Benefits of technology

1.本发明基于异构数据源的配置信息开展全面源特征分析,生成适配数据源特性的动态采集策略,精准识别关键标识字段与核心数据负载字段,稳定获取标准化、结构化的待融合数据对象,同时为待融合数据对象生成专属数据特征标识,以此搭建初始数据血缘图谱,完整建立数据对象的来源关联与结构依赖关系,保障数据采集的精准性与数据血缘的可追溯性。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122777601A_ABST
    Figure CN122777601A_ABST
Patent Text Reader

Abstract

This invention relates to the field of data fusion technology and proposes a method and system for heterogeneous data fusion processing based on big data acquisition. The method includes: performing source feature analysis on the configuration information of heterogeneous data sources to generate dynamic acquisition strategies for heterogeneous data sources, thereby obtaining data objects to be fused from the heterogeneous data sources; generating a data feature identifier for the data objects to be fused, and constructing an initial data lineage graph using the data feature identifier as nodes; performing content semantic parsing on the data objects to be fused to generate enhanced data objects carrying semantic tags, and updating the initial data lineage graph using the semantic tags to obtain a comprehensive data graph; responding to an external data fusion request, locating target nodes based on the comprehensive data graph and a preset fusion topology, and performing association aggregation to obtain fused data objects containing association relationships. This invention can improve the efficiency of heterogeneous data fusion processing.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of data fusion technology, and in particular to a method and system for heterogeneous data fusion processing based on big data collection. Background Technology

[0002] Current data acquisition processes for heterogeneous data sources generally employ fixed parameters and static configurations. They fail to conduct comprehensive source feature decomposition and in-depth analysis based on the configuration information of different data sources. They cannot formulate adaptive acquisition rules by combining the interface type, data update frequency, field structure parameters, and historical data samples of the data source. They cannot reasonably set the acquisition time window based on the interface transmission method, determine the window duration based on the data update frequency, or accurately identify key identifier fields and core data load fields through field attributes and data distribution. As a result, the acquisition strategy lacks dynamic adaptation and self-adjustment capabilities. The data extraction process suffers from both redundant acquisition and loss of effective information. It is impossible to stably generate standardized, structured data objects that can be directly used for fusion processing, which significantly reduces the accuracy and overall efficiency of data acquisition.

[0003] The existing data fusion processing flow lacks a systematic data lineage construction mechanism based on data feature identification. It cannot generate unique identifiers and build an initial lineage graph based on data sources, collected information, and structural features. At the same time, it lacks in-depth semantic analysis of the data objects to be fused. It cannot generate enhanced data objects with semantic tags through lexical segmentation, entity mapping, and feature extraction. It also cannot rely on semantic association to update and upgrade the lineage graph to form a comprehensive data graph. The source dependencies, business associations, and semantic connections between data objects cannot be fully sorted out and visualized. When responding to external data fusion requests, it cannot rely on the preset fusion topology to quickly and accurately locate target nodes. Data association and aggregation lack clear graph guidance and standardized topology support. Problems such as association confusion and disordered aggregation are prone to occur during the fusion process. The overall automation level, accuracy, and response speed of data fusion processing cannot meet the actual application needs in big data scenarios. Summary of the Invention

[0004] This invention provides a heterogeneous data fusion processing method and system based on big data acquisition to solve the problems mentioned in the background art.

[0005] To achieve the above objectives, the present invention provides a heterogeneous data fusion processing method based on big data acquisition, comprising: Pt.1. Perform source feature analysis on the configuration information of heterogeneous data sources to generate dynamic acquisition strategies for the heterogeneous data sources, so as to obtain the data objects to be fused from the heterogeneous data sources. Pt.2. Generate a data feature identifier for the data object to be fused, and construct an initial data lineage map using the data feature identifier as a node; Pt.3. Perform content semantic parsing on the data object to be fused to generate an enhanced data object carrying semantic tags, and use the semantic tags to update the initial data lineage map to obtain a comprehensive data map; Pt.4 In response to an external data fusion request, the target node is located based on the comprehensive data map and the preset fusion topology, and the association aggregation is performed to obtain a fusion data object containing the association relationship.

[0006] In a preferred embodiment, the step of performing source feature analysis on the configuration information of heterogeneous data sources to generate a dynamic acquisition strategy for the heterogeneous data sources includes: Extract the interface type, data update frequency, field structure parameters, and historical data samples of the heterogeneous data source from its configuration information; Based on the interface type and the data update frequency, set the collection time window and collection priority for the heterogeneous data source; The data transmission method indicated by the interface type is used to determine the start time of the acquisition time window, and the data update frequency is used to determine the duration of the acquisition time window. Based on the field structure parameters and the historical data samples, the key identification fields and core data load fields in the heterogeneous data source are identified, and field location identifiers are generated. The acquisition time window, acquisition priority, key identifier field, and core data load field are encapsulated to generate a dynamic acquisition strategy for the heterogeneous data source.

[0007] In a preferred embodiment, the step of identifying key identifier fields and core data load fields in the heterogeneous data source based on the field structure parameters and the historical data samples, and generating field location identifiers, includes: Extract the name, data type, and primary / foreign key relationships between each field from the field structure parameters to generate a list of field attributes; The frequency of occurrence of non-repeating values ​​corresponding to each field is counted from the historical data sample to generate a field value distribution record; Based on the list of field attributes and the field value distribution record, the discrimination weight value of the field is determined according to the following formula: ; In the formula, The discrimination weight value is... The number of unique values ​​for the field in the historical data sample. The total number of values ​​for the field in the historical data sample. The number of times the field appears in the field attribute list as a primary key or foreign key. This represents the maximum number of primary and foreign key references for all fields in the aforementioned list of field attributes. Fields whose discrimination weight values ​​exceed a preset threshold are identified as key identifier fields, and fields with the longest data length among those whose discrimination weight values ​​are lower than the preset threshold are identified as core data load fields, thus obtaining field location identifiers for the key identifier fields and the core data load fields.

[0008] In a preferred embodiment, the dynamic acquisition strategy for generating the heterogeneous data source to obtain the data object to be fused from the heterogeneous data source includes: The collection time window, the collection priority, the list of field names of the key identifier field, and the list of field names of the core data load field are written into the corresponding field positions of the strategy template in a preset encapsulation order. In the strategy template, a timer start flag is associated with the collection time window, and an execution queue sorting flag is associated with the collection priority; Write the list of field names of the key identifier fields into the index key area of ​​the strategy template, as the basis for record matching during subsequent data crawling; Write the list of field names of the core data load fields into the data extraction area of ​​the strategy template, as an identifier for reading target data content from heterogeneous data sources during subsequent data crawling; The completed strategy template is instantiated into a dynamic collection strategy object that can be parsed and executed by the data crawler, thus obtaining the data object to be merged from the heterogeneous data source.

[0009] In a preferred embodiment, generating a data feature identifier for the data object to be fused, and constructing an initial data lineage map using the data feature identifier as nodes, includes: The data source address of the data object to be fused is hashed to derive an address digest, and a collection timestamp is generated anchored to the collection time. Traverse the field structure of the data object to be merged, extract the data type and field length of the fields, and serialize and concatenate them according to the order in which the fields appear to obtain the data structure outline signature of the heterogeneous data source. The address digest, the collection timestamp, and the data structure outline signature are encapsulated and combined to obtain the data feature identifier of the heterogeneous data source. The data feature identifier is registered as a node in the graph storage area, and dependency entries are retrieved from the preset data source dependency rule base. The dependency entries include upstream data source identifiers and downstream data source identifiers. The address digest of the node is compared with the downstream data source identifier in the dependency entry. When a match is found, a directed association is established between the node and the existing node of the upstream data source identifier. The direction of the directed association is from the existing node to the node. After traversing all dependency entries, the initial data lineage map of the heterogeneous data source is obtained.

[0010] In a preferred embodiment, the step of performing content semantic parsing on the data object to be fused to generate an enhanced data object carrying semantic tags, and using the semantic tags to update the initial data lineage map to obtain a comprehensive data map, includes: Lexical segmentation is performed on the core data load field of the data object to be fused, and the segmented words are semantically mapped to the entity lexicon. Hit words are marked as key entities. By extracting the co-occurrence relationships and order positions among the missed words, the attribute feature pairs and operation association pairs of the heterogeneous data source are obtained; The key entities, attribute feature pairs, and operation association pairs are arranged into semantic tags, and the semantic tags are attached to the data objects to be fused to obtain the enhanced data objects of the heterogeneous data sources; Using the node of the enhanced data object as the current node, iteratively scan other nodes in the initial data lineage graph. When the key entity verification of the current node is consistent with that of other nodes, establish an undirected semantic association edge between the current node and the other nodes. Next, the attribute feature pairs and operation association pairs of the current node are similar to the corresponding features of other nodes. When the verification passes, another undirected semantic association edge is established between the current node and the other nodes to obtain the comprehensive data graph of the heterogeneous data source.

[0011] In a preferred embodiment, locating the target node based on the integrated data map and the preset fused topology includes: The metadata tags of the enhanced data objects in the comprehensive data graph are parsed out, and the key identifier field values ​​are extracted from the metadata tags; Configure the key identifier field value as the query key, and perform a deep traversal scan within the node attribute area of ​​the preset fusion topology structure; The node attribute values ​​encountered during the depth traversal scan are compared with the query key for equivalence verification. When the verification result is equivalent, the node currently being verified is marked as the target node of the heterogeneous data source.

[0012] In a preferred embodiment, locating the target node further includes: When no node attribute value matching the query key is found within the node attribute area of ​​the fused topology, a new node is dynamically derived in the fused topology based on the overall metadata tag of the enhanced data object. Write the core data payload field of the enhanced data object into the corresponding attribute slot of the new node; The temporal marker and the reference relationship marker are parsed from the metadata tags of the enhanced data object, and the temporal marker and the reference relationship marker are recorded as the edge attributes between the new node and other existing nodes. The new node is identified as the target node of the heterogeneous data source.

[0013] In a preferred embodiment, the step of performing association aggregation to obtain a fused data object containing association relationships includes: Using the target node as the aggregation center, adjacent nodes that have direct directed edges or undirected semantic edges with the target node are collected in the comprehensive data graph to form an aggregated node set; Extract the core data payload field from the enhanced data objects corresponding to the nodes in the aggregated node set; Based on the type and direction of the edges connecting the target node and its adjacent nodes in the comprehensive data graph, determine the splicing order and nesting level of the core data load fields; The core data payload fields are merged and assembled according to the splicing order and nesting level to obtain the fused data object of the heterogeneous data source.

[0014] To address the aforementioned problems, the present invention also provides a heterogeneous data fusion processing system based on big data acquisition, the system comprising: The source domain acquisition strategy module is used to perform source feature analysis on the configuration information of heterogeneous data sources and generate dynamic acquisition strategies for the heterogeneous data sources to obtain the data objects to be fused from the heterogeneous data sources. The bloodline foundation module is used to generate a data feature identifier for the data object to be fused, and to construct an initial data bloodline map using the data feature identifier as a node. The semantic spectrum enhancement module is used to perform content semantic parsing on the data object to be fused, generate an enhanced data object carrying semantic tags, and use the semantic tags to update the initial data lineage map to obtain a comprehensive data map. The graph fusion module is used to respond to external data fusion requests, locate target nodes based on the comprehensive data graph and the preset fusion topology, and perform association aggregation to obtain fusion data objects containing association relationships.

[0015] Compared with the prior art, the present invention has the following beneficial effects: 1. This invention performs comprehensive source feature analysis based on the configuration information of heterogeneous data sources, generates dynamic acquisition strategies adapted to the characteristics of data sources, accurately identifies key identifier fields and core data load fields, stably acquires standardized and structured data objects to be merged, and generates exclusive data feature identifiers for the data objects to be merged, thereby building an initial data lineage map, fully establishing the source association and structural dependency relationship of data objects, and ensuring the accuracy of data acquisition and the traceability of data lineage.

[0016] 2. This invention performs deep semantic parsing on the data objects to be fused, generates enhanced data objects with semantic tags, and updates and upgrades the initial data lineage graph based on the semantic tags to form a comprehensive data graph. It comprehensively sorts out the semantic relationships and business logic between data. When responding to external data fusion requests, it can quickly locate target nodes based on the comprehensive data graph and the preset fusion topology, and complete data association and aggregation in an orderly manner, significantly improving the automation level, accurate matching ability and overall execution efficiency of data fusion processing. Attached Figure Description

[0017] Figure 1 This is a flowchart illustrating a heterogeneous data fusion processing method based on big data acquisition, provided in an embodiment of the present invention. Figure 2 A functional block diagram of a heterogeneous data fusion processing system based on big data acquisition, provided in an embodiment of the present invention; The realization of the objective, functional features and advantages of the present invention will be further explained in conjunction with the embodiments and with reference to the accompanying drawings. Detailed Implementation

[0018] It should be understood that the specific embodiments described herein are merely illustrative of the invention and are not intended to limit the invention.

[0019] This application provides a heterogeneous data fusion processing method based on big data acquisition. The execution entity of this method includes, but is not limited to, at least one of the following electronic devices that can be configured to execute the method provided in this application: a server, a terminal, etc. In other words, the heterogeneous data fusion processing method based on big data acquisition can be executed by software or hardware installed on a terminal device or a server device. The server includes, but is not limited to, a single server, a server cluster, a cloud server, or a cloud server cluster. The server can be an independent server or a cloud server that provides basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communication, middleware services, domain name services, security services, content delivery networks (CDN), and big data and artificial intelligence platforms.

[0020] Reference Figure 1 The diagram shown is a flowchart illustrating a heterogeneous data fusion processing method based on big data acquisition, according to an embodiment of the present invention. In this embodiment, the heterogeneous data fusion processing method based on big data acquisition includes: Pt.1. Perform source feature analysis on the configuration information of heterogeneous data sources to generate dynamic acquisition strategies for the heterogeneous data sources, so as to obtain the data objects to be fused from the heterogeneous data sources. In this embodiment of the invention, the step of performing source feature analysis on the configuration information of heterogeneous data sources to generate a dynamic acquisition strategy for the heterogeneous data sources includes: Extract the interface type, data update frequency, field structure parameters, and historical data samples of the heterogeneous data source from its configuration information; Based on the interface type and the data update frequency, set the collection time window and collection priority for the heterogeneous data source; The data transmission method indicated by the interface type is used to determine the start time of the acquisition time window, and the data update frequency is used to determine the duration of the acquisition time window. Based on the field structure parameters and the historical data samples, the key identification fields and core data load fields in the heterogeneous data source are identified, and field location identifiers are generated. The acquisition time window, acquisition priority, key identifier field, and core data load field are encapsulated to generate a dynamic acquisition strategy for the heterogeneous data source.

[0021] The step of identifying key identifier fields and core data load fields in the heterogeneous data source based on the field structure parameters and the historical data samples, and generating field location identifiers, includes: Extract the name, data type, and primary / foreign key relationships between each field from the field structure parameters to generate a list of field attributes; The frequency of occurrence of non-repeating values ​​corresponding to each field is counted from the historical data sample to generate a field value distribution record; Based on the list of field attributes and the field value distribution record, the discrimination weight value of the field is determined according to the following formula: ; In the formula, The discrimination weight value is... The number of unique values ​​for the field in the historical data sample. The total number of values ​​for the field in the historical data sample. The number of times the field appears in the field attribute list as a primary key or foreign key. This represents the maximum number of primary and foreign key references for all fields in the aforementioned list of field attributes. Fields whose discrimination weight values ​​exceed a preset threshold are identified as key identifier fields, and fields with the longest data length among those whose discrimination weight values ​​are lower than the preset threshold are identified as core data load fields, thus obtaining field location identifiers for the key identifier fields and the core data load fields.

[0022] The dynamic acquisition strategy for generating the heterogeneous data source, to obtain the data object to be merged from the heterogeneous data source, includes: The collection time window, the collection priority, the list of field names of the key identifier field, and the list of field names of the core data load field are written into the corresponding field positions of the strategy template in a preset encapsulation order. In the strategy template, a timer start flag is associated with the collection time window, and an execution queue sorting flag is associated with the collection priority; Write the list of field names of the key identifier fields into the index key area of ​​the strategy template, as the basis for record matching during subsequent data crawling; Write the list of field names of the core data load fields into the data extraction area of ​​the strategy template, as an identifier for reading target data content from heterogeneous data sources during subsequent data crawling; The completed strategy template is instantiated into a dynamic collection strategy object that can be parsed and executed by the data crawler, thus obtaining the data object to be merged from the heterogeneous data source.

[0023] The system iterates through the complete storage structure and all configuration items of the heterogeneous data source configuration information, disassembles and reads the configuration information one by one, accurately extracts the specific transmission protocol type of the interface type, the periodic value of the statistical data update according to the fixed time dimension, the complete extraction of the text name, data type, primary and foreign key relationship and field length of the fields, and retrieves the set of historical data records that have been persistently stored within the specified time range. After completing the full information extraction operation, the system obtains the interface type, data update frequency, field structure parameters and historical data samples of the heterogeneous data source.

[0024] The start time of the acquisition time window is accurately determined based on the synchronous or asynchronous transmission method corresponding to the interface type. The duration of the acquisition time window is determined based on the specific period value of the data update frequency. The acquisition priority levels are divided into 1 to 5 levels according to the actual value of the interface transmission efficiency and the length gradient of the data update cycle. After completing all parameter settings, the acquisition time window and acquisition priority of the heterogeneous data source are obtained.

[0025] The synchronous transmission mode corresponding to the interface type is precisely matched to the start time of the collection time window that starts immediately after the data is ready. The asynchronous transmission mode corresponding to the interface type is precisely matched to the start time of the collection time window that starts after a fixed delay after the data callback is completed. Short periods with a data update frequency of less than 1 hour are matched to a collection time window duration of 30 minutes. Long periods with a data update frequency of 1 hour or more are matched to a collection time window duration of 60 minutes. This completes the accurate determination of the start time and duration of the collection time window.

[0026] Based on the primary key and foreign key association relationships in the field structure parameters, and combined with the distribution characteristics of the unique values ​​of fields in historical data samples, the fields that can uniquely identify a single data record are determined as key identification fields, and the fields that store the core business content and have the largest data volume are determined as core data load fields. After binding globally unique numerical codes to the key identification fields and core data load fields respectively, field location identifiers are generated.

[0027] The specific time and duration of the collection time window, the collection priority level value, the complete information of the key identifier field, and the complete information of the core data load field are sequentially integrated and encapsulated according to the preset hierarchical configuration format to form a standardized strategy configuration set. The strategy configuration set is then transformed into an entity configuration object that can be directly parsed and executed by the data collection component, generating a dynamic collection strategy for heterogeneous data sources.

[0028] It iterates through all the field entries contained in the field structure parameters one by one, reads the text name of each field, stores the corresponding data type identifier, and the primary key association and foreign key association relationships between fields in turn. All the above information is organized into a structured list document according to the original arrangement order of the fields in the data source, and a list of field attributes is generated.

[0029] Read all data record entries stored in the historical data sample line by line, screen and remove duplicate values ​​for each field, count the total number of unique values ​​that remain, and organize the total number of unique values ​​for each field into standardized record entries according to the original field order to generate field value distribution records.

[0030] For each field, first count the total number of unique values ​​for that field in the historical data sample, then count the total number of all values ​​for that field in the historical data sample. Divide the total number of unique values ​​by the total number of all values ​​to obtain the first proportion value. Next, count the total number of times that field appears as a primary key or foreign key in the field attribute list. Then, count the maximum number of primary and foreign key references for all fields in the field attribute list. Divide the number of primary and foreign key occurrences of that field by the maximum number of primary and foreign key references for all fields to obtain the second proportion value. Multiply the first proportion value and the second proportion value to obtain the final calculation result, which is the field's discriminant weight value.

[0031] Fields with a discrimination weight value greater than the preset threshold of 0.7 are uniformly identified as key identification fields. Among all fields with a discrimination weight value less than the preset threshold of 0.7, the character length of the stored data in each field is compared one by one. The field with the largest character length value is selected as the core data load field. A globally unique identification code is bound to the key identification field and the core data load field respectively to obtain the field positioning identifier of the key identification field and the core data load field.

[0032] The number of unique values ​​in a field is obtained by iterating through each record of the historical data sample, comparing the field values ​​one by one, removing identical duplicate values, and then counting the total number of unique values ​​remaining.

[0033] The total number of all values ​​for a field is obtained by counting the total number of all valid data records for the corresponding field in the historical data sample, row by row, and then removing records with null values.

[0034] The number of times a field appears as a primary key or foreign key is obtained by checking the primary key and foreign key relationships of the field in the field attribute list line by line, and counting the total number of times the field is marked as a primary key or foreign key.

[0035] The maximum value of primary and foreign key reference counts for all fields is determined by comparing the occurrence counts of primary and foreign keys for each field in the field attribute list and selecting the result with the largest value.

[0036] The field uniqueness ratio, which represents the uniqueness of the field data, is calculated by dividing the number of non-repeating values ​​of the field by the total number of all values ​​of the field. The field association importance ratio, which represents the business association importance of the field, is calculated by dividing the number of occurrences of the primary and foreign keys of the field by the maximum number of references of the primary and foreign keys of all fields. The field uniqueness ratio and the field association importance ratio are multiplied to obtain a quantitative value used to measure the overall distinguishing ability of the field.

[0037] This value comprehensively reflects the field's unique data identification capability and its importance in business context, providing a unified quantitative standard for determining key identification fields and core data load fields, and ensuring the accuracy and logical rationality of the field positioning identification generation process.

[0038] Following a preset fixed encapsulation order, the specific configuration information of the collection time window is written into the time configuration bit of the policy template, the level value of the collection priority is written into the priority configuration bit of the policy template, the complete list of field names of key identifier fields is written into the identifier field configuration bit of the policy template, and the complete list of field names of core data load fields is written into the load field configuration bit of the policy template, thus completing the precise location and writing of all configuration information.

[0039] Bind a dedicated timer start flag to the time configuration bit of the strategy template. This flag is used to trigger the timing start action of the collection time window when the collection start time is reached. Bind a dedicated execution queue sort flag to the priority configuration bit of the strategy template. This flag is used to determine the fixed order of the collection task in the execution queue according to the priority level.

[0040] Write the complete list of key identifier field names into the index key area of ​​the strategy template. The content stored in this area will serve as the sole basis for accurate matching of data records in the subsequent data crawling process.

[0041] Write the complete list of field names of the core data payload fields into the data extraction area of ​​the strategy template. The content stored in this area serves as a fixed pointer to accurately read the target data content from heterogeneous data sources in the subsequent data crawling process.

[0042] The strategy template, which has been filled with all configuration information and bound to function tags, is transformed into an entity configuration object that can be directly parsed and executed by the data crawler. This results in a dynamic data collection strategy object that can be run directly. After performing data collection operations according to the configuration rules of the dynamic data collection strategy object, a standardized data object from heterogeneous data sources to be merged is obtained.

[0043] Pt.2. Generate a data feature identifier for the data object to be fused, and construct an initial data lineage map using the data feature identifier as a node; In this embodiment of the invention, generating a data feature identifier for the data object to be fused, and constructing an initial data lineage map using the data feature identifier as a node, includes: The data source address of the data object to be fused is hashed to derive an address digest, and a collection timestamp is generated anchored to the collection time. Traverse the field structure of the data object to be merged, extract the data type and field length of the fields, and serialize and concatenate them according to the order in which the fields appear to obtain the data structure outline signature of the heterogeneous data source. The address digest, the collection timestamp, and the data structure outline signature are encapsulated and combined to obtain the data feature identifier of the heterogeneous data source. The data feature identifier is registered as a node in the graph storage area, and dependency entries are retrieved from the preset data source dependency rule base. The dependency entries include upstream data source identifiers and downstream data source identifiers. The address digest of the node is compared with the downstream data source identifier in the dependency entry. When a match is found, a directed association is established between the node and the existing node of the upstream data source identifier. The direction of the directed association is from the existing node to the node. After traversing all dependency entries, the initial data lineage map of the heterogeneous data source is obtained.

[0044] The data source address of the data object to be merged is subjected to character-by-character parsing and fixed-rule mapping digest conversion processing. The complete data source address character content is converted into a fixed-length, globally unique and non-repeating character sequence to form an address digest for uniquely identifying the data source address. The precise time point of data collection is locked synchronously, and the complete time information of year, month, day, hour, minute, second and millisecond is converted into continuous and uninterrupted digital code to form a collection timestamp that accurately marks the collection time.

[0045] A hierarchical and comprehensive traversal reading operation is performed on all fields of the data object to be merged. The data type identifier, actual storage byte length, and field validity attribute of each field are extracted in sequence. All extracted field feature information is concatenated and integrated into a continuous structured character combination in strict accordance with the original order of appearance of the fields in the data object to be merged, so as to obtain a unique data structure outline signature that represents the overall field composition characteristics of the heterogeneous data source.

[0046] The address digest, collection timestamp, and data structure outline signature are arranged sequentially in a preset three-segment fixed combination format, redundant encapsulation is performed, and the encoding format and character length are unified to form a unique identifier that is exclusive to the current data object to be fused, without duplication and can be uniquely identified, thus obtaining the data feature identifier of the heterogeneous data source.

[0047] Data feature identifiers are treated as independent graph nodes and entered into a dedicated graph storage area to complete node information registration and persistent storage. Dependency entries containing explicit upstream and downstream data source identifiers are retrieved one by one from a preset and pre-configured data source dependency rule base.

[0048] Perform a full character-by-character consistency check on the address digest corresponding to the registered graph node and the downstream data source identifier recorded in the dependency entry. If the check result shows that every character is completely matched, establish a one-way directed association between the current registered node and the existing graph node corresponding to the upstream data source identifier. The association direction is strictly from the existing upstream node to the currently registered downstream node.

[0049] After retrieving, reconciling, and establishing directed relationships for all dependency entries in the data source dependency rule base, all registered nodes and all valid directed relationships are integrated to form a graph framework with complete nodes, clear relationships, and a standardized structure, thus obtaining the initial data lineage graph of heterogeneous data sources.

[0050] Pt.3. Perform content semantic parsing on the data object to be fused to generate an enhanced data object carrying semantic tags, and use the semantic tags to update the initial data lineage map to obtain a comprehensive data map; In this embodiment of the invention, the step of performing content semantic parsing on the data object to be fused to generate an enhanced data object carrying semantic tags, and using the semantic tags to update the initial data lineage map to obtain a comprehensive data map, includes: Lexical segmentation is performed on the core data load field of the data object to be fused, and the segmented words are semantically mapped to the entity lexicon. Hit words are marked as key entities. By extracting the co-occurrence relationships and order positions among the missed words, the attribute feature pairs and operation association pairs of the heterogeneous data source are obtained; The key entities, attribute feature pairs, and operation association pairs are arranged into semantic tags, and the semantic tags are attached to the data objects to be fused to obtain the enhanced data objects of the heterogeneous data sources; Using the node of the enhanced data object as the current node, iteratively scan other nodes in the initial data lineage graph. When the key entity verification of the current node is consistent with that of other nodes, establish an undirected semantic association edge between the current node and the other nodes. Next, the attribute feature pairs and operation association pairs of the current node are similar to the corresponding features of other nodes. When the verification passes, another undirected semantic association edge is established between the current node and the other nodes to obtain the comprehensive data graph of the heterogeneous data source.

[0051] According to the preset business semantic unit splitting rules, the text content of the core data load field inside the data object to be merged is lexically segmented segment by segment. Based on field boundaries, semantic breakpoints, and punctuation separation, the complete content is split into independent and unambiguous word units. All the segmented words are matched one by one with the standard business terms and object terms in the pre-built and stored entity lexicon. The matching words whose character content is completely consistent with the standard terms are uniformly marked as key entities.

[0052] For any missing word that does not match any standard term in the entity lexicon, enumerate all pairs of word combinations one by one and count the common frequency and fixed order of each combination. Extract feature combinations that describe data attributes, states, and types from word combinations as attribute feature pairs, and extract logical combinations that describe data processing actions, execution logic, and related behaviors as operation association pairs.

[0053] Key entities, attribute feature pairs, and operation association pairs are arranged in a fixed hierarchical order, with key entities first, attribute feature pairs in the middle, and operation association pairs last. These are then uniformly encapsulated and arranged into a semantic tag combination with complete structure and information. The arranged semantic tag combination is then bound and attached to the identifier layer area of ​​the data object to be merged, resulting in an enhanced data object from a heterogeneous data source carrying complete semantic information.

[0054] Set the data lineage node corresponding to the enhanced data object as the current node. Perform a traversal scan on all nodes except the current node within the entire range of the initial data lineage graph. Perform a character-by-character full consistency check on the key entities carried by the current node and the key entities carried by other nodes. When the check result is that all characters are completely identical, establish a bidirectional undirected semantic association edge between the current node and the corresponding other nodes.

[0055] The attribute feature pairs and operation association pairs carried by the current node are checked for complete consistency in terms of execution content, structure, and order with the attribute feature pairs and operation association pairs carried by other nodes. When the check result shows that all information is completely identical, a second bidirectional undirected semantic association edge is established between the current node and the corresponding other nodes. This completes the node association supplementation and structural upgrade of the initial data lineage graph, resulting in a comprehensive data graph of heterogeneous data sources.

[0056] Pt.4 In response to an external data fusion request, the target node is located based on the comprehensive data map and the preset fusion topology, and the association aggregation is performed to obtain a fusion data object containing the association relationship.

[0057] In this embodiment of the invention, the step of locating the target node based on the integrated data map and the preset fusion topology includes: The metadata tags of the enhanced data objects in the comprehensive data graph are parsed out, and the key identifier field values ​​are extracted from the metadata tags; Configure the key identifier field value as the query key, and perform a deep traversal scan within the node attribute area of ​​the preset fusion topology structure; The node attribute values ​​encountered during the depth traversal scan are compared with the query key for equivalence verification. When the verification result is equivalent, the node currently being verified is marked as the target node of the heterogeneous data source.

[0058] The target node for location also includes: When no node attribute value matching the query key is found within the node attribute area of ​​the fused topology, a new node is dynamically derived in the fused topology based on the overall metadata tag of the enhanced data object. Write the core data payload field of the enhanced data object into the corresponding attribute slot of the new node; The temporal marker and the reference relationship marker are parsed from the metadata tags of the enhanced data object, and the temporal marker and the reference relationship marker are recorded as the edge attributes between the new node and other existing nodes. The new node is identified as the target node of the heterogeneous data source.

[0059] The process of performing correlation aggregation to obtain a fused data object containing correlation relationships includes: Using the target node as the aggregation center, adjacent nodes that have direct directed edges or undirected semantic edges with the target node are collected in the comprehensive data graph to form an aggregated node set; Extract the core data payload field from the enhanced data objects corresponding to the nodes in the aggregated node set; Based on the type and direction of the edges connecting the target node and its adjacent nodes in the comprehensive data graph, determine the splicing order and nesting level of the core data load fields; The core data payload fields are merged and assembled according to the splicing order and nesting level to obtain the fused data object of the heterogeneous data source.

[0060] The multi-level metadata tags carried by the augmented data objects in the comprehensive data map are disassembled layer by layer. The full content of the tag identification layer, attribute layer and association layer is disassembled in sequence. The exclusive storage location of the key identification field is accurately located from the disassembled tag content, and the specific character and numerical content corresponding to the field is completely extracted to obtain the standardized and directly searchable key identification field value.

[0061] The extracted key identifier field values ​​are set as the query keys for data retrieval. Following the fixed traversal path rule that extends layer by layer from the top root node of the fused topology to the lower branch nodes, a layer-by-layer, thorough, and comprehensive scanning operation is performed in all node attribute areas of the preset fused topology.

[0062] Throughout the entire deep traversal scan, for each node's attribute area, all node attribute values ​​are compared character by character and matched across all dimensions of data type in sync with the query key to achieve full consistency verification.

[0063] When every character of a node's attribute value is exactly the same as the query key and the corresponding data type is completely matched, the node that has completed the full verification is officially marked as the target node of the heterogeneous data source.

[0064] After a full traversal scan of all node attribute regions in the fused topology, if no node attribute value is found that completely matches the character content and data type of the query key, an independent node unit is created in the preset blank storage location of the fused topology based on the source characteristics, structural characteristics, and semantic characteristics carried by all metadata tags of the enhanced data object, and a new node is obtained by completing the dynamic derivation operation.

[0065] Extract all business data content carried by the core data load field in the enhanced data object, and accurately write the content of the core data load field into the exclusive attribute slot of the new node according to the fixed attribute storage format and field correspondence preset by the new node, thus completing the filling of all attribute content of the new node.

[0066] The metadata tags of the enhanced data objects are disassembled layer by layer, and the time sequence tags used to identify the data generation time and the reference relationship tags used to identify the data association are accurately extracted. The time sequence tags and reference relationship tags are used as connection feature information between nodes, and the edge attributes between new nodes and existing nodes in the fused topology are fully recorded.

[0067] The new node that has completed the attribute content filling and edge attribute configuration will be officially identified as the target node of the heterogeneous data source as a valid node that meets the requirements of data fusion retrieval.

[0068] The target node is identified as the data aggregation center. A complete traversal is performed across all nodes in the comprehensive data graph to select all valid nodes that have direct directed or undirected semantic connections with the target node. All selected nodes are then uniformly collected and organized to form a complete and clearly defined aggregated node set.

[0069] For each independent node in the aggregation node set, the corresponding enhanced data object is retrieved one by one. The exclusive core data payload field is completely extracted from the fixed semantic storage area of ​​the enhanced data object, and the full extraction operation of the core data payload fields of all adjacent nodes is completed.

[0070] Based on the specific types of connecting edges between the target node and each adjacent node in the comprehensive data graph, as well as the unidirectional or bidirectional pointing information of the edges, the order in which each core data load field participates in the data combination and its nesting level in the overall fusion structure are accurately determined.

[0071] All extracted core data payload fields are arranged sequentially according to a predetermined concatenation order. At the same time, layered nesting integration is performed according to a determined nesting level to complete the standardized merging and assembly of all fields, resulting in a fused data object of heterogeneous data sources containing complete relationships and business logic.

[0072] like Figure 2 The diagram shown is a functional block diagram of a heterogeneous data fusion processing system based on big data acquisition, provided by an embodiment of the present invention.

[0073] The heterogeneous data fusion processing system based on big data acquisition described in this invention can be installed in electronic devices. Depending on the functions implemented, the system may include a source domain acquisition module, a lineage-based foundation module, a semantic spectrum enhancement module, and a graph fusion module. The modules described in this invention can also be referred to as units, which are a series of computer program segments that can be executed by the processor of an electronic device and perform a fixed function, stored in the memory of the electronic device.

[0074] In this embodiment, the functions of each module / unit are as follows: The source domain acquisition strategy module is used to perform source feature analysis on the configuration information of heterogeneous data sources and generate dynamic acquisition strategies for the heterogeneous data sources in order to obtain the data objects to be fused from the heterogeneous data sources. The bloodline foundation module is used to generate a data feature identifier for the data object to be fused, and to construct an initial data bloodline map using the data feature identifier as a node. The semantic spectrum enhancement module is used to perform content semantic parsing on the data object to be fused, generate an enhanced data object carrying semantic tags, and use the semantic tags to update the initial data lineage map to obtain a comprehensive data map. The graph fusion module is used to respond to external data fusion requests, locate target nodes based on the comprehensive data graph and the preset fusion topology, and perform association aggregation to obtain fusion data objects containing association relationships.

[0075] In the several embodiments provided by this invention, it should be understood that the disclosed methods and systems can be implemented in other ways. For example, the system embodiments described above are merely illustrative; for instance, the division of modules is only a logical functional division, and other division methods may be used in actual implementation.

[0076] The modules described as separate components may or may not be physically separate. The components shown as modules may or may not be physical units; that is, they may be located in one place or distributed across multiple network units. Some or all of the modules can be selected to achieve the purpose of this embodiment according to actual needs.

[0077] Furthermore, the functional modules in the various embodiments of the present invention can be integrated into one processing unit, or each unit can exist physically separately, or two or more units can be integrated into one unit. The integrated unit can be implemented in hardware or in the form of hardware plus software functional modules.

[0078] It will be apparent to those skilled in the art that the present invention is not limited to the details of the exemplary embodiments described above, and that the present invention can be implemented in other specific forms without departing from the spirit or essential characteristics of the present invention.

[0079] This application embodiment can acquire and process relevant data based on artificial intelligence technology. Artificial intelligence is the theory, method, technology, and application system that uses digital computers or machines controlled by digital computers to simulate, extend, and expand human intelligence, perceive the environment, acquire knowledge, and use that knowledge to obtain optimal results.

[0080] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention and are not intended to limit it. Although the present invention has been described in detail with reference to preferred embodiments, those skilled in the art should understand that modifications or equivalent substitutions can be made to the technical solutions of the present invention without departing from the spirit and scope of the technical solutions of the present invention.

Claims

1. A heterogeneous data fusion processing method based on big data acquisition, characterized in that, The method includes: Pt.

1. Perform source feature analysis on the configuration information of heterogeneous data sources to generate dynamic acquisition strategies for the heterogeneous data sources, so as to obtain the data objects to be fused from the heterogeneous data sources. Pt.

2. Generate a data feature identifier for the data object to be fused, and construct an initial data lineage map using the data feature identifier as a node; Pt.

3. Perform content semantic parsing on the data object to be fused to generate an enhanced data object carrying semantic tags, and use the semantic tags to update the initial data lineage map to obtain a comprehensive data map; Pt.4 In response to an external data fusion request, the target node is located based on the comprehensive data map and the preset fusion topology, and the association aggregation is performed to obtain a fusion data object containing the association relationship.

2. The heterogeneous data fusion processing method based on big data acquisition as described in claim 1, characterized in that, The step of performing source feature analysis on the configuration information of heterogeneous data sources to generate dynamic acquisition strategies for the heterogeneous data sources includes: Extract the interface type, data update frequency, field structure parameters, and historical data samples of the heterogeneous data source from its configuration information; Based on the interface type and the data update frequency, set the collection time window and collection priority for the heterogeneous data source; The data transmission method indicated by the interface type is used to determine the start time of the acquisition time window, and the data update frequency is used to determine the duration of the acquisition time window. Based on the field structure parameters and the historical data samples, the key identification fields and core data load fields in the heterogeneous data source are identified, and field location identifiers are generated. The acquisition time window, acquisition priority, key identifier field, and core data load field are encapsulated to generate a dynamic acquisition strategy for the heterogeneous data source.

3. The heterogeneous data fusion processing method based on big data acquisition as described in claim 2, characterized in that, The step of identifying key identifier fields and core data load fields in the heterogeneous data source based on the field structure parameters and the historical data samples, and generating field location identifiers, includes: Extract the name, data type, and primary / foreign key relationships between each field from the field structure parameters to generate a list of field attributes; The frequency of occurrence of non-repeating values ​​corresponding to each field is counted from the historical data sample to generate a field value distribution record; Based on the list of field attributes and the field value distribution record, the discrimination weight value of the field is determined according to the following formula: ; In the formula, The discrimination weight value is... The number of unique values ​​for the field in the historical data sample. The total number of values ​​for the field in the historical data sample. The number of times the field appears in the field attribute list as a primary key or foreign key. This represents the maximum number of primary and foreign key references for all fields in the aforementioned list of field attributes. Fields whose discrimination weight values ​​exceed a preset threshold are identified as key identifier fields, and fields with the longest data length among those whose discrimination weight values ​​are lower than the preset threshold are identified as core data load fields, thus obtaining field location identifiers for the key identifier fields and the core data load fields.

4. The heterogeneous data fusion processing method based on big data acquisition as described in claim 3, characterized in that, The dynamic acquisition strategy for generating the heterogeneous data source, to obtain the data object to be merged from the heterogeneous data source, includes: The collection time window, the collection priority, the list of field names of the key identifier field, and the list of field names of the core data load field are written into the corresponding field positions of the strategy template in a preset encapsulation order. In the strategy template, a timer start flag is associated with the collection time window, and an execution queue sorting flag is associated with the collection priority; Write the list of field names of the key identifier fields into the index key area of ​​the strategy template, as the basis for record matching during subsequent data crawling; Write the list of field names of the core data load fields into the data extraction area of ​​the strategy template, as an identifier for reading target data content from heterogeneous data sources during subsequent data crawling; The completed strategy template is instantiated into a dynamic collection strategy object that can be parsed and executed by the data crawler, thus obtaining the data object to be merged from the heterogeneous data source.

5. The heterogeneous data fusion processing method based on big data acquisition as described in claim 1, characterized in that, The step of generating a data feature identifier for the data object to be fused, and constructing an initial data lineage map using the data feature identifier as a node, includes: The data source address of the data object to be fused is hashed to derive an address digest, and a collection timestamp is generated anchored to the collection time. Traverse the field structure of the data object to be merged, extract the data type and field length of the fields, and serialize and concatenate them according to the order in which the fields appear to obtain the data structure outline signature of the heterogeneous data source. The address digest, the collection timestamp, and the data structure outline signature are encapsulated and combined to obtain the data feature identifier of the heterogeneous data source. The data feature identifier is registered as a node in the graph storage area, and dependency entries are retrieved from the preset data source dependency rule base. The dependency entries include upstream data source identifiers and downstream data source identifiers. The address digest of the node is compared with the downstream data source identifier in the dependency entry. When a match is found, a directed association is established between the node and the existing node of the upstream data source identifier. The direction of the directed association is from the existing node to the node. After traversing all dependency entries, the initial data lineage map of the heterogeneous data source is obtained.

6. The heterogeneous data fusion processing method based on big data acquisition as described in claim 1, characterized in that, The process of performing content semantic parsing on the data objects to be fused, generating enhanced data objects carrying semantic tags, and updating the initial data lineage map using the semantic tags to obtain a comprehensive data map includes: Lexical segmentation is performed on the core data load field of the data object to be fused, and the segmented words are semantically mapped to the entity lexicon. Hit words are marked as key entities. By extracting the co-occurrence relationships and order positions among the missed words, the attribute feature pairs and operation association pairs of the heterogeneous data source are obtained; The key entities, attribute feature pairs, and operation association pairs are arranged into semantic tags, and the semantic tags are attached to the data objects to be fused to obtain the enhanced data objects of the heterogeneous data sources; Using the node of the enhanced data object as the current node, iteratively scan other nodes in the initial data lineage graph. When the key entity verification of the current node is consistent with that of other nodes, establish an undirected semantic association edge between the current node and the other nodes. Next, the attribute feature pairs and operation association pairs of the current node are similar to the corresponding features of other nodes. When the verification passes, another undirected semantic association edge is established between the current node and the other nodes to obtain the comprehensive data graph of the heterogeneous data source.

7. The heterogeneous data fusion processing method based on big data acquisition as described in claim 1, characterized in that, The method of locating target nodes based on the integrated data map and the preset fusion topology includes: The metadata tags of the enhanced data objects in the comprehensive data graph are parsed out, and the key identifier field values ​​are extracted from the metadata tags; Configure the key identifier field value as the query key, and perform a deep traversal scan within the node attribute area of ​​the preset fusion topology structure; The node attribute values ​​encountered during the depth traversal scan are compared with the query key for equivalence verification. When the verification result is equivalent, the node currently being verified is marked as the target node of the heterogeneous data source.

8. The heterogeneous data fusion processing method based on big data acquisition as described in claim 7, characterized in that, The target node for location also includes: When no node attribute value matching the query key is found within the node attribute area of ​​the fused topology, a new node is dynamically derived in the fused topology based on the overall metadata tag of the enhanced data object. Write the core data payload field of the enhanced data object into the corresponding attribute slot of the new node; The temporal marker and the reference relationship marker are parsed from the metadata tags of the enhanced data object, and the temporal marker and the reference relationship marker are recorded as the edge attributes between the new node and other existing nodes. The new node is identified as the target node of the heterogeneous data source.

9. The heterogeneous data fusion processing method based on big data acquisition as described in claim 8, characterized in that, The process of performing correlation aggregation to obtain a fused data object containing correlation relationships includes: Using the target node as the aggregation center, adjacent nodes that have direct directed edges or undirected semantic edges with the target node are collected in the comprehensive data graph to form an aggregated node set; Extract the core data payload field from the enhanced data objects corresponding to the nodes in the aggregated node set; Based on the type and direction of the edges connecting the target node and its adjacent nodes in the comprehensive data graph, determine the splicing order and nesting level of the core data load fields; The core data payload fields are merged and assembled according to the splicing order and nesting level to obtain the fused data object of the heterogeneous data source.

10. A heterogeneous data fusion processing system based on big data acquisition, characterized in that, The system for implementing the heterogeneous data fusion processing method based on big data acquisition as described in claim 1 includes: The source domain acquisition strategy module is used to perform source feature analysis on the configuration information of heterogeneous data sources and generate dynamic acquisition strategies for the heterogeneous data sources to obtain the data objects to be fused from the heterogeneous data sources. The bloodline foundation module is used to generate a data feature identifier for the data object to be fused, and to construct an initial data bloodline map using the data feature identifier as a node. The semantic spectrum enhancement module is used to perform content semantic parsing on the data object to be fused, generate an enhanced data object carrying semantic tags, and use the semantic tags to update the initial data lineage map to obtain a comprehensive data map. The graph fusion module is used to respond to external data fusion requests, locate target nodes based on the comprehensive data graph and the preset fusion topology, and perform association aggregation to obtain fusion data objects containing association relationships.