Mass data association fusion method, equipment and medium

Through AI association fusion models and hierarchical dynamic knowledge graphs, the problem of balancing data quality and real-time performance in traditional technologies has been solved, and efficient, accurate, and real-time fusion of massive heterogeneous data has been achieved. This adapts to the data association needs of different scenarios and improves the efficiency of data processing and the reliability of decision-making.

CN120805082AActive Publication Date: 2025-10-17DIANKEYUN (BEIJING) TECH CO LTD

Patent Information

Application Number
CN202511299816.3
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-09-12
Publication Date
2025-10-17
Estimated Expiration
2045-09-12

AI Technical Summary

Technical Problem

Traditional batch processing models find it difficult to balance data quality and real-time performance, resulting in delayed data association and inaccurate conflict resolution, which in turn affects decision-making bias. Existing data fusion solutions also lack scenario-based adjustment capabilities and cannot meet the real-time association needs of massive heterogeneous data.

Method used

An AI association fusion model is adopted to perform association fusion through hierarchical dynamic knowledge graph and dual-tower sub-model, combined with consistency judgment rules to achieve real-time semantic mapping and optimization, including preprocessing, annotated data stream generation, two-dimensional semantic vector and dynamic quality label processing, and semantic feature extraction using convolutional neural network and Transformer encoder, real-time monitoring and optimization of model performance.

Benefits of technology

It achieves efficient, accurate and real-time integration of massive data, adapts to the needs of different scenarios, improves cross-source correlation efficiency and decision-making reliability, reduces decision-making bias, and adapts to the growth of data sources and changes in scenarios.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120805082A_ABST
    Figure CN120805082A_ABST
Patent Text Reader

Abstract

The invention provides a mass data association fusion method and device and a medium, and the method comprises the steps: firstly receiving mass data from a plurality of heterogeneous data sources, and carrying out the preprocessing of the data; then, the data stream is integrated into an annotated data stream; wherein the labeled data stream comprises a two-dimensional semantic vector and a dynamic quality label; and finally, inputting the data into an AI association fusion model to obtain a fusion result. Wherein the association fusion model is provided with an association fusion and semantic mapping layer, and is used for carrying out association fusion according to the layered dynamic knowledge graph and a preset fusion rule, and carrying out semantic mapping according to the twin-tower sub-model and a consistency judgment rule. According to the method, real-time updating of the incidence relation is achieved through the layered dynamic knowledge graph, and the real-time scene requirement is met; the association complexity of mass data is reduced by using a double-tower model, and the processing speed is increased; and in combination with the judgment rule, the conflict resolution precision is improved. Finally, efficient, accurate and real-time fusion of mass data is realized, and a reliable basis is provided for analysis and decision of various data.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the technical field of data processing, and in particular to a mass data association fusion method, device and medium. BACKGROUND

[0002] With the development of digital economy, the mass data generated in many fields presents the characteristics of multi-source heterogeneity, and the data sources cover databases, real-time interfaces, log files, etc. Different identifiers are often used for the same entity in different data sources, resulting in a significant semantic gap in cross-source association. The traditional manual configuration of mapping rules is difficult to adapt to the rapid growth of data sources, and cannot meet the association efficiency requirements in dynamic scenarios.

[0003] At the same time, mass data generally have quality problems such as missing values, abnormal values, and duplicate records, while real-time transactions, risk control early warning, and user behavior analysis require strict data processing delay. The traditional batch processing mode is difficult to balance data quality and real-time performance, and is prone to decision bias due to data association lag and inaccurate conflict resolution. In addition, the existing data fusion scheme has the problems of static knowledge graph update, poor adaptability of bucketing strategy, and lack of scene adjustment ability of fusion rules, which further restricts the efficiency and reliability of mass data association fusion. Therefore, it is urgent to build an association fusion technology system that takes into account real-time performance, accuracy, and scene adaptability. SUMMARY

[0004] The embodiments of the present application provide a mass data association fusion method, device and medium to solve the problem that the traditional batch processing mode at the present stage is difficult to balance data quality and real-time performance, and is prone to decision bias due to data association lag and inaccurate conflict resolution.

[0005] In a first aspect, the embodiments of the present application provide a mass data association fusion method, comprising: receiving mass data from a plurality of heterogeneous data sources, and preprocessing the data; integrating the preprocessed data into labeled data stream; wherein the labeled data stream comprises a double-dimensional semantic vector and a dynamic quality label; inputting the labeled data stream into an AI association fusion model to obtain a fusion result; The AI association fusion model is provided with an association fusion and semantic mapping layer; the association fusion and semantic mapping layer is used for association fusion according to a hierarchical dynamic knowledge graph and a preset fusion rule, and semantic mapping according to a double-tower sub-model and a consistency resolution rule.

[0006] In a possible implementation, inputting the labeled data stream into the AI association fusion model to obtain the fusion result comprises: The two-dimensional semantic vector in the labeled data stream is dimensionally normalized to match the vector dimension with the preset input dimension of the model; and based on the credibility level of the data source, the dynamic quality label is weighted and initialized to obtain an adapted data stream; The preset fusion rule and the adapted data stream are transmitted to the associated fusion and semantic mapping layer to obtain a preliminary association set and an initial semantic mapping result to determine the fusion result. The associated fusion and semantic mapping layer extracts associated features corresponding to the adapted data stream from each level based on the hierarchical structure of the hierarchical dynamic knowledge graph, and combines the preset fusion rule to preliminarily associate and aggregate homologous data and associated data in the adapted data stream to generate a preliminary association set. The associated fusion and semantic mapping layer calls a double-tower sub-model to perform semantic mapping processing on the preliminary association set; the first sub-tower is used to receive the two-dimensional semantic vector in the preliminary association set and extract local semantic features through a convolutional neural network, and the second sub-tower is used to receive a semantic representation vector of a corresponding associated node in the hierarchical dynamic knowledge graph and extract global semantic features through a Transformer encoder; and the initial semantic mapping result is determined according to the cosine similarity of the output features of the two sub-towers.

[0007] In one possible implementation, the hierarchical dynamic knowledge graph includes three levels of a basic data layer, a feature association layer, and a semantic fusion layer, which are dynamically connected through associated edges; the basic data layer stores original entity and attribute information of each data source, the feature association layer stores feature-level association relationships between entities, and the semantic fusion layer stores semantic-level association relationships between entities and domain knowledge rules; the adapted data stream is transmitted to the associated fusion and semantic mapping layer to obtain a preliminary association set and an initial semantic mapping result, including: The reference entity information corresponding to each entity in the adapted data stream is extracted from the basic data layer, an entity recognition algorithm is used to recognize the entity type in the adapted data stream, and an entity alignment algorithm is used to match the to-be-fused entity with the reference entity to mark homologous entities and suspected associated entities; Based on the preset fusion rule and the feature association relationship stored in the feature association layer, the feature-level association aggregation is performed on the marked homologous entities and suspected associated entities to obtain a preliminary association set. The domain knowledge rules in the semantic fusion layer are called to perform semantic enhancement on the preliminary association set output by the feature association layer to obtain an initial semantic mapping result.

[0008] In a possible implementation, the preset fusion rule includes a basic rule set and scene adaptation parameters; the basic rule set includes a data priority rule, an attribute fusion rule, and an association strength calculation rule; the scene adaptation parameters include a rule triggering threshold, a feature weight coefficient, and a scene type identifier, and are used to dynamically adjust the execution logic of the basic rule set according to a specific application scene; based on the preset fusion rule and the feature association relationship stored in the feature association layer, the feature-level association aggregation is performed on the labeled homologous entities and suspected associated entities, to obtain a preliminary association set, including: For the labeled homologous entities, according to the attribute types thereof, the attribute fusion rule corresponding to the basic rule set is called to perform fusion processing on the multi-source attribute data, to generate a unified entity attribute description; For the labeled suspected associated entities, the association strength calculation rule in the basic rule set and the association relationship stored in the feature association layer are used to calculate the association strength between the entities; According to the association threshold set in the scene adaptation parameters, the suspected associated entities whose association strength satisfies the threshold condition are determined to exist association, and the feature information thereof is aggregated to form the preliminary association set.

[0009] In a possible implementation, the method includes: The data-driven semantic vector of the entity is extracted from the preliminary association set as the input of the first sub-tower, and the graph-driven semantic vector corresponding to the entity is extracted from the hierarchical dynamic knowledge graph as the input of the second sub-tower; The first sub-tower and the second sub-tower respectively perform feature coding on the input vectors thereof, to output corresponding semantic feature vectors; The semantic feature vectors output by the first sub-tower and the semantic feature vectors output by the second sub-tower are subjected to similarity calculation, to generate an initial semantic mapping result; Based on a preset consistency adjudication rule, the initial semantic mapping result is verified and corrected, to identify and solve semantic conflicts; According to the adjudication result, a final semantic mapping output is generated, and the conflict cases and / or artificial correction results generated in the adjudication process are used to iteratively optimize the double-tower sub-model and / or the consistency adjudication rule.

[0010] In a possible implementation, based on a preset consistency adjudication rule, the initial semantic mapping result is verified and corrected, to identify and solve semantic conflicts, including: The semantic similarity score is compared with a first threshold value, if the semantic similarity score is higher than the first threshold value, it is directly determined that the semantic mapping is successful, and the output result is output; if the semantic similarity score is lower than a second threshold value, it is directly determined that the semantic mapping fails, and the mapping is rejected; if the semantic similarity score is between the first threshold value and the second threshold value, a secondary adjudication is triggered; the first threshold value is higher than the second threshold value. For entities that trigger a second-level adjudication, the hierarchical dynamic knowledge graph is queried to obtain their associated relationships in their respective data sources and calculate the consistency score of the associated relationships. The semantic similarity score and the relationship consistency score are weighted and summed to obtain a comprehensive adjudication score. If the comprehensive adjudication score is higher than the third threshold, the semantic mapping is considered successful; otherwise, a third-level adjudication is triggered. For entities that trigger the third-level adjudication, conflict detection is performed using a preset logical rule library. If a logical conflict is detected, the semantic mapping is judged to have failed. If no logical conflict is detected, the comprehensive adjudication score is corrected in combination with the dynamic quality label, and whether the mapping is successful is determined based on the corrected final score.

[0011] In one possible implementation, the AI ​​association fusion model is further provided with a dynamic optimization and scheduling layer, and the method further includes: Through the dynamic optimization and scheduling layer, the operating status and output results of the association fusion and semantic mapping layers are monitored in real time to collect performance data; Analyze performance data based on preset performance indicators and resource constraints, identify performance bottlenecks or optimization opportunities, and generate optimization instructions accordingly; The optimization instructions are sent to the association fusion and semantic mapping layer to dynamically adjust and optimize the computing tasks, model parameters and fusion rules.

[0012] In one possible implementation, based on preset performance indicators and resource constraints, performance data is analyzed to identify performance bottlenecks or optimization opportunities, and optimization instructions are generated accordingly, including: Based on performance data, through preset optimization algorithms, key bottlenecks affecting performance are identified and optimization instructions are generated; among them, optimization instructions include: adjusting the network weights of the dual-tower sub-model, updating the threshold parameters in the consistency judgment rules, or adding, deleting, and modifying entities or relationships in the hierarchical dynamic knowledge graph.

[0013] An embodiment of the present invention provides an electronic device including a memory and a processor. The memory stores a computer program, and the processor implements the above-mentioned method for associating and fusing massive data when executing the computer program.

[0014] An embodiment of the present invention provides a computer-readable storage medium storing a computer program. When the computer program is executed by a processor, the computer program implements the above-mentioned method for associating and fusing massive data.

[0015] Compared with the prior art, the embodiment of the present application provides a mass data association fusion method, which first receives mass data from multiple heterogeneous data sources, pre-processes the data, then integrates the pre-processed data into labeled data stream, wherein the labeled data stream includes a two-dimensional semantic vector and a dynamic quality label, finally inputs the labeled data stream into an AI association fusion model to obtain a fusion result, wherein the AI association fusion model is provided with an association fusion and semantic mapping layer, the association fusion and semantic mapping layer is used for association fusion according to a hierarchical dynamic knowledge graph and a preset fusion rule, and simultaneously performs semantic mapping according to a double-tower sub-model and a consistency adjudication rule. The hierarchical dynamic knowledge graph is used to realize real-time updating of the association relationship, meet the real-time scene requirement, the double-tower model is used to reduce the mass data association complexity and speed up the processing speed, and the consistency adjudication rule containing the urgency is used to improve the conflict resolution accuracy. Finally, the mass data is efficiently, accurately and real-timely fused, and a reliable basis is provided for various data analysis and decision-making. BRIEF DESCRIPTION OF DRAWINGS

[0016] Figure 1 is an implementation flowchart of the mass data association fusion method provided by the embodiment of the present application. DETAILED DESCRIPTION

[0017] The embodiment of the present application will be described in detail below with reference to the drawings.

[0018] Figure 1 is an implementation flowchart of the mass data association fusion method provided by the embodiment of the present application. As shown in Figure 1 , the method can include: S110, receiving mass data from multiple heterogeneous data sources, pre-processing the data; S120, integrating the pre-processed data into labeled data stream; wherein the labeled data stream includes a two-dimensional semantic vector and a dynamic quality label; S130, inputting the labeled data stream into an AI association fusion model to obtain a fusion result; , wherein the AI association fusion model is provided with an association fusion and semantic mapping layer; the association fusion and semantic mapping layer is used for association fusion according to a hierarchical dynamic knowledge graph and a preset fusion rule, and simultaneously performs semantic mapping according to a double-tower sub-model and a consistency adjudication rule.

[0019] In the embodiment of the present application, first, a stream-batch integrated access mechanism is used to synchronously receive mass heterogeneous data from different systems, which specifically covers three types of core data sources: Structured data sources: such as user table, order table in relational database, establish stable connection through JDBC interface, pull data according to preset frequency (such as real-time transaction data 1 time per second, historical data 1 time per hour), ensure that the structured data field is complete; Semi-structured / unstructured data sources: such as JSON format logs returned by API interface, text log files (such as user behavior logs, system running logs), interface through Kafka, Flink and other stream processing frameworks, adopt "sharding parsing + format verification" strategy. Large volume log files (such as more than 10GB) are divided into 1MB slices, and the format legality is verified piece by piece; Third-party data sources: such as Excel data table provided by cooperative institutions, CSV file, access through OSS, FTP and other file transmission protocols, automatically identify file encoding (UTF-8, GBK, etc.) and separator (comma, tab, etc.), realize "no manual configuration format" automatic access.

[0020] The preprocessing stage is carried out around "data cleanliness improvement" and "subsequent association adaptation", including three key operations: Preliminary filtering of redundant and abnormal data: adopt "rule filtering + statistical analysis" dual mechanism. First, eliminate obviously invalid data through business rules; then, identify extreme abnormal values in numerical data through 3sigma criterion, and eliminate low value data through null value proportion statistics to reduce redundancy in subsequent processing; Data format standardization: unify the format of the same type of fields in different data sources. For example, convert different formats of time fields to a standard format; for numerical fields, uniformly retain 2 decimal places to avoid association failure caused by format differences; Core feature extraction: for each data, extract the core features required for subsequent semantic modeling and quality evaluation.

[0021] In the embodiment of the present application, "general semantics + domain semantics" dual-dimensional coding is adopted to generate semantic vectors that can be associated across sources: General semantic vector: adopt pre-trained word vector model (such as Word2Vec, Glove) to encode entity identifiers and attribute values in the data, convert text / numerical information into low-dimensional dense vectors (such as 128 dimensions), and ensure that the vector similarity of "same meaning information" in different data sources is high; Domain semantic vector: introduce domain-specific glossary to "fine-tune" the general semantic vector. Avoid association errors caused by "same name different meaning"; Vector integration: concatenate the general semantic vector and the domain semantic vector according to the weight to form the final dual-dimensional semantic vector, each vector carries "data meaning + domain attribute" dual information, providing semantic basis for cross-source association.

[0022] In the embodiments of the present application, based on the analysis result of the preprocessed data, a dynamic quality label containing "health status" and "correlation potential" is generated: Health score (1-100 points): calculated from three dimensions of data integrity (number of missing fields / total number of fields), timeliness (current time-data collection time), volatility (deviation of field value from historical mean).

[0023] Correlation potential score (1-10 points): based on the integrity of entity identification in the data (such as whether it contains a unique user identifier), the clarity of the semantic vector (vector modulus fluctuation range) is evaluated.

[0024] Label integration: combine health score and correlation potential score into dynamic quality label (such as "health 92+correlation 10" "health 65+correlation 5"), label real-time refresh with data update, provide quantitative basis for subsequent screening of high-quality data.

[0025] Bind each piece of preprocessed data with the corresponding "two-dimensional semantic vector + dynamic quality label" to form a structured labeled data stream, ensuring that each piece of data carries "correlation, evaluation" key information.

[0026] In the embodiments of the present application, the two-dimensional semantic vector enables automatic matching of heterogeneous labels from different data sources through vector similarity, avoiding the complexity of traditional manual configuration of mapping rules, and improving the accuracy of cross-source entity correlation from 65%.

[0027] In some embodiments, the labeled data stream is input into an AI correlation fusion model to obtain a fusion result, including: Dimensional normalization is performed on the two-dimensional semantic vector in the labeled data stream to match the vector dimension with the preset input dimension of the model; at the same time, based on the credibility level of the data source, the dynamic quality label is weighted initialized to obtain an adapted data stream; The preset fusion rule and the adapted data stream are transmitted to the correlation fusion and semantic mapping layer to obtain a preliminary correlation set and an initial semantic mapping result to determine the fusion result; Among them, the correlation fusion and semantic mapping layer is based on the hierarchical structure of the hierarchical dynamic knowledge graph, extracts the correlation features corresponding to the adapted data stream from each layer, combines the preset fusion rule, and preliminarily correlates and aggregates the homologous data and correlated data in the adapted data stream to generate a preliminary correlation set; The double-tower sub-model in the association fusion and semantic mapping layer is called to perform semantic mapping processing on the preliminary association set; the first sub-tower is used to receive the two-dimensional semantic vectors in the preliminary association set, and local semantic features are extracted through a convolutional neural network, and the second sub-tower is used to receive the semantic representation vectors of the corresponding association nodes in the hierarchical dynamic knowledge graph, and global semantic features are extracted through a Transformer encoder; and the initial semantic mapping result is determined according to the cosine similarity of the output features of the two sub-towers.

[0028] In the embodiment of the application, first, the labeled data stream is processed: on the one hand, the dimension normalization operation is performed on the two-dimensional semantic vectors therein, and the dimension size of the vector is adjusted through a standardization algorithm to strictly match the input dimension requirement of the model preset, so as to ensure that the vector data can be effectively received and processed by the model; on the other hand, according to the credibility level of the data source (such as the credibility difference of different sources such as core database and third-party interface), the health score and association potential score in the dynamic quality label are initialized, and different weights are given to data from different sources, so as to finally form an adapted data stream and improve the matching degree of data and model.

[0029] Then, the preset fusion rule (including scenario-based association logic, conflict processing priority, etc.) is transmitted to the association fusion and semantic mapping layer together with the adapted data stream, and the preliminary association set and the initial semantic mapping result are generated through the cooperative processing of the layer, which serves as the basis for determining the final fusion result.

[0030] The processing of the association fusion and semantic mapping layer includes two core links: Preliminary association set generation: the layer extracts the association features (such as hierarchical relationships between entities, attribute associations, etc.) corresponding to the adapted data stream from each level of the knowledge graph according to the hierarchical structure of the hierarchical dynamic knowledge graph (the hierarchical division of the core relationship layer and the non-core relationship layer); combined with the association logic defined in the preset fusion rule, the same source data and associated data with potential association in the adapted data stream are preliminarily associated and aggregated, and the data with direct or indirect association is classified and integrated to form a preliminary association set, reducing the data dispersion in subsequent processing.

[0031] Initial semantic mapping processing: a double-tower sub-model is called in the association fusion and semantic mapping layer to perform deep semantic mapping on the preliminary association set. Among them, the first sub-tower focuses on processing the two-dimensional semantic vector in the preliminary association set, and extracts local detailed features (such as specific attributes of entities, local semantic associations) in the vector through a convolutional neural network; the second sub-tower receives the semantic representation vector of the association node corresponding to the preliminary association set in the hierarchical dynamic knowledge graph, and captures global semantic features (such as the position of the entity in the whole knowledge network, the global association with other entities) by means of a Transformer encoder; finally, the cosine similarity between the output features of the two sub-towers is calculated to determine the initial semantic mapping result, quantifying the degree of semantic association between data and providing accurate basis for the final fusion result.

[0032] By constructing an AI association fusion system suitable for massive heterogeneous data, the semantic gap problem of multi-source data cross-domain association can be effectively solved. Relying on the two-dimensional semantic vector, the identification difference between different data sources is eliminated, the tedious operation of manually configuring mapping rules is reduced, and the cross-source association efficiency is improved. The combination of intelligent cleaning and hierarchical knowledge graph technology can accurately identify and handle data quality problems and conflict risks, reducing the decision bias caused by data quality defects. The application of the double-tower sub-model and the dynamic optimization mechanism greatly improves the efficiency of massive data association fusion, adapting to the rapid response demand in real-time scenarios. The full-link dynamic optimization capability ensures that the model can be continuously adjusted with the growth of data sources and changes in scenarios, adapting to the business characteristics of different fields, and finally realizing efficient collaborative processing of massive heterogeneous data from access to fusion. It provides high-quality and reliable data support for subsequent accurate decision-making, significantly improves the efficiency and accuracy of data fusion, and meets the core needs of massive data association fusion in various industries.

[0033] In some embodiments, the hierarchical dynamic knowledge graph includes three levels of basic data layer, feature association layer and semantic fusion layer, which are dynamically connected through association edges; wherein the basic data layer stores the original entity and attribute information of each data source, the feature association layer stores the feature-level association relationship between entities, and the semantic fusion layer stores the semantic-level association relationship between entities and domain knowledge rules; the adapted data stream is transmitted to the association fusion and semantic mapping layer to obtain the preliminary association set and the initial semantic mapping result, including: Extract the reference entity information corresponding to each entity in the adapted data stream from the basic data layer, identify the entity type in the adapted data stream using an entity recognition algorithm, and match the to-be-fused entities with the reference entities by an entity alignment algorithm to mark homologous entities and suspected associated entities; Based on the preset fusion rules and the feature association relationship stored in the feature association layer, the feature-level association of the marked homologous entities and suspected associated entities is aggregated to obtain the preliminary association set; The domain knowledge rules in the semantic fusion layer are called to perform semantic enhancement on the preliminary association set output by the feature association layer, to obtain an initial semantic mapping result.

[0034] In the embodiment of the present application, the layer dynamic knowledge graph adopts a three-level architecture design, including a basic data layer, a feature association layer and a semantic fusion layer, and information interaction and association update are realized between each layer through an association edge with dynamic attributes, forming a complete knowledge network system. Among them, the basic data layer serves as the bottom support of the knowledge graph, responsible for storing the original entities and their attribute information of each data source without deep processing, covering the basic data such as the basic identification and inherent characteristics of the entity, providing original information basis for the upper layer processing; the feature association layer is in the middle layer, focusing on storing the association relationship between entities based on the feature dimension, such as the direct association of entities in attribute values, feature vectors, etc., reflecting the surface feature association between entities; the semantic fusion layer as the top layer, stores the deep association relationship between entities based on the semantic level, and integrates the professional knowledge rules in the field, providing semantic interpretation and rule constraints for the association relationship.

[0035] The process of transmitting the adapted data stream to the association fusion and semantic mapping layer to obtain the preliminary association set and the initial semantic mapping result is specifically divided into three progressive links: Firstly, entity matching and marking are completed based on the basic data layer. The reference entity information corresponding to each entity in the adapted data stream (i.e. the stored, verified standard entity data) is retrieved and extracted from the basic data layer; the type of each entity in the adapted data stream is determined through an entity recognition algorithm to ensure the consistency of entity classification; then, an entity alignment algorithm is used to accurately compare the to-be-fused entities with the reference entities from multiple dimensions such as attribute similarity and identification consistency, mark entities of the same nature as homologous entities (such as entities with different identifications but pointing to the same user), and mark entities with potential association clues as suspected associated entities (such as entities sharing part of the characteristics but not explicitly confirmed to be associated), laying a foundation for subsequent association aggregation.

[0036] Secondly, the preliminary association set is generated relying on the feature association layer. Under the guidance of the preset fusion rules (which define the logic of feature association, priority, etc.), combined with the feature-level association relationship between entities stored in the feature association layer, the homologous entities and suspected associated entities marked are subjected to association aggregation processing: for homologous entities, their attribute information is directly merged to eliminate redundancy; for suspected associated entities, the association strength is judged according to the feature association relationship, and entities reaching the association threshold are classified into the same association group, finally forming a preliminary association set containing multiple association entities, realizing the surface feature-level integration of entities.

[0037] The third step is to implement semantic enhancement and mapping with the help of the semantic fusion layer. The domain knowledge rules stored in the semantic fusion layer are used to perform deep semantic processing on the preliminary association set output by the feature association layer. Domain knowledge rules are used to verify the rationality of entity associations in the preliminary association set and to correct any unreasonable associations. Furthermore, semantic interpretations are added to entity associations based on semantic-level association relationships, upgrading the preliminary association set from feature-level aggregation to semantic-level fusion. Ultimately, an initial semantic mapping result containing semantic association information is obtained, providing a deep semantic basis for subsequent precise fusion.

[0038] In some embodiments, the preset fusion rules include a basic rule set and scenario adaptation parameters; wherein the basic rule set includes data priority rules, attribute fusion rules, and association strength calculation rules; the scenario adaptation parameters include rule triggering thresholds, feature weight coefficients, and scenario type identifiers, which are used to dynamically adjust the execution logic of the basic rule set according to specific application scenarios; based on the preset fusion rules and the feature association relationships stored in the feature association layer, feature-level association aggregation is performed on the marked homologous entities and suspected associated entities to obtain a preliminary association set, including: For the marked homologous entities, according to their attribute types, the corresponding attribute fusion rules in the basic rule set are called to fuse the multi-source attribute data and generate a unified entity attribute description; For the marked suspected related entities, the association strength between the entities is calculated based on the association strength calculation rules in the basic rule set and the association relationships stored in the feature association layer; According to the association threshold set in the scene adaptation parameters, suspected associated entities whose association strength meets the threshold condition are determined to be associated, and their feature information is aggregated to form a preliminary association set.

[0039] In the embodiment of the present invention, the preset fusion rules adopt a two-layer design of "basic rules + scenario adaptation", which not only ensures the universality of the fusion logic but also adapts to the differentiated needs of different business scenarios: Basic rule set: As the core framework of the rule system, it covers three key rules and provides general logical support for data association and fusion. Among them, the data priority rule is used to define the priority of data from different data sources (for example, core database data has higher priority than third-party interface data), avoiding disordered selection when multi-source data conflicts; the attribute fusion rule specifies the fusion method of attribute data for different entity attribute types (such as deduplication and merging of text attributes, taking the mean / maximum value of numerical attributes, and taking the earliest / latest value of time attributes); the association strength calculation rule provides a method to quantify the degree of association between entities. By integrating indicators such as entity attribute similarity and feature vector matching, it outputs comparable association strength values, providing a basis for determining whether entities are associated.

[0040] Scene adaptation parameters: As a dynamic adjustment component of the rule system, it is used to flexibly optimize the execution logic of the basic rule set according to the characteristics of the specific application scene. Among them, the rule triggering threshold defines the critical condition for the rule to take effect (such as only when the data health score is higher than a certain threshold, the attribute fusion rule is executed); the feature weight coefficient adjusts the contribution proportion of different features in the correlation strength calculation; the scene type identifier is directly associated with the scene-specific rule configuration, ensuring that the rule system accurately matches the scene requirements.

[0041] For the marked homologous entities and suspected associated entities, through type processing and step-by-step calculation, a preliminary association set is finally generated, and the specific process is divided into three core links: 1. Attribute fusion processing of homologous entities Homologous entities refer to entities that are essentially the same but come from different data sources (such as user entities of the same user in system A and system B), and the core of their association and aggregation is to eliminate attribute redundancy and generate a unified description: First, identify the attribute type of each marked homologous entity; then, call the corresponding attribute fusion rule from the basic rule set according to the attribute type. For example, for the text attribute "user name", if there are slight differences in multi-source data, call the text deduplication and merging rule to retain the standard expression; for the numerical attribute "account balance", if there are deviations in multi-source data, call the numerical attribute mean value rule to calculate the unified balance value; for the time attribute "last login time", call the time attribute latest value rule to determine the unique login time; finally, integrate all the fused attributes to form a unified attribute description of the homologous entity, avoiding information confusion caused by scattered multi-source attributes.

[0042] 2. Association strength calculation of suspected associated entities Suspected associated entities refer to entities that have potential association clues but have not been explicitly confirmed to be associated, and the core of their association and aggregation is to quantify the association strength to determine whether they are truly associated: First, call the association strength calculation rule from the basic rule set to define the calculation dimensions and formula of the association strength (such as association strength = attribute similarity × 40% + feature vector matching degree × 60%); second, extract the association relationship data of the suspected associated entity pair from the feature association layer; then, substitute the association relationship data into the association strength calculation rule, combined with the adjusted feature weight coefficient in the scene adaptation parameters, calculate the specific association strength value of the entity pair, and convert "suspected association" into a quantifiable numerical index.

[0043] 3. Association determination and aggregation based on association threshold Through the association strength threshold, effective associations are screened to complete the construction of the preliminary association set: First, the correlation threshold value set for the current scene in the scene adaptation parameter is read; then, the correlation strength value of the suspected correlation entity is compared with the threshold value. If the correlation strength is greater than or equal to the threshold value, it is determined that the entity pair has an effective correlation; if the correlation strength is less than the threshold value, it is determined that the correlation is insufficient, and the entity pair is temporarily excluded from the correlation range; finally, for the entity pair determined to have an effective correlation, the feature information of each entity is aggregated and integrated, and is supplemented to the unified attribute description of the homologous entity to form a preliminary correlation set containing a homologous entity group and an effective correlation entity pair, thereby laying a foundation for subsequent depth fusion at the semantic level.

[0044] In some embodiments, the double-dimensional semantic vector includes a data-driven semantic vector and a graph-driven semantic vector; the labeled data stream is input into an AI correlation fusion model to obtain a fusion result, including: The data-driven semantic vector of the entity extracted from the preliminary correlation set is taken as the input of the first sub-tower, and the graph-driven semantic vector corresponding to the entity extracted from the hierarchical dynamic knowledge graph is taken as the input of the second sub-tower; The first sub-tower and the second sub-tower respectively perform feature coding on the respective input vectors to output corresponding semantic feature vectors; The semantic feature vectors output by the first sub-tower and the semantic feature vectors output by the second sub-tower are subjected to similarity calculation to generate an initial semantic mapping result; Based on a preset consistency adjudication rule, the initial semantic mapping result is checked and corrected to identify and solve semantic conflicts; According to the adjudication result, a final semantic mapping output is generated, and the conflict cases and / or artificial correction results generated in the adjudication process are used to iteratively optimize the double-tower sub-model and / or the consistency adjudication rule.

[0045] In some embodiments, based on a preset consistency adjudication rule, the initial semantic mapping result is checked and corrected to identify and solve semantic conflicts, including: The semantic similarity score is compared with a first threshold value, if higher than the first threshold value, it is directly determined that the semantic mapping is successful, and the output result is output; if lower than a second threshold value, it is directly determined that the semantic mapping fails, and the mapping is rejected; if between the first threshold value and the second threshold value, a secondary adjudication is triggered; wherein the first threshold value is higher than the second threshold value; For the entity triggering the secondary adjudication, the correlation relationship of the entity in the respective data source is obtained by querying the hierarchical dynamic knowledge graph, and the consistency score of the correlation relationship is calculated; the semantic similarity score and the relationship consistency score are weighted and summed to obtain a comprehensive adjudication score; if the comprehensive adjudication score is higher than a third threshold value, it is determined that the semantic mapping is successful, otherwise a tertiary adjudication is triggered; For the entity triggering the third-level adjudication, a preset logic rule base is used for conflict detection; if a logical conflict is detected, it is determined that semantic mapping fails; if no logical conflict is detected, the comprehensive adjudication score is corrected in combination with a dynamic quality label, and whether the mapping is successful is determined according to the final score after correction.

[0046] In the embodiment of the application, the process of the labeled data stream inputting the AI association fusion model to generate a fusion result takes a double-tower sub-model as the core, combines consistency adjudication and iterative optimization mechanism, realizes accurate conversion from preliminary association to final fusion result, and specifically includes five progressive links: 1. Double-tower sub-model input vector extraction.

[0047] In order to realize bidirectional semantic alignment of data and knowledge graph, two types of core vectors are extracted as double-tower inputs: First sub-tower input: Extract the "data-driven semantic vector" of the entity from the preliminary association set. The vector is generated based on the two-dimensional semantic vector (general + domain) of the entity itself, directly reflects the semantic features learned by the entity from the original data, and does not contain external association information of the knowledge graph; Second sub-tower input: Extract the "graph-driven semantic vector" of the entity from the hierarchical dynamic knowledge graph. The vector integrates the hierarchical association relationships related to the entity in the knowledge graph (such as attribute association in the basic data layer, feature association in the feature association layer, and semantic association in the semantic fusion layer), and reflects the semantic positioning of the entity in the global knowledge network.

[0048] 2. Feature encoding of double-tower sub-model.

[0049] Through the double-tower sub-model, the two types of input vectors are subjected to deep feature extraction, and the original semantic vector is converted into a more discriminative semantic feature vector: The first sub-tower focuses on the semantic coding of data itself, and extracts the local detail features (such as the exclusive attributes of the entity and the short-term feature changes) of the entity from the data-driven semantic vector through convolutional neural network and other structures; The second sub-tower focuses on the semantic coding of knowledge graph association, and extracts the global association features (such as the hierarchical relationship between the entity and other nodes and the semantic positioning under the domain rule constraint) of the entity from the graph-driven semantic vector through the Transformer encoder and other structures; Finally, the two sub-towers respectively output the encoded "data-side semantic feature vector" and "graph-side semantic feature vector", providing a high-quality feature basis for subsequent similarity calculation.

[0050] 3. Initial semantic mapping result generation.

[0051] By quantifying the association degree of the features output by the two sub-towers, an initial semantic mapping result is generated: The similarity calculation (such as cosine similarity) is performed on the "data side semantic feature vector" output by the first sub-tower and the "graph side semantic feature vector" output by the second sub-tower, to obtain a semantic similarity score. The higher the score, the more consistent the semantics learned from the data itself and the semantics learned from the knowledge graph; Based on the similarity score, an initial semantic mapping result is generated to preliminarily determine whether the semantic association between entities is valid.

[0052] 4. Verification and correction based on consistency adjudication rules.

[0053] To solve the semantic conflicts (such as ambiguous similarity scores, inconsistent data and graph semantics) that may exist in the initial mapping result, a multi-level verification and correction is required through pre-set consistency adjudication rules: The rationality of the initial semantic mapping result is verified in combination with the association relationship of the hierarchical dynamic knowledge graph, the logic rule library and the dynamic quality label; The conflicting results (such as moderate similarity scores and contradictory association relationships) are corrected to determine whether the semantic mapping is valid, ensuring the accuracy of the results.

[0054] 5. Final result generation and model iteration optimization.

[0055] Based on the consistency adjudication result, the final semantic mapping result is output as the core component of the fusion result; At the same time, the semantic conflict cases (such as misjudged mapping relationships) and manual correction results generated in the adjudication process are collected as training data and fed back to the system: The double-tower sub-model is iteratively trained to optimize the feature encoding capability and improve the accuracy of semantic similarity calculation; The threshold and weight parameters of the consistency adjudication rules are adjusted to enhance the adaptability of the rules to complex scenarios, realizing the continuous evolution of the model and the rules.

[0056] In the embodiments of the present application, the verification and correction based on the pre-set consistency adjudication rules adopt a "three-level adjudication" mechanism, which accurately identifies and solves semantic conflicts through multi-level screening and verification, ensures the reliability of the semantic mapping result, and the specific process is as follows: 1. First-level adjudication: preliminary screening based on semantic similarity score.

[0057] Two key thresholds (the first threshold > the second threshold) are set to preliminarily determine the semantic similarity score of the initial semantic mapping result: If the similarity score is higher than the first threshold: it indicates that the data side and the graph side semantics are highly consistent, and there is no obvious conflict, so it is directly determined as "semantic mapping success" and the result is output; If the similarity score is below the second threshold: it indicates that the semantic difference between the data side and the graph side is significant, and the conflict risk is extremely high. Directly determine as "semantic mapping failure" and reject the mapping; If the similarity score is between the first threshold and the second threshold: it indicates that there is uncertainty in the semantic association (such as partial feature matching but overall consistency is insufficient), and trigger secondary arbitration for further verification.

[0058] 2. Secondary arbitration: comprehensive determination combined with association relationship consistency.

[0059] For the entities in the primary arbitration "to be verified", introduce the association relationship data in the knowledge graph for deep verification: Association relationship consistency score calculation: query the hierarchical dynamic knowledge graph, extract the association relationship of the entity in each data source (such as "transaction association" of entity A and entity B in data source 1, "browsing association" in data source 2), and calculate the "association relationship consistency score" by comparing the coincidence degree and logical consistency of multi-source association relationship; Comprehensive arbitration score generation: the "semantic similarity score" in the primary arbitration and the "association relationship consistency score" in the secondary arbitration are weighted and summed according to the preset weight (such as 50% each) to obtain the comprehensive arbitration score; If the comprehensive arbitration score is higher than the third threshold: it indicates that after combining the association relationship, the reliability of the semantic association meets the standard, and it is determined as "semantic mapping success"; If the comprehensive arbitration score is lower than the third threshold: it indicates that there is still a significant conflict risk, and trigger tertiary arbitration.

[0060] 3. Tertiary arbitration: final determination based on logical rules and quality labels.

[0061] For the entities still in doubt in the secondary arbitration, introduce logical rules and data quality information for final verification: Logical conflict detection: call the preset logical rule library (such as common sense rules and business constraint rules in the field), and perform logical compliance detection on entity association; if obvious logical conflict is detected (such as "entity A and entity B association violates the rule of 'one account corresponds to one user'"), directly determine as "semantic mapping failure"; Quality label correction and final determination: if no logical conflict is detected, combine the dynamic quality label of the entity (such as health score and association potential score) to modify the comprehensive arbitration score (such as increasing the weight of the association relationship of the entity with high health score); according to the final score after modification, if it is higher than the third threshold, it is determined as "semantic mapping success", otherwise it is determined as "semantic mapping failure".

[0062] Through the layer-by-layer screening and verification of three-level adjudication, high-credibility semantic mapping can be quickly confirmed, and conflicting relationships can be accurately identified and excluded, and finally reliable semantic mapping results can be output, thereby providing core support for the accurate fusion of massive data.

[0063] In some embodiments, the AI correlation fusion model is also provided with a dynamic optimization and scheduling layer, and the method further comprises: Through the dynamic optimization and scheduling layer, the running state and output result of the correlation fusion and semantic mapping layer are monitored in real time to collect performance data; Based on the preset performance indicators and resource constraints, the performance data are analyzed to identify performance bottlenecks or optimization opportunities, and optimization instructions are generated accordingly; The optimization instructions are issued to the correlation fusion and semantic mapping layer to dynamically adjust and optimize the calculation tasks, model parameters and fusion rules therein.

[0064] In some embodiments, based on the preset performance indicators and resource constraints, the performance data are analyzed to identify performance bottlenecks or optimization opportunities, and optimization instructions are generated accordingly, comprising: Based on the performance data, the key bottlenecks affecting performance are identified through a preset optimization algorithm, and optimization instructions are generated; wherein the optimization instructions include: adjusting the network weight of the double-tower sub-model, updating the threshold parameter in the consistency adjudication rule, or performing entity or relationship addition, deletion or modification operation on the hierarchical dynamic knowledge graph.

[0065] In the embodiments of the present application, the dynamic optimization and scheduling layer serves as the "intelligent control center" of the AI correlation fusion model, and through real-time monitoring, intelligent analysis and dynamic adjustment, it ensures that the model always maintains efficient and stable operation in the process of processing massive data, and the core process thereof includes three key links: 1. Real-time monitoring and performance data collection.

[0066] The dynamic optimization and scheduling layer continuously tracks the full-link running state of the correlation fusion and semantic mapping layer, and focuses on collecting two types of core data: Running state data: including the processing time of each module (such as the feature encoding time of the double-tower sub-model and the consistency adjudication time), resource occupancy rate (such as CPU / GPU utilization rate and memory occupancy), task queue length (such as the number of correlation sets to be processed), etc., reflecting the real-time running load of the system; Output result data: including the correlation accuracy of the fusion result, the success rate of conflict resolution, the semantic mapping precision, etc., reflecting the core performance indicators of the model.

[0067] Through real-time collection and summarization of these data, comprehensive basis is provided for subsequent analysis and optimization.

[0068] 2. Performance analysis and optimization instruction generation.

[0069] Based on the preset performance indicators (such as association accuracy threshold, processing delay upper limit) and resource constraints (such as maximum computing power allocation, memory limit), the collected performance data is analyzed in multiple dimensions: If it is found that the processing time consumption of a certain module exceeds the threshold (such as the long coding time of the double-tower sub-model), it is determined to be a "performance bottleneck"; If it is found that the resource utilization is too low (such as high GPU idle rate) or part of the rule adaptability is insufficient (such as the consistency arbitration threshold leading to the increase of misjudgment rate), it is identified as an "optimization opportunity".

[0070] In combination with the preset optimization algorithm (such as resource scheduling algorithm based on reinforcement learning, parameter tuning algorithm based on statistical analysis), precise optimization instructions are generated for the bottleneck or opportunity, and the objects and specific methods that need to be adjusted are clearly defined.

[0071] 3. Dynamic adjustment and full-link optimization.

[0072] The generated optimization instructions are issued to the associated fusion and semantic mapping layer to trigger targeted adjustment: Scheduling of computing tasks (such as assigning high-priority association aggregation tasks to idle computing power nodes); Model parameter correction (such as adjusting the feature weight of the double-tower sub-model); Update the fusion rules (such as optimizing the threshold of consistency arbitration).

[0073] Through real-time feedback and adjustment, the model adapts to data distribution changes, resource fluctuations and scene requirements, and always maintains the optimal operating state.

[0074] In the embodiments of the present application, based on the collected performance data, the core factors affecting system performance are located through preset optimization algorithms (such as decision tree analysis, regression analysis): if the semantic mapping accuracy decreases, it may be due to the insufficient feature encoding ability of the double-tower sub-model; if the conflict resolution efficiency is low, it may be due to the unreasonable threshold setting of the consistency arbitration rule; if the association aggregation time is too long, it may be due to the redundancy of the entity / relationship storage structure of the hierarchical dynamic knowledge graph.

[0075] In the embodiments of the present application, three types of core optimization instructions are generated for the identified bottlenecks: Model parameter adjustment: such as adjusting the network weight of the double-tower sub-model (enhancing the ability to capture key semantic features), optimizing the activation function parameters (improving the discrimination of feature encoding), to improve the accuracy of semantic similarity calculation; Rule parameter update: modify the first / second / third threshold in the consistency adjudication rule (loosen or tighten the judgment standard according to the scene requirement), adjust the feature weight of the correlation strength calculation (adapt to the correlation focus of different scenes), in order to improve the accuracy of conflict resolution; Knowledge graph maintenance: add, delete or modify entities or relationships in the hierarchical dynamic knowledge graph (such as deleting redundant associations, supplementing high-frequency core relationships, and correcting incorrect semantic associations), in order to reduce the time consumption of association query and improve the support efficiency of knowledge graph.

[0076] Through the above process, the dynamic optimization and scheduling layer realizes the transformation of the model from "passive running" to "active evolution", ensuring that the AI association fusion model always maintains efficient, accurate and stable fusion capability when facing massive heterogeneous data and complex scenes, providing continuous and reliable technical support for data value mining.

[0077] The embodiment of the application provides an electronic device, including a memory and a processor, the memory stores a computer program, and the processor realizes the method shown in the above embodiment when executing the computer program.

[0078] The embodiment of the application provides a computer readable storage medium, the computer readable storage medium stores a computer program, and the computer program is executed by the processor to realize the method shown in the above embodiment.

[0079] In the above embodiments, the description of each embodiment has its own emphasis, and the parts not described or recorded in a certain embodiment can be referred to the related description of other embodiments. If there is no special description and logical conflict, the terms and / or descriptions of different embodiments are consistent and can be mutually referred to, and the technical features in different embodiments can be combined to form new embodiments according to their inherent logical relationship.

[0080] The above-described embodiments are only used to illustrate the technical solutions of the application, rather than limit them; although the application has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that the technical solutions recorded in the foregoing embodiments can be modified, or some technical features can be replaced by equivalents; and these modifications or replacements do not make the essence of the corresponding technical solutions deviate from the spirit and scope of the technical solutions of the embodiments of the application, and should be included in the protection scope of the application.

Claims

1. A method for associating and fusing massive data, characterized in that: include: Receiving massive amounts of data from multiple heterogeneous data sources and preprocessing the data; Integrate the preprocessed data into annotated data streams; The annotated data stream includes a two-dimensional semantic vector and a dynamic quality label; Inputting the annotated data stream into the AI ​​association fusion model to obtain a fusion result; Among them, the AI ​​association fusion model is provided with an association fusion and semantic mapping layer; the association fusion and semantic mapping layer is used to perform association fusion according to the hierarchical dynamic knowledge graph and preset fusion rules, and perform semantic mapping according to the dual-tower sub-model and consistency judgment rules.

2. The method for associating and fusing massive data according to claim 1, characterized in that: The annotated data stream is input into the AI ​​association fusion model to obtain the fusion results, including: Performing dimensional normalization on the two-dimensional semantic vectors in the annotated data stream so that the vector dimensions match the preset input dimensions of the model; and performing weight initialization on the dynamic quality labels based on the credibility level of the data source to obtain an adapted data stream; Transmitting the preset fusion rules and the adapted data stream to the association fusion and semantic mapping layer to obtain a preliminary association set and an initial semantic mapping result to determine a fusion result; The association fusion and semantic mapping layer is based on the hierarchical structure of the hierarchical dynamic knowledge graph, extracts the association features corresponding to the adapted data stream from each layer, and combines the preset fusion rules to perform preliminary association aggregation on the homologous data and related data in the adapted data stream to generate a preliminary association set; The association fusion and semantic mapping layer calls a dual-tower sub-model to perform semantic mapping processing on the preliminary association set; the first sub-tower is used to receive the two-dimensional semantic vector in the preliminary association set and extract local semantic features through a convolutional neural network, and the second sub-tower is used to receive the semantic representation vector of the corresponding association node in the hierarchical dynamic knowledge graph and extract global semantic features through a Transformer encoder; the initial semantic mapping result is determined based on the cosine similarity of the output features of the two sub-towers.

3. The method for associating and fusing massive data according to claim 2, characterized in that: The hierarchical dynamic knowledge graph includes three layers: a basic data layer, a feature association layer, and a semantic fusion layer. Each layer is dynamically connected through association edges. The basic data layer stores the original entities and attribute information of each data source, the feature association layer stores the feature-level association relationships between entities, and the semantic fusion layer stores the semantic-level association relationships and domain knowledge rules between entities. The adapted data stream is transmitted to the association fusion and semantic mapping layer to obtain a preliminary association set and initial semantic mapping results, including: Extract the reference entity information corresponding to each entity in the adapted data stream from the basic data layer, use the entity recognition algorithm to identify the entity type in the adapted data stream, match the entity to be fused with the reference entity through the entity alignment algorithm, and mark homologous entities and suspected related entities; Based on the preset fusion rules and the feature association relationships stored in the feature association layer, the marked homologous entities and suspected related entities are aggregated at the feature level to obtain a preliminary association set; The domain knowledge rules in the semantic fusion layer are called to semantically enhance the preliminary association set output by the feature association layer to obtain the initial semantic mapping result.

4. The method for associating and fusing massive data according to claim 3, characterized in that: The preset fusion rules include a basic rule set and scenario adaptation parameters; wherein the basic rule set includes data priority rules, attribute fusion rules, and association strength calculation rules; the scenario adaptation parameters include rule triggering thresholds, feature weight coefficients, and scenario type identifiers, which are used to dynamically adjust the execution logic of the basic rule set according to specific application scenarios; based on the preset fusion rules and the feature association relationships stored in the feature association layer, feature-level association aggregation is performed on the marked homologous entities and suspected associated entities to obtain a preliminary association set, including: For the marked homologous entities, according to their attribute types, the corresponding attribute fusion rules in the basic rule set are called to fuse the multi-source attribute data to generate a unified entity attribute description; For the marked suspected related entities, calculating the association strength between the entities according to the association strength calculation rules in the basic rule set and the association relationships stored in the feature association layer; According to the association threshold set in the scenario adaptation parameter, the suspected associated entities whose association strength meets the threshold condition are determined to be associated, and their feature information is aggregated to form the preliminary association set.

5. The method for associating and fusing massive data according to claim 2, characterized in that: The dual-dimensional semantic vector includes a data-driven semantic vector and a graph-driven semantic vector; The annotated data stream is input into the AI ​​association fusion model to obtain the fusion results, including: Extract the data-driven semantic vector of the entity from the preliminary association set as the input of the first sub-tower, and extract the graph-driven semantic vector corresponding to the entity from the hierarchical dynamic knowledge graph as the input of the second sub-tower; Perform feature encoding on the respective input vectors through the first sub-tower and the second sub-tower, and output corresponding semantic feature vectors; Calculate the similarity between the semantic feature vector output by the first sub-tower and the semantic feature vector output by the second sub-tower to generate an initial semantic mapping result; Based on preset consistency adjudication rules, the initial semantic mapping result is verified and corrected to identify and resolve semantic conflicts; The final semantic mapping output is generated according to the adjudication result, and the conflict cases and / or manual correction results generated during the adjudication process are used to iteratively optimize the dual-tower sub-model and / or consistency adjudication rules.

6. The method for associating and fusing massive data according to claim 5, characterized in that: Based on preset consistency adjudication rules, the initial semantic mapping results are verified and corrected to identify and resolve semantic conflicts, including: Compare the semantic similarity score with a first threshold; if it is higher than the first threshold, it is directly determined that the semantic mapping is successful and the result is output; if it is lower than the second threshold, it is directly determined that the semantic mapping fails and the mapping is rejected; if it is between the first threshold and the second threshold, a secondary ruling is triggered; wherein the first threshold is higher than the second threshold; For entities that trigger a second-level adjudication, the hierarchical dynamic knowledge graph is queried to obtain their associations in their respective data sources, and the consistency scores of the associations are calculated; the semantic similarity score and the consistency score are weighted and summed to obtain a comprehensive adjudication score; if the comprehensive adjudication score is higher than the third threshold, the semantic mapping is determined to be successful; otherwise, a third-level adjudication is triggered; For entities that trigger the third-level adjudication, conflict detection is performed using a preset logic rule library; if a logical conflict is detected, it is determined that the semantic mapping has failed; if no logical conflict is detected, the comprehensive adjudication score is corrected in combination with the dynamic quality label, and whether the mapping is successful is determined based on the corrected final score.

7. The method for associating and fusing massive data according to claim 2, characterized in that: The AI ​​association fusion model is further provided with a dynamic optimization and scheduling layer, and the method further comprises: Through the dynamic optimization and scheduling layer, the operating status and output results of the association fusion and semantic mapping layer are monitored in real time to collect performance data; Analyze the performance data based on preset performance indicators and resource constraints, identify performance bottlenecks or optimization opportunities, and generate optimization instructions accordingly; The optimization instructions are sent to the association fusion and semantic mapping layer to dynamically adjust and optimize the computing tasks, model parameters and fusion rules therein.

8. The method for associating and fusing massive data according to claim 7, characterized in that: Based on preset performance indicators and resource constraints, the performance data is analyzed to identify performance bottlenecks or optimization opportunities, and optimization instructions are generated accordingly, including: Based on the performance data, a preset optimization algorithm is used to identify key bottlenecks affecting performance and generate optimization instructions; wherein, the optimization instructions include: adjusting the network weights of the dual-tower sub-model, updating the threshold parameters in the consistency judgment rules, or adding, deleting, and modifying entities or relationships in the hierarchical dynamic knowledge graph.

9. An electronic device, characterized in that: The method comprises a memory and a processor, wherein the memory stores a computer program, and when the processor executes the computer program, the method according to any one of claims 1 to 8 is implemented.

10. A computer-readable storage medium, characterized in that The computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, the method according to any one of claims 1 to 8 is implemented.

Citation Information

Patent Citations

  • Multi-source data fusion method and system based on cloud computing

    CN119989267A

  • Real-time health risk prediction method and system based on dynamic knowledge graph

    CN120280136A

  • Multi-source heterogeneous data fusion method and system based on cloud computing

    CN120524422A

  • Knowledge fusion method and apparatus, computer device, and storage medium

    WO2020143184A1

Cited By

  • Agricultural heterogeneous data fusion verification method based on multi-dimensional semantic alignment operator

    CN121030275A

  • Method for verifying fusion of agricultural heterogeneous data based on multi-dimensional semantic alignment operator

    CN121030275B

  • Multi-level data intelligent statistical analysis system based on machine learning

    CN121502159A

  • Cognitive intelligent communication signal identification method and system

    CN121980311A

  • A Multi-Source Heterogeneous Data Fusion Method and System Based on Semantic Understanding

    CN122673987A