Cross-domain data association mining method based on knowledge graph
Patent Information
- Application Number
- CN202610898851.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2026-06-22
- Publication Date
- 2026-08-18
AI Technical Summary
然而,现有的图对齐技术多依赖于实体名称或少量属性的硬匹配,缺乏对深层语义的自动理解与对齐能力,导致融合后的知识图谱存在大量的语义冲突与冗余实体
本发明提供的一种基于知识图谱的跨域数据关联挖掘方法,通过设置跨域语义对齐单元与动态知识图谱构建单元,并建立二者之间的迭代反馈连接,实现了从异构数据到统一知识表示的一体化自动构建,解决了传统方法中语义对齐误差累积导致图谱质量低下的问题,显著提升了跨域数据融合的准确性与效率。本发明通过设计图增强关联特征提取单元与跨域关联强度评估单元,创新性地引入关系图卷积网络与多通道融合打分函数,特别是通过低频关系增强因子的设计,有效避免了关联挖掘对高频关系的偏好,能够发现隐含的、深层次的跨域关联模式,并借助元路径实例提供了内生的可解释性。此外,本发明构建了由关联模式解释与可视化单元至其他单元的反馈优化闭环,使得关联挖掘结果能够反向影响语义对齐的权重与特征提取网络的参数,赋予了系统持续学习与自适应演化的能力,从而在面对动态变化的数据分布与业务场景时,能够长期保持挖掘结果的准确性与鲁棒性。
Smart Images

Figure CN122595249A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of big data, and in particular to a method for cross-domain data association mining based on knowledge graphs. Background Technology
[0002] With the rapid development of big data technology, massive amounts of highly heterogeneous data have accumulated across different business domains (such as finance, healthcare, e-commerce, and social networks). This data is often scattered across independently built systems, forming a typical "data silo" phenomenon. However, complex objects or events in the real world often span multiple data domains. For example, a user's credit risk may be simultaneously linked to their borrowing records in the financial domain, their consumption behavior in the e-commerce domain, and their interaction network in the social domain. Therefore, uncovering the implicit and deep-seated relationships between these cross-domain data is crucial for applications such as precise recommendation, risk control, and knowledge discovery. Knowledge graphs, as an effective knowledge representation and reasoning tool, can structure data through entities and relationships, and have become the mainstream technical approach for handling complex relationship analysis.
[0003] Currently, existing methods for data association mining mainly fall into two categories. The first category is traditional feature matching methods based on statistics or rules. These methods typically require pre-defining field mapping relationships across data domains or manually constructing features. When the data source patterns change frequently or there are differences in domain semantics, the maintenance cost of these mapping rules is extremely high, and it is difficult to discover implicit associations beyond the predefined rules. The second category is analysis methods based on static knowledge graphs. These methods first construct subgraphs from each data source and then perform simple graph merging or alignment. However, existing graph alignment techniques mostly rely on hard matching of entity names or a few attributes, lacking the ability to automatically understand and align deep semantics, resulting in a large number of semantic conflicts and redundant entities in the merged knowledge graph. More importantly, once an existing knowledge graph is constructed, its structure and relationships are usually static, unable to evolve and logically correct itself based on subsequent association mining results or newly inflowing data. Meanwhile, existing association strength assessment methods mostly use single graph structure features (such as the number of paths and co-occurrence frequency), ignoring the semantic differences between different association paths and the fusion of entity attribute features. This results in insufficient accuracy and interpretability of the discovered associations, making it difficult to adapt to the real business environment where data distribution is constantly changing.
[0004] Therefore, how to achieve automated semantic alignment between multi-source heterogeneous data domains, construct a cross-domain knowledge graph that can dynamically evolve and has deep semantic understanding capabilities, and on this basis provide a high-precision, interpretable cross-domain data association mining method with feedback optimization mechanism, has become a technical problem that urgently needs to be solved by those skilled in the art. Summary of the Invention
[0005] The purpose of this invention is to provide a cross-domain data association mining method based on knowledge graphs to solve the problems existing in the prior art.
[0006] To achieve the above objectives, the present invention provides the following solution: This invention provides a method for cross-domain data association mining based on knowledge graphs, the method comprising: The acquired business data from the multi-source heterogeneous data domain is standardized and instantiated to generate a standardized data stream; Semantic alignment processing is performed on cross-domain entities and attributes in the standardized data stream to generate a semantic alignment annotation set; Based on the standardized data stream and the semantically aligned annotation set, a dynamic knowledge graph is constructed to generate a unified cross-domain knowledge graph; The cross-domain knowledge graph is subjected to graph-enhanced association feature extraction to generate an entity feature vector matrix and a set of candidate association paths; Using the entity feature vector matrix and the candidate association path set, cross-domain association strength is evaluated, and an association ranking list is generated. The associated sorting list is interactively corrected based on the received feedback signals, and the correction results are used to optimize the extraction process of the entity feature vector matrix.
[0007] Preferably, the generation of the semantic alignment annotation set includes: A cross-domain encoder based on a pre-trained language model is used to calculate the semantic vector similarity between entity names and attribute names in different data sources; By combining edit distance with domain dictionary matching, candidate alignment pairs are generated; The candidate alignment pairs are filtered by a preset similarity threshold to obtain a semantic alignment annotation set.
[0008] Preferably, the generation of a unified cross-domain knowledge graph includes: Using the semantic alignment annotation set as a bridge, triples from various data sources are mapped to a unified namespace; the context neighbor intersection algorithm is used to disambiguate and merge the mapped entities. A knowledge graph embedding model is used to pre-complete missing relationships, and a source domain label and timestamp are attached to each triple to generate a cross-domain knowledge graph.
[0009] Preferably, the generated entity feature vector matrix and candidate association path set include: A predefined set of cross-domain metapath templates; For each entity in the cross-domain knowledge graph, sample all meta-path instances within a preset hop count range; A relational graph convolutional network is used to aggregate features of the sampling results, generate an entity feature vector matrix, and record the meta-path instances connecting entity pairs as a set of candidate association paths; The layer update rule of the relational graph convolutional network is expressed as follows: ; In the formula, Representing entities In the Layer embedding vectors, For a set of relations, To pass through relationships With entity A set of connected neighbors. The normalization constant is For relationship In the The trainable weight matrix of the layer, The weight matrix is a self-connected matrix. This is the activation function.
[0010] Preferably, generating the associated sorted list includes: Entity pairs with different source domain labels in the cross-domain knowledge graph are selected as candidate cross-domain entity pairs; For each candidate cross-domain entity pair, calculate the meta-path weighted association score and the semantic-structural fusion association score; The weighted association score of the meta-path is weighted and fused with the semantic-structure fusion association score to obtain the total association score. Based on the total association score, all candidate cross-domain entity pairs are sorted in descending order, and the top K are selected as the association sorting list.
[0011] Preferably, the calculation method of the meta-path weighted association score is expressed as follows: ; In the formula, and For candidate cross-domain entity pairs, For connection and The collection of all metapath instances, For path Attention weights For path length, The cosine similarity function is used. and The embedding vectors of adjacent entities along the path. For relationship Frequency of occurrence in the global graph This is the temperature coefficient.
[0012] Preferably, the semantic-structural fusion association score is obtained by concatenating and element-wise multiplying the embedding vectors of entity pairs and then inputting them into a multilayer feedforward network. The calculation method of the total association score is expressed as follows: ; In the formula, To integrate weights, It is a multi-layer feedforward network. and This is the final embedding vector of the entity. This represents a vector concatenation operation. This represents element-wise product.
[0013] Preferably, the method further includes: Visualize and render the highly relevant results in the association ranking list to generate a visualization report containing interactive subgraphs and path descriptions; Obtain user feedback on the manually annotated results in the visualization report, as the feedback signal.
[0014] Preferably, the feedback optimization of the extraction process of the entity feature vector matrix using the correction results includes: When the feedback signal indicates that a specific association result is incorrect, the attention weight of the meta-path that generated the association result is reduced. The reduced attention weights are used to trigger incremental training of the relational graph convolutional network to update the entity feature vector matrix.
[0015] This invention also provides a cross-domain data association mining system based on knowledge graphs, comprising: The multi-source heterogeneous data access unit is used to standardize and instantiate the business data acquired from the multi-source heterogeneous data domain, and generate a standardized data stream. A cross-domain semantic alignment unit is used to perform semantic alignment processing on cross-domain entities and attributes in the standardized data stream to generate a semantic alignment annotation set. The dynamic knowledge graph construction unit is used to construct a dynamic knowledge graph based on the standardized data stream and the semantic alignment annotation set, and generate a unified cross-domain knowledge graph. The graph-enhanced association feature extraction unit is used to extract graph-enhanced association features from the cross-domain knowledge graph and generate an entity feature vector matrix and a set of candidate association paths. The cross-domain association strength assessment unit is used to assess the cross-domain association strength using the entity feature vector matrix and the candidate association path set, and generate an association ranking list. The association pattern interpretation and visualization unit is used to interactively correct the association sorting list based on the received feedback signals, and to use the correction results to optimize the extraction process of the entity feature vector matrix.
[0016] The present invention achieves the following beneficial technical effects compared to the prior art: This invention provides a cross-domain data association mining method based on knowledge graphs. By setting up a cross-domain semantic alignment unit and a dynamic knowledge graph construction unit, and establishing an iterative feedback connection between the two, it achieves integrated automatic construction from heterogeneous data to a unified knowledge representation. This solves the problem of low graph quality caused by the accumulation of semantic alignment errors in traditional methods, significantly improving the accuracy and efficiency of cross-domain data fusion. This invention innovatively introduces a relational graph convolutional network and a multi-channel fusion scoring function by designing a graph-enhanced association feature extraction unit and a cross-domain association strength evaluation unit. In particular, the design of a low-frequency relationship enhancement factor effectively avoids the bias of association mining towards high-frequency relationships, enabling the discovery of implicit, deep-level cross-domain association patterns, and providing endogenous interpretability through meta-path instances. Furthermore, this invention constructs a feedback optimization closed loop from the association pattern interpretation and visualization unit to other units, allowing the association mining results to influence the weights of semantic alignment and the parameters of the feature extraction network, endowing the system with the ability to continuously learn and adaptively evolve. This ensures that the mining results maintain accuracy and robustness in the face of dynamically changing data distributions and business scenarios over the long term. Attached Figure Description
[0017] To more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the drawings used in the embodiments will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0018] Figure 1 The flowchart of the cross-domain data association mining method based on knowledge graph provided by the present invention is shown. Detailed Implementation
[0019] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.
[0020] The purpose of this invention is to provide a cross-domain data association mining method based on knowledge graphs. This method constructs a cross-domain semantic alignment mechanism and a dynamic knowledge graph evolution system, and combines graph neural networks and multi-channel association strength evaluation to achieve high-precision automated mining and interpretable output of implicit associations between multi-source heterogeneous data domains.
[0021] To make the above-mentioned objects, features and advantages of the present invention more apparent and understandable, the present invention will be further described in detail below with reference to the accompanying drawings and specific embodiments.
[0022] Example 1: Reference Figure 1 The present invention specifically includes the following steps.
[0023] Step 1: Standardize and instantiate the business data obtained from the multi-source heterogeneous data domain to generate a standardized data stream.
[0024] In practical applications, the system first connects to raw data from different business domains through a multi-source heterogeneous data access unit. These data sources can include structured transaction records from financial systems, semi-structured electronic medical record logs from the medical field, user behavior tracking data from e-commerce platforms, and unstructured text information from social networks. Based on pre-configured data source parsing rules, such as path expressions for JSON-formatted logs or SQL mapping templates for relational databases, the system extracts key fields from each record in batches. For each raw record, the system converts it into a unified intermediate format, such as JSONLD entries, and performs field type inference, null value filling strategies, and timestamp standardization during the conversion process. Each standardized data record is encapsulated as a standardized data unit, and all units form a standardized data stream in the order of inflow. This standardized data stream retains the domain labels and timestamps of the original data source, providing a foundation for subsequent cross-domain semantic alignment.
[0025] Step 2: Perform semantic alignment processing on cross-domain entities and attributes in the standardized data stream to generate a semantic alignment annotation set.
[0026] After obtaining the standardized data stream, the system initiates the cross-domain semantic alignment unit. The core of this unit is a cross-domain encoder based on a pre-trained language model, such as BERT or RoBERTa. For entity names or attribute names from different data sources, the system inputs them into the encoder to obtain semantic vectors of fixed dimensions. Then, the cosine similarity between each pair of cross-domain names is calculated. To improve alignment accuracy, the system simultaneously introduces an edit distance algorithm to measure string-level similarity and combines it with dictionary matching rules pre-built by domain experts. The similarity across these three dimensions is weighted and summed to obtain a comprehensive similarity score. The system iterates through all cross-domain entity and attribute name pairs to generate a candidate alignment set. Subsequently, based on a preset similarity threshold (default set to 0.75), only alignment pairs with a comprehensive similarity greater than or equal to this threshold are retained, forming the final semantic alignment annotation set. This annotation set contains entity and attribute pairs with the same or equivalent semantics identified from different data domains; for example, labeling a personal customer number in the financial domain and a user ID in the social domain as the same entity.
[0027] Step 3: Based on the standardized data stream and the semantic alignment annotation set, construct a dynamic knowledge graph to generate a unified cross-domain knowledge graph.
[0028] After obtaining the standardized data stream and semantic alignment annotation set, the dynamic knowledge graph construction unit performs a fusion operation. First, using the semantic alignment annotation set as a bridge, the original triple representations from each data source—namely, head entity, relation, and tail entity—are mapped to a unified global namespace. For example, if a customer entity in the financial domain and a user entity in the e-commerce domain are labeled as equivalent in the alignment set, they share the same entity identifier in the global graph. Next, the system performs entity disambiguation using a context neighbor intersection algorithm. For entities with similar names but different semantics, the system compares their one-hop neighbor sets in the graph. If the Jaccard similarity coefficient of the neighbor sets of two entities is greater than 0.5, they are determined to be the same entity and merged; otherwise, they are retained as different entities. To address the sparsity problem of the knowledge graph, the system employs a knowledge graph embedding model, such as TransE, to pre-complete missing relation links. Specifically, for any pair of entities, if there is a potential semantic relationship between them but no explicit relation edge, the TransE model calculates a score based on the principle that the vector of the head entity plus the relation approximates the tail entity. When the score is higher than a preset threshold, the system automatically adds the relation edge. Finally, the system adds a source domain label and a timestamp to each triple in the knowledge graph, thereby generating a unified cross-domain knowledge graph with meta-information.
[0029] Step 4: Extract graph-enhanced association features from the cross-domain knowledge graph to generate an entity feature vector matrix and a set of candidate association paths.
[0030] After the cross-domain knowledge graph is constructed, the graph-enhanced association feature extraction unit starts processing. The system first predefines a set of cross-domain meta-path templates. Each meta-path template describes the possible semantic connection patterns between entities in different domains. For example, a user purchasing goods belongs to a category, where the user belongs to the social or financial domain, and the goods and category belong to the e-commerce domain. For each entity in the knowledge graph, the system performs random walk sampling within a preset number of hops (e.g., 3 to 4 hops), collecting all instance paths that conform to the meta-path templates. Each instance path records the sequence of entities and relationships along the way. Subsequently, the system uses a relational graph convolutional network to aggregate features from the sampled path information. The layer update rule of this relational graph convolutional network is expressed by Equation 1: ; In the formula, Representing entities In the Layer embedding vectors, For a set of relations, To pass through relationships With entity A set of connected neighbors. The normalization constant is For relationship In the The trainable weight matrix of the layer, The weight matrix is a self-connected matrix. is the activation function. After passing through multiple stacked relational graph convolutional networks, each entity finally obtains a fixed-dimensional embedding representation. The embedding vectors of all entities together constitute the entity feature vector matrix. At the same time, the system saves the meta-path instances between all connected entity pairs recorded during the sampling process, forming a candidate association path set.
[0031] Step 5: Using the entity feature vector matrix and the candidate association path set, perform cross-domain association strength evaluation and generate an association ranking list.
[0032] After feature extraction is complete, the cross-domain association strength evaluation unit begins calculating the association strength between entity pairs. The system first filters all entity pairs with different source domain labels from the cross-domain knowledge graph, forming a candidate cross-domain entity pair set. For each candidate entity pair... and The system calculates scores across two dimensions. The first dimension is the meta-path weighted association score, which reflects the structural strength of the connection between two entities through different semantic paths. Its calculation method is expressed by Formula 2: ; In the formula, and For candidate cross-domain entity pairs, For connection and The collection of all metapath instances, For path Attention weights For path length, The cosine similarity function is used. and The embedding vectors of adjacent entities along the path. For relationship Frequency of occurrence in the global graph The value is the temperature coefficient. The exponential term in the denominator of the formula makes the fraction corresponding to low-frequency relationships higher, thus enhancing the contribution of rare paths.
[0033] The second dimension is the semantic structure fusion association score, which is achieved by non-linearly fusing the embedding features of entities through a multi-layer feedforward network. Specifically, the system combines the final embedding vectors of the two entities... and A vector concatenation operation is performed to obtain a 2d-dimensional vector. Simultaneously, the element-wise product of the two vectors is calculated to obtain an interaction feature vector of dimension d. The concatenated result is then concatenated with the product result again to form a 3d-dimensional input feature, which is fed into a multi-layer feedforward network with two hidden layers. The hidden layer dimensions of this network are 128 and 64 respectively, and the output layer dimension is 1, ultimately yielding a semantic structure fusion association score.
[0034] The system then weights and fuses the scores from the two dimensions to obtain entity pairs. and The total score related to the relationship is expressed by Formula 3: ; In the formula, To integrate weights, It is a multi-layer feedforward network. and This is the final embedding vector of the entity. This represents a vector concatenation operation. This represents element-wise multiplication. After calculating the total association score for all candidate cross-domain entity pairs, the system sorts them in descending order of score and selects the top K results as the final association ranking list. K can be set according to the actual application scenario, for example, 100 or 200.
[0035] Step 6: Based on the received feedback signals, interactively correct the associated sorting list, and use the correction results to optimize the extraction process of the entity feature vector matrix.
[0036] After generating the association ranking list, the system outputs the list to the association pattern interpretation and visualization unit. This unit visualizes the high-association results in the list. For example, for the top 10 entity pairs, the system extracts key meta-path instances connecting them from the cross-domain knowledge graph, generates a visualization report containing interactive subgraphs and path description text, and displays it to the user. The user can interact with the interface, such as confirming a correlation is correct or marking a correlation as incorrect. When a user marks a correlation as incorrect, the system records this feedback signal. The system further analyzes the meta-path based on the incorrect correlation and reduces the attention weight corresponding to that path. The reduction operation can use a decay factor, such as multiplying the original weight by 0.8. Subsequently, the system uses the adjusted attention weights to trigger incremental training of the relationship graph convolutional network. Incremental training only uses the entity pairs corresponding to the paths with reduced weights as negative samples to perform a small amount of gradient updates, thereby updating the entity feature vector matrix. This feedback optimization process allows the system to learn from manual annotations, gradually correct unreasonable path preferences, and improve the accuracy of subsequent mining.
[0037] Furthermore, in a preferred embodiment, the method also includes continuous self-optimization of the association ranking list. The system periodically adds high-confidence association results confirmed by the user and their corresponding meta-path instances as positive samples to the historical knowledge cache for initial weight settings during the next full training of the relational graph convolutional network. Simultaneously, for path types that have not been confirmed by users for a long time or are frequently marked as incorrect, the system automatically reduces the sampling priority of their meta-path templates, thereby achieving adaptive adjustment of the meta-path template set.
[0038] The following is a specific application example to illustrate the above method. Assume the system accesses the customer transaction data domain of a financial institution and the user interaction data domain of a social media platform. The financial institution's data contains customer entities with attributes such as annual income and loan records; the social media platform's data contains user entities with attributes such as posting frequency and number of friends. In step two, the system, through semantic alignment, finds that the entity similarity between the financial institution's customer Zhang San and the social media platform's user ID user_123 reaches 0.92, exceeding the threshold of 0.75, and therefore marks them as an aligned entity pair. In the knowledge graph constructed in step three, these two entities are merged into the same global entity node. In step four, the system samples a meta-path instance: Zhang San's loan default bank risk control report mentions negative comments from user user_123. In step five, when the system calculates the total association score between Zhang San and user_123, this path contributes a high S_path value, and the MLP part also gives a high score due to the similarity of the two entities' embedding vectors. Ultimately, Zhang San and user_123 enter the top of the association ranking list. In the visualization report of step six, the system presents this association to the risk control analyst and provides an explanation of the path. If the analyst confirms the association is correct, the system feeds it back as a positive sample to the feature extraction unit, enhancing the attention weight of this type of path. Conversely, if the analyst finds an association to be incorrect, such as misaligning entities of different users, the system will reduce the weight of the corresponding path and trigger incremental training.
[0039] In summary, the cross-domain data association mining method based on knowledge graphs provided by this invention can effectively solve the problem of implicit association mining between multi-source heterogeneous data domains through iterative fusion of semantic alignment and graph construction, feature extraction based on relational graph convolutional networks, multi-channel association strength evaluation, and feedback optimization mechanism. It has the advantages of high accuracy, high interpretability, and adaptive evolution.
[0040] This invention has illustrated its principles and implementation methods using specific examples. The descriptions of these embodiments are merely illustrative of the method and its core ideas; furthermore, those skilled in the art will recognize that modifications may be made to the specific implementation methods and application scope based on the principles of this invention. Therefore, the content of this specification should not be construed as limiting the invention.
Claims
1. A cross-domain data association mining method based on knowledge graphs, characterized in that, The method includes: The acquired business data from the multi-source heterogeneous data domain is standardized and instantiated to generate a standardized data stream; Semantic alignment processing is performed on cross-domain entities and attributes in the standardized data stream to generate a semantic alignment annotation set; Based on the standardized data stream and the semantically aligned annotation set, a dynamic knowledge graph is constructed to generate a unified cross-domain knowledge graph; The cross-domain knowledge graph is subjected to graph-enhanced association feature extraction to generate an entity feature vector matrix and a set of candidate association paths; Using the entity feature vector matrix and the candidate association path set, cross-domain association strength is evaluated, and an association ranking list is generated. The associated sorting list is interactively corrected based on the received feedback signals, and the correction results are used to optimize the extraction process of the entity feature vector matrix.
2. The method for cross-domain data association mining based on knowledge graphs according to claim 1, characterized in that, The generated semantic alignment annotation set includes: A cross-domain encoder based on a pre-trained language model is used to calculate the semantic vector similarity between entity names and attribute names in different data sources; By combining edit distance with domain dictionary matching, candidate alignment pairs are generated; The candidate alignment pairs are filtered by a preset similarity threshold to obtain a semantic alignment annotation set.
3. The method for cross-domain data association mining based on knowledge graphs according to claim 1, characterized in that, The generation of a unified cross-domain knowledge graph includes: Using the semantic alignment annotation set as a bridge, triples from various data sources are mapped to a unified namespace; the context neighbor intersection algorithm is used to disambiguate and merge the mapped entities. A knowledge graph embedding model is used to pre-complete missing relationships, and a source domain label and timestamp are attached to each triple to generate a cross-domain knowledge graph.
4. The method for cross-domain data association mining based on knowledge graphs according to claim 1, characterized in that, The generated entity feature vector matrix and candidate association path set include: A predefined set of cross-domain metapath templates; For each entity in the cross-domain knowledge graph, sample all meta-path instances within a preset hop count range; A relational graph convolutional network is used to aggregate features of the sampling results, generate an entity feature vector matrix, and record the meta-path instances connecting entity pairs as a set of candidate association paths; The layer update rule of the relational graph convolutional network is expressed as follows: ; In the formula, Representing entities In the Layer embedding vectors, For a set of relations, To pass through relationships With entity Connected neighbor set, The normalization constant is For relationship In the The trainable weight matrix of the layer, The weight matrix is a self-connected matrix. This is the activation function.
5. The method for cross-domain data association mining based on knowledge graphs according to claim 4, characterized in that, The generation of the associated sorted list includes: Entity pairs with different source domain labels in the cross-domain knowledge graph are selected as candidate cross-domain entity pairs; For each candidate cross-domain entity pair, calculate the meta-path weighted association score and the semantic-structural fusion association score; The weighted association score of the meta-path is weighted and fused with the semantic-structure fusion association score to obtain the total association score. Based on the total association score, all candidate cross-domain entity pairs are sorted in descending order, and the top K are selected as the association sorting list.
6. The method for cross-domain data association mining based on knowledge graphs according to claim 5, characterized in that, The calculation method for the meta-path weighted association score is expressed as follows: ; In the formula, and For candidate cross-domain entity pairs, For connection and The collection of all metapath instances, For path Attention weights For path length, The cosine similarity function is used. and The embedding vectors of adjacent entities along the path. For relationship Frequency of occurrence in the global graph This is the temperature coefficient.
7. The method for cross-domain data association mining based on knowledge graphs according to claim 5, characterized in that, The semantic-structural fusion association score is obtained by concatenating and element-wise multiplying the embedding vectors of entity pairs, and then inputting the result into a multi-layer feedforward network. The calculation method for the total association score is as follows: ; In the formula, To integrate weights, It is a multi-layer feedforward network. and This is the final embedding vector of the entity. This represents a vector concatenation operation. This represents element-wise product.
8. The method for cross-domain data association mining based on knowledge graphs according to claim 1, characterized in that, The method further includes: Visualize and render the highly relevant results in the association ranking list to generate a visualization report containing interactive subgraphs and path descriptions; Obtain user feedback on the manually annotated results in the visualization report, as the feedback signal.
9. The method for cross-domain data association mining based on knowledge graphs according to claim 1, characterized in that, The process of using the correction results to optimize the extraction of the entity feature vector matrix includes: When the feedback signal indicates that a specific association result is incorrect, the attention weight of the meta-path that generated the association result is reduced. The reduced attention weights are used to trigger incremental training of the relational graph convolutional network to update the entity feature vector matrix.
10. A cross-domain data association mining system based on knowledge graphs, characterized in that, include: The multi-source heterogeneous data access unit is used to standardize and instantiate the business data acquired from the multi-source heterogeneous data domain, and generate a standardized data stream. A cross-domain semantic alignment unit is used to perform semantic alignment processing on cross-domain entities and attributes in the standardized data stream to generate a semantic alignment annotation set. The dynamic knowledge graph construction unit is used to construct a dynamic knowledge graph based on the standardized data stream and the semantic alignment annotation set, and generate a unified cross-domain knowledge graph. The graph-enhanced association feature extraction unit is used to extract graph-enhanced association features from the cross-domain knowledge graph and generate an entity feature vector matrix and a set of candidate association paths. The cross-domain association strength assessment unit is used to assess the cross-domain association strength using the entity feature vector matrix and the candidate association path set, and generate an association ranking list. The association pattern interpretation and visualization unit is used to interactively correct the association sorting list based on the received feedback signals, and to use the correction results to optimize the extraction process of the entity feature vector matrix.