Data blood relationship analysis method and system for power system

By combining the knowledge graph embedding model and the binary classifier, the reliability and accuracy issues of power system data lineage analysis are solved, and efficient data lineage relationship identification and tracking are achieved, which is suitable for large-scale and complex data flow scenarios.

CN120634004APending Publication Date: 2025-09-12STATE GRID HUNAN ELECTRIC POWER COMPANY LIMITED +2
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510700503.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-05-28
Publication Date
2025-09-12

AI Technical Summary

Technical Problem

Existing power system data lineage analysis methods have poor reliability and accuracy, high computing and storage pressure, and rely on manual intervention, resulting in inefficient solutions.

Method used

A knowledge graph embedding model is adopted to train analogy functions and non-analogy functions in combination with a binary classifier to achieve reliability and accuracy analysis of the blood relationship of power system data, including a method flow of steps S1 to S8.

Benefits of technology

It improves the reliability and accuracy of power system data lineage analysis, enhances analysis efficiency, and supports real-time lineage tracking and verification in large-scale and complex data flow scenarios.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120634004A_ABST
    Figure CN120634004A_ABST
Patent Text Reader

Abstract

The invention discloses a data consanguinity analysis method for a power system. The data consanguinity analysis method comprises the following steps: selecting a knowledge graph embedding model as a base model; obtaining an analogy object entity set and a non-analogy object entity set; training an analogy function and a non-analogy function; training query comparison representation; training a binary classifier; determining a dependency relationship of the prediction task according to the obtained binary classifier; calculating a blood relationship score of the prediction task according to the determined dependency relationship; and completing data blood relationship analysis for the power system according to the obtained blood relationship score. The invention also discloses a system for realizing the data blood relationship analysis method for the power system. Through dual functions of the analogy relation and the non-analogy relation, data consanguinity analysis for the power system is achieved, and the method is higher in reliability, better in accuracy and higher in efficiency.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the field of electrical automation, and in particular relates to a data lineage analysis method and system for a power system. Background Art

[0002] With the development of economy and technology and the improvement of people's living standards, electricity has become an indispensable secondary energy source in people's production and life, bringing endless convenience to people's production and life. Therefore, ensuring a stable and reliable supply of electricity has become one of the most important tasks of the power system.

[0003] Currently, the amount of data in power systems is growing exponentially, and power systems are facing challenges such as data silos, complex dependencies, and uneven data quality. Data lineage analysis is crucial to address key issues such as data quality monitoring, resource optimization, and data security.

[0004] Currently, researchers have proposed several solutions for data lineage analysis, including blockchain-based solutions, solutions based on multivariate fusion frameworks, solutions based on metadata management, and solutions based on similarity calculation. However, blockchain-based solutions are severely constrained by performance and cost; solutions based on multivariate fusion frameworks face significant computational and storage pressures due to the complex graphs and algorithms involved; solutions based on metadata management are extremely dependent on metadata quality and can be easily ineffective due to data defects; and solutions based on similarity calculations require manual intervention, which directly leads to subjective biases and poor reliability and accuracy. Summary of the Invention

[0005] One of the purposes of the present invention is to provide a data lineage analysis method for a power system with high reliability, good accuracy and high efficiency.

[0006] A second object of the present invention is to provide a system for implementing the data lineage analysis method for power systems.

[0007] The data lineage analysis method for a power system provided by the present invention comprises the following steps:

[0008] Training phase:

[0009] S1. Select the knowledge graph embedding model as the base model;

[0010] S2. Obtaining analogous object and non-analogous object entity sets;

[0011] S3. Training analogical functions and non-analogous functions;

[0012] S4. Training query contrast representation;

[0013] S5. Train a binary classifier;

[0014] Reasoning stage:

[0015] S6. Determine the dependency of the prediction task based on the obtained binary classifier;

[0016] S7. Calculate the kinship score of the prediction task based on the determined dependency relationship;

[0017] S8. Based on the lineage score obtained in step S7, complete the data lineage analysis for the power system.

[0018] The step S1 of selecting the knowledge graph embedding model as the base model specifically includes the following steps:

[0019] Select the knowledge graph embedding model as the base model;

[0020] The selected knowledge graph embedding models include the TransE model and the RotatE model.

[0021] The step S2 of obtaining the entity set of analog objects and non-analog objects specifically includes the following steps:

[0022] The data structure of the knowledge graph is represented by a triple (s, p, o), where s is the head entity, p is the relationship, and o is the tail entity. The set of all triples of the knowledge graph G is the dataset. The dataset is divided into training dataset, test dataset, and validation dataset.

[0023] In the training dataset, randomly select a triplet, mask the corresponding head entity to obtain an incomplete triplet q(?,p,o), or mask the corresponding tail entity to obtain an incomplete triplet q(s,p,?), and use it as the prediction task;

[0024] All entities in the dataset are used to replace the missing parts of the incomplete triples to obtain new triples;

[0025] In the set of new triplets obtained, remove the triplets that already exist in the training dataset, test dataset, and validation dataset;

[0026] Select a base model and use the scoring function of the base model to score the new triples;

[0027] Sort the new triples according to the scoring results;

[0028] Select the top N with the highest scores anl The triples are taken as analogy triples, and the replaced entities are taken as analogy entities of this query. The set of analogy entities is represented as analogy entity set E anl, the set of analogy triples is expressed as analogy triple set; excluding the top N with the highest score anl The remaining triples outside the triples are regarded as non-analogous triples, and the replaced entities are regarded as non-analogous entities of this query. The set of non-analogous entities is represented as non-analogous entity set E noanl ,The set of non-analogous triples is denoted as non-analogous triple set.

[0029] The training of analog functions and non-analog functions in step S3 specifically includes the following steps:

[0030] Given the entity embedding table E and relationship embedding table R of the knowledge graph G;

[0031] Obtain the embedding vectors of the missing original entities of the incomplete triples from the entity embedding table E, and at the same time obtain the embedding vectors of the original entities of the incomplete triples from the relation embedding table R;

[0032] Create analog entity projection vector Analogy Relationship Projection Vector Non-analog entity projection vector and non-analogous relation projection vector The projection vectors are all parameters to be trained and are initialized using random initialization;

[0033] The following formula is used as the analogy function f anl (p,o) and the non-analog function f noanl (p,o):

[0034]

[0035]

[0036] Where o anl Embedding for analog objects; is the element product; o is the tail entity; λ anl is the first weight; M anl is the first transformation matrix; p is the head entity; o noanl Embedding for non-analogous objects; λ noanl is the second weight; M noanl is the second transformation matrix;

[0037] The following formula is used to calculate the aggregate analog object embedding and aggregate non-analogous object embedding

[0038]

[0039] Where E anl Embed the table for analog object entities; oi is the i-th object in the entity embedding table; S() is the softmax function; E noanl Embedding table for non-analog object entities;

[0040] The following formula is used as the loss function for analog objects and the loss function for non-analog objects:

[0041]

[0042] In the formula is the loss function of the analog object; σ() is the sigmoid function; γ anl is the weight of the analogy loss function; || ||2 is the Euclidean norm; is the loss function for non-analog objects; γ noanl is the weight of the non-analog loss function;

[0043] According to the loss function of the analog object and the loss function of the non-analog object, the analog function and the non-analog function are trained;

[0044] The loss functions of analog objects and non-analog objects are optimized using the following steps:

[0045] Forward propagation calculates the corresponding Euclidean distance; calculates the current loss value of the loss function; backpropagates to obtain the corresponding gradient value; after applying gradient clipping, the model parameters are updated through the Adam optimizer; the gradient update is performed using the following formula:

[0046]

[0047] Where grad' is the updated gradient; grad is the gradient before the update; threshold is the maximum allowed norm of the gradient;

[0048] The following formula is used to update the learning rate:

[0049]

[0050] Where η t is the updated learning rate; η min is the minimum learning rate set; η max is the maximum learning rate set; A cur is the epoch number of the current training; A max is the total number of cycles.

[0051] The training query comparison representation described in step S4 specifically includes the following steps:

[0052] According to the trained analogy function and non-analogy function, the analogy object and non-analogy object of the original entity missing part of the incomplete triple are obtained;

[0053] Put the obtained analog objects and non-analog objects into triples to obtain analog object triples and non-analog object triples;

[0054] Calculate the scores of analog object triples and non-analog object triples through the scoring function of the base model;

[0055] The triple with higher score is regarded as dependent triple, and the corresponding score is dependency score F; judgment is made: for the missing entity of query q, if the score of the triple after replacing with analogous object is higher than the score of the triple after replacing with non-analogous object, then the query q is high analogy dependency, and the dependency mark I is set. q The value of is 1; for the missing entity of query q, if the score of the triple after replacing it with a non-analogous object is higher than the score of the triple after replacing it with an analogous object, then the query q is a high non-analogous dependency, and the dependency mark I q The value of is 0;

[0056] Embed the query q into the vector space and get the embedding vector v of q q for Where MLP() is a multi-layer perceptron, tanh() is an activation function, and W F are the parameters to be trained;

[0057] The minimum data set is represented as M; repeat the above steps according to the dependency marker I q Classify all queries in the dataset: Q(q) is the dependency label I in M ​​except q q The set of queries identical to q is represented as

[0058] During the training process, the following formula is used as the loss function:

[0059]

[0060] Where L sup is the loss function value; Q(q) is the dependency marker I in M ​​except q q The set of queries that are the same as q; v k is the embedding vector of the kth query in the query set Q(q); T is the set temperature parameter; v a is the embedding vector of the ath query in the minimum dataset except the current query q.

[0061] The training of the binary classifier described in step S5 specifically includes the following steps:

[0062] Freeze the parameters and weight values ​​obtained through training;

[0063] V qInput linear layer and follow the ground truth Train a binary classifier with a set loss; where The expression is w is the weight to be trained, b is the bias to be trained;

[0064] During training, the following loss function is used as the setting loss:

[0065]

[0066] Where L cls is the set loss value; |M| is the number of elements in the minimum dataset; I q Dependency marker.

[0067] Determining the dependency relationship of the prediction task based on the obtained binary classifier in step S6 specifically includes the following steps:

[0068] For the prediction task, the trained binary classifier is used to determine the dependency of the prediction task on the analog object or non-analog object, and the corresponding two-dimensional mask vector is obtained. in is the indicator function, if True like If false

[0069] The step S7 of calculating the blood relationship score of the prediction task based on the determined dependency relationship specifically includes the following steps:

[0070] All entities in the dataset are used to replace the missing parts of the incomplete triples of the prediction task and construct new triples;

[0071] Remove the triples that already exist in the original data set from the new triple set to obtain the replaced triples;

[0072] For the replaced triples obtained, the score of the replaced triples is calculated using the following formula:

[0073] Score(s,p,o)=f kge (s,p,o)+[α anl ·f kge (·),α noanl ·f kge (·)]·softmax(B)

[0074] Where Score(s,p,o) is the score of the replaced triple; f kge(·) is the score function value of the base model of the incomplete triplet of the prediction task after replacement; α anl is the first adaptive weight parameter, and α anl =min(|{(·)∈U}| / N anl ,1)×δ,N anl is the number of similar objects retrieved, (·) is the incomplete triplet of the prediction task after replacement, δ is the basic weight hyperparameter; α noanl is the second adaptive weight parameter, and α noanl =min(|{(·)∈U}| / N noanl ,1)×δ,N noanl is the number of non-analog objects retrieved.

[0075] Step S8, based on the blood relationship score obtained in step S7, completes the data blood relationship analysis for the power system, and specifically includes the following steps:

[0076] According to the lineage relationship score obtained in step S7, the triple with the highest score is selected as the result of the final prediction task to complete the data lineage analysis for the power system.

[0077] The present invention also provides a system for implementing the data lineage analysis method for power systems, comprising a model selection module, an object acquisition module, a function training module, a representation training module, a classification training module, a relationship determination module, a score calculation module and a lineage analysis module; the model selection module, the object acquisition module, the function training module, the representation training module, the classification training module, the relationship determination module, the score calculation module and the lineage analysis module are connected in series in sequence; the model selection module is used to select a knowledge graph embedding model as a base model and upload data information to the object acquisition module; the object acquisition module is used to obtain analog object and non-analog object entity sets based on the received data information, and upload the data information to the function training module; the function training module is used to train analog functions and non-analog objects based on the received data information. Analogy function, and upload the data information to the representation training module; the representation training module is used to train the query comparison representation according to the received data information, and upload the data information to the classification training module; the classification training module is used to train the binary classifier according to the received data information, and upload the data information to the relationship determination module; the relationship determination module is used to determine the dependency relationship of the prediction task according to the received data information and the obtained binary classifier, and upload the data information to the score calculation module; the score calculation module is used to calculate the blood relationship score of the prediction task according to the received data information and the determined dependency relationship, and upload the data information to the blood relationship analysis module; the blood relationship analysis module is used to complete the data blood relationship analysis for the power system according to the received data information and the obtained blood relationship score.

[0078] The data lineage analysis method and system for power systems provided by the present invention not only realizes data lineage analysis for power systems through the dual effects of analogical relationships and non-analogous relationships, but also has higher reliability, better accuracy and higher efficiency. BRIEF DESCRIPTION OF THE DRAWINGS

[0079] Figure 1 Schematic diagram of the process of the present invention.

[0080] Figure 2 Schematic diagram of the functional modules of the system of the present invention. DETAILED DESCRIPTION

[0081] like Figure 1 The method flow diagram of the present invention is shown as follows: The data lineage analysis method for power systems disclosed in the present invention includes the following steps:

[0082] Training phase: The main work done in the training phase is to train the model required for the inference phase and some related parameters to ensure better inference in the future;

[0083] S1. Select the knowledge graph embedding model as the base model; specifically, the following steps are included:

[0084] Select the knowledge graph embedding model as the base model;

[0085] The selected knowledge graph embedding models include TransE model, RotatE model, etc.

[0086] S2. Obtaining analogous object and non-analogous object entity sets;

[0087] This step is mainly based on the prediction task to obtain the analog object entity set and the non-analog object entity set, so as to facilitate the subsequent training task. The acquisition of analog entities and non-category entities draws on the idea of ​​the k-nearest neighbor algorithm. Analog entities are strongly related entities to missing entities, while non-analog entities are weakly related entities to missing entities.

[0088] The specific steps include:

[0089] The data structure of the knowledge graph is represented by a triple (s, p, o), where s is the head entity, p is the relationship, and o is the tail entity. The set of all triples of the knowledge graph G is the dataset. The dataset is divided into training dataset, test dataset, and validation dataset.

[0090] In the training dataset, randomly select a triplet, mask the corresponding head entity to obtain an incomplete triplet q(?,p,o), or mask the corresponding tail entity to obtain an incomplete triplet q(s,p,?), and use it as the prediction task;

[0091] All entities in the dataset are used to replace the missing parts of the incomplete triples to obtain new triples;

[0092] In the set of new triplets obtained, remove the triplets that already exist in the training dataset, test dataset, and validation dataset;

[0093] Select the base model and use the scoring function f of the base model kge (s, p, o) scores the obtained new triples; in specific implementation, the base model can adopt traditional knowledge graph embedding models based on geometric assumptions such as TransE, RotatE, HAKE, PairRE, DistMult, ANALOGY, Rot-Pro, DualE, etc.; as an optimization solution, HAKE is selected as the base model in this application;

[0094] Sort the new triples according to the scoring results;

[0095] Select the top N with the highest scores anl The triples are taken as analogy triples, and the replaced entities are taken as analogy entities of this query. The set of analogy entities is represented as analogy entity set E anl , the set of analogy triples is expressed as analogy triple set; excluding the top N with the highest score anl The remaining triples outside the triples are regarded as non-analogous triples, and the replaced entities are regarded as non-analogous entities of this query. The set of non-analogous entities is represented as non-analogous entity set E noanl ,The set of non-analogous triples is expressed as non-analogous triple set;

[0096] S3. Training analogical functions and non-analogous functions;

[0097] The purpose of training the analogy function is to find an analogy object for the entity replaced by the missing part of the incomplete triple, so that the analogy score can be calculated through this analogy object in the future. The purpose of training the non-analogy function is to find a representative non-analogy object for the entity replaced by the missing part of the incomplete triple, so that the non-analogy score can be calculated through this non-analogy object in the future.

[0098] The specific steps include:

[0099] Given the entity embedding table E and relationship embedding table R of the knowledge graph G;

[0100] Obtain the embedding vectors of the missing original entities of the incomplete triples from the entity embedding table E, and at the same time obtain the embedding vectors of the original entities of the incomplete triples from the relation embedding table R;

[0101] Create analog entity projection vector Analogy Relationship Projection Vector Non-analog entity projection vector and non-analogous relation projection vector The projection vectors are all parameters to be trained and are initialized using random initialization;

[0102] Considering that entities are usually associated with multiple relations, the analogy function takes the original entity embedding and the original relation embedding as input and outputs the analogy object embedding; the non-analogy function takes the original entity embedding and the original relation embedding as input and outputs the non-analogy object embedding; therefore, the following formula is used as the analogy function f anl (p,o) and the non-analog function f noanl (p,o):

[0103]

[0104] Where o anl Embedding for analog objects; is the element product; o is the tail entity; λ anl is the first weight; M anl is the first transformation matrix; p is the head entity; o noanl Embedding for non-analogous objects; λ noanl is the second weight; M noanl is the second transformation matrix; the first weight λ anl and the second weight λ noanl , when the data set contains less relation-level information, it can be set to 0 (i.e., ignoring the transformation of the relation part); when the data set contains more relation-level information, it can be set to 1; M anl and M anl It can be initialized by random initialization and automatically learned during training by gradient descent method;

[0105] The following formula is used to calculate the aggregate analog object embedding and aggregate non-analogous object embedding

[0106]

[0107] Where E anl Embed the table for analog object entities; o i is the i-th object in the entity embedding table; S() is the softmax function; E noanl Embedding table for non-analog object entities;

[0108] The following formula is used as the loss function for analog objects and the loss function for non-analog objects:

[0109]

[0110] In the formula is the loss function of the analog object; σ() is the sigmoid function; γ anl is the weight of the analogy loss function; ||||2 is the Euclidean norm; is the loss function for non-analog objects; γ noanl is the weight of the non-analog loss function; where, Used to continuously optimize the analogy function and related parameters, narrowing the gap between the embedding of analog objects and the embedding of aggregated analog objects; Used to continuously optimize non-analogous functions and related parameters, narrowing the gap between non-analogous object embeddings and aggregated non-analogous object embeddings;

[0111] According to the loss function of the analog object and the loss function of the non-analog object, the analog function and the non-analog function are trained;

[0112] The loss functions of analog objects and non-analog objects are optimized using the following steps:

[0113] The forward propagation calculates the corresponding Euclidean distance; the current loss value of the loss function is calculated; the backpropagation obtains the corresponding gradient value; after applying gradient clipping, the model parameters are updated through the Adam optimizer; the Adam optimizer is used, which combines the characteristics of momentum and adaptive learning rate and is suitable for non-stationary objective functions and sparse gradient problems; the loss function is minimized so that the distance between the analog object embedding and the aggregated analog object embedding approaches zero; the gradient update is performed using the following formula:

[0114]

[0115] Where grad' is the updated gradient; grad is the gradient before the update; threshold is the maximum allowed norm of the gradient;

[0116] After each epoch, the validation set performance is evaluated and the learning rate is adjusted. The initial learning rate is set to 0.001, and the cosine annealing strategy is used to dynamically adjust the learning rate. The cycle is 10 epochs, and the minimum learning rate is 0.0001. The following formula is used to update the learning rate:

[0117]

[0118] Where η t is the updated learning rate; η min is the minimum learning rate set; η max is the maximum learning rate set; A cur is the epoch number of the current training; A max is the total number of cycles;

[0119] S4. Training query contrast representation;

[0120] The purpose of training query contrast representations is to classify queries based on the dependency of entities with incomplete triples on analogical entities and non-analogous entities. Specifically, if entities with incomplete triples are highly dependent on analogical entities, then the query is an analogical query; if entities with incomplete triples are highly dependent on non-analogous entities, then the query is a non-analogous query. When all queries are mapped into a vector space, vector representations of queries of the same type are closer, while vector representations of queries of different types are farther apart.

[0121] The specific steps include:

[0122] According to the trained analogy function and non-analogy function, the analogy object and non-analogy object of the original entity missing part of the incomplete triple are obtained;

[0123] Put the obtained analog objects and non-analog objects into triples to obtain analog object triples and non-analog object triples;

[0124] Calculate the scores of analog object triples and non-analog object triples through the scoring function of the base model;

[0125] The triple with higher score is regarded as dependent triple, and the corresponding score is dependency score F; judgment is made: for the missing entity of query q, if the score of the triple after replacing with analogous object is higher than the score of the triple after replacing with non-analogous object, then the query q is high analogy dependency, and the dependency mark I is set. q The value of is 1; for the missing entity of query q, if the score of the triple after replacing it with a non-analogous object is higher than the score of the triple after replacing it with an analogous object, then the query q is a high non-analogous dependency, and the dependency mark I q The value of is 0;

[0126] Embed the query q into the vector space and get the embedding vector v of q q for Where MLP() is a multi-layer perceptron, tanh() is an activation function, and W F are the parameters to be trained;

[0127] The minimum data set is represented as M; repeat the above steps according to the dependency marker I q Classify all queries in the dataset: Q(q) is the dependency label I in M ​​except q q The set of queries identical to q is represented as

[0128] During the training process, the following formula is used as the loss function:

[0129]

[0130] Where L sup is the loss function value; Q(q) is the dependency marker I in M ​​except q q The set of queries that are the same as q; v k is the embedding vector of the kth query in the query set Q(q); T is the set temperature parameter; v a is the embedding vector of the ath query in the minimum dataset except the current query q;

[0131] S5. Train a binary classifier;

[0132] The purpose of training a binary classifier is to obtain the query's dependence on analogous objects and non-analogous objects in the reasoning stage, so as to know which type of entity should be paid more attention to.

[0133] The specific steps include:

[0134] Freeze the parameters and weight values ​​obtained through training;

[0135] V q Input linear layer and follow the ground truth Train a binary classifier with a set loss; where The expression is w is the weight to be trained, b is the bias to be trained;

[0136] During training, the following loss function is used as the setting loss:

[0137]

[0138] Where L cls is the set loss value; |M| is the number of elements in the minimum dataset; I q Mark for dependency;

[0139] Reasoning stage:

[0140] S6. Determine the dependency of the prediction task based on the obtained binary classifier; specifically, the steps include:

[0141] For the prediction task, the dependency of the prediction task on the analog object or non-analog object is determined through the trained binary classifier, and the corresponding two-dimensional mask vector B(η) is obtained as in is the indicator function, if True like If false

[0142] S7. Calculate the blood relationship score of the prediction task based on the determined dependency relationship; specifically, the steps include:

[0143] All entities in the dataset are used to replace the missing parts of the incomplete triples of the prediction task and construct new triples;

[0144] Remove the triples that already exist in the original data set from the new triple set to obtain the replaced triples;

[0145] For the replaced triples obtained, the score of the replaced triples is calculated using the following formula:

[0146] Score(s,p,o)=f kge (s,p,o)+[α anl ·f kge (·),α noanl ·f kge (·)]·softmax(B)

[0147] Where Score(s,p,o) is the score of the replaced triple; f kge (·) is the score function value of the base model of the incomplete triplet of the prediction task after replacement; α anl is the first adaptive weight parameter, and α anl =min(|{(·)∈U}| / N anl ,1)×δ,N anl is the number of analog objects to be retrieved (determined by grid search), (·) is the incomplete triplet of the prediction task after replacement, δ is the basic weight hyperparameter, determined by grid search; α noanl is the second adaptive weight parameter, and α noanl =min(|{(·)∈U}| / N noanl ,1)×δ,N noanl is the number of non-analog objects to be retrieved (determined by grid search); the adaptive weight parameter α anl and α noanl The training dataset is used to determine whether the test triple is suitable for anti-logical reasoning; when neither analogical reasoning nor non-analogical reasoning is suitable, the scoring function degenerates into a basic knowledge graph embedding model;

[0148] Taking the prediction task q(s,p,?) as an example, the dependency of query q on analog objects or non-analog objects is determined by a binary classifier, and a two-dimensional mask vector B(η) is obtained as All entities in the dataset are used to replace the missing parts of the incomplete triples of the prediction task to construct new triples. The triples that already exist in the original dataset are removed from the new triple set to obtain replaced triples. For the replaced triples, the following formula is used to calculate the score of the replaced triples:

[0149] Score(s,p,o)=f kge (s,p,o)+[α anl ·f kge (s,p,o anl ),α noanl ·f kge (s,p,o noanl )]·softmax(B) where o anl and o noanl Proportional to the number of triplets with the same (s, p) in the training set; α anl The calculation formula is α anl =min(|{(s,p,o i )∈U}| / N anl ,1)×δ,α noanl The calculation formula is α noanl =min(|{(s,p,o i )∈U}| / N noanl ,1)×δ;

[0150] S8. Based on the blood relationship score obtained in step S7, complete the data blood relationship analysis for the power system; specifically comprising the following steps:

[0151] According to the lineage relationship score obtained in step S7, the triple with the highest score is selected as the result of the final prediction task to complete the data lineage analysis for the power system.

[0152] The method of the present invention realizes the accurate identification of field-level data lineage relationships through contrastive learning knowledge graph embedding enhancement technology, solving the inefficiency problem caused by traditional methods relying on manual rules or insufficient metadata quality; combined with the global optimization capability of contrastive learning, it significantly improves the scalability of data lineage analysis and supports real-time lineage tracking and verification in large-scale complex data flow scenarios.

[0153] The present invention introduces a global optimization mechanism through a contrastive learning framework, and dynamically adjusts the embedding representation by combining analogical reasoning and non-analogous reasoning. By generating and screening high-confidence analogical and non-analogous triples, the model can mine the potential semantic associations between entities. This helps to complete incomplete triples and improve the robustness of the knowledge graph.

[0154] The present invention enhances the embedding discrimination of entities and relationships through contrastive learning. By aggregating the embedding representations of analogous objects and non-analogous objects, it forces similar entities to cluster in the vector space, separates heterogeneous entities, and enhances semantic boundaries. This helps to improve the reasoning ability of knowledge graphs and avoids the limitation of reasoning ability due to insufficient entity discrimination or redundant relationship embedding.

[0155] The present invention supports fine-grained lineage tracking through hierarchical triple modeling of knowledge graphs; uses a contrastive loss function to perform end-to-end training on the graph to ensure that high-confidence connections in lineage relationships have a shorter Euclidean distance in the embedding space; this helps to improve the credibility of data lineage and enhance the data lineage structure.

[0156] The present invention can achieve cross-scenario migration through a universal enhancement framework and the design of adaptive weights. By dynamically adjusting the weights of analogy and non-analogy reasoning, it can effectively avoid performance fluctuations caused by differences in data distribution.

[0157] like Figure 2 The figure shows a schematic diagram of the functional modules of the system of the present invention: the system disclosed by the present invention for realizing the data lineage analysis method for the power system comprises a model selection module, an object acquisition module, a function training module, a representation training module, a classification training module, a relationship determination module, a scoring calculation module and a lineage analysis module; the model selection module, the object acquisition module, the function training module, the representation training module, the classification training module, the relationship determination module, the scoring calculation module and the lineage analysis module are connected in series in sequence; the model selection module is used to select the knowledge graph embedding model as the base model and upload the data information to the object acquisition module; the object acquisition module is used to obtain the analog object and non-analog object entity set according to the received data information, and upload the data information to the function training module; the function training module is used to, according to the received data information, Training analogy functions and non-analogy functions, and uploading data information to the representation training module; the representation training module is used to train query comparison representation based on the received data information, and upload the data information to the classification training module; the classification training module is used to train the binary classifier based on the received data information, and upload the data information to the relationship determination module; the relationship determination module is used to determine the dependency relationship of the prediction task based on the received data information and the obtained binary classifier, and upload the data information to the score calculation module; the score calculation module is used to calculate the blood relationship score of the prediction task based on the received data information and the determined dependency relationship, and upload the data information to the blood relationship analysis module; the blood relationship analysis module is used to complete the data blood relationship analysis for the power system based on the received data information and the obtained blood relationship score.

[0158] The present invention enhances the reasoning ability of the knowledge graph embedding model through the dual effects of analogy and non-analogy relationships; analogy relationships mine the potential semantic associations between entities and improve the ability to complete missing information; non-analogy relationships enhance the ability to distinguish heterogeneous entities and avoid reasoning errors; the combination of the two optimizes the semantic structure of the embedding space and significantly improves reasoning accuracy and robustness.

[0159] The present invention uses contrastive learning to optimize knowledge graph embedding, and improves the model's reasoning and generalization capabilities through a global optimization mechanism; contrastive learning dynamically adjusts the embedding representation to ensure that similar entities are closer in the embedding space and heterogeneous entities are further away, enhancing the clarity of semantic boundaries, while significantly improving the scalability of the model in large-scale complex data scenarios.

[0160] The present invention solves the problem of performance fluctuation of traditional models in different scenarios through a dynamic adaptive weight adjustment mechanism; the adaptive weights dynamically adjust the weights of analogy and non-analogy reasoning according to the distribution of training data, avoiding reasoning bias caused by data distribution differences, while enhancing the robustness of the model under data noise and sparsity problems, and improving cross-scenario generalization capabilities.

[0161] This paper proposes a universal knowledge graph embedding enhancement framework with a modular design, which is adaptable to most knowledge graph embedding models (such as TransE, RotatE, etc.); the framework dynamically adjusts the inference weight to quickly adapt to the data distribution characteristics in different scenarios, supports real-time lineage tracking and verification in large-scale complex data flow scenarios, and facilitates the integration of new embedding models or inference algorithms to further expand application scenarios.

Claims

1. A data lineage analysis method for a power system, comprising the following steps: Training phase: S1. Select the knowledge graph embedding model as the base model; S2. Obtaining analogous object and non-analogous object entity sets; S3. Training analogical functions and non-analogous functions; S4. Training query contrast representation; S5. Train a binary classifier; Reasoning stage: S6. Determine the dependency of the prediction task based on the obtained binary classifier; S7. Calculate the kinship score of the prediction task based on the determined dependency relationship; S8. Based on the lineage score obtained in step S7, complete the data lineage analysis for the power system.

2. The data lineage analysis method for power system according to claim 1 is characterized in that The step S1 of selecting the knowledge graph embedding model as the base model specifically includes the following steps: Select the knowledge graph embedding model as the base model; The selected knowledge graph embedding models include the TransE model and the RotatE model.

3. The data lineage analysis method for power system according to claim 2 is characterized in that The step S2 of obtaining the entity set of analog objects and non-analog objects specifically includes the following steps: The data structure of the knowledge graph is represented by a triple (s, p, o), where s is the head entity, p is the relationship, and o is the tail entity. The set of all triples of the knowledge graph G is the dataset. The dataset is divided into training dataset, test dataset, and validation dataset. In the training dataset, randomly select a triplet, mask the corresponding head entity to obtain an incomplete triplet q(?,p,o), or mask the corresponding tail entity to obtain an incomplete triplet q(s,p,?), and use it as the prediction task; All entities in the dataset are used to replace the missing parts of the incomplete triples to obtain new triples; In the set of new triplets obtained, remove the triplets that already exist in the training dataset, test dataset, and validation dataset; Select a base model and use the scoring function of the base model to score the new triples; Sort the new triples according to the scoring results; Select the top N with the highest scores anl The triples are taken as analogy triples, and the replaced entities are taken as analogy entities of this query. The set of analogy entities is represented as analogy entity set E anl , the set of analogy triples is expressed as analogy triple set; excluding the top N with the highest score anl The remaining triples outside the triples are regarded as non-analogous triples, and the replaced entities are regarded as non-analogous entities of this query. The set of non-analogous entities is represented as non-analogous entity set E noanl ,The set of non-analogous triples is denoted as non-analogous triple set.

4. The data lineage analysis method for power system according to claim 3 is characterized in that The training of analog functions and non-analog functions in step S3 specifically includes the following steps: Given the entity embedding table E and relationship embedding table R of the knowledge graph G; Obtain the embedding vectors of the missing original entities of the incomplete triples from the entity embedding table E, and at the same time obtain the embedding vectors of the original entities of the incomplete triples from the relation embedding table R; Create analog entity projection vector Analogy Relationship Projection Vector Non-analog entity projection vector and non-analogous relation projection vector The projection vectors are all parameters to be trained and are initialized using random initialization; The following formula is used as the analogy function f anl (p,o) and the non-analog function f noanl (p,o): Where o anl Embedding for analog objects; is the element product; o is the tail entity; λ anl is the first weight; M anl is the first transformation matrix; p is the head entity; o noanl Embedding for non-analogous objects; λ noanl is the second weight; M noanl is the second transformation matrix; The following formula is used to calculate the aggregate analog object embedding and aggregate non-analogous object embedding Where E anl Embed the table for analog object entities; o i is the i-th object in the entity embedding table; S() is the softmax function; E noanl Embedding table for non-analog object entities; The following formula is used as the loss function for analog objects and the loss function for non-analog objects: In the formula is the loss function of the analog object; σ() is the sigmoid function; γ anl is the weight of the analogy loss function; || ||2 is the Euclidean norm; is the loss function for non-analog objects; γ noanl is the weight of the non-analog loss function; According to the loss function of the analog object and the loss function of the non-analog object, the analog function and the non-analog function are trained; The loss functions of analog objects and non-analog objects are optimized using the following steps: Forward propagation calculates the corresponding Euclidean distance; calculates the current loss value of the loss function; backpropagates to obtain the corresponding gradient value; after applying gradient clipping, the model parameters are updated through the Adam optimizer; the gradient update is performed using the following formula: Where grad' is the updated gradient; grad is the gradient before the update; threshold is the maximum allowed norm of the gradient; The following formula is used to update the learning rate: Where η t is the updated learning rate; η min is the minimum learning rate set; η max is the maximum learning rate set; A cur is the epoch number of the current training; A max is the total number of cycles.

5. The data lineage analysis method for power system according to claim 4 is characterized in that The training query comparison representation described in step S4 specifically includes the following steps: According to the trained analogy function and non-analogy function, the analogy object and non-analogy object of the original entity missing part of the incomplete triple are obtained; Put the obtained analog objects and non-analog objects into triples to obtain analog object triples and non-analog object triples; Calculate the scores of analog object triples and non-analog object triples through the scoring function of the base model; The triple with higher score is regarded as dependent triple, and the corresponding score is dependency score F; judgment is made: for the missing entity of query q, if the score of the triple after replacing with analogous object is higher than the score of the triple after replacing with non-analogous object, then the query q is high analogy dependency, and the dependency mark I is set. q The value of is 1; for the missing entity of query q, if the score of the triple after replacing it with a non-analogous object is higher than the score of the triple after replacing it with an analogous object, then the query q is a high non-analogous dependency, and the dependency mark I q The value of is 0; Embed the query q into the vector space and get the embedding vector v of q q for Where MLP() is a multi-layer perceptron, tanh() is an activation function, and W F are the parameters to be trained; The minimum data set is represented as M; Repeat the above steps, according to the dependency mark I q Classify all queries in the dataset: Q(q) is the dependency label I in M ​​except q q The set of queries identical to q is represented as During the training process, the following formula is used as the loss function: Where L sup is the loss function value; Q(q) is the dependency marker I in M ​​except q q The set of queries that are the same as q; v k is the embedding vector of the kth query in the query set Q(q); T is the set temperature parameter; v a is the embedding vector of the ath query in the minimum dataset except the current query q.

6. The data lineage analysis method for power system according to claim 5, characterized in that The training of the binary classifier described in step S5 specifically includes the following steps: Freeze the parameters and weight values ​​obtained through training; V q Input linear layer and follow the ground truth Train a binary classifier with a set loss; where The expression is w is the weight to be trained, b is the bias to be trained; During training, the following loss function is used as the setting loss: Where L cls is the set loss value; |M| is the number of elements in the minimum dataset; I q Dependency marker.

7. The data lineage analysis method for power system according to claim 6, characterized in that Determining the dependency relationship of the prediction task based on the obtained binary classifier in step S6 specifically includes the following steps: For the prediction task, the dependency of the prediction task on the analog object or non-analog object is determined through the trained binary classifier, and the corresponding two-dimensional mask vector B(η) is obtained as in is the indicator function, if True like If false 8. The data lineage analysis method for power system according to claim 7, characterized in that The step S7 of calculating the blood relationship score of the prediction task based on the determined dependency relationship specifically includes the following steps: All entities in the dataset are used to replace the missing parts of the incomplete triples of the prediction task and construct new triples; Remove the triples that already exist in the original data set from the new triple set to obtain the replaced triples; For the replaced triples obtained, the score of the replaced triples is calculated using the following formula: Score(s,p,o)=f kge (s,p,o)+[α anl ·f kge (·),α noanl ·f kge (·)]·soft max(B) Where Score(s,p,o) is the score of the replaced triple; f kge (·) is the score function value of the base model of the incomplete triplet of the prediction task after replacement; α anl is the first adaptive weight parameter, and α anl =min({(·)∈U}| / N anl ,1)×δ,N anl is the number of similar objects retrieved, (·) is the incomplete triplet of the prediction task after replacement, δ is the basic weight hyperparameter; α noanl is the second adaptive weight parameter, and α noanl =min({(·)∈U}| / N noanl ,1)×δ,N noanl is the number of non-analog objects retrieved.

9. The data lineage analysis method for power system according to claim 8, characterized in that Step S8, based on the blood relationship score obtained in step S7, completes the data blood relationship analysis for the power system, and specifically includes the following steps: According to the lineage relationship score obtained in step S7, the triple with the highest score is selected as the result of the final prediction task to complete the data lineage analysis for the power system.

10. A system for implementing the data lineage analysis method for a power system according to any one of claims 1 to 9, characterized in that It includes a model selection module, an object acquisition module, a function training module, a representation training module, a classification training module, a relationship determination module, a score calculation module and a lineage analysis module; the model selection module, the object acquisition module, the function training module, the representation training module, the classification training module, the relationship determination module, the score calculation module and the lineage analysis module are connected in series in sequence; the model selection module is used to select the knowledge graph embedding model as the base model and upload the data information to the object acquisition module; the object acquisition module is used to obtain the analog object and non-analog object entity set based on the received data information, and upload the data information to the function training module; The function training module is used to train analog functions and non-analog functions based on the received data information, and upload the data information to the representation training module; the representation training module is used to train query comparison representation based on the received data information, and upload the data information to the classification training module; The classification training module is used to train the binary classifier based on the received data information and upload the data information to the relationship determination module; The relationship determination module is used to determine the dependency relationship of the prediction task based on the received data information and the obtained binary classifier, and upload the data information to the score calculation module; The scoring calculation module is used to calculate the blood relationship score of the prediction task based on the received data information and the determined dependency relationship, and upload the data information to the blood relationship analysis module; The lineage analysis module is used to complete data lineage analysis for the power system based on the received data information and the obtained lineage relationship score.