Data anomaly detection method based on knowledge graph

By collecting multi-dimensional features in the knowledge graph and using logistic regression models to calculate the data correlation coefficient and dynamically adjusting the connection rules, the problem of not promptly reflected changes in the data correlation in the knowledge graph is solved, and the accuracy and efficiency of data connections are improved.

CN120069043AActive Publication Date: 2025-05-30MAILEFENG (XIAMEN) E-COMMERCE CO LTD

Patent Information

Application Number
CN202510547772.5
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-28
Publication Date
2025-05-30
Estimated Expiration
2045-04-28

AI Technical Summary

Technical Problem

In the knowledge graph, due to the diverse sources of data input and frequent updates, overload correlations are caused, and correlation changes are not reflected in time, resulting in an increase in the probability of invalid correlation, reducing data utilization efficiency, and causing knowledge graph distortion.

Method used

By collecting the semantic similarity, entity attribute matching degree, relationship link hierarchy, and entity interactions of each entity, the logistic regression model is used to calculate the data correlation coefficient, and compare it with the preset threshold, dynamically adjust the connection rules to ensure the scientificity of the number of standard connections.

Benefits of technology

It improves the accuracy of data connections, avoids the increase of invalid connections, improves the efficiency of data utilization, and enhances the overall structure of the knowledge graph and the dynamic adaptability of data correlation.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120069043A_ABST
    Figure CN120069043A_ABST
Patent Text Reader

Abstract

The invention discloses a data anomaly detection method based on a knowledge graph, relates to the technical field of data detection, and is used for solving the problems of graph entity relationship overload association and invalid connection probability increase, and accurately obtaining a data association degree coefficient through multi-dimensional feature acquisition and logistic regression model calculation. The correlation degree coefficient is compared with the preset threshold value, it is ensured that the screened correlation entities have high correlation, the accuracy of data connection is improved, meanwhile, the number of alternative connections and the number of screened entities are counted, and the connection rule is dynamically adjusted through weighted calculation of the deviation rate and the change rate, so that the accuracy of data connection is improved. The association degree coefficients in the alternative connection number are sorted and optimized, a group of fuzzy rules are formulated for fuzzy reasoning according to the individual connection relation, the time association degree and the data sparseness, and the weight result is calculated according to the data association degree coefficients, so that the data utilization efficiency is improved, and the anomaly detection accuracy is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of data detection, and more specifically, to a method for detecting data anomalies based on a knowledge graph. Background Art

[0002] A knowledge graph is a technology for structured storage and expression of knowledge. By constructing a semantic network through entities (nodes) and the relationships (edges) between them, and through semantic rules and relationship constraints, it can express complex logical relationships, support the update of dynamic data and the supplementation of knowledge, and can infer unexplicit data based on logical rules and context information.

[0003] The prior art has the following deficiencies:

[0004] Currently, in a knowledge graph, the relevance of data determines the number of connections between entities and their dynamic changes. The strength of the relevance directly affects the rationality of the connections and the overall structure of the knowledge graph. However, due to diverse data input sources and frequent updates, it leads to overloaded associations, and the changes in relevance are not reflected in the graph in a timely manner, increasing the probability of invalid associations, reducing the data utilization efficiency, and causing the knowledge graph to be distorted. Therefore, a method for detecting data anomalies based on a knowledge graph is proposed.

[0005] The above information disclosed in the background art section is only used to enhance the understanding of the background of the present disclosure, and thus it may include information that does not constitute the prior art known to those of ordinary skill in the art. Summary of the Invention

[0006] In order to overcome the above defects of the prior art, the embodiments of the present invention provide a method for detecting data anomalies based on a knowledge graph, and solve the problems proposed in the above background art by applying different product inspection methods.

[0007] To achieve the above object, the present invention provides the following technical solution, a method for detecting data anomalies based on a knowledge graph. To achieve the above object, the present invention provides the following technical solution, a method for detecting data anomalies based on a knowledge graph, including:

[0008] S1: Collect the semantic similarity, entity attribute matching degree, relationship connection level number, and entity interaction times corresponding to each entity, and obtain the data association degree coefficient through substitution into a logistic regression calculation;

[0009] S2: Obtain the data correlation degree coefficient. Taking one entity as an individual, count the data correlation degree coefficients corresponding to all other entities related to it, compare them with a preset correlation threshold, select the other entities with data correlation degree coefficients marked as associated entities and count the quantity as the alternative connection quantity. At the same time, count the quantity of screened entities, compare them with the abnormal threshold, and according to the comparison result, enable the adjustment connection rule, collect the deviation rate and the change ratio of the alternative connection quantity for weighted calculation to obtain the adjustment ratio, and determine the standard connection quantity;

[0010] S3: Obtain the standard connection quantity and the alternative connection quantity, and make a comparison to generate a connection quantity adjustment plan. Sort the data correlation degree coefficients in the alternative connection quantity related to the individual according to the numerical value from small to large, and delete them from small to large until the alternative connection quantity reaches the standard connection quantity to obtain the connection relationship of the individual;

[0011] S4: Obtain the connection relationship of the individual to get the directly connected entities and indirectly connected entities, record the interaction event timestamp sequence between the individual and each connected entity to obtain the time correlation degree and data sparsity;

[0012] S5: Use fuzzy logic for the time correlation degree and data sparsity to determine the weight result of calculating the data correlation degree coefficient.

[0013] In a preferred embodiment, through the unique identifier of each entity, obtain the name, label, attribute value and annotation of the entity, as well as the adjacent nodes and their connection relationships of the entity, convert the entity and its information into a vector form, use a pre-trained word embedding model to encode the text description of the entity to generate an entity vector, and substitute the entity vector into the cosine similarity to calculate the semantic similarity between two entities , Denote the i-th comparison, then there are i comparison entities;

[0014] Extract the attributes of each entity from the knowledge graph. According to the different types of attributes, define different matching rules. For discrete and definite attribute values, compare whether the attribute values are the same, that is, calculate the ratio by matching the number of attributes with the total number of attributes to obtain the discrete matching degree. For continuous numerical attributes, use the difference to calculate the continuous matching degree; set different weights for the discrete matching degree and the continuous matching degree, and substitute them into the weighted calculation to obtain the entity attribute matching degree ;

[0015] By constructing paths in the knowledge graph, use breadth-first search to calculate the shortest path. Starting from the starting entity node, expand layer by layer outward until the target entity node is found, and the path length is the number of relationship connection levels ;

[0016] By distinguishing direct interactions and indirect interactions, and recording each time an entity interacts through a connection relationship as a direct interaction. By searching for all paths passing through the current entity and calculating the number of paths, on each path, every time the current entity is passed through, it is recorded as an indirect interaction. The direct interactions and indirect interactions are accumulated to obtain the entity interaction count. ;

[0017] Substitute the subject semantic similarity, entity attribute matching degree, relationship connection level number, and entity interaction count into the logistic regression formula to calculate the data association degree coefficient.

[0018] In a preferred embodiment, count the data association degree coefficients corresponding to all other entities related to it;

[0019] After obtaining the data association degree coefficients corresponding to all other entities, compare and analyze the data association degree coefficients with the continuously iterated association threshold;

[0020] If the data association degree coefficient is greater than or equal to the association threshold, mark the other entity corresponding to the data association degree coefficient as an associated entity and generate an association signal;

[0021] If the data association degree coefficient is less than the association threshold, mark the other entity corresponding to the data association degree coefficient as a screened entity and generate a screening signal;

[0022] Count the data association degree coefficients of the entities marked as associated entities as the alternative connection quantity;

[0023] Count the data association degree coefficients of the entities marked as screened entities and determine the quantity, and compare it with the abnormal threshold;

[0024] If the quantity of the data association degree coefficients of the entities marked as screened entities is greater than or equal to the abnormal threshold, generate a data anomaly signal. After deleting the screened entities, enable the adjusted connection rule. Otherwise, generate a data normalcy signal.

[0025] In a preferred embodiment, enable the adjusted connection rule. The specific steps are as follows:

[0026] Collect the deviation rate and the change ratio of the alternative connection quantity;

[0027] The deviation rate is used to measure the degree of difference between the quantity of screened entities and the abnormal threshold. Its acquisition logic is to calculate the deviation rate P by subtracting the abnormal threshold from the quantity of the data association degree coefficients of the entities marked as screened entities and taking the ratio with the abnormal threshold.

[0028] The change ratio of the number of alternative connections reflects the degree of fluctuation of the current number of alternative connections compared to the average number of alternative connections, and is used to measure the stability of the number of alternative connections. Its acquisition logic is to subtract the historical average number of alternative connections from the current number of alternative connections and calculate the ratio with the historical average number of alternative connections to obtain the change ratio R of the number of alternative connections.

[0029] In a preferred embodiment, the deviation rate and the change ratio of the number of alternative connections are substituted into the weighted formula for calculation to obtain the adjustment ratio, and the corresponding connection rules are added: ;

[0030] In the formula, is the adjustment ratio, and are the influence weights of the control deviation rate and the change ratio of the number of alternative connections on the adjustment ratio respectively;

[0031] Among them, , ensuring that the adjustment ratio will not cause excessive correction due to excessive weight;

[0032] Multiply the adjustment ratio by the number of alternative connections to obtain the standard number of connections B;

[0033] For the number of data association degree coefficients marked as deleted entities being less than the abnormal threshold, the default adjustment ratio is 1, and the default number of alternative connections is the standard number of connections.

[0034] In a preferred embodiment, if the standard number of connections is the same as the number of alternative connections, directly use the number of alternative connections as the standard number of connections to complete the adjustment of the connection relationship of the individual;

[0035] If there is a difference between the standard number of connections and the number of alternative connections, a connection number adjustment plan is generated. Subtract the standard number of connections from the number of alternative connections to obtain the difference connection number. Delete the set of data association degree coefficients in the number of alternative connections from small to large so that the number of alternative connections is the same as the standard number of connections, and connect the connection relationship of the individual according to the adjusted number of alternative connections.

[0036] In a preferred embodiment, through the connection relationship of the individual, extract the interaction timestamp sequence with each directly or indirectly connected entity, calculate the standard deviation of the timestamp sequence, and introduce a normalization formula to map the time correlation to the [0,1] interval to obtain the time correlation degree;

[0037] Statistically count the number of directly connected and indirectly connected entities of an individual, calculate the ratio of the total number of connections of all entities to the number of all entities in the knowledge graph to obtain the average number of connections of entities in the knowledge graph, and calculate the ratio of the number of connections of the individual to the average number of connections of entities in the knowledge graph to obtain the data sparsity.

[0038] In a preferred embodiment, the time correlation degree and the data sparsity are defined as input variables and divided into different fuzzy sets respectively;

[0039] Define the calculation weight result of the data association degree coefficient as the output variable and divide it into a fuzzy set;

[0040] Formulate fuzzy rules to describe the influence of the time correlation degree and the data sparsity on the calculation weight result of the data association degree coefficient;

[0041] Perform fuzzy reasoning according to the fuzzy rules to determine the calculation weight result of the data association degree coefficient.

[0042] Technical effects and advantages of the present invention:

[0043] 1. Through multi-dimensional feature collection and calculation using a logistic regression model, the present invention accurately obtains the data association degree coefficient. By comparing the association degree coefficient with a preset threshold, it is ensured that the selected associated entities have a high correlation, thereby improving the accuracy of data connection. At the same time, count the number of alternative connections and the number of screened entities, and through the weighted calculation of the deviation rate and the change ratio, dynamically adjust the connection rules to ensure the scientific nature of the standard connection quantity. Through the sorting and optimization of the association degree coefficients in the alternative connection quantities, fine control of individual connections is achieved, which not only avoids the increase of invalid connections but also improves the efficiency of data utilization.

[0044] 2. By obtaining the individual connection relationship, the present invention gets the directly connected entities and the indirectly connected entities, records the time stamp sequence of the interaction events between the individual and each connected entity, obtains the time correlation degree and the data sparsity, formulates a set of fuzzy rules for fuzzy reasoning, determines the calculation weight result of the data association degree coefficient, improves the dynamic adaptation ability of data correlation, enhances the flexibility and stability of the connection rules, and increases the accuracy of anomaly detection. Brief description of the drawings

[0045] Figure 1 It is the method flow chart of the data anomaly detection method based on the knowledge graph of the present invention. Detailed implementation manners

[0046] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.

[0047] In fact, data anomalies regarding knowledge graphs are issues of data correlation. In data correlation, data processing includes but is not limited to data redundancy and repetition, data missingness and incompleteness, data conflicts and contradictions, abnormal associations and noisy data, as well as polysemy and semantic ambiguity. Among them, due to the excessive number of data sources, the relationship connections become complex and difficult to process, even causing relationship contradictions and conflicts, resulting in data anomalies.

[0048] Embodiment 1

[0049] Please refer to Figure 1 , the data anomaly detection method based on a knowledge graph, and the specific operation process is as follows:

[0050] S1: Collect the semantic similarity, entity attribute matching degree, relationship connection level number, and entity interaction times corresponding to each entity, and obtain the data correlation degree coefficient through substitution into logistic regression calculation.

[0051] Among them, the knowledge graph contains multiple entities. Each entity forms corresponding data in the graph through different attributes and relationships, and at the same time, each entity represents a piece of data.

[0052] Semantic similarity refers to an index that measures the similarity degree of two entities in the semantic level in a knowledge graph. It reflects the high or low correlation between two entities in terms of description, context, or logical meaning. Its acquisition logic is to obtain the name, label, attribute value, and annotation of the entity through the unique identifier of each entity, as well as the adjacent nodes of the entity and their connection relationships, convert the entity and its information into vector form, use a pre-trained word embedding model to encode the text description of the entity, generate an entity vector, and substitute the entity vector into the cosine similarity to calculate the semantic similarity between two entities.

[0053] It should be noted that the pre-trained word embedding model is a language model pre-trained based on large-scale text data. Its core purpose is to map text (words, phrases, sentences, etc.) into a continuous vector space, so that texts with similar semantics are closer in the vector space. These models utilize the context relationships of natural language to learn the semantic representations of words or sentences and are the basis of semantic computing. For example, the static word embedding model Word2Vec is first trained through the Skip-Gram or CBOW method to find that words with similar semantics appear in similar contexts, and each word has a fixed vector representation. Then, the vocabulary with vector representations is substituted into the cosine similarity calculation to obtain;

[0054] Specifically, the calculation formula of cosine similarity is expressed as:

[0055] Let entities A and B, and the entity vectors be represented by and respectively. The specific cosine similarity formula is as follows: ;

[0056] In the formula, is the dot product of the two entity vectors, is the modulus of entity vector A, is the modulus of entity vector B; according to the formula, the semantic similarity is obtained; where i is the i-th comparison, and the comparison entities are i;

[0057] The entity attribute matching degree refers to the consistency matching of two entities in the knowledge graph based on attribute information (such as name, type, feature value, etc.), which is used to measure the strength of their relevance at the attribute level. Its acquisition logic is to extract the attributes of each entity from the knowledge graph, and according to the different attribute types, different matching rules are defined. For discrete and definite attribute values, compare whether the attribute values are the same, that is, calculate the ratio of the number of matching attributes to the total number of attributes to obtain the discrete matching degree. For continuous numerical attributes, use the difference to calculate the continuous matching degree; set different weights for the discrete matching degree and the continuous matching degree, and substitute them into the weighted calculation to obtain the entity attribute matching degree ;

[0058] Among them, for continuous numerical attributes, the difference is used to calculate the continuous matching degree. The specific calculation formula is: ;

[0059] In the formula, is the continuous matching degree, A and B are the values of the corresponding attributes of the two entities, is the maximum value of the corresponding attributes of the two entities;

[0060] Specifically, the attributes of each entity include, but are not limited to, basic attributes (such as entity name, identifier, type), numerical attributes (such as quantity, time, etc.), classification attributes (such as classification labels, status), text attributes (such as description, definition, etc.), and relationship attributes (such as relationships with other entities, etc.), which will not be elaborated here;

[0061] It should be noted that different weights are set for the discrete matching degree and the continuous matching degree. The setting of the weights is based on the more critical degree of discrete attributes (such as type, classification) in the field set by the experimenter, or by analyzing data to evaluate the contribution of discrete and continuous attributes to the matching result to set the weights. For example, calculating the correlation between the two types of attributes and the matching result, etc., which will not be elaborated here;

[0062] Then, as can be seen from the above, the weights for different entities are also different;

[0063] The relationship connection level number refers to the minimum number of hops (i.e., the level depth) required for two entities to be connected through a relationship (relationship edge) in the knowledge graph, which is used to measure the strength of their structural association. The smaller the level number, the closer the connection between the two entities. Its acquisition logic is to construct a path in the knowledge graph, use breadth-first search to calculate the shortest path, start from the starting entity node, expand layer by layer outward until the target entity node is found, and the path length is the relationship connection level number ;

[0064] Specifically, breadth-first search is an algorithm for traversing a graph. It visits all nodes in the graph in hierarchical order (starting from the starting node and expanding layer by layer). BFS first visits the nodes closer to the starting node, and then gradually visits the nodes farther away from the starting node until all reachable nodes in the graph are traversed. The specific implementation steps are as follows: starting from the starting node in the graph, first visit this node, and then visit all unvisited nodes directly connected to the current node (i.e., adjacent nodes). For each visited node, add its adjacent nodes to the queue to be visited, and visit them in the order of the nodes taken out from the queue until the queue is empty. This process expands the nodes of the graph layer by layer, ensuring that the path from the starting node to the target node is the shortest;

[0065] It should be noted that the starting entity node is the entity node compared with the other entity nodes in the above content, and the target entity node is the other corresponding entity nodes to be compared;

[0066] Among them, the relationship connection level number is the shortest path length from one entity to another entity;

[0067] The number of entity interactions refers to the number of interactions (or connections) that occur between two entities in a knowledge graph through direct or indirect relationships. The number of interactions is an important indicator for measuring the strength of the connection between two entities. The more interactions there are, usually the stronger the correlation between the two entities. Its acquisition logic is to distinguish direct interactions and indirect interactions, and record each time an entity interacts through a connection relationship as a direct interaction. By searching all paths passing through the current entity and calculating the number of paths, each time the current entity is passed on each path, it is recorded as an indirect interaction. The direct interactions and indirect interactions are accumulated to obtain the number of entity interactions ;

[0068] It should be noted that the number of direct interactions refers to the number of interactions that occur between two entities through a direct single relationship, and the number of indirect interactions refers to the number of interactions that occur between two entities through multiple relationship chains, that is, indirect connections;

[0069] Optionally, the experimenter can add more detailed multi-level interaction rules on the above basis. The direct relationship between two entities is set as the first-level interaction, the indirect relationship through an intermediate node is set as the second-level interaction, the relationship chain through two intermediate nodes is set as the third-level interaction, and so on. Generally, the deeper the level of interaction, the weaker the connection, which is used as the basis for evaluating the strength of the connection;

[0070] Furthermore, different weights are set for the second and third-level interactions of indirect relationships with multiple interactions. Different weights are set according to the number of interactions to enhance the possibility of hidden data interactions. The specific weight rules set are not limited and will not be elaborated here;

[0071] The semantic similarity, entity attribute matching degree, relationship connection level number, and number of entity interactions are normalized. All input variables will be converted to the same range to ensure the balanced contribution of each input to the model. Specifically, the normalization method is to normalize through the [0,1] interval, and the specific formula is expressed as: ;

[0072] In the formula, is the semantic similarity, is the normalized semantic similarity, is the minimum value of the semantic similarity, is the maximum value of the semantic similarity;

[0073] Among them, the entity attribute matching degree, relationship connection level number, and number of entity interactions also adopt the above formula for normalization, which will not be elaborated here;

[0074] After normalization, the values of the subject semantic similarity, entity attribute matching degree, relationship connection level number, and entity interaction times are all within the range of [0, 1]. The system can compare and make decisions on a unified scale, thereby improving the accuracy and efficiency of the logistic regression model;

[0075] Substitute the subject semantic similarity, entity attribute matching degree, relationship connection level number, and entity interaction times into the logistic regression formula to calculate the data association degree coefficient. The specific formula is expressed as follows: ;

[0076] In the formula, L is the result of logistic regression calculation, that is, the data association degree coefficient, e is the natural base, y is the linear combination term of the logistic regression model. Specifically, y can be set as: ;

[0077] In the formula, is the bias term, , , and are the regression coefficients of the subject semantic similarity, entity attribute matching degree, relationship connection level number, and entity interaction times respectively;

[0078] S2: Obtain the data association degree coefficient. Taking one entity as an individual, count the data association degree coefficients corresponding to all other entities related to it, compare them with the preset association threshold, select the data association degree coefficients of the other entities marked as associated entities and count the quantity as the alternative connection quantity. At the same time, count the quantity of screened entities, compare them with the abnormal threshold, and according to the comparison result, enable the adjustment connection rule, collect the deviation rate and the change ratio of the alternative connection quantity for weighted calculation to obtain the adjustment ratio, and determine the standard connection quantity;

[0079] Among them, in the process of calculating the subject semantic similarity, entity attribute matching degree, relationship connection level number, and entity interaction times in the above step S1, at least two individuals are required to calculate. Therefore, it can be inferred that taking one entity as an individual, this entity is the entity node compared with the other entity nodes;

[0080] Count the data association degree coefficients corresponding to all other entities related to it, and record the quantity of the data association degree coefficients corresponding to all other entities as n;

[0081] After obtaining the data association degree coefficients corresponding to all other entities, compare and analyze the data association degree coefficients with the continuously iterated association threshold;

[0082] If the data correlation degree coefficient is greater than or equal to the correlation threshold, mark other entities corresponding to the data correlation degree coefficient as associated entities and generate an association signal;

[0083] If the data correlation degree coefficient is less than the correlation threshold, mark other entities corresponding to the data correlation degree coefficient as entities to be screened out and generate a screening signal;

[0084] It should be noted that the correlation threshold is obtained from the data correlation degree coefficients corresponding to all other current entities and the historical relevance analysis set data, which will not be elaborated here;

[0085] Statistically analyze the data correlation degree coefficients of the entities marked as associated entities and determine the quantity as m, which is used as the alternative connection quantity;

[0086] Statistically analyze the data correlation degree coefficients of the entities marked as entities to be screened out and determine the quantity as n - m, and compare it with the abnormal threshold;

[0087] If the quantity of the data correlation degree coefficients of the entities marked as entities to be screened out is greater than or equal to the abnormal threshold, generate a data anomaly signal. After deleting the entities to be screened out, enable the adjusted connection rule. Otherwise, generate a data normal state signal;

[0088] It should be noted that the abnormal threshold is obtained by the experimenter based on the historical data distribution law and analyzing the distribution characteristics of the data in the current experiment. At the same time, the abnormal threshold is dynamically updated;

[0089] Specifically, to enable the adjusted connection rule, the specific steps are as follows:

[0090] Collect the deviation rate and the change ratio of the alternative connection quantity;

[0091] The deviation rate is used to measure the difference degree between the quantity of the entities to be screened out and the abnormal threshold. Its acquisition logic is to calculate the deviation rate P by subtracting the abnormal threshold from the quantity of the data correlation degree coefficients of the entities marked as entities to be screened out and then taking the ratio with the abnormal threshold;

[0092] The change ratio of the alternative connection quantity reflects the fluctuation degree of the current alternative connection quantity compared with the average alternative connection quantity and is used to measure the stability of the alternative connection quantity. Its acquisition logic is to calculate the change ratio R of the alternative connection quantity by subtracting the historical average alternative connection quantity from the current alternative connection quantity and then taking the ratio with the historical average alternative connection quantity;

[0093] Among them, the historical average alternative connection quantity can be statistically obtained based on the historical database or the historical average alternative connection quantity determined based on different time domains, etc., which will not be elaborated here;

[0094] Substitute the deviation rate and the change ratio of the number of alternative connections into the weighted formula for calculation to obtain the adjustment ratio, and add the corresponding connection rules: ;

[0095] In the formula, is the adjustment ratio, and respectively represent the influence weights of the control deviation rate and the change ratio of the number of alternative connections on the adjustment ratio;

[0096] Among them, , ensuring that the adjustment ratio will not be over-corrected due to excessive weight;

[0097] Multiply the adjustment ratio by the number of alternative connections to obtain the standard number of connections B;

[0098] Furthermore, for the data association degree coefficient quantity of the data marked as the entity to be screened is less than the abnormal threshold, the default adjustment ratio is 1. At this time, the default number of alternative connections is the standard number of connections;

[0099] S3: Obtain the standard number of connections and the number of alternative connections, compare them, generate a connection number adjustment plan, sort each data association degree coefficient in the alternative connections related to the individual according to the numerical value, and delete them from small to large until the number of alternative connections reaches the standard number of connections to obtain the connection relationship of the individual;

[0100] Compare the standard number of connections with the number of alternative connections to obtain the following results:

[0101] If the standard number of connections is the same as the number of alternative connections, no subsequent processing is required. Directly use the number of alternative connections as the standard number of connections to complete the adjustment of the connection relationship of the individual;

[0102] If there is a difference between the standard number of connections and the number of alternative connections, generate a connection number adjustment plan, specifically:

[0103] Sort each data association degree coefficient in the alternative connections related to the individual according to the numerical value and integrate them into a set expressed as: ; Among them, is the maximum value of the data association degree coefficient in the alternative connections, is the second maximum value of the data association degree coefficient in the alternative connections, is the minimum value of the data association degree coefficient in the alternative connections; m is the total number of each data association degree coefficient in the alternative connections related to the individual;

[0104] Subtract the standard connection quantity from the alternative connection quantity to obtain the difference connection quantity. Delete the set of data correlation degree coefficients in the alternative connection quantity in ascending order, and the adjusted set of alternative connection quantities is as follows: ; where is the data correlation degree coefficient in the Bth alternative connection quantity;

[0105] At this time, the alternative connection quantity is the same as the standard connection quantity. Connect the connection relationships of the individuals according to the adjusted alternative connection quantity;

[0106] Through multi-dimensional feature collection and logistic regression model calculation, the present invention accurately obtains the data correlation degree coefficient. By comparing the correlation degree coefficient with a preset threshold, it is ensured that the selected associated entities have a high degree of correlation, thereby improving the accuracy of data connection. At the same time, the alternative connection quantity and the number of screened entities are counted, and the connection rules are dynamically adjusted through weighted calculation of the deviation rate and the change ratio, ensuring the scientific nature of the standard connection quantity. Through sorting and optimization of the correlation degree coefficients in the alternative connection quantity, fine control of data connection is achieved, avoiding both the increase of invalid connections and improving the efficiency of data utilization.

[0107] Embodiment 2

[0108] In Embodiment 1 of the present invention, it is mainly illustrated by way of example that through multi-dimensional feature collection and logistic regression model calculation, the data correlation degree coefficient is accurately obtained. By comparing the correlation degree coefficient with a preset threshold, it is ensured that the selected associated entities have a high degree of correlation, thereby improving the accuracy of data connection. At the same time, the alternative connection quantity and the number of screened entities are counted, and the connection rules are dynamically adjusted through weighted calculation of the deviation rate and the change ratio, and the operation strategy of sorting and optimization of the correlation degree coefficients in the alternative connection quantity; however, in Embodiment 1, only starting from the connection relationships of individuals, the processing and sorting of the connection relationships are completed. However, it is impossible to accurately adjust the data correlation in the dynamic change process, reducing the flexibility and stability of the connection rules and increasing the latency of anomaly detection; in view of the above problems, Embodiment 2 of the present invention is further refined;

[0109] S4: Obtain the connection relationships of individuals to obtain directly connected entities and indirectly connected entities, record the time stamp sequence of interaction events between the individuals and each connected entity, and obtain the time correlation degree and data sparsity;

[0110] Specifically, directly connected entities refer to the set of entities that have a one-layer relationship with the individual, and indirectly connected entities refer to the set of entities that are connected to the individual through multiple relationship chains. The specific explanations of direct and indirect have been described in the above-mentioned description of the number of entity interactions in Embodiment 1 and will not be elaborated here;

[0111] The acquisition logic of time correlation is to extract the interaction timestamp sequence of an individual with each directly or indirectly connected entity through the connection relationship of the individual, calculate the standard deviation of the timestamp sequence, introduce a normalization formula to map the time correlation to the interval [0, 1], and obtain the time correlation;

[0112] It should be noted that the smaller the standard deviation, the more concentrated the time distribution and the higher the time correlation;

[0113] Among them, the introduced normalization formula is expressed as: ;

[0114] In the formula, is the time correlation between the individual and each connected entity, is the standard deviation of the timestamp sequence, is the maximum standard deviation of the connection time distribution in the entire knowledge graph;

[0115] Specifically, for the acquisition rule of Embodiment 2, it is set based on the detection of the individual in Embodiment 1. That is, after processing the individual connection relationship and until the next detection of the individual, the acquisition of Embodiment 2 is enabled to adjust the calculation weight of the data correlation degree coefficient and complete the dynamic closed-loop of data anomaly detection in the knowledge graph;

[0116] The acquisition logic of data sparsity is to count the number of directly and indirectly connected entities of an individual, calculate the ratio of the total number of connections of all entities to the number of all entities in the knowledge graph to obtain the average number of connections of entities in the knowledge graph, and calculate the ratio of the number of connections of the individual to the average number of connections of entities in the knowledge graph to obtain the data sparsity;

[0117] Optionally, the experimenter can consider the local connection density of the category to which the individual belongs, and perform weighted calculation on the overall sparsity of the knowledge graph and the average sparsity in the category to which the individual belongs to determine the correction value of the data sparsity;

[0118] S5: Use fuzzy logic to determine the calculation weight result of the data correlation degree coefficient based on the time correlation and data sparsity;

[0119] For example, "High", "Low", "Medium" for time correlation, and "Sparse", "Dense", "Moderate" for data sparsity;

[0120] Formulate a set of fuzzy rules to describe the influence of different input variables on the output variable. The definition of the rules can be based on professional knowledge or obtained through data analysis and experiments. For example:

[0121] Mark the time correlation as X, the data sparsity as U, and the calculation weight result of the data correlation degree coefficient as C_results;

[0122] Then it can be defined as: Rule 1: IF (X is High) AND (U is Dense) THEN (C_results is High) Rule 2: IF (U is Low) AND (U is Sparse) THEN (C_results is Low) ...

[0123] Perform fuzzy reasoning according to the fuzzy rules to determine the calculation weight result of the data correlation degree coefficient;

[0124] It should be noted that the division of the fuzzy set can be adjusted according to the actual situation. For example, although three fuzzy sets are taken as an example in this embodiment, in fact, the time correlation and data sparsity can be divided into more than three sets to facilitate more accurate adjustment according to different time spans.

[0125] Furthermore, for the judgment of high, medium, and low of the time correlation and data sparsity, thresholds can be set according to the actual situation for judgment. For example, when the time correlation exceeds 75%, it is marked as "High", and when the data sparsity is higher than 70%, it is marked as "Sparse", etc., which will not be elaborated here.

[0126] According to the calculation weight result of the data correlation degree coefficient, substitute it into the next round of step S1 to complete the dynamic closed-loop of the data anomaly detection of the knowledge graph;

[0127] The present invention obtains the direct connection entities and indirect connection entities by acquiring the individual connection relationships, records the interaction event timestamp sequences of the individual and each connection entity, obtains the time correlation and data sparsity, formulates a set of fuzzy rules for fuzzy reasoning, determines the calculation weight result of the data correlation degree coefficient, improves the dynamic adaptation ability of the data correlation, enhances the flexibility and stability of the connection rules, and increases the accuracy of anomaly detection.

[0128] The above formulas are all dimensionless and take their numerical values for calculation. The formulas are obtained by collecting a large amount of data for software simulation to obtain a formula closest to the real situation. The preset parameters in the formulas are set by those skilled in the art according to the actual situation.

[0129] The above embodiments can be implemented in whole or in part by software, hardware, firmware, or any combination thereof. When implemented using software, the above embodiments can be implemented in whole or in part in the form of a computer program product. The computer program product includes one or more computer instructions or computer programs. When the computer instructions or computer programs are loaded or executed on a computer, the processes or functions described in the embodiments of the present application are generated in whole or in part. The computer can be a general-purpose computer, a special-purpose computer, a computer network, or other programmable devices. The computer instructions can be stored in a computer-readable storage medium, or transmitted from one computer-readable storage medium to another. For example, the computer instructions can be transmitted from one website, computer, server, or data center to another website, computer, server, or data center in a wired or wireless (such as infrared, wireless, microwave, etc.) manner. The computer-readable storage medium can be any available medium that can be accessed by a computer, or a data storage device such as a server or data center that contains one or more collections of available media. The available media can be magnetic media (such as floppy disks, hard disks, magnetic tapes), optical media (such as DVDs), or semiconductor media. The semiconductor media can be a solid-state drive.

[0130] It should be understood that in various embodiments of the present application, the sequence numbers of the above processes do not imply the order of execution. The order of execution of each process should be determined by its function and internal logic, and should not constitute any limitation to the implementation process of the embodiments of the present application.

[0131] Those of ordinary skill in the art can realize that the units and algorithm steps of the examples described in combination with the embodiments disclosed herein can be implemented by electronic hardware, or a combination of computer software and electronic hardware. Whether these functions are executed in a hardware or software manner depends on the specific application and design constraints of the technical solution. Professional technicians can use different methods to implement the described functions for each specific application, but such implementation should not be considered to exceed the scope of the present application.

[0132] Those skilled in the art can clearly understand that for the convenience and brevity of description, the specific working processes of the systems, devices, and units described above can refer to the corresponding processes in the foregoing method embodiments, and will not be repeated here.

[0133] In several embodiments provided in the present application, it should be understood that the disclosed systems, devices, and methods can be implemented in other ways. For example, the device embodiments described above are merely illustrative. For example, the division of the units is only a logical function division. In actual implementation, there may be other division methods. For example, multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the displayed or discussed couplings or direct couplings or communication connections to each other can be through some interfaces. The indirect couplings or communication connections of the devices or units can be in electrical, mechanical, or other forms.

[0134] The units described as separate components may or may not be physically separated. The components shown as units may or may not be physical units, that is, they can be located in one place or distributed to multiple network units. Some or all of the units can be selected according to actual needs to achieve the purpose of the solution of this embodiment.

[0135] In addition, in each embodiment of the present application, the functional units can be integrated in a processing unit, or each unit can exist physically alone, or two or more units can be integrated in one unit.

[0136] If the functions are implemented in the form of software function units and sold or used as independent products, they can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present application, in essence, or the part that contributes to the prior art or part of this technical solution can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions for causing a computer device (which can be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the methods described in each embodiment of the present application. The foregoing storage medium includes: various media such as USB flash drives, mobile hard disks, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical discs that can store program codes.

[0137] As described above, the above are only the specific implementation manners of the present application, but the protection scope of the present application is not limited thereto. Any person skilled in the art within the technical scope disclosed in the present application can easily think of changes or substitutions, which should all be covered within the protection scope of the present application. Therefore, the protection scope of the present application should be subject to the protection scope of the claims.

Claims

1. A data anomaly detection method based on knowledge graph, characterized by: include: S1: Collect the semantic similarity, entity attribute matching degree, relationship connection level and entity interaction number of each entity, and obtain the data association coefficient by substituting it into logistic regression calculation; S2: Obtain the data association coefficient, take one entity as an individual, count the data association coefficients corresponding to all other entities related to it, and compare them with the preset association threshold, select other entities marked as the data association coefficients of the associated entities and count their number as the number of candidate connections, and at the same time count the number of screened entities and compare them with the abnormal threshold. According to the comparison results, enable the connection adjustment rule, collect the deviation rate and the change ratio of the number of candidate connections for weighted calculation to obtain the adjustment ratio, and determine the standard number of connections; S3: Obtain the standard connection number and the candidate connection number, compare them, generate a connection number adjustment plan, sort the data correlation coefficients in the candidate connection number related to the individual according to the numerical value, delete them from small to large, until the candidate connection number reaches the standard connection number, and obtain the connection relationship of the individual; S4: Obtain the connection relationship of individuals, obtain directly connected entities and indirectly connected entities, record the timestamp sequence of interaction events between individuals and each connected entity, and obtain time correlation and data sparsity; S5: Use fuzzy logic to determine the data association degree coefficient and calculate the weight result based on the time association degree and data sparsity.

2. The data anomaly detection method based on knowledge graph according to claim 1 is characterized in that: Through the unique identification of each entity, we get the entity's name, label, attribute value and annotation, as well as the entity's adjacent nodes and their connection relationships. We convert the entity and its information into vector form, use the pre-trained word embedding model to encode the entity's text description, generate entity vectors, and substitute the entity vectors into the cosine similarity to calculate the semantic similarity between the two entities. , represents the i-th alignment, then there are i aligned entities; Extract the attributes of each entity from the knowledge graph, define different matching rules according to the different attribute types, and compare whether the discrete and definite attribute values ​​are the same, that is, calculate the ratio of the number of matching attributes to the total number of attributes to obtain the discrete matching degree. For continuous numerical attributes, use the difference to calculate the continuous matching degree. Set different weights for discrete matching and continuous matching, and substitute them into weighted calculation to get the matching degree of entity attributes. ; By constructing the path in the knowledge graph, using breadth-first search to calculate the shortest path, starting from the starting entity node, expanding outward layer by layer until the target entity node is found, the path length is the number of relationship levels ; By distinguishing direct interactions from indirect interactions, and recording each time entities interact through connection relationships, it is recorded as a direct interaction. By finding all the paths that the current entity passes through, the number of paths is calculated. On each path, each time the current entity passes through, it is recorded as an indirect interaction. The direct interactions and indirect interactions are accumulated and calculated to get the number of entity interactions. ; The main semantic similarity, entity attribute matching degree, relationship connection level and entity interaction times are substituted into the logistic regression formula to calculate the data association degree coefficient.

3. The data anomaly detection method based on knowledge graph according to claim 2 is characterized in that: Calculate the data correlation coefficient corresponding to all other related entities; After obtaining the data correlation coefficients corresponding to all other entities, the data correlation coefficients are compared and analyzed with the continuously iterated correlation thresholds; If the data association degree coefficient is greater than or equal to the association threshold, other entities corresponding to the data association degree coefficient are marked as associated entities, and an association signal is generated; If the data association degree coefficient is less than the association threshold, other entities corresponding to the data association degree coefficient are marked as screened entities, and a screened signal is generated; The correlation coefficient of the data marked as the associated entity is counted as the number of candidate connections; The data correlation coefficients marked as screened entities are counted and the quantity is determined, and compared with the abnormal threshold; If the number of data association degree coefficients marked as screened entities is greater than or equal to the abnormal threshold, a data abnormality signal is generated, the screened entities are deleted, and the adjustment connection rules are enabled. Otherwise, a data normality signal is generated.

4. The data anomaly detection method based on knowledge graph according to claim 3 is characterized in that: Enable the connection adjustment rules. The specific steps are as follows: The acquisition deviation rate and the change ratio of the number of alternative connections; The deviation rate is used to measure the difference between the number of screened entities and the abnormal threshold. The logic of obtaining it is to calculate the deviation rate P by subtracting the abnormal threshold from the number of data association coefficients marked as screened entities and then comparing the coefficients with the abnormal threshold. The change ratio of the number of alternative connections reflects the degree of fluctuation of the current number of alternative connections compared to the average number of alternative connections. It is used to measure the stability of the number of alternative connections. Its acquisition logic is to subtract the historical average number of alternative connections from the current number of alternative connections and calculate the ratio of change R of the number of alternative connections with the historical average number of alternative connections.

5. The method for data anomaly detection based on knowledge graph according to claim 4 is characterized in that: Substitute the deviation rate and the change ratio of the number of candidate connections into the weighted formula to obtain the adjustment ratio, and add the corresponding connection rules: ; In the formula, To adjust the ratio, and are the weights of the control deviation rate and the change ratio of the number of alternative connections on the adjustment ratio; in, , ensuring that the adjustment ratio does not lead to over-correction due to excessive weight; The adjustment ratio is multiplied by the number of alternative connections to obtain the standard number of connections B; For data whose correlation coefficient number is less than the abnormal threshold and is marked as a screened entity, the default adjustment ratio is 1, and the default number of candidate connections is the standard number of connections.

6. The method for data anomaly detection based on knowledge graph according to claim 5, characterized in that: If the number of standard connections is the same as the number of candidate connections, the number of candidate connections is directly used as the number of standard connections to complete the adjustment of the individual connection relationship; If there is a difference between the standard number of connections and the number of alternative connections, a connection number adjustment plan is generated. The standard number of connections is subtracted from the number of alternative connections to obtain the difference number of connections. The set of data association degree coefficients in the number of alternative connections is deleted from small to large so that the number of alternative connections is the same as the standard number of connections. The individual connection relationships are connected according to the adjusted number of alternative connections.

7. The method for data anomaly detection based on knowledge graph according to claim 6, characterized in that: Through the individual connection relationship, the interaction timestamp sequence with each directly or indirectly connected entity is extracted, the standard deviation of the timestamp sequence is calculated, and the normalization formula is introduced to map the time correlation to the [0,1] interval to obtain the time correlation degree; Count the number of directly connected and indirectly connected entities of an individual, calculate the ratio of the total number of connections of all entities to the number of all entities in the knowledge graph to obtain the average number of connections of entities in the knowledge graph, and calculate the ratio of the number of connections of an individual to the average number of connections of entities in the knowledge graph to obtain data sparsity.

8. The method for data anomaly detection based on knowledge graph according to claim 7, characterized in that: The time correlation and data sparsity are defined as input variables and divided into different fuzzy sets respectively; The weighted result of the data association coefficient calculation is defined as the output variable and divided into fuzzy sets; Formulate fuzzy rules to describe the influence of time correlation and data sparsity on the weight calculation results of data correlation coefficient; Fuzzy reasoning is performed according to fuzzy rules to determine the data association degree coefficient and calculate the weight result.

Citation Information

Patent Citations

  • Knowledge graph determination method and device, equipment and storage medium

    CN115774789A

  • Knowledge graph correction method based on data correlation and causal mode

    CN117709449A

  • Method and device for judging relevance between service entities

    CN117852641A

  • Large model multi-round dialogue method and device, computer equipment and storage medium

    CN118036756A

  • Visual development system for knowledge graph

    CN118093895A

Cited By

  • Software system development method based on multi-modal AI large model

    CN120762638A

  • Abnormality detection method and device for image data and storage medium

    CN120852791A