A quality optimization method, device and equipment for intelligence knowledge graph
By pre-processing, abnormal detection and confidence calculation of intelligence data, combined with the correction of GPT large model, the problems of poor environmental perception and low efficiency of traditional methods when dealing with the dynamic and time-consuming needs of intelligence data are solved, and efficient optimization of knowledge graphs and improvement of intelligence analysis are achieved.
Patent Information
- Application Number
- CN202510437339.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-09
- Publication Date
- 2025-06-06
- Estimated Expiration
- 2045-04-09
AI Technical Summary
When traditional knowledge graph quality optimization methods respond to the dynamic, multi-source heterogeneity and high timeliness of intelligence data, there are problems such as poor environmental perception, high error rate and low efficiency.
By acquiring multi-source intelligence data, performing pre-processing and preliminary anomaly detection, calculating confidence to determine decision-making plans, performing secondary anomaly detection and using GPT large model for abnormal correction, and optimizing the knowledge graph.
It realizes intelligent processing of intelligence knowledge graphs, improves data quality, enhances the accuracy and reliability of intelligence analysis, adapts to the dynamic changes and high-time efficiency of intelligence data, reduces manual intervention and improves processing efficiency.
Smart Images

Figure CN119961761B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of knowledge graph processing technology, and in particular to a quality optimization method, device and equipment for an intelligence knowledge graph. Background Art
[0002] In the field of intelligence analysis, knowledge graphs, as core tools, present entities and their relationships in a structured manner, providing solid support for key tasks such as intelligent retrieval, threat correlation analysis, and event reasoning. At present, traditional knowledge graph quality optimization methods mainly rely on rule engines, statistical learning models, and manual review. However, as the complexity and scale of intelligence data continue to expand, these traditional methods have gradually exposed many insurmountable limitations when dealing with data dynamics, multi-source heterogeneity, and high timeliness requirements.
[0003] The defects and shortcomings of existing technologies mainly include: extremely limited environmental perception capabilities, the system is unable to dynamically perceive subtle changes in the distribution of intelligence data in real time. For example, when faced with new entities or emergencies, it is often difficult to respond in a timely manner, resulting in a delay of up to 72 hours in the identification of new entities; the decision-making mechanism is rigid and inflexible, the rule system lacks the flexibility of adaptive adjustment, and the misjudgment rate is high; the execution efficiency is low, and it is unable to cope with large-scale data. Traditional methods have slow processing speeds and are unable to cope with the high throughput requirements of intelligence data. Summary of the invention
[0004] In view of this, the purpose of the present invention is to propose a quality optimization method, device and equipment for intelligence knowledge graph, aiming to solve the problems of poor environmental perception, high error rate and low efficiency in traditional methods when dealing with the dynamic nature, multi-source heterogeneity and high timeliness requirements of intelligence data.
[0005] To achieve the above object, the present invention provides a method for optimizing the quality of an intelligence knowledge graph, the method comprising:
[0006] Acquire multi-source intelligence data and pre-process them to obtain initial triplet data;
[0007] Performing preliminary anomaly detection on the initial triplet data to obtain temporal feature anomalies and semantic feature anomalies;
[0008] Calculating the confidence of the initial triple data corresponding to the abnormal time series features and the abnormal semantic features, and determining the optimal decision plan according to the confidence;
[0009] According to the optimal decision-making scheme, secondary anomaly detection is performed on the initial triple data corresponding to the abnormal time series features and the abnormal semantic features to obtain abnormal triple data;
[0010] The GPT big model is used to perform anomaly correction on the abnormal triple data to obtain an updated optimized knowledge graph.
[0011] Preferably, the performing preliminary anomaly detection on the initial triplet data to obtain time series feature anomalies and semantic feature anomalies includes:
[0012] Extracting features from the initial triplet data to obtain temporal features and semantic features;
[0013] Performing anomaly detection on the time series feature based on the Prophet algorithm to obtain anomalies of the time series feature;
[0014] The cosine similarity between the semantic feature and the stock triple data is calculated. If it is lower than a preset threshold, the corresponding semantic feature is marked as abnormal, and the semantic feature abnormality is obtained.
[0015] Preferably, the confidence includes a first confidence and a second confidence; the confidence of the initial triple data corresponding to the abnormal temporal feature and the abnormal semantic feature is calculated, and the optimal decision plan is determined according to the confidence, including:
[0016] Calculating using the preset rule confidence and model confidence respectively to obtain a first confidence and a second confidence;
[0017] When it is determined that the first confidence level is greater than the second confidence level, the rule library is called to process as the optimal decision solution; otherwise, the model library is called to process as the optimal decision solution.
[0018] Preferably, the calculation using the preset rule confidence and model confidence respectively to obtain the first confidence and the second confidence includes:
[0019] The first confidence is obtained by calculating rule confidence = total number of rule triggering times / number of successful rule matching times × rule complexity attenuation factor, wherein the value range of the rule complexity attenuation factor is 0.8 to 1.0;
[0020] The second confidence is obtained by calculating the model confidence = 0.4×XGBootst_AUC+0.6×GAT_F1_score, where:
[0021] , where TPR i+1 Indicates the true rate of the current point i, TPR i Represents the false positive rate of the current point i, FPR i+1 Indicates the true rate of the next point of the current point i, FPR i Indicates the false positive rate of the next point of the current point i;
[0022] ,
[0023] Precision = TP / (TP + FP), Recall = TP / (TP + FN), where TP represents true positive examples, FP represents false positive examples, and FN represents false negative examples.
[0024] Preferably, performing secondary anomaly detection on the initial triplet data corresponding to the abnormal temporal feature and the abnormal semantic feature according to the optimal decision scheme to obtain abnormal triplet data includes:
[0025] The pre-built rule base is used to match the SHACL constraint rules on the initial triple data corresponding to the timing feature anomaly and the semantic feature anomaly, and the matched data is used as the abnormal triple data.
[0026] Preferably, performing secondary anomaly detection on the initial triplet data corresponding to the abnormal temporal feature and the abnormal semantic feature according to the optimal decision scheme to obtain abnormal triplet data includes:
[0027] The pre-trained XGBoost classifier and GAT model are used to detect the initial triple data corresponding to the time series feature anomaly and the semantic feature anomaly, and the data with existing relationship anomalies is used as the abnormal triple data.
[0028] Preferably, the use of the GPT large model to perform anomaly correction on the abnormal triple data to obtain an updated optimized knowledge graph includes:
[0029] Generate multiple candidate correction suggestions based on the abnormal triple data using the GPT large model;
[0030] Each candidate correction suggestion is verified by a preset SecurityBert model to obtain a corresponding correction rationality result;
[0031] The candidate correction suggestion corresponding to the correction rationality result being greater than the similarity threshold is selected to perform anomaly correction on the abnormal triple data to obtain the optimized knowledge graph.
[0032] Preferably, the method further comprises:
[0033] After the abnormal triplet data is corrected using the GPT large model, the corrected triplet data is added to the training data set to update the iterative GAT model.
[0034] To achieve the above object, the present invention also provides a quality optimization device for an intelligence knowledge graph, the device comprising:
[0035] A data acquisition unit is used to acquire multi-source intelligence data and perform preprocessing to obtain initial triplet data;
[0036] A first detection unit is used to perform preliminary anomaly detection on the initial triple data to obtain time series feature anomalies and semantic feature anomalies;
[0037] A calculation unit, used to calculate the confidence of the initial triple data corresponding to the abnormal time series feature and the abnormal semantic feature, and determine the optimal decision plan according to the confidence;
[0038] A second detection unit is used to perform secondary anomaly detection on the initial triple data corresponding to the abnormal time series feature and the abnormal semantic feature according to the optimal decision scheme to obtain abnormal triple data;
[0039] The correction unit is used to use the GPT large model to perform anomaly correction on the abnormal triple data to obtain an updated optimized knowledge graph.
[0040] In order to achieve the above-mentioned objectives, the present invention also proposes a quality optimization device for an intelligence knowledge graph, comprising a processor, a memory, and a computer program stored in the memory, wherein the computer program is executed by the processor to implement the steps of a quality optimization method for an intelligence knowledge graph as described in the above-mentioned embodiment.
[0041] In order to achieve the above objectives, the present invention also proposes a computer-readable storage medium, on which a computer program is stored, and the computer program is executed by a processor to implement the steps of a quality optimization method of an intelligence knowledge graph as described in the above embodiment.
[0042] In order to achieve the above objectives, the present invention also proposes a computer program product, including a computer program / instruction, which, when executed by a processor, implements the steps of a quality optimization method for an intelligence knowledge graph as described in the above embodiment.
[0043] Beneficial effects:
[0044] The above scheme obtains multi-source intelligence data and pre-processes it to obtain initial triple data, performs preliminary anomaly detection, time series feature anomalies, and semantic feature anomalies, and then determines the optimal decision plan based on the calculated confidence level. The abnormal data is secondary detected to obtain abnormal triple data, and finally the GPT large model is used to correct it to obtain the optimized knowledge graph. This method realizes the intelligent processing of the entire process from data acquisition, anomaly detection, decision optimization to anomaly correction, effectively improves the quality of the intelligence knowledge graph, enhances the accuracy and reliability of intelligence analysis, can adapt to the dynamic changes and high timeliness requirements of intelligence data, reduces manual intervention, and improves processing efficiency.
[0045] By using the Prophet algorithm to detect anomalies in time series features and calculating the cosine similarity between semantic features and stock data to determine semantic feature anomalies, the accuracy and reliability of anomaly detection are improved. In terms of confidence calculation, rule confidence and model confidence are used for calculation respectively, and the corresponding rule library or model library is called for processing based on the comparison results of the two, realizing the intelligent switching of rules and models, taking into account both the stability of rules and the flexibility of models, optimizing the formulation of decision-making plans, further improving the effect of anomaly detection and processing, and making the entire quality optimization method more intelligent and efficient.
[0046] By using the pre-built rule library to match SHACL constraint rules, abnormal triplet data can be accurately located based on a mature rule system, ensuring the standardization and consistency of the detection results; using the pre-trained XGBoost classifier and GAT model for detection, the advantages of machine learning models in complex relationship recognition and classification are fully utilized, and potential complex relationship anomalies can be effectively discovered. The two methods complement each other, ensuring the comprehensiveness of detection, while improving the efficiency and accuracy of detection, providing a solid foundation for subsequent anomaly correction.
[0047] Multiple candidate correction suggestions are generated through the GPT large model and verified with the help of the SecurityBert model to ensure the rationality and accuracy of the correction suggestions. At the same time, the similarity threshold is introduced as a screening criterion to further ensure the quality of the correction. The corrected triple data is added to the training set, and the GAT model is incrementally updated and iterated, so that the model can continuously adapt to new intelligence data and changing threat scenarios, improve the system's continuous learning ability and intelligence level, and ensure the long-term effectiveness and adaptability of the entire quality optimization method. BRIEF DESCRIPTION OF THE DRAWINGS
[0048] In order to more clearly illustrate the embodiments of the present invention or the technical solutions in the prior art, the drawings required for use in the embodiments or the description of the prior art will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying creative work.
[0049] Figure 1 A flowchart of a method for optimizing the quality of an intelligence knowledge graph provided in accordance with an embodiment of the present invention.
[0050] Figure 2 A schematic diagram of the overall process of quality optimization of an intelligence knowledge graph is provided for one embodiment of the present invention.
[0051] Figure 3 A schematic diagram of the structure of a quality optimization device for an intelligence knowledge graph provided by one embodiment of the present invention.
[0052] The realization of the purpose, functional features and advantages of the invention will be further explained in conjunction with embodiments and with reference to the accompanying drawings. DETAILED DESCRIPTION
[0053] In order to make the purpose, technical solutions and advantages of the embodiments of the present invention clearer, the technical solutions in the embodiments of the present invention will be clearly and completely described below in conjunction with the drawings in the embodiments of the present invention. Obviously, the described embodiments are part of the embodiments of the present invention, rather than all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without making creative work are within the scope of protection of the present invention. Therefore, the following detailed description of the embodiments of the present invention provided in the drawings is not intended to limit the scope of the invention claimed for protection, but merely represents the selected embodiments of the present invention. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without making creative work are within the scope of protection of the present invention.
[0054] The present invention is described in detail below with reference to the embodiments.
[0055] Reference Figure 1 The figure is a flow chart of a quality optimization method of an intelligence knowledge graph provided by an embodiment of the present invention.
[0056] In this embodiment, the method includes:
[0057] S11, obtain multi-source intelligence data and preprocess it to obtain initial triplet data.
[0058] In this embodiment, the multi-source heterogeneous triple data are integrated and cleaned, and stored in a temporary knowledge graph, which may include historical triple data or stock triple data. Specifically, data access: supports HTTP / HTTPS / AMQP protocols, and is compatible with multi-source intelligence data such as triple data streams, social media texts, and graph databases (such as neo4j); stream and batch processing: including the use of the Flink engine to parse data in real time (delay <50ms), and the format converter automatically processes JSON-LD / RDF / CSV formats. Data cleaning: align entities in heterogeneous data based on the TransE embedding model to achieve entity alignment; use the SimHash algorithm to eliminate duplicate triples, and unify entity naming to achieve deduplication and standardization.
[0059] S12, performing preliminary anomaly detection on the initial triplet data to obtain temporal feature anomalies and semantic feature anomalies.
[0060] Furthermore, in step S12, the initial triplet data is subjected to preliminary anomaly detection to obtain time series feature anomalies and semantic feature anomalies, including:
[0061] S12-1, extracting features from the initial triplet data to obtain temporal features and semantic features;
[0062] S12-2, performing anomaly detection on the time series feature based on the Prophet algorithm to obtain anomalies of the time series feature;
[0063] S12-3, calculating the cosine similarity between the semantic feature and the stock triple data, if it is lower than a preset threshold, marking the corresponding semantic feature as abnormal, and obtaining the semantic feature abnormality.
[0064] In this embodiment, by performing time series anomaly detection and semantic anomaly detection on the initial triple data, time series feature anomaly and semantic feature anomaly are obtained. Specifically:
[0065] Time series anomaly detection: Use sliding windows (dynamically adjust window size) and the Prophet algorithm to monitor data stream mutations, such as detecting a surge in entities related to "conflict events in a certain region" (threshold: >500 new additions within 1 hour).
[0066] 1) Time series feature extraction method: The initial triple data is predicted by the constructed Prophet algorithm to obtain time series features, where the Prophet algorithm is obtained after modeling based on historical triple data or stock triple data. Specifically, the historical triple data or stock triple data is segmented according to a fixed or dynamic sliding window, and the statistical features of each window are extracted. Then, the Prophet algorithm is used to model the time series, extract the time pattern (such as trend, seasonality, mutation point), and finally calculate the frequency, time interval distribution, sudden increase rate, etc. of the event to obtain the time series features. For example: count the number of times a certain entity appears within 1 hour to form a time series feature.
[0067] 2) Time series anomaly detection: Detect according to the set sliding window to obtain the actual value (such as counting the number of new events of an entity within 1 hour is 500), and use the Prophet algorithm to predict the trend of the sliding window to obtain the predicted value. , where y actual is the actual value, y pred is the predicted value. C is set according to actual needs. The value range of C is 100-200. , it means the deviation is large and is judged as abnormal.
[0068] Semantic anomaly detection, including:
[0069] 1) Semantic feature extraction: BERT is first used to extract the text semantic features of the initial triple data, and the text semantic features are converted into high-dimensional vectors. At the same time, TransE is used for knowledge graph embedding to capture the structural information of entity-relationships. Finally, the BERT semantic vector and the TransE structural vector are fused to form a joint embedding representation.
[0070] 2) Anomaly detection: Compare the cosine similarity of the average values of all existing triple vectors in the initial triple data (entity-relationship-entity) corresponding to the semantic features and the existing triple data (the system's original triple graph data or historical triple data). If it is lower than the threshold (such as 0.7), it is marked as an anomaly.
[0071] S13, calculating the confidence of the initial triple data corresponding to the temporal feature anomaly and the semantic feature anomaly, and determining the optimal decision plan according to the confidence.
[0072] Further, the confidence includes a first confidence and a second confidence; in step S13, the confidence of the initial triple data corresponding to the abnormal temporal feature and the abnormal semantic feature is calculated, and the optimal decision plan is determined according to the confidence, including:
[0073] S13-1, respectively using the preset rule confidence and model confidence to perform calculations to obtain a first confidence and a second confidence;
[0074] S13-2, when it is determined that the first confidence level is greater than the second confidence level, the rule library is called to process as the optimal decision solution; otherwise, the model library is called to process as the optimal decision solution.
[0075] Further, in step S13-1, the calculation is performed using the preset rule confidence and model confidence respectively to obtain the first confidence and the second confidence, including:
[0076] S13-1-1, calculate using rule confidence = total number of rule triggering times / number of successful rule matching times × rule complexity attenuation factor to obtain the first confidence, wherein the value range of the rule complexity attenuation factor is 0.8 to 1.0;
[0077] S13-1-2, using model confidence = 0.4 × XGBootst_AUC + 0.6 × GAT_F1_score to calculate, to obtain the second confidence, wherein,
[0078] , where TPR i+1 Indicates the true rate of the current point i, TPR i Represents the false positive rate of the current point i, FPR i+1 Indicates the true rate of the next point (i+1) of the current point, FPR i Indicates the false positive rate of the next point (i+1) of the current point;
[0079] ,
[0080] Precision = TP / (TP + FP), Recall = TP / (TP + FN), where TP represents true positive examples, FP represents false positive examples, and FN represents false negative examples.
[0081] Reference Figure 2 As shown, in this embodiment, the rule confidence and model confidence of the input abnormal triple data are calculated according to the preset confidence formula, and the size of the two is further determined to achieve intelligent switching between the rule base and the model base. That is,
[0082] If the rule confidence > model confidence, the SHACL constraint rule matching is triggered, and the constraint type (integrity / consistency, etc.) and the matched triples are output;
[0083] If the model confidence > rule confidence, call the XGBoost classifier and GAT model for multi-level classification and output high-risk abnormal triple data;
[0084] If the two are equal, start the hybrid mode, execute the rule base (i.e. trigger the SHACL constraint rule matching) and the model base (i.e. call the XGBoost classifier and the GAT model) in parallel, and take the union result of the two.
[0085] Rule confidence: Calculated based on the historical matching accuracy of SHACL rules. The formula is:
[0086] Rule confidence = total number of rule triggers / number of successful rule matches × rule complexity attenuation factor,
[0087] The rule complexity attenuation factor (value range 0.8~1.0) is used to balance the potential misjudgment risk of complex rules. For example, among 5 pieces of data, 3 of them are matched by the SHACL constraint rule, and the rule complexity attenuation factor is 0.8, then the corresponding rule confidence is: 3 / 5×0.8=0.48.
[0088] Model confidence: Dynamically adjusted according to the classification performance of XGBoost and GAT models. The formula is:
[0089] Model confidence = 0.4×XGBootst_AUC+0.6×GAT_F1_score, where AUC (Area Under the ROC Curve) in XGBootst_AUC measures the ability of the XGBootst model to distinguish between positive and negative samples by calculating the area under the ROC curve. The specific test process is as follows:
[0090] 1. Train the XGBoost model;
[0091] 2. Make predictions for the test set and output the probability that the sample belongs to the positive class;
[0092] 3. Calculate the ROC curve;
[0093] 4. Use the trapezoidal method to calculate the area under the ROC curve as follows:
[0094] The AUC under the ROC curve is calculated using the trapezoidal rule. The ROC curve drawing process includes:
[0095] 1) Set multiple thresholds between 0 and 1: 0.1, 0.2, 0.3...;
[0096] 2) Calculate TPR (True Positive Rate) and FPR (False Positive Rate) under different thresholds. TPR = TP / (TP+FN) represents the proportion of all actual positive examples that are correctly classified as positive examples, and FPR = FP / (FP+TN) represents the proportion of all actual negative examples that are misclassified as positive examples.
[0097] 3) With FPR as the horizontal axis and TPR as the vertical axis, plot the points under different thresholds, and connect these points to get the ROC curve;
[0098] 4) Use the trapezoidal method to calculate AUC (which is actually the area under the ROC curve);
[0099] In the ROC curve, each point (TPR i , FPR i ) represents the false positive rate (FPR) and true positive rate (TPR) under a threshold, and i represents the index of the current point. When the trapezoidal method calculates the area under the ROC curve, two adjacent points (FPR i , TPR i ) and (FPR i+1 , TPR i+1 ) forms a trapezoid.
[0100] The calculation formula is:
[0101] , where TPR is the true positive rate of the test set predicted by the trained XGBootst model, and FPR is the false positive rate; GAT_F1_score is the harmonic mean of precision and recall, which comprehensively measures the accuracy and coverage of the GAT model. Its calculation formula is:
[0102] ,
[0103] Precision = TP / (TP + FP),
[0104] Recall = TP / (TP + FN), where TP (True Positive): actually a positive example that is correctly classified as a positive example; FN (False Negative): actually a positive example that is mistakenly classified as a negative example; FP (False Positive): actually a negative example that is mistakenly classified as a positive example; TN (True Negative): actually a negative example that is correctly classified as a negative example.
[0105] Furthermore, the construction of the rule base includes:
[0106] (1) Define intelligence domain ontology, including:
[0107] Entity: Intelligence report, threat, intelligence source, event, time;
[0108] Relationship: the source of the intelligence report, the threat described in the intelligence report, indicators related to the threat, entities involved in the intelligence, time of the incident, threat level (such as 1-5), credibility (1-100%), and related intelligence / incidents.
[0109] (2) Define SHACL constraint rules, including: data integrity (for example, each intelligence report must have a source), data consistency (for example, the time format must comply with ISO 8601), data type constraints (for example, the threat level must be an integer, 1-5), uniqueness constraints (for example, intelligence report IDs cannot be repeated), and custom business constraints (for example, events cannot be directly associated with threats).
[0110] (3) The generation of the rule base adopts a combination of top-down and bottom-up methods. On the one hand, the basic ontology and initial rule base defined based on the intelligence domain knowledge is a top-down approach, which is to define the intelligence domain ontology, including entities and relationships, and then manually enter them according to the SHACL constraint rules. On the other hand, data-driven discovery is used, that is, a bottom-up approach. First, the existing triple data is parsed to construct a graph structure, and then the similarity analysis of the graph structure is performed, which includes using Node2Vec to generate entity / relationship low-dimensional embedding vectors, and then using K-Means to cluster nodes and identify the structure of the same cluster. Finally, the graph neural network (GNN) model is trained, taking the graph structure of the same cluster as input and outputting potential SHACL constraint rules, so as to mine valuable rules from the data. After that, rule verification and conflict resolution are carried out, which are key steps in closed-loop optimization. Specifically, by counting the matching frequency and successful detection rate of each SHACL constraint rule in historical data, SHACL constraint rules with low coverage (such as less than 5%) are filtered out; the Apache Jena reasoning engine is used to detect rule conflicts. For example, rule A requires "event time must be filled in", while rule B allows "event time to be empty". For such conflicts, manual methods are prioritized over data-driven methods to ensure the accuracy and practicality of the rule base and improve the quality optimization effect of the intelligence knowledge graph.
[0111] Furthermore, the model library includes a pre-trained XGBoost classifier for quickly screening low-risk anomalies (92% accuracy) and a pre-trained GAT model (graph attention network) for detecting complex relationship anomalies.
[0112] S14, performing secondary anomaly detection on the initial triple data corresponding to the temporal feature anomaly and the semantic feature anomaly according to the optimal decision plan to obtain abnormal triple data.
[0113] Further, in step S14, the initial triplet data corresponding to the abnormal time series features and the abnormal semantic features are subjected to secondary abnormality detection according to the optimal decision scheme to obtain abnormal triplet data, including:
[0114] The pre-built rule base is used to match the SHACL constraint rules on the initial triple data corresponding to the timing feature anomaly and the semantic feature anomaly, and the matched data is used as the abnormal triple data.
[0115] Further, in step S14, the initial triplet data corresponding to the abnormal time series features and the abnormal semantic features are subjected to secondary abnormality detection according to the optimal decision scheme to obtain abnormal triplet data, including:
[0116] The pre-trained XGBoost classifier and GAT model are used to detect the initial triple data corresponding to the time series feature anomaly and the semantic feature anomaly, and the data with existing relationship anomalies is used as the abnormal triple data.
[0117] In this embodiment, the initial triple data corresponding to the abnormal timing features and semantic features are matched one by one with the SHACL constraint rules, and the matched SHACL constraint rules (such as integrity, consistency, etc.) and the triple data matched by the SHACL constraint rules in the initial triple data corresponding to the abnormal timing features and semantic features are obtained as abnormal triple data. Alternatively, the initial triple data corresponding to the abnormal timing features and semantic features are first subjected to preliminary screening by the XGBoost classifier (the low-risk anomalies of the above data are screened out by the pre-trained XGBoost classifier, and the low-risk abnormal data are ignored), and the remaining non-low-risk data are screened by the GAT model to obtain abnormal triple data with complex relationships.
[0118] S15, using the GPT big model to perform anomaly correction on the abnormal triple data to obtain an updated optimized knowledge graph.
[0119] Furthermore, in step S15, the abnormal triple data is corrected using the GPT large model to obtain an updated optimized knowledge graph, including:
[0120] S15-1, using the GPT large model to generate multiple candidate correction suggestions based on the abnormal triple data;
[0121] S15-2, verifying each candidate correction suggestion through a preset SecurityBert model to obtain a corresponding correction rationality result;
[0122] S15-3, select the candidate correction suggestion corresponding to the correction rationality result greater than the similarity threshold to perform abnormal correction on the abnormal triple data to obtain the optimized knowledge graph.
[0123] In this embodiment, a large model (such as GPT-4) is used to generate multiple candidate correction suggestions based on abnormal triple data, where the prompt word template is as follows:
[0124] [Abnormal triplet data]: [X→activity location→Y city];
[0125] [Context]: Y city entity surges by 300%, semantic similarity 0.55;
[0126] [Quest]: Generates possible relationship modifiers (e.g. "Planning an Attack", "Recruiting Members");
[0127] Then, a domain-specific model (such as the SecurityBert model) is used to verify the rationality of the candidate correction suggestions (if the similarity is > 0.85, it will be adopted). If the conditions are met, the candidate correction suggestions that meet the conditions will be updated to the knowledge graph; otherwise, they will be marked as "pending manual review".
[0128] Finally, the correction process is traced, including recording the timestamp of the correction operation, the basis for decision-making (rule base / model base) and the chain of evidence, to support audit backtracking.
[0129] Furthermore, the method further comprises:
[0130] S16, after using the GPT large model to perform anomaly correction on the abnormal triple data, the obtained corrected triple data is added to the training data set to update the iterative GAT model.
[0131] In this embodiment, the modified triplet data obtained by the above modification is added to the training set to achieve incremental update and iteration of the GAT model. Specifically, the decision weights are dynamically optimized based on the PPO algorithm (such as adjusting the accuracy vs. timeliness weights); the GAT model is automatically trained every week and the newly added intelligence data is embedded.
[0132] The following is described by a specific embodiment, specifically:
[0133] Access to real-time data streams of dark web forums (AMQP protocol), processing 100,000 open source intelligence pieces per day, including JSON-LD and CSV formats. Use TransE model to map to a unified entity, SimHash to remove duplicates, and standardize triples to RDF format.
[0134] Abnormal feature detection, including:
[0135] a) Time series anomaly: Monitor the "Y event" entity in the sliding window (dynamically adjusted 1 to 6 hours) and obtain 500 event entities, while the Prophet algorithm predicts 250 event entities. Set C to 100. According to Calculate 500-250=250, 250>100, that is, the actual value and the predicted value have a large deviation, and it is marked as abnormal data;
[0136] b) Semantic anomaly: The cosine similarity of "X organization → event location → Y city" is calculated by BERT-TransE joint embedding = 0.45 (threshold 0.7). The specific calculation includes:
[0137] According to the formula Calculation is performed, A represents the embedding vector of the triple “X organization → activity location → Y city”, A=[0.9, 0.1, 0.3, 0.7, 0.2], B represents the average embedding vector of the stock triples,
[0138] B=[0.2,0.8,0.6,0.1,0.5], thus
[0139] = (0.9×0.2)+(0.1×0.8)+(0.3×0.6)+(0.7×0.1)+(0.2×0.5) =0.61;
[0140] ;
[0141] ;
[0142] The calculated cosine similarity = 0.45 < 0.7 is marked as abnormal data.
[0143] Dynamic decision processing, including:
[0144] Confidence calculation: Rule confidence: Match the "event-location" constraint in the SHACL rule base (historical accuracy of the rule base SHACL = total number of rule triggers / number of successful rule matches = 92% = 0.92), complexity attenuation factor = 0.85, then rule confidence = historical accuracy × complexity attenuation factor = 0.92 × 0.85 = 0.782; Model confidence:
[0145] Xgboost model predictions:
[0146] Assume that true positives (TP) = 80, false positives (FP) = 30, false negatives (FN) = 20, and true negatives (TN) = 70, then TPR = TP / (TP+FN) = 80 / (80+20) = 0.8, FPR = FP / (FP+TN) = 30 / (30+70) = 0.3;
[0147] Calculate AUC: Assume that the (FPR, TPR) points of the model at multiple thresholds are as follows: (0,0), (0.1,0.4), (0.2,0.65), (0.3,0.8), (1,1), then
[0148] AUC=(0.1-0)×(0+0.4) / 2+(0.2-0.1)×(0.4+0.64) / 2+(0.3-0.2)×(0.65+ 0.8) / 2+(1-0.3)×(0.8+1) / 2=0.1×0.2+0.1×0.525+0.1×0.725+0.7×0.9= 0.775.
[0149] Calculate GAT_F1_score:
[0150] The classification results of the GAT model are as follows:
[0151] True positives (TP) = 85, false positives (FP) = 25, false negatives (FN) = 15,
[0152] Precision = TP / (TP+FP)=85 / (85+25)=85 / 110≈0.7727,
[0153] Recall rate recall = TP / (TP+FN)=85 / (85+15)=85 / 100=0.85,
[0154] F1=2×(Precision×Recall) / (Precision+Recall)=
[0155] 2×(0.7727×0.85) / (0.7727+0.85)=0.8096; therefore, XGBoost_AUC=0.78, GAT_F1_score=0.81, and the model confidence is calculated as 0.4×0.76+0.6×0.81=0.79; the model confidence is greater than the rule confidence.
[0156] Decision execution: By calling the XGBoost and GAT models, it was detected that "X organization → activity location → Y city" had a complex association anomaly (risk score 0.91).
[0157] Anomaly correction and tracing, including:
[0158] GPT-4 model output: The candidate correction suggestions are [“X organization → planned attack → Y city”, “X organization → planned activities → Y city”, “X organization → suspected actions → Y city”].
[0159] SecurityBert verification: Calculate the semantic similarity between the candidate correction suggestion and the existing triple data = 0.89 (threshold 0.85), and adopt the correction suggestion.
[0160] Traceability tracking record: timestamp 2023-10-05T14:23:45Z, decision-making model weight 1.87, GAT risk score 0.91, external evidence association related operation log relationship similarity before correction 0.55 - after correction 0.89, operation type = automatic.
[0161] Online Model Evolution:
[0162] Incremental training: Data injection adds the revised “X organization → planned attack → Y city” to the GAT training set and generates adversarial samples (such as “X organization → charity event → Y city”); fine-tune GAT every week, with a learning rate of 0.001 (initial value 0.01), and retain 90% of the historical weights.
[0163] Reinforcement learning optimization: PPO strategy: Construct a reward function with correction accuracy (94%), response speed (delay < 5 minutes), and resource consumption (GPU < 80%):
[0164] R=0.6×0.94+0.3×0.9−0.1×0.7=0.815. The model weight coefficient β of the decision engine is improved according to the R value (from 0.6→0.65).
[0165] Reference Figure 3 Shown is a schematic diagram of the structure of a quality optimization device for an intelligence knowledge graph provided by an embodiment of the present invention.
[0166] In this embodiment, the device 20 includes:
[0167] A data acquisition unit 21 is used to acquire multi-source intelligence data and perform preprocessing to obtain initial triplet data;
[0168] A first detection unit 22 is used to perform preliminary anomaly detection on the initial triple data to obtain time series feature anomalies and semantic feature anomalies;
[0169] A calculation unit 23 is used to calculate the confidence of the initial triple data corresponding to the abnormal time series feature and the abnormal semantic feature, and determine the optimal decision plan according to the confidence;
[0170] A second detection unit 24 is used to perform secondary anomaly detection on the initial triple data corresponding to the abnormal time series feature and the abnormal semantic feature according to the optimal decision scheme to obtain abnormal triple data;
[0171] The correction unit 25 is used to use the GPT large model to perform anomaly correction on the abnormal triple data to obtain an updated optimized knowledge graph.
[0172] Each unit module of the device 20 can respectively execute the corresponding steps in the above method embodiment, so each unit module will not be described in detail here. Please refer to the description of the above corresponding steps for details.
[0173] The embodiment of the present invention further provides a quality optimization device for an intelligence knowledge graph, the device comprising the quality optimization device for the intelligence knowledge graph as described above, wherein the quality optimization device for the intelligence knowledge graph can adopt Figure 3 The structure of the embodiment can be executed accordingly. Figure 1 The technical solution of the method embodiment shown has similar implementation principles and technical effects. For details, please refer to the relevant records in the above embodiments, which will not be repeated here.
[0174] The device includes: a mobile phone, a digital camera, a tablet computer, or other devices with a camera function, or a device with an image processing function, or a device with an image display function. The device may include components such as a memory, a processor, an input unit, a display unit, and a power supply.
[0175] Among them, the memory can be used to store software programs and modules, and the processor executes various functional applications and data processing by running the software programs and modules stored in the memory. The memory may mainly include a program storage area and a data storage area, wherein the program storage area may store an operating system, an application required for at least one function (such as an image playback function, etc.), etc.; the data storage area may store data created according to the use of the device, etc. In addition, the memory may include a high-speed random access memory, and may also include a non-volatile memory, such as at least one disk storage device, a flash memory device, or other volatile solid-state storage devices. Accordingly, the memory may also include a memory controller to provide the processor and the input unit with access to the memory.
[0176] The input unit can be used to receive input digital or character or image information, and generate keyboard, mouse, joystick, optical or trackball signal input related to user settings and function control. Specifically, the input unit of this embodiment includes not only a camera, but also a touch-sensitive surface (such as a touch display) and other input devices.
[0177] The display unit can be used to display information input by the user or information provided to the user and various graphical user interfaces of the device, which can be composed of graphics, text, icons, videos and any combination thereof. The display unit may include a display panel, and optionally, the display panel may be configured in the form of LCD (Liquid Crystal Display), OLED (Organic Light-Emitting Diode), etc. Further, the touch-sensitive surface may cover the display panel, and when the touch-sensitive surface detects a touch operation on or near it, it is transmitted to the processor to determine the type of touch event, and then the processor provides a corresponding visual output on the display panel according to the type of touch event.
[0178] The embodiment of the present invention further provides a computer-readable storage medium, which may be a computer-readable storage medium included in the memory in the above embodiment; or a computer-readable storage medium that exists independently and is not installed in a device. The computer-readable storage medium stores at least one instruction, which is loaded and executed by a processor to implement Figure 1 The quality optimization method of the intelligence knowledge graph shown in the figure. The computer-readable storage medium can be a read-only memory, a disk or an optical disk, etc.
[0179] The embodiment of the present invention also provides a computer program product, including a computer program / instruction, which is loaded and executed by a processor to implement Figure 1 A quality optimization method for an intelligence knowledge graph is shown.
[0180] It should be noted that each embodiment in this specification is described in a progressive manner, and each embodiment focuses on the differences from other embodiments, and the same or similar parts between the embodiments can be referred to each other. For the device embodiment, equipment embodiment and storage medium embodiment, since they are basically similar to the method embodiment, the description is relatively simple, and the relevant parts can be referred to the partial description of the method embodiment.
[0181] Furthermore, in this document, the terms "comprises," "comprising," or any other variation thereof are intended to encompass non-exclusive inclusion, such that a process, method, article, or device that includes a series of elements includes not only those elements, but also other elements not explicitly listed, or elements inherent to such process, method, article, or device. In the absence of further limitations, an element defined by the phrase "comprising a ..." does not exclude the presence of additional identical elements in the process, method, article, or device that includes the element.
[0182] The above description shows and describes the preferred embodiments of the present invention. It should be understood that the present invention is not limited to the form disclosed herein, and should not be regarded as excluding other embodiments, but can be used in various other combinations, modifications and environments, and can be modified within the scope of the invention, through the above teachings or the technology or knowledge of the relevant field. The changes and modifications made by those skilled in the art shall not depart from the spirit and scope of the present invention, and shall be within the scope of protection of the claims attached to the present invention.
Claims
1. A quality optimization method for an intelligence knowledge graph, characterized in that: The method comprises: Acquire multi-source intelligence data and pre-process them to obtain initial triplet data; Performing preliminary anomaly detection on the initial triplet data to obtain temporal feature anomalies and semantic feature anomalies; Calculating the confidence of the initial triple data corresponding to the abnormal time series features and the abnormal semantic features, and determining the optimal decision plan according to the confidence; Further, the confidence includes a first confidence and a second confidence; the confidence of the initial triple data corresponding to the abnormal time series feature and the abnormal semantic feature is calculated, and the optimal decision plan is determined according to the confidence, including: Calculating using the preset rule confidence and model confidence respectively to obtain a first confidence and a second confidence; When it is determined that the first confidence level is greater than the second confidence level, the rule library is called to process as the optimal decision solution; otherwise, the model library is called to process as the optimal decision solution; Further, the calculation is performed using the preset rule confidence and model confidence respectively to obtain the first confidence and the second confidence, including: The first confidence is obtained by calculating rule confidence = total number of rule triggering times / number of successful rule matching times × rule complexity attenuation factor, wherein the value range of the rule complexity attenuation factor is 0.8 to 1.0; The second confidence is calculated using model confidence = 0.4*XGBootst_AUC+0.6*GAT_F1_score, where XGBootst_AUC represents the index value of the XGBootst model. AUC in XGBootst_AUC measures the ability of the XGBootst model to distinguish between positive and negative samples by calculating the area under the ROC curve; GAT_F1_score represents the harmonic mean of precision and recall; where, XGBootst_AUC= , Where, TPR i+1 Indicates the true rate of the next point of the current point, TPR i Indicates the true rate of the current point, FPR i+1 Indicates the false positive rate of the next point of the current point, FPR i Indicates the false positive rate of the current point. In the ROC curve, each point (TPR i , FPR i ) represents the true positive rate and false positive rate under a threshold i, i represents the index of the current point, and the value range of i is 0-1; GAT_F1_score= , Precision = TP / (TP+FP), Recall = TP / (TP+FN), where TP represents true positives, FP represents false positives, and FN represents false negatives; According to the optimal decision-making scheme, secondary anomaly detection is performed on the initial triple data corresponding to the abnormal time series features and the abnormal semantic features to obtain abnormal triple data; The GPT big model is used to perform anomaly correction on the abnormal triple data to obtain an updated optimized knowledge graph.
2. The quality optimization method of the intelligence knowledge graph according to claim 1 is characterized in that: The preliminary anomaly detection is performed on the initial triple data to obtain time series feature anomalies and semantic feature anomalies, including: Extracting features from the initial triplet data to obtain temporal features and semantic features; Performing anomaly detection on the time series feature based on the Prophet algorithm to obtain anomalies of the time series feature; The cosine similarity between the semantic feature and the stock triple data is calculated. If it is lower than a preset threshold, the corresponding semantic feature is marked as abnormal, and the semantic feature abnormality is obtained.
3. The quality optimization method of the intelligence knowledge graph according to claim 1 is characterized in that: The performing secondary anomaly detection on the initial triplet data corresponding to the abnormal time series feature and the abnormal semantic feature according to the optimal decision scheme to obtain abnormal triplet data includes: The pre-built rule base is used to match the SHACL constraint rules on the initial triple data corresponding to the timing feature anomaly and the semantic feature anomaly, and the matched data is used as the abnormal triple data.
4. The quality optimization method of the intelligence knowledge graph according to claim 1 is characterized in that: The performing secondary anomaly detection on the initial triplet data corresponding to the abnormal time series feature and the abnormal semantic feature according to the optimal decision scheme to obtain abnormal triplet data includes: The pre-trained XGBoost classifier and GAT model are used to detect the initial triple data corresponding to the time series feature anomaly and the semantic feature anomaly, and the data with existing relationship anomalies is used as the abnormal triple data.
5. The quality optimization method of the intelligence knowledge graph according to claim 1 is characterized in that: The method of using the GPT large model to perform abnormal correction on the abnormal triple data to obtain an updated optimized knowledge graph includes: Generate multiple candidate correction suggestions based on the abnormal triple data using the GPT large model; Each candidate correction suggestion is verified by a preset SecurityBert model to obtain a corresponding correction rationality result; The candidate correction suggestion corresponding to the correction rationality result being greater than the similarity threshold is selected to perform anomaly correction on the abnormal triple data to obtain the optimized knowledge graph.
6. The quality optimization method of the intelligence knowledge graph according to claim 1 is characterized in that: The method further comprises: After the abnormal triplet data is corrected using the GPT large model, the corrected triplet data is added to the training data set to update the iterative GAT model.
7. A quality optimization device for an intelligence knowledge graph, characterized in that: The device comprises: A data acquisition unit is used to acquire multi-source intelligence data and perform preprocessing to obtain initial triplet data; A first detection unit is used to perform preliminary anomaly detection on the initial triple data to obtain time series feature anomalies and semantic feature anomalies; A calculation unit, used to calculate the confidence of the initial triple data corresponding to the abnormal time series feature and the abnormal semantic feature, and determine the optimal decision plan according to the confidence; Further, the confidence level includes a first confidence level and a second confidence level; and the calculation unit is used to: Calculating using the preset rule confidence and model confidence respectively to obtain a first confidence and a second confidence; When it is determined that the first confidence level is greater than the second confidence level, the rule library is called to process as the optimal decision solution; otherwise, the model library is called to process as the optimal decision solution; Further, the calculation is performed using the preset rule confidence and model confidence respectively to obtain the first confidence and the second confidence, including: The first confidence is obtained by calculating rule confidence = total number of rule triggering times / number of successful rule matching times × rule complexity attenuation factor, wherein the value range of the rule complexity attenuation factor is 0.8 to 1.0; The second confidence is calculated using model confidence = 0.4*XGBootst_AUC+0.6*GAT_F1_score, where XGBootst_AUC represents the index value of the XGBootst model. AUC in XGBootst_AUC measures the ability of the XGBootst model to distinguish between positive and negative samples by calculating the area under the ROC curve; GAT_F1_score represents the harmonic mean of precision and recall; where, XGBootst_AUC= , Where, TPR i+1 Indicates the true rate of the next point of the current point, TPR i Indicates the true rate of the current point, FPR i+1 Indicates the false positive rate of the next point of the current point, FPR i Indicates the false positive rate of the current point. In the ROC curve, each point (TPR i , FPR i ) represents the true positive rate and false positive rate under a threshold i, i represents the index of the current point, and the value range of i is 0-1; GAT_F1_score= , Precision = TP / (TP+FP), Recall = TP / (TP+FN), where TP represents true positives, FP represents false positives, and FN represents false negatives; A second detection unit is used to perform secondary anomaly detection on the initial triple data corresponding to the abnormal time series feature and the abnormal semantic feature according to the optimal decision scheme to obtain abnormal triple data; The correction unit is used to use the GPT large model to perform anomaly correction on the abnormal triple data to obtain an updated optimized knowledge graph.
8. A quality optimization device for an intelligence knowledge graph, characterized in that: It includes a processor, a memory and a computer program stored in the memory, wherein the computer program is executed by the processor to implement the steps of a quality optimization method of an intelligence knowledge graph as described in any one of claims 1 to 6.
Citation Information
Patent Citations
Fault detection method based on analytic hierarchy process and weighted vote decision fusion
CN106355030A
Multi-task recommendation algorithm fusing user behaviors and knowledge graph
CN117370674A