Knowledge graph data intelligent management method and system based on semantic web technology

Through the intelligent management method of knowledge graph data based on semantic web technology, the problems of real-time processing and multi-source noise interference of high-frequency dynamic data streams are solved, the dynamic update and rapid response of the knowledge graph are realized, and its application efficiency in the fields of intelligent search, decision support and semantic reasoning is improved.

CN120706519AInactive Publication Date: 2025-09-26GUIZHOU XIAOQI TECHNOLOGY CO LTD
View PDF 0 Cites 8 Cited by

Patent Information

Application Number
CN202510817591.X
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-06-18
Publication Date
2025-09-26
Estimated Expiration
Not applicable · inactive patent

AI Technical Summary

Technical Problem

Existing technologies have difficulty achieving real-time performance when processing high-frequency dynamic data streams, and systematically resolve multi-source noise interference and semantic conflicts, which affects the accuracy and consistency of knowledge graphs. In addition, priority sorting and credibility assessment are insufficient, restricting their application effectiveness in actual scenarios.

Method used

Adopting the knowledge graph data intelligent management method based on semantic web technology, through the steps of real-time text data collection, semantic information extraction, timestamp processing, noise filtering, semantic resolution, knowledge mapping, graph embedding and credibility calculation, it realizes the real-time processing of high-frequency data streams, precise filtering of multi-source data and dynamic conflict resolution.

Benefits of technology

It realizes the dynamic update and rapid response of knowledge graphs, improves their application efficiency in intelligent search, decision support and semantic reasoning, ensures the accuracy, consistency and reliability of knowledge graphs, and adapts to the efficiency and stability of large-scale dynamic data processing.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120706519A_ABST
    Figure CN120706519A_ABST
Patent Text Reader

Abstract

The invention relates to a knowledge graph data intelligent management method and system based on a semantic web technology, and the method comprises the steps: obtaining text data from a high-frequency data flow in real time, and generating a first semantic set through segmentation processing and semantic extraction; performing noise filtering and sorting on the multi-source data to generate a second semantic set; constructing semantic representation compatible with the knowledge graph; utilizing a graph embedding algorithm to generate graph updating data through source weight optimization; based on a historical conflict mode and a credibility weighting model, intelligent resolution of semantic conflicts is completed; and generating dynamic situation awareness data through real-time incremental loading and multi-dimensional association analysis. According to the method, through time sequence priority dynamic weighting, multi-source noise accurate filtering and cross-modal credibility evaluation, the problems of response lag, redundancy accumulation and insufficient conflict resolution during high-frequency dynamic data processing of a traditional method can be solved, and therefore real-time updating and consistency maintenance of the knowledge graph are achieved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of knowledge graph data management, and specifically to a method and system for intelligent management of knowledge graph data based on semantic web technology. Background Art

[0002] In today's data-driven information society, knowledge graphs, as the core carrier of structured knowledge, have become a crucial infrastructure for intelligent search, decision support, and semantic reasoning. With the rapid development of technologies such as the Internet of Things and social media, the rate of data generation is increasing exponentially. In particular, the multi-source, heterogeneous information in high-frequency, dynamic data streams urgently requires automated means to efficiently integrate and update it in real time. Against this backdrop, knowledge graph management methods based on semantic web technologies have become a research hotspot. Through techniques such as semantic modeling and relational reasoning, these methods aim to enhance the dynamic adaptability and interpretability of knowledge systems.

[0003] However, existing technologies still face significant challenges in dealing with high-frequency dynamic data streams. On the one hand, traditional methods often rely on static rules or offline batch processing mechanisms, which are difficult to adapt to the real-time requirements of data streams, resulting in delayed knowledge updates and the accumulation of redundant information. On the other hand, there is a lack of systematic solutions to the noise interference, semantic conflicts, and entity ambiguity problems of multi-source data, which seriously affect the accuracy and consistency of knowledge graphs. In addition, when processing large-scale dynamic data, existing technologies often find it difficult to balance timeliness and reliability due to problems such as insufficient priority sorting mechanisms and lack of credibility assessments, further restricting the application effectiveness of knowledge graphs in practical scenarios. Summary of the Invention

[0004] Based on this, the purpose of the present invention is to provide a knowledge graph data intelligent management method and system based on semantic web technology that can realize real-time processing of high-frequency data streams, precise filtering of multi-source noise and intelligent resolution of dynamic conflicts.

[0005] The purpose of the present invention is achieved by the following scheme:

[0006] In a first aspect, the present invention provides a method for intelligent management of knowledge graph data based on semantic web technology, comprising the following steps:

[0007] S1: acquiring real-time text data from a high-frequency data stream based on a preset sampling frequency, and segmenting and extracting semantic information from the real-time text data based on a natural language processing method to generate a first semantic set;

[0008] S2: extracting the timestamp attribute of each data in the first semantic set based on a predefined time priority rule, and performing time sorting on the data with duplicate or missing timestamps to generate a second semantic set;

[0009] S3: Based on semantic integration technology, the second semantic set is subjected to noise filtering, semantic resolution, and knowledge mapping processing to generate a graph semantic representation that is compatible with the knowledge graph;

[0010] S4: Based on the graph embedding knowledge fusion algorithm, data is extracted from the graph semantic representation and integrated with the nodes and edges in the knowledge graph. When duplicate entities appear, the retained objects are judged based on the preset source weights to generate graph update data.

[0011] S5: Based on the probability model trained with historical data and the conflict resolution algorithm, the graph update data is subjected to credibility calculation and data resolution processing to generate a resolution semantic set.

[0012] S6: Based on dynamic loading technology and semantic web version control mechanism, the data in the semantic set is monitored and adjusted in real time to generate semantic situation awareness data; semantic situation awareness data is used to reflect the changes in entity relationships and semantic conflict resolution status in the knowledge graph in real time.

[0013] In one embodiment, S3 of a method for intelligent management of knowledge graph data based on semantic web technology provided by the present invention specifically includes the following steps:

[0014] S31: performing noise data identification processing on the second semantic set based on keyword matching and context analysis, applying a preset relevance threshold to filter low-relevance content, and generating a graph semantic segment;

[0015] S32: Vectorize the graph semantic segments based on the semantic vector generation model to generate a graph feature vector. The calculation formula of the graph feature vector is:

[0016]

[0017] Among them, v is the graph feature vector, w i is the i-th word in the semantic segment of the graph, Embedding(w i ) is the word vector generated by the word embedding model, and n is the number of words in the graph semantic segment;

[0018] S33: Perform structured transformation on the graph feature vector based on the knowledge representation method, and map and align it with the entities and relationships of the knowledge graph to generate a graph semantic representation.

[0019] In one embodiment, S31 of a method for intelligent management of knowledge graph data based on semantic web technology provided by the present invention specifically includes the following steps:

[0020] S311: performing noise keyword matching processing on the second semantic set based on a preset keyword blacklist to generate an initial filtering result;

[0021] S312: Quantify the context relevance of the initial filtering result to generate a relevance score. The calculation formula of the relevance score is:

[0022]

[0023] Among them, RelevanceScore is the relevance score, TF(t) is the frequency of text word t in the text, DF(t) is the document frequency of text word t in the knowledge graph, N is the total number of documents in the knowledge graph, and T is the vocabulary set of the current text;

[0024] S313: Filter the relevance scores based on a preset relevance threshold to generate graph semantic segments.

[0025] In one embodiment, S4 of a method for intelligent management of knowledge graph data based on semantic web technology provided by the present invention specifically includes the following steps:

[0026] S41: Vectorize the entities and relationships in the graph semantic representation based on the graph embedding algorithm to generate the entity set to be fused;

[0027] S42: Calculate the cosine similarity between the entity set to be fused and the existing entities based on the embedding vectors of the knowledge graph nodes to generate a candidate entity similarity set; identify whether each entity in the candidate entity similarity set exceeds a preset duplication threshold, and if so, mark the entities exceeding the threshold as duplicate entities to generate an optimized set to be fused containing duplicate entities;

[0028] S43: Based on the preset source weights, duplicate entity attributes in the optimized set to be fused are merged to generate graph update data.

[0029] In one embodiment, S5 of a method for intelligent management of knowledge graph data based on semantic web technology provided by the present invention specifically includes the following steps:

[0030] S51: Calculate the source credibility and time decay factor of the entity in the graph update data to generate an initial credibility score;

[0031] S52: Perform historical conflict detection based on entity identifiers in the graph update data and historical versions of the knowledge graph to generate a set of high-risk conflict entities;

[0032] S53: performing weighted calculation on the score values ​​in the initial credibility score set, the conflict marks of the high-risk conflict entity set, and the entity confidence of the graph update data to generate a final credibility score;

[0033] S54: Based on the final credibility score, the high-risk conflict entity set is identified and processed to determine whether the final credibility score of each entity is lower than the preset credibility threshold. If so, the entity is deleted from the high-risk conflict entity set; otherwise, the attribute with the highest source weight among the entities in the high-risk conflict entity set is retained to generate a resolution semantic set.

[0034] In one embodiment, S6 of a method for intelligent management of knowledge graph data based on semantic web technology provided by the present invention specifically includes the following steps:

[0035] S61: Based on real-time monitoring technology, the data in the resolved semantic set is loaded in batches and incremental update instructions are generated;

[0036] S62: Based on the semantic web-driven RDF triple update rules, the instructions in the incremental update instruction set are parsed and processed, and the nodes and edges of the knowledge graph are added, deleted, and modified to generate an updated knowledge graph;

[0037] S63: Based on the updated knowledge graph, the entity-relationship-time dimension analysis and processing are performed through the multi-dimensional semantic association matrix construction technology to generate semantic situational awareness data that supports real-time decision-making.

[0038] In one embodiment, S2 of a method for intelligent management of knowledge graph data based on semantic web technology provided by the present invention specifically includes the following steps:

[0039] S21: Based on the data source identifier of each piece of data in the first semantic set, extract the corresponding credibility score from a preset historical data credibility library to generate a data source credibility value;

[0040] S22: Calculate the difference between the timestamp of each data item in the first semantic set and the current time to generate a time decay factor. The calculation formula of the time decay factor is:

[0041]

[0042] Among them, TimeDecay is the time decay factor, α is the time decay coefficient, t current is the current time, t record Record timestamps for data;

[0043] S23: generating entity confidence based on the entity recognition probability value output by the natural language processing in the first semantic set;

[0044] S24: Prioritize the data source credibility, time decay factor, and entity confidence to generate a priority weight. The priority weight calculation formula is:

[0045] P=w1*SourceCred+w2*TimDecay+w3*EntityConf

[0046] Among them, P is the priority weight, SourceCred is the data source credibility, TimeDecay is the time decay factor, EntityConf is the entity confidence, w1, w2, w3 are dynamic weight coefficients;

[0047] S25: Sort each piece of data in the first semantic set in descending order based on the priority weight to generate a second semantic set.

[0048] In a second aspect, the present invention provides a knowledge graph data intelligent management system based on semantic web technology, comprising:

[0049] a data acquisition processing module, configured to acquire real-time text data from a high-frequency data stream based on a preset sampling frequency, and segment and extract semantic information from the real-time text data based on a natural language processing method to generate a first semantic set;

[0050] A time attribute processing module is used to extract the timestamp attribute of each data in the first semantic set based on a predefined time priority rule, and to perform time sorting processing on the data with repeated or missing timestamps to generate a second semantic set;

[0051] A graph semantic integration module is used to perform noise filtering, semantic resolution, and knowledge mapping on the second semantic set based on semantic integration technology to generate a graph semantic representation that is compatible with the knowledge graph;

[0052] The graph data integration module is used to extract data based on the graph semantic representation of the graph embedding knowledge fusion algorithm, and integrate it with the nodes and edges in the knowledge graph. When duplicate entities appear, the retained objects are judged according to the preset source weights to generate graph update data;

[0053] The credibility calculation and resolution module is used to perform credibility calculation and data resolution processing on the graph update data based on the probability model trained with historical data and the conflict resolution algorithm, and generate a resolution semantic set;

[0054] The real-time monitoring and adjustment module is used to monitor and adjust the data in the semantic set in real time based on dynamic loading technology and semantic web version control mechanism to generate semantic situation awareness data; the semantic situation awareness data is used to reflect the changes in entity relationships and the status of semantic conflict resolution in the knowledge graph in real time.

[0055] In a third aspect, the present application provides a computer device comprising a memory and a processor, wherein the memory stores a computer program, and when the processor executes the computer program, it implements any of the above-mentioned methods for intelligent management of knowledge graph data based on semantic web technology.

[0056] In a fourth aspect, the present application provides a computer-readable storage medium on which a computer program is stored. When the computer program is executed by a processor, it implements any of the above-mentioned knowledge graph data intelligent management methods based on semantic web technology.

[0057] In summary, the knowledge graph data intelligent management method based on semantic web technology provided by the present invention can solve the lag problem of traditional methods in responding to real-time requirements by real-time collection and processing of high-frequency dynamic data streams, thereby realizing dynamic updating and rapid response of knowledge graphs; at the same time, in the process of integrating multi-source heterogeneous data, the method systematically solves problems such as noise interference, semantic conflict and entity ambiguity through technologies such as noise filtering, semantic resolution and knowledge mapping, thereby improving the accuracy, consistency and reliability of the knowledge graph; and uses the graph-embedded knowledge fusion algorithm and credibility evaluation mechanism to balance the timeliness and reliability of the system, thereby ensuring the efficiency and stability of the knowledge graph in large-scale dynamic data processing. Finally, through semantic situation awareness and real-time adjustment mechanism, the knowledge graph can better adapt to the ever-changing data environment, thereby enhancing its application efficiency in fields such as intelligent search, decision support and semantic reasoning, providing users with more accurate, timely and reliable knowledge services, and promoting the widespread application and development of knowledge graph technology in actual scenarios.

[0058] For better understanding and implementation, the present invention is described in detail below with reference to the accompanying drawings. BRIEF DESCRIPTION OF THE DRAWINGS

[0059] Figure 1 A flowchart of a method for intelligent management of knowledge graph data based on semantic web technology provided in an embodiment of the present application;

[0060] Figure 2 A schematic diagram of the process of generating a graph semantic representation provided in an embodiment of the present application;

[0061] Figure 3 A structural diagram of a knowledge graph data intelligent management system based on semantic web technology provided in another embodiment of the present application. DETAILED DESCRIPTION

[0062] To facilitate understanding of the present invention, the present invention will be described more fully below with reference to the accompanying drawings. The drawings illustrate preferred embodiments of the present invention. However, the present invention may be implemented in many different forms and is not limited to the embodiments described herein. Rather, these embodiments are provided to provide a more thorough and comprehensive understanding of the disclosure.

[0063] Unless otherwise defined, all technical and scientific terms used herein have the same meanings as those commonly understood by those skilled in the art to which this invention pertains. The terms used in this specification are for the purpose of describing specific embodiments only and are not intended to limit the invention. The term "and / or" as used herein includes any and all combinations of one or more of the associated listed items.

[0064] In one embodiment, Figure 1 As shown, a method for intelligent management of knowledge graph data based on semantic web technology is provided. This embodiment uses the method applied to a terminal as an example. It is understandable that the method can also be applied to a server, or to a system including a terminal and a server, and implemented through the interaction between the terminal and the server. In this embodiment, the method includes the following steps:

[0065] S1: Acquire real-time text data from a high-frequency data stream based on a preset sampling frequency, and segment and extract semantic information from the real-time text data based on a natural language processing method to generate a first semantic set.

[0066] Specifically, the system captures text data from high-frequency dynamic data streams in real time through a preset sampling frequency (for example, 100 samples per second). The sampling frequency setting is dynamically optimized according to the data stream generation rate and the system's processing capabilities to ensure the real-time and integrity of data collection. The captured text data first undergoes preliminary cleaning, including removing noise characters, formatting text structure, and standardizing encoding formats (such as UTF-8) to ensure the accuracy of subsequent processing. The system then segments the cleaned text data through natural language processing technology. The segmentation algorithm can use a word segmentation method based on syntactic analysis (such as the maximum matching method) or a sequence annotation model based on deep learning (such as BERT). The segmented text unit serves as the basic unit for subsequent semantic extraction.

[0067] Natural language processing technology is the core of the system's text data processing. By simulating how humans understand language, natural language processing performs operations such as word segmentation, part-of-speech tagging, and syntactic analysis on text, thereby extracting key information from the text. For example, word segmentation divides continuous text along word boundaries to facilitate subsequent semantic analysis. Part-of-speech tagging classifies each word into its part of speech, such as noun or verb, providing a foundation for syntactic analysis. Syntactic analysis constructs a grammatical structure for a sentence, identifying its subject, predicate, object, and other components, thereby understanding the sentence's structure and semantics. These technologies work together to enable the system to extract valuable semantic information from text.

[0068] After data segmentation, the system further extracts semantic features from the segmented text units, including entity recognition, relationship extraction, and semantic role labeling. Entity recognition can be implemented using a named entity recognition model based on conditional random fields (CRF) or a transformer; relationship extraction can be implemented using dependency parsing or a graph neural network (GNN); and semantic role labeling is used to clarify the semantic roles of each entity in the text. The extracted semantic information is stored in a structured form, generating a first semantic set that provides a foundation for subsequent processing.

[0069] S2: extracting the timestamp attribute of each data in the first semantic set based on a predefined time priority rule, and performing time sorting processing on the data with repeated or missing timestamps to generate a second semantic set.

[0070] Specifically, predefined time priority rules prioritize data based on the credibility of the data source, the contextual consistency of the timestamp, and the system's preset time base. For example, data from highly reliable data sources will have its timestamp prioritized higher. For data with timestamps that clearly don't match the contextual description, the system will adjust the timestamp based on the context.

[0071] Specifically, the system extracts the timestamp attribute for each data in the first semantic set, and the timestamp extraction is achieved through regular expression matching or rule-based time expression parsing. If the timestamp is not explicitly included in the data, the system infers the time information through contextual semantics, such as a description based on the relative time of the event. For data with repeated or missing timestamps, the system uses predefined time priority rules for processing. The time priority rules are based on the credibility of the data source, the contextual consistency of the timestamp, and the time benchmark preset by the system for sorting. For data with repeated timestamps, the system gives priority to retaining data from high-credibility sources; for data with missing timestamps, the system supplements it through contextual semantic inference or the system's default timestamp. The processed timestamp data is sorted in chronological order to ensure the temporality of the data. The sorting algorithm uses quick sort or merge sort to ensure sorting efficiency, and generates a second semantic set through the sorted data to provide support for subsequent semantic integration.

[0072] S3: Based on semantic integration technology, the second semantic set is subjected to noise filtering, semantic resolution and knowledge mapping processing to generate a graph semantic representation that is compatible with the knowledge graph.

[0073] Specifically, the predefined noise filtering rules include low-frequency word filtering based on word frequency statistics, redundant information removal based on semantic similarity, and outlier detection based on context consistency checking. Through these rules, the system can effectively remove noise from the data and retain the core semantic information. Semantic resolution technology is the key to dealing with semantic conflicts or ambiguity. The system resolves ambiguous entities through contextual reasoning, synonym replacement, or semantic constraints based on the knowledge base. For example, when the word "apple" appears in the text, the system will determine whether it refers to fruit or a technology company based on the context. Contextual reasoning infers the most reasonable semantics by analyzing other entities and relationships in the text; synonym replacement replaces ambiguous words with more clear expressions by searching the synonym dictionary; semantic constraints based on the knowledge base use the entity relationships in the existing knowledge base to constrain the possible values ​​of semantics.

[0074] Specifically, the system further maps the resolved semantic data into the ontology structure of the knowledge graph, matching the entities, relationships, and attributes in the semantic data with corresponding elements in the knowledge graph based on predefined ontology mapping rules. Ontology mapping rules define how entities and relationships in different data sources correspond to ontology elements in the knowledge graph, ensuring that the data is accurately integrated into the knowledge graph structure. The resulting mapping generates a graph semantic representation compatible with the knowledge graph, providing a foundation for subsequent knowledge fusion.

[0075] S4: The graph semantic representation of the graph is used for data extraction based on the graph embedding knowledge fusion algorithm, and the data is integrated with the nodes and edges in the knowledge graph. When repeated entities appear, the retained objects are judged according to the preset source weights to generate graph update data.

[0076] Graph embedding algorithms are key technologies for converting semantic data into vector form. The system uses a random walk-based graph embedding algorithm (Node2Vec) or a graph neural network (GNN)-based GraphSAGE (Graph Sample and Aggregate) algorithm for graph embedding. Node2Vec captures the adjacency relationships and structural features between entities by performing random walks in the graph, and converts these features into low-dimensional vector representations. GraphSAGE generates an embedding vector for each node by aggregating the feature information of neighboring nodes, which can better capture the global structural information in the graph. These algorithms enable the system to convert complex semantic relationships into computable vector form, providing a foundation for knowledge fusion.

[0077] Specifically, the system integrates the embedded semantic data with the existing nodes and edges in the knowledge graph. The integration process includes entity matching, relationship expansion, and attribute updates. Entity matching is achieved through algorithms based on string similarity or semantic similarity, which can accurately identify the same entity in different data sources; relationship expansion is completed through inference rules or graph pattern matching, which is used to discover new relationships or verify the correctness of existing relationships; attribute updates are based on the comparison of old and new versions of the data to ensure that the attribute information in the knowledge graph is up to date. When duplicate entities are detected, the system determines which objects to retain based on the preset source weights. Source weights are dynamically adjusted based on the credibility, update frequency, and historical accuracy of the data source. Entities corresponding to data sources with higher weights are retained first, while low-weight entities are merged or marked as redundant. The integrated data generates graph update data, including adding new nodes, updating edges, and modifying attributes, to provide support for the dynamic update of the knowledge graph.

[0078] S5: Based on the probability model trained with historical data and the conflict resolution algorithm, the graph update data is subjected to credibility calculation and data resolution processing to generate a resolution semantic set.

[0079] Specifically, probabilistic models trained on historical data are central to the system's assessment of data credibility. The system employs machine learning algorithms, such as Bayesian networks or logistic regression models, to train on historical data, learning characteristics such as the credibility of the data source, data consistency, and the rationality of contextual semantics. Bayesian networks construct probabilistic graphical models to represent the conditional probabilistic relationships between different features, enabling updated judgments on data credibility based on new evidence. Logistic regression models predict the credibility of new data by learning the linear relationship between features and credibility. These models enable the system to accurately assess the credibility of graph updates. The credibility score, expressed as a probability value, is used to measure data reliability.

[0080] Conflict resolution algorithms are key to handling semantic conflicts in graph update data. These algorithms are based on the principles of credibility priority, time priority, or semantic consistency. For example, when different data sources provide inconsistent descriptions of the attributes of the same entity, the system prioritizes information provided by the more credible data source; or, based on the timestamp, selects more recent data as a reference; or, through a semantic consistency check, selects the description that is most consistent with the contextual semantics. The conflict resolution process ensures the consistency and accuracy of the knowledge graph. Data processed through credibility calculation and conflict resolution generates a resolved semantic set containing verified semantic information, which supports the subsequent dynamic updates of the knowledge graph and semantic situational awareness.

[0081] S6: Based on dynamic loading technology and semantic web version control mechanism, the data in the semantic set is monitored and adjusted in real time to generate semantic situation awareness data; semantic situation awareness data is used to reflect the changes in entity relationships and semantic conflict resolution status in the knowledge graph in real time.

[0082] Specifically, dynamic loading technology is key to the system's real-time data updates. Similar to the dynamic loading of AARs in Android apps, this technology loads data from resolved semantic sets into the knowledge graph in real time. Dynamic loading ensures real-time data availability while avoiding frequent reconstruction of the knowledge graph's overall structure. The loading process is based on an incremental update mechanism, processing only new or updated data, thereby improving system efficiency and responsiveness.

[0083] The semantic web version control mechanism is an important means for the system to manage the knowledge graph update process. It records the metadata of each update (such as update time, update content, and update source) for traceability and rollback. Version control ensures the stability and maintainability of the knowledge graph. The system monitors the loaded data in real time to detect problems such as semantic conflicts, data anomalies, or update failures. The monitoring system triggers adjustment mechanisms based on predefined thresholds and rules, such as automatically repairing semantic conflicts or marking abnormal data. The adjustment process ensures the real-time and accuracy of the knowledge graph. The monitored and adjusted data generates semantic situational awareness data, which is used to reflect in real time the changes in entity relationships in the knowledge graph, the resolution status of semantic conflicts, and the timeliness of updates, providing support for the dynamic management and intelligent application of the knowledge graph.

[0084] In summary, the knowledge graph data intelligent management method based on semantic web technology provided by the present invention can solve the lag problem of traditional methods in responding to real-time requirements by real-time collection and processing of high-frequency dynamic data streams, thereby realizing dynamic updating and rapid response of knowledge graphs; at the same time, in the process of integrating multi-source heterogeneous data, the method systematically solves problems such as noise interference, semantic conflict and entity ambiguity through technologies such as noise filtering, semantic resolution and knowledge mapping, thereby improving the accuracy, consistency and reliability of the knowledge graph; and uses the graph-embedded knowledge fusion algorithm and credibility evaluation mechanism to balance the timeliness and reliability of the system, thereby ensuring the efficiency and stability of the knowledge graph in large-scale dynamic data processing. Finally, through semantic situation awareness and real-time adjustment mechanism, the knowledge graph can better adapt to the ever-changing data environment, thereby enhancing its application efficiency in fields such as intelligent search, decision support and semantic reasoning, providing users with more accurate, timely and reliable knowledge services, and promoting the widespread application and development of knowledge graph technology in actual scenarios.

[0085] In one embodiment, Figure 2 As shown, S3 of the knowledge graph data intelligent management method based on semantic web technology provided by the present invention specifically includes the following steps:

[0086] S31: Perform noise data identification processing on the second semantic set based on keyword matching and context analysis, apply a preset relevance threshold to filter low-relevance content, and generate graph semantic fragments.

[0087] Specifically, the system uses keyword matching and context analysis techniques to identify noise in the data in the second semantic set. Keyword matching involves scanning the text using a predefined list of keywords to identify key information relevant to the target semantics. Context analysis assesses the semantic integrity and consistency of the data by analyzing the context of the text. For example, the system can identify redundant information unrelated to the target semantics or semantically conflicting content.

[0088] Presetting a relevance threshold is key to the system's ability to filter out low-relevance content. This threshold is dynamically adjusted based on keyword frequency, contextual semantic consistency, and data credibility. The system calculates a relevance score for each piece of data. If the score falls below the preset threshold, it is marked as noise and filtered out. This filtered data generates graph semantic fragments, providing the foundation for subsequent processing.

[0089] In this embodiment, the graph semantic fragments are generated through the following steps:

[0090] S311: performing noise keyword matching processing on the second semantic set based on a preset keyword blacklist to generate an initial filtering result.

[0091] Specifically, the system scans each piece of data in the second semantic set using a predefined keyword blacklist, identifying noise data containing blacklisted keywords. The keyword blacklist is generated based on domain knowledge and historical data statistics, and includes common noise words and irrelevant vocabulary. The system marks and filters out this noise data, generating an initial filtering result.

[0092] S312: Quantify the context relevance of the initial filtering result to generate a relevance score.

[0093] Specifically, the calculation formula for the relevance score is:

[0094]

[0095] Among them, RelevanceScore is the relevance score, TF(t) is the frequency of text word t in the text, DF(t) is the document frequency of text word t in the knowledge graph, N is the total number of documents in the knowledge graph, and T is the vocabulary set of the current text; this formula quantifies the importance of a text word in the current context by combining word frequency and inverse document frequency.

[0096] S313: Filter the relevance scores based on a preset relevance threshold to generate graph semantic segments.

[0097] Specifically, the system filters relevance scores based on a preset relevance threshold. This threshold is dynamically adjusted based on historical data and domain knowledge to distinguish between highly relevant and less relevant content. Data with relevance scores below the threshold is marked as noise and filtered out, ultimately generating graph semantic fragments. Graph semantic fragments contain highly relevant semantic information, providing a foundation for subsequent processing.

[0098] S32: Vectorize the graph semantic segments based on the semantic vector generation model to generate a graph feature vector.

[0099] Specifically, the system uses a semantic vector generation model to vectorize the semantic segments of the graph and generate a graph feature vector. The semantic vector generation model uses word embedding technology to convert words in the text into low-dimensional vector representations, capturing the semantic features and contextual relationships of the words. For example, the system can use word embedding models such as Word2Vec or GloVe to map each word into a vector space of fixed dimension. The calculation formula for the graph feature vector is:

[0100]

[0101] Among them, v is the graph feature vector, w i is the i-th word in the semantic segment of the graph, Embedding(w i ) is the word vector generated by the word embedding model, and n is the number of words in the graph semantic segment. Using this formula, the system averages all the word vectors in the graph semantic segment to generate a comprehensive feature vector that represents the semantic features of the entire semantic segment.

[0102] S33: Perform structured transformation on the graph feature vector based on the knowledge representation method, and map and align it with the entities and relationships of the knowledge graph to generate a graph semantic representation.

[0103] Specifically, the system converts the graph feature vector into a structured form through knowledge representation methods. Knowledge representation methods include technologies such as entity recognition, relationship extraction, and attribute mapping. Entity recognition identifies entities in the text by analyzing the semantic information in the feature vector; relationship extraction identifies the relationship between entities by analyzing the semantic associations between entities; and attribute mapping matches the attribute information in the feature vector with the corresponding attributes in the knowledge graph. The system further maps and aligns the structured feature vector with the entities and relationships in the knowledge graph. Mapping alignment matches the entities, relationships, and attributes in the feature vector with the corresponding elements in the knowledge graph through predefined ontology mapping rules. For example, the system can match entities in the feature vector with entities in the knowledge graph through string similarity or semantic similarity algorithms to ensure that the data can be accurately integrated into the structure of the knowledge graph. The resulting semantic representation of the graph provides support for subsequent knowledge fusion and updating.

[0104] The above-mentioned intelligent management method of knowledge graph data based on semantic web technology can accurately identify and filter out low-correlation noise data through noise keyword matching and context association quantification processing, ensuring that the subsequently processed data has high semantic relevance; at the same time, the system converts semantic fragments into representations in high-dimensional vector space through semantic vectorization processing, which can better capture the semantic and grammatical characteristics of words and provide richer semantic information for the construction of knowledge graphs; through knowledge representation and mapping alignment processing, the graph feature vectors are converted into structured representations compatible with the knowledge graph, ensuring that the semantic fragments can be accurately mapped to entities and relationships in the knowledge graph, thereby improving the accuracy and consistency of the knowledge graph. Through the implementation of the above steps, the present invention can not only improve the construction efficiency of the knowledge graph, but also enhance the application effectiveness of the knowledge graph in fields such as intelligent search, decision support and semantic reasoning, providing users with a more accurate, efficient and reliable knowledge service.

[0105] In one embodiment, S4 of a method for intelligent management of knowledge graph data based on semantic web technology provided by the present invention specifically includes the following steps:

[0106] S41: Vectorize the entities and relationships in the graph semantic representation based on the graph embedding algorithm to generate a set of entities to be fused.

[0107] Specifically, the system vectorizes the entities and relationships in the semantic representation of the graph based on a graph embedding algorithm to generate a set of entities to be fused. The graph embedding algorithm is a technology that converts graph structure data into a low-dimensional vector representation. It can capture the structural characteristics and semantic information of entities and relationships in the knowledge graph. The graph embedding algorithms used by the system, such as Node2Vec or GraphSAGE, learn the adjacency relationships and global structures of entities and relationships in the graph through random walks or graph neural networks. The Node2Vec algorithm combines the advantages of depth-first search and breadth-first search, and can perform flexible random walks in the graph, thereby capturing local and global relationships between entities.

[0108] The GraphSAGE algorithm aggregates feature information from neighboring nodes to generate an embedding vector for each node, enabling better processing of large-scale graph data. Through this graph embedding algorithm, the system transforms complex graph structures into computable vectors, providing a foundation for subsequent similarity calculations and fusion processing. The vectorized entities and relationships constitute the entity set to be fused, providing the foundation for subsequent similarity calculations and fusion processing.

[0109] S42: Based on the embedding vector of the knowledge graph node, the cosine similarity of the entity set to be fused and the existing entities is calculated to generate a candidate entity similarity set; each entity in the candidate entity similarity set is identified to see whether it exceeds the preset duplication threshold. If so, the entities exceeding the threshold are marked as duplicate entities to generate an optimized set to be fused containing duplicate entities.

[0110] Specifically, cosine similarity is a method to measure the similarity between two vectors by calculating the cosine value of the angle between the two vectors to evaluate their similarity. The calculation formula of cosine similarity is:

[0111]

[0112] Where A and b are the embedding vectors of the entity to be fused and the existing entity, respectively. The system calculates the similarity between each entity to be fused and all existing entities, generating a candidate entity similarity set. A preset duplication threshold is determined based on domain knowledge and historical data statistics to distinguish duplicate entities.

[0113] Specifically, the system compares the similarity of each entity in the candidate entity similarity set against a preset duplication threshold to identify duplicate entities that exceed the threshold. For entities that exceed the threshold, the system marks them as duplicates and generates an optimized set to be fused containing the duplicates. This optimized set to be fused retains the parts of the set to be fused that are not duplicates of existing entities, providing support for subsequent attribute merging and graph updates.

[0114] S43: Based on the preset source weights, duplicate entity attributes in the optimized set to be fused are merged to generate graph update data.

[0115] Specifically, the preset source weights are dynamically adjusted based on the credibility, update frequency, and historical accuracy of the data source. The system merges the attribute values ​​of duplicate entities using a weighted average method, calculated as follows:

[0116]

[0117] Among them, S i is the i-th data source of the repeated entity, Weight(S i ) is the data source S i The weight of AttributeValue(S i ) is the data source S iThe system then merges and processes the data to generate graph updates, including adding new nodes, updating edges, and modifying attributes. This updated data is then applied to the knowledge graph to ensure its timeliness and accuracy. This updated data is loaded into the knowledge graph via an incremental update mechanism, avoiding frequent reconstruction of the overall structure and improving system efficiency and responsiveness. Ultimately, by dynamically updating the knowledge graph, the system achieves efficient integration and real-time updates of high-frequency dynamic data streams.

[0118] In one embodiment, S5 of a method for intelligent management of knowledge graph data based on semantic web technology provided by the present invention specifically includes the following steps:

[0119] S51: Calculate the source credibility and time decay factor of the entity in the graph update data to generate an initial credibility score.

[0120] Specifically, source credibility is dynamically assessed based on the historical accuracy, update frequency, and domain authority of the data source. The time decay factor reflects the timeliness of the data; as time passes, the credibility of old data gradually decreases. By combining these factors, the system calculates the initial credibility score InitialTrust(e) for each entity using the following formula:

[0121] InitialTrust(e)=α*SourceTrust(e)+(1-α)*TimeDecay(e)

[0122] Where α is the weight coefficient, SourceTrust(e) is the source trustworthiness of entity e, and TimeDecay(e) is the time decay factor of entity e. Source trustworthiness is assessed based on the historical performance of the data source. For example, a data source that consistently provides accurate information is assigned a higher source trustworthiness. The time decay factor is calculated using an exponential decay function to ensure that the timeliness of the data is properly reflected.

[0123] S52: Perform historical conflict detection based on entity identifiers in the graph update data and historical versions of the knowledge graph to generate a set of high-risk conflict entities.

[0124] Specifically, historical conflict detection identifies potential conflicts by analyzing changes in the attributes and relationships of entities across different versions. For example, the system can detect that an entity has been assigned different attribute values ​​or relationships across different versions. This process involves version control mechanisms. By comparing the differences between different versions, the system identifies entities that have conflicted in historical versions and marks them as high-risk conflict entities. The management of historical versions ensures that the system can trace the historical changes of entities, thereby effectively detecting potential conflicts.

[0125] S53: Perform weighted calculation on the score values ​​in the initial credibility score set, the conflict marks of the high-risk conflict entity set, and the entity confidence of the graph update data to generate a final credibility score.

[0126] Specifically, the calculation formula for the final credibility score is:

[0127] FT(e)=β*InitialTrust(e)+(1-β)*CM(e)+γ*EC(e)

[0128] Where β and γ are weight coefficients, CM(e) is the conflict flag for entity e, and EC(e) is the confidence score for entity e. This formula integrates multiple factors to ensure that the final credibility score fully reflects the trustworthiness of the entity. The conflict flag is based on the results of historical conflict detection, indicating whether the entity has encountered conflicts in previous versions. The entity confidence score reflects the data provider's confidence in the entity information and is typically based on the data's source and verification process.

[0129] S54: Based on the final credibility score, the high-risk conflict entity set is identified and processed to determine whether the final credibility score of each entity is lower than the preset credibility threshold. If so, the entity is deleted from the high-risk conflict entity set; otherwise, the attribute with the highest source weight among the entities in the high-risk conflict entity set is retained to generate a resolution semantic set.

[0130] Specifically, the preset credibility threshold is determined based on domain knowledge and historical data statistics to distinguish the credibility of entities. The system removes entities whose final credibility score is lower than the threshold from the set of high-risk conflict entities. For the retained entities, the system selects the attributes with the highest source weight to retain and generates a set of resolved semantics. The source weight is dynamically adjusted based on the credibility, update frequency and historical accuracy of the data source. In this way, the system can effectively resolve semantic conflicts and ensure the accuracy and consistency of the knowledge graph. This process involves a comprehensive evaluation of the data source to ensure that only the most reliable information is retained, thereby improving the overall quality of the knowledge graph.

[0131] In one embodiment, S6 of a method for intelligent management of knowledge graph data based on semantic web technology provided by the present invention specifically includes the following steps:

[0132] S61: Based on real-time monitoring technology, the data in the resolved semantic set is loaded in batches to generate incremental update instructions.

[0133] Specifically, the system loads the data in the resolved semantic set in batches based on real-time monitoring technology, generating incremental update instructions. This real-time monitoring technology ensures that the system can promptly detect data changes and generate corresponding incremental update instructions based on these changes. Batch loading ensures system efficiency and stability, avoiding system overload caused by loading large amounts of data all at once.

[0134] Specifically, the system divides the data in the semantically resolved set into multiple batches for processing by setting the batch size and loading frequency. Each batch of data is loaded separately into the knowledge graph, generating incremental update instructions to ensure that the knowledge graph can reflect the latest semantic information in a timely manner. Real-time monitoring technology ensures that any data changes can be detected in a timely manner by establishing a continuous communication channel between the data source and the knowledge graph. This technology usually involves an event-driven architecture. When the data source changes, the system will receive a notification and take immediate action. The incremental update mechanism ensures that only the changed data is loaded into the knowledge graph by recording a log of data changes, thereby improving the efficiency and responsiveness of the system.

[0135] S62: Based on the semantic web-driven RDF triple update rules, the instructions in the incremental update instruction set are parsed and processed, and the nodes and edges of the knowledge graph are added, deleted, and modified to generate an updated knowledge graph.

[0136] Specifically, RDF (Resource Description Framework) triples are the core data model of the Semantic Web. They consist of a subject, a predicate, and an object, and are used to represent semantic relationships between entities. The system parses incremental update instructions, identifies the RDF triples that need to be updated, and performs corresponding operations based on predefined update rules. For example, when an instruction requires adding a new entity, the system creates a new node in the knowledge graph and establishes connections with other nodes based on the predicate and object information in the triple. For deletion operations, the system removes the specified node or edge and adjusts related connections to maintain the integrity of the graph. Modification operations involve updating the properties of nodes or the relationship types of edges. This process ensures the semantic consistency and structural integrity of the knowledge graph during the update process by strictly adhering to the semantic specifications and update rules of RDF triples, enabling it to accurately reflect the latest semantic information.

[0137] S63: Based on the updated knowledge graph, the entity-relationship-time dimension analysis and processing are performed through the multi-dimensional semantic association matrix construction technology to generate semantic situational awareness data that supports real-time decision-making.

[0138] Specifically, a multidimensional semantic association matrix is ​​a mathematical model used to capture complex semantic relationships between entities. It not only considers the static associations between entities and relationships but also introduces a temporal dimension to reflect how these associations change over time. By constructing such a matrix, the system can comprehensively analyze the strength, directionality, and dynamic trends of semantic associations between entities in the knowledge graph. In the entity dimension, the matrix records the attributes and states of each entity; in the relationship dimension, it captures the various semantic relationships between entities and their strength; and in the temporal dimension, it reflects the evolution of these entities and relationships over time. Through multidimensional analysis, the system can identify influential entities and relationships, as well as their changing patterns at different points in time. Ultimately, the system generates semantic situational awareness data that intuitively presents key semantic situations within the knowledge graph, providing strong support for real-time decision-making. This process, through continuous optimization of the matrix construction algorithm and analysis model, improves the accuracy and timeliness of semantic situational awareness data, ensuring the reliability and effectiveness of decision support.

[0139] In one embodiment, S2 of a method for intelligent management of knowledge graph data based on semantic web technology provided by the present invention specifically includes the following steps:

[0140] S21: Based on the data source identifier of each data in the first semantic set, the corresponding credibility score is extracted from a preset historical data credibility library to generate a data source credibility value.

[0141] Specifically, based on the data source identifier of each data in the first semantic set, the system extracts the corresponding credibility score from the preset historical data credibility library to generate a data source credibility value. The data source identifier is a unique identifier used to distinguish different data sources. The system uses this identifier to search for the corresponding credibility score in the historical data credibility library. The historical data credibility library is a database that stores the historical performance of each data source, which records the accuracy, consistency and reliability of the data provided by each data source in the past. The credibility score is dynamically calculated based on the historical performance of the data source, usually using a weighted average method, which comprehensively considers the performance of multiple dimensions of the data source, such as accuracy, update frequency and domain authority.

[0142] S22: Calculate the difference between the timestamp of each data in the first semantic set and the current time to generate a time attenuation factor.

[0143] Specifically, the calculation formula of the time attenuation factor is:

[0144]

[0145] Among them, TimeDecay is the time decay factor, α is the time decay coefficient, t current is the current time, t record Record timestamps for data. The time decay factor reflects the timeliness of data; as time passes, the credibility of older data decreases. The time decay coefficient α determines the rate of decay and is typically adjusted based on the data type and domain knowledge.

[0146] S23: Generate entity confidence based on the entity recognition probability value output by natural language processing in the first semantic set.

[0147] Specifically, entity recognition is an important task in natural language processing. It uses machine learning models to identify entities in text and outputs a probability value to indicate the confidence level of the recognition. This probability value reflects the model's confidence in the entity recognition result and is usually between 0 and 1, with higher values ​​indicating higher confidence.

[0148] S24: Prioritize the data source credibility, time decay factor, and entity confidence to generate a priority weight.

[0149] Specifically, the calculation formula for priority weight is:

[0150] P=w1*SourceCred+w2*TimeDecay+w3*EntityConf

[0151] Where P is the priority weight, SourceCred is the data source credibility, TimeDecay is the time decay factor, EntityConf is the entity confidence, and w1, w2, and w3 are dynamic weight coefficients. Dynamic weight coefficients are adjusted according to different application scenarios and requirements to ensure that the priority weight accurately reflects the overall quality of the data.

[0152] S25: Sort each piece of data in the first semantic set in descending order based on the priority weight to generate a second semantic set.

[0153] Specifically, descending sorting ensures that data with higher priority weights are ranked first, allowing them to be prioritized in subsequent processing, improving system efficiency and accuracy. The sorting process uses efficient sorting algorithms, such as quick sort or merge sort, to ensure efficiency and stability. In this way, the system ensures that the processed data is rigorously screened and sorted, thereby improving the quality and efficiency of knowledge graph updates.

[0154] The above-mentioned method for intelligent management of knowledge graph data based on semantic web technology can achieve efficient and accurate knowledge graph data processing, ensuring the credibility and timeliness of the data. By extracting credibility scores from a preset historical data credibility library, a reliable basic assessment is provided for each data item, ensuring the reliability of data processing. Calculating the time decay factor can dynamically adjust the credibility of the data, ensuring that the timeliness of the data is accurately considered, making data processing more in line with actual needs. At the same time, using the entity recognition probability value output by natural language processing to generate entity confidence further enhances the credibility assessment of the data and makes data processing more accurate. Through the priority weight calculation formula, the data source credibility, time decay factor and entity confidence are comprehensively considered to generate the priority weight of each data item, ensuring the scientific and dynamic nature of data processing and making the data priority allocation more reasonable. Finally, by sorting the data in descending order to generate the second semantic set, the efficiency and orderliness of data processing are ensured, providing a solid foundation for the subsequent knowledge graph construction. This series of steps can effectively solve the noise interference, semantic conflict and entity ambiguity problems of multi-source heterogeneous data, enhance the adaptability and stability of knowledge graphs in dynamic data environments, provide high-quality knowledge services for applications such as intelligent search, decision support and semantic reasoning, and significantly improve the actual application efficiency of knowledge graphs.

[0155] Preferably, if Figure 3 As shown, the present invention also provides a knowledge graph data intelligent management system 700 based on semantic web technology, which is configured with the following modules:

[0156] A data acquisition processing module 710 is configured to acquire real-time text data from a high-frequency data stream based on a preset sampling frequency, and segment and extract semantic information from the real-time text data based on a natural language processing method to generate a first semantic set;

[0157] The time attribute processing module 720 is used to extract the timestamp attribute of each data in the first semantic set based on the predefined time priority rule, and perform time sorting processing on the data with repeated or missing timestamps to generate a second semantic set;

[0158] A graph semantic integration module 730 is configured to perform noise filtering, semantic resolution, and knowledge mapping on the second semantic set based on semantic integration technology to generate a graph semantic representation compatible with the knowledge graph;

[0159] Graph data integration module 740, which is used to extract data based on the graph semantic representation of the graph embedded knowledge fusion algorithm, integrate it with the nodes and edges in the knowledge graph, and determine the retained objects based on the preset source weight when duplicate entities appear, and generate graph update data;

[0160] Credibility calculation and resolution module 750, used to perform credibility calculation and data resolution processing on the graph update data based on the probability model trained with historical data and the conflict resolution algorithm, and generate a resolution semantic set;

[0161] The real-time monitoring and adjustment module 760 is used to monitor and adjust the data in the semantic set in real time based on dynamic loading technology and semantic web version control mechanism to generate semantic situation awareness data; the semantic situation awareness data is used to reflect the changes in entity relationships and the status of semantic conflict resolution in the knowledge graph in real time.

[0162] In summary, the knowledge graph data intelligent management system based on semantic web technology provided by the present invention can solve the lag problem of traditional methods in responding to real-time requirements by real-time collection and processing of high-frequency dynamic data streams, thereby realizing dynamic updating and rapid response of knowledge graphs; at the same time, in the process of integrating multi-source heterogeneous data, the method systematically solves problems such as noise interference, semantic conflict and entity ambiguity through technologies such as noise filtering, semantic resolution and knowledge mapping, thereby improving the accuracy, consistency and reliability of the knowledge graph; and uses graph-embedded knowledge fusion algorithms and credibility assessment mechanisms to balance the timeliness and reliability of the system, thereby ensuring the efficiency and stability of the knowledge graph in large-scale dynamic data processing. Finally, through semantic situational awareness and real-time adjustment mechanisms, the knowledge graph can better adapt to the ever-changing data environment, thereby enhancing its application efficiency in fields such as intelligent search, decision support and semantic reasoning, providing users with more accurate, timely and reliable knowledge services, and promoting the widespread application and development of knowledge graph technology in practical scenarios.

[0163] Preferably, the graph semantic integration module 730 is configured with the following units:

[0164] A noise filtering unit 731 is configured to perform noise data identification processing on the second semantic set based on keyword matching and context analysis, filter low-relevance content by applying a preset relevance threshold, and generate a graph semantic segment;

[0165] Preferably, the noise filtering unit 731 is configured with the following subunits:

[0166] The keyword matching filtering subunit 7311 is configured to perform noise keyword matching processing on the second semantic set based on a preset keyword blacklist to generate an initial filtering result;

[0167] The context relevance quantification subunit 7312 is used to quantify the context relevance of the initial filtering results to generate a relevance score;

[0168] The semantic segment generation subunit 7313 is used to filter the relevance scores based on a preset relevance threshold to generate graph semantic segments.

[0169] A graph vector generation unit 732 is used to perform vectorization processing on the graph semantic segments based on a semantic vector generation model to generate a graph feature vector;

[0170] The semantic representation generation unit 733 is used to perform structured transformation processing on the graph feature vector based on the knowledge representation method, and to perform mapping and alignment processing with the entities and relationships of the knowledge graph to generate a graph semantic representation.

[0171] Preferably, the atlas data integration module 740 is configured with the following units:

[0172] A vectorization processing unit 741 is used to perform vectorization processing on entities and relationships in the graph semantic representation based on a graph embedding algorithm to generate a set of entities to be fused;

[0173] Similarity calculation and marking unit 742 is used to calculate the cosine similarity between the entity set to be fused and the existing entities based on the embedding vector of the knowledge graph node to generate a candidate entity similarity set; identify whether each entity in the candidate entity similarity set exceeds a preset duplication threshold, and if so, mark the entities exceeding the threshold as duplicate entities to generate an optimized set to be fused containing duplicate entities;

[0174] The duplicate entity processing unit 743 is used to merge the duplicate entity attributes in the optimized set to be merged based on the preset source weights to generate graph update data.

[0175] Preferably, the credibility calculation and resolution module 750 is configured with the following units:

[0176] An initial score calculation unit 751 is used to calculate the source credibility and time decay factor of the entity in the graph update data to generate an initial credibility score;

[0177] A historical conflict detection unit 752 is used to perform historical conflict detection based on entity identifiers in the graph update data and historical versions of the knowledge graph to generate a set of high-risk conflict entities;

[0178] A final score calculation unit 753 is configured to perform weighted calculation based on the score values ​​in the initial credibility score set, the conflict flags of the high-risk conflict entity set, and the entity confidence of the graph update data to generate a final credibility score;

[0179] The conflict entity processing unit 754 is used to identify and process the high-risk conflict entity set based on the final credibility score, and determine whether the final credibility score of each entity is lower than the preset credibility threshold. If so, the entity is deleted from the high-risk conflict entity set; otherwise, the attribute with the highest source weight among the entities in the high-risk conflict entity set is retained to generate a resolution semantic set.

[0180] Preferably, the real-time monitoring and adjustment module 760 is configured with the following units:

[0181] The data loading unit 761 is used to load the data in the resolved semantic set in batches based on real-time monitoring technology and generate incremental update instructions;

[0182] The instruction parsing and graph updating unit 762 is used to parse the instructions in the incremental update instruction set based on the RDF triple update rules driven by the semantic web, add, delete, and modify the nodes and edges of the knowledge graph, and generate an updated knowledge graph;

[0183] The situation awareness data generation unit 763 is used to perform entity-relationship-time dimension analysis and processing based on the updated knowledge graph through multi-dimensional semantic association matrix construction technology to generate semantic situation awareness data that supports real-time decision-making.

[0184] Preferably, the time attribute processing module 720 is configured with the following units:

[0185] The source credibility extraction unit 721 is used to extract the corresponding credibility score from the preset historical data credibility library based on the data source identifier of each data in the first semantic set to generate a data source credibility value;

[0186] A time decay factor calculation unit 722 is configured to calculate the difference between the timestamp of each data item in the first semantic set and the current time to generate a time decay factor;

[0187] An entity confidence generating unit 723, configured to generate an entity confidence based on the entity recognition probability value output by the natural language processing in the first semantic set;

[0188] The priority weight calculation unit 724 is used to perform priority calculation on the data source credibility, time decay factor and entity confidence to generate a priority weight;

[0189] The semantic set sorting unit 725 is configured to sort each piece of data in the first semantic set in descending order based on the priority weight to generate a second semantic set.

[0190] In one embodiment, the present application also provides a computer device including a memory and a processor, wherein the memory stores a computer program, and the processor implements the above-mentioned method for intelligent management of knowledge graph data based on semantic web technology when executing the computer program.

[0191] In one embodiment, the present application also provides a computer-readable storage medium on which a computer program is stored. When the computer program is executed by a processor, it implements the above-mentioned knowledge graph data intelligent management method based on semantic web technology.

[0192] In the description of this specification, the reference terms "one embodiment," "some embodiments," "example," "specific example," or "some examples" mean that the specific features, structures, materials, or characteristics described in conjunction with the embodiment or example are included in at least one embodiment or example of the present application. Moreover, the specific features, structures, materials, or characteristics described may be combined in any appropriate manner in any one or more embodiments or examples. In addition, those skilled in the art may combine and integrate different embodiments or examples described in this specification, as well as features of different embodiments or examples, unless they are mutually inconsistent.

[0193] For the device embodiments, since they basically correspond to the method embodiments, the relevant parts can be referred to the partial description of the method embodiments. The device embodiments described above are merely illustrative, wherein the components described as separate parts may or may not be physically separated, and the parts displayed as units may or may not be physical units, that is, they may be located in one place, or they may be distributed on multiple network units. Some or all of the modules can be selected according to actual needs to achieve the purpose of the disclosed solution. A person of ordinary skill in the art can understand and implement it without expending creative work.

[0194] The above description is merely a specific embodiment of the present application, but the scope of protection of the present application is not limited thereto. Any person skilled in the art can easily conceive of various modifications or substitutions within the technical scope disclosed in this application, and such modifications or substitutions should be included within the scope of protection of the present application. Therefore, the scope of protection of the present application should be based on the scope of protection of the claims.

Claims

1. A method for intelligent management of knowledge graph data based on semantic web technology, characterized in that: The following steps are involved: S1: acquiring real-time text data from a high-frequency data stream based on a preset sampling frequency, and segmenting and extracting semantic information from the real-time text data based on a natural language processing method to generate a first semantic set; S2: extracting the timestamp attribute of each data in the first semantic set based on a predefined time priority rule, and performing time sorting processing on the data with duplicate or missing timestamps to generate a second semantic set; S3: performing noise filtering, semantic resolution, and knowledge mapping on the second semantic set based on semantic integration technology to generate a graph semantic representation compatible with the knowledge graph; S4: The graph semantic representation of the graph is used for data extraction based on the graph embedding knowledge fusion algorithm, and the data is integrated with the nodes and edges in the knowledge graph. When repeated entities appear, the retained objects are judged according to the preset source weights to generate graph update data. S5: Perform credibility calculation and data resolution processing on the graph update data based on the probability model trained with historical data and the conflict resolution algorithm to generate a resolution semantic set; S6: Based on dynamic loading technology and semantic web version control mechanism, the data in the resolved semantic set is monitored and adjusted in real time to generate semantic situation awareness data; the semantic situation awareness data is used to reflect the entity relationship changes and semantic conflict resolution status in the knowledge graph in real time.

2. The method according to claim 1, characterized in that The S3 includes: S31: performing noise data identification processing on the second semantic set based on keyword matching and context analysis, applying a preset relevance threshold to filter low-relevance content, and generating a graph semantic segment; S32: Vectorize the graph semantic segment based on the semantic vector generation model to generate a graph feature vector. The calculation formula of the graph feature vector is: Among them, v is the graph feature vector, w i is the i-th word in the semantic segment of the graph, Embedding(w i ) is the word vector generated by the word embedding model, and n is the number of words in the semantic segment of the graph; S33: Performing structured transformation processing on the graph feature vector based on the knowledge representation method, and mapping and aligning it with the entities and relationships of the knowledge graph to generate the graph semantic representation.

3. The method according to claim 2, characterized in that The S31 includes: S311: performing noise keyword matching processing on the second semantic set based on a preset keyword blacklist to generate an initial filtering result; S312: Quantify the context relevance of the initial filtering result to generate a relevance score. The calculation formula of the relevance score is: Among them, RelevanceScore is the relevance score, TF(t) is the frequency of text word t in the text, DF(t) is the document frequency of text word t in the knowledge graph, N is the total number of documents in the knowledge graph, and T is the vocabulary set of the current text; S313: Filter the relevance scores based on a preset relevance threshold to generate graph semantic segments.

4. The method according to claim 1, wherein The S4 includes: S41: performing vectorization processing on entities and relationships in the graph semantic representation based on a graph embedding algorithm to generate a set of entities to be fused; S42: Calculating the cosine similarity between the entity set to be fused and the existing entities based on the embedding vectors of the knowledge graph nodes to generate a candidate entity similarity set; identifying whether each entity in the candidate entity similarity set exceeds a preset duplication threshold, and if so, marking the entities exceeding the threshold as duplicate entities to generate an optimized set to be fused containing duplicate entities; S43: merging repeated entity attributes in the optimized set to be fused based on preset source weights to generate graph update data.

5. The method according to claim 1, wherein The S5 includes: S51: Calculating the source credibility and time decay factor of the entity in the graph update data to generate an initial credibility score; S52: Perform historical conflict detection based on entity identifiers in the graph update data and historical versions of the knowledge graph to generate a set of high-risk conflict entities; S53: performing weighted calculation on the score values ​​in the initial credibility score set, the conflict marks of the high-risk conflict entity set, and the entity confidence of the graph update data to generate a final credibility score; S54: Based on the final credibility score, the high-risk conflict entity set is identified and processed to determine whether the final credibility score of each entity is lower than a preset credibility threshold. If so, the entity is deleted from the high-risk conflict entity set; otherwise, the attribute with the highest source weight among the entities in the high-risk conflict entity set is retained to generate a resolution semantic set.

6. The method according to claim 5, characterized in that The S6 includes: S61: Based on real-time monitoring technology, the data in the resolved semantic set is loaded in batches to generate incremental update instructions; S62: Based on the semantic web-driven RDF triple update rule, the instructions in the incremental update instruction set are parsed and processed, and the nodes and edges of the knowledge graph are added, deleted, and modified to generate an updated knowledge graph; S63: Based on the updated knowledge graph, entity-relationship-time dimension analysis and processing are performed through multi-dimensional semantic association matrix construction technology to generate semantic situation awareness data that supports real-time decision-making.

7. The method according to any one of claims 1 to 6, characterized in that The S2 includes: S21: Based on the data source identifier of each piece of data in the first semantic set, extract the corresponding credibility score from a preset historical data credibility library to generate a data source credibility value; S22: Calculate the difference between the timestamp of each data item in the first semantic set and the current time to generate a time decay factor. The calculation formula of the time decay factor is: Among them, TimeDecay is the time decay factor, α is the time decay coefficient, t current is the current time, t record Record timestamps for data; S23: Generate entity confidence based on the entity recognition probability value output by natural language processing in the first semantic set; S24: Priority calculation is performed on the data source credibility, the time decay factor, and the entity confidence to generate a priority weight. The calculation formula of the priority weight is: P=w1*SourceCred+w2*TimeDecay+w3*EntityConf Among them, P is the priority weight, SourceCred is the data source credibility, TimeDecay is the time decay factor, EntityConf is the entity confidence, w1, w2, w3 are dynamic weight coefficients; S25: Sort each piece of data in the first semantic set in descending order based on the priority weight to generate a second semantic set.

8. A knowledge graph data intelligent management system based on semantic web technology, characterized by: The system comprises: a data acquisition processing module, configured to acquire real-time text data from a high-frequency data stream based on a preset sampling frequency, and segment and extract semantic information from the real-time text data based on a natural language processing method to generate a first semantic set; a time attribute processing module, configured to extract a timestamp attribute from each piece of data in the first semantic set based on a predefined time priority rule, and perform time sorting processing on data with duplicate or missing timestamps to generate a second semantic set; A graph semantic integration module, configured to perform noise filtering, semantic resolution, and knowledge mapping processing on the second semantic set based on semantic integration technology to generate a graph semantic representation compatible with the knowledge graph; A graph data integration module is used to extract data based on the graph semantic representation of the graph embedded knowledge fusion algorithm, integrate it with the nodes and edges in the knowledge graph, and determine the retained objects based on the preset source weights when duplicate entities appear, and generate graph update data; A credibility calculation and resolution module is used to perform credibility calculation and data resolution processing on the graph update data based on a probability model trained with historical data and a conflict resolution algorithm, and generate a resolution semantic set; A real-time monitoring and adjustment module is used to monitor and adjust the data in the resolved semantic set in real time based on dynamic loading technology and semantic web version control mechanism to generate semantic situation awareness data; the semantic situation awareness data is used to reflect the entity relationship changes and semantic conflict resolution status in the knowledge graph in real time.

9. A computer device comprising a memory and a processor, wherein the memory stores a computer program, wherein: When the processor executes the computer program, the method according to any one of claims 1 to 7 is implemented.

10. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the method according to any one of claims 1 to 7 is implemented.

Citation Information

Cited By

  • Power equipment fault prediction knowledge graph updating method

    CN121031766A

  • Mineral processing equipment alarm method and system based on semantic analysis

    CN121052258A

  • A beneficiation equipment alarm method and system based on semantic analysis

    CN121052258B

  • Method and device for constructing interactive electric power material knowledge unit

    CN121094097A

  • Self-adaptive dynamic knowledge updating method and system in combination with spatio-temporal data

    CN121144328A