Large language model retrieval enhancement generation method for cyberspace security emergency intelligent analysis
By building a search and generation mechanism enhanced by professional field knowledge graph and design knowledge, the limitations of large language models in professional field query are solved, and a professional field knowledge Q&A with high accuracy, coverage and interpretability is achieved.
Patent Information
- Application Number
- CN202510164020.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-14
- Publication Date
- 2025-06-06
AI Technical Summary
Large language models show limitations in professional fields or highly specialized queries, including limited knowledge timeliness, lack of professional knowledge system and data segmentation, and insufficient context perception.
By building a professional field knowledge graph, designing a knowledge-enhanced search strategy and generation mechanism to achieve structured expression, intelligent retrieval and accurate generation of professional field knowledge. Specific steps include multi-source data acquisition and preprocessing, knowledge graph construction, graph-based hierarchical retrieval strategies, and a knowledge-enhanced generation framework.
It significantly improves the accuracy of knowledge questions and answers in professional fields, improves the coverage of knowledge retrieval, and enhances the interpretability and professionalism of answers, realizing the structured expression and intelligent correlation of knowledge.
Smart Images

Figure CN120104649A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of natural language processing, and in particular to the field of retrieval enhancement and generation technology based on knowledge graphs, and specifically to a method for enhancing the retrieval and generation capabilities of a large language model in professional knowledge question and answering by using knowledge graphs. Background Art
[0002] Large language models have excellent capabilities in the field of natural language processing (NLP), including reasoning, question answering, programming, and text generation (He Zhe, Zeng Runxi, Qin Wei, et al. Social impact and governance of new generation artificial intelligence technologies such as ChatGPT [J]. E-Government, 2023, (04): 2-24.). However, large language models often encounter many obstacles when handling practical tasks. Large language models often encounter limitations in context length and tend to ignore text at the center of the context compared to text at the beginning or conclusion. Large language models require a lot of time and computing resources during each training iteration, resulting in delayed knowledge updates.
[0003] In order to further improve the performance of large language models and reduce computational costs, researchers have proposed a variety of optimization techniques. Hu et al. (Hu EJ, Shen Y, Wallis P, et al. Lora: Low-rank adaptation of large language models [J]. arXiv preprint arXiv: 2106.09685, 2021.) proposed the LoRA (low-rank adaptation) technology, which significantly reduced the number of trainable parameters in downstream tasks by injecting a trainable low-rank decomposition matrix into each layer of the Transformer architecture. In terms of prompt engineering, Wei et al. (Wei J, Wang X, Schuurmans D, et al. Chain-of-thought prompting elicits reasoning in large language models [J]. Advances in neural information processing systems, 2022, 35: 24824-24837.) proposed the Chain-of-Thought (CoT) technology, which significantly improved the accuracy of the model in multiple reasoning benchmarks by adding thought chain prompts. Self-consistency proposed by Wang et al. (Wang X, Wei J, Schuurmans D, et al. Self-consistency improves chain of thought reasoning in language models[J].arXivpreprint arXiv:2203.11171, 2022.) aims to "replace the naive greedy decoding method used in chain thought prompts". The idea is to sample multiple different reasoning paths through a few-shot thought chain and use the generated results to select the most consistent answer. This helps improve the performance of thought chain prompts in tasks involving arithmetic and common sense reasoning. For complex tasks that require exploration or anticipatory strategies, traditional or simple prompting techniques are not enough.Yao et al. (Yao S, Yu D, Zhao J, et al. Tree of thoughts: Deliberate problem solving with large language models [J]. Advances in Neural Information Processing Systems, 2024, 36.) proposed a Tree of Thoughts (ToT) framework, which is based on a summary of thought chain prompts to guide the language model to explore thinking as an intermediate step to solve general problems. Despite these optimization techniques, large language models still show obvious limitations, especially in dealing with specific domains or highly specialized queries (Kandpal N, Deng H, Roberts A, et al. Large language models struggle to learn long-tail knowledge [C] / / International Conference on Machine Learning. PMLR, 2023: 15696-15707.). A common problem is the generation of incorrect information or "hallucinations" (Huang L, Yu W, Ma W, et al. Survey on hallucination in large language models: Principles, taxonomy, challenges, and open questions [J]. ACM Transactions on Information Systems, 2023.), especially when queries exceed the model's training data or require the latest information. These shortcomings highlight the impracticality of deploying large language models as black-box solutions in actual production environments without additional safeguards. A promising approach to mitigate these limitations is Retrieval-Augmented Generation (RAG), which integrates external data retrieval into the generation process, thereby enhancing the model's ability to provide accurate and relevant responses.
[0004] The NaiveRAG research paradigm proposed by Ma et al. (Ma X, Gong Y, He P, et al. Query rewriting for retrieval-augmented large language models[J].arXiv preprint arXiv:2305.14283,2023.) represents the earliest approach and gained attention shortly after ChatGPT was widely adopted. Naive RAG follows the traditional process, including indexing, retrieval, and generation. It is also known as the "retrieval-read" framework. However, Naive RAG faces significant challenges in three key areas: "retrieval", "generation", and "augmentation". Advanced Retrieval Augmented Generation (Advanced RAG) (Ilin I. Advanced rag techniques: an illustrated overview[EB / OL]. (2023)) is proposed with targeted enhancements to address the shortcomings of Naive RAG. In terms of retrieval quality, Advanced RAG implements pre-retrieval and post-retrieval strategies. To address the indexing challenges encountered by Naive RAG, Advanced RAG improves its indexing method using techniques such as sliding windows, fine-grained segmentation, and metadata. It also introduces various methods to optimize the retrieval process.
[0005] In order to improve the capabilities of traditional retrieval enhancement systems. Edge et al. proposed GraphRAG (Edge D, Trinh H, Cheng N, et al. From local to global: A graph rag approach to query-focused summarization [J]. arXiv preprint arXiv: 2404.16130, 2024.), which uses a large language model to build an entity-based knowledge graph and pre-generates community summaries of related entity groups, thereby capturing local and global relationships in a document collection, thereby enhancing the query-centric summarization (QFS) task. Guo et al. (Guo Z, Xia L, Yu Y, et al. Lightrag: Simple and fast retrieval-augmented generation [J]. 2024.) proposed a graph-based RAG system, LightRAG, which significantly improved the performance of information retrieval and generation through graph-enhanced text indexing, dual-layer retrieval paradigm, and incremental update algorithm. Yang et al. (Yang R, Yang B, Feng A, et al. Graphusion: A RAG Framework for Knowledge Graph Construction with a Global Perspective [J]. arXiv preprint arXiv: 2410.17600, 2024.) proposed the Graphusion framework, which constructs a globally consistent scientific knowledge graph from free text through three steps: seed entity generation, candidate triple extraction, and knowledge graph fusion. It significantly improves the performance of question-answering tasks, especially in complex question-answering tasks in educational scenarios, and provides new ideas and solutions for the construction and application of knowledge graphs in the scientific field.
[0006] Although large language models have developed from the basic Transformer architecture to the present day, significant progress has been made in model architecture, training methods, and application technologies. However, when applied in professional fields, they still face challenges such as limited knowledge timeliness and lack of professional knowledge systems. Although retrieval-enhanced generation technology provides a feasible solution to these problems, existing methods still have limitations such as data segmentation and insufficient context awareness when processing professional field data. Summary of the invention
[0007] The purpose of the present invention is to overcome the shortcomings of the prior art and provide a large language model retrieval enhancement generation method for cyberspace security emergency intelligent analysis, which realizes the structured expression, intelligent retrieval and accurate generation of professional domain knowledge by constructing a professional domain knowledge graph and designing a knowledge enhancement retrieval strategy and generation mechanism. To achieve the above purpose, the technical solution adopted by the present invention is as follows:
[0008] A large language model retrieval enhancement generation method for cyberspace security emergency intelligent analysis includes the following steps:
[0009] Step 1: Knowledge base construction
[0010] Establish basic data sets through multi-source data collection and preprocessing. The collection sources include threat intelligence, attack logs, research reports, etc., and use API collection, web crawlers, log aggregation and other mechanisms to achieve automated collection; clean and preprocess the collected data, including data quality inspection, data cleaning, data standardization and feature processing, to ensure data quality and consistency;
[0011] Step 2: Knowledge graph construction
[0012] First, we define the entity type system and divide entities into five core concepts: threat entity, attack entity, vulnerability entity, asset entity, and defense entity. At the same time, we define four core relationship types: attack relationship, threat relationship, vulnerability relationship, and defense relationship. We improve the quality of knowledge graph construction through entity recognition and relationship extraction methods based on large language models. Finally, we achieve effective knowledge integration through entity merging and relationship fusion mechanisms.
[0013] Step 3: Knowledge Retrieval
[0014] Design a graph-based hierarchical retrieval strategy. First, the user's natural language query is converted into a structured query representation through the query understanding and conversion module. Then, the path retrieval algorithm is used to retrieve relevant knowledge paths in the knowledge graph, involving the optimization settings of key parameters such as path depth, number of retrieval results, and similarity threshold.
[0015] Step 4: Answer Generation
[0016] Develop a knowledge-enhanced generation framework that transforms the retrieved discrete knowledge into coherent and accurate answers through contextual organization and multi-round reasoning. At the same time, design a multi-dimensional quality assessment and optimization mechanism, including three key components: consistency verification, completeness assessment, and explainability enhancement, to ensure the professionalism and reliability of the generated answers.
[0017] Preferably, the data preprocessing in step 1 includes four stages: data quality check, data cleaning, data standardization and feature processing, and the data is checked in a cycle until the quality standard is met.
[0018] Preferably, the entity merging in step 2 is performed by a similarity calculation function: Sim(e 1 ,e 2 )=LLM(Φ(e 1 ,e 2 )), where Φ(e 1 ,e 2 ) is a specially designed prompt template function.
[0019] Preferably, the relational fusion process in step 2 can be formalized as a three-stage mapping function: R' = I(C(R 1 ∪R 2 ),ctx), where R 1 and R 2 is the set of relations to be fused, ctx is the context information, C is the conflict resolution function, and I is the relation reasoning function.
[0020] The beneficial effects of the present invention are:
[0021] Significantly improved the accuracy of professional knowledge question answering. In the experimental evaluation, the accuracy of the method of the present invention in professional fields reached 86.7%, which is 18.5 percentage points higher than that of directly using a large language model and 12.4 percentage points higher than that of traditional retrieval enhancement generation methods;
[0022] Effectively improve the coverage of knowledge retrieval. The retrieval knowledge coverage of the present invention reaches 81.2%, which is 11.7 percentage points higher than the traditional method, indicating that the retrieval strategy based on knowledge graph can more comprehensively acquire and integrate relevant professional knowledge;
[0023] Enhanced the explainability and professionalism of answers. The present invention achieved a high score of 4.35 in the explanation quality score. Through a multi-dimensional quality assessment and optimization mechanism, it ensured the professionalism and reliability of the generated answers; and realized the structured expression and intelligent association of knowledge. By constructing a professional domain knowledge graph, the present invention overcomes the limitations of traditional methods in flat data representation and context awareness; it has good scalability and adaptability. The modular design adopted by the present invention enables the system to quickly adapt to different professional fields. By adjusting the entity type system and relationship definition, it can flexibly respond to the rapid update of professional domain knowledge and changes in conceptual relationships. BRIEF DESCRIPTION OF THE DRAWINGS
[0024] Figure 1 Retrieval enhancement generation system framework diagram;
[0025] Figure 2 Data cleaning and preprocessing process;
[0026] Figure 3 Entity recognition and extraction prompt word design;
[0027] Figure 4 Design of prompt words for entity relationship recognition;
[0028] Figure 5 Entity similarity prompt template;
[0029] Figure 6 Entity merger prompt template;
[0030] Figure 7 Relationship conflict reminder template;
[0031] Figure 8 Relational reasoning prompt template;
[0032] Fig. 9 Query understanding and conversion hint word structure;
[0033] Fig.10 Enhanced generation of prompt word templates;
[0034] Fig.11 Cyber attack detection case;
[0035] Fig.12 Information security governance case. DETAILED DESCRIPTION
[0036] The present invention is further described below in conjunction with the accompanying drawings and embodiments.
[0037] This patent proposes a retrieval enhancement generation method based on knowledge graph, which enhances the retrieval effect by establishing a professional field knowledge graph and designs a special knowledge enhancement response generation mechanism. The specific implementation of the present invention is described in detail below in conjunction with the accompanying drawings.
[0038] like Figure 1 As shown, the overall framework of the present invention includes two modules: knowledge base construction and knowledge enhanced answer generation. The knowledge base construction module is responsible for the collection, preprocessing and knowledge graph construction of professional field data; the knowledge enhanced answer generation module includes two sub-modules: knowledge retrieval and answer generation, which complete the retrieval of relevant knowledge and the generation of professional answers respectively.
[0039] In terms of knowledge base construction, the present invention mainly faces challenges such as the multi-source heterogeneity of professional data and the complexity of conceptual relationships. This requires the system to be able to effectively process professional data from different sources and formats, and accurately express the complex relationships between entities through knowledge graphs. To this end, the present invention designs a complete data processing flow and a knowledge graph construction method based on a large language model to build a high-quality professional knowledge base.
[0040] In terms of knowledge-enhanced answer generation, the core challenges of this patent include accurate retrieval of professional knowledge, integration of multi-source knowledge, and professional guarantee of answer generation. The system needs to design an efficient knowledge retrieval strategy while ensuring the relevance of the retrieval results to the user's query. In addition, the system needs to maintain the accuracy of professional terminology during answer generation and provide clear explanations and reasoning processes.
[0041] In response to the above challenges, this patent has taken a number of key technical measures. First, a large language model is introduced to assist entity recognition and relationship extraction in the knowledge base construction process, which improves the quality of knowledge graph construction. Secondly, a graph-based hierarchical retrieval strategy is designed to enhance the retrieval effect of professional knowledge. Finally, through the optimized prompt engineering design, the professionalism and explainability of the answers generated by the system are improved.
[0042] Data preprocessing module
[0043] In the retrieval enhancement generation system, building a high-quality cybersecurity knowledge base is a core step. Knowledge base construction enhances the retrieval system by segmenting documents into smaller, more manageable segments. This strategy can quickly identify and access relevant information without analyzing the entire document.
[0044] Cybersecurity data collection is the first step in building a knowledge base, laying an important data foundation for retrieval enhancement generation. Specifically, a good data source architecture can help ensure the diversity, timeliness and coverage of the knowledge base content, and meet the needs of cybersecurity incident processing. After determining the above approaches, use different collection mechanisms to automatically collect data from different channels, and perform preliminary cleaning of the data during collection.
[0045] For network security data, data cleaning and preprocessing is one of the important links in data processing, which aims to remove noise, redundant information and unnecessary data in the data. In addition, this link also needs to fill and calibrate the missing data to ensure the accuracy and completeness of the data. The core process of data cleaning and preprocessing includes data quality inspection, data cleaning, data standardization and feature processing. The data cleaning and preprocessing process is as follows: Figure 2 shown.
[0046] The data cleaning and preprocessing process begins with data input, followed by a data quality check to assess the integrity and validity of the data. If the data quality is unqualified, it will enter the data cleaning stage for deduplication and noise filtering, and will be checked repeatedly until the quality standard is met; when the data quality is qualified, it will enter the data standardization stage to unify the format and perform normalization, and finally complete the data classification, labeling and feature extraction in the feature processing stage. The final output result is a clean data set with high and consistent data quality after removing noise, redundancy and invalid data. At the same time, all data have been standardized in a unified format to support subsequent analysis and retrieval.
[0047] Knowledge base building module
[0048] The present invention adopts a graph-based knowledge base construction method, and the formulaic definition of the knowledge graph can be expressed as: G = (E, R, F) where E = {e 1 ,e 2 ,...,e n} is an entity set, R = {r 1 ,r 2 ,...,r m} is a set of relations, and F is a set of constraints and characteristic functions defined on entities and relations.
[0049] In specific implementation, the present invention divides entities into five categories:
[0050] Threat entities: including malicious IP, malicious domain names, malware and other specific threat vectors
[0051] Attack entity: describes the specific attack activities, including event records, attack techniques and tools used Vulnerability entity: records security vulnerability related information, including vulnerability number, impact scope and repair plan
[0052] Asset entity: describes the IT assets that need to be protected, including hardware devices, operating systems, and applications
[0053] Defense entity: includes various security protection measures, such as security recommendations, defense measures and response plans
[0054] Relationship types are defined as four core relationships:
[0055] Attack relationship: describes the relationship between attack entities and attack targets
[0056] Threat relationship: indicates the relationship between different threat indicators
[0057] Vulnerability relationship: describe the impact of the vulnerability and the solution to fix it
[0058] Defense Relationship: Connecting the Correspondence between Defense Measures and Threats
[0059] Entity Recognition and Relation Extraction
[0060] The present invention performs entity recognition and relationship extraction through a method based on a large language model:
[0061] (1) Entity recognition uses Figure 3 The prompt word design structure shown in the figure includes modules such as task description, entity category definition, output format, etc. The prompt word design adopts the principles of hierarchical type definition, open category division, and semantic boundary constraint.
[0062] (2) Relation extraction is performed using Figure 4 The prompt word structure shown in the figure contains four core components: task definition, input specification, relationship classification and output mode. This method combines explicit relationship mapping and implicit relationship discovery.
[0063] Entity merging and relationship fusion
[0064] The present invention adopts a fusion framework based on a large language model to perform entity merging and relationship fusion:
[0065] (1) Entity merging is performed through a similarity calculation function: Sim(e 1 ,e 2 )=LLM(Φ(e 1 ,e 2 )), where Φ(e 1 ,e 2 ) is a specially designed prompt template function used to analyze entity similarity from multiple dimensions.
[0066] like Figure 5 As shown in the figure, the prompt template for entity similarity calculation adopts a hierarchical structure design. The prompt template first gives the basic information of the two entities to be compared, including entity values, categories, and attributes. Then the similarity analysis is performed from four dimensions: the literal similarity of the entity names, the matching degree of the categories, the semantic similarity of the attribute values, and the expertise in the field of network security. The system requires a similarity score between 0 and 1 to be returned without any explanation.
[0067] like Figure 6 As shown in the figure, the entity merge prompt template designs a complete entity information fusion framework. The template first lists the complete information of the two entities to be merged, including entity value, category, attribute, confidence and timestamp. Then five specific merge rules are formulated: select more accurate or specific entity name, retain higher confidence category, merge complementary attribute information, retain the latest timestamp, and select higher confidence. Finally, the system is required to return the merge result in JSON format.
[0068] (2) The relational fusion process can be formalized as a three-stage mapping function: R' = I(C(R1 ∪R 2 ),ctx), where R 1 and R 2 is the set of relations to be fused, ctx is the context information, C is the conflict resolution function, and I is the relation reasoning function.
[0069] like Figure 7 As shown in the figure, the relationship conflict resolution prompt template adopts a systematic conflict handling mechanism. The template first clearly lists the source entity and target entity information of the conflict, as well as a list of all related relationships. On this basis, conflict resolution is carried out through four key dimensions: analyzing the confidence of the relationship, evaluating the degree of context support, considering the time order, and weighing the specificity of the relationship. The template requires the most appropriate relationship to be selected or merged based on these factors, and the results are returned in JSON format.
[0070] like Figure 8 As shown in the figure, the relational reasoning prompt template constructs a rigorous reasoning rule system. The template takes the existing relationship set and context information as input and sets four core reasoning requirements: there must be sufficient context support, conform to network security domain knowledge, be consistent with existing relationships, and provide confidence for each inferred relationship. Through this design, the reliability of the reasoning process and the accuracy of the reasoning results are ensured.
[0071] Knowledge retrieval module
[0072] The knowledge retrieval module is one of the core components of the present invention, which mainly includes two key parts: query understanding and conversion, and path retrieval. Fig. 9 As shown in the figure, query understanding and transformation adopts a modular design concept, which consists of six core components: system role definition, task description, input format, output format, and specific examples. This design ensures the accuracy and consistency of query understanding by clearly defining role positioning and task boundaries.
[0073] The knowledge retrieval module of the present invention includes two key parts: query understanding and conversion, and path retrieval. In the query understanding and conversion link, the present invention designs a special prompt word structure system, which includes system role definition, task description, input format, output format and example components. Through this system, the system can accurately convert the user's natural language query into a structured query representation, laying the foundation for subsequent knowledge retrieval.
[0074] In terms of path retrieval strategy, the present invention adopts a path-based retrieval method, which is mainly aimed at searching for single-hop and multi-hop relationships. The retrieval process involves three key parameters: knowledge graph path depth, number of retrieval results, and similarity threshold. For each target entity e, its path retrieval process can be expressed as: P(e, d) = {p|p = (e, r 1 ,e1 ,r 2 ,...,r x ,e j )∈G,len(p)≤d}, where G represents the knowledge graph, r represents the relationship type, and len(p) represents the path length. The system filters the search path through the designed type matching algorithm to ensure the relevance of the search results.
[0075] Answer generation module
[0076] The answer generation module transforms the retrieved discrete knowledge into coherent and accurate answers through a knowledge-enhanced generation framework. Fig.10 As shown in the figure, this module adopts a multi-stage generation scheme based on prompt words, which includes core components such as system role definition, task description, input format and output format. The prompt word design ensures the professionalism and standardization of the generated content.
[0077] The answer generation framework of the present invention transforms the retrieved discrete knowledge into a coherent and accurate answer through context organization and multi-round reasoning. Specifically, the context organization process can be formalized as:
[0078] C=Φ(P)={c 1 ,c 2 ,...,c m}, where P is the retrieved path set and Φ is the context transformation function. The context priority sorting uses a comprehensive scoring function: Priority(c i )=ω 1 ·Relevance(c i )+ω 2 ·Recency(c i )+ω 3 ·Authority(c i )
[0079] To ensure the quality of generated answers, the present invention designs a multi-dimensional quality assessment and optimization mechanism, including:
[0080] (1) Consistency check: Calculate the consistency between the generated answer and the contextual knowledge, and evaluate the answer quality through the semantic consistency measurement function.
[0081] (2) Completeness assessment: Evaluate whether the answer completely covers all aspects of the query requirements and quantify it by calculating the answer coverage.
[0082] (3) Enhanced explainability: The explainability of answers is improved through hierarchical processing, including extracting supporting evidence from context, building reasoning chains, and generating natural language explanations.
[0083] The answer generation module of the present invention also designs an automated process for quality optimization. The process first extracts supporting evidence from the context, then generates a standardized reference description, then builds a reasoning chain from evidence to conclusion, and finally generates a natural language explanation based on the reasoning chain. This hierarchical design not only ensures the reliability of the answer, but also enhances the comprehensibility of the system output.
[0084] Example:
[0085] In the part of presenting the results of the invention method, the experiment was evaluated using the CyberMetric dataset, which contains 10,000 professional questions in multiple fields such as penetration testing, cryptography, and network security.
[0086] (1) The present invention has achieved the best results in three core indicators:
[0087] The domain accuracy (DA) reached 86.7%, which is 18.5 percentage points higher than directly using the large language model and 12.4 percentage points higher than the traditional RAG.
[0088] The retrieval knowledge coverage rate (KCR) reached 81.2%, which was 11.7 and 5.9 percentage points higher than that of traditional RAG and GraphRAG respectively;
[0089] The explanation quality score (EQS) received a high score of 4.35, indicating that the system is able to provide professional and logically clear explanations.
[0090] (2) Typical case analysis
[0091] This paper selects two typical cases for in-depth analysis:
[0092] Network attack detection example: Fig.11 As shown in the figure, the system can accurately identify the key features of DDoS attacks and provide professional and logically clear explanations. This case achieved high scores in the three dimensions of professional domain accuracy (DA), retrieval knowledge coverage (KCR) and explanation quality score (EQS), which were 1.0, 0.92 and 4.5 respectively.
[0093] Information security governance case: Fig.12 As shown in the figure, the system demonstrated excellent professional analysis capabilities in data classification and grading management, accurately pointing out that data value is the primary consideration. This case also received high expert ratings, reaching 1.0, 0.90 and 4.5 points in the three evaluation dimensions respectively.
[0094] Through the above experimental evaluation and case analysis, the superior performance of the present invention in professional knowledge question answering tasks is fully demonstrated, especially in terms of knowledge retrieval quality, answer generation accuracy and professional interpretation ability. These results verify the effectiveness and practical value of the retrieval enhancement generation method based on knowledge graph.
Claims
1. A large language model retrieval enhancement generation method for cyberspace security emergency intelligent analysis, characterized by , including the following steps: Step 1: Knowledge base construction Establish a basic data set through multi-source data collection and preprocessing. The collection sources include threat intelligence, attack logs, research reports, etc., and use API collection, web crawlers, log aggregation and other mechanisms to achieve automated collection; clean and preprocess the collected data, including data quality inspection, data cleaning, data standardization and feature processing, to ensure data quality and consistency; Step 2: Knowledge graph construction First, we define the entity type system and divide entities into five core concepts: threat entity, attack entity, vulnerability entity, asset entity, and defense entity. At the same time, we define four core relationship types: attack relationship, threat relationship, vulnerability relationship, and defense relationship. We improve the quality of knowledge graph construction through entity recognition and relationship extraction methods based on large language models. Finally, we achieve effective knowledge integration through entity merging and relationship fusion mechanisms. Step 3: Knowledge Retrieval Design a graph-based hierarchical retrieval strategy. First, the user's natural language query is converted into a structured query representation through the query understanding and conversion module. Then, the path retrieval algorithm is used to retrieve relevant knowledge paths in the knowledge graph, involving the optimization settings of key parameters such as path depth, number of retrieval results, and similarity threshold. Step 4: Answer Generation Develop a knowledge-enhanced generation framework that transforms the retrieved discrete knowledge into coherent and accurate answers through contextual organization and multi-round reasoning; at the same time, design a multi-dimensional quality assessment and optimization mechanism, including three key components: consistency verification, completeness assessment, and explainability enhancement, to ensure the professionalism and reliability of the generated answers.
2. The method according to claim 1, characterized in that ,The entity merging in step 2 is performed through a similarity calculation function: Sim(e1,e2)=LLM(Φ(e1,e2)), where Φ(e1,e2) is a specially designed prompt template function.
3. The method according to claim 1, characterized in that ,The relation fusion process in the step 2 can be formalized as a three-stage mapping function: R'=I(C(R1∪R2),ctx), where R1 and R2 are the relation sets to be fused, ctx is the context information, C is the conflict resolution function, and I is the relation reasoning function.
4. The method according to claim 1, characterized in that , the answer generation framework in step 4 can be formalized as follows through the context organization process: C = Φ(P) = {c1, c2, ..., c m }, where P is the retrieved path set, Φ is the context transformation function; the context priority sorting uses a comprehensive scoring function: Priority(c i )=ω1·Relevance(c i )+ω2·Recency(c i )+ω3·Authority(c i ).
Citation Information
Cited By
Digital emergency plan management method based on knowledge graph
CN120911994A
Large model intelligent analysis and suggestion system and method for urban management field
CN120930773A
Power grid fault handling plan knowledge retrieval method and related device
CN121071197A
Key information infrastructure association evaluation system and method based on large language model
CN121283762A
Safety and robustness automatic testing method for large language model in medical field
CN121327624A