System operation security dialogue method generated by large model fine tuning and two-stage retrieval

Through the method of fine-tuning of large-models and two-stage search generation, the search accuracy and response speed problems in the field of complex system operation security are solved, intelligent and efficient knowledge management is realized, multi-condition combination query and dynamic semantic understanding are supported, and structured answers are provided.

CN120471172AActive Publication Date: 2025-08-12CHONGQING UNIV +1

Patent Information

Application Number
CN202510608578.3
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-05-13
Publication Date
2025-08-12
Estimated Expiration
2045-05-13

AI Technical Summary

Technical Problem

The search accuracy of existing complex systems in the field of operational security is limited, it cannot handle complex queries, lacks intelligent decision-making support, has poor user experience, slow response speed, and cannot meet real-time requirements.

Method used

The method of large-model fine-tuning and two-stage retrieval generation is adopted, including data acquisition and preprocessing, question-and-answer data collection construction, base model fine-tuning, knowledge base construction, vector recall and answer generation. Through the fine-tuning of GPT-4 and InternLM2.5-Chat-7B models, semantic similarity calculation and fine-tuning are carried out in combination with Dual-Encoder and Cross-Encoder architectures, a knowledge base covering the operation safety knowledge of complex systems is constructed.

Benefits of technology

It realizes intelligence, efficiency and precision in the field of operation safety of complex systems, improves retrieval accuracy and response speed, supports multi-condition combination query and dynamic semantic understanding, provides structured answers, and meets emergency decision-making and training needs.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120471172A_ABST
    Figure CN120471172A_ABST
Patent Text Reader

Abstract

The invention discloses a system operation security dialogue method generated by large model fine tuning and two-stage retrieval, and relates to the technical field of complex system operation security. Constructing a question and answer pair data set; finely adjusting the base model; constructing a knowledge base; carrying out vector recall; selecting and rearranging; and generating answers and the like. According to the system operation security dialogue method generated through large model fine tuning and two-stage retrieval, an automatic question answering and rapid retrieval mechanism based on a knowledge model in complex system operation security is achieved, and intellectualization, high efficiency and precision in the field of complex system operation security are promoted.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of complex system operation security technology, and in particular to a system operation security dialogue method based on large model fine-tuning and two-stage retrieval generation. Background Art

[0002] In today's society, complex systems are widely used in infrastructure sectors such as transportation, energy, and communications. These infrastructures are the cornerstone of the normal operation of modern society. With the continuous advancement of science and technology, the scale and complexity of various systems are increasing. The development status of existing technologies is as follows: 1. Digitization of knowledge resources: Knowledge resources in the field of complex system operational safety, such as laws and regulations, operating procedures, and technical documentation, have been gradually digitized into electronic documents and databases, facilitating storage and initial retrieval. 2. Simple search functions: Some systems provide basic search functions that enable keyword searches within digitized knowledge resources, helping users quickly locate documents or fragments containing specific keywords. 3. Preliminary classification and organization: Preliminary classification of complex system operational safety knowledge, such as by professional field or business process, is performed to enable users to more specifically find the information they need. 4. Application of expert systems: Expert systems are often used in complex scenarios, such as incident handling and emergency decision-making. These systems incorporate the experience and knowledge of domain experts, providing guidance and advice to novice personnel.

[0003] At the same time, there are also some defects and shortcomings:

[0004] 1. Limited retrieval accuracy: (1) Limitations of keyword matching: Retrieval methods that rely on keyword matching are difficult to understand the user's true intentions and the semantic content of the document, resulting in a large amount of irrelevant or low-quality information in the retrieval results, and users need to spend extra time and energy to screen; (2) Unable to handle complex queries: For query requirements involving multiple conditions and complex logical relationships, simple retrieval functions often cannot accurately understand and respond, and cannot provide accurate answers.

[0005] 2. Lack of intelligent decision-making support: (1) Insufficient information integration: When faced with complex security issues, it is impossible to automatically integrate multi-source and multi-type knowledge resources to provide decision makers with comprehensive and in-depth analysis and suggestions; (2) Inability to predict and reason: There is a lack of predictive analysis capabilities for potential risks and trends, and it is impossible to reason and simulate based on existing knowledge to support preventive and forward-looking decision-making.

[0006] 3. Poor user experience: (1) Single interaction method: The human-computer interaction method is relatively primitive, mainly relying on text input and simple commands. It lacks natural language processing capabilities and cannot understand users' natural language questions, affecting query efficiency; (2) Slow response speed: When processing large-scale knowledge bases or complex queries, the system's response speed is slow, unable to meet real-time requirements, and delaying decision-making. Summary of the Invention

[0007] The purpose of the present invention is to provide a system operation security dialogue method generated by large model fine-tuning and two-stage retrieval. It adopts a two-stage retrieval architecture based on large model enhancement and lightweight fine-tuning technology, breaking through the difficulties of traditional complex system operation security knowledge management systems in retrieval accuracy, response speed and model deployment cost.

[0008] To achieve the above objectives, the present invention provides a system operation security dialogue method for large model fine-tuning and two-stage retrieval generation, comprising the following steps:

[0009] S1. Data collection and preprocessing: Collect data in the field of complex system operation security and preprocess the data;

[0010] S2. Question-answer pair data set construction: Based on the pre-processed data in S1, the question-answer pair data set is constructed using the GPT-4 large model;

[0011] S3, base model fine-tuning: The InternLM2.5-Chat-7B model is selected as the base model. According to the question-answer data set in S2, the XTuner framework combined with the QLoRA method is used to perform lightweight fine-tuning on the base model.

[0012] S4. Knowledge base construction: Based on the llamaindex framework and using the bce-embedding-base_v1 embedding model, a complex system operation security knowledge base is constructed;

[0013] S5, Vector Recall: Based on the Dual-Encoder architecture embedding model and the HNSW algorithm, the first stage of rapid retrieval of the knowledge base built in S4 is performed to recall knowledge blocks;

[0014] S6, Selected Reranking: Based on the knowledge blocks recalled by the vectors in S5, the bce-reranker-base_v1 semantic refined ranking model is used for the second stage of precise screening;

[0015] S7, answer generation: Combine the fine-tuned base model in S3 with the retrieval results in S6 to generate the answer.

[0016] Preferably, the specific steps of S1 are as follows:

[0017] S11. Collect data on complex system operational safety: This data includes professional literature, laws, regulations, and standards, academic resources, and industry knowledge bases. Professional literature includes books and authoritative works. Laws, regulations, and standards include relevant national and industry laws, regulations, operating procedures, and technical standards. Academic resources include papers and patents. Industry knowledge bases include practitioner examination question banks and typical accident case databases.

[0018] S12. Preprocess the data: Preprocessing includes data cleaning, data deduplication and text formatting. Data cleaning includes removing redundancy and format standardization.

[0019] Preferably, the specific steps of S2 are as follows:

[0020] S21, prompt word design: Use templates to guide GPT-4 to extract questions from the preprocessed data in S1, with one answer data corresponding to five questions;

[0021] S22. Semantic consistency verification: By calculating the cosine similarity between the question and the answer, screen the question-answer pairs with similarity greater than 0.8;

[0022] S23. Format conversion: convert the question and answer data into JSON format.

[0023] Preferably, the specific formula of cosine similarity in S22 is as follows:

[0024]

[0025] Among them, Q and A are the word vector representations of questions and answers respectively.

[0026] Preferably, the specific steps of S3 are as follows:

[0027] S31, efficient parameter fine-tuning: freeze the backbone parameters of the base model and only adjust the low-rank matrix of the adapter module. QLoRA quantizes the weight matrix.

[0028] S32. Prompt word design: Design prompt word templates based on the requirements of complex system operation security scenarios;

[0029] S33. Loss function design: Use the cross-entropy loss function to optimize the base model's ability to generate complex system security domain knowledge;

[0030] S34, Training strategy: Use the AdamW optimizer, set the learning rate to 2e-4, the batch size to 16, and perform early stopping on the domain validation set. During training, the domain validation set is used to evaluate the base model. If the validation set performance does not improve for multiple epochs, stop training early to prevent overfitting of the base model.

[0031] S35. Model conversion: Use the xtuner convert pth_to_hf command to convert the model weight file originally trained using Pytorch into the currently common HuggingFace format file;

[0032] S36, Model Merge: Based on the additional layers fine-tuned by QLoRA, use the xtuner convert merge command to merge the trained layers with the original base model.

[0033] Preferably, the cross entropy loss function in S33 is as follows:

[0034]

[0035] Where T is the sequence length, θ is the adapter parameter, P is the conditional probability distribution, y1,…,y t-1 Output elements for the history in the sequence.

[0036] Preferably, the specific steps of S4 are as follows:

[0037] S41, knowledge chunking: split long text into chunks of 200 words based on punctuation marks;

[0038] S42. Vector generation: Use the bce-embedding-base_v1 model to encode each knowledge block into a 768-dimensional semantic vector p i ;

[0039] S43. Storage and Indexing: The vector storage of the llamaindex framework uses the Chroma database to store text vectors in the Chroma vector database, and performs similarity search in large-scale data sets through the core HNSW algorithm of the Chroma vector database.

[0040] Preferably, the specific steps of S5 are as follows:

[0041] S51. Query vectorization: Use the bce-embedding-base_v1 model to convert user questions into query vectors q;

[0042] S52, similarity calculation: In the vector database, the cosine similarity formula in S22 is used to calculate the similarity between q and all semantic vectors p in the knowledge base. i The top-N knowledge blocks are recalled according to the similarity, where N=10.

[0043] Preferably, the specific steps of S6 are as follows:

[0044] S61, Cross Encoding: Concatenate the query and the passage and input them into the bce-reranker-base_v1 model to capture the semantic interaction features of the query and the passage;

[0045] S62, Semantic Score: Based on the semantic interaction features of Query and Passage, the semantic relevance score s of Query and Passage is calculated through the Transformer model i , and call the sigmoid function in torch to normalize the scores and then sort them;

[0046] S63. Result screening: After re-arranging according to the scores, the relevance score s is screened out. i For fragments with a probability of >0.6, we filter out low-quality fragments and select the top-M highly relevant knowledge blocks as the basis for answer generation, where M=5.

[0047] Preferably, the specific steps of S7 are as follows:

[0048] S71, context fusion: concatenate the selected M knowledge blocks with the user query to form a model input sequence;

[0049] S72. Hint Engineering: Design hint templates to guide the fine-tuned base model to generate structured answers.

[0050] Therefore, the present invention adopts the above-mentioned large model fine-tuning and two-stage retrieval to generate a system operation security dialogue method, which has the following beneficial effects:

[0051] (1) Deep integration of domain knowledge: By integrating multi-source heterogeneous data such as professional literature, regulations and standards, and accident cases, a knowledge system covering different scenarios of complex system operation safety is constructed to solve the data fragmentation problem of traditional methods. Using GPT-4 to generate question-answer pairs and combining them with semantic verification, we break through the bottleneck of high cost and narrow coverage of manual labeling and significantly improve the diversity and accuracy of training data. The question-answer pairs generated on this basis are improved from one answer data corresponding to one question to one answer data corresponding to five questions, expanding the semantic coverage of the training data, enabling the model to learn multi-dimensional language expressions of the same knowledge, and improving the generalization ability of complex queries.

[0052] (2) Lightweight model fine-tuning technology: QLoRA quantization technology based on the XTuner framework enables efficient domain adaptation of large models under limited computing power, avoiding the hardware resource consumption of full parameter fine-tuning. By freezing the main parameters and only fine-tuning the adapter layer, the general capabilities of the base model are retained while giving it the ability to accurately generate professional terminology and safety regulations in the field of complex system operation safety. Prompt words are designed according to actual application needs to guide the model to learn specific expression patterns and knowledge structures.

[0053] (3) Combining rapid coarse screening with precise fine ranking: In the first stage (vector recall), a dual encoder model is used to independently encode user questions and the knowledge base into semantic vectors. Vector similarity retrieval is performed on the offline vector library to quickly recall knowledge fragments with similar semantics. In the second stage (semantic fine ranking), a cross-encoder model is used to conduct in-depth semantic interaction analysis on candidate fragments, effectively identifying the implicit association between user intent and knowledge fragments, filtering out low-quality content, and improving the accuracy of relevant content retrieval. The two-stage division of labor balances speed and accuracy, avoiding the drawbacks of "missed detection" or "false detection" in a single retrieval mode.

[0054] (4) Dynamic semantic understanding capability: Automatically associate technical terms (such as "traction power supply system failure") with fuzzy expressions (such as "sudden power outage of the train") to improve the robustness of intent parsing in complex scenarios; support multi-condition combination queries (such as "track circuit fault handling process in rainy days"), and accurately match complex knowledge fragments including environmental factors, equipment types, and operating steps.

[0055] (5) Multi-scenario adaptive capabilities: Emergency command: quickly generate a structured plan that includes handling steps, responsible departments, and technical basis to assist on-site decision-making; Personnel training: improve training efficiency through interactive dialogue simulation of troubleshooting, procedure learning, and other scenarios; Hidden danger warning: match potential risk patterns based on the historical case library to achieve proactive safety warnings.

[0056] The technical solution of the present invention is further described in detail below through the accompanying drawings and embodiments. BRIEF DESCRIPTION OF THE DRAWINGS

[0057] Figure 1 A framework diagram of a complex system operation security real-time dialogue method embodiment of the system operation security dialogue method for large model fine-tuning and two-stage retrieval generation of the present invention;

[0058] Figure 2 A two-stage retrieval generation framework diagram for an embodiment of the system operation security dialogue method for large model fine-tuning and two-stage retrieval generation of the present invention. DETAILED DESCRIPTION

[0059] The technical solution of the present invention is further described below with reference to the accompanying drawings and embodiments.

[0060] Unless otherwise defined, the technical or scientific terms used in the present invention shall have the usual meanings understood by persons of ordinary skill in the field to which the present invention belongs. The words "first", "second" and similar terms used in the present invention do not indicate any order, quantity or importance, but are only used to distinguish different components. Words such as "include" or "comprise" mean that the elements or objects preceding the word include the elements or objects listed after the word and their equivalents, without excluding other elements or objects. Words such as "connect" or "connected" are not limited to physical or mechanical connections, but may include electrical connections, whether direct or indirect. "Up", "down", "left", "right" and the like are only used to indicate relative positional relationships. When the absolute position of the object being described changes, the relative positional relationship may also change accordingly.

[0061] Example

[0062] See also Figure 1-2 The present invention provides a system operation safety dialogue method based on large model fine-tuning and two-stage retrieval generation. The following is a detailed solution for the high-speed rail operation safety scenario in complex system operation safety to better illustrate the dialogue method proposed by the present invention.

[0063] The present invention provides a system operation security dialogue method for large model fine-tuning and two-stage retrieval generation, including the following steps:

[0064] S1. Data collection and preprocessing: Collect data in the field of complex system operation security and preprocess the data. The specific steps are as follows:

[0065] S11. The present invention collects real high-speed rail operation safety data through multiple channels:

[0066] (1) Professional Literature: 52 books published in the past 20 years on high-speed rail transport organization, operational safety assurance, traction power supply system and other fields, including authoritative works such as "High-speed Railway Operation Safety Management" and "Technical Regulations for High-speed Railway Traction Power Supply System";

[0067] (2) Regulations and Standards: 101 laws, regulations, operating procedures, and technical standards related to high-speed rail operation safety issued by the state and the industry, including the "High-Speed Railway Safety Management Regulations" and the "Railway Technical Management Regulations";

[0068] (3) Academic resources: 500 papers and patents on high-speed rail track circuits, turnout machines, safety accident analysis, etc. from platforms such as IEEE and CNKI;

[0069] (4) Industry knowledge base: including the high-speed rail practitioners’ entry examination question bank and typical accident case database;

[0070] S12. Preprocess the data: Preprocessing includes data cleaning (removing redundancy and format standardization), data deduplication and text formatting.

[0071] S2. Question-Answer Pair Dataset Construction: Based on the pre-processed data in S1, the GPT-4 large model is used to construct question-answer pairs for fine-tuning the high-speed rail operation safety model. The specific steps are as follows:

[0072] S21. Prompt Design: Using the template "Please carefully design five questions based on the {input_value} I will provide below. The answer to each of these five questions must be the complete content I provide, and these questions must accurately correspond to the content. Ensure that the answer is complete, without partial answers or requiring additional information. Ensure that each question directly outputs the original content I provide. The format of the five questions output is: que1: que2: que3: que4: que5 ...

[0073] S22. Semantic consistency verification: By calculating the cosine similarity between the question and the answer, screen the question-answer pairs with similarity greater than 0.8. The specific formula for cosine similarity is:

[0074]

[0075] Among them, Q (question) and A (answer) are the word vector representations of the question and answer respectively;

[0076] S23. Format conversion: Convert the question and answer data into JSON format. The specific format is as follows:

[0077]

[0078]

[0079] Unlike existing technologies that rely on a single data source and a single round of question-and-answer pairs, this paper utilizes GPT-4 to generate an expanded "one data, five questions" question-and-answer strategy, combined with a cosine similarity verification mechanism, to enhance the semantic coverage of the training data. This multi-dimensional semantic representation approach enables the model to learn different language representations of the same knowledge, enhancing its generalization capabilities for complex queries and addressing the problem of redundant search results caused by insufficient semantic understanding in traditional systems.

[0080] S3. Fine-tuning the base model: The InternLM2.5-Chat-7B model is selected as the base model. Based on the question-answer data set in S2, the XTuner framework is used in combination with the QLoRA method to perform lightweight fine-tuning on the base model. The specific steps are as follows:

[0081] S31, efficient parameter fine-tuning: Freeze the backbone parameters of the base model and only adjust the low-rank matrix of the adapter module to reduce the amount of calculation. QLoRA quantizes the weight matrix (4-bit quantization) to reduce memory usage and improve inference efficiency.

[0082] S32. Prompt Word Design: Based on the requirements of high-speed rail operation safety scenarios, design a prompt word template: You are a professional assistant in the field of high-speed rail operation safety. Your task is to answer the user's question {input} based on your professional knowledge in the field of high-speed rail safety. The answer must be accurate, comprehensive, logically clear, and in line with industry standards. If the user's question involves multiple aspects, please answer them one by one;

[0083] S33. Loss function design: The cross entropy loss function is used to optimize the model's ability to generate knowledge in the field of high-speed rail safety. The cross entropy loss function is specifically:

[0084]

[0085] Where T is the sequence length, θ is the adapter parameter, P is the conditional probability distribution, y1,…,y t-1 Output elements for the history in the sequence;

[0086] S34, Training strategy: Use the AdamW optimizer, set the learning rate to 2e-4, the batch size to 16, and perform early stopping on the in-domain validation set. That is, during training, the model is evaluated on the in-domain validation set. If the validation set performance does not improve for several consecutive epochs, stop training early to prevent overfitting of the model.

[0087] S35. Model conversion: Use the xtuner convert pth_to_hf command to convert the model weight file originally trained using Pytorch into the currently common HuggingFace format file;

[0088] S36. Model merging: Based on the additional layer (Adapter) fine-tuned by QLoRA, use the xtuner convertmerge command to merge the trained layer (Adapter) with the original model.

[0089] The present invention uses the XTuner framework in combination with the QLoRA method to reduce the model memory usage while maintaining the excellent performance of the model by freezing the base model backbone parameters (only adjusting <1% of the adapter parameters) and 4-bit quantization technology. During the fine-tuning process, the present invention designs prompt words according to actual application requirements to guide the model to learn specific expression patterns and knowledge structures. This technology enables the model to be quickly deployed on ordinary hardware, greatly improving the scalability and real-time response capabilities of the system, meeting the timeliness requirements of complex system operation safety emergency scenarios, and further enhancing the model's ability to understand complex queries through the optimized design of prompt words, thereby improving the accuracy and professionalism of the answers.

[0090] S4. Knowledge base construction: Based on the llamaindex framework and using the bce-embedding-base_v1 embedding model, a high-speed rail operation safety knowledge vector library is constructed. The specific steps are as follows:

[0091] S41, knowledge chunking: split long text into chunks of 200 words based on punctuation marks, ensuring that each knowledge chunk contains complete technical points, balancing semantic integrity and retrieval efficiency;

[0092] S42. Vector generation: Use the bce-embedding-base_v1 model (dual-encoder architecture) to encode each knowledge block into a 768-dimensional semantic vector p i ;

[0093] S43. Storage and Indexing: llamaindex's vector storage uses the Chroma database to store text vectors. The Chroma database supports 768-dimensional vector storage. With the help of the core HNSW (Hierarchical Navigable Small World) algorithm of the Chroma vector database (using a pre-built navigation graph to implement approximate nearest neighbor search (ANN)), fast and efficient similarity search is achieved in large-scale data sets.

[0094] S5, Vector Recall: Based on the Dual-Encoder architecture embedding model and the HNSW algorithm, the first stage of rapid retrieval of the knowledge base built in S4 is performed to recall knowledge blocks. The specific steps are as follows:

[0095] S51. Query vectorization: Use the bce-embedding-base_v1 model to convert the user question into a query vector q. The bce-embedding-base_v1 model uses a dual encoder to independently encode the query and passage into 768-dimensional vectors.

[0096] S52, similarity calculation: In the vector database, the cosine similarity formula in S22 is used to calculate the similarity between q and all semantic vectors p in the knowledge base. i The top-N (N=10) knowledge blocks are recalled according to the similarity.

[0097] S6, Selected Reranking: Based on the knowledge blocks recalled by the vectors in S5, the bce-reranker-base_v1 semantic refined ranking model (based on the Cross-Encoder architecture) is used to perform the second stage of precise screening. The specific steps are as follows:

[0098] S61, cross coding: concatenate Query and Passage into "Query <sep>Passage" format and input into the bce-reranker-base_v1 model (based on the Transformer architecture), using multi-head attention to capture the semantic interaction features between Query and Passage;

[0099] S62, Semantic Score: Based on the semantic interaction features of Query and Passage, the Transformer model calculates the semantic relevance score s between Query and Passage i , and call the sigmoid function in torch to normalize the scores and then sort them;

[0100] S63. Result screening: After re-arranging according to the scores, the relevance score s is screened out. i >0.6, filter out low-quality fragments, and select Top-M (M=5) highly relevant knowledge blocks as the basis for answer generation, which improves the accuracy compared to single-stage retrieval.

[0101] The present invention uses a two-stage retrieval generation technology, adopting the embedding model bce-embedding-base_v1 and the semantic reranking model bce-reranker-base_v1. The two-stage retrieval generation architecture includes two major stages. First, through offline preprocessing, the knowledge block is processed by the embedding model (Embedding), and a semantic vector is generated and stored in the vector database, completing the vectorization construction of the knowledge base and providing basic support for subsequent retrieval. In the online user interaction stage, after the user asks a question, the embedding model first generates a query vector from the question; through retrieval recall, the Embedding model based on the Dual-Encoder (dual encoder structure) architecture adopts a dual encoder structure, and the query (user question) and passage (knowledge base text fragment) are respectively input into independent encoders to generate their own semantic vectors. There is no information interaction between the two during the encoding process. Vector similarity retrieval is performed in the vector database to achieve rapid recall of knowledge fragments with similar semantics. Subsequently, the score re-ranking module activates the Reranker model of the Cross-Encoder architecture. It uses a cross-encoder structure to splice the Query and Passage and input them into the model, so that the two can fully exchange information during the encoding process. The Transformer model is used to extract semantic relationships from the recalled fragments, calculate the semantic relevance between the Query and Passage, and put highly relevant fragments in front and filter out low-quality content.

[0102] Different from the low accuracy of traditional single-stage retrieval, this invention adopts a two-stage architecture combining Dual-Encoder and Cross-Encoder: in the first stage, Dual-Encoder performs vector similarity retrieval on the offline vector library to quickly recall semantically similar knowledge fragments, and in the second stage, Cross-Encoder is used to capture semantic interaction features for precise rearrangement. The hierarchical retrieval strategy greatly improves the retrieval accuracy and effectively filters out low-quality content through threshold screening (relevance score > 0.6). In addition, the synergy between the dual encoder and the cross encoder enables the system to handle complex semantic associations, such as accurately identifying the potential connection between "turnout machine anomaly" and "track circuit failure", which is impossible with traditional keyword matching technology.

[0103] S7, answer generation: Combine the fine-tuned InternLM2.5-Chat-7B model in S3 with the search results in S6 to generate the answer. The specific steps are as follows:

[0104] S71, context fusion: combine the selected M knowledge blocks with the user query to form the model input sequence "Query <sep>Context1 <sep>Context2 <sep>...”;

[0105] By designing "Query <sep>Contexts" context fusion prompt template, building context-aware input, <sep>As a delimiter, the model is guided to distinguish between questions and background knowledge, integrating search results with user queries and generating structured answers that adhere to the norms of complex system operational safety. This prompt strategy, combined with domain qualifiers, ensures the professionalism and accuracy of answers, for example by automatically linking regulatory clauses with technical standards. Existing technologies often lack the ability to integrate domain knowledge of complex system operational safety.

[0106] S72. Prompt Engineering: Design a prompt template: "Based on professional knowledge in the field of high-speed rail operation safety, answer the following question: {Query}. Related knowledge is as follows: {Contexts}" to guide the fine-tuned InternLM2.5-Chat-7B model to generate structured answers.

[0107] The variables {Query} and {Contexts} are dynamically populated with content, combined with the domain qualifier "knowledge of complex system operational security domain". Ultimately, the fine-tuned InternLM2.5-Chat-7B model generates and outputs structured answers that meet industry standards based on the high-quality knowledge fragments that have been re-ranked after scoring.

[0108] This real-time dialogue method for complex system operation safety first collects and preprocesses multi-source data such as professional literature, regulations and standards in the field of complex system operation safety, generates question-answer pairs with the help of GPT-4, and constructs a dataset through cosine similarity verification; based on InternLM2.5-Chat-7B, it uses XTuner combined with the QLoRA method for fine-tuning (freezing backbone parameters, 4-bit quantization, etc.) and completes model conversion and merging; based on the Llamaindex framework, the preprocessed data is divided into blocks and semantic vectors are generated through the embedding model and stored in the Chroma database. After a two-stage retrieval, highly relevant text blocks are screened out, finally integrated into the query and constructed as prompt words, which are input into the fine-tuned model to finally generate structured answers.

[0109] Therefore, the present invention adopts the system operation safety dialogue method generated by the above-mentioned large model fine-tuning and two-stage retrieval, and adopts a two-stage retrieval architecture and lightweight fine-tuning technology based on large model enhancement, breaking through the difficulties of traditional complex system operation safety knowledge management systems in retrieval accuracy, response speed and model deployment cost, and realizing automated question-answering and rapid retrieval mechanism based on knowledge models in complex system operation safety, and is committed to promoting the intelligence, efficiency and precision in the field of complex system operation safety.

[0110] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention rather than to limit the same. Although the present invention has been described in detail with reference to the preferred embodiments, those skilled in the art should understand that they can still modify or replace the technical solutions of the present invention with equivalents, and these modifications or equivalent replacements cannot cause the modified technical solutions to deviate from the spirit and scope of the technical solutions of the present invention.< / sep> < / sep> < / sep> < / sep> < / sep> < / sep>

Claims

1. A system-wide secure dialogue method for large-scale model fine-tuning and two-stage retrieval generation, characterized by: The following steps are involved: S1. Data collection and preprocessing: Collect data in the field of complex system operation security and preprocess the data; S2. Question-answer pair data set construction: Based on the pre-processed data in S1, the question-answer pair data set is constructed using the GPT-4 large model; S3, base model fine-tuning: The InternLM2.5-Chat-7B model is selected as the base model. According to the question-answer data set in S2, the XTuner framework combined with the QLoRA method is used to perform lightweight fine-tuning on the base model. S4. Knowledge base construction: Based on the llamaindex framework and using the bce-embedding-base_v1 embedding model, a complex system operation security knowledge base is constructed; S5, Vector Recall: Based on the Dual-Encoder architecture embedding model and the HNSW algorithm, the first stage of rapid retrieval of the knowledge base built in S4 is performed to recall knowledge blocks; S6, Selected Reranking: Based on the knowledge blocks recalled by the vectors in S5, the bce-reranker-base_v1 semantic refined ranking model is used for the second stage of precise screening; S7, answer generation: Combine the fine-tuned base model in S3 with the retrieval results in S6 to generate the answer.

2. The system operation security dialogue method for large model fine-tuning and two-stage retrieval generation according to claim 1 is characterized in that: The specific steps of S1 are as follows: S11. Collect data on complex system operational safety: This data includes professional literature, laws, regulations, and standards, academic resources, and industry knowledge bases. Professional literature includes books and authoritative works. Laws, regulations, and standards include relevant national and industry laws, regulations, operating procedures, and technical standards. Academic resources include papers and patents. Industry knowledge bases include practitioner examination question banks and typical accident case databases. S12. Preprocess the data: Preprocessing includes data cleaning, data deduplication and text formatting. Data cleaning includes removing redundancy and format standardization.

3. The system operation security dialogue method for large model fine-tuning and two-stage retrieval generation according to claim 2 is characterized in that: The specific steps of S2 are as follows: S21, prompt word design: Use templates to guide GPT-4 to extract questions from the preprocessed data in S1, with one answer data corresponding to five questions; S22. Semantic consistency verification: By calculating the cosine similarity between the question and the answer, screen the question-answer pairs with similarity greater than 0.8; S23. Format conversion: convert the question and answer data into JSON format.

4. The system operation security dialogue method for large model fine-tuning and two-stage retrieval generation according to claim 3 is characterized in that: The specific formula of cosine similarity in S22 is as follows: Among them, Q and A are the word vector representations of questions and answers respectively.

5. The system operation security dialogue method for large model fine-tuning and two-stage retrieval generation according to claim 4 is characterized in that: The specific steps of S3 are as follows: S31, efficient parameter fine-tuning: freeze the backbone parameters of the base model and only adjust the low-rank matrix of the adapter module. QLoRA quantizes the weight matrix. S32. Prompt word design: Design prompt word templates based on the requirements of complex system operation security scenarios; S33. Loss function design: Use the cross-entropy loss function to optimize the base model's ability to generate complex system security domain knowledge; S34, Training strategy: Use the AdamW optimizer, set the learning rate to 2e-4, the batch size to 16, and perform early stopping on the domain validation set. During training, the domain validation set is used to evaluate the base model. If the validation set performance does not improve for multiple epochs, stop training early to prevent overfitting of the base model. S35. Model conversion: Use the xtuner convert pth_to_hf command to convert the model weight file originally trained using Pytorch into the currently common HuggingFace format file; S36, Model Merge: Based on the additional layers fine-tuned by QLoRA, use the xtuner convert merge command to merge the trained layers with the original base model.

6. The system operation security dialogue method for large model fine-tuning and two-stage retrieval generation according to claim 5 is characterized in that: The cross entropy loss function in S33 is as follows: Where T is the sequence length, θ is the adapter parameter, P is the conditional probability distribution, y1,…,y t-1 Output elements for the history in the sequence.

7. The system operation security dialogue method for large model fine-tuning and two-stage retrieval generation according to claim 6 is characterized in that: The specific steps of S4 are as follows: S41, knowledge chunking: split long text into chunks of 200 words based on punctuation marks; S42. Vector generation: Use the bce-embedding-base_v1 model to encode each knowledge block into a 768-dimensional semantic vector p i ; S43. Storage and Indexing: The vector storage of the llamaindex framework uses the Chroma database to store text vectors in the Chroma vector database, and performs similarity search in large-scale data sets through the core HNSW algorithm of the Chroma vector database.

8. The system operation security dialogue method for large model fine-tuning and two-stage retrieval generation according to claim 7 is characterized in that: The specific steps of S5 are as follows: S51. Query vectorization: Use the bce-embedding-base_v1 model to convert user questions into query vectors q; S52, similarity calculation: In the vector database, the cosine similarity formula in S22 is used to calculate the similarity between q and all semantic vectors p in the knowledge base. i The top-N knowledge blocks are recalled according to the similarity, where N=10.

9. The system operation security dialogue method for large model fine-tuning and two-stage retrieval generation according to claim 8 is characterized in that: The specific steps of S6 are as follows: S61, Cross Encoding: Concatenate the query and the passage and input them into the bce-reranker-base_v1 model to capture the semantic interaction features of the query and the passage; S62, Semantic Score: Based on the semantic interaction features of Query and Passage, the semantic relevance score s of Query and Passage is calculated through the Transformer model i , and call the sigmoid function in torch to normalize the scores and then sort them; S63. Result screening: After re-arranging according to the scores, the relevance score s is screened out. i For fragments with a probability of >0.6, we filter out low-quality fragments and select the top-M highly relevant knowledge blocks as the basis for answer generation, where M=5.

10. The system operation security dialogue method for large model fine-tuning and two-stage retrieval generation according to claim 9 is characterized in that: The specific steps of S7 are as follows: S71, context fusion: concatenate the selected M knowledge blocks with the user query to form a model input sequence; S72. Prompt Engineering: Design prompt templates to guide the fine-tuned base model to generate structured answers.

Citation Information

Patent Citations

  • Knowledge question-answering system based on large language model

    CN119396975A

  • Operation and maintenance question-answering system and method based on large language model workflow

    CN119441410A

  • Power document intelligent question and answer method and system based on large language model

    CN119577082A

  • Model knowledge base-based question and answer method and device, storage medium and electronic equipment

    CN119719276A

Cited By

  • Multi-round reasoning question answering method and system based on query graph driving

    CN120687579A