Device fault intelligent question-answering method based on large model enhancement

By applying a large-model-enhanced intelligent question-and-answer method to equipment fault diagnosis, combining the fault knowledge graph and multi-agent collaborative inference model, the problem that traditional methods are difficult to deal with complex faults is solved, and efficient and accurate fault diagnosis and analysis are achieved.

CN120086341APending Publication Date: 2025-06-03TONGJI UNIV
View PDF 0 Cites 7 Cited by

Patent Information

Application Number
CN202510176859.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-02-18
Publication Date
2025-06-03

AI Technical Summary

Technical Problem

Traditional equipment fault diagnosis methods are difficult to deal with complex cross-device and cross-system faults, and lack flexibility and real-time adaptability, making it difficult to deal with changing production environments and complex fault scenarios.

Method used

Using the intelligent question-and-answer method of equipment failure based on large-model enhancement, by constructing a fault knowledge graph and a multi-agent collaborative inference model, different inference mechanisms are called for single-hop, short-hop and multi-hop problems to generate equipment failure analysis in natural language form.

Benefits of technology

It improves the efficiency and accuracy of fault diagnosis, provides intelligent Q&A with high reliability and strong readability, and can effectively solve equipment failure problems in complex industrial scenarios.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120086341A_ABST
    Figure CN120086341A_ABST
Patent Text Reader

Abstract

The invention relates to an equipment fault intelligent question-answering method based on large model enhancement, and the method comprises the following steps: obtaining steel production line data, and constructing a fault knowledge graph; user questions are obtained and classified through semantic analysis, and classification results comprise single-hop questions, short-hop questions and multi-hop questions; calling a corresponding reasoning mechanism according to the category of the user question, generating an answer according to the fault knowledge graph, and outputting equipment fault analysis in a natural language form through a retrieval enhancement generation technology; for a single-hop problem, core entities and relationships are matched based on an Aho-Corasick automaton, and answers are directly retrieved from the fault knowledge graph; for a short jump problem, deriving an answer from the fault knowledge graph by adopting a single-agent inference model; for a multi-hop problem, a multi-agent collaborative reasoning model is adopted to generate a multi-step reasoning path based on the fault knowledge graph. Compared with the prior art, intelligent question answering with high reliability and high readability can be efficiently realized in a complex industrial scene.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of industrial intelligence, and in particular relates to an intelligent device fault answering method enhanced by a large model. Background Art

[0002] Improving the effectiveness of equipment fault diagnosis directly promotes production efficiency and ensures the safe operation of equipment. However, traditional fault diagnosis methods usually rely on manual experience or simple data analysis tools, and it is difficult to cope with the complex dynamic associations between equipment and diverse fault scenarios. For example, when multiple devices in a complex system experience collaborative failures, it may be difficult to accurately locate the root cause solely through manual analysis. In addition, traditional methods are usually based on preset rules and static models, lacking flexibility and real-time adaptation capabilities, and showing obvious limitations in the face of changing production environments. Especially in the context of the rapid growth of data scale and the increasing complexity of fault causes, traditional diagnosis methods are gradually showing their deficiencies in terms of efficiency and accuracy.

[0003] As a technology that can integrate and express structured knowledge, knowledge graphs have been widely used in the field of industrial intelligence in recent years. In the fault diagnosis of steel production line equipment, knowledge graphs can form a complete knowledge network by modeling equipment, sensor data, and process flows, providing basic support for the reasoning of complex problems. The question-answering system based on knowledge graphs can discover potential knowledge relationships from structured data through logical reasoning, providing more diverse information for equipment fault diagnosis. Although the question-answering system based on knowledge graphs performs better than traditional methods in equipment fault diagnosis, there are still the following problems in practical applications: 1) Traditional single-layer reasoning-based knowledge graphs are difficult to efficiently handle complex problems involving cross-equipment and cross-systems. Limited by the diversity and quality of data sources, there are generally sparsity and incompleteness problems in the construction of knowledge graphs, resulting in interrupted reasoning paths or insufficient coverage of reasoning results, further restricting their application potential. 2) The expression methods of user questions are diverse, which may involve semantic metaphors or complex sentence patterns. The semantic understanding ability of traditional systems is insufficient, and it is easy to produce misunderstandings or ignore key information when dealing with such questions, significantly affecting the accurate parsing of user questions by the system; and it is difficult to dynamically adjust the reasoning strategy according to the complexity of the question, thus unable to balance the efficiency of single-hop questions and the deep reasoning requirements of multi-hop questions, lacking flexibility. 3) The answers generated by traditional systems are mostly presented in structured data, lacking the coherence and readability of natural language, making it difficult for users to understand. Therefore, it is necessary to improve the existing methods to efficiently achieve intelligent question-answering with high reliability and strong readability in complex industrial scenarios, providing a new solution for equipment fault diagnosis. Summary of the Invention

[0004] The object of the present invention is to overcome the defects existing in the above-mentioned prior art and provide an intelligent device fault Q&A method enhanced by a large model, which can efficiently achieve intelligent Q&A with high reliability and strong readability in complex industrial scenarios.

[0005] The object of the present invention can be achieved by the following technical solutions:

[0006] The present invention provides an intelligent device fault Q&A method enhanced by a large model, including the following steps:

[0007] Obtain steel production line data and construct a fault knowledge graph. The steel production line data includes equipment operation data, historical fault records, maintenance and repair records, operation logs, and production line environment data;

[0008] Obtain the user's question and classify it. The classification results include single-hop questions, short-hop questions, and multi-hop questions;

[0009] According to the category to which the user's question belongs, call the corresponding reasoning mechanism, generate an answer based on the fault knowledge graph, and output a device fault analysis in natural language form through retrieval enhancement generation technology; among them, for single-hop questions, match the core entities and relationships based on the Aho-Corasick automaton, and directly retrieve the answer from the fault knowledge graph; for short-hop questions, use a single-agent reasoning model to deduce the answer from the fault knowledge graph; for multi-hop questions, use a multi-agent collaborative reasoning model to generate a multi-step reasoning path based on the fault knowledge graph.

[0010] Further, use the F-Qwen2-7B model to perform semantic analysis on the user's question to obtain the classification result. The specific process is as follows:

[0011] Embed the user's question through the F-Qwen2-7B model to obtain a high-dimensional vector representation of the user's question. Calculate the cosine similarity between the high-dimensional vector representation and a predefined question set in the field of steel production lines. If the calculated cosine similarity is higher than the set threshold, proceed to the next step; otherwise, return a standardized prompt message;

[0012] Use the F-Qwen2-7B model to generate a set of questions with similar semantics to the user's question to assist in analyzing the complexity of the user's question, and then obtain the classification result.

[0013] Further, use the QLoRA strategy to fine-tune the F-Qwen2-7B model.

[0014] Further, for single-hop questions, first use the Aho-Corasick automaton to match the core entities and relationships, and then select a predefined Cypher statement to retrieve relevant triples from the fault knowledge graph as the answer.

[0015] Furthermore, the single-agent reasoning model includes an embedding module, a knowledge graph reasoning module, and an answer decision module;

[0016] The embedding module includes a question embedding sub-module and a knowledge graph embedding sub-module. The question embedding sub-module uses the Chinese-RoBERTa-wwm-ext model to embed the user's question into a low-dimensional vector space and generate a hidden representation of the user's question. The knowledge graph embedding sub-module uses the ConvE model to learn the embeddings of all entities and relationships in the fault knowledge graph through convolution operations;

[0017] The knowledge graph reasoning module includes an agent and an external environment, and the external environment is modeled as a Markov decision process, represented by the tuple (S, A, p, R a ), where S represents the state space, which contains all entities and relationships in the fault knowledge graph, A represents the action space, corresponding to the relationship edges selected by the agent in the fault knowledge graph, p is the state transition probability, used to describe the probability of transitioning to the next state given the current state and the selected action, and R a represents the reward function, used to feedback the decision-making quality of the agent; the knowledge graph reasoning module is driven by the Actor-Critic reinforcement learning method and is used to generate multiple reasoning paths and corresponding candidate answers;

[0018] The answer decision module is used to select the candidate answer with the highest score as the output according to the hidden representation of the user's question through the attention mechanism and the multi-path aggregation technology.

[0019] Furthermore, the process of the answer decision module selecting the candidate answer with the highest score is specifically as follows:

[0020] Calculate the attention score between the user's question q and the i-th reasoning path p i and the attention score between the user's question q and the i-th candidate answer a i

[0021]

[0022] where ATT is the dot product attention operation, and h q , are the hidden representations of the user's question q, the i-th reasoning path p i , and the i-th candidate answer a i respectively;

[0023] Calculate the aggregated path representation h agg : ​​

[0024]

[0025] where w i is the weight of the i-th inference path .

[0026] Connect h q and h agg , and obtain the probability P(a i |q) that the candidate answer a i is the correct answer for the given question q through a linear classifier:

[0027]

[0028]

[0029] where W is the weight matrix of the linear classifier and b is the bias term.

[0030] Furthermore, the multi-agent collaborative inference model includes a relation selection agent, an entity selection agent, and a fact selection agent. The relation selection agent and the entity selection agent are used to collaboratively infer the inference path according to the fault knowledge graph, and the fact selection agent is used to expand the fault knowledge graph through dynamic knowledge completion technology.

[0031] Furthermore, the fact selection agent is defined as a quadruple (S FE , A FE , R FE , π FE ), where S FE , A FE , R FE , π FE represent the state space, action space, reward, and policy of the fact selection agent respectively;

[0032] When the relation selection agent reaches a certain entity e i , the fact selection agent extracts the fact triples F i related to the entity e RS from the external corpus and merges them into the fault knowledge graph. At this time, the state is defined as where b(e i ) represents the set of candidate facts related to the current entity e i , and a tti represents the attention embedding, which is calculated by the following formula:

[0033]

[0034] where αij is the attention score calculated based on the graph attention network, which is used to measure the relevance between the current inference path and the candidate triple, f ij is the concatenation of the vector embeddings of the entities and relations of each triple;

[0035] When the entity selection agent reaches the relation r i the fact selection agent extracts the fact triples F i-1 , r i ) related to (e ES from the external corpus and merges them into the fault knowledge graph. The state at this time is defined as

[0036] Furthermore, when training the multi-agent collaborative inference model, an alternating training and joint optimization strategy is adopted.

[0037] Furthermore, the loss function of the multi-agent collaborative inference model is constructed based on the loss functions of the relation selection agent, the entity selection agent, and the fact selection agent. An entropy regularization term is introduced into the loss function of each agent.

[0038] Compared with the prior art, the present invention has the following beneficial effects:

[0039] 1. In the device fault intelligent question answering method proposed by the present invention, a hierarchical inference strategy is applied. By dividing user questions into single-hop questions, short-hop questions, and multi-hop questions, and adopting a differentiated inference mechanism to generate answers for questions of different complexities, the efficiency of inference and the accuracy of answers can be significantly improved. Specifically, for single-hop questions, the present invention matches the core entity and relation based on the Aho-Corasick automaton and quickly retrieves answers from the fault knowledge graph; for short-hop questions, the present invention uses a single-agent inference model to deduce answers from the fault knowledge graph and provides an efficient short-range inference path; for multi-hop questions, the present invention uses a multi-agent collaborative inference model to generate multi-step inference paths based on the fault knowledge graph, which can provide more accurate answers to complex questions within a limited number of inference steps, thereby efficiently realizing intelligent question answering with high reliability and strong readability in complex industrial scenarios and providing a new solution for device fault diagnosis.

[0040] 2. The present invention combines the semantic understanding and natural language generation capabilities of large language models, significantly improving the coherence and fluency of question answering. Traditional knowledge graph question answering systems usually return answers in the form of structured data or triples, making it difficult for users to understand or directly apply; the present invention combines retrieval-enhanced generation technology, which can convert the inference path into an inference chain described in natural language, and finally generates a logically clear and natural answer, avoiding the problem of information fragmentation, enabling users to quickly obtain key information, and significantly improving the user experience.

[0041] 3. The dynamic knowledge completion technology of the present invention expands the fault knowledge graph, solving the problem of broken reasoning paths caused by sparsity and incompleteness in traditional knowledge graphs. By introducing a fact extraction mechanism, the present invention can dynamically extract relevant facts from external corpora and update the fault knowledge graph, effectively filling knowledge gaps, ensuring that the system always has the latest domain knowledge, and enhancing the applicability and robustness of the model.

[0042] 4. The present invention uses the QLoRA strategy to fine-tune the F-Qwen2-7B model, which can greatly reduce the computational resource requirements for model training and inference while maintaining high-performance performance. QLoRA achieves parameter-efficient fine-tuning by optimizing the low-rank representation of key matrices in the model and only adjusting a small number of parameters in the attention mechanism. This method not only reduces the video memory occupancy but also ensures the integrity of the model's basic knowledge and avoids the occurrence of catastrophic forgetting. In addition, QLoRA combined with quantization technology enables the model to run on ordinary hardware, providing broader possibilities for the industrial deployment and practical application of the model. Compared with traditional full-parameter fine-tuning methods, the present invention can significantly reduce the computing cost and meet the application requirements of large models in complex industrial scenarios. BRIEF DESCRIPTION OF THE DRAWINGS

[0043] Figure 1 is a flowchart of the method of the present invention;

[0044] Figure 2 is a schematic diagram of the overall structure of the LLMs-KGR-QA model;

[0045] Figure 3 is a schematic diagram of the inference process of the TA-KGR model;

[0046] Figure 4 is a schematic diagram of the output result of the LLMs-KGR-QA model. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0047] The present invention will be described in detail below with reference to the accompanying drawings and specific embodiments. This embodiment is implemented on the premise of the technical solution of the present invention, and detailed implementation manners and specific operation processes are given, but the protection scope of the present invention is not limited to the following embodiments.

[0048] Embodiment:

[0049] This embodiment provides a method for intelligent equipment fault answering based on large model enhancement, which is implemented through the LLMs-KGR-QA model as shown in Figure 2 . This method is as shown in Figure 1 and includes the following steps:

[0050] S1. Obtain steel production line data and construct a fault knowledge graph.

[0051] Obtain multi-source data from the steel production line, including equipment operation data, historical failure records, maintenance and repair records, operation logs, and production line environment data. Clean, format, and extract features from the data to ensure the integrity and consistency of the data input into the model.

[0052] S2. Obtain the user's question, perform domain judgment and question classification, and the classification results include single-hop questions, short-hop questions, and multi-hop questions.

[0053] First, it is necessary to determine whether the question input by the user belongs to the steel production line domain. For this purpose, in this embodiment, through an embedded post-similarity matching mechanism, it is ensured that the semantics of the user's question is highly consistent with the knowledge base of the steel production line domain. Specifically, the user's question is embedded through Chinese-RoBERTa-wwm-ext to obtain a high-dimensional vector representation q of the user's question, and then this high-dimensional vector representation is compared with the embedded representation d of the predefined question set in the steel production line domain. i Perform cosine similarity calculation, and the calculation formula is: If the calculated cosine similarity is higher than the set threshold τ (such as 0.7), it is considered that the question belongs to the steel production line domain and proceed to the next step; otherwise, return a standardized prompt message, such as "This question does not belong to the steel production line domain and no relevant answers can be provided", to avoid generating misleading answers.

[0054] To further optimize the setting of the similarity threshold τ, the threshold can be dynamically adjusted based on the distribution of questions within the domain. Specifically, an optimal threshold can be selected by maximizing the difference between the similarity distribution of questions within the domain and the threshold:

[0055]

[0056] Among them, is an indicator function, which is counted as 1 when the similarity is greater than or equal to the threshold, and 0 otherwise. Through the embedded post-similarity matching mechanism, the system can focus on the question-and-answer tasks in the steel production line domain, and at the same time ensure that when the question is not within the domain, the system will not generate irrelevant or low-quality answers, thereby enhancing the overall stability and credibility of the question-and-answer system.

[0057] Once the problem is determined to be domain-related, the model further analyzes the complexity of the problem. At this time, the F-Qwen2-7B model is used to generate a set of questions with similar semantics to the original user question. These similar questions help the model understand the problem from different perspectives and assist in the construction of the reasoning chain. For example, if the user question is: "When commissioning the SP404 reducer on the ironmaking production line, it is found that the input shaft bearing is damaged. What is the cause of the failure?", the list of similar questions generated based on the F-Qwen2-7B model is: [1. When commissioning the SP404 reducer on the ironmaking production line, there is abnormal noise on the input shaft. What are the possible causes of bearing damage?, 2. What are the possible causes of the damage to the input shaft bearing of the SP404 reducer?, 3. Why does the input shaft bearing of the SP404 reducer get damaged during commissioning?]. Then, the questions are classified according to complexity, divided into single-hop questions, short-hop questions or multi-hop questions, and the classification result determines which reasoning mechanism will be used next.

[0058] In a preferred embodiment, the QLoRA strategy is used to fine-tune the F-Qwen2-7B model to significantly reduce the computational resource requirements.

[0059] S3. According to the category to which the user question belongs, call the corresponding reasoning mechanism, construct a reasoning path based on the fault knowledge graph to generate an answer. The specific reasoning mechanisms corresponding to each category are as follows:

[0060] (1) Single-hop questions

[0061] For single-hop questions, the rule matching method is used to recall relevant knowledge from the fault knowledge graph. For the question raised by the user, first use the Aho-Corasick automaton to match the core entity and relationship, and then select the predefined Cypher statement to retrieve relevant triples from the fault knowledge graph as the answer. For example, when the user question is "What is the cause of the coiler roll failure?", the F-Qwen2-7B model will locate the fault instance entity "coiler roll" and the relationship "cause of failure", and then return the answer through Cypher query: The roller body has a fatigue fracture.

[0062] The above process can be formalized as a simple relationship retrieval problem:

[0063] Retrieve(e,r)={(e,r,t)|e=e and r=r}

[0064] Where e is the core entity, r is the relationship, and a set of relevant triples is retrieved from the fault knowledge graph through Cypher query. The relevant Cypher statements are as follows:

[0065]

[0066] In the core process of single-hop problem processing, the Aho-Corasick automaton combines natural language processing and automaton technology to achieve precise identification and semantic analysis of key entities. This mechanism is based on the trie structure and is extended by constructing mismatch pointers (fail pointers) and using the breadth-first search algorithm, enabling efficient handling of multi-pattern matching problems. When the user asks a question, the Aho-Corasick automaton first uses its built-in semantic analysis ability to locate the core entities in the question. These core entities are keywords closely related to the semantics of the question. In this way, the Aho-Corasick automaton filters out irrelevant information and thus more accurately understands the user's intention. This process is similar to scanning the text in the Aho-Corasick automaton and identifying the pattern strings. Similar to the process where the Aho-Corasick automaton uses the fail pointer to continue matching other pattern strings after successful pattern matching. By this method, the model can efficiently identify and deduce the key relationships in the question. For example, the model will identify the core entities in the user's question and combine the structured information in the steel production line fault knowledge graph to find the potential relationships between these entities. This process of identifying semantic associations, like the state transition in the Aho-Corasick automaton, can identify deep relationships in complex questions. After identifying the core entities and relationships, a predefined Cypher query statement is selected to retrieve triple data in the steel production line fault knowledge graph. By this method, the model can transform natural language questions into executable query statements and provide answers for users by recalling relevant triples and subgraphs in the steel production line fault knowledge graph.

[0067] (2) Short-hop problem

[0068] For short-hop problems, a single-agent reasoning model (SA-KGR model) is used for reasoning. The single-agent reasoning model (SA-KGR model) includes three core modules: an embedding module, a knowledge graph reasoning module, and an answer decision module.

[0069] 1) Embedding module: It mainly includes a question embedding sub-module and a knowledge graph embedding sub-module. Through these two sub-modules, the SA-KGR model can construct a semantically aligned embedding space, providing a basis for subsequent reasoning and answer selection. The question embedding sub-module uses the Chinese-RoBERTa-wwm-ext model to embed the user's question into a low-dimensional vector space, generating a hidden representation of the user's question to fully capture the semantic information in the user's question; the knowledge graph embedding sub-module uses the ConvE model to learn the embeddings of all entities and relationships in the fault knowledge graph through convolutional operations, which can effectively handle the complex information of steel production line faults.

[0070] 2) Knowledge Graph Reasoning Module: Driven by the Actor-Critic reinforcement learning method, it is used to generate multiple reasoning paths and corresponding candidate answers, including an agent and an external environment. The external environment is modeled as a Markov decision process MDP, represented by a tuple (S, A, p, R a ). Among them, S represents the state space, which contains all possible entities and relationships in the fault knowledge graph. A represents the action space, corresponding to the relationship edges selected by the agent in the fault knowledge graph (such as "fault impact", "preventive measures", etc.). p is the state transition probability, used to describe the probability of transitioning to the next state given the current state and the selected action. R a represents the reward function, which is used to evaluate the quality of the agent's execution of an action in a given state and is a key factor driving the agent's learning.

[0071] The agent selects actions based on the current state. The action space consists of all possible relationship types (such as "fault impact", "preventive measures", etc.). The goal of the agent is to optimize its behavior strategy by maximizing the cumulative reward.

[0072] The Actor-Critic policy network is a reinforcement learning algorithm that combines the advantages of policy optimization and value estimation and is widely used to solve tasks in continuous action spaces or high-dimensional state spaces. The Actor-Critic algorithm has two main components: the Actor network and the Critic network. The Actor network is a neural network used to generate the agent's policy, that is, to generate an action distribution. In reinforcement learning, an episode refers to a complete process from when the agent starts interacting with the environment until a certain termination condition is reached. In each episode, the agent explores paths by sampling actions. The policy function π(a|s; θ) represents the probability of selecting action a in the given state s, and the parameter θ is the learnable parameter of the policy network. The policy network learns the optimal policy by optimizing the parameters to maximize the cumulative reward. Assume the state dimension is d and the size of the action space is n. The policy network has three layers: the input layer, the hidden layer, and the output layer. The mapping from the input layer to the hidden layer:

[0073] h 1 = ReLU(W 1 s + b 1 )

[0074] Among them, is the weight matrix, is the bias, and ReLU is the activation function.

[0075] The mapping from the hidden layer to the output layer:

[0076] h 2 = ReLU(W 2 h1 +b 2 )

[0077] π(a|s; θ) = softmax(W 3 h 2 +b 3 )

[0078] where π(a|s; θ) represents the probability distribution of actions, is the weight matrix, b 2 and b 3 are biases.

[0079] The Critic network is used to estimate the value of a given state (i.e., the expected return that the agent can obtain in that state), and it consists of an input layer, a hidden layer, and an output layer. In the forward propagation process, a multi-layer perceptron is used to process the input state s, and the value estimate V(s) of this state is output:

[0080] V(s) = W 3 ·ReLU(W 2 ·ReLU(W 1 ·s + b 1 ) + b 2 ) + b 3

[0081] where W 1 , W 2 , W 3 are the weight matrices of each layer respectively, and b 1 , b 2 , b 3 are bias vectors. The Critic network affects the Actor network by feeding back an advantage function A(s, a). In the Actor-Critic method, the Actor policy network is used to generate the action distribution, and the Critic value network is used to estimate the state value function. The Actor updates the policy based on the feedback provided by the Critic to improve the efficiency and stability of the policy.

[0082] In this embodiment, the policy network (Actor network) generates the probability distribution of each action, and the agent randomly samples an action according to these probabilities. After an action is selected, the agent calls the interact function to obtain feedback to determine whether the action can form a valid path, and returns different reward values according to the following three situations: ① If the new position is the same as the target position, the reward is +1, indicating that the path is found, and the system records the path and updates the state; ② If the path is valid but the target has not been reached, the reward is 0, and the agent continues to make decisions until the target is found or the step limit is reached; ③ If the path is invalid, the reward is -1, and the state of the agent remains unchanged. At the same time, a counter die is incremented to record the number of consecutive selections of invalid paths. This helps the agent identify invalid paths, optimize future decisions, and make it more likely to choose paths that may be valid.

[0083] During the training process, if the agent has difficulty finding a valid path from the current state to the target, a path search function, i.e., the Breadth-First Search (BFS) strategy, is introduced to systematically traverse all possible paths and preferentially find valid paths. The introduction of BFS provides more useful training samples, avoids getting stuck in infinite loops or repeatedly selecting invalid paths, thereby improving the learning effect.

[0084] 3) Answer decision module: Integrates technologies such as embedded representation, attention mechanism, and multi-path aggregation to achieve efficient answer decision-making. The answer decision module selects the candidate answer with the highest score as the output. The specific process is as follows:

[0085] Calculate the attention score between the user question q and the i-th reasoning path p i of

[0086]

[0087] where h q , are the hidden representations of the user question q and the i-th reasoning path p i respectively, ATT is the dot product attention operation, used to represent the similarity between h q and . By calculating the similarity between the user question and the reasoning path, the system can assign higher weights to more relevant reasoning paths, so that it pays more attention to those paths that are more likely to provide useful information during the reasoning process. This can prevent the model from blindly choosing the wrong path, especially when multiple reasoning paths exist simultaneously, and helps the model preferentially select the reasoning path that best matches the semantics of the question.

[0088] Calculate the attention score between the user question q and the i-th candidate answer a iAttention score

[0089]

[0090] Among them, is the hidden representation of the i-th candidate answer a i . By calculating the similarity between the user's question and the candidate answers, this can help the model judge which candidate answer best fits the current question context according to the semantic information of the question.

[0091] During the inference process, multiple inference paths are obtained. Different paths may provide different information, and the importance of each path may be different. By assigning weights to each path, the information of multiple paths can be combined, so that the model is not limited to the results of a single inference path, thereby enhancing the robustness of the inference process. To better utilize the information of these paths, the answer decision module performs weighted aggregation on them to obtain the aggregated path representation h agg :

[0092]

[0093] Among them, w i is the weight of the representation of the i-th inference path . The way of weighted aggregation enables the model to flexibly adjust the contribution of each path according to factors such as the length of the path and the relevance to the question. This can prevent the model from relying only on a single path, but comprehensively consider the influence of all possible paths, thereby improving the accuracy of inference.

[0094] Then, connect the candidate answer representation the question representation h q and the aggregated path representation h agg to form a more complete feature representation to select the final answer. This can not only consider the direct relevance between the question and the candidate answer, but also fully combine the intermediate process in the inference path. This way of information fusion can enhance the model's ability to handle complex inference scenarios.

[0095] Finally, through the linear classifier softmax, obtain the probability P(a i is the correct answer) for the given question q: i |q):

[0096]

[0097] Among them, W is the weight matrix of the linear classifier, and b is the bias term.

[0098] 4) Training and Optimization of the Single-Agent Inference Model (SA-KGR Model): The task of the Critic network is to estimate the state value V(s) of a given state s, and this estimated value is optimized by comparing it with the actual return. The loss function of the Critic network uses the Mean Squared Error (MSE) loss, which is defined as:

[0099]

[0100] where r is the immediate reward at the current time step, γ is the discount factor used to balance short-term and long-term rewards, and r + γV(s ′ ) is a single-step estimate of the cumulative reward for the Temporal Difference (TD) update method. It only considers the current immediate reward r and the value estimate V(s ′ ) of the next state. V(s) is the value estimate of the current state s by the Critic network. V(s ′ ) is the value estimate of the next state s ′ by the Critic network.

[0101] The task of the Actor network is to learn a policy π(a|s; θ) and output the probability distribution of each action. Its update process is based on the advantage function A(s,a), which represents how good an action a is on average under the current policy, and is defined as:

[0102] A(s,a) = r + γV(s ′ ) - V(s)

[0103] The Critic network affects the Actor network by feeding back an advantage function A(s,a). The policy gradient method updates the network parameters by minimizing the loss function to maximize the cumulative reward. The loss function L of the Actor network is defined as:

[0104] L actor = -logπ(a|s; θ)·A(s,a)

[0105] In each iteration, the Actor and Critic networks calculate their respective losses and then update their parameters through independent optimizers, jointly optimizing in the same training process. This joint training method ensures that the Actor learns a better policy, while the Critic continuously improves the accuracy of the state value evaluation, thereby jointly improving the effect of reinforcement learning.

[0106] (3) Multi-hop Problem

[0107] For multi-hop problems, a multi-agent collaborative reasoning model (TA-KGR model) is used for reasoning. The multi-agent collaborative reasoning model includes a relation selection agent, an entity selection agent, and a fact selection agent. Among them, the relation selection agent navigates the reasoning path by selecting appropriate relations, the entity selection agent is responsible for finding the correct entity in the knowledge graph for the next step of reasoning, and the fact selection agent extends the fault knowledge graph through dynamic knowledge completion technology and supports multi-hop reasoning. They collaborate through reinforcement learning to explore complex reasoning paths.

[0108] Given (e q , r q , e t ), where e q represents the query entity, r q represents the query relation, and e t represents the target entity. The relation selection agent and the entity selection agent cooperate to infer the path of the target entity e q related to the query entity e q and the query relation r t . This inference path is used for prediction and provides interpretability. At each time step i, the relation selection agent and the entity selection agent attempt to select an edge based on the available information, and the fact selection agent is trained to suggest the most relevant facts related to the current reasoning step i. Suppose the entity selection agent reaches the entity e i on the steel production line fault knowledge graph G at time step i. The fact selection agent retrieves fact triples (e i , r i , e k ) from the external corpus C and merges them into the steel production line fault knowledge graph G. Therefore, the relation selection agent obtains other options for extending the reasoning path. The specific settings of each agent are as follows:

[0109] 1) Relation selection agent

[0110] The relation selection agent is defined as a 4-tuple (S RS , A RS , R RS , π RS ), which is responsible for determining the relation based on the entity e i at time step i.

[0111] State: The state at step i is represented by a tuple , where e i ∈ ε represents the current entity, R represents all relation spaces, represents the fact triples extracted by the fact selection agent related to the current entity e i . The inference history path embedding vector representing the relationship.

[0112]

[0113] Among them, r i-1 represents the previous relationship connecting e i . The initial historical path vector is:

[0114] h 0 = LSTM([0; r q )

[0115] The initial state is After taking an action, the relationship selection agent transfers to the next state. Action: For the state at the i-th step e i any outgoing edge can be used as a candidate action Using to represent the action space at the i-th step, after executing the action, the relationship selection agent enters the next state.

[0116] Policy network: The purpose of the Actor policy network is to guide the agent to select a suitable action from all available actions . Combine the current entity embedding e i , the historical relationship embedding and the query embedding (e q , r q ) to form the input representation of the policy network:

[0117]

[0118] Process the concatenated input features through the self-attention mechanism to generate weights that focus on relevant entities:

[0119] q = W q ·X r

[0120] k = W k ·X r

[0121] v = W v ·X r

[0122] Calculate the attention scores:

[0123]

[0124] Calculate the attention weights through the softmax function:

[0125] att_probs = softmax(att_scores)

[0126] Weighted sum of values v through attention weights:

[0127]

[0128] Then, process the state representation through a two-layer fully connected network:

[0129]

[0130] X r2 = W r2 (ReLU(X r1 )) + b r2

[0131] Take the dot product with the embedding of the candidate relation and calculate the attention weight x through the softmax function to calculate the probability distribution:

[0132] p(r next | s t ) = softmax(R · X r2 )

[0133] where W r1 and W r2 represent the trainable weights of the policy network, R is the embedding matrix of the candidate relation, and the probability distribution p(r next | s t ) generated by softmax, and the relation selection agent samples actions according to this distribution to determine the relation selected in the next step. Determine the probability that the relation selection agent takes corresponding actions at time step i. The history-dependent policy of the relation agent is designed as The policy π RS represents the set of policies of the agent at each time step. The policy at each time step i will output a probability distribution p(r next | s t ), guiding the agent to select a relation based on historical information and the current state.

[0134] Reward: The reward function R Rs(s,a) is used to evaluate the effect of the agent performing an action in a specific state. The agent optimizes its behavior strategy by maximizing the cumulative reward. The action space consists of various relationship types, such as "fault impact" and "preventive measures", etc. The policy network generates the probability distribution of actions. The agent randomly selects an action according to the probability and calls the interact function to obtain feedback to determine whether the path is valid. The reward mechanism is as follows: when the target path is found, the reward is +1, the path is recorded and the state is updated; when the path is valid but the target is not reached, the reward is 0 and the decision-making continues; when the path is invalid, the reward is -1, the state remains unchanged, and the invalid path counter is incremented. This helps the agent optimize its decision-making and reduce the selection of invalid paths.

[0135] 2) Entity selection agent

[0136] State: The state at step i is defined as a tuple where r i represents the current relationship, represents the fact triple related to the current relationship extracted by the fact selection agent from the external corpus, represents the historical embedding vector of the inference path. Update the historical embedding vector according to the historical information:

[0137]

[0138] In the first step, the entity e i is equal to the initial query entity e q , and the starting historical embedding vector is defined:

[0139]

[0140] The starting state is represented as

[0141] Action: In the state at step i , any out-edge of r i can be used as a candidate action The action space at step i is defined as where A ES represents the entire action space. After executing an action, the agent advances to the next state.

[0142] Policy network: The policy network is used to guide the selection of the optimal action from all available actions. The policy network combines the embedding r of the current relationship i , the historical embedding and the query embedding (e q , r q ) to form the current state representation:

[0143]

[0144] Similarly, the self-attention mechanism is introduced:

[0145]

[0146] Then, the state representation is processed through two layers of fully connected networks:

[0147]

[0148] X e2 = W e2 (ReLU(X e1 )) + b e2

[0149] Perform a dot product with the candidate entity embedding and calculate the probability distribution using softmax:

[0150] p(e next | s t ) = softmax(E · X e2 )

[0151] where W e1 and W e2 represent the trainable weights of the policy network, E is the candidate entity embedding matrix, and the probability distribution output by softmax is used to select the next entity e next , that is, according to sampling actions determine the probability that the entity selection agent takes corresponding actions at time step i. The historical dependence strategy of the entity selection agent is designed as

[0152] Reward: The reward function R RS (s, a) is used to evaluate the effect of the agent executing actions in a specific state. The reward mechanism is the same as that of the relation selection agent.

[0153] In the above way, the agent can gradually approach the goal according to the output of the policy network in different states.

[0154] 3) Fact selection agent

[0155] When the collaborative reasoning of the entity and relation selection agents does not reach the target entity, the fact selection agent will be enabled. The goal of the fact selection agent is to provide the most relevant fact triples for the entity e i or relation r i in the steel production line fault knowledge graph G at time step i. These relevant facts will be temporarily added to the steel production line fault knowledge graph G to expand the action space of the relation and entity selection agents. The structure of the fact selection agent can be defined as a quadruple (SFE , A FE , R FE , π FE ).

[0156] Status: When the relation selection agent reaches an entity e i , the fact selection agent will extract the fact triples F i related to e RS from the external corpus. At this time, the status is defined as where S FE represents the entire state space, b(e i ) represents the set of candidate facts related to the current entity e i , a tti represents the attention embedding, which helps the agent select the facts most relevant to the current entity. When the entity selection agent reaches r i , the fact selection agent will extract the fact triples F i-1 related to (e i , r ES ) from the external corpus. At this time, the status is defined as

[0157] The purpose of introducing the graph attention network is to help the fact selection agent select the fact triples related to the current reasoning path. The GAT network calculates the matching scores between each candidate triple and the current reasoning path through the self-attention mechanism. Specifically, the attention score α ij is used to measure the correlation between the current reasoning path and the candidate triple:

[0158]

[0159] where the attention vector a T dynamically assigns weights between different nodes, ensuring that important nodes have a greater influence on the final output. represents the historical representation of the current reasoning path, e q and r q represent the embeddings of the query entity and relation respectively, r ij represents the relation embedding in the candidate triple, and W is the linear transformation matrix. By concatenating the current state information (including historical path, query entity, relation, etc.) and the embeddings of the candidate triple, and using linear transformation and LeakyReLU activation to calculate the attention score.

[0160] Action: The role of the fact selection agent is to select the triples related to the reasoning process from the external corpus. For each fact (e i ) extracted from b(e i , r i , ei+1 ) to obtain an action The action space at the i-th step is denoted as For each triple (e i , r i , e k ), the embedding representation f ij According to the attention score α ij perform weighted summation to generate the attention embedding a tti :

[0161]

[0162] where f ij is the concatenation of the vector embeddings of the entities and relations of each triple.

[0163] Through the above steps, the fact selection agent can focus its attention on the more relevant triples, thereby generating the comprehensive feature a tti . The attention score α ij measures the relevance between the candidate triple and the current inference path (represented by the historical representation query entity e q and relation r q ). This score determines the degree of influence of each candidate triple on the current inference. The attention embedding a tti represents a comprehensive result containing the relevance of each candidate triple after being processed by GAT. As a state feature, it is input into the policy network together with other state information (such as the embeddings of entities and relations) for further decision-making.

[0164] The goal of the policy network is to combine the calculated weighted feature a tti with the current state feature to generate the final action probability distribution. The design of the policy network is as follows.

[0165] First, the state feature a tti is processed through a two-layer fully connected neural network:

[0166] X e1 = ReLU(W e1 ·a tti + b e1 )

[0167] X e2 = W e2 ·X e1 + b e2

[0168] Then, the output X e2Take the dot product with the embedding matrix of the candidate relation or entity, and use the softmax function to calculate the probability distribution:

[0169] p(e next |s t ) = softmax(E · X e2 )

[0170] where W e1 and W e2 are the trainable weights of the policy network, and E is the embedding matrix of the candidate entity or relation. Finally, sample according to the policy to select the next action

[0171] Reward: When the fact selection agent successfully extracts relevant fact triples and infers the correct target path, the path is considered valid. In this case, the reward mechanism is the same as that of the relation selection agent and the entity selection agent.

[0172] 4) Training and optimization of the multi-agent collaborative inference model (TA-KGR model):

[0173] Each agent in the training can temporarily freeze the parameters of other agents and focus on its own optimization. During the training process, the following steps can effectively improve the overall performance.

[0174] Initial training: First, independently train the relation selection agent to select the optimal relation related to the current entity at each time step. This stage is mainly optimized through a random sampling strategy, that is, the agent collects experience while exploring various possible paths, continuously adjusts the strategy, and enables it to select the optimal relation from a wide range of relation spaces.

[0175] Alternating training: Once the relation selection agent has the basic ability, the alternating training stage will follow. In this stage, first freeze the parameters of the relation selection agent, and start training the fact selection agent and the entity selection agent to maximize their respective rewards on the given external knowledge base and inference path. During the alternating training process, the system only optimizes one agent each time, and the other agents keep their parameters unchanged. This can ensure that each agent can effectively learn at different stages.

[0176] Joint Optimization: As the individual optimization processes of each agent progress, the system gradually introduces a joint optimization strategy that allows the agents to share inference path information and influence each other. In this way, each agent can not only maximize its own reward but also enhance its cooperation in the overall path inference process, maximizing the global reward of the model. Specifically, joint optimization regulates the comprehensive weights of the reward function to make the interactions between agents have a positive effect on improving the model performance.

[0177] The advantage of this alternating training and joint optimization strategy is that it not only ensures that each agent can gradually optimize its own strategy but also enables them to learn to weigh and choose in cooperation. Through this progressive optimization process, the model can gradually improve the efficiency of global path inference and thus achieve better performance in complex steel production line fault knowledge graph inference tasks.

[0178] In path search, the goal of the entity selection agent is to select the next entity e at each time step t+1 , and the probability of this selection is given by the policy network, i.e., P(e t+1 | s t ). The goal is to maximize the expected cumulative reward R t of selecting entities under this policy. That is, the objective function of the entity selection agent is:

[0179]

[0180] where R t is the cumulative reward from the current time step t to future time steps, defined as:

[0181] R t = r t + γr t+1 + γ 2 r t+2 + …

[0182] That is, the cumulative reward R t is the sum of the immediate reward r t at the current time step t and the discounted future time step rewards. The discount factor γ controls the importance of future rewards. The closer the value is to 1, the more important future rewards are, and the closer it is to 0, the weaker the importance of future rewards.

[0183] To encourage the agent to search for diverse paths and prevent it from falling into local optima, an entropy regularization term is added to the loss function. The definition of the entropy H(π(s t )) is:

[0184]

[0185] where H(π(st ) is the entropy of policy π in state s t By introducing entropy regularization, the agent will be encouraged to explore more possible paths rather than just choosing known high-reward paths.

[0186] The loss function of the entity selection agent after entropy regularization is

[0187]

[0188] where β e is the weight of entropy regularization, controlling the influence degree of entropy on the overall objective.

[0189] Similarly, the loss function of the relation selection agent is obtained:

[0190]

[0191] The loss function of the fact selection agent:

[0192]

[0193] To effectively train the three agents, a joint loss function is adopted. The loss functions of each agent are combined into the final objective function:

[0194] L joint (θ) = α entity L entity (θ) + α relation L relation (θ) + α fact L fact (θ)

[0195] where α entity , α relation and α fact are the weight coefficients of the loss functions. The optimal combination is determined through grid search experiments. α entity = 0.4, α relatioh = 0.4, α fact = 0.2, ensuring that each agent can effectively cooperate in different tasks.

[0196] To update the parameters of the agent, the Adam optimizer is used. During the training process, the policy gradient mechanism is combined, and the agent's policy is adjusted through reward reshaping to ensure that each agent can effectively learn how to maximize its reward in its respective task. By optimizing the above joint loss function, the overall inference ability of the model will be improved.

[0197] S4. Output the device fault analysis in the form of natural language through the Retrieval-Augmented Generation (RAG) technology.

[0198] LLMs are very sensitive to the transformation method of input prompt templates. To solve this problem, this embodiment chooses to follow RAG prompt tuning. Based on the extracted path, the F-Qwen2-7B model transforms the generated reasoning path into an easy-to-understand natural language description and outputs the answer to the user in a structured form. At the same time, it supports displaying the reasoning path through a visualization tool to help users intuitively understand the cause and solution of the problem. For example: Reasoning chain description: "The oil cylinder of the coiler leaks oil, resulting in abnormal hydraulic system and further causing the equipment to stop running. It is recommended to replace the gland bolts, clean the oil leakage area and check the sealing status." After the reasoning is completed, the LLMs-KGR-QA model can generate a fault diagnosis result containing the following information:

[0199] Fault type: The specific fault type of the clearly diagnosed equipment or system, such as "hydraulic system failure".

[0200] Fault cause: The direct cause of the fault obtained through reasoning, such as "oil leakage".

[0201] Fault location: The specific location or component where the fault occurs clearly diagnosed during the diagnosis, such as "coiler oil cylinder".

[0202] Reasoning path: The complete reasoning chain, including all key entities and relationships during the reasoning process, such as "hydraulic system -> oil leakage -> gland bolt loosening".

[0203] Solution: Fault solution measures recommended based on the reasoning path, such as "replace the gland bolts".

[0204] Example: For a given prompt

[0205]

[0206] The F-Qwen2-7B model can transform the reasoning path into a reasoning chain:

[0207]

[0208] Through the natural language description of the reasoning chain, users can clearly understand the complete reasoning process and answer of the problem.

[0209] To verify the effectiveness of the above method, this embodiment conducts experimental verification on a dataset of a certain steel production line. The experimental environment is shown in Table 1. The experiment uses Python and is implemented using libraries such as Transformers and PyTorch. These libraries provide powerful support and frameworks for the training of large language models. In the experiment, the QLoRA fine-tuning method is selected to fine-tune the large model, and the fine-tuning parameters are shown in Table 2. To construct a multi-label attribute graph, Neo4j is used, and its Cypher graph query language is used to efficiently operate graph data. To demonstrate the effectiveness of the model, it is evaluated on the already constructed steel production line fault dataset.

[0210] In the experiment, to comprehensively evaluate the performance of different models in the steel production line fault knowledge answering task, multiple evaluation criteria such as Hits@10, MRR, and response speed are adopted. These criteria not only cover the accuracy and comprehensive performance of the model but also involve the inference efficiency and resource consumption of the model.

[0211] Table 1 Experimental Environment

[0212]

[0213] Table 2 Parameter Settings

[0214]

[0215]

[0216] In the comparative experiments of the LLMs-KGR-QA model, this embodiment selected large language models of various scales, including Chinese-llama-2-7b, GPT4O, the quantized Qwen-72B, and Qwen-7B. Among them, Chinese-llama-2-7b is a Chinese pre-trained large language model optimized for Chinese natural language processing tasks. Its training on Chinese data enables it to perform excellently in tasks such as Chinese text generation, semantic understanding, and question-answering reasoning, and is suitable for reasoning tasks with high requirements for Chinese scenarios. Its relatively small parameter scale also allows for efficient deployment with limited computing resources. GPT4O is a large-scale general large language model with powerful natural language understanding and generation capabilities, capable of handling multiple languages and complex tasks. Qwen-72B has 7.2 billion parameters. Through the QLoRA quantization technology, its model size is compressed, and the video memory requirement after quantization is reduced to about 36GB, which can run on hardware such as NVIDIA A100 80GB, thus achieving efficient reasoning with low resource consumption. Although Qwen-7B has fewer parameters, it performs well in small-scale tasks or resource-constrained situations, especially maintaining high reasoning accuracy in scenarios that require strong semantic understanding capabilities. In addition, to prove the role of the single-agent reasoning model and the multi-agent collaborative reasoning model in the entire framework, this embodiment also conducted ablation studies, and the ablation models are shown in Table 3.

[0217] Table 3 Ablation Models

[0218]

[0219] This embodiment verified the performance of the LLMs-KGR-QA model in processing steel production line fault knowledge question-answering tasks through multiple experiments. The experiments were carried out respectively in question classification and fault knowledge question-answering tasks, and compared with other comparative models to comprehensively evaluate the advantages of the model.

[0220] (1) Ablation Experiment

[0221] The ablation experiment aims to evaluate the contribution of each component of the LLMs-KGR-QA model to the inference performance. Specifically, by removing the rule matching method, single-agent (SA-KGR) and multi-agent (TA-KGR) modules, the impact on Hits@10, MRR, and response speed is observed. As can be seen from the experimental results in Table 4, the rule matching method shows an extremely fast response speed (0.05 seconds). However, due to relying on predefined rules, its Hits@10 and MRR are relatively low, 79.49% and 71.02% respectively, and it is suitable for handling simple single-hop questions. In contrast, SA-KGR and TA-KGR introduce a reasoning mechanism of reinforcement learning, resulting in a significant improvement in accuracy. Especially, TA-KGR performs excellently in complex multi-hop questions, with Hits@10 reaching 97.36% and MRR being 91.75%. Although the response speed decreases slightly, compared with the rule matching method, these two models have better adaptability in complex reasoning scenarios. The LLMs-KGR-QA model performs the best, with Hits@10 reaching 98.67% and MRR being 95.36%, and the response speed being 4.3 seconds. The experimental results show that the combination of the generation ability of large language models and knowledge graph reasoning can provide high-quality natural language answers while maintaining the inference effect.

[0222] Table 4 Comparison test results of ablation study (%)

[0223]

[0224] (2) Case analysis of fault question and answer

[0225] The purpose of this experiment is to verify the improvement of the effect of the question and answer task after the combination of the large model and the steel production line fault knowledge graph. The experiment compares the performance of different large models (Qwen2-7b, Chinese-llama-2-7b, Qwen-72b, GPT4O) and the LLMs-KGR-QA model in dealing with real steel production line fault diagnosis problems. The experimental results and analysis are as follows.

[0226] Question 1: At the site of the 1800 acid continuous rolling mill, it is found that the automatic step of the welder stops abnormally, and an alarm of abnormal cooling water flow (insufficient) is reported, resulting in the stop of the automatic step of the welder. What is the cause of the fault? How to solve it?

[0227] Answer result of Qwen2-7b model

[0228]

[0229]

[0230] The output result of Chinese-llama-2-7b model is as follows:

[0231] The output results of the Qwen-72B model are as follows:

[0232]

[0233]

[0234] The output results of the GPT4O model are as follows:

[0235]

[0236]

[0237] Figure 3 It shows the reasoning process of the TA-KGR model in the steel production line fault knowledge graph. Based on the user's question, the TA-KGR model starts from the entity node of the 1800 acid continuous rolling mill unit and gradually reasons to obtain the answer related to the question along the path composed of equipment, faults, fault phenomena, fault causes, and solutions. Based on the reasoning path obtained by the TA-KGR model and the user's question, the output results of the LLMs-KGR-QA model are as Figure 4 shown.

[0238] It can be seen from the results that the answers of the Qwen2-7b model and the Chinese-llama-2-7b model are relatively vague and lack targeted analysis of specific problems. For example, it is not clearly explained whether the cooling water temperature really constitutes a problem in the current scenario. In contrast, the Qwen-72B model provides a more detailed analysis, covering various components of the cooling water system, including water pumps, pipelines, sensors, etc., but the explanation of how the control system of the welder itself affects the abnormal stop of the automatic step is not deep enough. In addition, the answer of the GPT4O model has a wide coverage, proposing various possibilities such as cooling water flow sensor failure, cooling water pipeline blockage, and cooling water pump failure, but the model does not clearly give the priority of the solution, and the relevance between multiple potential problems and solutions is not clear enough, which may lead to confusion in the priority order for users in actual operation. The LLMs-KGR-QA model combines the reasoning path of the steel production line fault knowledge graph and proposes the most targeted fault causes and solutions. Compared with other models, this model first proposes the problem of the flow limiting valve being blocked by foreign objects, which is not mentioned by other models. In addition, LLMs-KGR-QA gives a more accurate reasoning path by combining the steel production line fault knowledge graph and the user's question, covering the full-process solution from insufficient cooling water flow to alarm system optimization. The model can gradually lock the root cause of the problem through multi-hop reasoning, and the technical discussion and improvement suggestions for the alarm system also provide a guarantee for the complete solution of the problem.

[0239] The above description of the embodiments is to enable those of ordinary skill in the art to understand and use the invention. It is obvious that those skilled in the art can easily make various modifications to these embodiments and apply the general principles described herein to other embodiments without creative efforts. Therefore, the present invention is not limited to the above embodiments, and the improvements and modifications made by those skilled in the art without departing from the scope of the present invention according to the disclosure of the present invention should be within the protection scope of the present invention.

Claims

1. An intelligent question-answering method for equipment failure based on large model enhancement, characterized in that: The following steps are involved: Obtain steel production line data and build a fault knowledge graph, wherein the steel production line data includes equipment operation data, historical fault records, maintenance and repair records, operation logs, and production line environment data; Obtain user questions and classify them. The classification results include single-hop questions, short-hop questions, and multi-hop questions. According to the category of the user's question, the corresponding reasoning mechanism is called, the answer is generated according to the fault knowledge graph, and the equipment fault analysis in natural language form is output through the retrieval enhancement generation technology; for single-hop questions, the answer is directly retrieved from the fault knowledge graph based on the Aho-Corasick automaton matching core entities and relationships; for short-hop questions, the single-agent reasoning model is used to derive the answer from the fault knowledge graph; For multi-hop problems, a multi-agent collaborative reasoning model is used to generate a multi-step reasoning path based on the fault knowledge graph.

2. According to claim 1, a method for intelligent question-answering of equipment failure based on large model enhancement is characterized in that: The F-Qwen2-7B model is used to perform semantic analysis on user questions to obtain classification results. The specific process is as follows: The user questions are embedded through the F-Qwen2-7B model to obtain a high-dimensional vector representation of the user questions. The cosine similarity between the high-dimensional vector representation and the predefined question set in the steel production line field is calculated. If the calculated cosine similarity is higher than the set threshold, the next step is entered. Otherwise, a standardized prompt message is returned. The F-Qwen2-7B model is used to generate a set of questions with similar semantics to user questions to assist in analyzing the complexity of user questions and then obtain classification results.

3. According to claim 2, a method for intelligent question-answering of equipment failure based on large model enhancement is characterized in that: The QLoRA strategy was used to fine-tune the F-Qwen2-7B model.

4. According to the method of intelligent question-answering of equipment failure based on large model enhancement in claim 1, it is characterized in that: For single-hop problems, the Aho-Corasick automaton is first used to match core entities and relations, and then a predefined Cypher statement is selected to retrieve relevant triples from the fault knowledge graph as answers.

5. According to claim 1, a method for intelligent question-answering of equipment failure based on large model enhancement is characterized in that: The single-agent reasoning model includes an embedding module, a knowledge graph reasoning module and an answer decision module; The embedding module includes a question embedding submodule and a knowledge graph embedding submodule. The question embedding submodule uses the Chinese-RoBERTa-wwm-ext model to embed user questions into a low-dimensional vector space to generate a hidden representation of the user questions. The knowledge graph embedding submodule uses the ConvE model to learn the embedding of all entities and relationships in the fault knowledge graph through convolution operations. The knowledge graph reasoning module includes an intelligent agent and an external environment, wherein the external environment is modeled as a Markov decision process with a tuple (S, A, p, R a ), where S represents the state space, including all entities and relations in the fault knowledge graph, A represents the action space, corresponding to the relationship edge selected by the agent in the fault knowledge graph, p is the state transition probability, which is used to describe the probability of transitioning to the next state under the given state and selected action, and R a represents a reward function, which is used to feedback the decision quality of the intelligent agent; the knowledge graph reasoning module is driven by the Actor-Critic reinforcement learning method to generate multiple reasoning paths and corresponding candidate answers; The answer decision module is used to select the candidate answer with the highest score as output based on the hidden representation of the user's question through the attention mechanism and multi-path aggregation technology.

6. According to claim 5, a method for intelligent question-answering of equipment failure based on large model enhancement is characterized in that: The process of selecting the candidate answer with the highest score by the answer decision module is as follows: Calculate the user question q and the i-th reasoning path p i Attention score And the user question q and the i-th candidate answer a i Attention score Among them, ATT is the dot product attention operation, h q , They are user question q and the i-th reasoning path p respectively. i , the i-th candidate answer a i Hidden representation of ; Calculate the aggregated path representation h agg : Among them, w i is the representation of the i-th reasoning path The weight of connect h q and h agg , through the linear classifier to obtain for a given question q, candidate answer a i The probability of the correct answer is P(a i |q): Among them, W is the weight matrix of the linear classifier and b is the bias term.

7. According to claim 1, a method for intelligent question-answering of equipment failure based on large model enhancement is characterized in that: The multi-agent collaborative reasoning model includes a relationship selection agent, an entity selection agent and a fact selection agent. The relationship selection agent and the entity selection agent are used to collaboratively infer reasoning paths based on the fault knowledge graph, and the fact selection agent is used to expand the fault knowledge graph through dynamic knowledge completion technology.

8. The method for intelligent question-answering of equipment failure based on large model enhancement according to claim 7 is characterized in that: The fact selection agent is defined as a four-tuple (S FE ,A FE ,R FE ,π FE ), where S FE , A FE , R FE , π FE Respectively represent the state space, action space, reward, and strategy of the fact selection agent; When the relation selection agent reaches an entity e i When the fact selection agent extracts the entity e from the external corpus i Related fact triples F RS And merged into the fault knowledge graph, the state is defined as Among them, b(e i ) indicates that the current entity e i The set of relevant candidate facts, a tti represents the attention embedding, which is calculated by the following formula: Among them, α ij is the attention score calculated based on the graph attention network, which is used to measure the relevance of the current reasoning path and the candidate triple. ij The concatenation of the entity and relation vector embeddings for each triple; When the entity selection agent reaches the relation r i When the fact selection agent extracts the relevant information from the external corpus, i-1 ,r i ) related fact triples F ES And merged into the fault knowledge graph, the state at this time is defined as 9. The method for intelligent question-answering of equipment failure based on large model enhancement according to claim 7 is characterized in that: When training the multi-agent collaborative reasoning model, an alternating training and joint optimization strategy is adopted.

10. The method for intelligent question-answering of equipment failure based on large model enhancement according to claim 7 is characterized in that: The loss function of the multi-agent collaborative reasoning model is constructed based on the loss functions of the relationship selection agent, the entity selection agent and the fact selection agent, and an entropy regularization term is introduced into the loss function of each agent.

Citation Information

Cited By

  • Intelligent operation and maintenance question-answering system for cable manufacturing equipment

    CN120561253A

  • Obtaining method and device of troubleshooting knowledge graph, equipment and medium

    CN120725115A

  • Expressway toll auxiliary question and answer method and system based on knowledge graph

    CN120744141A

  • A highway tolling auxiliary question and answer method and system based on a knowledge graph

    CN120744141B

  • Event clue knowledge enhancement generation method and system based on deep learning

    CN120893554A