Mixed reality mine ventilation equipment maintenance guidance method

The Markov decision-making process constructed through multi-source sensors and knowledge graphs, combined with deep reinforcement learning, realizes intelligent maintenance guidance for mine ventilation equipment, solves the problems of relying on experience and lack of dynamic strategies in the existing technology, and improves the standardization and efficiency of maintenance.

CN120338560AInactive Publication Date: 2025-07-18NUOWENKE BLOWER FAN BEIJING

Patent Information

Application Number
CN202510822916.3
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-06-19
Publication Date
2025-07-18
Estimated Expiration
Not applicable · inactive patent

AI Technical Summary

Technical Problem

The existing mine ventilation equipment maintenance guidance methods rely on the experience of operation and maintenance personnel, lack standardization and optimization, cannot cope with complex failure scenarios, and lack dynamic adjustment strategies, making it difficult to achieve a closed-loop interaction of diagnosis-guidance-verification.

Method used

Real-time state data is collected through multi-source sensors, and the entity-relationship-attribute triplets are constructed based on the natural language processing of historical maintenance data and the knowledge graph, which are mapped to Markov decision-making process (MDP), and constructed agents through improved deep reinforcement learning to output optimal maintenance strategies.

Benefits of technology

It realizes dynamic simulation of equipment failures and automated guidance for optimal maintenance actions, adapts to equipment diversity and uncertainty, and improves the standardization and efficiency of maintenance.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120338560A_ABST
    Figure CN120338560A_ABST
Patent Text Reader

Abstract

The invention discloses a mixed reality mine ventilation equipment maintenance guidance method. The method comprises the following steps: S1, collecting real-time state data of mine ventilation equipment based on a multi-source sensor; s2, collecting historical maintenance data, performing natural language preprocessing on the historical maintenance data, constructing an entity-relation-attribute triple for the preprocessed data through a BERT model, and obtaining an optimized maintenance knowledge graph based on the entity-relation-attribute triple; s3, taking the maintenance process of the mine equipment as a Markov decision process, and taking the optimized maintenance knowledge graph as a state space to carry out Markov decision process modeling to obtain an MDP prediction model; and S4, based on the MDP prediction model, constructing intelligent agent learning through improved deep reinforcement learning, selecting an optimal maintenance action strategy for different equipment states under the MDP prediction model, and outputting an optimal maintenance guidance strategy.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of mine intelligent prediction, and particularly to a method for guiding the maintenance of mine ventilation equipment using mixed reality. Background Art

[0002] With the transformation of the operation and maintenance requirements of mine ventilation equipment from "after-sales maintenance" to "predictive maintenance", the maturity of mixed reality (MR) technology provides a new path for maintenance guidance. Through the interactive method of virtual-real fusion, it can intuitively superimpose maintenance knowledge on the actual scene of the equipment, improving the on-site operation efficiency. At the same time, the application of deep learning and knowledge graph technology in the industrial field is gradually deepening, providing theoretical support for the structured modeling and intelligent decision-making of maintenance data. How to transform multi-source heterogeneous data (such as sensor real-time data, historical maintenance texts) into a computable decision-making model has become the key technical direction for realizing the intelligent maintenance of mine equipment.

[0003] Currently, the methods for guiding the maintenance of mine ventilation equipment usually rely on the experience of operation and maintenance personnel, with subjective judgments, making it difficult to ensure the standardization and optimization of maintenance effects. Moreover, the existing mine ventilation equipment maintenance guidance systems are mostly static process guides (such as fixed-step animation demonstrations), lacking the ability to dynamically adjust strategies according to the real-time state of the equipment, unable to handle complex fault scenarios (such as multi-component collaborative failures), and not deeply binding the maintenance guidance information (such as operation steps, parameter thresholds) in the scenario with the underlying decision-making model, making it difficult to achieve a closed-loop interaction of "diagnosis - guidance - verification". Therefore, a method for guiding the maintenance of mine ventilation equipment using mixed reality is proposed herein. Summary of the Invention

[0004] In order to overcome the above-mentioned defects of the prior art and achieve the above object, the present invention proposes the following technical solutions: A method for guiding the maintenance of mine ventilation equipment using mixed reality, comprising: S1: Collecting real-time state data of mine ventilation equipment based on multi-source sensors; S2: Collecting historical maintenance data, performing natural language preprocessing on the historical maintenance data, and constructing entity-relationship-attribute triples for the preprocessed data through a BERT model, and obtaining an optimized maintenance knowledge graph based on the entity-relationship-attribute triples; S3: Regarding the maintenance process of mine equipment as a Markov decision process, and modeling the Markov decision process with the optimized maintenance knowledge graph as the state space to obtain an MDP prediction model; S4: Based on the MDP prediction model, constructing an intelligent agent through improved deep reinforcement learning to learn the strategy of selecting the optimal maintenance action for different equipment states under the MDP prediction model, and outputting the optimal maintenance guidance strategy; The improved deep reinforcement learning adopts the Twin Delayed Deep Deterministic Policy Gradient algorithm, including two critic networks Q1 and Q2 to fit the state-action value, and one actor network to generate actions.

[0005] The multi-source sensors include vibration sensors, temperature sensors, and wind pressure sensors.

[0006] The BERT model is based on the Transformer encoder architecture, and through the multi-head attention mechanism, it performs context semantic encoding on each word in the preprocessed historical maintenance data to generate the context semantic representation of each word.

[0007] The process of constructing entity-relationship-attribute triples through the BERT model is as follows: Taking the preprocessed historical maintenance data as the input of the BERT model, entity recognition is performed based on the context semantic representation. The entities include equipment entities, fault entities, and operation entities. Determine the start position and end position of the entity in the text. For two entities, based on the start position and end position of the two entities, perform mean pooling operation on the semantic representation of the BERT model. Concatenate and input the pooled representations of the two entities into the fully connected layer of the BERT model, calculate the probability of the existence of a relationship between the entities through the softmax function, and form a triple set with the entities themselves based on the output probability of the existence of a relationship between the entities.

[0008] The optimized maintenance knowledge graph is obtained by storing the triples according to the storage structure of the knowledge graph.

[0009] The process of taking the optimized maintenance knowledge graph as the state space is as follows: The states in the state space are obtained based on the subgraph representations in the optimized maintenance knowledge graph, and each state corresponds to a subgraph in the optimized maintenance knowledge graph.

[0010] The process of obtaining the MDP prediction model is as follows: Map the graph structure of the knowledge graph to the quadruple of the MDP, including the state space S, the action space A, the state transition probability , and the reward function R; The actions in the action space A are maintenance operations, corresponding to the maintenance action nodes in the optimized maintenance knowledge graph. The state transition probability represents the probability of transitioning to state after performing action a in state s, which is calculated through historical maintenance data statistics or knowledge graph relationship weight calculation, and the formula is expressed as: ; Among them, is the total number of times the action a is executed in the state s in history, is the number of times of transferring to the state ; represents the new state transferred from the current state s after executing the action a; The reward function R defines the reward value according to the action execution result: ; where, >0 is the positive reward, <0 is the negative reward, is the intermediate state reward; The final MDP prediction model is expressed as .

[0011] The process of constructing an agent by improving deep reinforcement learning is as follows: Randomly initialize the parameters of two critic networks and the parameters of one actor network. The network structure adopts a multi-layer perceptron, the activation function of the hidden layer is the Swish function, and configure two target critic networks and one target actor network, and the initialization of the parameters of the two target critic networks and the parameters of one target actor network is the same as the current network; Set the key hyperparameters, the discount factor , the exploration noise , the delayed update step d = 2, the experience replay buffer, and initialize the capacity to be The experience replay buffer D, and the storage format is where d is the termination flag.

[0012] The process of the policy for selecting the optimal maintenance action is as follows: For the state subgraph mapped by the knowledge graph, extract the node features and edge features, and encode them into a vector through the graph attention network , where t is the dimension of the state vector; Output the deterministic action through the actor network, and add the exploration noise to obtain the actual executed action , execute the actual executed action After that, the MDP model transfers to the new state according to the transition probability, and directly obtains the immediate reward through the reward function ; Put into the buffer D, where indicates that the fault is solved, otherwise , is the next action to be executed, randomly sample a batch of experiences from the buffer D, and for the current state and the next state Perform normalization processing, update the target Q value through the Critic network, update the actor network by maximizing the output of the critic network, synchronize the target network parameters every time the actor network is updated, and when the loss of the critic network is continuously stable below 0.1 for 1000 steps and the update amplitude of the actor network policy is less than , it is determined that the policy converges, and the deterministic policy output by the actor network is the optimal maintenance policy.

[0013] The present invention has the following beneficial effects: In the present invention, first, the device status data is collected in real time through multi-source sensors such as vibration, temperature, and wind pressure. Combining natural language processing and knowledge graph construction of historical maintenance data, the device faults are abstracted into entity-relationship-attribute triples, forming a complete knowledge network including device components, fault phenomena, and maintenance actions, solving the problems of traditional manual diagnosis relying on experience and lacking data support; Secondly, the knowledge graph is mapped to the state space of the Markov decision process (MDP), and the uncertainty of the maintenance process is modeled through state transition probabilities and reward functions to realize the dynamic simulation of "fault state - maintenance action - result feedback"; Finally, the agent is trained through the twin-delayed deep deterministic policy gradient algorithm, and the optimal maintenance actions for different device states are output through the actor network (such as "replace the bearing + adjust the viscosity of the grease"), replacing traditional manual experience decision-making, and encouraging the agent to explore potential better strategies by adding Gaussian noise and action clipping, avoiding falling into local optima and adapting to the diversity and uncertainty of mine equipment faults. BRIEF DESCRIPTION OF THE DRAWINGS

[0014] Figure 1 It is a method step diagram of a mixed reality mine ventilation equipment maintenance guidance method proposed by the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0015] The technical solutions in the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all of the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.

[0016] Embodiment: As Figure 1 shown, a mixed reality mine ventilation equipment maintenance guidance method proposed by the present invention includes: S1: Collect real-time status data of mine ventilation equipment based on multi-source sensors; Deploy multi-source sensors at key parts of the mine ventilation equipment, including vibration sensors, temperature sensors, and wind pressure sensors, to collect real-time status data of the equipment in real time; Specifically, the vibration sensor, temperature sensor, and wind pressure sensor all sample the equipment vibration at a fixed sampling frequency, convert the continuous time-domain signal into discrete time-series data, and the data collected by the sensor is transmitted to the edge computing node through wireless communication. During the transmission process, the CRC (Cyclic Redundancy Check) algorithm is used to check the data, generate a check code and append it to the end of the data frame. The receiving end judges whether the data has an error during the transmission process by recalculating the check code and comparing it with the received check code. If an error occurs, the error data is deleted; Finally, the data collected by the multi-source sensors is combined into real-time status data, expressed as .

[0017] S2: Collect historical maintenance data, perform natural language preprocessing on the historical maintenance data, and construct entity-relationship-attribute triples through the BERT model for the preprocessed data, and obtain an optimized maintenance knowledge graph based on the entity-relationship-attribute triples; Natural language preprocessing process: Based on the equipment manual, historical maintenance work orders, and text data, remove noise information such as special characters, HTML tags, and extra spaces in the text. Let the original text be D, and process the original text D into pure text ; Perform word segmentation on the pure text and use the jieba word segmentation tool to split the sentence into a pure word sequence; For example: For the sentence "The temperature of the fan bearing is too high, resulting in equipment shutdown", after word segmentation, we get ["fan", "bearing", "temperature", "too high", "resulting in", "equipment", "shutdown"]; Based on the pure word sequence, perform named entity recognition to identify entities in the text, including equipment component names (fan, bearing), fault phenomena (temperature too high), etc.; After completing the text preprocessing, the preprocessed historical maintenance data is obtained ; The process of constructing entity-relationship-attribute triples based on the BERT model is as follows: For the processed historical maintenance data , directly extract entities and attributes through regular expression matching rules; Specifically, the preprocessed historical maintenance data It contains structured text and unstructured text. For the structured text, the fan model is FJ-100 and the rated power is 100kW. By using regular expressions to match, the entity fan can be extracted, the corresponding value of the attribute model is FJ-100, and the corresponding value of the attribute rated power is 100kW. For unstructured text (such as historical maintenance records, equipment technical documents, etc.), regular expressions are used to clean special symbols in the text (such as special codes in the mine environment, irrelevant tags, etc.); The preprocessed historical maintenance data is used as the input of the BERT model; The BERT model is based on the Transformer encoder architecture and encodes the context semantics of each word in through the multi-head attention mechanism. During the encoding process, the BERT model considers all the content before and after the word and generates the context semantic representation h of each word; For example, for the sentence "Replace the bearing to solve the problem of overheating temperature", the h of words such as "replace the bearing" and "temperature" will incorporate each other's semantic associations, and the semantic representation of "replace" will combine information such as "bearing" (operation object) and "temperature" (problem association); Specifically, traditional word embeddings (such as Word2Vec) can only obtain the static semantics of words, while the context semantic encoding of the BERT model can dynamically capture the true meaning of words in different text contexts, providing a high-precision semantic basis for accurately identifying entity relationships; Based on the context semantic representation h, entity recognition is performed. Entity j includes equipment entities (such as fan bearings), fault entities (such as overheating temperature), and operation entities (such as replacement), and the starting position u and ending position of the entity in the text are determined ; For two entities and , based on the starting positions and ending positions of the two entities and , the mean pooling operation is performed on the semantic representation h of the BERT model; , , for the BERT model's semantic representation h; Let the semantic representation covered by entity be , and its pooled representation , where i is the index; For , the mean pooling operation is performed to obtain ; The pooled representations of the two entities , are concatenated and input into the fully connected layer (classifier) of the BERT model, and the probability of the existence of relationship r between the entities is calculated through the softmax function: ; Among them, W is the classifier weight matrix, and b is the bias vector, which is used to fit the mapping from the entity representation to the relationship category; Specifically, the pooling operation can compress the semantic representations of multiple words covered by the entity into a single vector, extract the overall semantics of the entity, and facilitate subsequent calculations. The fully connected layer composed of the fully connected layer and softmax can map the entity semantic representation to a predefined relationship category (such as maintenance operation - for - faulty component - belongs to - equipment), realize the quantification of relationship probability, and provide a basis for judging the relationship between entities; Based on the probability r of the relationship existing between the output entities, a triple set is formed with the entities themselves ; These triples are stored according to the storage structure of the knowledge graph (based on the graph database Neo4j) to construct an optimized maintenance knowledge graph; First, node creation is performed. Traverse all triples, extract entities and remove duplicates, and create nodes with corresponding labels for each entity. For example, extract the entities in the triples (such as equipment components, fault phenomena, maintenance actions), remove duplicates, and create labeled nodes. The labels of node V include equipment, component, fault, and maintenance action, and each node contains an attribute set (such as model number, temperature value); Secondly, relationship creation is performed. According to the relationships in the triples, relationship edges are created between the corresponding nodes. For example, according to the relationship type in the triples, directed edges are created between nodes. The relationship types include (equipment - component), (cause - fault), (action - fault), and each edge contains an attribute set (such as occurrence time, validity); Finally, attribute filling is performed. The attribute - value information in the triples is supplemented to the corresponding nodes or relationships. For example, node attributes and relationship attributes are stored in the form of key - value pairs. For example, the attributes of the equipment node , relationship attributes ; The optimized maintenance knowledge graph is represented as G=(V, E, P). The entire graph is a directed graph composed of nodes, relationships, and attributes, where: V represents the set of nodes. Each node represents an entity (equipment, component, fault, maintenance operation), and the type is distinguished by labels. The label is the definition of the semantic category of the node, which is convenient for classification query and management; E is the set of directed relationship edges. Each edge connects two nodes and represents the semantic relationship between entities (such as the inclusion relationship between equipment and components, and the causal relationship between faults and causes); P is a set of attributes that attach to nodes or relationships and are used to describe the characteristics of entities or relationships (such as the model number and production date of a device node, and the creation time of a relationship edge). Specifically, constructing a knowledge graph based on triples is to transform the scattered knowledge extraction results (triples) into a structured and extensible property graph. In the graph of the knowledge graph, entities are classified and managed by node labels, and edges are characterized by relationship types, directions, and attributes to accurately depict the semantic associations between entities, ultimately forming a complete knowledge network of equipment-component-fault-maintenance operations.

[0018] S3: Take the maintenance process of mine equipment as a Markov decision process, and use the optimized maintenance knowledge graph as the state space to model the Markov decision process to obtain an MDP prediction model; Map the graph structure of the knowledge graph to the quadruple of MDP, including the state space S, the action space A, the state transition probability , and the reward function R; Construct the state space S. The states in the state space are obtained based on the subgraph representation in the optimized maintenance knowledge graph, that is, the combination of nodes and edges, which describe the current fault state or maintenance progress of the equipment. Each state s corresponds to a subgraph in the optimized maintenance knowledge graph , where are state-related entities, are state-related relationships; Specifically, the state space S is an abstract description of the current state of the equipment by MDP. In the mine maintenance scenario, using the subgraph of the knowledge graph to represent the state is essentially using a semantic network of "entity + relationship" to depict information such as the faults, components, and maintenance progress of the equipment; Construct the action space A. The action a is a maintenance operation (such as "inspect", "replace", "supplement"), which corresponds to the maintenance action node in the optimized maintenance knowledge graph , and the action space is represented as ; Specifically, the action space A is a set of operations executable by the MDP model, covering all maintenance behaviors that may change the state of the equipment. In the mine maintenance scenario, the action directly maps to the maintenance action node in the knowledge graph to ensure that the operation is consistent with the knowledge system; Obtain the state transition probability , that is, the probability of transferring from state s to state after executing action a, which is calculated by statistical analysis of historical maintenance data or the relationship weights of the knowledge graph: ; where is the total number of times action a is executed in state s in history, is the number of times of transferring to state , Denotes the new state transferred from the current state s after executing action a; Specifically, the state transition probability Describes the probability that the agent in the MDP model transfers to state after executing action a in state s, which reflects the true evolution law of equipment failures. In the mine maintenance scenario, probability calculation can be achieved by combining historical data statistics and knowledge graph reasoning; Construction of the reward function : Define the reward value according to the action execution result:

[0019] Among them, >0 is a positive reward, such as fault resolution, <0 is a negative reward, such as ineffective operation, is the intermediate state reward, such as partial repair; Specifically, the reward function R is a quantitative evaluation of the quality of actions in the MDP model, guiding the agent to learn the optimal strategy (maximizing long-term rewards). In the mine maintenance scenario, the reward reflects the balance between fault resolution progress and maintenance costs; The final MDP prediction model is expressed as ; Specifically, the construction of the quadruple in the MDP model essentially transforms the semantic associations of the knowledge graph into a mathematical decision model. S is used to describe the current fault state in the knowledge graph subgraph to ensure the semantic integrity of the state. A is used to reuse the maintenance action nodes of the knowledge graph to ensure the feasibility of operations, is used to fuse historical data and graph reasoning to ensure the authenticity of the transition probability. R is based on the fault criteria of the graph to ensure the objectivity of the reward. By S3, the mine equipment maintenance knowledge is transformed into a computable and optimizable decision model, providing a basis for subsequent deep reinforcement learning.

[0020] S4: Based on the MDP prediction model, construct an agent to learn the strategy of selecting the optimal maintenance action for different equipment states under the MDP prediction model through improved deep reinforcement learning, and output the optimal maintenance guidance strategy; Based on the MDP prediction model output by S3 , construct an improved deep reinforcement learning agent. The agent uses the MDP model as the interaction environment. Its goal is to learn the strategy of selecting the optimal action a for different states S in the state space S according to the state transition probability P and the reward function R; Using the Twin Delayed Deep Deterministic Policy Gradient (TD3) algorithm to solve the problem of overestimation in traditional deep Q algorithms, including two critic networks Q1 and Q2 to fit the state-action value, and one actor network Generate actions; First, perform the initialization configuration of the agent, and randomly initialize the parameters of the two critic networks 、 、and the parameters of one actor network The network structure uses a multi-layer perceptron, the activation function of the hidden layer is the Swish function, and two target critic networks and one target actor network are configured. Moreover, the parameters of the two target critic networks 、 、and the parameters of one target actor network Are initialized to be the same as the current network; Then set the key hyperparameters, the discount factor , the exploration noise (add random noise to the actions generated by the actor), the delay update step d = 2 (update the actor network once every d updates of the critic network), and the experience replay buffer; Initialize the experience replay buffer D with a capacity of , and the storage format is (s, a, r, , d) (including the termination flag d, used to distinguish whether the selected action is for fault resolution or not); After the construction of the agent based on the improved deep reinforcement learning is completed, the construction of the agent is complete. The agent interacts in a loop according to the following process in the mine maintenance scenario simulated by the MDP: For the state subgraph mapped by the knowledge graph , extract the node features (such as equipment model, fault type) and edge features (such as relationship type, confidence), and encode them into a vector through the graph attention network , where t is the dimension of the state vector: Output the deterministic action through the actor network, and add exploration noise to obtain the actual executed action:

[0021] Among them, is the action clipping function, which limits the network output deterministic action within the valid interval , are the minimum and maximum valid values of the action (the boundaries of the action space), is the Gaussian noise with a mean of 0 and a standard deviation of ; Execute the actual execution action After that, the MDP model transfers to a new state according to the transition probability and directly obtains the immediate reward through the reward function ; Store in buffer D, where represents the fault resolution (terminal state), otherwise , is the next execution action; Finally, randomly sample a batch of experiences from buffer D (B = 256, i is the index), and normalize the current state and the next state . Update the target Q value through the Critic network, update the actor network by maximizing the output of the critic network. Every time the actor network is updated, synchronize the target network parameters. When the loss of the critic network stabilizes below 0.1 for 1000 consecutive steps and the update amplitude of the actor network policy is less than , it is determined that the policy converges, and the deterministic policy output by the actor network is the optimal maintenance policy: , where is the optimal policy in the Markov decision process; Example: Device status and real-time data mapping of the knowledge graph: The sensor collects the bearing temperature of fan #1 as 85°C (normal range 20 - 40°C) and the vibration amplitude as 0.8 mm / s (normal < 0.3 mm / s); Knowledge graph subgraph: s = (nodes: {fan #1 bearing, overheating temperature, excessive vibration}, edges: {(bearing, appears, overheating temperature), (bearing, appears, excessive vibration)}; Encoded as a vector s = [0.85, 0.92, 0.11,...] through the Graph Attention Network (GAT); Actor network inference and policy output: The input of the actor network is the state vector s = [0.85, 0.92, 0.11,...]; Deterministic action generation: The trained actor network calculates: = "Replace the bearing + adjust the grease viscosity to 320 cSt (in the real scenario, the action will be encoded as a numerical vector and then mapped to an executable instruction); Optimal policy representation: This action is the optimal strategy in the current state s, that is: = Replace the bearing + adjust the grease viscosity to 320 cSt.

[0022] In the application, several formulas involved are calculated by taking their numerical values after dimensionlessization. The establishment of the formulas is obtained by collecting a large amount of data for software simulation to get a formula closest to the real situation. Some coefficients or weights in the formulas are set by those skilled in the art according to the actual situation, so no more details will be given here.

[0023] The above embodiments can be implemented in whole or in part by software, hardware, firmware or any other combination. When implemented using software, the above embodiments can be implemented in whole or in part in the form of a computer program product. Those of ordinary skill in the art can realize that the units and algorithm steps of each example described in combination with the embodiments disclosed herein can be implemented by electronic hardware or a combination of computer software and electronic hardware. Whether these functions are executed in a hardware or software manner depends on the specific application and design constraints of the technical solution.

[0024] Although the embodiments of the present invention have been shown and described, for those of ordinary skill in the art, it can be understood that various changes, modifications, substitutions and variations can be made to these embodiments without departing from the principles and spirit of the present invention. The scope of the present invention is defined by the appended claims and their equivalents.

Claims

1. A method for guiding the maintenance of mixed reality mine ventilation equipment, characterized in that, Including: S1: Collect the real-time status data of mine ventilation equipment based on multi-source sensors; S2: Collect historical maintenance data, perform natural language preprocessing on the historical maintenance data, and construct entity-relationship-attribute triples for the preprocessed data through the BERT model, and obtain an optimized maintenance knowledge graph based on the entity-relationship-attribute triples; S3: Take the maintenance process of mine equipment as a Markov decision process, and model the Markov decision process with the optimized maintenance knowledge graph as the state space to obtain an MDP prediction model; S4: Based on the MDP prediction model, construct an agent through improved deep reinforcement learning to learn the strategy of selecting the optimal maintenance action for different equipment states under the MDP prediction model, and output the optimal maintenance guidance strategy; The improved deep reinforcement learning adopts the Twin Delayed Deep Deterministic Policy Gradient algorithm, which includes two critic networks Q1 and Q2 to fit the state-action value, and one actor network to generate actions.

2. The method for guiding the maintenance of a mixed reality mine ventilation device according to claim 1, characterized in that, The multi-source sensors include vibration sensors, temperature sensors, and air pressure sensors.

3. A method for guiding the maintenance of a mixed reality mine ventilation device according to claim 2, characterized in that, The BERT model is based on the Transformer encoder architecture, and encodes the context semantics of each word in the preprocessed historical maintenance data through the multi-head attention mechanism to generate the context semantic representation of each word.

4. A method for guiding the maintenance of a mixed reality mine ventilation device according to claim 1, characterized in that, The process of constructing entity-relationship-attribute triples through the BERT model is as follows: Take the preprocessed historical maintenance data as the input of the BERT model, and perform entity recognition based on the context semantic representation. The entities include equipment entities, fault entities, and operation entities; Determine the start position and end position of the entity in the text. For two entities, based on the start position and end position of the two entities, perform a mean pooling operation on the semantic representation of the BERT model; Concatenate and input the pooled representations of the two entities into the fully connected layer of the BERT model, calculate the probability of the relationship between the entities through the softmax function, and form a triple set with the entities themselves based on the output probability of the relationship between the entities.

5. A method for guiding the maintenance of a mixed reality mine ventilation device according to claim 4, characterized in that, The optimized maintenance knowledge graph is obtained by storing the triples according to the storage structure of the knowledge graph.

6. A method for guiding the maintenance of a mixed reality mine ventilation device according to claim 1, characterized in that, The process of taking the optimized maintenance knowledge graph as the state space is as follows: The states in the state space are obtained based on the subgraph representation in the optimized maintenance knowledge graph, and each state corresponds to a subgraph in the optimized maintenance knowledge graph.

7. A method for guiding the maintenance of a mixed reality mine ventilation device according to claim 6, characterized in that, The process of obtaining the MDP prediction model is as follows: Map the graph structure of the knowledge graph to the quadruple of the MDP, including the state space S, the action space A, the state transition probability , and the reward function R; The actions in the action space A are maintenance operations, corresponding to the maintenance action nodes in the optimized maintenance knowledge graph. The state transition probability represents the probability of transitioning to state after executing action a in state s, which is calculated through historical maintenance data statistics or knowledge graph relationship weights, and is expressed by the formula: ; Among them, is the total number of times action a is executed in state s in history, is the number of times transferred to state and represents the new state transferred from the current state s after executing action a; The reward function R defines the reward value according to the action execution result: ; Among them, >0 is a positive reward, <0 is a negative reward, is the intermediate state reward; The final MDP prediction model is expressed as .

8. A method for guiding the maintenance of a hybrid reality mine ventilation device according to claim 1, characterized in that, The process of constructing an agent through improved deep reinforcement learning is as follows: Randomly initialize the parameters of the two critic networks and the parameters of one actor network. The network structure adopts a multi-layer perceptron, and the activation function of the hidden layer is the Swish function. Configure two target critic networks and one target actor network, and initialize the parameters of the two target critic networks and the parameters of one target actor network to be the same as the current network; Set key hyperparameters, the discount factor , exploration noise , the number of delayed update steps d = 2, the experience replay buffer, the initial capacity is the experience replay buffer D, the storage format is where d is the termination flag.

9. A method for guiding the maintenance of a mixed reality mine ventilation device according to claim 8, characterized in that, The process of selecting the strategy of the optimal maintenance action is as follows: Extract node features and edge features from the state subgraph of the knowledge graph mapping, and encode them into vectors through a graph attention network , where t is the dimension of the state vector; Output a deterministic action through the actor network, and add exploration noise to obtain the actually executed action , execute the actually executed action After that, the MDP model transfers to a new state according to the transition probability, and directly obtains the immediate reward through the reward function ; Store in buffer D, where indicates that the fault is resolved, otherwise , is the next execution action. Randomly sample a batch of experiences from buffer D, and normalize the current state and the next state . Update the target Q-value through the Critic network, update the actor network by maximizing the output of the critic network. Every time the actor network is updated once, synchronize the target network parameters. When the loss of the critic network is continuously stable below 0.1 for 1000 steps, and the update amplitude of the actor network policy is less than , it is determined that the policy converges, and the deterministic policy output by the actor network is the optimal maintenance policy.

Citation Information

Patent Citations

  • A data mining-based multi-critic reinforcement learning electric power economic dispatching method

    CN112381359A

  • Reinforcement learning-based power grid regulation and control strategy optimization method

    CN113988508A

  • Article recommendation method based on generative adversarial network model and deep reinforcement learning, electronic equipment and medium

    CN114202061A

  • Steel production line equipment diagnosis method based on self-learning entity relationship joint extraction

    CN114756687A

  • Substation equipment AR auxiliary maintenance method based on three-dimensional knowledge graph and terminal

    CN115510253A

Cited By

  • Intelligent granary ventilation and energy consumption optimization decision-making method based on reinforcement learning

    CN120725247A

  • Government affair problem processing method and device based on large model driving

    CN121052388A