Mine rescue training and danger dynamic simulation method based on virtual reality

By combining virtual reality modeling and deep Q-network with rule-based graph reasoning, the dynamic modeling and decision-making problems of the mine rescue training system in complex environments were solved, a highly flexible and intelligent training system was realized, and emergency response capabilities and task execution efficiency were improved.

CN120633414AInactive Publication Date: 2025-09-12BEIJING SLINTE TECH CO LTD
View PDF 0 Cites 4 Cited by

Patent Information

Application Number
CN202510746301.7
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-06-05
Publication Date
2025-09-12
Estimated Expiration
Not applicable · inactive patent

AI Technical Summary

Technical Problem

The existing mine rescue training system has limitations in reproducing diverse scenarios, improving emergency response capabilities, and experiencing dangerous environments. It lacks the ability to deeply model the evolution process of rescue missions, behavioral decision-making paths, and dynamic risk control. The rule system has poor adaptability and scalability, the task recommendation logic is fixed, and there is a lack of multi-source information fusion and strategy fusion judgment.

Method used

Using virtual reality modeling technology, deep Q network and rule graph reasoning mechanism, we construct task state diagram, intelligent task jump decision, rule graph structure self-evolution and scoring fusion mechanism. By collecting training personnel's operating behavior, environmental parameters and physiological responses, combined with deep Q network and rule graph, task node evaluation and recommendation are carried out to achieve dynamic training control.

Benefits of technology

It has improved the environmental restoration, task decision-making intelligence and training process adaptability of mine rescue training, enhanced emergency response capabilities and operational execution capabilities, and achieved real-time optimization and efficient execution of task processes.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120633414A_ABST
    Figure CN120633414A_ABST
Patent Text Reader

Abstract

The invention discloses a mine rescue training and danger dynamic simulation method based on virtual reality, and the method comprises the following steps: S1, modeling a rescue process into a directed graph structure, and outputting a task state diagram bound with a virtual reality scene; s2, collecting an operation behavior, a task state, an environmental parameter and a physiological response, and constructing a training state vector; s3, inputting the state vector into the deep Q network to calculate a jump node evaluation value, and outputting a jump node; s4, modeling the control rule as a rule node, constructing a directed acyclic graph, binding a weight, and outputting a rule graph; s5, performing self-evolution on the rule atlas according to training feedback, and outputting an evolution structure; s6, reasoning in the evolution rule map, fusing the rule recommendation node and the jump node, and outputting a final task node; and S7, executing jump control according to the final node, loading a virtual scene, and pushing environment parameters, task contents and risk information. According to the invention, intelligent path decision and training process adaptive optimization of the mine rescue task are realized.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of virtual reality and intelligent decision-making technology, and in particular to a mine rescue training and dangerous dynamic simulation method based on virtual reality. Background Art

[0002] Traditional mine rescue training relies heavily on a combination of paper-based materials, static diagrams, and experience-based training, often combined with field exercises. While this approach can provide some operational guidance and on-site awareness, it has significant limitations in recreating diverse scenarios, enhancing emergency response capabilities, and providing exposure to hazardous environments. Field exercises struggle to fully replicate the complex environments and high-pressure psychological pressures of a real disaster. Furthermore, limited space, equipment, and safety considerations make some high-risk training tasks difficult to conduct, resulting in insufficient training coverage and task completion.

[0003] In recent years, the development of virtual reality (VR) technology has provided new insights into mine rescue training. VR-based training systems can create immersive three-dimensional mining scenarios, recreating a variety of complex working conditions and typical accident cases. This allows trainees to conduct interactive operational training within a simulated environment, thereby improving their emergency response capabilities and awareness of operational standards. Previous research has explored some mining training systems that utilize VR to recreate real-world mining environments and combine simple behavioral logic to simulate and rehearse training tasks. However, most of these systems remain at the "scenario construction + static task loading" level, lacking the ability to deeply model the evolution of rescue tasks, behavioral decision paths, and dynamic risk control.

[0004] In addition, some existing systems have introduced artificial intelligence models to identify operational behaviors or evaluate training performance. For example, some studies have used neural networks to identify whether operational actions during training are compliant, or to judge the quality of task completion based on scoring of training results. However, these models are often independent of the training task path itself and do not participate in the selection and recommendation of task nodes, making it difficult to support the dynamic planning and strategy evolution of the task path. Due to the lack of comprehensive modeling capabilities for the complex logical relationships between training process status, behavioral feedback, environmental factors, and task rules, existing systems generally have problems such as fixed task recommendation logic, weak path evolution capabilities, and low decision-making intelligence levels, making it difficult to adapt to the suddenness and uncertainty characteristics of the evolution of real accidents.

[0005] Existing technologies also have certain shortcomings in rule modeling. Some systems attempt to import rule libraries into training platforms to perform rule matching on behavioral outcomes during training. However, these rules are mostly expressed in linear processes or static conditions, lacking structured expression and graph reasoning capabilities, and are unable to model complex dependencies between tasks. More critically, current rule systems are mostly statically constructed and lack the ability to dynamically adjust their structure based on training feedback. This results in poor adaptability and scalability, and inability to achieve self-evolutionary optimization based on training history.

[0006] Furthermore, in terms of task recommendation decision-making mechanisms, there is still a lack of methods for integrating AI reasoning results with rule-based logic. Most systems still use predefined paths or single model results to make decisions, ignoring the need for multi-source information fusion in behavioral decision-making. For example, while some models can output ratings for multiple candidate nodes, they are unable to comprehensively consider factors such as task urgency, path risk, and strategy differences, making it difficult to make the optimal choice of jump node. This "model-rule separation" training task recommendation mechanism directly limits the VR training system's intelligent response capabilities and task adaptability in complex dynamic environments.

[0007] In terms of behavioral control, existing systems often employ scripted or state-machine control structures, meaning they jump directly to the next task node based on the task node currently completed by the trainee. This process lacks comprehensive assessment of factors such as behavioral strategies, node scores, and deviations from training objectives, and lacks mechanisms for strategy integration and dynamic score adjustment. While this control approach is highly efficient, it exhibits significant rigidity and inflexibility in the face of multi-factor driven task evolution.

[0008] Therefore, how to provide a virtual reality-based mine rescue training and dangerous dynamic simulation method is a problem that technical personnel in this field urgently need to solve. Summary of the Invention

[0009] One purpose of the present invention is to propose a mine rescue training and hazard dynamic simulation method based on virtual reality. The present invention makes full use of virtual reality modeling technology, deep Q network and rule graph reasoning mechanism, and describes in detail the dynamic task control process of fusion task state graph modeling, intelligent task jump decision, rule graph structure self-evolution and scoring fusion mechanism. It has the advantages of high environmental restoration, strong task decision-making intelligence and high adaptability of training process.

[0010] A method for mine rescue training and dangerous dynamic simulation based on virtual reality according to an embodiment of the present invention includes the following steps:

[0011] S1. Model the mine rescue training process as a directed graph structure and output a task state graph with node binding to the virtual reality scene;

[0012] S2. During the training process, the operator's operation behavior, current task status, virtual environment parameters, and physiological response indicators are collected to construct a training state vector;

[0013] S3. Input the training state vector into the deep Q network model, extract the behavior and environment characteristics, calculate the jump node evaluation value in the task state diagram, and output the mine rescue task jump node;

[0014] S4. Model the training control rules as rule nodes. Rule nodes are connected by logical edges to form a directed acyclic graph structure. Weight values ​​are bound to the edges to output the rule graph structure.

[0015] S5. Based on the task execution results and behavioral feedback recorded during the training process, the rule graph structure is self-evolved and the evolved rule graph structure is output;

[0016] S6. Based on the training state vector, perform reasoning operations in the evolved rule graph structure to obtain the rule graph recommended task node, fuse the rule graph recommended task node with the mine rescue task jump node, and output the final task node;

[0017] S7. Execute the jump control of the task state diagram according to the final task node, load the virtual reality training scene corresponding to the final task node, push the environmental parameters, training content and risk information, and complete the training process.

[0018] Optionally, the S1 specifically includes:

[0019] S11. Collecting mine training area structural data and historical mine accident data, wherein the mine training area structural data includes tunnel topology, key node locations, ventilation paths, and safe haven locations; and the historical mine accident data includes accident type, accident triggering conditions, accident location, rescue path, and failed mission node records, to construct a mine training basic dataset.

[0020] S12. Construct a mine scene topology map based on the mine training basic data set, wherein the mine scene topology map includes a set of scene nodes and a set of spatially connected paths between the scene nodes, and output the mine scene topology map;

[0021] S13. Define a set of mine rescue training task nodes, where each training task node is a specific rescue training task. The training tasks include a hazard source identification task, an alarm task, an evacuation judgment task, an on-site rescue task, and a personnel transfer task.

[0022] S14. Construct a task state graph based on the mine rescue training task node set. The task state graph is a directed graph structure, including a training task node set and a state transition path set between the nodes. Each state transition path is a directed connection relationship from one training task node to another training task node.

[0023] S15. Assign a unique task number, task objective description, task triggering condition, task required environment parameters and virtual reality scene mapping instructions to each training task node in the task status diagram, and output a task node attribute information set;

[0024] S16. Bind each training task node to the corresponding virtual reality three-dimensional scene area according to the task node attribute information set and the task status graph, load it when the virtual reality system is initialized, and output the task status graph bound to the virtual reality scene.

[0025] Optionally, the training state vector includes task node number, operation behavior code, environmental state parameters and physiological state parameters; the environmental state parameters include smoke concentration, temperature, light intensity, toxic gas index, and obstructions; the physiological state parameters include heart rate, body movement frequency, reaction delay, and voice intensity.

[0026] Optionally, the S3 specifically includes:

[0027] S31, input the training state vector as input to the input layer of the deep Q network model;

[0028] S32. Set multiple fully connected hidden layers in the deep Q network model, extract features from the training state vector, and generate intermediate representations of task behavior and environment features;

[0029] S33: Input the intermediate expression into the output layer and output the basic action value function of all jumpable actions in the current training state;

[0030] S34, according to action a in the task state diagram i The target node v j , construct a structured path value function:

[0031] Q * (X s ,a i )=Q(X s ,a i )·ω d (v j )-λ c ·C(a i );

[0032] Among them, Q * (X s,a i ) is the structured path value evaluation function, Q(X s ,a i ) is the basic action value function, X s is the training state vector, a i is a jump action starting from the current task node in the current task state diagram, ω d (v j ) is the target task node v j The task importance factor, v j For jump action a i The target node, λ c is the risk control weight coefficient, C(a i ) is jump action a i The corresponding path cost indicator;

[0033] S35. Select the action with the largest structured path value function from all actions as the optimal jump action a under the current training state. * ;

[0034] S36, according to the optimal jump action a * Corresponding jump relationship, determine the target task node v t+1 , the target task node v t+1 The target node is defined as the mine rescue mission jump node.

[0035] Optionally, the S4 specifically includes:

[0036] S41. Based on the characteristics of mine rescue training tasks, construct a rule node set N = {R1, R2, ..., R n}, each rule node R k Defined as a triplet R k =(C k ,A k ,O k ), where C k Indicates the rule triggering condition, A k Indicates the rule execution action, O k Indicates the expected result after the action is performed;

[0037] S42. According to the conditional dependency between the rule nodes, a directed edge set E between the rule nodes is constructed. R ={e ij ∣e ij =(R i ,R j )};

[0038] S43, based on the rule node set N and directed edge set E R, construct the regular graph structure G R ;

[0039] S44. Assign edge weights to each rule connection path in the rule graph structure:

[0040] W(e ij )=α·S c (R i ,R j )+β·H f (P ij )+γ·T r (R j );

[0041] Among them, W(e ij ) is from the rule node R i To rule node R j The edge weight, e ij represents a directed edge in the rule graph, α is the weight coefficient of the structural similarity score item, S c (R i ,R j ) is a rule node R i The output result of the rule node R j The structural similarity score between the trigger conditions, R i , R j is a rule node, i represents the number of the current rule node in the rule graph, and j represents the node with the same i The target rule node number with dependency relationship, β is the weight coefficient of historical reasoning validity scoring item, H f (P ij ) represents the rule path R i →R j The frequency score of being effectively triggered in the past training process, γ is the weight coefficient of the node task urgency score item, T r (R j ) is a rule node R j The urgency score of the corresponding task in the task status diagram;

[0042] S45. Based on the edge weight results of all rule paths, a rule graph edge weight matrix is ​​constructed, and a rule graph structure including a rule node set, a rule connection structure and an edge weight matrix is ​​output.

[0043] Optionally, the construction of the directed edge set between the rule nodes includes determining the connection logic between the rule nodes, comparing the output result of each rule node with the triggering condition of other rule nodes, and when the output result of one rule node meets the triggering condition of another rule node, establishing a directed connection path between the two rule nodes, and sequentially establishing all rule paths that meet the dependency relationship to form a directed edge set E between the rule nodes. R .

[0044] Optionally, the S5 specifically includes:

[0045] S51. Collect rule triggering records during the training process and construct a historical reasoning trajectory set H;

[0046] S52. Based on the historical reasoning trajectory set H, calculate the trigger frequency of the rule node, and filter the low-frequency redundant node set according to the node contribution, remove all nodes and associated edges belonging to the low-frequency redundant node set from the rule graph, and output the simplified rule node set N ' With edge set E ' R ;

[0047] S53. Update the edge weights of the rule path according to the training feedback score and the historical performance of the rule path to obtain an updated edge weight matrix;

[0048] S54, normalize the updated edge weight matrix to form the edge weight matrix after the regular graph evolution And with the simplified rule node set N ' With edge set E ' R Together they form the evolved regular graph structure G ' R .

[0049] Optionally, the S6 specifically includes:

[0050] S61, obtain the mine rescue task jump node output by the deep Q network model, and the evolved rule graph structure G' R The rule graph obtained by inference recommends task nodes v r ;

[0051] S62. Calculate the structured path value function of the deep Q network direction based on the training state vector:

[0052] F DQN (v d )=Q * (X s ,a d )·φ(v d );

[0053] Among them, F DQN (v d ) is the comprehensive score of the jump node in the deep Q network direction, v d Indicates the jump node of the mine rescue mission, Q * (X s ,a d ) is the structured path value evaluation function, X s is the training state vector, φ(v d ) is the mine rescue mission jump node v d Attribute weight function of

[0054] S63. Based on the rule graph path reasoning results, calculate the comprehensive scoring function of the rule graph recommendation node:

[0055]

[0056] Among them, F RULE (v r ) is the rule graph recommendation task node v r The comprehensive score value, v r is the recommended task node for the rule graph, L is the total number of hops from the current task node to the target node along the rule graph reasoning path, l is the hop index, k is the edge index within the path, and W ′ (e k ) is the weight of the kth edge after evolution, e k represents the kth directed edge on the path in the regular graph structure, δ l is the path length penalty factor, ψ(v k ) is the structural importance scoring function;

[0057] S64, calculating a confidence difference index between the deep Q network score and the rule score;

[0058] S65. Construct the final fusion scoring function:

[0059] F fused (v)=σ(λ d ·F DQN (v d )+λ r ·F RULE (v r )+θ·Δ conf (v d ,v r ));

[0060] Among them, F fused (v) is the fusion score function, σ(·) is the activation function, and λ dis the scoring weight coefficient of the deep Q network direction, λ r is the scoring weight coefficient of the rule graph direction, θ is the difference adjustment coefficient, Δ conf (v d ,v r ) is the confidence difference index;

[0061] S66. Select the node with the highest fusion score from the candidate nodes as the final task node in the task state diagram under the current training state.

[0062] Optionally, the calculation of the confidence difference index includes obtaining the comprehensive score value of the jump node in the direction of the deep Q network and the score value of the recommended node output by the rule graph structure, calculating the score difference between the two, and forming a node score difference term; further calculating the difference between the task attribute score of the jump node and the structural importance score of the rule node to form a node structure attribute deviation term; simultaneously obtaining the action strategy probability distribution output by the deep Q network and the normalized weight distribution of the jump action in the rule graph reasoning path, and calculating the KL divergence between the two to form a strategy distribution deviation term; the above three types of difference metrics are weighted and superimposed to construct a confidence difference index Δ conf (v d ,v r ).

[0063] Optionally, the S7 specifically includes:

[0064] S71. Call the final task node in the task state diagram as the task execution target of the current training phase, and retrieve the virtual reality scene binding information corresponding to the final task node;

[0065] S72: Load the virtual reality 3D training scene resources bound to the final task node, including the 3D space structure model, initial environment state parameters, and interactive control components, to complete the initialization of the training sub-scene;

[0066] S73. Based on the environmental parameter configuration of the final task node, push the scene perception data required for task execution, including gas concentration value, visibility level, obstruction location, heat source distribution, and audio and video warning device status;

[0067] S74: Push the training task content information configured in the final task node, including the task objective description, interactive guidance text, operation time limit, risk warning level, and corresponding operation feedback mechanism, to start the virtual reality task process;

[0068] S75. After the training task is completed, collect the behavioral path, response time, operation accuracy, task completion status and interaction record data during the training process to calculate the training feedback score of the current stage.

[0069] The beneficial effects of the present invention are:

[0070] First, unlike existing mine rescue training systems that rely heavily on static task process orchestration, manual rule execution, or single-model path recommendation, this paper proposes an integrated intelligent training method that integrates virtual reality modeling, deep reinforcement learning, and rule graph self-evolution reasoning. This method systematically constructs a closed-loop dynamic training and deduction system encompassing "task modeling—behavior collection—path evaluation—rule evolution—strategy fusion—scenario loading." In terms of task modeling, by modeling the mine rescue process as a directed graph structure and constructing a task state graph with node attribute descriptions, state transition logic, and VR scene binding instructions, this method achieves a precise mapping between virtual tasks and three-dimensional interactive spaces, effectively enhancing the flexibility of training path control and the realism of the immersive experience.

[0071] Secondly, in terms of intelligent decision-making mechanisms, this invention innovatively introduces a deep Q-network model. Using the training state vector as input, it comprehensively collects multi-source state information such as the trainee's operational behavior, physiological reactions, and environmental indicators, constructs an intermediate feature representation of behavior and environment, and dynamically scores the jumpable nodes in the task state graph using a structured path value function. This evaluates the action benefit, path cost, and node importance of each candidate node, enabling intelligent recommendation of the optimal task jump node. Unlike traditional fixed-path or expert rule-based path planning, this method enables adaptive jump selection within complex task structures, significantly improving the responsiveness and accuracy of task decision-making.

[0072] Again, in terms of logical rule modeling, the present invention constructs a structured rule graph, which represents the logical rule nodes in the training task process as triples and constructs a directed dependency edge structure between nodes to form a directed acyclic graph structure with reasoning capabilities. At the same time, multi-dimensional scoring factors (structural similarity, historical trigger frequency, task urgency) are introduced to construct an edge weight model to form an edge weight matrix that can be used for path deduction. This not only improves the logical expression ability of rule modeling, but also supports the structured absorption and rule evolution of training behavior feedback. With the help of rule trigger trajectory analysis and edge weight self-update mechanism, the system can gradually optimize the rule structure during the training process and realize self-learning and dynamic adaptation of the logical graph.

[0073] In addition, in terms of strategy fusion, the present invention proposes a fusion scoring mechanism. By constructing the scoring functions output by the deep Q network and the rule graph, combined with the scoring confidence difference measurement index, a fusion scoring function is constructed to uniformly evaluate the multi-source recommendation nodes. This mechanism not only improves the synergy between AI path recommendation and rule reasoning, but also enhances the robustness and interpretability of the node selection results. Finally, the system selects the jump target node based on the fusion scoring results, controls the task state diagram jump, completes the loading of the next stage training task and pushes the environment parameters, task content, and risk prompts, thereby achieving deep optimization of the continuity and responsiveness of the virtual training process.

[0074] Overall, this invention overcomes several limitations of existing technologies in dynamic task modeling, intelligent recommendation, and coupled logical rule deduction, establishing an intelligent mine rescue training system encompassing "state perception - strategy learning - rule evolution - path fusion - scenario response." Compared to traditional approaches characterized by rigid task processes, slow training response, and a disconnect between models and rules, this invention offers the combined advantages of highly dynamic paths, intelligent task recommendations, evolving rule logic, and visualized training execution. This effectively enhances mine workers' emergency response and operational execution capabilities in complex environments, offering significant engineering and safety management benefits. BRIEF DESCRIPTION OF THE DRAWINGS

[0075] The accompanying drawings are used to provide a further understanding of the present invention and constitute a part of the specification. Together with the embodiments of the present invention, they are used to explain the present invention and do not constitute a limitation of the present invention. In the accompanying drawings:

[0076] Figure 1 This is a flow chart of the virtual reality-based mine rescue training and dangerous dynamic simulation method proposed by the present invention;

[0077] Figure 2 This is a flowchart of the decision-making process of the mine rescue task jump node based on the deep Q network model in the virtual reality-based mine rescue training and dangerous dynamic simulation method proposed by the present invention;

[0078] Figure 3 This is a schematic diagram of the rule graph structure modeling and edge weight allocation mechanism in the virtual reality-based mine rescue training and dangerous dynamic simulation method proposed in the present invention. DETAILED DESCRIPTION

[0079] The present invention will now be described in further detail with reference to the accompanying drawings, which are simplified schematic diagrams that illustrate the basic structure of the present invention in a schematic manner.

[0080] refer to Figure 1-3,The method of mine rescue training and dangerous dynamic simulation based on virtual reality ,includes the following steps:

[0081] S1. Model the mine rescue training process as a directed graph structure and output a task state graph with node binding to the virtual reality scene;

[0082] S2. During the training process, the operator's operation behavior, current task status, virtual environment parameters, and physiological response indicators are collected to construct a training state vector;

[0083] S3. Input the training state vector into the deep Q network model, extract the behavior and environment characteristics, calculate the jump node evaluation value in the task state diagram, and output the mine rescue task jump node;

[0084] S4. Model the training control rules as rule nodes. Rule nodes are connected by logical edges to form a directed acyclic graph structure. Weight values ​​are bound to the edges to output the rule graph structure.

[0085] S5. Based on the task execution results and behavioral feedback recorded during the training process, the rule graph structure is self-evolved and the evolved rule graph structure is output;

[0086] S6. Based on the training state vector, perform reasoning operations in the evolved rule graph structure to obtain the rule graph recommended task node, fuse the rule graph recommended task node with the mine rescue task jump node, and output the final task node;

[0087] S7. Execute the jump control of the task state diagram according to the final task node, load the virtual reality training scene corresponding to the final task node, push the environmental parameters, training content and risk information, and complete the training process.

[0088] The present invention provides a mine rescue training and hazard dynamic simulation method based on virtual reality. By modeling the mine rescue process as a directed graph structure and combining deep Q-network with rule graph self-evolution reasoning, it breaks through the static task and fixed path limitations of traditional mine training. By collecting the operational behavior, physiological response and environmental parameters of the trainees, constructing a dynamic training state vector, and combining multi-dimensional data for intelligent recommendation of task nodes and path jump control, real-time optimization of the task process and improvement of emergency response capabilities are achieved. This method not only enhances the flexibility and intelligence level of mine rescue training, but also improves the accuracy of task decision-making and training efficiency. It has strong adaptability and scalability, can effectively improve the emergency response capabilities of mine rescue personnel in complex environments, and has important engineering promotion value and application prospects.

[0089] In this embodiment, S1 specifically includes:

[0090] S11. Collecting mine training area structural data and historical mine accident data, wherein the mine training area structural data includes tunnel topology, key node locations, ventilation paths, and safe haven locations; and the historical mine accident data includes accident type, accident triggering conditions, accident location, rescue path, and failed mission node records, to construct a mine training basic dataset.

[0091] S12. Construct a mine scene topology map based on the mine training basic data set, wherein the mine scene topology map includes a set of scene nodes and a set of spatially connected paths between the scene nodes, and output the mine scene topology map;

[0092] S13. Define a set of mine rescue training task nodes, where each training task node is a specific rescue training task. The training tasks include a hazard source identification task, an alarm task, an evacuation judgment task, an on-site rescue task, and a personnel transfer task.

[0093] S14. Construct a task state graph based on the mine rescue training task node set. The task state graph is a directed graph structure, including a training task node set and a state transition path set between the nodes. Each state transition path is a directed connection relationship from one training task node to another training task node.

[0094] S15. Assign a unique task number, task objective description, task triggering condition, task required environment parameters and virtual reality scene mapping instructions to each training task node in the task status diagram, and output a task node attribute information set;

[0095] S16. Bind each training task node to the corresponding virtual reality three-dimensional scene area according to the task node attribute information set and the task status graph, load it when the virtual reality system is initialized, and output the task status graph bound to the virtual reality scene.

[0096] The present invention constructs a basic data set for mine training by collecting mine training area structural data and historical accident data, and constructs a mine scene topology map based on this. Combined with virtual reality technology, it improves the practicality and emergency response capabilities of mine rescue training. By defining a set of task nodes and constructing a task status diagram, dynamic planning and intelligent decision-making of mine rescue tasks are realized. The attribute information of each training task node is accurately mapped to the virtual reality scene, ensuring the authenticity and efficiency of the training content. This method not only improves the task execution accuracy and environmental simulation realism during the training process, but also enhances the flexibility of the training path and the responsiveness of emergency tasks. It has high efficiency, scalability and a wide range of engineering application value.

[0097] In this embodiment, the training state vector includes a task node number, an operation behavior code, environmental state parameters, and physiological state parameters; the environmental state parameters include smoke concentration, temperature, light intensity, toxic gas index, and obstructions; the physiological state parameters include heart rate, body movement frequency, reaction delay, and voice intensity.

[0098] The present invention achieves comprehensive monitoring and intelligent analysis of mine rescue training by collecting training state vectors, including task node numbers, operational behavior codes, environmental state parameters, and physiological state parameters. The combination of environmental state parameters such as smoke concentration and temperature with physiological state parameters such as heart rate and reaction delay provides an accurate basis for task decision-making. This method not only improves real-time feedback and emergency response capabilities during training, but also enhances the intelligence and adaptability of task node selection, enabling more accurate assessment of trainees' operational performance and environmental risks. It has high real-time performance and accuracy, and has significant engineering application value.

[0099] In this embodiment, S3 specifically includes:

[0100] S31, input the training state vector as input to the input layer of the deep Q network model;

[0101] S32. Set multiple fully connected hidden layers in the deep Q network model, extract features from the training state vector, and generate intermediate representations of task behavior and environment features;

[0102] S33: Input the intermediate expression into the output layer and output the basic action value function of all jumpable actions in the current training state;

[0103] S34, according to action a in the task state diagram i The target node v j , construct a structured path value function:

[0104] Q * (X s ,a i )=Q(X s ,a i )·ω d (v j )-λ c ·C(a i );

[0105] Among them, Q * (X s ,a i ) is the structured path value evaluation function, Q(X s ,a i ) is the basic action value function, X s is the training state vector, ai is a jump action starting from the current task node in the current task state diagram, ω d (v j ) is the target task node v j The task importance factor, v j For jump action a i The target node, λ c is the risk control weight coefficient, C(a i ) is jump action a i The corresponding path cost indicator;

[0106] The structured path value evaluation function is used to evaluate the value of each candidate jump node in the task state graph. In the implementation of the present invention, the structured path value function is generated by calculating the basic action value of each jump action and combining it with the dynamic characteristics of the training state vector, such as the task node number, operation behavior, environmental parameters, and physiological state. This function not only considers the action benefits of the jump node, but also includes the path cost and node importance, thereby intelligently optimizing the jump path selection.

[0107] S35. Select the action with the largest structured path value function from all actions as the optimal jump action a under the current training state. * ;

[0108] S36, according to the optimal jump action a * Corresponding jump relationship, determine the target task node v t+1 , the target task node v t+1 The target node is defined as the mine rescue mission jump node.

[0109] The present invention uses a deep Q-network model to extract features from the training state vector, and combines it with a structured path value evaluation function to intelligently evaluate the jump path of the mine rescue task. This method converts the operational behavior and environmental characteristics of the task node into an intermediate expression, and further calculates the basic action value and path risk of each jump action, and finally selects the optimal jump action with the maximum structured path value. This mechanism significantly improves the accuracy and dynamic response capability of task path decision-making, not only enhances the intelligent decision-making capability during the training process, but also can adjust the task execution path in real time, optimizing training efficiency and effect. Compared with traditional static path planning methods, the present invention has stronger adaptability and scalability in emergency response, task adjustment and safety control, and has broad practical application value.

[0110] In this embodiment, the S4 specifically includes:

[0111] S41. Based on the characteristics of mine rescue training tasks, construct a rule node set N = {R1, R2, ..., Rn}, each rule node R k Defined as a triplet R k =(C k ,A k ,O k ), where C k Indicates the rule triggering condition, A k Indicates the rule execution action, O k Indicates the expected result after the action is performed;

[0112] S42. According to the conditional dependency between the rule nodes, a directed edge set E between the rule nodes is constructed. R ={e ij ∣e ij =(R i ,R j )};

[0113] S43, based on the rule node set N and directed edge set E R , construct the regular graph structure G R ;

[0114] S44. Assign edge weights to each rule connection path in the rule graph structure:

[0115] W(e ij )=α·S c (R i ,R j )+β·H f (P ij )+γ·T r (R j );

[0116] Among them, W(e ij ) is from the rule node R i To rule node R j The edge weight, e ij represents a directed edge in the rule graph, α is the weight coefficient of the structural similarity score item, S c (R i ,R j ) is a rule node R i The output result of the rule node R j The structural similarity score between the trigger conditions, R i , R j is a rule node, i represents the number of the current rule node in the rule graph, and j represents the node with the same i The target rule node number with dependency relationship, β is the weight coefficient of historical reasoning validity scoring item, H f (P ij ) represents the rule path Ri →R j The frequency score of being effectively triggered in the past training process, γ is the weight coefficient of the node task urgency score item, T r (R j ) is a rule node R j The urgency score of the corresponding task in the task status diagram;

[0117] S45. Based on the edge weight results of all rule paths, a rule graph edge weight matrix is ​​constructed, and a rule graph structure including a rule node set, a rule connection structure and an edge weight matrix is ​​output.

[0118] The present invention realizes dynamic decision-making and intelligent reasoning of mine rescue training tasks by constructing a rule graph structure based on a set of rule nodes. Each rule node is defined in the form of a triple, including triggering conditions, execution actions and expected results, and a directed edge set is constructed according to the dependency relationship between nodes. By assigning edge weights to each rule path and combining scoring factors such as structural similarity, historical reasoning validity and task urgency, the reasoning ability of the rule graph is optimized. This method selects the optimal rule path by evaluating the edge weights of the rule path, which significantly improves the decision-making accuracy and response speed during the training process. Compared with traditional methods, the present invention has greater flexibility and intelligence in terms of dynamic updating of rules, task adjustment and adaptability in complex environments, providing an efficient and accurate solution for mine rescue training, and has a wide range of engineering application value.

[0119] In this embodiment, the construction of the directed edge set between the rule nodes includes determining the connection logic between the rule nodes, comparing the output result of each rule node with the triggering conditions of other rule nodes, and when the output result of one rule node meets the triggering conditions of another rule node, establishing a directed connection path between the two rule nodes, and sequentially establishing all rule paths that meet the dependency relationship to form a directed edge set E between the rule nodes. R .

[0120] This method constructs a set of directed edges between rule nodes and dynamically establishes connection paths based on the dependencies between rule nodes, achieving automated reasoning and path optimization of rule graphs. By comparing the output results of rule nodes with the trigger conditions, the system can adaptively adjust path connections to ensure the flexibility and accuracy of task execution. This method significantly improves the intelligent decision-making ability and task adaptability during training, providing a more efficient and reliable rule-based reasoning mechanism for mine rescue training, and has strong engineering application value.

[0121] In this embodiment, the S5 specifically includes:

[0122] S51. Collect rule triggering records during the training process and construct a historical reasoning trajectory set H;

[0123] S52. Based on the historical reasoning trajectory set H, calculate the trigger frequency of the rule node, and filter the low-frequency redundant node set according to the node contribution, remove all nodes and associated edges belonging to the low-frequency redundant node set from the rule graph, and output the simplified rule node set N ' With edge set E ' R ;

[0124] S53. Update the edge weights of the rule path according to the training feedback score and the historical performance of the rule path to obtain an updated edge weight matrix;

[0125] S54, normalize the updated edge weight matrix to form the edge weight matrix after the regular graph evolution And with the simplified rule node set N ' With edge set E ' R Together they form the evolved regular graph structure G ' R .

[0126] The present invention effectively optimizes the rule graph structure by collecting rule trigger records during the training process and combining them with a set of historical reasoning trajectories. By calculating the trigger frequency of the rule nodes, low-frequency redundant nodes are screened out and removed, which reduces unnecessary computational burden and improves the simplicity and operational efficiency of the rule graph. At the same time, based on the training feedback score and historical performance, the edge weights of the rule path are updated to achieve dynamic optimization of the rule graph. The normalized edge weight matrix ensures accuracy and stability in the rule reasoning process. This method not only improves the response speed and intelligence of the training system, but also enhances the adaptability and scalability of the rule reasoning process, helps to provide more accurate training decisions in complex environments, and has strong engineering application value.

[0127] In this embodiment, S6 specifically includes:

[0128] S61, obtain the mine rescue task jump node output by the deep Q network model, and the evolved rule graph structure G ' R The rule graph obtained by inference recommends task nodes v r ;

[0129] S62. Calculate the structured path value function of the deep Q network direction based on the training state vector:

[0130] F DQN (v d )=Q* (X s ,a d )·φ(v d );

[0131] Among them, F DQN (v d ) is the comprehensive score of the jump node in the deep Q network direction, v d Indicates the jump node of the mine rescue mission, Q * (X s ,a d ) is the structured path value evaluation function, X s is the training state vector, φ(v d ) is the mine rescue mission jump node v d Attribute weight function of

[0132] S63. Based on the rule graph path reasoning results, calculate the comprehensive scoring function of the rule graph recommendation node:

[0133]

[0134] Among them, F RULE (v r ) is the rule graph recommendation task node v r The comprehensive score value, v r is the recommended task node for the rule graph, L is the total number of hops from the current task node to the target node along the rule graph reasoning path, l is the hop index, k is the edge index within the path, and W ′ (e k ) is the weight of the kth edge after evolution, e k represents the kth directed edge on the path in the regular graph structure, δ l is the path length penalty factor, ψ(v k ) is the structural importance scoring function;

[0135] S64, calculating a confidence difference index between the deep Q network score and the rule score;

[0136] S65. Construct the final fusion scoring function:

[0137] F fused (v)=σ(λ d ·F DQN (v d )+λ r ·F RULE (v r )+θ·Δ conf (v d ,v r ));

[0138] Among them, Ffused (v) is the fusion score function, σ(·) is the activation function, and λ d is the scoring weight coefficient of the deep Q network direction, λ r is the scoring weight coefficient of the rule graph direction, θ is the difference adjustment coefficient, Δ conf (v d ,v r ) is the confidence difference index;

[0139] S66. Select the node with the highest fusion score from the candidate nodes as the final task node in the task state diagram under the current training state.

[0140] The present invention realizes intelligent task recommendation in mine rescue training by combining the deep Q network and the evolved rule graph. By calculating the comprehensive score of the recommended nodes of the deep Q network and the rule graph, combining the structured path value evaluation function and the rule graph path reasoning results, the task jump path is dynamically adjusted. By calculating the confidence difference index between the two, the effective fusion of the deep Q network score and the rule reasoning results is ensured, thereby improving the accuracy and flexibility of task recommendation. Finally, by constructing a fusion scoring function, the system can select the optimal jump task node, significantly improving the decision response speed and task adaptability of mine rescue training. This method is efficient, intelligent and scalable, can provide optimized decision support for complex training tasks, and has broad practical application prospects.

[0141] In this embodiment, the calculation of the confidence difference index includes obtaining the comprehensive score value of the jump node in the direction of the deep Q network and the score value of the recommended node output by the rule graph structure, calculating the score difference between the two, and forming a node score difference term; further calculating the difference between the task attribute score of the jump node and the structural importance score of the rule node to form a node structure attribute deviation term; at the same time, obtaining the action strategy probability distribution output by the deep Q network and the normalized weight distribution of the jump action in the rule graph reasoning path, and calculating the KL divergence between the two to form a strategy distribution deviation term; the above three types of difference metrics are weighted and superimposed to construct a confidence difference index Δ conf (v d ,v r ).

[0142] This method achieves efficient fusion of multi-source information by calculating a confidence difference index, combining deep Q-network scoring with rule-based graph recommendation scoring. By measuring score differences, node structure attribute deviations, and policy distribution deviations, the decision-making basis for task node selection is dynamically adjusted. This method significantly improves the intelligence and flexibility of task recommendations, ensuring adaptive adjustment of path selection in complex training environments. Through the weighted superposition of confidence difference indicators, the accuracy and response speed of task jumps are further optimized, demonstrating high efficiency, intelligence, and broad engineering application prospects.

[0143] In this embodiment, the S7 specifically includes:

[0144] S71. Call the final task node in the task state diagram as the task execution target of the current training phase, and retrieve the virtual reality scene binding information corresponding to the final task node;

[0145] S72: Load the virtual reality 3D training scene resources bound to the final task node, including the 3D space structure model, initial environment state parameters, and interactive control components, to complete the initialization of the training sub-scene;

[0146] S73. Based on the environmental parameter configuration of the final task node, push the scene perception data required for task execution, including gas concentration value, visibility level, obstruction location, heat source distribution, and audio and video warning device status;

[0147] S74: Push the training task content information configured in the final task node, including the task objective description, interactive guidance text, operation time limit, risk warning level, and corresponding operation feedback mechanism, to start the virtual reality task process;

[0148] S75. After the training task is completed, collect the behavioral path, response time, operation accuracy, task completion status and interaction record data during the training process to calculate the training feedback score of the current stage.

[0149] The present invention achieves high intelligence and real-time response in mine rescue training through the combination of task state diagram and virtual reality technology. The corresponding virtual reality three-dimensional scene resources are loaded according to the final task node, and the environmental perception data and training task content information related to the task execution are pushed to provide a real training experience. During the training process, not only are environmental parameters such as gas concentration and visibility monitored in real time, but the interactivity and practicality of the training are also improved through interactive control components and task feedback mechanisms. By collecting data such as behavioral paths and response time during the training process, the training effect is dynamically evaluated based on the feedback results, and accurate optimization suggestions are provided for subsequent training. This intelligent training mechanism effectively improves the emergency response capabilities and operational accuracy of mine rescue personnel in complex environments, and has broad practical application value and engineering promotion prospects.

[0150] Example 1:

[0151] To verify the feasibility of this invention, it was applied to an immersive safety training base constructed by a mining enterprise. This base, equipped with a typical mine tunnel structure, ventilation system, and 3D scanning data, is equipped to deploy a virtual reality training platform and collect operational data. Using "simulated rescue training for gas leak emergencies" as the primary task scenario, the system constructed a mine task state diagram, training state vectors, rule graphs, and a scoring fusion mechanism to compare and analyze the performance differences between traditional training methods and the system of this invention in key indicators such as task completion efficiency, decision accuracy, behavioral deviation rate, and task response time.

[0152] In the simulation scenario, the system first constructs a task state diagram based on the mine site's structural information. It sets up multiple task nodes, including operational links such as gas leak identification, alarm issuance, evacuation route determination, blockade area establishment, ventilation equipment activation, and personnel guidance. These nodes are then bound to various functional areas in the virtual reality three-dimensional scene. After trainees wear VR equipment and enter the simulation environment, the system collects their operational behavior, environmental perception responses (such as lighting identification and path selection), physiological status (such as heart rate and body movement frequency), and system status (such as smoke concentration and visibility) in real time, dynamically generating a training state vector.

[0153] During training, the deep Q-network model performs a structured path scoring of all jumpable task nodes based on the current state vector. It then recommends the next execution node based on task urgency, path risk, and jump cost. Simultaneously, the system performs inference operations based on a pre-set rule graph structure, proposes rule-recommended nodes, and uses a scoring fusion function to weight node recommendations from both sources to select the final jump node. The system then loads the corresponding virtual scene, including a dynamic simulation of smoke diffusion, gas concentration warnings, voice guidance, and multimedia interactive prompts, to guide the user through the next task.

[0154] Through comparative data collection in three typical task rehearsal cycles, it was found that compared with the traditional method of executing training tasks in the order of the script, the solution of the present invention significantly improves the dynamic adjustment capability of the training task path and the efficiency of responding to sudden tasks. In traditional systems, users must complete operations in sequence according to fixed steps, and the path adjustment capability is extremely low. Once there is a deviation from the task behavior or an unreasonable path is selected, it is difficult for the system to provide timely feedback and corrections, resulting in a long time to complete the task. In the system of the present invention, due to the introduction of a deep reinforcement learning decision-making and rule graph structure reasoning fusion mechanism based on the behavior state vector, even if the trainees have behavioral deviations at certain task nodes, the system can still recommend a suitable jump path based on the current state and guide the user to re-enter the main process.

[0155] For example, in the early stages of a gas leak, if a trainee fails to promptly identify the risk of gas diffusion and enters a high-concentration area, the system detects environmental parameter changes and behavioral deviation data, outputs real-time task jump correction suggestions, and overloads the ventilation escape path task node to guide the trainee through the emergency evacuation and guidance tasks. After training, the system calculates operation completion time, path rationality score, number of incorrect responses, and node deviation rate. It found that the application of the system reduced the average task completion time from 18.7 minutes to 12.3 minutes, increased the jump node accuracy from 71.5% to 92.6%, shortened the critical task response delay by an average of approximately 4.2 seconds, and reduced the incorrect path selection rate by over 53%.

[0156] The following is the comparative data of key performance indicators during training:

[0157] Table 1 Comparison of key performance indicators of the system of the present invention and the traditional system in mine rescue training

[0158]

[0159]

[0160] From the comparative data in Table 1 above, it can be seen that in the actual application of mine rescue training and mission deduction, the present invention is significantly superior to traditional virtual training systems in terms of task completion efficiency, jump path decision accuracy, key node response timeliness, and behavioral deviation correction capabilities. Specifically, in terms of task completion efficiency, the present invention breaks the rigid structure of traditional process linear execution by constructing a dynamic task state diagram based on virtual reality and introducing a deep Q network and rule graph fusion mechanism, thereby realizing real-time path optimization and intelligent jump control. The data in the table show that the average task completion time has been significantly shortened from 18.7 minutes of the traditional system to 12.3 minutes, and the task execution efficiency has increased by more than 34%, reflecting the advantages of the present invention in path dynamic planning and rapid response capabilities.

[0161] In terms of jump path decision accuracy, this invention leverages a state-vector-driven reinforcement learning model and a structured rule-based graph inference mechanism to achieve optimal node recommendations based on multi-source information fusion. This effectively avoids the traditional system's decision-making issues of task jumps based on empirical assumptions and isolated, unconnected rules. The data in the table shows that the proposed system achieves a jump node selection accuracy of 92.6%, an improvement of over 20 percentage points compared to the traditional system's 71.5%. This demonstrates that the proposed system's recommended paths are more reasonable and task-appropriate in complex scenarios.

[0162] From the perspective of behavioral deviation control capabilities, traditional training systems lack a targeted feedback mechanism. Once trainees deviate from the preset operating route, it is difficult for the system to intervene in time, resulting in reduced training efficiency. However, during the execution of the task, the present invention calculates the behavioral offset in real time and dynamically corrects the jump path based on the scoring mechanism, successfully reducing the node behavior deviation rate from 24.3% to 9.7%, and the number of erroneous responses from 2.9 to 1.2, significantly improving operational standardization and task recovery capabilities. At the same time, the system also has a path correction function, which automatically triggers correction 2.4 times per training on average, effectively reducing the impact of human bias.

[0163] Regarding the timeliness of response at key nodes, this invention utilizes a deep Q-network to evaluate action costs and benefits, and weights jump paths based on the urgency score of the rule graph. This enables the system to react more quickly to critical emergencies such as gas leaks and personnel guidance. Data shows that the average response latency of the proposed system is 2.6 seconds, compared to 6.8 seconds for traditional systems, a response speed improvement of over 60%. This is of great significance for simulating escape decisions and emergency deployment in real-world mining accidents.

[0164] Furthermore, in terms of high-risk path judgment, the present invention reduces the high-risk path misselection rate from 13.4% to 6.2% through score fusion and structural evolution mechanisms, reducing the error selection rate by approximately 54%. Combined with the system's coordinated monitoring of environmental parameters, user operations, and task maps, the training path is guaranteed to be both optimal and safe, enhancing fault tolerance and robustness.

[0165] In terms of system loading and interactive experience, this invention leverages an efficient binding mechanism between task nodes and VR scenes, along with a content push module, to enhance the consistency and realism of training scenarios. Data shows that the scene loading accuracy rate has increased to 98.5%, and the user satisfaction score has reached 4.6 points, both higher than the 87.2% and 3.4 points of traditional systems, demonstrating the improvements this invention has made to the user experience.

[0166] In summary, the present invention not only effectively solves the problems of existing mine rescue training systems in terms of fixed task paths, fragmented rules, slow response, and uncontrollable behavioral deviations, but also builds an integrated intelligent training closed loop of "modeling-decision-evolution-control" through the deep integration of reinforcement learning and structural graphs. It has obvious advantages in multiple dimensions such as efficiency, accuracy, response speed, and adaptability, and has the prospect of promotion in high-risk operating environments such as mines, tunnels, and subways, and can inject new intelligent capabilities into the emergency training system.

[0167] The above description is only a preferred specific embodiment of the present invention, but the scope of protection of the present invention is not limited thereto. Any technician familiar with the technical field, within the technical scope disclosed by the present invention, who makes equivalent replacements or changes based on the technical solution and inventive concept of the present invention, should be covered by the scope of protection of the present invention.

Claims

1. A mine rescue training and dangerous dynamic simulation method based on virtual reality, characterized in that: The steps include: S1. Model the mine rescue training process as a directed graph structure and output a task state graph with node binding to the virtual reality scene; S2. During the training process, the operator's operation behavior, current task status, virtual environment parameters, and physiological response indicators are collected to construct a training state vector; S3. Input the training state vector into the deep Q network model, extract the behavior and environment characteristics, calculate the jump node evaluation value in the task state diagram, and output the mine rescue task jump node; S4. Model the training control rules as rule nodes. Rule nodes are connected by logical edges to form a directed acyclic graph structure. Weight values ​​are bound to the edges to output the rule graph structure. S5. Based on the task execution results and behavioral feedback recorded during the training process, the rule graph structure is self-evolved and the evolved rule graph structure is output; S6. Based on the training state vector, perform reasoning operations in the evolved rule graph structure to obtain the rule graph recommended task node, fuse the rule graph recommended task node with the mine rescue task jump node, and output the final task node; S7. Execute the jump control of the task state diagram according to the final task node, load the virtual reality training scene corresponding to the final task node, push the environmental parameters, training content and risk information, and complete the training process.

2. The method for mine rescue training and dangerous dynamic simulation based on virtual reality according to claim 1, characterized in that: Said S1 specifically includes: S11. Collecting mine training area structural data and historical mine accident data, wherein the mine training area structural data includes tunnel topology, key node locations, ventilation paths, and safe haven locations; and the historical mine accident data includes accident type, accident triggering conditions, accident location, rescue path, and failed mission node records, to construct a mine training basic dataset. S12. Construct a mine scene topology map based on the mine training basic data set, wherein the mine scene topology map includes a set of scene nodes and a set of spatially connected paths between the scene nodes, and output the mine scene topology map; S13. Define a set of mine rescue training task nodes, where each training task node is a specific rescue training task. The training tasks include a hazard source identification task, an alarm task, an evacuation judgment task, an on-site rescue task, and a personnel transfer task. S14. Construct a task state graph based on the mine rescue training task node set. The task state graph is a directed graph structure, including a training task node set and a state transition path set between the nodes. Each state transition path is a directed connection relationship from one training task node to another training task node. S15. Assign a unique task number, task objective description, task triggering condition, task required environment parameters and virtual reality scene mapping instructions to each training task node in the task status diagram, and output a task node attribute information set; S16. Bind each training task node to the corresponding virtual reality three-dimensional scene area according to the task node attribute information set and the task status graph, load it when the virtual reality system is initialized, and output the task status graph bound to the virtual reality scene.

3. The method for mine rescue training and dangerous dynamic simulation based on virtual reality according to claim 1, characterized in that: The training state vector includes a task node number, an operation behavior code, environmental state parameters, and physiological state parameters; the environmental state parameters include smoke concentration, temperature, light intensity, toxic gas index, and obstructions; the physiological state parameters include heart rate, body movement frequency, reaction delay, and voice intensity.

4. The method for mine rescue training and dangerous dynamic simulation based on virtual reality according to claim 1, characterized in that: The S3 specifically includes: S31, input the training state vector as input to the input layer of the deep Q network model; S32. Set multiple fully connected hidden layers in the deep Q network model to extract features from the training state vector and generate an intermediate representation of the task behavior and the environment characteristics; S33: Input the intermediate expression into the output layer and output the basic action value function of all jumpable actions in the current training state; S34, according to action a in the task state diagram i The target node v j , construct a structured path value function: Q * (X s ,a i )=Q(X s ,a i )·ω d (v j )-λ c ·C(a i ); Among them, Q * (X s ,a i ) is the structured path value evaluation function, Q(X s ,a i ) is the basic action value function, X s is the training state vector, a i is a jump action starting from the current task node in the current task state diagram, ω d (v j ) is the target task node v j The task importance factor, v j For jump action a i The target node, λ c is the risk control weight coefficient, C(a i ) is jump action a i The corresponding path cost indicator; S35. Select the action with the largest structured path value function from all actions as the optimal jump action a under the current training state. * ; S36, according to the optimal jump action a * Corresponding jump relationship, determine the target task node v t+1 , the target task node v t+1 The target node is defined as the mine rescue mission jump node.

5. The method for mine rescue training and dangerous dynamic simulation based on virtual reality according to claim 1, characterized in that: The S4 specifically includes: S41. Based on the characteristics of mine rescue training tasks, construct a rule node set N = {R1, R2, ..., R n }, each rule node R k Defined as a triplet R k =(C k ,A k ,O k ), where C k Indicates the rule triggering condition, A k Indicates the rule execution action, O k Indicates the expected result after the action is performed; S42. According to the conditional dependency between the rule nodes, a directed edge set E between the rule nodes is constructed. R ={e ij ∣e ij =(R i ,R j )}; S43, based on the rule node set N and directed edge set E R , construct the regular graph structure G R ; S44. Assign edge weights to each rule connection path in the rule graph structure: W(e ij )=α·S c (R i ,R j )+β·H f (P ij )+γ·T r (R j ); Among them, W(e ij ) is from the rule node R i To rule node R j The edge weight, e ij represents a directed edge in the rule graph, α is the weight coefficient of the structural similarity score item, S c (R i ,R j ) is a rule node R i The output result of the rule node R j The structural similarity score between the trigger conditions, R i , R j is a rule node, i represents the number of the current rule node in the rule graph, and j represents the node with the same i The target rule node number with dependency relationship, β is the weight coefficient of historical reasoning validity scoring item, H f (P ij ) represents the rule path R i →R j The frequency score of being effectively triggered in the past training process, γ is the weight coefficient of the node task urgency score item, T r (R j ) is a rule node R j The urgency score of the corresponding task in the task status diagram; S45. Based on the edge weight results of all rule paths, a rule graph edge weight matrix is ​​constructed, and a rule graph structure including a rule node set, a rule connection structure and an edge weight matrix is ​​output.

6. The method for mine rescue training and dangerous dynamic simulation based on virtual reality according to claim 5, characterized in that: The method of constructing a directed edge set between rule nodes includes determining the connection logic between rule nodes, comparing the output result of each rule node with the triggering conditions of other rule nodes, and establishing a directed connection path between the two rule nodes when the output result of one rule node meets the triggering conditions of another rule node. All rule paths that meet the dependency relationship are established in turn to form a directed edge set E between rule nodes. R .

7. The method for mine rescue training and dangerous dynamic simulation based on virtual reality according to claim 1, characterized in that: The S5 specifically includes: S51. Collect rule triggering records during the training process and construct a historical reasoning trajectory set H; S52. Based on the historical reasoning trajectory set H, calculate the trigger frequency of the rule node, and filter the low-frequency redundant node set according to the node contribution, remove all nodes and associated edges belonging to the low-frequency redundant node set from the rule graph, and output the simplified rule node set N ' With edge set E ' R ; S53. Update the edge weights of the rule path according to the training feedback score and the historical performance of the rule path to obtain an updated edge weight matrix; S54, normalize the updated edge weight matrix to form the edge weight matrix after the regular graph evolution And with the simplified rule node set N ' With edge set E ' R Together they form the evolved regular graph structure G ' R .

8. The method for mine rescue training and dangerous dynamic simulation based on virtual reality according to claim 1, characterized in that: The S6 specifically includes: S61, obtain the mine rescue task jump node output by the deep Q network model, and the evolved rule graph structure G ' R The rule graph obtained by inference recommends task nodes v r ; S62. Calculate the structured path value function of the deep Q network direction based on the training state vector: F DQN (v d )=Q * (X s ,a d )·φ(v d ); Among them, F DQN (v d ) is the comprehensive score of the jump node in the deep Q network direction, v d Indicates the jump node of the mine rescue mission, Q * (X s ,a d ) is the structured path value evaluation function, X s is the training state vector, φ(v d ) is the mine rescue mission jump node v d Attribute weight function of S63. Based on the rule graph path reasoning results, calculate the comprehensive scoring function of the rule graph recommendation node: Among them, F RULE (v r ) is the rule graph recommendation task node v r The comprehensive score value, v r is the recommended task node for the rule graph, L is the total number of hops from the current task node to the target node along the rule graph reasoning path, l is the hop index, k is the edge index within the path, and W ′ (e k ) is the weight of the kth edge after evolution, e k represents the kth directed edge on the path in the regular graph structure, δ l is the path length penalty factor, ψ(v k ) is the structural importance scoring function; S64, calculating a confidence difference index between the deep Q network score and the rule score; S65. Construct the final fusion scoring function: F fused (v)=σ(λ d ·F DQN (v d )+λ r ·F RULE (v r )+θ·D conf (v d ,v r )); Among them, F fused (v) is the fusion score function, σ(·) is the activation function, and λ d is the scoring weight coefficient of the deep Q network direction, λ r is the scoring weight coefficient of the rule graph direction, θ is the difference adjustment coefficient, Δ conf (v d ,v r ) is the confidence difference index; S66. Select the node with the highest fusion score from the candidate nodes as the final task node in the task state diagram under the current training state.

9. The method for mine rescue training and dangerous dynamic simulation based on virtual reality according to claim 8, characterized in that: The calculation of the confidence difference index includes obtaining the comprehensive score value of the jump node in the direction of the deep Q network and the score value of the recommended node output by the rule graph structure, calculating the score difference between the two, and forming a node score difference term; further calculating the difference between the task attribute score of the jump node and the structural importance score of the rule node, and forming a node structure attribute deviation term; at the same time, obtaining the action strategy probability distribution output by the deep Q network and the normalized weight distribution of the jump action in the rule graph reasoning path, and calculating the KL divergence between the two, and forming a strategy distribution deviation term; The above three types of difference metrics are weighted and superimposed to construct the confidence difference index Δ conf (v d ,v r ).

10. The method for mine rescue training and dangerous dynamic simulation based on virtual reality according to claim 1, characterized in that: The S7 specifically includes: S71. Call the final task node in the task state diagram as the task execution target of the current training phase, and retrieve the virtual reality scene binding information corresponding to the final task node; S72: Load the virtual reality 3D training scene resources bound to the final task node, including the 3D space structure model, initial environment state parameters, and interactive control components, to complete the initialization of the training sub-scene; S73. Based on the environmental parameter configuration of the final task node, push the scene perception data required for task execution, including gas concentration value, visibility level, obstruction location, heat source distribution, and audio and video warning device status; S74: Push the training task content information configured in the final task node, including the task objective description, interactive guidance text, operation time limit, risk warning level, and corresponding operation feedback mechanism, to start the virtual reality task process; S75. After the training task is completed, collect the behavioral path, response time, operation accuracy, task completion status and interaction record data during the training process to calculate the training feedback score of the current stage.

Citation Information

Cited By

  • Virtual reality immersive practical training content generation method

    CN120852118A

  • A method for generating immersive virtual reality training content

    CN120852118B

  • Mining self-rescuer training examination system and method

    CN121073726A

  • Evaluation system for college student occupational scene simulation training

    CN121437232A