Multimodal internet of things perception driven dangerous scene virtual training method and system
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2026-04-02
- Publication Date
- 2026-08-11
AI Technical Summary
[0003]然而,这类常规方法存在明显局限
本方法能够实现对真实危险场景的精细化、动态化数字孪生构建。基于分布式物联感知节点采集的多模态数据,通过跨模态关联性挖掘与危险因素识别,生成的场景语义描述模型不仅包含静态的环境要素,更深度嵌入了危险因素的时空演化逻辑,为虚拟实训提供了高度拟真的语义基础。所构建的虚拟危险场景具有物理一致性约束,环境状态数据与虚拟空间节点精确绑定,确保了虚拟环境中物理规律与真实世界的一致性,显著提升了受训者的沉浸感与训练代入感。
Smart Images

Figure CN122551635A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to multimodal sensing technology, and more particularly to a method and system for virtual training in hazardous scenarios driven by multimodal IoT sensing. Background Technology
[0002] Virtual training technology for hazardous scenarios aims to provide operators with a safe and controllable training experience by simulating real hazardous environments. Existing technologies typically rely on pre-built static 3D virtual scene models. These models are created based on historical accident data, design drawings, or limited environmental scan data, striving to visually reproduce the spatial layout and key equipment of the target scene. To enhance immersion and training relevance, the system pre-sets a series of fixed hazardous event trigger points and scripted evolution paths within the virtual scene based on expert experience. When the trainee's interactive behavior meets the preset conditions, the corresponding hazardous state change and feedback are triggered, thus completing one training cycle.
[0003] However, these conventional methods have significant limitations. Static models struggle to reflect the dynamics and uncertainties of real-world hazardous scenarios. Their pre-defined logic for hazard evolution is rigid, failing to simulate complex and unexpected risk situations arising from real-time changes in environmental parameters, cascading equipment failures, or human error. This results in insufficient alignment between training content and real-world conditions, significantly diminishing training effectiveness. Furthermore, existing methods often rely on single or a few data sources to construct scenarios, lacking deep integration and utilization of multi-dimensional and multi-type perceptual information from the field. This renders virtual scenarios lacking physical consistency and spatiotemporal realism, further weakening the sense of presence and effectiveness of training. Summary of the Invention
[0004] This invention provides a method and system for virtual training in hazardous scenarios driven by multimodal IoT perception, which can solve the problems in the prior art.
[0005] A first aspect of this invention provides a virtual training method for hazardous scenarios driven by multimodal IoT perception, comprising: Multimodal sensing data of the target hazardous scene is collected by distributed IoT sensing nodes. Based on the spatiotemporal alignment rules, cross-modal correlation mining and hazard factor identification are performed on the multimodal sensing data to generate a scene semantic description model. A virtual scene topology is constructed based on the scene semantic description model, and data binding is performed based on the spatial position association between the environmental state data in the multimodal perception data and the spatial nodes in the virtual scene topology to form a virtual dangerous scene with physical consistency constraints. Based on the interactive operations of the training subjects in the virtual dangerous scenario, a behavioral feature sequence is generated. Then, based on the evolution logic of the risk factors in the scenario semantic description model, a temporal correlation analysis and safety assessment are performed on the behavioral feature sequence to obtain a risk response strategy. The risk response strategy is applied to the virtual scene topology to trigger the dynamic evolution of risk factors and generate feedback signals. The feedback signals are then used to adaptively correct the evolution logic in the scene semantic description model.
[0006] Based on spatiotemporal alignment rules, cross-modal correlation mining and risk factor identification are performed on the multimodal sensing data to generate a scene semantic description model, including: A unified spatiotemporal reference framework is established to align and calibrate the temporal reference and spatial coordinates of different modal data in multimodal sensing data, and a spatiotemporal co-occurrence relationship diagram of multimodal data is constructed under the unified spatiotemporal reference framework. Based on the response intensity distribution in the spatiotemporal co-occurrence relationship diagram, the co-evolution mode of cross-modal data within a specific spatiotemporal window is identified. By analyzing the similarity and deviation between the co-evolution mode and the evolution law of known dangerous scenarios, the activation state and spatial diffusion direction of potential dangerous factors in the current scenario are determined. Based on the activation state and the spatial diffusion direction, the temporal dependency constraints and causal propagation paths of the hazard factors are extracted. Based on the temporal dependency constraints and the causal propagation paths, a scene semantic description model is constructed, which includes hazard factor identification, spatial distribution boundaries, evolutionary stage division, and propagation link topology.
[0007] A unified spatiotemporal reference framework is established to align and calibrate the temporal reference and spatial coordinates of different modal data in multimodal sensing data. A spatiotemporal co-occurrence relationship diagram of the multimodal data is then constructed within this unified spatiotemporal reference framework, including: The timestamp sequence and spatial coordinate distribution of each modality data are extracted from multimodal sensing data. A cross-correlation function of cross-modal time series and a registration error field of spatial coordinates are constructed. The systematic time delay deviation and periodic drift characteristics of each modality data relative to the reference time axis are identified by the peak position of the cross-correlation function. The rigid transformation parameters and nonlinear distortion parameters of each modality data relative to the reference coordinate system are identified by the gradient distribution of the registration error field. Based on the systematic time delay deviation, the periodic drift characteristics, the rigid transformation parameters and the nonlinear distortion parameters, adaptive spatiotemporal calibration of each modality data is performed to establish a unified spatiotemporal reference framework. Under the unified spatiotemporal reference framework, the calibrated multimodal data is projected onto a multi-scale spatiotemporal grid hierarchy. Local observation manifolds of different modal data are extracted within each scale spatiotemporal grid, and the topological consistency metric and dynamic evolution synchronization metric between the local observation manifolds are calculated. Using a multi-scale spatiotemporal grid as nodes and the weighted fusion value of the topological consistency metric and the dynamic evolution synchronicity metric as edge weights, a spatiotemporal co-occurrence relationship graph of multimodal data is constructed.
[0008] A virtual scene topology is constructed based on the scene semantic description model, and data binding is performed based on the spatial position association between the environmental state data in the multimodal perception data and the spatial nodes in the virtual scene topology to form a virtual hazardous scene with physical consistency constraints, including: The causal propagation path and spatial distribution boundary of the risk factors are extracted from the scene semantic description model. A directed topological connection between spatial nodes is constructed based on the directionality of the transmission link in the causal propagation path. The influence domain range of each spatial node is determined based on the spatial distribution boundary, forming a virtual scene topology structure that includes the node influence domain and directed transmission constraints. For environmental state data in multimodal perception data, calculate the spatial inclusion relationship between the spatial coordinates of each environmental state data and the influence domain of each spatial node in the virtual scene topology. Bind environmental state data whose spatial coordinates fall within the same influence domain to the corresponding spatial node, and determine the constraint transmission weight of each spatial node's state change on adjacent nodes based on the directed transit constraint. Based on the constraint transmission weights and the evolution logic in the scene semantic description model, physical consistency constraints for state evolution are configured for each spatial node, and a virtual dangerous scene is constructed based on the virtual scene topology and the physical consistency constraints.
[0009] Based on the interactive operations of the training subjects in the virtual dangerous scenario, a behavioral feature sequence is generated. Then, based on the hazard factor evolution logic in the scenario semantic description model, a temporal correlation analysis and safety assessment are performed on the behavioral feature sequence to obtain risk response strategies, including: Record the interaction timestamps of the training objects in the virtual dangerous scenario and construct a behavioral feature sequence with the operation object identifier. Analyze the temporal dependency constraints of the dangerous factors in the scenario semantic description model. Timely align the operation time interval in the behavioral feature sequence with the evolution time window in the temporal dependency constraints. Calculate the temporal deviation between the operation timing of the training objects and the evolution time of the dangerous factors. Based on the temporal deviation, the operation segment corresponding to the evolution time of the risk factor is identified from the behavioral feature sequence. The spatial consistency between the operation action direction and the risk factor propagation direction is calculated by comparing the matching relationship between the operation object identifier in the operation segment and the node identifier on the causal propagation path in the scene semantic description model. Based on the temporal deviation and spatial consistency, it is determined whether the operation of the training object meets the requirements for blocking risk factors. For operations that do not meet the blocking requirements, the temporal adjustment amount and spatial redirection target are identified, and a risk response strategy including operation timing correction suggestions and operation object selection suggestions is generated.
[0010] Based on the temporal deviation, operational segments corresponding to the evolution time of risk factors are identified from the behavioral feature sequence. The spatial consistency between the direction of operational action and the direction of risk factor propagation is calculated by comparing the matching relationship between the operational object identifier in the operational segment and the node identifiers on the causal propagation path in the scene semantic description model. This includes: Using the time sequence deviation as the time window screening threshold, operations whose timestamps fall within the preset time window range before and after the evolution time of the hazard factor are selected from the behavioral feature sequence. The selected operations are organized into operation segments, and the operation object identifiers of each operation in the operation segment are parsed. The node identifier sequence carrying the hazard propagation on the causal propagation path is parsed from the scene semantic description model, and the spatial position mapping relationship between the operation object identifier and the node identifier sequence is established. Based on the spatial location mapping relationship, the topological position of the operation object corresponding to each operation in the operation segment is identified on the causal propagation path. The direction of change of the topological position of the operation object between adjacent operations on the causal propagation path is analyzed. The vector angle between the direction of change of the topological position and the preset propagation direction of the causal propagation path is calculated. The spatial consistency between the direction of operation and the direction of propagation of the hazard factor is quantified based on the vector angle.
[0011] Applying the risk response strategy to the virtual scene topology triggers the dynamic evolution of risk factors and generates feedback signals. Adaptively correcting the evolution logic in the scene semantic description model using these feedback signals includes: The operational suggestions in the risk response strategy are transformed into state intervention instructions in the virtual scene. The state intervention instructions are applied to the target space nodes of the virtual scene topology. The state changes are transmitted through the topological connection relationship in the virtual scene topology, triggering the dynamic evolution of the risk factors along the causal propagation path. The actual evolution trajectory of the risk factors is recorded. The actual evolution trajectory is compared with the preset evolution trajectory in the scene semantic description model to identify the trajectory deviation position and deviation magnitude and generate feedback signals. The trajectory deviation position in the feedback signal is analyzed to locate the logical defects in the corresponding temporal dependency constraints or causal propagation paths in the scene semantic description model. The deviation magnitude in the feedback signal is analyzed to quantify the impact of the logical defects on evolution prediction. Based on the location result and impact of the logical defects, the evolution logic in the scene semantic description model is adaptively corrected using the feedback signal.
[0012] A second aspect of the present invention provides a multimodal IoT-driven virtual training system for hazardous scenarios, comprising: The perception data unit is used to collect multimodal perception data of the target hazardous scene through distributed IoT perception nodes, and to perform cross-modal correlation mining and hazard factor identification on the multimodal perception data based on spatiotemporal alignment rules to generate a scene semantic description model. The virtual scene unit is used to construct a virtual scene topology based on the scene semantic description model, and to bind data according to the spatial position association between the environmental state data in the multimodal perception data and the spatial nodes in the virtual scene topology to form a virtual dangerous scene with physical consistency constraints. The behavioral safety unit is used to generate a sequence of behavioral features based on the interactive operations of the training object in the virtual dangerous scenario, and to perform temporal correlation analysis and safety assessment on the sequence of behavioral features based on the evolution logic of the danger factors in the scenario semantic description model, so as to obtain a risk response strategy. The dynamic evolution unit is used to apply the risk response strategy to the virtual scene topology, trigger the dynamic evolution of risk factors and generate feedback signals, and use the feedback signals to adaptively correct the evolution logic in the scene semantic description model.
[0013] A third aspect of the present invention provides an electronic device, comprising: processor; Memory used to store processor-executable instructions; The processor is configured to invoke instructions stored in the memory to execute the aforementioned method.
[0014] A fourth aspect of the present invention provides a computer-readable storage medium having stored thereon computer program instructions that, when executed by a processor, implement the aforementioned method.
[0015] The beneficial effects of this application are as follows: This method enables the construction of refined and dynamic digital twins for real-world hazardous scenarios. Based on multimodal data collected from distributed IoT sensing nodes, through cross-modal correlation mining and hazard factor identification, the generated scene semantic description model not only includes static environmental elements but also deeply embeds the spatiotemporal evolution logic of hazard factors, providing a highly realistic semantic foundation for virtual training. The constructed virtual hazardous scenarios have physical consistency constraints, and environmental state data is precisely bound to virtual space nodes, ensuring the consistency between the physical laws in the virtual environment and the real world, significantly enhancing the trainees' immersion and training engagement.
[0016] This method enables deep, intelligent assessment and dynamic risk guidance of trainees' behavior. By analyzing the behavioral feature sequences generated by trainees in virtual scenarios and performing temporal correlation analysis based on the evolutionary logic in the scenario semantic model, it can accurately assess the safety of their operations and generate targeted risk response strategies. This assessment goes beyond simple right or wrong judgment; it can correlate preceding and following operation sequences to identify behavioral patterns that lead to the accumulation or triggering of dangers.
[0017] The system possesses dynamic evolution and self-optimization capabilities. The generated risk response strategies directly impact the virtual scene, triggering the dynamic evolution of hazardous factors and thereby altering the training context and challenge difficulty in real time. Simultaneously, the system utilizes feedback signals generated during the evolution process to adaptively correct the evolutionary logic in the scene semantic description model, making the hazardous evolution process of the virtual scene more closely resemble the complexity and uncertainty of the real world. This enables intelligent adjustment of training difficulty and continuous enrichment of training content. Attached Figure Description
[0018] Figure 1 This is a flowchart illustrating the multimodal IoT perception-driven virtual training method for hazardous scenarios according to an embodiment of the present invention. Figure 2 This is a flowchart of the adaptive calibration process for cross-modal data based on multi-scale spatiotemporal grids, according to an embodiment of the present invention. Detailed Implementation
[0019] To make the objectives, technical solutions, and advantages of the embodiments of the present invention clearer, the technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.
[0020] The technical solution of the present invention will be described in detail below with reference to specific embodiments. These specific embodiments can be combined with each other, and the same or similar concepts or processes may not be described again in some embodiments.
[0021] Figure 1 This is a flowchart illustrating the multimodal IoT perception-driven virtual training method for hazardous scenarios according to an embodiment of the present invention. Figure 1 As shown, the method includes: Multimodal sensing data of the target hazardous scene is collected by distributed IoT sensing nodes. Based on the spatiotemporal alignment rules, cross-modal correlation mining and hazard factor identification are performed on the multimodal sensing data to generate a scene semantic description model. A virtual scene topology is constructed based on the scene semantic description model, and data binding is performed based on the spatial position association between the environmental state data in the multimodal perception data and the spatial nodes in the virtual scene topology to form a virtual dangerous scene with physical consistency constraints. Based on the interactive operations of the training subjects in the virtual dangerous scenario, a behavioral feature sequence is generated. Then, based on the evolution logic of the risk factors in the scenario semantic description model, a temporal correlation analysis and safety assessment are performed on the behavioral feature sequence to obtain a risk response strategy. The risk response strategy is applied to the virtual scene topology to trigger the dynamic evolution of risk factors and generate feedback signals. The feedback signals are then used to adaptively correct the evolution logic in the scene semantic description model.
[0022] In one optional implementation, cross-modal correlation mining and hazard factor identification are performed on the multimodal sensing data based on spatiotemporal alignment rules to generate a scene semantic description model, including: A unified spatiotemporal reference framework is established to align and calibrate the temporal reference and spatial coordinates of different modal data in multimodal sensing data, and a spatiotemporal co-occurrence relationship diagram of multimodal data is constructed under the unified spatiotemporal reference framework. Based on the response intensity distribution in the spatiotemporal co-occurrence relationship diagram, the co-evolution mode of cross-modal data within a specific spatiotemporal window is identified. By analyzing the similarity and deviation between the co-evolution mode and the evolution law of known dangerous scenarios, the activation state and spatial diffusion direction of potential dangerous factors in the current scenario are determined. Based on the activation state and the spatial diffusion direction, the temporal dependency constraints and causal propagation paths of the hazard factors are extracted. Based on the temporal dependency constraints and the causal propagation paths, a scene semantic description model is constructed, which includes hazard factor identification, spatial distribution boundaries, evolutionary stage division, and propagation link topology.
[0023] When establishing a unified spatiotemporal reference framework, it is necessary to align and calibrate the time reference and spatial coordinates of different modal data in multimodal sensing data. Multimodal sensing data typically includes visual image streams, depth point clouds, inertial measurement unit outputs, environmental acoustic signals, and physical quantity sensing data such as temperature and pressure. These data are generated by different sensors under their own independent sampling clocks and coordinate systems. Time reference alignment requires identifying the local clock offset and drift rate of each sensor. By collecting the timestamp differences of the synchronous trigger signals of each sensor, a clock mapping function is established to uniformly transform the timestamps of all modal data to the reference clock domain. Spatial coordinate alignment requires determining the pose transformation matrix of each sensor relative to the scene reference coordinate system. By calibrating the correspondence of common feature points of the target in the fields of view of multiple sensors, the rotation matrix and translation vector are solved to uniformly transform the spatial coordinates of each modal data to the scene reference coordinate system. After alignment and calibration are completed, a spatiotemporal co-occurrence relationship graph of multimodal data is constructed under a unified spatiotemporal reference framework. This graph uses spatiotemporal units as nodes, and each spatiotemporal unit corresponds to a set of multimodal data in a specific spatial region of the scene within a specific time window. The edges between nodes represent the adjacency relationship and data flow between spatiotemporal units, and the edge weights reflect the correlation strength of data features in adjacent spatiotemporal units.
[0024] Constructing a spatiotemporal co-occurrence graph requires grouping and aggregating aligned multimodal data according to a spatial grid and a temporal window. The spatial grid divides the scene into several spatial voxels, with voxel sizes determined by the scene's extent and sensor resolution, typically ranging from 0.5 m³ to 2 m³. The temporal window length is determined based on the minimum observable timescale of the hazard evolution, typically ranging from 0.1 sec to 1 sec. For each spatiotemporal unit, statistical features of each modality are extracted as node attributes: color histogram and texture energy for the visual modality; surface normal vector distribution and curvature statistics for the depth modality; acceleration amplitude and angular velocity direction for the inertial modality; spectral peak positions and energy concentration for the acoustic modality; and mean and variance for the physical quantity modality. Edge weights between adjacent spatiotemporal units are quantified by calculating the cosine similarity or mutual information of the node attribute vectors. Connections are established between unit pairs with similarity greater than 0.7 or mutual information greater than 0.5.
[0025] Based on the response intensity distribution in the spatiotemporal co-occurrence graph, we identify the co-evolutionary patterns of cross-modal data within a specific spatiotemporal window. The response intensity distribution is determined by the activation degree of multimodal data at each node in the spatiotemporal co-occurrence graph. The activation degree is calculated using the norm or weighted sum of the node attribute vectors, and the weight coefficients reflect the contribution of each modality to the characterization of risk factors. The co-evolutionary pattern manifests as a spatiotemporal propagation path of activation intensity formed by several nodes in the spatiotemporal co-occurrence graph within a continuous time window. This path exhibits a specific spatial diffusion direction and temporal dependency. Identifying the co-evolutionary pattern requires performing spatiotemporal clustering and path tracing on the spatiotemporal co-occurrence graph. The clustering algorithm uses a density-based spatial clustering method, grouping nodes with response intensities higher than a threshold and spatial distances less than two voxel sizes into the same cluster. The path tracing algorithm establishes temporal associations between clusters within a continuous time window, forming a propagation link from the initial activated node to the current activated node.
[0026] This study analyzes the similarity and deviation between co-evolutionary patterns and known hazardous scenario evolution patterns to determine the activation state and spatial diffusion direction of potential hazards in the current scenario. Known hazardous scenario evolution patterns are stored as a template library. Each template records typical spatiotemporal propagation path characteristics of a specific type of hazard, including path length distribution, diffusion velocity range, node activation intensity variation curves, and multimodal response ratios. Similarity calculation employs a dynamic time warping algorithm, aligning the spatiotemporal path of the current co-evolutionary pattern with the paths of each template in the library. The Euclidean distance of the aligned path feature vectors is calculated; templates with a distance less than a preset threshold are identified as similar hazardous scenarios. Deviation calculation focuses on local differences between the current path and the most similar template path, particularly path branch points, propagation velocity abrupt change points, and modal response anomalies. Path segments with a deviation greater than 0.3 are marked as potential hazard evolution anomaly regions. Activation state determination is based on the number of nodes and response intensity of the co-evolutionary pattern at the current moment. A hazard is considered active when the number of nodes exceeds 5 and the average response intensity exceeds a threshold of 0.6. The spatial diffusion direction is determined by calculating the displacement vector of the centroid of the active node within a continuous time window. The direction of the displacement vector is the main diffusion direction of the hazard factor.
[0027] Based on activation state and spatial diffusion direction, the temporal dependency constraints and causal propagation paths of hazard factors are extracted. Temporal dependency constraints describe the temporal sequence and duration constraints between different stages in the evolution of hazard factors, extracted by analyzing the temporal variation patterns of node activation intensity in a co-evolutionary model. The node activation intensity time series is segmented, with each segment corresponding to the zero point of the first derivative or the extreme point of the second derivative of the activation intensity. Each segment corresponds to one stage of hazard factor evolution. The time interval between adjacent stages constitutes the temporal dependency constraint, the mean and variance of which are statistically obtained from historical hazard scenario data. Causal propagation paths describe the diffusion links of hazard factors in space and the causal influence relationships between nodes, extracted by analyzing the temporal order of node activation and spatial adjacency relationships in the spatiotemporal co-occurrence graph. For any two nodes, if the activation time of the preceding node is earlier than that of the following node and the spatial distance is less than 3 voxel dimensions, a directed causal edge is established between the two nodes. The weight of the edge is determined by the peak value of the cross-correlation of the time delay functions of the response intensities of the two nodes. The causal propagation path is formed by connecting a series of causal edges. The starting point of the path is the earliest activated node, and the ending point of the path is the boundary node activated at the current moment.
[0028] Based on temporal dependency constraints and causal propagation paths, a scene semantic description model is constructed, including hazard factor identification, spatial distribution boundaries, evolutionary stage divisions, and propagation link topology. Hazard factor identification is determined by the matched hazardous scene template category, while simultaneously recording the similarity score and main deviation features between the current scene and the template. The spatial distribution boundary is determined by the convex hull of the current set of activated nodes, with the vertex coordinates of the convex hull represented in the scene reference coordinate system. The boundary update frequency is consistent with the time window length. The evolutionary stage division divides the hazard factor evolution process into an initial germination stage, a rapid diffusion stage, a stable propagation stage, and a decay and retreat stage. Each stage corresponds to a specific segment of the node activation intensity time series, and the segment boundary time and duration interval are recorded in the model. The propagation link topology stores the causal propagation path in the form of a directed graph. Nodes in the graph correspond to spatiotemporal units, and edges correspond to causal relationships. Node attributes include spatial coordinates, activation time, and multimodal response features, while edge attributes include causal strength, propagation delay, and propagation probability. The scene semantic description model is stored in a structured data format, supporting real-time querying and incremental updates. The model update cycle is synchronized with the time window length.
[0029] In one optional implementation, a unified spatiotemporal reference framework is established to align and calibrate the temporal reference and spatial coordinates of different modal data in multimodal sensing data, and a spatiotemporal co-occurrence relationship diagram of multimodal data is constructed under the unified spatiotemporal reference framework, including: The timestamp sequence and spatial coordinate distribution of each modality data are extracted from multimodal sensing data. A cross-correlation function of cross-modal time series and a registration error field of spatial coordinates are constructed. The systematic time delay deviation and periodic drift characteristics of each modality data relative to the reference time axis are identified by the peak position of the cross-correlation function. The rigid transformation parameters and nonlinear distortion parameters of each modality data relative to the reference coordinate system are identified by the gradient distribution of the registration error field. Based on the systematic time delay deviation, the periodic drift characteristics, the rigid transformation parameters and the nonlinear distortion parameters, adaptive spatiotemporal calibration of each modality data is performed to establish a unified spatiotemporal reference framework. Under the unified spatiotemporal reference framework, the calibrated multimodal data is projected onto a multi-scale spatiotemporal grid hierarchy. Local observation manifolds of different modal data are extracted within each scale spatiotemporal grid, and the topological consistency metric and dynamic evolution synchronization metric between the local observation manifolds are calculated. Using a multi-scale spatiotemporal grid as nodes and the weighted fusion value of the topological consistency metric and the dynamic evolution synchronicity metric as edge weights, a spatiotemporal co-occurrence relationship graph of multimodal data is constructed.
[0030] Extracting the timestamp sequences and spatial coordinate distributions of each modality from multimodal sensing data requires parsing and preprocessing the raw data streams acquired by different sensors. Timestamp sequence extraction involves extracting the sampling time information recorded in each modal data stream, organizing the timestamps recorded by the local clocks of each sensor into a time series according to the sampling order. The timestamp accuracy is typically on the order of microseconds to milliseconds. Spatial coordinate distribution extraction involves extracting the spatial location information carried in each modal data. For the visual modality, the spatial coordinates correspond to the 3D projection position of image pixels in the camera coordinate system; for the depth modality, they correspond to the 3D Cartesian coordinates of the point cloud in the sensor coordinate system; for the inertial modality, they correspond to the installation position of the measurement unit in the carrier coordinate system; and for the environmental sensor, they correspond to the deployment position of the node in the scene coordinate system.
[0031] Constructing the cross-correlation function for cross-modal time series requires selecting a reference mode and a mode to be calibrated. The reference mode is typically selected from sensor data with the highest clock stability or the highest sampling frequency. For the timestamp sequences of the reference mode and the mode to be calibrated, the curves of the data features changing over time in both sequences are extracted. The correlation coefficients of the two feature curves at different time offsets are calculated. The time offset scanning range covers the expected maximum time delay difference, and the scanning step size is 0.1 to 1 times the timestamp accuracy. The cross-correlation function records the relationship between the correlation coefficient and the time offset. The time offset corresponding to the peak value of the function is the systematic time delay deviation between the two modes. Constructing the registration error field of spatial coordinates requires identifying the correspondence between different modal data in overlapping spatial regions. Significant feature points or surface fragments are extracted within the common viewing region, a matching relationship is established, and the spatial coordinate differences of the matched pairs in their respective coordinate systems are calculated. The difference vectors are then interpolated in the spatial domain to form the error field.
[0032] When identifying systematic time delay bias and periodic drift characteristics by analyzing the peak position of the cross-correlation function, the systematic time delay bias is directly determined by the time offset corresponding to the global maximum peak of the cross-correlation function. Periodic drift characteristics are identified by analyzing the time intervals and amplitude variations of multiple local peaks in the cross-correlation function. If local peaks appear with an approximately constant period and their amplitudes fluctuate regularly, it indicates that the clock mode to be calibrated has periodic drift. The drift period is equal to the time interval between adjacent local peaks, and the drift amplitude is determined by the fluctuation range of the peak amplitudes. When identifying rigid transformation parameters and nonlinear distortion parameters by analyzing the gradient distribution of the registration error field, the rigid transformation parameters include a 3D rotation matrix and a 3D translation vector. These are solved by minimizing the registration error of the matching feature point pairs, using singular value decomposition or iterative nearest-point algorithms. Nonlinear distortion parameters are obtained by fitting a high-order polynomial or spline function to the registration error field. The fitting order is determined based on the complexity of the error field, typically ranging from order 2 to 5.
[0033] When performing adaptive spatiotemporal calibration of modal data based on systematic time delay bias, periodic drift characteristics, rigid transformation parameters, and nonlinear distortion parameters, time calibration involves subtracting the systematic time delay bias from the original timestamp of the mode to be calibrated to obtain a coarse calibration timestamp. Further dynamic compensation is then performed based on the periodic drift characteristics, with the compensation amount calculated from the phase position and drift amplitude function within the drift period at the current moment. After calibration, the alignment error between the timestamp and the reference time axis is less than twice the timestamp accuracy. Spatial calibration involves first rotating and translating the original spatial coordinates of the mode to be calibrated using rigid transformation parameters, and then performing local deformation correction on the transformed coordinates based on nonlinear distortion parameters. The correction amount is determined by the distortion function value at the coordinate position. After calibration, the registration error between the spatial coordinates and the reference coordinate system is less than 1.5 times the sensor's spatial resolution. After completing spatiotemporal calibration, all modal data are aligned to a unified spatiotemporal reference frame in both the time and spatial dimensions.
[0034] Within a unified spatiotemporal reference framework, when projecting calibrated multimodal data onto a multi-scale spatiotemporal grid hierarchy, the multi-scale spatiotemporal grid divides the 4-dimensional spatiotemporal domain into spatiotemporal units of different granularities. The spatial dimension employs a grid hierarchy of levels 3 to 5, with each level's voxel size being 0.5 times that of the level above. The coarsest voxel size is determined based on the scene's extent, typically ranging from 5 to 20 meters, while the finest voxel size is determined based on the sensor resolution, typically ranging from 0.2 to 1 meter. The temporal dimension employs a window hierarchy of levels 2 to 4, with each window length being 0.3 to 0.5 times that of the level above. The coarsest window length typically ranges from 10 to 60 seconds, while the finest window length typically ranges from 0.5 to 2 seconds. The projection operation assigns each modal data point to its corresponding spatiotemporal grid unit according to its spatiotemporal coordinates.
[0035] When extracting local observation manifolds of different modalities within spatiotemporal grids at various scales, for each spatiotemporal grid cell, the data points of each modality belonging to that cell are collected, and the statistical feature vectors of each modality are extracted. For the visual modality, the first 10 principal components of the color histogram and the first 8 principal components of the edge orientation histogram are extracted; for the depth modality, the mean vector and eigenvalues of the surface normal vector distribution and the covariance matrix are extracted; for the inertial modality, the mean and variance of acceleration and angular velocity and the first 5 peak frequencies of the power spectral density are extracted; and for environmental sensors, the mean, variance, and time derivative of physical quantities such as temperature and pressure are extracted. The statistical feature vectors of each modality are concatenated into a single point in a high-dimensional feature space, which constitutes the embedded representation of the local observation manifold in the feature space.
[0036] When calculating the topological consistency and dynamic evolution synchronicity measures between locally observed manifolds, the topological consistency measure focuses on the geometric proximity and structural similarity of locally observed manifolds in different spatiotemporal grid units within the feature space. For any two locally observed manifolds in spatiotemporal grid units, the Euclidean distance between their embedding points in the feature space is calculated. Manifold pairs with a distance less than a threshold are considered topologically consistent. Further analysis of the local neighborhood structure around the two manifolds is performed by comparing the overlap of their k-nearest neighbor sets, where k ranges from 5 to 15. The overlap is calculated using the Jaccard coefficient, and manifold pairs with an overlap greater than 0.6 exhibit high topological consistency. The dynamic evolution synchronicity measure focuses on the synergy of the evolutionary trajectories of locally observed manifolds in different spatiotemporal grid units over time. For any two spatially adjacent spatiotemporal grid units, the embedding point sequences of their locally observed manifolds within a continuous time window are extracted. The dynamic time-warped distance or cross-correlation coefficient between the two sequences is calculated. Trajectory pairs with a distance less than a threshold or a correlation coefficient greater than 0.7 indicate high synchronicity in the multimodal data evolution of the two units.
[0037] Using multi-scale spatiotemporal grids as nodes and a weighted fusion of topological consistency and dynamic evolution synchronicity measures as edge weights, a spatiotemporal co-occurrence graph for multimodal data is constructed. The graph's node set includes all spatiotemporal grid cells across all scales. Node attributes record the spatial coordinate range, time window boundaries, scale level, and embedded representation of the local observation manifold for each spatiotemporal grid cell. Edges connect node pairs that satisfy spatial adjacency or temporal continuity. Spatial adjacency requires that the spatial volumes corresponding to two nodes share at least one face or edge, while temporal continuity requires that the time windows corresponding to two nodes are either end-to-end or overlap. Edge weights are calculated as a weighted sum of the topological consistency and dynamic evolution synchronicity measures. The weight coefficients are determined based on the application scenario; for scenarios emphasizing spatial propagation, the topological consistency weight is between 0.6 and 0.8, and the dynamic evolution synchronicity weight is between 0.2 and 0.4. The spatiotemporal co-occurrence graph is stored in adjacency list or sparse matrix format, supporting node query, neighborhood traversal, and subgraph extraction operations.
[0038] In one optional implementation, a virtual scene topology is constructed based on the scene semantic description model, and data binding is performed based on the spatial position association between environmental state data in the multimodal perception data and spatial nodes in the virtual scene topology to form a virtual hazardous scene with physical consistency constraints, including: The causal propagation path and spatial distribution boundary of the risk factors are extracted from the scene semantic description model. A directed topological connection between spatial nodes is constructed based on the directionality of the transmission link in the causal propagation path. The influence domain range of each spatial node is determined based on the spatial distribution boundary, forming a virtual scene topology structure that includes the node influence domain and directed transmission constraints. For environmental state data in multimodal perception data, calculate the spatial inclusion relationship between the spatial coordinates of each environmental state data and the influence domain of each spatial node in the virtual scene topology. Bind environmental state data whose spatial coordinates fall within the same influence domain to the corresponding spatial node, and determine the constraint transmission weight of each spatial node's state change on adjacent nodes based on the directed transit constraint. Based on the constraint transmission weights and the evolution logic in the scene semantic description model, physical consistency constraints for state evolution are configured for each spatial node, and a virtual dangerous scene is constructed based on the virtual scene topology and the physical consistency constraints.
[0039] In constructing virtual hazardous scenarios, the causal propagation paths of hazardous factors are extracted by structurally analyzing the scenario semantic description model. These paths describe the evolutionary chain of a hazardous event from its source to its spread, such as the transmission process of gas diffusion and poisoning risk in a chemical scenario caused by a leak. Based on the sequence and direction of influence of each link in this propagation path, corresponding spatial nodes are established in a three-dimensional virtual space, and directed connections are established between nodes according to the directionality of the causal link. Simultaneously, the spatial distribution boundary information of hazardous factors is obtained from the scenario semantic description model. This boundary is determined by measured data from IoT sensing nodes, such as the detection range of toxic gas concentrations or the thermal imaging contours of high-temperature areas. Based on these spatial distribution boundaries, an influence domain is defined for each spatial node. This domain is typically represented by a sphere, ellipsoid, or irregular geometry, with its radius or boundary parameters determined by the statistical characteristics of the actual sensing data.
[0040] After forming a virtual scene topology containing node influence domains and directed transitive constraints, spatial positioning processing is performed on the environmental state data from the multimodal sensing data. Environmental state data includes measured values of physical quantities such as temperature, humidity, gas concentration, and light intensity, with each data point carrying spatial coordinate information at the time of acquisition. By calculating the geometric relationship between these spatial coordinates and the influence domains of each spatial node, it is determined whether a data point is located within the influence domain of a certain node. Specifically, distance calculation from the point to the center of the influence domain or boundary inclusion checks are used. When the spatial coordinates of a certain environmental state data point meet the condition of falling within a certain influence domain, the data is bound to the corresponding spatial node as its state attribute value. For cases where multiple measurement points exist within the same influence domain, a spatial weighted average or maximum value selection strategy is used to comprehensively determine the node state value.
[0041] After data binding is completed, the influence weights of state changes of each spatial node on adjacent nodes are calculated based on the directed transitive constraints in the virtual scene topology. These constraint transitive weights consist of two parts: first, a basic weight determined by the physical mechanisms in the causal propagation path, such as gas diffusion following a concentration gradient law, where the weight decreases exponentially with increasing distance; and second, a correction coefficient obtained by analyzing the temporal correlation of node state changes in historical data, based on statistical analysis of actual perceived data. For spatial nodes with multiple incoming connections, a weighted summation method is used to comprehensively calculate the state evolution trend after being influenced by adjacent nodes.
[0042] To ensure the physical realism of the virtual scene, physical consistency constraints are configured for each spatial node based on the constraint propagation weights and the evolutionary logic recorded in the scene semantic description model. These constraints include limits on the range of state variables, change rate boundaries, time delay parameters for state propagation between nodes, and energy or matter conservation relationships. For example, in a fire scene, the temperature rise rate of temperature nodes must conform to the laws of combustion dynamics, and heat transfer between adjacent nodes must satisfy the constraints of the heat conduction equation. By encoding these constraints into the physical rules of the virtual scene engine, the evolutionary behavior of hazardous factors in the virtual hazardous scene is made consistent with real physical processes, thereby providing training subjects with a high-fidelity simulation experience of hazardous environments.
[0043] In one optional implementation, a behavioral feature sequence is generated based on the interactive operations of the training subjects in the virtual hazardous scenario. Then, based on the hazard factor evolution logic in the scenario semantic description model, a temporal correlation analysis and safety assessment are performed on the behavioral feature sequence to obtain risk response strategies, including: Record the interaction timestamps of the training objects in the virtual dangerous scenario and construct a behavioral feature sequence with the operation object identifier. Analyze the temporal dependency constraints of the dangerous factors in the scenario semantic description model. Timely align the operation time interval in the behavioral feature sequence with the evolution time window in the temporal dependency constraints. Calculate the temporal deviation between the operation timing of the training objects and the evolution time of the dangerous factors. Based on the temporal deviation, the operation segment corresponding to the evolution time of the risk factor is identified from the behavioral feature sequence. The spatial consistency between the operation action direction and the risk factor propagation direction is calculated by comparing the matching relationship between the operation object identifier in the operation segment and the node identifier on the causal propagation path in the scene semantic description model. Based on the temporal deviation and spatial consistency, it is determined whether the operation of the training object meets the requirements for blocking risk factors. For operations that do not meet the blocking requirements, the temporal adjustment amount and spatial redirection target are identified, and a risk response strategy including operation timing correction suggestions and operation object selection suggestions is generated.
[0044] like Figure 2As shown, the method includes: During the operation of the virtual hazardous scenario, interactive sensors mounted on the training subject's operating equipment capture input commands in real time. Each operation is recorded as a data tuple containing an operation timestamp, an operation object identifier, an operation type, and operation parameters. The operation timestamp uses UTC format synchronized with the virtual scenario clock, achieving millisecond-level accuracy to ensure the accuracy of subsequent time-series analysis. The operation object identifier corresponds to a specific interactive entity number in the virtual scenario. The operation type covers categories such as equipment switching, valve adjustment, and tool use. The operation parameters record quantitative information such as force, angle, and duration. The continuously recorded operation data tuples are organized chronologically into a behavioral feature sequence, which reflects the complete operational trajectory of the training subject in the hazardous scenario.
[0045] Temporal dependency constraints of hazard factors are extracted from the scene semantic description model. These constraints define the time window boundaries from the triggering of the hazard factor to its evolution into different stages. The time intervals between adjacent operations in the behavioral feature sequence are extracted and registered with the evolution time windows of the corresponding stages in the temporal dependency constraints. When calculating the temporal deviation, the time difference between the timestamp of the training object's operation and the timestamp of the hazard factor's evolution is measured, using the hazard factor's evolution time as the baseline. The absolute value of the difference characterizes the timeliness of the operation response. A negative difference indicates that the operation is performed prematurely, while a positive difference indicates that the operation is delayed. A difference exceeding a preset threshold is considered a significant temporal deviation.
[0046] Based on the distribution characteristics of temporal deviation, operational fragments with timestamps falling within the neighborhood of critical moments in the evolution of hazard factors are selected from the behavioral feature sequences. The operational object identifiers are extracted from these fragments and matched against the causal propagation paths recorded in the scene semantic description model. The causal propagation paths are stored in a directed graph structure, where nodes represent spatial entities affected or influenced by hazard factors, and edges represent propagation directions. The operational action direction vector is obtained by determining whether the operational object identifier is located at an upstream node, a critical propagation node, or a downstream isolation node in the causal propagation path. The cosine of the angle between this vector and the hazard factor propagation direction vector is calculated as a quantitative indicator of spatial consistency. A cosine value close to 1 indicates a high degree of consistency between the operational direction and the blocking requirement, while a value close to -1 indicates that the operational direction is the same as the hazard propagation direction, thus exacerbating the risk.
[0047] A safety assessment matrix is established by integrating temporal deviation and spatial consistency, setting a temporal tolerance threshold and a spatial consistency lower limit. When the temporal deviation exceeds the tolerance threshold or the spatial consistency falls below the lower limit, the operation is deemed not to meet the requirements for blocking hazardous factors. For temporal discrepancies, the difference between the ideal and actual operation time is calculated, and this difference is output as a temporal adjustment amount, indicating to the trainees the specific amount of time to advance or delay the operation. For spatial discrepancies, key nodes that can effectively cut off the propagation chain are identified in the upstream nodes of the causal propagation path, and the identifier of these nodes is output as spatial redirection targets. The risk response strategy is generated in structured data form, including operation timing correction suggestion fields and operation object selection suggestion fields. The former provides specific time adjustment amounts and directions, while the latter provides recommended operation object identifiers and their spatial coordinates in the virtual scenario. This strategy is fed back to the trainees in real time through a visualization interface, helping them understand the causes of operational deviations and master the correct emergency response timing and operation target selection logic.
[0048] In one optional implementation, based on the temporal deviation, the operation segment corresponding to the evolution time of the hazard factor is identified from the behavioral feature sequence. The spatial consistency between the direction of operation and the direction of hazard factor propagation is calculated by comparing the matching relationship between the operation object identifier in the operation segment and the node identifier on the causal propagation path in the scene semantic description model. Using the time sequence deviation as the time window screening threshold, operations whose timestamps fall within the preset time window range before and after the evolution time of the hazard factor are selected from the behavioral feature sequence. The selected operations are organized into operation segments, and the operation object identifiers of each operation in the operation segment are parsed. The node identifier sequence carrying the hazard propagation on the causal propagation path is parsed from the scene semantic description model, and the spatial position mapping relationship between the operation object identifier and the node identifier sequence is established. Based on the spatial location mapping relationship, the topological position of the operation object corresponding to each operation in the operation segment is identified on the causal propagation path. The direction of change of the topological position of the operation object between adjacent operations on the causal propagation path is analyzed. The vector angle between the direction of change of the topological position and the preset propagation direction of the causal propagation path is calculated. The spatial consistency between the direction of operation and the direction of propagation of the hazard factor is quantified based on the vector angle.
[0049] When performing operation segment identification and spatial consistency calculation, the numerical range of temporal deviation is first obtained, and this range is used as the width parameter of the time window. Specifically, when the temporal deviation is 0.3 seconds, it is extended by 0.3 seconds before and after the evolution of the hazard factor, forming a time window with a total width of 0.6 seconds. All operation records in the behavioral feature sequence are traversed, and the timestamp attribute of each operation record is extracted to determine whether the timestamp falls within the aforementioned time window range. The operation records that satisfy the time window constraints are arranged in ascending order of timestamp to form operation segments.
[0050] Each operation record in the operation segment is parsed to extract the operation object identifier field. The operation object identifier adopts the Uniform Resource Identifier (URI) format and contains a unique number of the spatial node in the virtual scene topology. The storage structure of the causal propagation path is read from the scene semantic description model. This structure is organized in the form of a directed graph, and each node in the graph carries a spatial node number attribute. All nodes in the causal propagation path are traversed, and the node numbers are extracted according to the connection order of the directed edges to form a node identifier sequence.
[0051] A mapping table is established between operation object identifiers and node identifier sequences. Each operation object identifier in an operation segment is matched against the node identifier sequence. When an operation object identifier matches a node number in the node identifier sequence, the index position of that node in the causal propagation path is recorded; this index position is the topological position. The operations in the operation segment are arranged in timestamp order. The topological position values of two adjacent operations are extracted, and the difference between the topological position of the later operation and the previous operation is calculated to obtain the change in topological position.
[0052] The topological position change is converted into a vector representation. In the directed graph structure of the causal propagation path, the direction of index increment is defined as the positive propagation direction. When the topological position change is greater than zero, the operation object moves forward along the causal propagation path, and the vector direction is a positive unit vector; when the topological position change is less than zero, the operation object moves backward along the causal propagation path, and the vector direction is a negative unit vector. The preset propagation direction vector of the causal propagation path is read from the scene semantic description model, and this vector points along the edge direction in the directed graph structure.
[0053] The angle between the direction vector of the operation and the direction vector of hazard propagation is calculated, and the cosine value of the angle is obtained by vector dot product operation. When the cosine value is close to 1, the two vectors are in the same direction; when the cosine value is close to negative 1, the two vectors are in opposite directions. The absolute value of the cosine value is used as a quantitative index of spatial consistency, with a value between 0 and 1. The spatial consistency values calculated for all adjacent operations in the operation segment are weighted and averaged. The weight coefficient is proportional to the proximity of the timestamps of the operations within the time window; the closer the operation is to the moment of hazard evolution, the higher the weight, thus obtaining the overall spatial consistency evaluation value of the operation segment.
[0054] In one optional implementation, the risk response strategy is applied to the virtual scene topology to trigger the dynamic evolution of risk factors and generate feedback signals. The feedback signals are then used to adaptively correct the evolution logic in the scene semantic description model, including: The operational suggestions in the risk response strategy are transformed into state intervention instructions in the virtual scene. The state intervention instructions are applied to the target space nodes of the virtual scene topology. The state changes are transmitted through the topological connection relationship in the virtual scene topology, triggering the dynamic evolution of the risk factors along the causal propagation path. The actual evolution trajectory of the risk factors is recorded. The actual evolution trajectory is compared with the preset evolution trajectory in the scene semantic description model to identify the trajectory deviation position and deviation magnitude and generate feedback signals. The trajectory deviation position in the feedback signal is analyzed to locate the logical defects in the corresponding temporal dependency constraints or causal propagation paths in the scene semantic description model. The deviation magnitude in the feedback signal is analyzed to quantify the impact of the logical defects on evolution prediction. Based on the location result and impact of the logical defects, the evolution logic in the scene semantic description model is adaptively corrected using the feedback signal.
[0055] After obtaining the risk response strategy, a state intervention mechanism is needed to realize the dynamic evolution and feedback correction of the virtual scenario. The risk response strategy includes operational suggestions for unsafe behaviors of the training subjects, such as specific actions like "closing the gas valve," "starting emergency ventilation," and "evacuating the danger zone." An operational suggestion parser transforms these natural language suggestions into executable state intervention instructions. These intervention instructions use structured encoding, containing five fields: intervention type identifier, target node spatial coordinates, state parameter change, execution priority, and timestamp. For example, "closing the gas valve" is transformed into an instruction structure with intervention type 0x01, target node N127, state parameter (flow rate = 0), priority P1, and timestamp T0.
[0056] When a state intervention command is applied to a target spatial node in the virtual scene topology, the node's attribute parameters are modified through a node state update mechanism. The virtual scene topology is stored using a directed graph data structure, with nodes connected by edges representing physical relationships or causal propagation relationships. When the state of a target node changes, the state change is propagated to adjacent nodes along the topological connections based on the edge weights and propagation delay parameters. The propagation process follows the constraints of energy and matter conservation; for example, when a pipeline node is closed, the upstream pressure node value accumulates, while the downstream flow node value decays. Hazardous factors dynamically evolve along predefined causal propagation paths, which are determined by logical rules in the scene semantic description model. For example, "combustible gas leak → air mixing → reaching the explosion limit → encountering an ignition source → explosion" constitutes a complete propagation chain.
[0057] During the evolution process, the state parameters of hazard factors are recorded in real time at each time step, including four dimensions: hazard type, spatial location, intensity value, and radius of influence, forming an actual evolution trajectory dataset. This dataset is compared point-by-point with the preset evolution trajectories stored in the scene semantic description model. The preset trajectories are trained based on historical IoT sensing data and represent the evolutionary patterns of typical hazardous scenarios. The comparison process uses a dynamic time warping algorithm to match two trajectories. When the difference in state parameters at a certain time point exceeds a threshold Δ_th, it is marked as a trajectory deviation position, and the corresponding timestamp and spatial node identifier are recorded. The deviation magnitude is defined as the Euclidean distance between the actual state vector and the preset state vector, used to quantify the severity of the deviation. The trajectory deviation position, deviation magnitude, and associated node identifier are encapsulated into a feedback signal structure.
[0058] The deviation position of the feedback signal trajectory includes a timestamp and a spatial node identifier. The timestamp is used to inversely index the temporal dependency constraints at the corresponding moment in the scene semantic description model, while the spatial node identifier is used to locate the relevant causal propagation path. If the deviation occurs in the early stages of evolution, it usually points to a flaw in the initial condition judgment rule; if the deviation occurs in the middle stages, it indicates an unreasonable setting of the state transition probability; if the deviation occurs in the final stages, it indicates that the dangerous consequence prediction model is inaccurate. For each logical defect in the location, a correction weight is assigned based on the magnitude of the deviation; the larger the deviation, the higher the correction weight.
[0059] The adaptive correction process employs an incremental learning mechanism, using the actual evolutionary trajectories in the feedback signals as new training samples to update the evolutionary logic parameters in the scene semantic description model. Corrections are made to the temporal dependency constraints by adjusting the state trigger threshold, to the causal propagation path by adjusting the inter-node propagation coefficients, and to the dangerous consequence prediction by adjusting the result mapping function. The correction process preserves the basic structure of the historical evolutionary logic, only fine-tuning local parameters that deviate from the relevant parameters, thus avoiding global model instability. After multiple training iterations, the evolutionary prediction accuracy of the scene semantic description model gradually improves.
[0060] A second aspect of the present invention provides a multimodal IoT-driven virtual training system for hazardous scenarios, comprising: The perception data unit is used to collect multimodal perception data of the target hazardous scene through distributed IoT perception nodes, and to perform cross-modal correlation mining and hazard factor identification on the multimodal perception data based on spatiotemporal alignment rules to generate a scene semantic description model. The virtual scene unit is used to construct a virtual scene topology based on the scene semantic description model, and to bind data according to the spatial position association between the environmental state data in the multimodal perception data and the spatial nodes in the virtual scene topology to form a virtual dangerous scene with physical consistency constraints. The behavioral safety unit is used to generate a sequence of behavioral features based on the interactive operations of the training object in the virtual dangerous scenario, and to perform temporal correlation analysis and safety assessment on the sequence of behavioral features based on the evolution logic of the danger factors in the scenario semantic description model, so as to obtain a risk response strategy. The dynamic evolution unit is used to apply the risk response strategy to the virtual scene topology, trigger the dynamic evolution of risk factors and generate feedback signals, and use the feedback signals to adaptively correct the evolution logic in the scene semantic description model.
[0061] A third aspect of the present invention provides an electronic device, comprising: processor; Memory used to store processor-executable instructions; The processor is configured to invoke instructions stored in the memory to execute the aforementioned method.
[0062] A fourth aspect of the present invention provides a computer-readable storage medium having stored thereon computer program instructions that, when executed by a processor, implement the aforementioned method.
[0063] This invention can be a method, apparatus, system, and / or computer program product. The computer program product may include a computer-readable storage medium having computer-readable program instructions loaded thereon for performing various aspects of the invention.
[0064] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, and not to limit them; although the present invention has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some or all of the technical features; and these modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the scope of the technical solutions of the embodiments of the present invention.
Claims
1. A virtual training method for hazardous scenarios driven by multimodal IoT perception, characterized in that, include: Multimodal sensing data of the target hazardous scene is collected by distributed IoT sensing nodes. Based on the spatiotemporal alignment rules, cross-modal correlation mining and hazard factor identification are performed on the multimodal sensing data to generate a scene semantic description model. A virtual scene topology is constructed based on the scene semantic description model, and data binding is performed based on the spatial position association between the environmental state data in the multimodal perception data and the spatial nodes in the virtual scene topology to form a virtual dangerous scene with physical consistency constraints. Based on the interactive operations of the training subjects in the virtual dangerous scenario, a behavioral feature sequence is generated. Then, based on the evolution logic of the risk factors in the scenario semantic description model, a temporal correlation analysis and safety assessment are performed on the behavioral feature sequence to obtain a risk response strategy. The risk response strategy is applied to the virtual scene topology to trigger the dynamic evolution of risk factors and generate feedback signals. The feedback signals are then used to adaptively correct the evolution logic in the scene semantic description model.
2. The method according to claim 1, characterized in that, Based on spatiotemporal alignment rules, cross-modal correlation mining and risk factor identification are performed on the multimodal sensing data to generate a scene semantic description model, including: A unified spatiotemporal reference framework is established to align and calibrate the temporal reference and spatial coordinates of different modal data in multimodal sensing data, and a spatiotemporal co-occurrence relationship diagram of multimodal data is constructed under the unified spatiotemporal reference framework. Based on the response intensity distribution in the spatiotemporal co-occurrence relationship diagram, the co-evolution mode of cross-modal data within a specific spatiotemporal window is identified. By analyzing the similarity and deviation between the co-evolution mode and the evolution law of known dangerous scenarios, the activation state and spatial diffusion direction of potential dangerous factors in the current scenario are determined. Based on the activation state and the spatial diffusion direction, the temporal dependency constraints and causal propagation paths of the hazard factors are extracted. Based on the temporal dependency constraints and the causal propagation paths, a scene semantic description model is constructed, which includes hazard factor identification, spatial distribution boundaries, evolutionary stage division, and propagation link topology.
3. The method according to claim 2, characterized in that, A unified spatiotemporal reference framework is established to align and calibrate the temporal reference and spatial coordinates of different modal data in multimodal sensing data. A spatiotemporal co-occurrence relationship diagram of the multimodal data is then constructed within this unified spatiotemporal reference framework, including: The timestamp sequence and spatial coordinate distribution of each modality data are extracted from multimodal sensing data. A cross-correlation function of cross-modal time series and a registration error field of spatial coordinates are constructed. The systematic time delay deviation and periodic drift characteristics of each modality data relative to the reference time axis are identified by the peak position of the cross-correlation function. The rigid transformation parameters and nonlinear distortion parameters of each modality data relative to the reference coordinate system are identified by the gradient distribution of the registration error field. Based on the systematic time delay deviation, the periodic drift characteristics, the rigid transformation parameters and the nonlinear distortion parameters, adaptive spatiotemporal calibration of each modality data is performed to establish a unified spatiotemporal reference framework. Under the unified spatiotemporal reference framework, the calibrated multimodal data is projected onto a multi-scale spatiotemporal grid hierarchy. Local observation manifolds of different modal data are extracted within each scale spatiotemporal grid, and the topological consistency metric and dynamic evolution synchronization metric between the local observation manifolds are calculated. Using a multi-scale spatiotemporal grid as nodes and the weighted fusion value of the topological consistency metric and the dynamic evolution synchronicity metric as edge weights, a spatiotemporal co-occurrence relationship graph of multimodal data is constructed.
4. The method according to claim 1, characterized in that, A virtual scene topology is constructed based on the scene semantic description model, and data binding is performed based on the spatial position association between the environmental state data in the multimodal perception data and the spatial nodes in the virtual scene topology to form a virtual hazardous scene with physical consistency constraints, including: The causal propagation path and spatial distribution boundary of the risk factors are extracted from the scene semantic description model. A directed topological connection between spatial nodes is constructed based on the directionality of the transmission link in the causal propagation path. The influence domain range of each spatial node is determined based on the spatial distribution boundary, forming a virtual scene topology structure that includes the node influence domain and directed transmission constraints. For environmental state data in multimodal perception data, calculate the spatial inclusion relationship between the spatial coordinates of each environmental state data and the influence domain of each spatial node in the virtual scene topology. Bind environmental state data whose spatial coordinates fall within the same influence domain to the corresponding spatial node, and determine the constraint transmission weight of each spatial node's state change on adjacent nodes based on the directed transit constraint. Based on the constraint transmission weights and the evolution logic in the scene semantic description model, physical consistency constraints for state evolution are configured for each spatial node, and a virtual dangerous scene is constructed based on the virtual scene topology and the physical consistency constraints.
5. The method according to claim 1, characterized in that, Based on the interactive operations of the training subjects in the virtual dangerous scenario, a behavioral feature sequence is generated. Then, based on the hazard factor evolution logic in the scenario semantic description model, a temporal correlation analysis and safety assessment are performed on the behavioral feature sequence to obtain risk response strategies, including: Record the interaction timestamps of the training objects in the virtual dangerous scenario and construct a behavioral feature sequence with the operation object identifier. Analyze the temporal dependency constraints of the dangerous factors in the scenario semantic description model. Timely align the operation time interval in the behavioral feature sequence with the evolution time window in the temporal dependency constraints. Calculate the temporal deviation between the operation timing of the training objects and the evolution time of the dangerous factors. Based on the temporal deviation, the operation segment corresponding to the evolution time of the risk factor is identified from the behavioral feature sequence. The spatial consistency between the operation action direction and the risk factor propagation direction is calculated by comparing the matching relationship between the operation object identifier in the operation segment and the node identifier on the causal propagation path in the scene semantic description model. Based on the temporal deviation and spatial consistency, it is determined whether the operation of the training object meets the requirements for blocking risk factors. For operations that do not meet the blocking requirements, the temporal adjustment amount and spatial redirection target are identified, and a risk response strategy including operation timing correction suggestions and operation object selection suggestions is generated.
6. The method according to claim 5, characterized in that, Based on the temporal deviation, operational segments corresponding to the evolution time of risk factors are identified from the behavioral feature sequence. The spatial consistency between the direction of operational action and the direction of risk factor propagation is calculated by comparing the matching relationship between the operational object identifier in the operational segment and the node identifiers on the causal propagation path in the scene semantic description model. This includes: Using the time sequence deviation as the time window screening threshold, operations whose timestamps fall within the preset time window range before and after the evolution time of the hazard factor are selected from the behavioral feature sequence. The selected operations are organized into operation segments, and the operation object identifiers of each operation in the operation segment are parsed. The node identifier sequence carrying the hazard propagation on the causal propagation path is parsed from the scene semantic description model, and the spatial position mapping relationship between the operation object identifier and the node identifier sequence is established. Based on the spatial location mapping relationship, the topological position of the operation object corresponding to each operation in the operation segment is identified on the causal propagation path. The direction of change of the topological position of the operation object between adjacent operations on the causal propagation path is analyzed. The vector angle between the direction of change of the topological position and the preset propagation direction of the causal propagation path is calculated. The spatial consistency between the direction of operation and the direction of propagation of the hazard factor is quantified based on the vector angle.
7. The method according to claim 1, characterized in that, Applying the risk response strategy to the virtual scene topology triggers the dynamic evolution of risk factors and generates feedback signals. Adaptively correcting the evolution logic in the scene semantic description model using these feedback signals includes: The operational suggestions in the risk response strategy are transformed into state intervention instructions in the virtual scene. The state intervention instructions are applied to the target space nodes of the virtual scene topology. The state changes are transmitted through the topological connection relationship in the virtual scene topology, triggering the dynamic evolution of the risk factors along the causal propagation path. The actual evolution trajectory of the risk factors is recorded. The actual evolution trajectory is compared with the preset evolution trajectory in the scene semantic description model to identify the trajectory deviation position and deviation magnitude and generate feedback signals. The trajectory deviation position in the feedback signal is analyzed to locate the logical defects in the corresponding temporal dependency constraints or causal propagation paths in the scene semantic description model. The deviation magnitude in the feedback signal is analyzed to quantify the impact of the logical defects on evolution prediction. Based on the location result and impact of the logical defects, the evolution logic in the scene semantic description model is adaptively corrected using the feedback signal.
8. A multimodal IoT-driven virtual training system for hazardous scenarios, used to implement the method of any one of claims 1-7, characterized in that, include: The perception data unit is used to collect multimodal perception data of the target hazardous scene through distributed IoT perception nodes, and to perform cross-modal correlation mining and hazard factor identification on the multimodal perception data based on spatiotemporal alignment rules to generate a scene semantic description model. The virtual scene unit is used to construct a virtual scene topology based on the scene semantic description model, and to bind data according to the spatial position association between the environmental state data in the multimodal perception data and the spatial nodes in the virtual scene topology to form a virtual dangerous scene with physical consistency constraints. The behavioral safety unit is used to generate a sequence of behavioral features based on the interactive operations of the training object in the virtual dangerous scenario, and to perform temporal correlation analysis and safety assessment on the sequence of behavioral features based on the evolution logic of the danger factors in the scenario semantic description model, so as to obtain a risk response strategy. The dynamic evolution unit is used to apply the risk response strategy to the virtual scene topology, trigger the dynamic evolution of risk factors and generate feedback signals, and use the feedback signals to adaptively correct the evolution logic in the scene semantic description model.
9. An electronic device, characterized in that, include: processor; Memory used to store processor-executable instructions; The processor is configured to invoke instructions stored in the memory to execute the method according to any one of claims 1 to 7.
10. A computer-readable storage medium having computer program instructions stored thereon, characterized in that, When the computer program instructions are executed by the processor, they implement the method described in any one of claims 1 to 7.