Coal Mine Disaster Risk Prevention and Control Platform Based on Multimodal Perception and AI Video Recognition
The coal mine disaster risk prevention and control platform, which combines multimodal perception and AI video recognition, solves the problem of misjudging minor abnormalities in underground scenarios in existing technologies, and achieves high-precision, real-time, and proactive prevention and control of coal mine disasters.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- JINGHANG IND TECHNOLOGY (SHANDONG) CO LTD
- Filing Date
- 2026-06-08
- Publication Date
- 2026-07-10
AI Technical Summary
Existing technologies often misjudge minor, highly obscuring anomalies that are prone to occur in underground coal mines as water mist, shadows, or normal thermal disturbances of equipment. There is a lack of means to continuously determine and coordinate the prevention and control of real micro-disasters after the pseudo-anomalies are stripped away, resulting in insufficient risk identification capabilities.
The coal mine disaster risk prevention and control platform based on multimodal perception and AI video recognition constructs a cross-modal causal consistency matrix through multi-source data collection, micro-scene construction, feature extraction, and pseudo-anomaly removal modules. It generates non-disaster disturbance baselines and removes pseudo-anomalies, outputs the location, propagation direction, and risk level of risk sources, and performs operations such as graded power outages, local ventilation adjustments, and personnel evacuation.
It significantly reduced the false alarm rate and the missed alarm rate, and improved the accuracy and stability of identifying small-scale and atypical risks such as localized gas accumulation, early overheating of idler rollers, gas release from cracks in the top plate, and personnel obstruction and stagnation, thus realizing real-time proactive intervention and adaptive prevention and control of risks.
Smart Images

Figure CN122365401A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of video recognition technology, specifically to a coal mine disaster risk prevention and control platform based on multimodal perception and AI video recognition. Background Technology
[0002] In coal mines, at return air corners, inside transfer points, and in the maintenance areas of tunneling face equipment, minute phenomena such as swaying dust curtains, reflective spray droplets, drifting coal dust clouds, and localized eddies are prone to occur. Small amounts of gas release from roof cracks, early frictional heat spots on idler rollers, or personnel being blocked and trapped are often misjudged by existing single-sensor monitoring or ordinary video algorithms as water mist, shadows, or normal thermal disturbances of equipment. These anomalies are small in scope, short in duration, remote in location, and uncommon. However, once coupled with localized gas accumulation or restricted ventilation, they can easily lead to sudden combustion, explosion, suffocation, or roof collapse. Current technology lacks the means to continuously determine and coordinate the prevention and control of real micro-hazards after the pseudo-anomalies have been stripped away. Summary of the Invention
[0003] The purpose of this invention is to provide a coal mine disaster risk prevention and control platform based on multimodal perception and AI video recognition, so as to solve the shortcomings of the prior art.
[0004] To achieve the above objectives, the present invention provides the following technical solution: a coal mine disaster risk prevention and control platform based on multimodal perception and AI video recognition, comprising:
[0005] Multi-source data acquisition module: acquires video stream V, gas concentration G, wind speed W, temperature and humidity T, acoustic and vibration signals A, and roadway topology parameters L of the target area;
[0006] Micro-scene construction module: Based on L, reflectivity suppression, sway occlusion separation and low light compensation are applied to V to construct micro-scene unit U and obtain visual perturbation feature Fv;
[0007] Feature extraction module: Asynchronously registers G, W, T, and A according to U to extract gas response feature Fg, thermal inertia feature Ft, and acoustic vibration abrupt change feature Fa;
[0008] False anomaly removal module: Construct a cross-modal causal consistency matrix C based on Fv, Fg, Ft and Fa, and generate the corresponding non-hazardous disturbance baseline B to obtain the residual risk characteristics R after removing false anomalies;
[0009] Linkage control module: Inputs R and L into the risk evolution model, outputs the risk source location, propagation direction and risk level result Y, and executes graded power outage, local ventilation adjustment, directional spraying, personnel evacuation and model self-updating based on Y.
[0010] Preferably, constructing a cross-modal causal consistency matrix C based on Fv, Fg, Ft, and Fa includes the following steps:
[0011] Using the micro-scene unit U as the index, Fv, Fg, Ft and Fa are segmented according to a unified sliding time window, and modal association candidate set M is generated by combining the airflow direction, sensor spacing and equipment position in the roadway topology parameters L;
[0012] Perform topologically constrained time-delay matching on each candidate feature pair in M, calculate the sequential relationship between visual disturbance leading, gas response lagging, thermal inertia persistence and acoustic vibration sudden change, and obtain the time-series coupling parameter D;
[0013] Based on D-screening, pseudo-association feature pairs that do not meet the propagation path, response delay threshold, and co-occurrence duration threshold are eliminated, and the effective causal edge set E with visual anomaly-gas / thermal / acoustic vibration response chain transmission relationship is retained.
[0014] A cross-modal causal consistency matrix C is generated by weighting the direction, intensity, and duration of E.
[0015] Preferably, performing topologically constrained time-delay matching on each candidate feature pair in M includes the following steps:
[0016] Based on the airflow direction, the distance between the sensors at both ends of the candidate feature pairs, and the equipment occlusion relationship in the roadway topology parameters L, a path constraint set P is constructed for each candidate feature pair, and a corresponding dynamic time delay search window Δt is generated according to the path length and wind speed W.
[0017] Within Δt, asymmetric sliding matching is performed on candidate feature pairs, and event anchoring is implemented for visual disturbance peaks, gas concentration inflection points, the starting point of temperature slow rise segments, and acoustic vibration pulse peaks to obtain time-stamped sequences Q for each modal event.
[0018] Based on Q-computation, visual lead time difference, gas lag time difference, thermal inertia duration and acoustic vibration instantaneous offset are calculated, and abnormal matches that do not meet the topological reachability, time difference upper limit and duration threshold are filtered out to obtain the effective temporal relation set H;
[0019] H is sequentially encoded and weighted by confidence to generate the temporal coupling parameter D.
[0020] Preferably, based on the temporal coupling parameter D, and combined with the airflow channels, bifurcation connections, and equipment obstruction locations in the roadway topology parameter L, a propagation path verification map G for candidate feature pairs is constructed to determine whether each candidate feature pair has an achievable propagation chain from upstream to downstream or from the risk source to the monitoring point. For candidate feature pairs that pass the propagation path verification, the visual lead time difference, gas lag time difference, thermal inertia duration, and acoustic vibration instantaneous offset in D are called up and compared with the preset response delay threshold intervals, respectively, to eliminate abnormal feature pairs with inverted timing, excessive lag, or instantaneous mismatch. For feature pairs after time delay screening, the co-occurrence time period of the corresponding abnormal event is extracted, and the overlap duration between the visual anomaly maintenance interval and the gas response interval, thermal duration interval, or acoustic vibration sudden interval is calculated to screen out short-term occasional coupling pairs that do not meet the minimum co-occurrence duration threshold. The remaining feature pairs are weighted and scored according to propagation reachability, time delay matching degree, and co-occurrence stability, and the effective causal edge set E is output.
[0021] Preferably, generating the corresponding non-hazard disturbance baseline B includes: based on the temporal coupling parameter D, combined with the roadway topology parameter L, equipment operating status, and spray start / stop records, extracting a reference sample set N from historical no-alarm periods that is consistent with the topological position and operating conditions of the current micro-scene unit U; clustering the disturbance sources of Fv, Fg, Ft, and Fa in N, and constructing non-hazard disturbance prototype vectors Z for spray reflection, dust curtain swaying, equipment idling vibration, and airflow fluctuations respectively; calculating the visual brightness amplitude range, gas fluctuation upper limit, thermal slow rise slope upper limit, and acoustic frequency band tolerance corresponding to each disturbance prototype based on Z, forming disturbance envelope parameters constrained by topological position and operating conditions; indexing and encoding the disturbance envelope parameters according to the micro-scene unit U to generate a non-hazard disturbance baseline B that corresponds one-to-one with the current scene.
[0022] Preferably, based on the current micro-scene unit U, the corresponding non-disaster disturbance baseline B is invoked, and Fv, Fg, Ft, and Fa are respectively mapped to the visual disturbance envelope, gas fluctuation envelope, thermal inertia envelope, and acoustic vibration tolerance envelope in B to obtain the baseline deviation set Δ for each mode. Combined with the effective causal edge set E, causal constraint residual decomposition is performed on each mode deviation in Δ, retaining the coupled deviation components that satisfy the chain relationship of subsequent gas response, thermal persistence, or acoustic vibration burst caused by visual anomalies, to obtain the candidate residual set Rc. Pseudo-anomaly stripping is performed on Rc, deducting the homologous components corresponding to spray reflection, dust curtain sway, equipment idling vibration, and short-term airflow fluctuation, and filtering out transient residuals with durations lower than the minimum risk maintenance threshold to obtain the net residual set Rn. Rn is weighted and encoded according to modal contribution, temporal consistency, and spatial clustering to generate the residual risk feature R after pseudo-anomaly stripping.
[0023] Preferably, the deviation set Δ is subjected to directed correlation expansion using the initial mode, response mode, and edge weight parameters in the effective causal edge set E as constraints, constructing a visual deviation-gas deviation-thermal deviation / acoustic vibration deviation residual transfer subgraph Gd, and marking isolated deviations not on the constraint path of E as components to be removed; based on the time delay parameters, intensity weights, and duration weights corresponding to each directed edge in Gd, the adjacent modal deviations within Δ are subjected to piecewise projection, separating the causal response components that satisfy the sequential triggering relationship and the synchronization disturbances that do not satisfy the triggering order. The dynamic component is used to obtain the initial residual pair Rc1; the coupling consistency of Rc1 is checked according to the spatial adjacency relationship and temporal overlap interval within the micro-scene unit U, and the residual chains that are continuously transmitted in the same local area and whose cross-modal overlap duration reaches the threshold are retained to obtain the candidate coupling residual set Rc2; chain-based cumulative encoding is performed on Rc2, and the visual driving quantity, gas diffusion quantity, thermal retention quantity and acoustic vibration mutation quantity in each residual chain are superimposed according to the edge weight, and the coupling deviation component that satisfies the causal constraint is output as the candidate residual set Rc for subsequent pseudo-anomaly stripping.
[0024] Preferably, the coupling consistency verification of Rc1 based on the spatial adjacency relationship and temporal overlap interval within the micro-scene unit U includes the following steps: mapping the landing points of each residual pair in Rc1 according to the spatial grid position of the micro-scene unit U, establishing an adjacent grid connectivity table based on the tunnel topology parameter L, and screening out candidate residual groups with spatial distances less than a preset adjacency threshold; extracting the start and end timestamps of the corresponding abnormal events for each candidate residual group, calculating the temporal overlap interval and overlap ratio between visual deviation, gas deviation, thermal deviation, or acoustic vibration deviation, and obtaining the spatiotemporal coupling parameters; performing consistency judgment on each candidate residual group based on the spatiotemporal coupling parameters, eliminating pseudo-coupled residual groups that are spatially adjacent but have excessive temporal misalignment, or that overlap temporally but cross non-connected boundaries, and retaining continuous transmission chains; jointly weighting the retained continuous transmission chains according to spatial adjacency strength and temporal overlap stability to generate a candidate coupling residual set Rc2 that satisfies the local co-domain and simultaneous coupling conditions.
[0025] Preferably, the linkage control module divides the target area into spatial nodes corresponding to the micro-scene unit U based on the roadway topology parameter L, and establishes directed propagation edges based on the roadway connectivity, airflow direction and equipment occlusion relationship between the nodes to form a risk evolution diagram; after mapping R to the corresponding nodes, the earliest appearing node that continuously exceeds the minimum risk maintenance threshold is identified in chronological order as the candidate node for the risk source location, and the source point confidence score is calculated to determine the risk source location.
[0026] Preferably, the linkage control module starts from the location of the risk source, searches for candidate paths of propagation direction along the directed propagation edge to adjacent nodes, and determines the propagation direction based on time-progressive scoring, intensity attenuation scoring, topology legality scoring, and spatial continuity scoring; after the location of the risk source and the propagation direction are determined, the risk level result Y is output, which includes the location of the risk source, the propagation direction, and the risk level.
[0027] The technical effects and advantages provided by the present invention in the above technical solution are as follows:
[0028] 1. This invention constructs a micro-scene unit U corresponding to a local high-risk area underground by synergistically fusing video stream V, gas concentration G, wind speed W, temperature and humidity T, acoustic and vibration signals A, and roadway topology parameters L. Furthermore, it extracts visual disturbance features Fv, gas response features Fg, thermal inertia features Ft, and acoustic and vibration mutation features Fa, thus overcoming the problem of insufficient ability to identify minor, rare, and highly obstructed disaster precursors in existing technologies that rely on single sensors or ordinary video recognition methods. In particular, by constructing a cross-modal causal consistency matrix C, generating a non-disaster disturbance baseline B matching the current scene, and obtaining residual risk features R after stripping away false anomalies, it can effectively distinguish the differences between non-disaster disturbances such as spray reflection, dust curtain swaying, equipment idling vibration, and short-term airflow fluctuations and real disaster precursors. This significantly reduces false alarm and false negative rates, and improves the accuracy and stability of identifying small-scale, atypical risks such as localized gas accumulation, early overheating of idler rollers, gas release from roof cracks, and personnel obstruction and stagnation.
[0029] 2. This invention inputs the residual risk characteristics R (after removing pseudo-anomalies) and roadway topology parameters L into the risk evolution model, enabling it to output the risk source location, propagation direction, and risk level result Y. Based on this, it executes graded power outages, local ventilation adjustments, directional spraying, personnel evacuation, and model self-updating, thus elevating risk prevention and control from traditional post-event alarm methods to proactive intervention methods oriented towards the risk evolution process. Compared to existing technologies, this invention not only combines underground spatial structure, airflow paths, and equipment distribution to achieve more targeted coordinated responses, avoiding large-scale false shutdowns and ineffective responses, but also continuously calibrates and optimizes the risk evolution model through the response results, giving the platform self-learning and adaptive capabilities, thereby improving the real-time performance, accuracy, and engineering applicability of coal mine disaster risk prevention and control. Attached Figure Description
[0030] To more clearly illustrate the technical solutions in the embodiments of this application or the prior art, the drawings used in the embodiments will be briefly introduced below. Obviously, the drawings described below are only some embodiments recorded in this invention. For those skilled in the art, other drawings can be obtained based on these drawings.
[0031] Figure 1 This is a flowchart of the modules of the coal mine disaster risk prevention and control platform based on multimodal perception and AI video recognition of the present invention.
[0032] Figure 2 This is a flowchart of the method for constructing the cross-modal causal consistency matrix of the present invention. Detailed Implementation
[0033] To make the objectives, technical solutions, and advantages of the embodiments of the present invention clearer, the technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.
[0034] For examples, please refer to Figure 1 and Figure 2 As shown in this embodiment, the coal mine disaster risk prevention and control platform based on multimodal perception and AI video recognition includes:
[0035] In one embodiment of the present invention, a multi-source data acquisition module is used to acquire video stream V, gas concentration G, wind speed W, temperature and humidity T, acoustic and vibration signals A, and roadway topology parameters L of the target area, and to form the basic input data required for subsequent risk identification. The target area is a local space in a coal mine where disasters are prone to occur or are difficult to identify, preferably at least one of the following: a tunneling face, a transfer point, a return air corner, a connecting roadway bend, an equipment maintenance area, and an area with local ventilation changes.
[0036] Specifically, video stream V is continuously acquired by industrial cameras positioned at the top, side, or near the equipment in the target area. These industrial cameras preferably possess low-light imaging capabilities and are dustproof, fogproof, and explosion-proof. They are used to acquire video data corresponding to personnel activities, equipment operation, coal dust disturbance, water mist reflection, smoke changes, and obstruction states within the target area. Gas concentration G is acquired by gas sensors deployed at the upper corner of the target area, the return air side, above the equipment, or at local gas accumulation locations. This data characterizes the local gas accumulation, release, or diffusion state. Wind speed W is acquired by a wind speed sensor, reflecting the airflow direction, strength, and eddy current changes within the target area. Temperature and humidity T are acquired by a temperature and humidity sensor, characterizing the environmental heat and humidity conditions and the impact of spray, water accumulation, and heat source disturbances on the environment. Acoustic and vibration signals A are acquired by acoustic sensors and / or vibration sensors, characterizing roller friction, abnormal equipment noise, micro-cracks in the roof, impact vibration, and abnormal mechanical disturbances.
[0037] The roadway topology parameter L is used to characterize the spatial structure relationship of the roadway in the target area, and includes at least one or more of the following: roadway direction, cross-sectional dimensions, turning positions, bifurcation relationships, equipment layout positions, sensor installation positions, ventilation structure positions, and connectivity relationships between adjacent areas. The roadway topology parameter L can be obtained from existing three-dimensional geodetic models of the mine, roadway design drawings, laser scanning results, inertial navigation mapping results, or manual calibration results.
[0038] To ensure the accuracy of subsequent multimodal data fusion, the multi-source data acquisition module can also uniformly timestamp video stream V, gas concentration G, wind speed W, temperature and humidity T, and acoustic vibration signal A, and establish a spatial mapping relationship based on sensor installation coordinates and roadway topology parameters L, so as to uniformly associate data from different sources with the same target area or the same local spatial unit. Preferably, the multi-source data acquisition module continuously acquires data according to a preset sampling period, and performs buffering, compensation, or removal processing on missing data, abnormal jump data, and communication delay data, thereby outputting a structured multi-source raw dataset, providing a data foundation for subsequent asynchronous spatiotemporal registration, feature extraction, and risk evolution analysis based on micro-scene units.
[0039] In one embodiment of the present invention, a micro-scene construction module is used to perform reflection suppression, sway occlusion separation, and low-light compensation on video stream V based on the roadway topology parameters L, constructing micro-scene units U to obtain visual disturbance features Fv, providing a visual basis for subsequent multimodal spatiotemporal registration and risk assessment. The roadway topology parameters L include at least one or more of the following: roadway direction, corner positions, cross-sectional dimensions, equipment layout positions, ventilation duct positions, dust curtain positions, conveyor belt positions, support component positions, and camera installation positions. The micro-scene construction module determines key areas related to disaster evolution in video stream V based on L and reduces interference from irrelevant backgrounds on the recognition results.
[0040] Specifically, the micro-scene construction module first establishes a mapping relationship between the video image and the underground space structure based on L, dividing the video image into several local areas corresponding to the roadway structure, equipment boundaries, or airflow channels, and using each local area as a candidate micro-scene range. Areas such as the inner side of transfer points, return air corners, near ventilation duct outlets, the shaded area under the conveyor belt, and the areas before and after the dust curtain are prioritized as highly sensitive candidate areas to enhance the detection capabilities for localized gas accumulation, smoke retention, personnel obstruction, and equipment thermal disturbance.
[0041] During the reflection suppression process, the micro-scene construction module targets false targets caused by spray droplets, water film adhesion, metal surface reflection, and bright patches on wet coal walls. It performs brightness saturation recognition, specular reflection discrimination, and edge consistency correction on bright areas in video stream V to reduce the prominence of non-real disaster targets in the image. Preferably, it can combine the brightness variation patterns between consecutive frames, spatial texture continuity, and regional motion consistency to distinguish between instantaneous flickering reflections and stable reflective backgrounds, thereby preserving effective information such as smoke diffusion, thermal disturbance edge drift, and abnormal target contours.
[0042] During the swing occlusion separation process, the micro-scene construction module performs temporal decomposition on the foreground motion region in video stream V to address periodic occlusions caused by dust curtain swings, cable swaying, ventilation duct vibrations, and the swinging of local suspended objects. It identifies the regular swing components related to ventilation disturbances and separates them from the irregular disturbance components corresponding to personnel lingering, smoke spread, coal dust cloud drift, and abnormal equipment displacement. This separation process reduces the impact of swing occlusion on the identification of abnormal personnel presence, tracking of visible hotspot areas, and extraction of weak smoke targets.
[0043] During low-light compensation, the micro-scene construction module addresses the loss of detail caused by night shift maintenance, insufficient local lighting, equipment backlighting, and dark areas at alley corners. It performs brightness enhancement, dark area texture restoration, and local contrast enhancement on the video stream V, while maintaining the relative stability of alley outlines, equipment edges, and personnel shapes to avoid noise amplification or false contour generation caused by simple brightening. Preferably, the low-light compensation and reflection suppression are performed jointly to balance dark area detail restoration and bright area suppression.
[0044] Based on the preprocessed video data, the micro-scene construction module constructs micro-scene units U according to the principle of "spatial location + structural boundary + temporal continuity". Each micro-scene unit U is a spatiotemporal segment corresponding to a specific local risk behavior, and at least includes a local spatial range, a duration range, and a corresponding visual state sequence. For example, the gas accumulation zone above the return air corner can be constructed as a gas accumulation type micro-scene unit, the area near the transfer point roller can be constructed as a frictional heat disturbance type micro-scene unit, and the blind spot behind the dust curtain can be constructed as an obstruction and retention type micro-scene unit. Each micro-scene unit U can independently characterize the formation, persistence, and diffusion process of visual anomalies within a local area.
[0045] Based on this, the micro-scene construction module extracts visual perturbation features Fv from each micro-scene unit U. The visual perturbation features Fv include at least one or more of the following: brightness abrupt change features, abnormally bright residual features, contour deformation features, persistent occlusion features, local drift features, texture blur diffusion features, smoke plume direction features, personnel or equipment boundary interruption features, and residual motion features after periodic oscillation suppression. The visual perturbation features Fv are used to characterize visual anomalies that may correspond to disaster precursors within the target area, and serve as inputs for subsequent asynchronous spatiotemporal registration and cross-modal causality verification with gas concentration G, wind speed W, temperature and humidity T, and acoustic vibration signals A.
[0046] In one embodiment of the present invention, a feature extraction module is used to asynchronously spatiotemporally register gas concentration G, wind speed W, temperature and humidity T, and acoustic vibration signal A according to micro-scene unit U, and extract gas response feature Fg, thermal inertia feature Ft, and acoustic vibration mutation feature Fa to form a multimodal environment characterization corresponding to visual disturbance feature Fv, providing input for subsequent cross-modal causal verification and residual risk identification. The micro-scene unit U is generated by the aforementioned micro-scene construction module based on the roadway topology parameter L and video stream V, and includes at least a local spatial range, a time segment range, and a corresponding visual state sequence. Therefore, this module uses U as a unified registration carrier to map sensor data with different sampling frequencies, different installation positions, and different response delays to the same local risk scene.
[0047] Specifically, the feature extraction module first establishes spatial correlations between gas concentration G, wind speed W, temperature and humidity T, and acoustic / vibration signals A and micro-scene units U based on the sensor installation locations, tunnel connectivity, equipment distribution locations, and airflow paths recorded in the tunnel topology parameters L. For sensors located upstream, downstream, top, sidewall, or near equipment in U, this module determines the contribution range and influence weight of each sensor's data to the micro-scene unit U based on its relative distance to U, obstruction relationship, and ventilation channel direction, thereby avoiding misassignment of sensor changes from irrelevant areas to the current local scene.
[0048] Furthermore, since the sampling frequencies of the video stream V and various sensor signals are different, and there is an inherent lag in gas diffusion, heat transfer, and mechanical vibration propagation, this module performs asynchronous temporal alignment processing on G, W, T, and A. This asynchronous temporal alignment processing includes: initial synchronization of data for each modality based on a unified timestamp; compensation for the delay in the arrival of gas concentration G in micro-scene unit U based on the wind speed W, which represents the airflow velocity and direction; compensation for the cumulative lag of thermal anomaly signals based on the slope of temperature and humidity T changes and environmental heat dissipation conditions; and time window matching for mechanical noise, impact vibration, and crack activity signals based on the correspondence between acoustic and vibration signals A and the occurrence time of visual events. Through this asynchronous spatiotemporal registration, previously misaligned abnormal changes in different modalities can be uniformly mapped to the same disaster evolution process corresponding to micro-scene unit U.
[0049] After registration, the feature extraction module extracts gas response features Fg from the gas concentration G and wind speed W. These gas response features Fg include at least one or more of the following: local concentration increase amplitude, concentration change rate, concentration fluctuation frequency, upstream-to-downstream concentration transfer trend, accumulation characteristics that change inversely with wind speed, local backflow and retention characteristics, and diffusion characteristics extending along the boundary of the micro-scene unit U. These gas response features Fg are used to characterize states such as gas release from roof cracks, gas accumulation in blind areas, and local gas retention or abnormal diffusion caused by ventilation disturbances. They are particularly useful for distinguishing between short-term concentration changes caused by normal airflow fluctuations and persistent abnormal responses caused by precursors to local disasters.
[0050] Simultaneously, the feature extraction module extracts thermal inertial features Ft from temperature and humidity T and time-series information related to equipment operating status. The thermal inertial features Ft include at least one or more of the following: gradual temperature rise characteristics, duration of temperature rise, post-heating attenuation characteristics, humidity-related characteristics, localized thermal disturbance aggregation characteristics, and gradual trend characteristics corresponding to visual hotspot areas. The thermal inertial features Ft are used to characterize phenomena such as early heat accumulation from roller friction, localized overheating at motor or cable connections, and the continued presence of abnormal heat sources after spraying stops. Unlike instantaneous thermal interference, thermal inertial features Ft emphasize the accumulation, maintenance, and attenuation patterns of abnormal heat sources over a certain period, thereby improving the ability to identify small-scale, early-stage, and difficult-to-detect thermal risks.
[0051] Furthermore, the feature extraction module extracts acoustic vibration abrupt change features Fa from the acoustic vibration signal A. These acoustic vibration abrupt change features Fa include at least one or more of the following: impact pulse features, broadband abnormal noise features, continuous friction band enhancement features, periodic vibration instability features, peak energy transition features, short-term sudden cracking noise features, and abnormal harmonic features deviating from the normal operating spectrum of the equipment. These acoustic vibration abrupt change features Fa are used to characterize states such as idler roller jamming, bearing abnormalities, micro-crack activity in the support structure, precursors of localized roof rupture, and foreign object friction in the equipment. They can also be corroborated by the contour changes, obstruction disturbances, and smoke diffusion phenomena in the visual disturbance features Fv.
[0052] Preferably, the feature extraction module can also normalize, calibrate, and compress the time window for the gas response feature Fg, thermal inertia feature Ft, and acoustic vibration mutation feature Fa to form a structured feature sequence that corresponds one-to-one with the micro-scene unit U. The structured feature sequence not only retains the intensity information of each modal anomaly change, but also retains the chronological order, duration, and propagation relationship of the anomalies, thus providing basic data support for subsequent construction of a cross-modal causal consistency matrix, generation of non-hazardous disturbance baselines, and residual risk features R after removing pseudo-anomalies.
[0053] In one embodiment of the present invention, a pseudo-anomaly stripping module is used to construct a cross-modal causal consistency matrix C based on Fv, Fg, Ft, and Fa, and generate a corresponding non-hazardous disturbance baseline B, thereby obtaining the residual risk feature R after stripping pseudo-anomalies. This process uses a micro-scene unit U as a unified data organization unit, enabling the visual disturbance feature Fv corresponding to video stream V, the gas response feature Fg corresponding to gas concentration G, the thermal inertia feature Ft corresponding to temperature and humidity T, and the acoustic vibration mutation feature Fa corresponding to acoustic vibration signal A to be correlated and analyzed within the same local space and the same evolution period.
[0054] Specifically, using the micro-scene unit U as an index, Fv, Fg, Ft, and Fa are segmented according to a unified sliding time window, and a modal association candidate set M is generated by combining the airflow direction, sensor spacing, and equipment position in the roadway topology parameters L. The unified sliding time window preferably adopts a continuous sliding mode with a length of 3 to 20 seconds and a step size of 0.5 to 5 seconds. The window length is determined according to the shortest identifiable anomaly duration in the target area. For persistent anomalies such as roller friction hot spots and roof crack gas release, the window length is preferably 8 to 12 seconds; for sudden acoustic and vibration anomalies, the window length is preferably 3 to 5 seconds. Within each sliding time window, spatial positioning is first completed according to the roadway topology parameters L, and then Fv, Fg, Ft, and Fa within the same micro-scene unit U that can form a forward and backward propagation relationship are combined into candidate feature pairs, and written into the modal association candidate set M according to modal category, spatial position, and time window number.
[0055] When performing topology-constrained time-delay matching on each candidate feature pair in M, the path constraint set P for each candidate feature pair is first constructed based on the airflow direction, the distance between the sensors at both ends of the candidate feature pair, and the equipment occlusion relationship in the roadway topology parameters L. Then, a corresponding dynamic time-delay search window Δt is generated according to the path length and wind speed W. Here, the path length is the actual propagation distance of the candidate feature pair on the roadway connection path, and the wind speed W is the average wind speed from upstream to downstream of the path. When there is equipment occlusion or a return air bend, an occlusion compensation time of 0.5 to 3 seconds is added to the propagation time estimated based on the path length and wind speed W.
[0056] In one embodiment of the present invention, when performing topology-constrained time-delay matching on each candidate feature pair in M, the dynamic time-delay search window Δt can be determined based on the propagation path length, wind speed W, and equipment occlusion relationship of the candidate feature pair in the roadway topology parameter L.
[0057] For the starting feature point and the response feature point in the candidate feature pair, the reachable propagation path between them is first determined based on the roadway topology parameter L, denoted as the target propagation path in the path constraint set P. If this propagation path consists of n connected paths, its total path length S is determined by the following formula: , k=1,2,…,n; where sk represents the actual length of the k-th connected path, in meters. The actual length is preferably determined based on a combination of the sensor installation location, the length of the roadway centerline, the equipment bypass boundary, and the corner transition length. If a candidate feature crosses a non-connected boundary, or if its propagation direction is completely opposite to the airflow direction and there is no return flow structure for support, then the candidate feature is determined not to be included in the dynamic time-delay search window Δt calculation.
[0058] After determining the total path length S, the effective average wind speed, denoted as WS, is calculated based on the wind speed W corresponding to each segment of the path. A path-weighted average method is preferred for determining this. Where Wk represents the wind speed on the k-th connected path, in meters per second; θk represents the angle between the path direction of the k-th segment and the local wind flow direction; cosθk is used to characterize the consistency of the wind flow direction. When the path direction is consistent with the wind flow direction, cosθk is close to 1; when there is deflection, cosθk is less than 1.
[0059] If WS is less than the preset minimum effective wind speed threshold Wmin, the candidate feature pair can be marked as a low-confidence propagation path, or WS can be set according to Wmin to avoid abnormal amplification of time delay due to an excessively small denominator. The preferred value for Wmin is 0.2 m / s to 0.5 m / s.
[0060] In actual downhole scenarios, equipment obstruction, return air bends, and localized flow around points can increase propagation delay. Therefore, an obstruction compensation time tb is introduced in addition to the basic propagation time. The obstruction compensation time tb can be determined by the following formula: tb = N1 × δ1 + N2 × δ2 + N3 × δ3; where N1 represents the number of equipment obstruction points on the propagation path; N2 represents the number of return air bends on the propagation path; N3 represents the number of localized flow around points or bifurcation switching points on the propagation path; δ1 represents the compensation time corresponding to a single equipment obstruction point; δ2 represents the compensation time corresponding to a single return air bend; and δ3 represents the compensation time corresponding to a single localized flow around point or bifurcation switching point. Preferably, δ1 is 0.5 to 2 seconds, δ2 is 0.5 to 1.5 seconds, and δ3 is 0.3 to 1 second. The specific values can be obtained through statistical calibration of historical labeled samples.
[0061] Based on the total path length S and the effective average wind speed WS, the propagation base duration t0 is obtained: Further combining the occlusion compensation time tb, the center time delay τc of the candidate feature pair is obtained: ; where τc represents the estimated center time delay of the propagation from the initial feature point to the response feature point under the constraints of the current roadway topology parameters L and wind speed W.
[0062] Considering the effects of downhole airflow fluctuations, sensor sampling errors, and local disturbances, relaxed boundaries are set on both sides of the central time delay τc to form a dynamic time delay search window Δt. The dynamic time delay search window Δt can be expressed as: Wherein, ε1 represents the forward relaxation time, and ε2 represents the backward relaxation time. Preferably, ε1 and ε2 are jointly determined based on the wind speed fluctuation amplitude, sampling period, and historical matching deviation; in one embodiment, they can be taken as: ; Wherein, β and γ are relaxation coefficients, preferably β is 0.1 to 0.2 and γ is 0.15 to 0.3. Generally, ε2 can be slightly larger than ε1 to cover the possible hysteresis phenomenon in the response mode.
[0063] For example, in a return air corner scenario, the total path length S between the initial feature point corresponding to a visual disturbance peak and the response feature point corresponding to the gas concentration inflection point is 8 meters. The two wind speeds along the path are 1.0 m / s and 0.8 m / s respectively, and both paths are basically aligned with the airflow direction. Therefore, the effective average wind speed WS can be approximated as 0.9 m / s. If there is one device obstruction and one return air corner along the propagation path, and δ1 is 1 second and δ2 is 0.8 seconds, then the obstruction compensation time tb is 1.8 seconds. At this time, the propagation base duration t0 is approximately 8.89 seconds, and the center time delay τc is approximately 10.69 seconds. If β is 0.1 and γ is 0.2, then the forward relaxation time ε1 is approximately 1.07 seconds, and the backward relaxation time ε2 is approximately 2.14 seconds. Therefore, the dynamic time delay search window Δt can be determined to be approximately 9.62 seconds to 12.83 seconds. Subsequent matching will only be performed on gas response features falling within this time range.
[0064] When performing asymmetric sliding matching on candidate feature pairs within Δt, event anchoring is implemented for visual disturbance peaks, gas concentration inflection points, the starting point of the temperature gradual rise segment, and acoustic vibration pulse peaks to obtain the time-stamped sequence Q of each modal event. Visual disturbance peaks are determined by the local maxima of brightness abrupt change amplitude, contour deformation intensity, or local drift intensity; gas concentration inflection points are determined by the position where the first-order rate of change of gas concentration G changes from low to high and exceeds the upper limit of background fluctuation; the starting point of the temperature gradual rise segment is determined by the position where three consecutive sampling points in temperature and humidity T maintain a consistent upward trend and the cumulative increase exceeds the 95th percentile of the non-alarm sample; acoustic vibration pulse peaks are determined by the position where the short-time energy peak and frequency band abrupt change of acoustic vibration signal A simultaneously cross the threshold. The above thresholds are preferably jointly calibrated based on historical non-alarm periods and marked risk periods, where non-alarm periods are used to determine the upper limit of normal fluctuations, and marked risk periods are used to determine the anomaly identification sensitivity.
[0065] Based on Q, the visual lead time difference, gas lag time difference, thermal inertia duration, and acoustic vibration instantaneous offset are calculated, and abnormal matches that do not meet the topological reachability, time difference upper limit, and persistence threshold are filtered out to obtain the effective time series relation set H. The visual lead time difference is defined as the time difference between the peak of the visual disturbance and the subsequent gas concentration inflection point, the start of the temperature rise phase, or the acoustic vibration pulse peak; the gas lag time difference is defined as the delay when the gas concentration G deviates significantly after the visual disturbance peak; the thermal inertia duration is defined as the duration from the start of the temperature rise phase to the return to the normal envelope; the acoustic vibration instantaneous offset is defined as the amount by which the acoustic vibration pulse peak moves forward or backward relative to the visual disturbance peak. The time difference upper limit is preferably determined by the 95th percentile value of historical sample statistics, and the persistence threshold is preferably set according to different risk types, where the thermal inertia duration can be set to 5 to 60 seconds, and the acoustic vibration instantaneous offset can be set to 0 to 3 seconds. After performing sequential encoding and confidence weighting on H, a temporal coupling parameter D is generated. The sequential encoding is used to describe the sequence type of "visual disturbance leading - gas response lag - thermal inertia persistence" or "visual disturbance leading - acoustic vibration sudden change instantaneous". The confidence weighting is calculated using a 0 to 1 normalization method. Specifically, the degree of time difference deviation, peak clarity and spatial position consistency are converted into scores in the range of 0 to 1, and then summed according to the initial weights of 0.4, 0.3 and 0.3. The initial weights can be calibrated through training with labeled risk samples, and the final value is determined based on the highest classification accuracy or the highest F1 value.
[0066] Based on the temporal coupling parameter D, and combined with the airflow channels, bifurcation connections, and equipment obstruction locations in the roadway topology parameter L, a propagation path verification graph G for candidate feature pairs is constructed to determine whether each candidate feature pair has an reachable propagation chain from upstream to downstream or from the risk source to the monitoring point. The nodes in the propagation path verification graph G correspond to the event points Fv, Fg, Ft, and Fa, and the edges correspond to the propagation directions allowed by the roadway topology parameter L. When a candidate feature pair crosses a non-connected boundary, flows against the airflow direction without return flow support, or is completely blocked by equipment obstruction, it is determined to be unreachable. For candidate feature pairs that pass the propagation path verification, the visual lead time difference, gas lag time difference, thermal inertia duration, and acoustic vibration instantaneous offset in D are called up and compared with the preset response delay threshold intervals. Abnormal feature pairs with inverted timing, excessive lag, or instantaneous mismatch are eliminated. Then, the co-occurrence time period of the corresponding abnormal events is extracted from the feature pairs after delay screening. The overlap duration between the visual anomaly maintenance interval and the gas response interval, thermal duration interval, or acoustic vibration sudden interval is calculated. Short-term sporadic coupling pairs that do not meet the minimum co-occurrence duration threshold are eliminated. The minimum co-occurrence duration threshold is preferably set according to the risk type: 2 to 8 seconds for gas anomalies, 5 to 20 seconds for thermal anomalies, and 0.5 to 3 seconds for acoustic vibration anomalies. The remaining features are weighted and scored according to propagation reachability, time delay matching degree, and co-occurrence stability, outputting a set of valid causal edges E. Propagation reachability is scored binaryly, with 1 for reachability and 0 for non-reachability. Time delay matching degree is converted to a score of 0 to 1 based on the degree to which the candidate feature's actual time difference is near the center of the threshold interval, from high to low. Co-occurrence stability is converted to a score of 0 to 1 based on the proportion of overlap duration to the smaller of the respective durations. The combination order of the three scores is as follows: first determine propagation reachability, then calculate time delay matching degree, and finally calculate co-occurrence stability. If any of the preceding terms is not met, the subsequent terms are not included in the scoring.
[0067] A cross-modal causal consistency matrix C is generated by weighting the direction, strength, and duration of the effective causal edge set E. The cross-modal causal consistency matrix C is preferably constructed as a 4x4 matrix with row and column indices Fv, Fg, Ft, and Fa, where rows represent the initial mode and columns represent the response mode. For example, the matrix element in row Fv and column Fg represents the degree of causal consistency for "visual disturbance leading - gas response lagging," and the matrix element in row Fv and column Ft represents the degree of causal consistency for "visual disturbance leading - thermal inertia persisting." The value of each matrix element is calculated sequentially from the direction validity, strength score, and duration score of the corresponding edge in the effective causal edge set E. Direction validity is used to confirm whether the causal edge conforms to the preset propagation order; a value of 1 is assigned if it conforms, and 0 if it does not. The strength score is given by the edge weight parameter, which is determined by a combination of propagation reachability, time delay matching degree, and co-occurrence stability scores. The duration score is obtained by normalizing the maintenance duration of the corresponding anomalous chain within a unified sliding time window. The preferred normalization method is linear normalization using minimum and maximum values. The minimum value is taken as the lower bound of similar edges in historical samples without alarms, and the maximum value is taken as the upper bound of similar edges in labeled risk samples. Values exceeding the upper bound are treated as 1, and values below the lower bound are treated as 0. The resulting cross-modal causal consistency matrix C not only characterizes whether different modes are correlated, but also whether the correlation satisfies the actual propagation path, time delay sequence, and continuous evolution law in coal mines. For example, in the early overheating scenario of the idler roller at the transfer point, if a visual disturbance peak appears first, followed by the start of a gradual temperature rise, accompanied by a short-duration acoustic pulse peak, then the matrix elements corresponding to Fv to Ft and Fv to Fa are higher, while the matrix elements corresponding to Fv to Fg are lower.
[0068] When generating the corresponding non-disaster disturbance baseline B, firstly, based on the time-series coupling parameter D, combined with the roadway topology parameter L, equipment operating status, and spray start / stop records, a reference sample set N is extracted from historical alarm-free periods that is consistent with the current micro-scene unit U in terms of topological location and operating conditions. "Consistent topological location" means that the reference sample and the current micro-scene unit U are located at the same roadway corner, in the same transfer point neighborhood, in the same ventilation duct outlet area, or in the same area before or after the same dust curtain. "Consistent operating conditions" means that the equipment operating status, spray start / stop status, shift time, and main ventilation status are the same or similar. The preferred sample selection order is first by roadway topology parameter L, then by equipment operating status, and finally by spray start / stop records. The number of reference sample sets N is preferably no less than 100 sliding time windows. When there are insufficient samples under the same micro-scene unit U, adjacent micro-scene units U with similar topological structures can be introduced as supplements, but the airflow direction and equipment layout must be consistent.
[0069] The disturbance sources Fv, Fg, Ft, and Fa in N are clustered, and the non-hazard disturbance prototype vectors Z are constructed for spray reflection, dust curtain sway, equipment idling vibration, and airflow fluctuation. Density-based clustering is preferred for disturbance source clustering. First, the brightness fluctuation amplitude, boundary sway frequency, and local drift velocity in Fv; the fluctuation amplitude and frequency in Fg; the gradual rise and fall slope in Ft; and the dominant frequency band position and peak energy in Fa are used as clustering dimensions. Then, the samples are clustered into multiple clusters according to their density distribution in the feature space. Clusters with high consistency between their centers and spray start / stop records, dust curtain sway records, or equipment idling records are labeled as the non-hazard disturbance prototype vectors Z corresponding to spray reflection, dust curtain sway, equipment idling vibration, and airflow fluctuation, respectively. Density-based clustering is used because the duration and amplitude distribution of non-hazard disturbances are often irregular, and density-based clustering is more suitable for preserving sparse boundary samples.
[0070] Based on Z, the visual brightness amplitude range, gas fluctuation upper limit, thermal rise slope upper limit, and acoustic frequency band tolerance corresponding to each disturbance prototype are calculated to form disturbance envelope parameters constrained by topological location and operating conditions. The visual brightness amplitude range is preferably taken from the 5th to 95th percentile of the corresponding disturbance prototype vector Z in the reference sample set N; the gas fluctuation upper limit is preferably taken as the 99th percentile to ensure tolerance for occasional airflow fluctuations; the thermal rise slope upper limit is preferably taken as the 95th percentile to avoid misidentifying short-term temperature rise noise as a risk heat source; the acoustic frequency band tolerance is preferably taken as a frequency band range with the center of the main frequency band as the reference, taking three standard deviations above and below the mean. The disturbance envelope parameters are indexed and encoded according to the micro-scene unit U to generate a non-hazardous disturbance baseline B corresponding one-to-one with the current scene. In other words, the non-hazardous disturbance baseline B is not a single fixed threshold, but a scene-based envelope set bound to the roadway topology parameters L, equipment operating status, and spray start / stop records. For example, in the return air corner area where the spray is on and the dust curtain is swinging, the non-hazardous disturbance baseline B allows for higher visual brightness fluctuations and a wider range of local swing frequencies; while in the maintenance area where the spray is off, the corresponding visual disturbance envelope is significantly narrowed.
[0071] Based on the current micro-scene unit U, the corresponding non-disaster disturbance baseline B is invoked, and Fv, Fg, Ft, and Fa are respectively mapped to the visual disturbance envelope, gas wave envelope, thermal inertia envelope, and acoustic vibration tolerance envelope in B to obtain the baseline deviation set Δ for each mode. The component mapping preferably adopts the mapping rule of prioritizing the same location, same operating condition, and same mode: first, the same mode envelope parameters are searched in the non-disaster disturbance baseline B corresponding to the current micro-scene unit U. If there are multiple candidate envelopes, the envelope consistent with the current equipment operating status and spray start / stop record is selected first; if there are still multiple candidate envelopes, the historical operating condition envelope closest to the start time of the current event is selected. The calculation order of the baseline deviation set Δ for each mode is as follows: first, it is determined whether the current feature value falls within the corresponding envelope range. If it falls within the range, it is recorded as 0 deviation; if it exceeds the range, it is converted into deviation amount according to the proportion of the amplitude exceeding the upper limit or falling below the lower limit to the width of the envelope. This proportion is preferably normalized to the interval between 0 and 1, and is truncated to 1 when it exceeds 1.
[0072] When performing causal constraint residual decomposition on each modal deviation in Δ using the effective causal edge set E, the initial mode, response mode, and edge weight parameters in the effective causal edge set E are used as constraints to perform directed correlation expansion on the deviation set Δ, constructing a visual deviation-gas deviation-thermal deviation / acoustic vibration deviation residual transfer subgraph Gd, and marking isolated deviations not on the constraint path of E as components to be removed. Subsequently, based on the time delay parameters, intensity weights, and duration weights corresponding to each directed edge in Gd, piecewise projection is performed on adjacent modal deviations in Δ to separate causal response components that satisfy the sequential triggering relationship and synchronous disturbance components that do not satisfy the triggering order, obtaining the initial residual pair Rc1. The segmented projection here refers to aligning the deviation of a certain mode along its corresponding effective time interval in the temporal coupling parameter D to the next mode, retaining only the deviation components that fall within the effective time interval and have the same direction of intensity change; for example, if a gas deviation occurs within 8 to 12 seconds after a visual deviation occurs, the gas deviation can be projected as the causal response component of the visual deviation, otherwise it is retained as a synchronization disturbance component to be removed.
[0073] When performing coupling consistency verification on Rc1 based on the spatial adjacency relationship and temporal overlap interval within the micro-scene unit U, the residual pairs in Rc1 are first mapped to their respective spatial grid positions within the micro-scene unit U. An adjacent grid connectivity table is then established based on the tunnel topology parameter L, and candidate residual groups with spatial distances less than a preset adjacency threshold are selected. The spatial grid of the micro-scene unit U is preferably divided with a grid side length of 0.5 to 2 meters, and the adjacency threshold is preferably 1 to 2 grid units. When used in narrow areas such as return air corners, the adjacency threshold can be 1 grid unit; when used in wider areas such as the outer edge of transfer points, the adjacency threshold can be 2 grid units. Next, the start and end timestamps of corresponding abnormal events are extracted from each candidate residual group, and the temporal overlap interval and overlap ratio between visual deviation, gas deviation, thermal deviation, or acoustic / vibration deviation are calculated to obtain the spatiotemporal coupling parameters. The temporal overlap ratio is preferably calculated by dividing the overlap duration by the smaller of the two durations. Consistency judgment is performed on each candidate residual group based on spatiotemporal coupling parameters. Pseudo-coupled residual groups that are spatially adjacent but have excessive temporal misalignment, or that overlap temporally but cross non-connected boundaries, are eliminated, retaining continuous transmission chains. Finally, the retained continuous transmission chains are jointly weighted according to spatial adjacency strength and temporal overlap stability to generate a candidate coupled residual set Rc2 that satisfies the conditions of local co-domain and simultaneous coupling. Spatial adjacency strength is preferably calculated by inversely converting the actual spatial distance to the adjacency threshold; the closer the distance, the higher the score. Temporal overlap stability is preferably determined by the overlap ratio and the degree of overlap fluctuation within the continuous window; the higher the overlap ratio and the smaller the fluctuation, the higher the score.
[0074] Chain-based cumulative encoding is performed on Rc2, superimposing the visual driving quantity, gas diffusion quantity, thermal retention quantity, and acoustic-vibration mutation quantity in each residual chain according to edge weights. The output is a coupled deviation component that satisfies the causal constraint, which serves as the candidate residual set Rc for subsequent pseudo-anomaly removal. Specifically, the visual driving quantity is the peak value and area of the visual deviation component in a continuous window; the gas diffusion quantity is the cumulative deviation of the gas deviation component within a dynamic time delay search window Δt; the thermal retention quantity is the cumulative duration and gradual rise intensity of the thermal deviation component within a sustained segment; and the acoustic-vibration mutation quantity is the peak energy and frequency band shift amplitude of the acoustic-vibration deviation component in a pulse event. Each quantity is first normalized to 0 to 1 according to the range of the reference sample set N and the labeled risk samples, and then superimposed according to the corresponding edge weight parameters in the effective causal edge set E. The initial values of the edge weight parameters can be determined by the scoring results of propagation reachability, time delay matching degree, and co-occurrence stability; when labeled samples are available, they can be adjusted item by item in the range of 0.1 to 0.8 through grid search to maximize the distinguishability between risk samples and non-alarm samples.
[0075] When performing pseudo-anomaly stripping on Rc, components corresponding to spray reflection, dust curtain swaying, equipment idling vibration, and short-term airflow fluctuations are deducted, and transient residuals with durations below the minimum risk maintenance threshold are screened out, resulting in the net residual set Rn. A component with a similar origin is one that is highly consistent with the prototype vector Z of the non-hazardous disturbance in terms of characteristic shape, duration, frequency band structure, and spatial location. The preferred determination order is: first, compare whether the spatial location is consistent with the location of the spray, dust curtain, or idling equipment; second, compare whether the morphological similarity reaches the threshold; and finally, compare whether the duration falls within the normal range of the corresponding non-hazardous disturbance. Morphological similarity can be calculated comprehensively based on peak difference, slope difference, frequency band difference, and time overlap ratio, and normalized to the 0-1 range. When the morphological similarity is greater than 0.85 and a continuous transmission chain constrained by a complete and effective causal edge set E is not formed, it is determined to be a component with a similar origin and is deducted. The minimum risk retention threshold is preferably determined based on the risk type. For gas-related risks, the minimum retention time is preferably 3 to 10 seconds; for thermal risks, the minimum retention time is preferably 8 to 30 seconds; and for sudden acoustic and vibration risks, the minimum retention time is preferably 1 to 5 seconds. Transient residuals below the corresponding threshold, even if the amplitude is high, are preferentially judged as occasional disturbances and not included in the final output.
[0076] Rn is weighted and encoded according to modal contribution, temporal consistency, and spatial clustering to generate residual risk features R after removing pseudo-anomalies. Modal contribution represents the dominance of Fv, Fg, Ft, and Fa in the current net residual set Rn, and its value is preferably given by the proportion of each modal net residual energy to the total net residual energy. Temporal consistency represents whether the net residual chain continuously satisfies the sequential relationship of the temporal coupling parameter D and the effective causal edge set E, and its value is preferably determined by the proportion of the number of times the sequential relationship is satisfied within a continuous window to the total number of windows. Spatial clustering represents whether the net residuals are concentrated in the same micro-scene unit U or adjacent grid regions, and its value is preferably determined by the clustering density of the net residual landing points on the spatial grid. The preferred weighting order of the three is to calculate modal contribution first, then temporal consistency, and finally spatial clustering, and combine them with initial weights of 0.35, 0.35, and 0.30. This initial weight can be calibrated through historical labeled samples. When a certain type of disaster has more obvious spatial clustering characteristics, the weight of spatial clustering can be appropriately increased. The final output residual risk feature R retains the part of the cross-modal continuous response in the real risk chain, and eliminates the pseudo-anomalies caused by spray reflection, dust curtain swaying, equipment idling vibration and short-term airflow fluctuations. Therefore, it is more suitable as the input for subsequent risk evolution models.
[0077] For example, in a scenario where there is spray droplet reflection and local airflow fluctuation at the return air corner, the video stream V may first show obvious bright flickering, followed by slight fluctuations in the gas concentration G. If this bright flickering is mainly concentrated in the spray area, lasts for only 2 seconds, and is consistent with the oscillation frequency of the dust curtain, then after comparison with the non-hazardous disturbance baseline B, although the visual disturbance produces a certain deviation, it will be subtracted during the causal constraint residual decomposition and pseudo-anomaly removal process due to the lack of a stable and effective causal edge set E for support, and its high consistency with the non-hazardous disturbance prototype vector Z corresponding to the spray reflection. Conversely, if a stable visual disturbance first appears near the transfer point idler roller, followed by a continuous increase in the thermal inertia feature Ft corresponding to temperature and humidity T, accompanied by a short-term pulse peak in the acoustic vibration mutation feature Fa corresponding to the acoustic vibration signal A, then this anomaly chain can form high consistency in the cross-modal causal consistency matrix C, and after removing pseudo-anomalies, it will be retained as a high-weight residual risk feature R.
[0078] In one embodiment of the present invention, the linkage control module is used to input R and L into the risk evolution model, output the risk source location, propagation direction, and risk level result Y, and execute graded power outage, local ventilation adjustment, directional spraying, personnel evacuation, and model self-updating based on Y. R represents the residual risk characteristics after removing pseudo-anomalies, including at least the spatial location, modal contribution, temporal consistency, and spatial clustering of the corresponding micro-scene unit U; L represents the roadway topology parameters, including at least the roadway direction, cross-sectional dimensions, bifurcation relationships, equipment layout locations, ventilation structure locations, and connectivity relationships between adjacent areas. The risk evolution model uses R to represent the local true anomaly state and L to define the risk propagation boundary and propagation direction.
[0079] Specifically, the target area is first divided into spatial nodes corresponding to micro-scene units U based on L. Directed propagation edges are then established based on the roadway connectivity, airflow direction, and equipment obstruction relationships between nodes, forming a risk evolution diagram. Each node records its own risk level R, and each directed propagation edge records the propagation distance, number of turns, and degree of obstruction. Subsequently, R is mapped to the corresponding nodes, and the earliest node that continuously exceeds the minimum risk maintenance threshold is identified in chronological order as a candidate node for risk source location. The minimum risk maintenance threshold is preferably between 3 and 30 seconds and is determined through calibration using historical labeled samples based on different disaster types. A source point credibility score is calculated for each candidate risk source location node. This source point credibility score includes at least the initial occurrence time score, residual peak score, local clustering score, and upstream counter-evidence score. The initial occurrence time score characterizes whether the node's risk level R appears first and most stably; the residual peak score characterizes the intensity of the node's risk level R; the local clustering score characterizes whether adjacent nodes form an outward diffusion distribution around the node; and the upstream counter-evidence score excludes false source points where earlier anomalies already exist upstream. All the above scores are normalized to the 0 to 1 range and then weighted and summed. The node with the highest comprehensive score is taken as the location of the risk source.
[0080] After determining the location of the risk source, the direction of risk propagation is inferred based on L. Specifically, starting from the location of the risk source, a search is performed along the directed propagation edge towards adjacent nodes to form several candidate propagation paths; then, a propagation consistency score is calculated for each candidate propagation path. The propagation consistency score includes at least a time progression score, an intensity decay score, a topological validity score, and a spatial continuity score. The time progression score is used to characterize whether the downstream node R appears within a reasonable time window after the upstream node R appears; the intensity decay score is used to characterize whether R gradually decays from near to far along the propagation direction; the topological validity score is used to characterize whether the candidate path conforms to the connectivity relationship and airflow direction defined by L; the spatial continuity score is used to characterize whether there are gaps between adjacent nodes on the candidate path that exceed a preset jump distance. The preferred scoring order is to first determine the topological validity score, and then calculate the time progression score, intensity decay score, and spatial continuity score; if the topological validity score does not meet the conditions, the corresponding candidate propagation path is directly eliminated. Finally, the path with the highest score that exceeds the propagation confirmation threshold is determined as the propagation direction; when multiple branch paths simultaneously exceed the propagation confirmation threshold, a multi-branch propagation direction can be output. The propagation confirmation threshold is preferably between 0.60 and 0.85, and is repeatedly calibrated using historical treatment samples.
[0081] After the location and propagation direction of the risk source are determined, the risk evolution model outputs a risk level result Y. Y includes at least the location of the risk source, the propagation direction, and the risk level. The risk level is preferably divided into four levels: Level 1 indicates that the anomaly is confined to a local node and has not yet formed a stable expansion; Level 2 indicates that the anomaly has propagated to adjacent nodes; Level 3 indicates that the anomaly continues to expand and is located near equipment or personnel activity areas; Level 4 indicates that the anomaly has the potential to rapidly amplify or couple into a disaster. The preferred order for determining the risk level is to first determine whether the risk source location is stable, then determine whether the propagation direction is valid, and then classify the risk level based on the scope of influence, the modal contribution of R, temporal consistency, and spatial clustering. To avoid false alarms caused by transient disturbances, it is preferable to require the same risk level to recur within two to five consecutive sliding time windows before confirmation.
[0082] When implementing graded power outages based on risk level Y, the location of the risk source and the coverage area of its propagation direction are mapped to the power supply circuits of the corresponding equipment. For level 1 risk, only an early warning is issued and the power of related equipment is limited; for level 2 risk, non-essential equipment in the local area where the risk source is located is powered off, while monitoring and ventilation equipment are retained; for level 3 risk, graded power outages are implemented for power equipment within the coverage area of the propagation direction; for level 4 risk, an expanded power outage is implemented in the hazardous area, retaining power only for emergency communication, monitoring, and necessary safety equipment.
[0083] When performing localized airflow adjustment based on Y, the target section for airflow adjustment is determined according to the location of the risk source, the direction of propagation, and the location of ventilation structures in L. The damper opening, duct angle, or local ventilation fan operating level is then adjusted accordingly. If the risk expands along the return air direction, it is preferable to guide the airflow to a safe dilution channel; if the risk lingers at equipment obstructions, it is preferable to create a bypass flow. The airflow adjustment range is preferably set to three levels: low, medium, and high, corresponding to adjustments of 5% to 10%, 10% to 20%, and 20% to 35% of the current airflow, respectively, and selected according to Y.
[0084] When conducting directional spraying based on risk source location, the spray initiation point, spray angle, and duration are determined by considering the direction of propagation and the location of the risk source. For gas or dust anomalies, it is preferable to deploy a spray barrier at 1 to 3 nodes ahead of the propagation direction; for thermal anomalies, it is preferable to directly target the area of heat accumulation. The spray duration can be set from 10 to 300 seconds according to the risk level and dynamically adjusted based on whether risk level (R) continues to decrease.
[0085] When evacuating personnel according to Y, first determine the danger zone, priority evacuation route, and prohibited entry area based on L. Then, evacuate personnel along connecting routes that are opposite to the direction of propagation and do not intersect with it. Level 1 risk level only indicates moving away from the local area; Level 2 risk level requires evacuating personnel from the local work area; Level 3 risk level requires evacuating personnel within the coverage area of the propagation direction and adjacent nodes; Level 4 risk level requires evacuating personnel from the entire danger zone and downstream connecting routes.
[0086] After the above actions are completed, the linkage control module performs a model self-update. Specifically, it records the changes in Y before and after the action, the content of the action instruction, the duration of the action, and the trajectory of R. It then generates an action effect label based on whether R drops below the release threshold within a predetermined time, whether the propagation direction terminates, and whether the risk source location disappears. If the action is effective, the sample is added to the forward update sample set to calibrate the propagation damping parameter, branch bias parameter, and risk level threshold in the risk evolution model. If the action is ineffective or misjudged, it is added to the correction sample set for reverse correction. Preferably, an incremental update is performed every 20 to 100 valid samples, and the change in parameters at one time is limited to 5% to 15% of the original value to prevent model drift.
[0087] For example, in a scenario where the residual risk characteristic R continuously increases near the transfer point idler, the risk evolution model first identifies the inner node of the transfer point as the risk source location based on L. It then finds that R continuously expands downstream along the conveyor belt to two adjacent nodes, and the risk level Y reaches level 3. At this point, the linkage control module first performs a graded power outage on the transfer equipment and adjacent power equipment, then performs medium-level localized ventilation and directional spraying on the area from the inner side of the transfer point to one downstream node, and organizes personnel in this localized work area to evacuate along the reverse safety passage. If R significantly decreases and the propagation direction disappears within five consecutive sliding time windows after the treatment, this sample is marked as a valid treatment sample and used for subsequent model self-updates. As another example, if R is detected at the return air corner with a high gas mode contribution and rapidly expanding along the return air direction, and the risk level Y reaches level 4, the linkage control module preferentially expands the power outage range, adopts high-level localized ventilation, and expands the personnel evacuation boundary to prevent further coupling of localized gas accumulation into a disaster.
[0088] The above description is merely a specific embodiment of this application, but the scope of protection of this application is not limited thereto. Any changes or substitutions that can be easily conceived by those skilled in the art within the scope of the technology disclosed in this application should be included within the scope of protection of this application.
Claims
1. A coal mine disaster risk prevention and control platform based on multimodal perception and AI video recognition, characterized in that: include: Multi-source data acquisition module: acquires video stream V, gas concentration G, wind speed W, temperature and humidity T, acoustic and vibration signals A, and roadway topology parameters L of the target area; Micro-scene construction module: Based on L, reflectivity suppression, sway occlusion separation and low light compensation are applied to V to construct micro-scene unit U and obtain visual perturbation feature Fv; Feature extraction module: Asynchronously registers G, W, T, and A according to U to extract gas response feature Fg, thermal inertia feature Ft, and acoustic vibration abrupt change feature Fa; False anomaly removal module: Construct a cross-modal causal consistency matrix C based on Fv, Fg, Ft and Fa, and generate the corresponding non-hazardous disturbance baseline B to obtain the residual risk characteristics R after removing false anomalies; Linkage control module: Inputs R and L into the risk evolution model, outputs the risk source location, propagation direction and risk level result Y, and executes graded power outage, local ventilation adjustment, directional spraying, personnel evacuation and model self-updating based on Y.
2. The coal mine disaster risk prevention and control platform based on multimodal perception and AI video recognition according to claim 1, characterized in that: Constructing a cross-modal causal consistency matrix C based on Fv, Fg, Ft, and Fa includes the following steps: Using the micro-scene unit U as the index, Fv, Fg, Ft and Fa are segmented according to a unified sliding time window, and modal association candidate set M is generated by combining the airflow direction, sensor spacing and equipment position in the roadway topology parameters L; Perform topologically constrained time-delay matching on each candidate feature pair in M, calculate the sequential relationship between visual disturbance leading, gas response lagging, thermal inertia persistence and acoustic vibration sudden change, and obtain the time-series coupling parameter D; Based on D-screening, pseudo-association feature pairs that do not meet the propagation path, response delay threshold, and co-occurrence duration threshold are removed, and the effective causal edge set E with visual anomaly-gas / thermal / acoustic vibration response chain transmission relationship is retained. A cross-modal causal consistency matrix C is generated by weighting the direction, intensity, and duration of E.
3. The coal mine disaster risk prevention and control platform based on multimodal perception and AI video recognition according to claim 2, characterized in that: Perform topologically constrained time-delay matching on each candidate feature pair in M, including the following steps: Based on the airflow direction, the distance between the sensors at both ends of the candidate feature pairs, and the equipment occlusion relationship in the roadway topology parameters L, a path constraint set P is constructed for each candidate feature pair, and a corresponding dynamic time-delay search window Δt is generated according to the path length and wind speed W. Within Δt, asymmetric sliding matching is performed on candidate feature pairs, and event anchoring is implemented for visual disturbance peak, gas concentration inflection point, temperature rise start point and acoustic pulse peak to obtain time-stamped sequence Q of each modal event; Based on Q-computation, visual lead time difference, gas lag time difference, thermal inertia duration and acoustic vibration instantaneous offset are calculated, and abnormal matches that do not meet the topological reachability, time difference upper limit and duration threshold are filtered out to obtain the effective temporal relation set H; H is sequentially encoded and weighted by confidence to generate the temporal coupling parameter D.
4. The coal mine disaster risk prevention and control platform based on multimodal perception and AI video recognition according to claim 3, characterized in that: Based on the temporal coupling parameter D, combined with the airflow channels, bifurcation connections and equipment obstruction locations in the roadway topology parameter L, a propagation path verification map G for candidate feature pairs is constructed to determine whether each candidate feature pair has an reachable propagation chain from upstream to downstream or from risk source to monitoring point. For candidate feature pairs that pass the propagation path verification, the visual lead time difference, gas lag time difference, thermal inertia duration, and acoustic vibration instantaneous offset in D are called and compared with the preset response delay threshold intervals respectively. Abnormal feature pairs with inverted timing, excessive lag, or instantaneous mismatch are eliminated. For feature pairs after time delay screening, the co-occurrence time period of the corresponding abnormal event is extracted, and the overlap duration between the visual abnormality maintenance interval and the gas response interval, thermal duration interval, or acoustic vibration sudden interval is calculated. Short-term occasional coupling pairs that do not meet the minimum co-occurrence duration threshold are eliminated. The remaining feature pairs are weighted and scored according to propagation reachability, time delay matching degree, and co-occurrence stability, and the effective causal edge set E is output.
5. The coal mine disaster risk prevention and control platform based on multimodal perception and AI video recognition according to claim 1, characterized in that: Generate a corresponding non-hazard disturbance baseline B, including: based on the temporal coupling parameter D, combined with the roadway topology parameter L, equipment operating status, and spray start / stop records, extract a reference sample set N from historical no-alarm periods that is consistent with the topological position and operating conditions of the current micro-scene unit U; cluster the disturbance sources of Fv, Fg, Ft, and Fa in N, and construct non-hazard disturbance prototype vectors Z for spray reflection, dust curtain sway, equipment idling vibration, and airflow fluctuations respectively; calculate the visual brightness amplitude range, gas fluctuation upper limit, thermal slow rise slope upper limit, and acoustic frequency band tolerance corresponding to each disturbance prototype based on Z, forming disturbance envelope parameters constrained by topological position and operating conditions; index and encode the disturbance envelope parameters according to the micro-scene unit U to generate a non-hazard disturbance baseline B that corresponds one-to-one with the current scene.
6. The coal mine disaster risk prevention and control platform based on multimodal perception and AI video recognition according to claim 5, characterized in that: Based on the current micro-scene unit U, the corresponding non-disaster disturbance baseline B is invoked, and Fv, Fg, Ft, and Fa are respectively mapped to the visual disturbance envelope, gas fluctuation envelope, thermal inertia envelope, and acoustic vibration tolerance envelope in B to obtain the baseline deviation set Δ for each mode. Combined with the effective causal edge set E, causal constraint residual decomposition is performed on each mode deviation in Δ, retaining the coupled deviation components that satisfy the chain relationship of subsequent gas response, thermal persistence, or acoustic vibration burst caused by visual anomalies, to obtain the candidate residual set Rc. Pseudo-anomaly stripping is performed on Rc, deducting the homologous components corresponding to spray reflection, dust curtain swaying, equipment idling vibration, and short-term airflow fluctuation, and filtering out transient residuals with durations lower than the minimum risk maintenance threshold to obtain the net residual set Rn. Rn is weighted and encoded according to modal contribution, temporal consistency, and spatial clustering to generate the residual risk feature R after pseudo-anomaly stripping.
7. The coal mine disaster risk prevention and control platform based on multimodal perception and AI video recognition according to claim 6, characterized in that: Using the initial mode, response mode, and edge weight parameters in the effective causal edge set E as constraints, the deviation set Δ is expanded by direction association to construct a visual deviation-gas deviation-thermal deviation / acoustic deviation residual transfer subgraph Gd, and isolated deviations not on the constraint path of E are marked as components to be removed. Based on the time delay parameters, intensity weights, and duration weights corresponding to each directed edge in Gd, the adjacent modal deviations in Δ are segmented and projected to separate the causal response components that satisfy the sequential triggering relationship and the synchronous disturbance components that do not satisfy the triggering order, resulting in the initial residual pair Rc1. The coupling consistency of Rc1 is checked according to the spatial adjacency relationship and temporal overlap interval within the micro-scene unit U, and residual chains that are continuously transmitted in the same local area and whose cross-modal overlap duration reaches the threshold are retained to obtain the candidate coupled residual set Rc2. Chain-based cumulative encoding is performed on Rc2, and the visual driving quantity, gas diffusion quantity, thermal retention quantity, and acoustic mutation quantity in each residual chain are superimposed according to the edge weights to output the coupled deviation components that satisfy the causal constraints, which serve as the candidate residual set Rc for subsequent pseudo-anomaly removal.
8. The coal mine disaster risk prevention and control platform based on multimodal perception and AI video recognition according to claim 7, characterized in that: The coupling consistency verification of Rc1 is performed according to the spatial adjacency relationship and temporal overlap interval within the micro-scene unit U, including the following steps: mapping the landing points of each residual pair in Rc1 according to the spatial grid position of the micro-scene unit U, establishing an adjacent grid connectivity table based on the roadway topology parameter L, and screening out candidate residual groups with spatial distances less than a preset adjacency threshold; extracting the start and end timestamps of the corresponding abnormal events for each candidate residual group, calculating the temporal overlap interval and overlap ratio between visual deviation, gas deviation, thermal deviation, or acoustic vibration deviation, and obtaining the spatiotemporal coupling parameters; performing consistency judgment on each candidate residual group based on the spatiotemporal coupling parameters, eliminating pseudo-coupled residual groups that are spatially adjacent but have excessive temporal misalignment, or that overlap temporally but cross non-connected boundaries, and retaining continuous transmission chains; jointly weighting the retained continuous transmission chains according to spatial adjacency strength and temporal overlap stability to generate a candidate coupling residual set Rc2 that satisfies the local co-domain and simultaneous coupling conditions.
9. The coal mine disaster risk prevention and control platform based on multimodal perception and AI video recognition according to claim 1, characterized in that: The linkage control module divides the target area into spatial nodes corresponding to micro-scene units U based on the roadway topology parameter L, and establishes directed propagation edges based on the roadway connectivity, airflow direction and equipment occlusion relationship between nodes to form a risk evolution diagram. After mapping R to the corresponding nodes, the earliest appearing node that continuously exceeds the minimum risk maintenance threshold is identified in chronological order as a candidate node for the risk source location, and the source point confidence score is calculated to determine the risk source location.
10. The coal mine disaster risk prevention and control platform based on multimodal perception and AI video recognition according to claim 9, characterized in that: The linkage control module starts from the location of the risk source, searches along the directed propagation edge to adjacent nodes to form candidate paths for propagation direction, and determines the propagation direction based on time-progressive scoring, intensity attenuation scoring, topology legality scoring, and spatial continuity scoring. After the location of the risk source and the propagation direction are determined, the risk level result Y is output, which includes the location of the risk source, the propagation direction, and the risk level.