Clustered mine-wide AI video analysis and linkage method
By constructing a shadow spatiotemporal feature dictionary and performing multi-view analysis, a confidence spectrum is generated, stable image features are extracted, and shadow misjudgments in mine video monitoring are eliminated. This achieves high stability and accuracy of the mine safety monitoring system and reduces the safety risks caused by misjudgments.
Patent Information
- Application Number
- CN202511455985.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-10-13
- Publication Date
- 2025-12-26
- Estimated Expiration
- 2045-10-13
AI Technical Summary
Existing mine video monitoring systems cannot distinguish between shadows and dynamic changes in the ground when mine cars pass by, leading to misjudgments that trigger large-scale evacuation orders, causing congestion among underground personnel and secondary accidents.
By constructing a spatiotemporal feature dictionary of shadows, combining multi-view geometric trajectories and polarization images, a shadow confidence spectrum is generated, stable image features are extracted, a feature fingerprint of misjudgment triggers is constructed, and consistency arbitration is performed in distributed nodes. Timestamp offset, viewpoint difference and light energy transition gradient are fused, the evacuation response threshold is calculated and the threshold is dynamically adjusted, a two-phase verification is introduced to eliminate misjudgment instructions, and finally the dynamic expansion of shadows is extinguished through a polarization metasurface illumination structure.
It significantly improves the stability and accuracy of the mine safety monitoring system, reduces the risk of ineffective evacuation and congestion caused by misjudgment, and enhances intelligent response capabilities.
Smart Images

Figure CN120932160B_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the field of artificial intelligence and mine safety monitoring, and particularly relates to a cluster type full-mine AI video analysis and linkage method. BACKGROUND
[0002] The "cluster type full-mine AI video analysis and linkage" refers to deploying multiple video monitoring points and intelligent analysis terminals in the full range of the mine, using artificial intelligence algorithms to identify and comprehensively analyze the information of personnel behavior, equipment operation state, environmental changes and the like in the monitoring picture in real time, and through the cluster computing and collaborative mechanism, the results of different monitoring points are interconnected and dynamically fused to form a unified risk perception network covering the full mine. On this basis, once an abnormal situation occurs in a certain area, the system can trigger linkage control in real time, such as alarm, start and stop of ventilation equipment, dispatch of personnel or linkage of other safety subsystems, to realize cross-regional intelligent response from dispersed monitoring to overall coordination, thereby significantly improving the real-time, global and intelligent level of mine safety management.
[0003] The prior art has the following disadvantages:
[0004] In the prior art, mine video monitoring relies on AI visual recognition algorithms to analyze the picture in real time for identifying sudden disasters such as collapse. However, when a mine car passes through the monitoring area, the car body will produce a large amount of stretching and fast moving shadow on the roadway wall or ground under strong light. This dynamic shadow is easily confused with the abnormal image features of stratum collapse in the video picture. Since the algorithm of the prior art cannot effectively distinguish between the transient change of the shadow and the real disaster signal of the actual stratum subsidence, it often leads to the system triggering a large-scale risk avoidance evacuation instruction by mistake. Under this misjudgment condition, a large number of underground workers will rush into the limited escape passage in a very short time, forming a short-time congestion, further causing secondary safety accidents such as stampede, thereby turning the misjudgment itself into a serious mine safety risk.
[0005] The above information disclosed in the background section is only intended to enhance the understanding of the background of the present disclosure, and therefore it can include information that does not constitute the prior art known to those of ordinary skill in the art. SUMMARY
[0006] The object of the present application is to provide a cluster type full-mine AI video analysis and linkage method to solve the problems in the background art.
[0007] In order to achieve the above object, the present application provides the following technical solution: a cluster type full-mine AI video analysis and linkage method, comprising the following steps:
[0008] Under the unified event time baseline, the light source modulation information and the mine car running trajectory are acquired, the time-space feature dictionary of the shadow is constructed, and the shadow profile reference for dynamic comparison is generated;
[0009] Based on the shadow profile reference, multi-view geometric trajectories are collected, polarization images and temperature images are fused, falling textures are stripped, and a shadow confidence spectrum is constructed;
[0010] Based on the shadow confidence spectrum, a counterfactual playback sequence is generated, image features that remain unchanged in multiple playbacks are extracted, and a misjudgment trigger feature fingerprint is formed;
[0011] Based on the feature fingerprint, consistency arbitration is performed in a distributed node, time stamp offset, view angle difference and light energy transition gradient are fused, consistency score is calculated, and a dynamic adjustment threshold of evacuation response threshold is generated;
[0012] Based on the dynamic adjustment threshold of evacuation response threshold, a linkage pre-confirmation chain is constructed, a two-phase verification window is introduced, a voice prompt is matched with a geomagnetic passage count, and a misjudgment instruction is eliminated;
[0013] Under the stable operation of the confirmation chain, a time reversal phase gate instruction is generated, a polarization metasurface lighting structure is driven to extinguish the dynamic expansion of the shadow, and a calibration matrix is written back to the event time baseline, forming a closed-loop control process.
[0014] Preferably, the shadow profile reference generation step is as follows:
[0015] Under the unified event time baseline, the light source modulation information and the mine car running trajectory data are acquired;
[0016] Based on the light source modulation information, the light intensity change, the beam diffusion angle and the illumination direction are collected, and the linear speed, acceleration, vehicle attitude and lighting projection area boundary of the mine car are combined;
[0017] The combination data of the light source and the mine car trajectory are used to construct the time-space path sample of the shadow, and the edge position and spatial form of the shadow in the continuous image frame are recorded;
[0018] According to the path samples generated by the combination of multiple light sources and mine cars, the time-space feature dictionary of the shadow is established, and the matching degree of the shadow area in the image and the dictionary path is compared in real-time image analysis. The shadow profile reference is dynamically generated for subsequent abnormal identification and judgment.
[0019] Preferably, the shadow confidence spectrum construction step is as follows:
[0020] Under the constraint of the shadow profile reference, multi-view images are collected and geometric trajectory information is extracted, spatial paths are calculated through edge detection and triangulation, and candidate shadow areas are calibrated;
[0021] Collecting polarization images and temperature images in the candidate shadow area, performing pixel-level alignment and fusing into three data streams of brightness, polarization and temperature;
[0022] Based on the edge intensity, polarization direction consistency and temperature gradient change rate of each image block in the fused image stream, joint determination is performed, and a confidence spectrum image is labeled with confidence;
[0023] The confidence spectrum image is output in time sequence, which is used for region confidence value extraction and recognition weight adjustment before subsequent image recognition.
[0024] Preferably, the feature fingerprint triggered by misjudgment is formed as follows:
[0025] Based on the shadow confidence spectrum, the time window with severe confidence value fluctuation is determined, and the corresponding vehicle and lighting parameters are extracted;
[0026] Based on the extracted behavior parameters, multiple counterfactual replay image sequences are reconstructed and lighting and temperature normalization processing is performed;
[0027] In each replay sequence, the stable structure area in the image is extracted and a stable image area feature table is established;
[0028] Based on the feature table, a three-layer expression structure of region structure index map, region content feature table and context association information set is constructed to form a feature fingerprint;
[0029] The feature fingerprint is embedded in the recognition judgment process for spatial matching comparison, which is used to identify misjudgment high-risk areas and suppress misjudgment trigger weight.
[0030] Preferably, the evacuation response threshold dynamic adjustment threshold is generated as follows:
[0031] Using the feature fingerprint as a constraint condition, image data consistent with the image frame time of the main recognition node is synchronously called in multiple monitoring nodes, and image feature matching is performed to construct an image matching response matrix;
[0032] Fusing the timestamp offset, perspective geometric difference and light energy transition feature of each node, the consistency score is calculated, and the consistency score is used as a quantitative basis for recognition confidence;
[0033] According to the level of consistency score, the determination threshold of evacuation response is automatically adjusted, and the threshold boundary is dynamically corrected combined with the historical score distribution curve, so as to realize the continuous optimization and adjustment of the trigger condition.
[0034] Preferably, the calculation of the consistency score takes time synchronization degree, spatial coincidence degree and lighting response consistency as indexes, respectively sets weights and performs weighted fusion, and when the score is lower than the preset lower limit, the voice prompt information and the passage count data are forcibly introduced for redundant comparison.
[0035] Preferably, the linkage pre-confirmation chain construction step is as follows:
[0036] Set the verification time window and start the pre-confirmation process within the scoring critical range, call the call center voice prompt signal and the geomagnetic sensor passage count information;
[0037] Extract the audio main frequency feature and the keyword content within the verification time window, complete the semantic matching, and synchronously analyze the passage frequency and passage direction generated by the geomagnetic induction to identify whether an effective personnel response behavior is formed;
[0038] The voice prompt signal and the passage count signal are used as decision inputs, and a joint judgment is made with the current score. If both signals pass the verification and the score is not lower than the trusted threshold, the linkage response is allowed to be executed, otherwise it is marked as a pending confirmation event.
[0039] Preferably, after the confirmation chain is stably running, the time reversal instruction is generated to drive the polarization metasurface lighting to extinguish the shadow expansion, and the calibration result is written back to the event time baseline step as follows:
[0040] According to the confirmation chain feedback result, the shadow core area is extracted and a spatial light field modulation target graph is generated to construct a target data for guiding the intervention waveform injection;
[0041] Start the programmable polarization metasurface lighting array to inject a phase conjugate micro-envelope waveform, and interfere to extinguish the dynamic expansion of the shadow;
[0042] Collect the post-intervention image and reconstruct the image feature vector to generate a lighting intervention response matrix and synchronously write it back to the unified event time baseline;
[0043] According to the response matrix, update the regional trustworthiness mapping table, and dynamically correct the future identification threshold.
[0044] In the above technical solutions, the technical effects and advantages provided by the present application are:
[0045] The application realizes the structural identification and confidence quantification of dynamic shadow for the first time by constructing a unified event time baseline and shadow profile reference, combining multi-view geometric trajectory acquisition, polarization and temperature information fusion processing; on this basis, the stable image features are extracted by using counterfactual replay to form misjudgment feature fingerprints, ensuring that the identification process has time consistency and image stability; through the cluster consistency arbitration mechanism, multi-source scores are calculated, and false trigger instructions are removed by combining a double on-site verification mechanism, realizing multi-dimensional confirmation of the identification result; finally, the polarization metasurface lighting structure is controlled by time reversal phase gating, actively extinguishing the dynamic growth of shadow at the physical level, realizing the breakthrough from image recognition to light field intervention, and completely eliminating the misjudgment source. The overall scheme constructs a closed-loop control chain from image recognition, result verification to active intervention and calibration write-back, greatly improving the stability, accuracy and intelligent response capability of the mine safety monitoring system, and significantly reducing the risk of invalid evacuation, congestion and secondary accidents caused by misjudgment. BRIEF DESCRIPTION OF DRAWINGS
[0046] In order to more clearly illustrate the technical solutions in the embodiments of the present application or the prior art, the drawings needed in the embodiments will be briefly introduced as follows. Obviously, the drawings in the following description only represent some embodiments described in the present application, and other drawings can also be obtained by those skilled in the art based on these drawings.
[0047] Figure 1 The method flowchart of the cluster type full mine AI video analysis and linkage method of the present application. DETAILED DESCRIPTION
[0048] Example implementations will now be described more fully with reference to the accompanying drawings. Example implementations, however, can be implemented in many different ways and should not be construed as limited to the examples set forth herein; rather, these examples are provided so that this disclosure will be thorough and complete, and will fully convey the inventive concept to those skilled in the art.
[0049] The present application provides a cluster type full mine AI video analysis and linkage method as shown in Figure 1 The method flowchart of the cluster type full mine AI video analysis and linkage method of the present application, comprising the following steps:
[0050] Under the unified event time baseline, the light source modulation prior information in the mine environment is acquired, the space-time feature dictionary of the shadow is established combined with the mine car running track data, and the corresponding relationship between the starting boundary and the termination boundary of the shadow is recorded in the dictionary, and the shadow profile reference for dynamic comparison is generated;
[0051] In the process of mine AI video analysis, to solve the problem of video misjudgment caused by dynamic interference of light and shadow, first of all, under the premise of unified event time baseline, a shadow profile reference for dynamic comparison needs to be constructed, which is based on the behavior of light source and the trajectory of mine car, and fuses space structure and time sequence. The specific implementation process is as follows:
[0052] A unified event time baseline is established, and prior information of light source modulation is collected. By deploying a time reference device in the ground control center, a high-precision time signal is generated, and the time signal is transmitted to all roadway video collection points through an optical fiber channel to ensure that the image frames collected by each monitoring device are consistent in millisecond-level time accuracy. After that, at each location where there is a lighting device, including underground ceiling lights, handheld maintenance lights, mine car headlights, rechargeable mobile light groups, and regional lighting equipment worn by personnel, the working voltage, brightness curve, on / off period, beam spread angle, and irradiation direction of the light source are recorded respectively. The specific method is to set up a multi-channel optical receiver near the device, and collect the irradiance intensity change value at a sampling interval of 20 milliseconds, and then calibrate the time stamp and normalize the reference brightness to obtain the brightness modulation curve of each light source in unit time. In addition, in order to restore the natural background light disturbance caused by non-human-controlled light changes, a three-axis light intensity detection device is deployed in an area of the mine without lighting equipment and vehicle passing, which continuously monitors the ambient light intensity and synchronously archives it as a baseline background for evaluating global brightness changes in subsequent modeling.
[0053] The trajectory data of the mine car is obtained and combined with the light source modulation information for analysis. An inertial measurement unit is installed on the mine car body, including a gyroscope and an accelerometer, and a real-time positioning system is configured, which receives the signal emitted by the positioning tag carried by the mine car through the ultra-wideband positioning base station to calculate the coordinate position of the mine car in the roadway space in real time. In each time slice, the linear speed, acceleration, running direction, body attitude, and steering angle of the mine car are recorded; combined with the lighting parameters of the overhead lighting fixture on the mine car, including the height of the fixture from the ground, the fixed angle, the light emission cone coverage range, and the direction vector of the light emitting surface center point, the accurate irradiation area boundary of the car light beam on the roadway wall, top and ground is calculated. By superimposing the trajectory change of the mine car and the moving trend of the light beam irradiation area, the shadow position change caused by the movement of the mine car can be calculated in each time period. Especially when the car speed suddenly changes or turns, the roll angle and yaw angle data should be additionally introduced to correct the lighting direction vector to eliminate the light beam deviation error caused by the inertia of the car body. Finally, the physical correspondence between the light source and the light receiving surface under the guidance of the mine car running is formed, providing a basis for subsequent shadow space boundary identification.
[0054] On the basis of combining the light source modulation curve with the mine car trajectory data, a shadow space-time feature dictionary is constructed. The dictionary is established according to a unified time baseline, and each light source-car combination is taken as an index unit. The time, illumination angle, irradiation surface feature and mine car attitude are set as four input conditions, and the spatial expansion path of the shadow in the image sequence is taken as the output content. The edge position coordinates of the shadow projected on the ground, wall and roof are recorded frame by frame. Through space-time compression coding of the shadow samples generated by multiple mine cars under different running speeds, turning radii, illumination intensities and roadway structure conditions, the shadow boundary changes of continuous frames are modeled as a continuous and predictable path in three-dimensional space. The path contains not only the starting boundary and ending boundary positions of the shadow, but also the specific form of the shadow change in each frame image between the starting and ending boundaries. In the specific coding process, in order to improve the recognition accuracy, each shadow path can be discretized into time interval points, and physical parameters such as actual illumination incident angle, roadway surface material reflectivity and surface inclination angle can be introduced as supplements to characterize the regularity of shadow deformation under different physical environments. In addition, in order to enhance the adaptability of the dictionary, an incremental acquisition strategy is adopted to automatically add new samples to the feature dictionary during the operation of the mine, and the weight of redundant data is reduced to maintain the accuracy and query efficiency of the dictionary.
[0055] Using the established shadow space-time feature dictionary, a shadow profile reference that can be used for real-time dynamic comparison is generated. In real-time image stream analysis, whenever a large boundary shift, sudden brightness drop or spatial texture change region appears in the monitoring image, the corresponding time baseline and vehicle trajectory state of the current image frame are immediately called to retrieve the corresponding space-time path sample in the shadow feature dictionary. By comparing the boundary position of the to-be-identified region in the real-time image with the typical shadow path recorded in the dictionary, the deviation value of the two in terms of spatial distribution and time advancement trend is calculated. If the deviation value is below the set dynamic error tolerance threshold, the deformation in the current picture is determined to be a normal shadow evolution process, and it is excluded from the subsequent abnormal identification process; if the deviation exceeds the threshold, or the starting position of the current shadow is obviously inconsistent with the light source propagation direction, it is recorded as a potential abnormal image for the next stage of processing. The shadow profile reference in this step combines physical light behavior, mine car space attitude change and time sequence image evolution in the construction process, realizes predictable modeling of the shadow behavior, significantly improves the accuracy and real-time performance of shadow misjudgment exclusion in video monitoring without increasing the consumption of too many computing resources, and provides a reliable reference basis for subsequent accurate identification of abnormal events.
[0056] Under the constraint of the shadow profile reference, geometric trajectory information from multiple monitoring angles is collected, and polarization image data and temperature image data are fused to separate the falling texture information similar to rock stratum collapse in the image, and a shadow confidence spectrum with quantitative reliability is constructed for continuous input of subsequent abnormal identification judgment;
[0057] After the completion of the shadow profile reference based on the unified event time baseline, in order to realize the accurate distinction between the shadow and the abnormal event image in the image recognition stage, a shadow confidence spectrum image with quantifiable reliability is constructed by introducing multi-view geometric analysis and multi-domain information fusion for subsequent continuous calling, and the specific operation includes the following steps:
[0058] According to the established shadow profile reference, image data of multiple fixed monitoring points with overlapping space positions of predicted shadow trajectories is collected, and geometric trajectory information in the corresponding time window is extracted. In the roadway structure, according to the lighting dead angle, the high-frequency vehicle entry and exit points and the geographical distribution of the possible misjudgment area, high-sensitivity cameras with fixed focal length, wide-angle lenses and image synchronous clocks are symmetrically deployed along the two sides and the top of the roadway, ensuring that there are at least three groups of overlapping angle acquisition devices every two meters. In order to establish the geometric correspondence of three-dimensional space, each group of devices needs to be spatially calibrated before deployment through a set of fixed reference materials, including a standard plate with known side length and high-contrast pattern. The intrinsic and extrinsic matrices of each camera are obtained through front shooting, rotating shooting and multi-angle imaging. Subsequently, during the shadow prediction period, image frames are synchronously extracted from multiple cameras, and the extractable structural elements such as mine car outlines, light boundaries, ground feature lines and support steel frames in the image coordinate system are subjected to edge detection, and the projection path of the object edge in the three-dimensional physical coordinate system is calculated by the binocular triangulation algorithm. In these space paths, any part that has a trajectory overlap with the starting boundary, the expansion path and the termination position recorded in the shadow profile reference is marked as a candidate shadow area, which is the processing range for subsequent spectral information recognition and locking.
[0059] In the candidate shadow area, polarization image data and temperature image data are synchronously collected and registered and fused. At each fixed video acquisition position, two types of special imaging devices are arranged adjacent to the visible light imaging device: one is a polarization imaging device with a omnidirectional linear polarization receiving array, which can obtain the main polarization direction and polarization degree value of the area at a specific time by placing polarization plates of different angles in front of each pixel subarray; the other is a high-sensitivity infrared thermal imaging device, which uses a non-cooled thermopile infrared sensor to continuously collect the radiation intensity changes of each image area in the eight to fourteen micrometer band and converts it into a temperature pseudo-color image in real time. After the collection of the two types of images, the polarization image and the thermal imaging image are geometrically registered along the space principal axis direction; then the three types of images (visible light, polarization map and thermal map) are aligned in the pixel-level space coordinates through image resampling. After fusion, in each candidate area, three three-dimensional matrices are constructed, corresponding to the brightness change distribution, the polarization direction mutation distribution and the temperature gradient change distribution, respectively, and arranged in time frame sequence to form a perception data stream along the time axis.
[0060] Based on the fused image perception data stream, texture recognition, feature analysis and layer-by-layer stripping processing are performed on the image blocks falling into the shadow prediction area to extract the reliable shadow area and exclude the high-risk texture features similar to the rock stratum falling but not real events. In specific operation, first, edge gradient calculation is performed on each image block in the visible light image to extract the boundary with high contour intensity; then, the polarization direction consistency test is performed on the polarization image at the corresponding position to determine whether there are multiple pixel sub-regions with vibration direction mutation or scattering direction difference; further, the temperature intensity change rate of the position point in the temperature image is combined to calculate whether the temperature of each image block on the time axis has an abrupt change exceeding the set threshold. The above three criteria are logically combined to determine that if the image block boundary contour is consistent with the shadow profile change, the polarization consistency is high and the temperature curve is smooth, it is determined as normal shadow behavior and is assigned a high confidence value; if the boundary in the image block is irregular, the polarization direction is disorderly and the local temperature drops sharply, it is marked as a potential misjudgment source area and is assigned a low confidence value. This confidence value is marked in the image coordinates in units of pixels and a single-frame shadow confidence spectrum image is generated in the form of a heat map. In order to enhance the spatial continuity and boundary smoothness, spatial Gaussian filtering and time median filtering are performed on the initial confidence spectrum to remove sudden noise or abnormal jump blocks.
[0061] The shadow confidence spectrum image generated by the processing is output in time sequence form as an input signal for subsequent image anomaly recognition, and participates in real-time linkage triggering decision weight adjustment. In actual use, before each frame of main monitoring image enters the recognition algorithm process, the corresponding shadow confidence spectrum image of the current frame is first called to map the position and extract the confidence value of the to-be-recognized area in the main image frame. If the confidence value of the to-be-recognized area in the corresponding position of the confidence spectrum image is higher than the confidence threshold, the recognition weight is automatically reduced to reduce the probability of entering the abnormal reporting path; if the confidence value of the area in the confidence spectrum shows a downward trend and there are corresponding jump features in the heat map and polarization map, the weight is maintained or enhanced, and the area is recorded as an abnormal area to be judged. By introducing the confidence spectrum as a pre-intervention means, the image feature interference caused by light and shadow dynamic evolution can be effectively eliminated before the recognition algorithm, and the false positive rate caused by light and shadow is reduced from the source.
[0062] Based on the constructed shadow confidence spectrum, a counterfactual playback sequence is generated to replay the monitoring picture under the same speed and lighting conditions as the simulation and actual conditions, extract the stable image features that remain unchanged in multiple playback, and form the feature fingerprint of misjudgment triggering;
[0063] On the basis of the constructed shadow confidence spectrum image sequence, in order to eliminate the influence of light dynamic or scene occasional changes on abnormal recognition and judgment, a counterfactual playback mechanism is proposed, which reproduces multiple monitoring pictures close to the actual conditions, extracts stable image features that repeatedly appear and are not affected by disturbances, and generates misjudgment trigger feature fingerprints with uniqueness and discriminability, including the following steps:
[0064] Based on the constructed shadow confidence spectrum, the time window with greater misjudgment risk in the target image sequence is determined, and the behavior parameters and light state information of the mine car in the corresponding time period are extracted. For the frame sequence with dramatic confidence value fluctuations, frequent spatial region boundary mutations, and large deviation from the expected shadow path in the confidence spectrum, the mine car running speed, motion direction, attitude angle, vehicle coordinate position, and vehicle body light source opening and closing state, light beam projection angle, illumination intensity change curve are extracted one by one according to the unified event time baseline index corresponding to each image. The data interval of five seconds before and after the time point is extended by 0.2 seconds to form a complete playback window of ten seconds. In this way, each group of image frames is associated with accurate vehicle state and light behavior, providing data basis for constructing simulation replay.
[0065] Based on the time window and the collected behavior parameters, a series of representative counterfactual playback sequences are constructed, and the image evolution process consistent with the actual scene is reconstructed. First, select the key parameter points generated by the mine car in actual operation, including uniform straight line, high speed passing, turning acceleration, lighting surge, light shielding, etc. Multiple typical combined states are selected as input boundary conditions for simulation replay. Under each simulation condition, the image space reconstruction algorithm is used to superimpose the current set of light projection path and shadow evolution trajectory according to the order of the original image frames collected by the actual monitoring device, and the image change process generated by the joint action of light and vehicle under the current parameter combination is reproduced. In each playback, the visible light image sequence, polarization image sequence and temperature image sequence are also restored, and their one-to-one correspondence in spatial coordinates, time stamps and viewing angles is maintained. In order to control the interference of external environmental variables on the recognition effect, all playback images are standardized in brightness based on the background brightness of the actual sampling time point during the reproduction process, and the image frames are color normalized and edge smoothed to ensure that each simulation result can be used for subsequent comparison.
[0066] After the reconstruction of the multi-group counterfactual replay image sequence is completed, the image structure features appearing in the image are extracted frame by frame, and the stability of their appearance in all replay rounds is statistically analyzed. The specific method is: taking the image frame as the unit, identifying the structural texture area existing in it, including the side projection contour of the mine car, the reflection area of the fixed object of the shaft wall, the fixed crack line on the ground, the shadow transition zone of the support beam, etc., and based on the consistency of edge gradient direction, the stability of brightness change and the temperature field average change range, the change degree of each region is quantitatively scored. When a region shows unchanged edges, approximately the same brightness curve, and no significant fluctuations in the temperature field under at least five different replay conditions, the region is recorded as a "stable structure area". At the same time, in each replay, the key parameters of the region such as spatial position, boundary shape, gray texture, and thermal feature curve are encoded and archived to form a stable image area feature table.
[0067] Based on the extracted stable image area feature table, the corresponding misjudgment trigger feature fingerprint is constructed, and the spatial matching model is constructed and the expression method is coded. Specifically, all regions that meet the stability conditions are spatially related to determine their relative distribution pattern, region overlap degree and adjacent relationship in the image coordinates, forming a two-dimensional spatial structure diagram; then generate a description vector for each region in the dimensions of gray distribution, texture direction, thermal imaging response, etc. and enter it as an attached data item in the structure diagram. In order to enhance the scene adaptability of the feature fingerprint, the position time of the first appearance of the image structure, the corresponding vehicle operating conditions, the current light disturbance type and the then confidence spectrum score and other behavior context data need to be recorded as a mapping chain of the causal relationship between the identification feature fingerprint and the misjudgment scene. The final completed feature fingerprint has a three-layer expression structure: first, the region structure index map, which indicates the spatial layout of all stable regions; second, the region content feature table, which encodes all image property indicators; third, the context association information set, which records environmental parameters and trigger logic. This structure can be used as a reference template for subsequent identification processes.
[0068] The generated feature fingerprint is written into a stability check sub-process in the identification and judgment process, which is used to identify the possible misjudgment source area in advance. In the actual image processing process, before each frame of input image enters the linkage judgment process, it is first matched and compared with the constructed feature fingerprint structure in space, and its area position, structural characteristics and environmental context information are compared. If the local area in the current image is highly consistent with a known misjudgment feature fingerprint, and its confidence spectrum score is relatively low, the area is determined as a misjudgment high-risk area, and a judgment inhibition weight is applied in the linkage decision stage to prevent false response triggered by such structure. Through the above mechanism, the present application not only realizes the inversion simulation and feature attribution of typical misjudgment behavior, but also can deposit it as high-risk reference information in the identification database, forming a complete closed-loop process from image reconstruction, stability evaluation, feature extraction to scene fingerprint modeling.
[0069] Using the feature fingerprint as a constraint condition, cluster consistency arbitration is performed in the distributed node, the timestamp offset, perspective geometric difference and light energy transition gradient of different monitoring nodes are fused, the mutual consistency score between multi-source inputs is calculated, and the dynamic adjustment threshold of the evacuation response threshold is generated according to the consistency score;
[0070] After completing the feature fingerprint extraction of the stable image area in the counterfactual playback, in order to ensure that the identification and judgment have cross-node credibility support, a cluster consistency arbitration method based on feature fingerprint constraint condition is proposed, and on this basis, a dynamic adjustment threshold of the response threshold is generated, which is used to adapt the linkage control precision management in the multi-node cooperative environment. The specific process includes the following steps:
[0071] Based on the constructed image feature fingerprint structure, a multi-monitoring node image calling and matching processing flow is started according to a unified event time baseline. In the mine environment, the fixed monitoring points at different positions receive a unified time signal through a preset time synchronization mechanism, ensuring that the image acquisition time accuracy is controlled within milliseconds. Taking the abnormal trigger frame detected by the main recognition node as a reference, all other monitoring points with overlapping coverage are searched, and the image data collected in the same time period is called from these points. Such image data includes but is not limited to visible light images, polarization images, and thermal imaging images, each frame of which carries spatial position identification and timestamp data. After the image calling is completed, the constructed feature fingerprint is applied to the corresponding area of each node image one by one, and a one-to-one matching comparison of image structure, texture edge, gray distribution, and temperature response is performed in units of image blocks. Each matching result produces a corresponding matching index set, including structure coincidence degree, gray similarity, texture direction consistency, and thermal image response overlap degree, and these data are filled into the fingerprint matching response matrix for consistency analysis in the next step. In this way, cross-area verification of the feature fingerprint in multiple node images can be realized, ensuring that the judgment result is not caused by occasional image disturbance at a certain fixed angle.
[0072] For the image matching response results of each monitoring point, multi-dimensional consistency score calculation is performed in combination with the timestamp offset, geometric difference of viewing angle, and light energy transition characteristics to determine whether the current feature has cross-node consensus. In the time dimension, the timestamp data attached to each image frame is first extracted, and the reference time offset value relative to the main recognition node is calculated to determine whether there is acquisition delay, synchronization loss, or other timing abnormalities. An offset value less than five milliseconds is considered to have high timing consistency, an offset value between five milliseconds and twenty milliseconds requires timing compensation, and an offset value exceeding twenty milliseconds is marked as time mismatch. In the spatial dimension, based on the physical position, lens orientation, and relative height data recorded during installation of the monitoring points, combined with known roadway structure information, a monitoring image space projection relationship is established to calculate the deformation degree of the same feature area in the image coordinate system under different viewing angles. Through affine transformation and image registration algorithm, the projection overlap area of the area in each viewing angle image is obtained, and the larger the overlap area, the higher the viewing angle consistency. In the energy dimension, the illumination curve generated by the change of the light source over time is extracted, and the brightness transition slope of the matching area in the image gray response is superimposed. The light energy transition point position is identified through curve fitting to determine whether the transition is caused by real lighting change or by shadow obstruction or mine car shading. Finally, the time synchronization degree, spatial coincidence degree, and lighting response consistency are respectively assigned weights, and the cluster consistency score of the feature fingerprint in the current recognition window is calculated, with the score range being 0 to 1, and the higher the value, the stronger the reliability of the recognition event.
[0073] After obtaining the consistency score, a dynamic threshold adjustment mechanism for evacuation response based on the current score value is established. This mechanism automatically adjusts the minimum decision value required to trigger the evacuation response according to the spatial support degree of the identified event, ensuring that the system can make differentiated responses when facing different levels of credible images. In the case where the score is higher than 0.85, the threshold value can be automatically lowered, and the identified event enters the response execution phase without the need for additional verification; when the score is between 0.60 and 0.85, the system keeps the threshold unchanged and prepares to introduce a redundant confirmation mechanism for further comparison and confirmation; when the score is lower than 0.60, the system raises the threshold and forcibly calls independent information sources such as voice announcements, personnel passage counts, and video stream frame dynamic change trends as additional comparison conditions to prevent low-consistency events from misleading the trigger link. To improve the response continuity of the dynamic threshold, a set of historical threshold regression curves is also set in this step, which is based on the statistical distribution of consistency scores and actual trigger results in the previous hour of identified events, dynamically adjusting the score interval boundaries to prevent frequent oscillation of the threshold due to frequent abnormal events. Finally, a dynamic response control logic closely coupled with the recognition credibility is formed, which can automatically adjust the sensitivity of the response trigger mechanism according to the consistency of the cluster node data structure, achieving two-way optimization of accurate response and false alarm suppression.
[0074] Based on the dynamic adjustment of the evacuation response threshold, a pre-link confirmation chain is constructed, with the consistency score as the input signal, and a millisecond-level two-phase verification time window is introduced, in which the voice prompt signal of the call station and the passage count information of the geomagnetic sensor are verified simultaneously to eliminate single-source false alarm instructions;
[0075] After the cluster consistency score and dynamic response threshold adjustment are completed, to avoid single-point false alarm triggering mine safety linkage behavior, a pre-link confirmation chain based on the score input is introduced, which uses a cross-domain dual-signal verification mechanism to perform redundant confirmation of the identification result at the physical layer. This confirmation chain operates within a millisecond time range and is driven by artificial voice instructions and personnel passage data, and includes the following steps:
[0076] After the consistency score is calculated and the dynamic response threshold is generated, the pre-confirmation process is started and a verification time window is set. After the current image recognition event is analyzed by image structure analysis, multi-angle comparison and light energy transition judgment, if its consistency score is within the critical range of the dynamic threshold (for example, between 0.65 and 0.85), the system enters the pre-confirmation process. In order to complete effective confirmation without delaying the response of the linkage, the system starts a verification time window based on the completion time of the current score, and the duration of the window is not more than 500 milliseconds. During this period, the system does not immediately start any physical linkage device, but waits for confirmation information from two independent signal channels, namely the real-time voice prompt signal from the dispatching call station and the personnel passage count information collected by the geomagnetic sensor. The length of the time window is determined by modeling the underground broadcast response delay, sensor data return period, and network delay tolerance, ensuring real-time and fault tolerance under actual production conditions.
[0077] During the verification time window, the voice prompt signal and the passage count signal are accurately collected and content verified to form the dual verification logic of the pre-linkage confirmation. In terms of voice verification, the dispatcher makes a voice announcement to the area where the recognition event occurs through the broadcast terminal. Common prompt contents include "please avoid", "abnormality occurs", "numbered area alarm", etc. The voice content is captured by the sound collector deployed at the roadway node at the same time as it is transmitted to the underground loudspeaker. The sound collector digitizes the received audio data, extracts its main frequency features and keyword content, and combines the number information of the event recognition area for semantic comparison and matching. If the relevant key prompt sentence related to the current recognition event is clearly identified within the verification time window, and the content structure matches the system preset template, the voice prompt is considered as an effective confirmation signal. At the same time, in the roadway section involved in the recognition event, the geomagnetic sensor installed at the ground passage main line also starts to collect the magnetic field disturbance data caused by the passing personnel. Each time a metal object (such as a miner's lamp, tools, safety helmet, etc.) passes through the geomagnetic induction coil, the device generates an electromagnetic change event with a time stamp accurate to milliseconds. The system compares the passage data within the time window with the historical average passage rate of the area, and if it detects that the passage frequency increases more than three times the average baseline within 500 milliseconds, and the passage direction is consistent with the evacuation direction, it is determined that the personnel response behavior has actually occurred, constituting the second confirmation signal.
[0078] After the two confirmation signals complete acquisition and verification, they are used as decision inputs combined with the current consistency score to jointly determine whether to allow the execution of the linkage instruction. If both the voice signal and the passage count signal meet the time window, semantic content, location number, and direction path verification conditions, and the current event score is not lower than the score confidence threshold (for example, not lower than 0.65), the current event is determined to be an effective linkage trigger event with image recognition basis and behavior response support, and the system will allow the execution of response instructions, such as controlling local ventilation, turning off power in a specific section, issuing an escape voice alarm, or dispatching personnel to evacuate, etc. If only one of the two confirmation signals is established, or neither of them meets the verification conditions, the recognition event is marked as "to be confirmed" and does not enter the linkage execution process. The system will continue to track the subsequent image frame evolution process and score changes, and re-initiate the scoring and verification process in the next recognition cycle. This mechanism ensures that every time the linkage response is executed, it has undergone a three-step authentication process based on image analysis, consistency verification, and physical behavior dual signal support, greatly reducing the risk of false trigger.
[0079] Under the condition that the linkage pre-confirmation chain runs stably, time reversal phase gate instructions are generated according to the feedback results of the confirmation chain to control the programmable polarization metasurface lighting structure to inject phase conjugate micro-envelope waveforms into the shadow center area of the monitoring picture, extinguish the dynamic expansion characteristics of the shadow, and write the calibrated matrix after lighting intervention to the unified event time baseline synchronously, forming a closed-loop reversible dynamic control process.
[0080] After the voice prompt and personnel passage data complete double verification, the recognition event is confirmed to be valid, and the active light field intervention link is entered. This step uses micro-envelope lighting waveforms with phase conjugate characteristics to extinguish the dynamic shadow area that poses a risk of false trigger in real time, and writes the intervention results to the time baseline to achieve closed-loop control. The specific steps are as follows:
[0081] After completing the voice prompt verification and geomagnetic count verification in the linkage pre-confirmation chain, and successfully judging that the current recognition event is a high-confidence event, the shadow core area of the recognition event in the image picture is extracted. This area is identified by the system during image processing and has complete edge information, dynamic evolution direction, and texture continuity. The system first reads the gray distribution curve, gradient direction graph, texture fracture position, illumination reflection path, and historical change trajectory of the area at the time of recognition, and combines the light source modulation prior information and mine car trajectory data to construct a "spatial light field modulation target graph" that accurately describes the current shadow structure. Each pixel unit in this target graph contains angle response, historical phase interference trajectory, and current wavefront disturbance trend, which is used to guide the next step of precise light field injection process. This data structure matches the high similarity fragments in the historical shadow dictionary to obtain the best intervention waveform model suitable for the current time and space conditions.
[0082] According to the target map, the phase driving units of the corresponding area in the programmable polarization metasurface lighting array are started, and the corresponding control panel translates the intervention waveform parameters of the target area into the phase delay, electrically controlled polarization direction and incident angle combination of a plurality of microscopic lighting units, generating a complete intervention matrix signal. The polarization metasurface is composed of independently driven electro-optical regulation units, each of which changes the scattering direction and phase delay capability of its nanostructure to light waves by applying voltage, thereby realizing real-time regulation of light waveforms at the microscopic scale. The system drives several lighting units to emit micro-envelope waveforms with phase conjugation characteristics synchronously according to the modulation frequency, envelope width, interference angle and phase inversion information of the required injection waveform, so that they are injected into the shadow formation core area along the optical reflection track in the reverse path, and destructively interfere with the original shadow wavefront, instantaneously canceling its intensity field. In order to ensure complete coverage of the intervention range, the system will perform a comparative analysis of the image frames within 25 milliseconds before and after the lighting area, confirm whether the original gray abnormal area has been restored to the average brightness of the environment at the image processing level, and if not, automatically perform a phase fine-tuning and injection again.
[0083] After the completion of the lighting intervention operation and the image change meets the shadow extinguishing condition, the system immediately collects the latest image frame and enters the resampling and feature reconstruction process. This image frame takes the intervention area as the core and performs pixel-level reanalysis on its gray distribution, texture arrangement direction and boundary contour clarity, and forms a new set of feature vectors after reconstruction. The system compares the set of vectors with the feature vectors of the corresponding area of the image before intervention, and generates a difference map before and after intervention. This difference map is used to judge: 1. whether the intervention area is completely extinguished from the original light and shadow disturbance characteristics; 2. whether new unexpected feature disturbances are introduced; 3. whether the intervention effect has an overflow effect in the adjacent area. After the comparison is completed, the intervention result is recorded in the form of "lighting intervention response matrix", including intervention start and end time, intervention waveform parameters, image feature difference before and after, shadow stripping success rate, light offset angle and inversion success identifier, etc. Detailed content, and the matrix is used as a new calibration factor, bound to the unified time baseline of the current event, replaces the feature state before intervention, and is called in the subsequent image recognition process.
[0084] The response matrix after the lighting intervention is linked and compared with the historical feature evolution model, the regional credibility mapping table of the image recognition system is updated, and finally the closed-loop regulation and control mechanism of recognition-verification-intervention-feedback is completed. The credibility mapping table is a dynamic atlas generated based on the intervention response results of the historical shadow area, which is used to reflect the potential misjudgment risk and processed state of each image area in the future recognition period. The identification trigger threshold of the intervention successful area is appropriately adjusted to reduce the probability of repeated triggering of the alarm in the same area due to residual disturbance in a short period of time. At the same time, for the area where the intervention fails to completely remove the misjudgment feature, it is set to "enhanced monitoring state", and higher frequency image analysis is included in the future monitoring period, and the voice and pass dual verification mechanism is preferentially triggered. At this point, the image recognition is no longer a linear judgment process, but a high-coupling closed-loop dynamic regulation and control mechanism from image input -> shadow recognition -> event verification -> active intervention -> image feedback -> calibration rewriting is established, and the engineering closed-loop management of the shadow misjudgment problem in the mine complex environment is truly realized.
[0085] The present application realizes the structural identification and confidence quantification of dynamic shadow for the first time by constructing a unified event time baseline and shadow profile reference, combining multi-view geometric trajectory acquisition, polarization and temperature information fusion processing. On this basis, the stable image features are extracted by using counterfactual playback to form the misjudgment feature fingerprint, ensuring that the identification process has time consistency and image stability. Through the cluster consistency arbitration mechanism, multi-source scores are calculated, and the false trigger instructions are removed by combining the double real-time verification mechanism, realizing multi-dimensional confirmation of the identification result. Finally, the polarization metasurface lighting structure is controlled by the time reversal phase gate, which actively extinguishes the dynamic growth of shadow, realizes the breakthrough from image recognition to light field intervention, and completely eliminates the misjudgment source. The overall scheme constructs a closed-loop control chain from image recognition, result verification to active intervention and calibration rewriting, greatly improves the stability, accuracy and intelligent response ability of the mine safety monitoring system, and significantly reduces the risk of invalid evacuation, congestion and secondary accidents caused by misjudgment.
[0086] The above only describes certain exemplary embodiments of the present application by way of illustration, and it is needless to say that those skilled in the art can modify the described embodiments in various ways without departing from the spirit and scope of the present application. Therefore, the above figures and description are illustrative in nature and should not be understood as limiting the scope of protection of the claims of the present application.
Claims
1. A cluster type of mine-wide AI video analysis and linkage method, characterized in that, Comprise the following steps: Under the unified event time baseline, obtain light source modulation information and mine car running track, construct shadow space-time feature dictionary, generate shadow profile reference for dynamic comparison; Based on the shadow profile reference, collect multi-view geometric trajectory, fuse polarization image and temperature image, strip falling texture, construct shadow confidence spectrum; Based on the shadow confidence spectrum, generate counterfactual replay sequence, extract image features that remain unchanged in multiple replays, form feature fingerprints of misjudgment triggering; Based on the feature fingerprint, perform consistency arbitration in distributed nodes, fuse timestamp offset, view angle difference and light energy transition gradient, calculate consistency score, generate evacuation response threshold dynamic adjustment threshold; Based on the evacuation response threshold dynamic adjustment threshold, construct the linkage front-end confirmation chain, introduce the biphasic verification window, match the voice prompt and the geomagnetic pass count, and eliminate the misjudgment instruction; Under the stable operation of the confirmation chain, generate time reversal phase gate instruction, drive the polarization metasurface lighting structure to extinguish the shadow dynamic expansion, and write the calibration matrix to the event time baseline.
2. The cluster type full-mine AI video analysis and linkage method according to claim 1, characterized in that, The shadow profile reference generation step is as follows: Under the unified event time baseline, obtain light source modulation information and mine car running track data; Based on the light source modulation information, collect the light intensity change, the light beam diffusion angle and the illumination direction, and combine the linear speed, acceleration, vehicle attitude and illumination projection area boundary of the mine car; Use the combined data of light source and mine car track to construct the space-time path sample of shadow, and record the edge position and spatial form of shadow in continuous image frames; According to the path sample generated by multiple light sources and mine cars, establish the space-time feature dictionary of shadow, and compare the matching degree of shadow area in image and dictionary path in real-time image analysis, dynamically generate shadow profile reference for subsequent abnormal identification and judgment.
3. The cluster type full-mine AI video analysis and linkage method according to claim 2, characterized in that, The shadow confidence spectrum construction step is as follows: Under the constraint of shadow profile reference, collect multi-view images and extract geometric trajectory information, calculate space path and label candidate shadow area through edge detection and triangulation; In the candidate shadow area, collect polarization image and temperature image, perform pixel-level alignment and fuse into three data streams of brightness, polarization and temperature; Based on the edge intensity, polarization direction consistency and temperature gradient change rate of each image block in the fused image stream, perform joint determination and label confidence to generate confidence spectrum image; Output the confidence spectrum image in time sequence for subsequent region confidence value extraction and recognition weight adjustment before image recognition.
4. The cluster type full-mine AI video analysis and linkage method according to claim 3, characterized in that, The feature fingerprint formation step of misjudgment triggering is as follows: Determine the time window with violent confidence value fluctuation based on the shadow confidence spectrum and extract the corresponding vehicle and illumination parameters; Based on the extracted behavior parameters, reconstruct multiple counterfactual replay image sequences and perform illumination and temperature normalization processing; Extract the stable structure area in each replay sequence and establish the stable image area feature table; Based on the feature table, construct a three-layer expression structure of region structure index graph, region content feature table and context association information set, form a feature fingerprint; Embed the feature fingerprint into the recognition and judgment process for spatial matching comparison, identify misjudgment high-risk areas and suppress misjudgment triggering weight.
5. The cluster type full-mine AI video analysis and linkage method according to claim 4, characterized in that, The evacuation response threshold dynamic adjustment threshold generation step is as follows: Synchronize the image data of the monitoring nodes with the image frame time of the main identification node, and perform image feature matching to construct an image matching response matrix, using the feature fingerprint as a constraint condition; Fuse the timestamp offset, perspective geometric difference, and light energy transition characteristics of each node, calculate the consistency score, and use the consistency score as a quantitative basis for identification reliability; Automatically adjust the judgment threshold of the evacuation response according to the level of the consistency score, and dynamically correct the threshold boundary combined with the historical score distribution curve.
6. The cluster type full-mine AI video analysis and linkage method according to claim 5, characterized in that, The calculation of the consistency score takes time synchronization, spatial coincidence, and light response consistency as indicators, sets weights and performs weighted fusion, and when the score is lower than the preset lower limit, forcibly introduces voice prompt information and passage count data for redundant comparison.
7. The cluster type full-mine AI video analysis and linkage method according to claim 5, characterized in that, The steps of the pre-linkage confirmation chain construction are as follows: Set the verification time window and start the pre-confirmation process within the score critical range, call the call center voice prompt signal and the passage count information of the geomagnetic sensor; Extract the audio main frequency feature and keyword content within the verification time window, complete semantic matching, and synchronously analyze the passage frequency and passage direction generated by the geomagnetic induction to identify whether an effective personnel response behavior is formed; Take the voice prompt signal and passage count signal as decision input, and jointly judge with the current score. If both signals pass the verification and the score is not lower than the reliability threshold, the linkage response is allowed to be executed, otherwise it is marked as a pending event.
8. The cluster type full-mine AI video analysis and linkage method according to claim 7, characterized in that, After the stable operation of the confirmation chain, the time reversal instruction drives the polarization metasurface lighting to extinguish the shadow expansion, and the calibration result is written back to the event time baseline steps as follows: According to the feedback result of the confirmation chain, extract the shadow core area and generate a spatial light field modulation target graph, and construct a target data for guiding the injection of intervention waveforms; Start the programmable polarization metasurface lighting array to inject phase-conjugate micro-envelope waveforms, and interfere to extinguish the dynamic expansion of the shadow; Collect the post-intervention image and reconstruct the image feature vector, generate the lighting intervention response matrix, and synchronously write back to the unified event time baseline; Update the regional reliability mapping table according to the response matrix, and dynamically correct the future identification threshold.
Citation Information
Patent Citations
Room state detection method and system based on image recognition
CN120612737A
Generating maps without shadows using geometry
US20190295315A1