Method and system for monitoring coal mine through linkage of 3D and video
Through the modularly designed 3D and video linkage monitoring method, the efficient integration of video and sensor data is achieved, and the problems of low accuracy, high cost and poor real-time performance of coal mine monitoring systems in the existing technology are solved, improving the efficiency and safety of coal mine production safety management.
Patent Information
- Application Number
- CN202510382346.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-28
- Publication Date
- 2025-08-12
AI Technical Summary
In the prior art, the coal mine monitoring system has low accuracy, high cost, poor real-time performance, low efficiency and low safety, making it difficult to achieve an effective combination of three-dimensional spatial perception and video surveillance, resulting in limited efficiency and accuracy of coal mine safety production management.
The modularly designed 3D and video linkage monitoring method is adopted. Through real-time data acquisition of monitoring modules and sensor modules, the video data processing module performs feature matching and depth estimation, combined with the sensor data processing module to perform data fusion, dynamically update the 3D model, and realizes adaptive fusion and efficient collaborative processing of video and sensor data.
It realizes the efficient combination of video and sensor data, improves positioning accuracy, security and real-time, reduces system costs, and is suitable for all-round three-dimensional monitoring of complex coal mine environments, provides three-dimensional situational awareness with second-level response, and reduces operation and maintenance costs.
Smart Images

Figure CN120472084A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the general field of image data processing or generation technology, and in particular to image data processing or generation. Background Art
[0002] In the field of coal mine safety monitoring, the iterative development of digital and intelligent technologies is crucial. The latest generation of 3D and video-linked monitoring systems for coal mines is gaining increasing attention. They integrate advanced 3D imaging and video surveillance technologies, providing powerful support for efficient monitoring of underground spaces and work areas in coal mines. Industry concerns have arisen regarding the difficulties of data integration and the lack of intuitive information presentation in traditional monitoring systems, which limit the efficiency and accuracy of coal mine safety management.
[0003] For example, Chinese patent publication number CN117333406A discloses a method for dynamic weighted fusion of multi-source sensor data in underground coal mines. The method provides the following technical solutions. This invention belongs to the field of intelligent coal mine technology, specifically relating to a method for dynamic weighted fusion of multi-source sensor data in underground coal mines. An image enhancement algorithm combining single-parameter homomorphic filtering and histogram equalization in HSV space is added to the visual image preprocessing process to enhance the brightness and contrast of underground images. A Mahalanobis distance-based consistency detection method is used to assess sensor data quality, detect sensor data degradation, and adaptively select sensor data suitable for the current environment for effective fusion. A LiDAR / IMU / Camera factor graph model is constructed based on the key parameters of each sensor. A dynamic combination model for multi-source sensor data weights is constructed based on data quality to dynamically adjust the weights of the sensor data fusion factors. Compared to LVI-SAM, this method achieves higher pose estimation accuracy and robustness, with a trajectory root mean square error of 0.19m. The average direct point cloud comparison distance between the point cloud map and the one stitched together using gantry-mounted 3D laser scanning is less than 0.13m, meeting the accuracy requirements for mine robot positioning and mapping, and providing theoretical reference and technical support for intelligent mining and safety inspections in coal mines. However, the aforementioned SLAM method for dynamically weighted fusion of multi-source sensor data in underground coal mines only involves sensor data processing, making it difficult to integrate 3D spatial perception and video surveillance technologies to achieve comprehensive, three-dimensional monitoring of the coal mine's internal environment. It also fails to combine surveillance video and sensor data to generate real-time, updated 3D images. Summary of the Invention
[0004] The present invention solves the problems of low precision, high cost, poor real-time performance, low efficiency and low safety in the existing technology, and proposes a method and system for 3D and video linkage monitoring of coal mines, achieving the purpose of combining monitoring video and sensor data, high safety, accurate positioning, good real-time performance, low cost and high efficiency.
[0005] To achieve the above object, the present invention adopts the following technical solutions: A method for monitoring a coal mine using 3D and video linkage includes the following steps: S1: The monitoring module and sensor module collect corresponding data in real time and transmit it to the backend; S2: The video data processing module completes 3D position mapping through feature matching and camera calibration, and then performs depth estimation. The sensor data processing module and the video data processing module process data synchronously; S3: By detecting, classifying and processing inconsistencies between video data and sensor data, adaptive fusion data is mapped into 3D space; S4: Create and dynamically update the status of the 3D model and display it on the front end.
[0006] The benefit of this design is that it enables efficient collaborative processing of multi-source data through modular division of labor and automated process design. The synchronous processing mechanism effectively avoids data time sequence misalignment, and the dynamic update strategy significantly reduces computing resource consumption, ensuring the real-time and stable operation of the system.
[0007] A system for monitoring coal mines through 3D and video linkage includes: a front-end and back-end separated architecture; the front-end displays monitoring images and 3D images, and the back-end includes a processing module and an acquisition module, the acquisition module includes a monitoring module and a sensor module, the processing module includes a video data processing module, a sensor data processing module and a data fusion module, the monitoring module is connected to the video data processing module, the sensor module is connected to the sensor data processing module, and the video data processing module and the sensor data processing module send the completed data to the data fusion module, and the data fusion module is connected to the 3D generation module.
[0008] Preferably, in step S2, the video data processing module includes the following steps: S2.1: Read the video stream through the video decoder, extract each frame in the video according to the timestamp, and perform preprocessing operations on each frame of the surveillance video; S2.2: Use feature detection algorithms to extract significant feature points in each frame. By performing convolution and pooling on each part of the video frame, a feature map is generated to identify key objects and their locations. S2.3: Use stereo vision methods to perform depth estimation. Calculate the disparity between two frames captured by the camera and then calculate the three-dimensional depth of the object. S2.4: The depth information obtained by stereo vision or monocular depth estimation is converted into 3D world coordinates using the camera’s intrinsic parameter matrix based on the pixel coordinates in the image, the optical center of the camera, and the focal length of the camera. The 2D coordinates of each pixel in the image are mapped to 3D space.
[0009] The benefit of this design is that it establishes a complete video data processing chain, forming a closed loop from data decoding to 3D reconstruction. A timestamp synchronization mechanism ensures frame sequence integrity, convolutional pooling enhances feature extraction, and a dual-path depth estimation strategy accommodates the needs of diverse scenarios, significantly improving the accuracy and adaptability of 3D modeling.
[0010] Preferably, the step S3 includes the following steps: S3.1: Data acquisition and initialization, specifically including extracting the 3D coordinates of objects from video frames and obtaining the 3D coordinates of objects from sensors; S3.2: Calculate the Euclidean distance between the three-dimensional coordinates of the video data and the sensor data, and set an inconsistency threshold; compare the Euclidean distance with the inconsistency threshold. If the Euclidean distance is greater than the inconsistency threshold, it is determined that there is an inconsistency between the video data and the sensor data, and the process proceeds to step S3.3. If the Euclidean distance is less than or equal to the inconsistency threshold, it is determined that the video data and the sensor data are consistent, and the process proceeds to step S4. S3.3: Set different levels of inconsistency thresholds, and classify them into slight inconsistency, moderate inconsistency, and severe inconsistency based on the Euclidean distance and their size; S3.4: Adaptively fuse the video and sensor 3D coordinates based on the inconsistent classification results; S3.5: Use Kalman filtering to smooth the final 3D coordinates and use an incremental update mechanism to update only the parts of the model that have changed; S3.6: Refine and optimize the 3D model using the Laplacian optimization algorithm and multi-resolution technology. Use the Laplacian optimization method to reduce the amount of calculation by minimizing the smoothness and error of the model surface to obtain the optimized 3D model.
[0011] The benefit of this design is that it builds an intelligent data fusion system, implementing differentiated processing strategies through a multi-level threshold judgment mechanism. Kalman filtering, combined with incremental updates, ensures data continuity while reducing computational load. Laplacian optimization effectively balances model accuracy and computational efficiency, forming a complete quality assurance chain.
[0012] Preferably, step S3.4 includes, for cases of slight inconsistency, using adaptive weight fusion to dynamically adjust the weights of video and sensor data so that data sources with larger deviations obtain lower weights, and the weight calculation is based on the distance between the data sources, specifically including dynamic weight calculation and weighted fusion calculation; for cases of moderate inconsistency, using Kalman filtering to correct the data sources with larger deviations, that is, the three-dimensional coordinates of the video data and the sensor data; for cases of severe inconsistency, directly ignoring the data sources with larger deviations, and only using data sources with higher consistency.
[0013] The benefit of this design lies in its innovative hierarchical processing mechanism, which employs optimal solutions for different deviation levels. A dynamic weight allocation algorithm ensures flexibility in data fusion, while Kalman filtering effectively eliminates sensor drift errors. An abnormal data rejection mechanism ensures the reliability of system output, forming a multi-layered data credibility assurance system.
[0014] Preferably, in step S4, when the video processing is completed, the data fusion module is notified to integrate the data, and finally sent to the 3D module to generate a 3D model, by bringing together the three-dimensional coordinates of each pixel to form point cloud data, and by triangulating the point cloud, the point cloud is converted into a three-dimensional mesh model; when the sensor data processing is completed, the data fusion module will be notified to integrate the data, and the status of the 3D model will be dynamically updated according to the sensor data.
[0015] The benefit of this design is that it establishes an event-driven model update mechanism and optimizes resource allocation through an asynchronous notification system. Point cloud triangulation technology balances model detail with computational efficiency, and a dynamic update strategy avoids the waste of full rendering resources, significantly improving system responsiveness.
[0016] Preferably, in step S2.1, the preprocessing operation includes noise removal, image enhancement and color correction; in step S2.3, the method of depth estimation using stereo vision method is that the depth of the object is equal to the product of the focal length of the camera and the baseline distance between the two cameras and the quotient of the parallax in the image.
[0017] The benefit of this design is that it establishes a standardized pre-processing pipeline and improves subsequent processing quality through multi-stage image optimization. A formulated depth calculation method ensures quantifiable measurement accuracy, and a parametric design facilitates system calibration and optimization, providing a reliable data foundation for 3D reconstruction.
[0018] Preferably, in step S2, the sensor data processing module performs data preprocessing on the data, specifically including denoising, interpolation and standardization processing.
[0019] The benefits of this design include establishing a sensor data quality assurance system, eliminating environmental interference through multi-stage filtering, compensating for data transmission delays through interpolation algorithms, and ensuring compatibility of multi-source data through standardization, creating a unified benchmark for data fusion.
[0020] Preferably, the step S1 includes: the monitoring module collects video data through real-time video monitoring, transmits it to the backend via RTSP for real-time frame image processing and object recognition, and the data collected by the sensor module is transmitted to the backend via the Internet of Things.
[0021] The advantage of this design is that it uses professional protocols to ensure data transmission quality, the RTSP protocol ensures low-latency transmission of video streams, the IoT architecture enables flexible expansion of sensor networks, and the dual-channel design takes into account both data real-time and reliability.
[0022] Preferably, the sensor module includes a gas sensor, a temperature and humidity sensor, and a pressure sensor. The data collected by the sensor module is transmitted to the sensor data processing module through the Internet of Things; the monitoring module performs real-time video monitoring and transmits the video data to the video data processing module through a video transmission protocol; the back-end integrates and processes the data and pushes it to the front-end for display.
[0023] Compared with the prior art, the present invention has the following beneficial effects.
[0024] 1. This invention constructs a three-dimensional monitoring network by innovatively integrating video surveillance data with multi-type sensor data. The video processing module uses advanced stereo vision algorithms to achieve millimeter-level depth perception, while networked monitoring of sensors such as gas, temperature, and humidity covers all aspects of environmental parameters. Through a visual perception fusion algorithm, the system can intelligently identify and correct data deviations, achieving centimeter-level positioning accuracy for the three-dimensional model. This multi-source data complementarity mechanism effectively overcomes the limitations of a single monitoring method, particularly in the complex environment of coal mines, and can accurately identify potential danger areas.
[0025] 2. This invention adopts an incremental model update strategy, automatically identifying data update areas through a change detection algorithm and re-rendering only the local model that has changed. Compared with the full update method, this mechanism significantly reduces the computational load. Combined with the smoothing processing of the Kalman filter, the image can be kept smooth even in the case of network fluctuations. This design is particularly suitable for coal mine tunnel extension scenarios, reflecting the latest status of the excavation face in real time and providing emergency command with three-dimensional situational awareness with a response time of seconds.
[0026] 3. By introducing the Laplacian optimization algorithm and multi-resolution modeling technology, the present invention significantly reduces the computing resource requirements while ensuring the accuracy of the model. The amount of data in the optimized grid model is greatly reduced, and the GPU memory usage is reduced, so that conventional servers can support the three-dimensional modeling needs of large-scale coal mines. The adaptive weight allocation mechanism dynamically adjusts the processing strategy according to the credibility of the data, and can also operate stably in underground substations with limited hardware resources. This intelligent resource management method reduces the overall operation and maintenance costs of the system and is particularly suitable for promotion and implementation in coal mining enterprises. BRIEF DESCRIPTION OF THE DRAWINGS
[0027] Figure 1 The present invention is a flow chart of a method for monitoring coal mines using 3D and video linkage.
[0028] Figure 2 This is a timing diagram of a 3D and video linkage system of a method for monitoring coal mines using 3D and video linkage according to the present invention. DETAILED DESCRIPTION
[0029] To make the objectives, technical solutions, and advantages of the present disclosure more apparent, embodiments of the present disclosure are described in further detail below with reference to the accompanying drawings. The proportions of the components herein are not drawn to scale, and the proportions and dimensions shown in the accompanying drawings are not intended to limit the essential technical solutions of the present disclosure. These embodiments do not describe all details in detail, nor do they limit the present disclosure to the specific embodiments described.
[0030] See also Figure 1-2 As shown, a method for monitoring a coal mine by 3D and video linkage includes the following steps: S1: The monitoring module and sensor module collect corresponding data in real time and transmit it to the backend; S2: The video data processing module completes 3D position mapping through feature matching and camera calibration, and then performs depth estimation. The sensor data processing module and the video data processing module process data synchronously; S3: By detecting, classifying and processing inconsistencies between video data and sensor data, adaptive fusion data is mapped into 3D space; S4: Create and dynamically update the status of the 3D model and display it on the front end.
[0031] A system for monitoring coal mines through 3D and video linkage includes: a front-end and back-end separated architecture; the front-end displays monitoring images and 3D images, and the back-end includes a processing module and an acquisition module, the acquisition module includes a monitoring module and a sensor module, the processing module includes a video data processing module, a sensor data processing module and a data fusion module, the monitoring module is connected to the video data processing module, the sensor module is connected to the sensor data processing module, and the video data processing module and the sensor data processing module send the completed data to the data fusion module, and the data fusion module is connected to the 3D generation module.
[0032] This invention combines three-dimensional spatial perception technology with video surveillance technology to achieve comprehensive, three-dimensional monitoring of the internal environment of coal mines. Through intelligent algorithms and real-time data processing, this invention can accurately monitor and analyze every corner of the coal mine, providing timely and accurate safety warnings and management decision support. The application of this invention will significantly improve coal mine production safety, reduce accident risks, and protect the lives and property of miners.
[0033] like Figure 1 In one embodiment shown, Figure 1 The present invention is a flow chart of a method for monitoring coal mines using 3D and video linkage.
[0034] This paper proposes a coal mine monitoring method that combines 3D technology with video surveillance technology. Through multi-level and multi-dimensional data integration, it is possible to achieve accurate monitoring of the coal mine environment. The workflow of the method of the present invention can be divided into the following main steps: S1: Real-time data collection and transmission between monitoring module and sensor module In practical applications in coal mines, monitoring systems primarily use cameras and various sensors to collect video and environmental data. The monitoring module captures dynamic conditions within the mine through real-time video surveillance and transmits this data to back-end systems via the RTSP protocol (or other video transmission protocols). Simultaneously, various sensors within the mine (such as temperature, gas, humidity, and pressure sensors) collect environmental data in real time and transmit it to the back-end system via the Internet of Things (IoT). During this process, all transmitted data is transmitted in real time to the back-end processing module, ensuring that the back-end system receives all dynamic information from the coal mine site immediately.
[0035] S2: Video Data Processing and 3D Position Mapping In the back-end system, the video data processing module first processes the real-time video data. Specifically, this process includes the following steps: S2.1 Video Frame Extraction and Preprocessing: Through the video decoder, the backend system extracts each frame based on the timestamp in the video stream. Each frame is first preprocessed with noise removal, image enhancement, and color correction to improve the accuracy of subsequent processing.
[0036] S2.2 Feature point extraction and camera calibration: Use feature detection algorithms to extract significant feature points in each frame of the image, and generate feature maps through convolution and pooling operations to identify key objects in the image and their locations.
[0037] S2.3 Depth estimation: Depth estimation is performed using stereo vision methods. The disparity between two frames of images obtained from two cameras is calculated, and the three-dimensional depth of the object is calculated using a formula.
[0038] S2.4 Three-dimensional coordinate mapping: Using the camera intrinsic parameter matrix, combined with information such as the pixel coordinates in the image, the focal length and optical center of the camera, the two-dimensional coordinates of each pixel in the image are mapped to three-dimensional space to obtain the position of the object in three-dimensional space.
[0039] Through the above processing steps, the video data is converted into three-dimensional position information that can reflect the mine environment, providing the necessary data support for the subsequent creation of the three-dimensional model.
[0040] S3: Fusion and update of video and sensor data The fusion of video data and sensor data is one of the core steps of the present invention. Since there may be some inconsistencies between video data and sensor data during the acquisition process, these data must be rigorously fused to ensure the accuracy and stability of the 3D model.
[0041] S3.1 Data Acquisition and Initialization: First, the 3D coordinates of the object are extracted from the video data, and the corresponding 3D coordinates are obtained through the sensor module. These data will be initialized and passed to the subsequent processing module.
[0042] S3.2 Inconsistency Detection: Calculate the 3D Euclidean distance between the video data and the sensor data and set an inconsistency threshold. If the Euclidean distance is greater than the set threshold, the data is determined to be inconsistent and the next step of processing is carried out.
[0043] S3.3 Inconsistency classification and handling: Based on the degree of inconsistency, it is divided into three situations: slight inconsistency, moderate inconsistency and severe inconsistency, and different handling strategies are adopted: For slight inconsistencies, adaptive weight fusion is used; For moderate inconsistency, Kalman filtering is used to correct data with large deviations; For serious inconsistencies, data sources with large deviations are directly ignored, and only data with high consistency are retained.
[0044] S3.4 Adaptive Fusion: Dynamically adjusts the fusion weights of video and sensor data based on the inconsistency classification results. For minor inconsistencies, the system adaptively adjusts the weight ratios. For moderate inconsistencies, Kalman filtering effectively corrects data sources with significant deviations, while severe inconsistencies are simply ignored.
[0045] S3.5 Incremental Update Mechanism and Data Smoothing: An incremental update mechanism is used to update only the changed parts to reduce the system's computational workload. At the same time, a Kalman filter is used to smooth the three-dimensional coordinates to ensure data continuity and stability.
[0046] S4: Creation and dynamic update of 3D models After data fusion is complete, the system generates a 3D model based on the synthesized 3D coordinates. By aggregating the 3D coordinates of each pixel, point cloud data is generated. Point cloud data contains the 3D coordinates of the points in the scene and additional information (such as color and depth). Triangulation techniques are then used to convert the point cloud data into a 3D mesh model.
[0047] To further optimize the model, the system employs the Laplacian optimization algorithm and multi-resolution technology to refine and refine the 3D model. This minimizes surface smoothness and errors, reducing computational effort and resulting in an optimized 3D model. This model is dynamically updated based on real-time sensor data, ensuring it consistently reflects the mine's current conditions.
[0048] The system part of the present invention specifically includes: a front-end and back-end separated architecture; the front-end displays monitoring images and 3D images, the back-end includes a processing module and an acquisition module, the acquisition module includes a monitoring module and a sensor module, the processing module includes a video data processing module, a sensor data processing module and a data fusion module, the monitoring module is connected to the video data processing module, the sensor module is connected to the sensor data processing module, and the video data processing module and the sensor data processing module are sent to the data fusion module after completion, and the data fusion module is connected to the 3D generation module. The sensor module includes a gas sensor, a temperature and humidity sensor and a pressure sensor, and the data collected by the sensor module is transmitted to the sensor data processing module via the Internet of Things; the monitoring module performs real-time video monitoring and transmits the video data to the video data processing module via a video transmission protocol; the back-end integrates and processes the data and pushes it to the front-end for display.
[0049] The present invention combines video monitoring with real-time acquisition and in-depth processing of sensor data, innovatively fuses video data with sensor data, and realizes three-dimensional and precise monitoring of the underground environment of coal mines. First, through real-time data acquisition of video monitoring and sensors, combined with depth estimation and three-dimensional position mapping, the monitoring system can accurately locate and track each key object in real three-dimensional space, thereby providing a more intuitive decision-making basis for the safety management of coal mines. Secondly, an adaptive fusion algorithm is adopted, which can dynamically adjust the weights of video and sensor data, and adjust the data source according to the degree of inconsistency, thereby ensuring the efficiency and stability of data fusion, and avoiding the error problems that are prone to occur in traditional methods. By introducing technologies such as incremental update and Kalman filtering, the system can effectively reduce the amount of calculation, improve real-time performance, and avoid unnecessary repeated rendering. Finally, the application of Laplacian optimization algorithm and multi-resolution technology further improves the accuracy and details of the three-dimensional model, making the dynamic monitoring of the coal mine working environment more accurate and reliable. In general, this invention not only improves the efficiency and accuracy of coal mine safety monitoring, reduces manual intervention, and reduces the risk of coal mine accidents, but also significantly improves the system's real-time response capability and sustainability by optimizing the utilization of computing resources, providing an innovative and efficient solution for safety management in the coal mining industry.
[0050] like Figure 2 In one embodiment shown, Figure 2 This is a timing diagram of the 3D and video linkage system for a method of monitoring coal mines using 3D and video linkage. This method uses a separate front-end and back-end architecture. The back-end is responsible for collecting video and sensor data, processing the data, and generating dynamically updated 3D images, while the front-end is responsible for displaying the monitoring and 3D images.
[0051] Real-time video surveillance is performed using cameras installed on-site at the coal mine. Video data is transmitted via RTSP (or other video transmission protocols) to the backend for real-time frame image processing and object recognition. Once integrated, the video data is pushed to the frontend for display.
[0052] Data collected by various sensors at coal mines (such as gas sensors, temperature and humidity sensors, and pressure sensors) is transmitted to a back-end system via the Internet of Things. This back-end system has two modules: a video data processing module and a sensor data processing module, which process data synchronously.
[0053] After receiving the video stream data, the video data processing module will first perform the following operations: 1. Video frame extraction and correction: The video stream is read through a video decoder, and each frame in the video is extracted according to the timestamp. In each frame of the surveillance video, preprocessing operations such as noise removal, image enhancement and color correction are first performed to improve the accuracy of subsequent image recognition and depth estimation.
[0054] I′Preprocessing(I) Where I is the original video frame and I′ is the processed image.
[0055] 2. Feature matching and camera calibration Object recognition in video frames relies on feature extraction and matching. Feature detection algorithms extract significant feature points from each frame. By performing convolution and pooling on each portion of the video frame, a feature map is generated, identifying key objects and their locations.
[0056] Obj i =CNN(I′) Among them, Obj i is the label and position coordinates of the detected i-th object, and its coordinates on the image plane are (u i ,υ i ).
[0057] 3. Depth Estimation Depth estimation is performed using stereo vision. The disparity between two frames captured by the camera is calculated to estimate the three-dimensional depth of the object. Assuming the disparity between the two frames is δ, the depth Z of the object can be calculated using the following formula: Where f is the focal length of the camera, B is the baseline distance between the two cameras, and δ is the disparity in the image.
[0058] 4. 3D Position Mapping Depth information obtained through stereo vision or monocular depth estimation can be used to map the two-dimensional coordinates of each pixel in the image to three-dimensional space. Assuming a pixel point [u, υ] in the image has a depth value of Z, it can be converted to three-dimensional world coordinates [X, Y, Z] using the camera's intrinsic parameter matrix.
[0059] Where (u,υ) is the pixel coordinate in the image, (c x , c y ) is the optical center of the camera, f x , f y is the focal length of the camera, and Z is the depth value.
[0060] When video processing is complete, the data fusion module is notified to integrate the data and ultimately send it to the 3D module to generate a 3D model. By aggregating the 3D coordinates of each pixel, point cloud data is formed. A point cloud is a dataset consisting of numerous 3D points in a scene, each with its own coordinates and additional information (such as color and depth). By triangulating the point cloud, it can be converted into a 3D mesh model.
[0061] The sensors in the coal mine (such as temperature, gas concentration, pressure, etc.) provide a wealth of environmental information that can be used to dynamically update the state of the 3D model. After receiving the sensor data, the sensor data processing module will perform the following operations: The received sensor data may contain noise, such as fluctuations in temperature sensors or errors in vibration sensors. The sensor data needs to be denoised, interpolated, and normalized. Assume that there are multiple sensor data S1, S2, ..., S n Each sensor has a position P1, P2, ..., P in space n The sensor data is first normalized and filtered: S′ i =Normalize(Filter(S i ))When the sensor data processing is completed, the data fusion module will be notified to integrate the data.
[0062] To ensure the consistency of video data and sensor data in a 3D scene, the present invention designs a visual perception fusion algorithm (VPFA), which is described as follows: When video and sensor data are inconsistent, directly performing weighted averaging fusion can lead to data distortion or bias. To address this inconsistency, the present invention designs a more adaptive fusion algorithm that dynamically adjusts the weights of video and sensor data, even excluding anomalous data when necessary, to more accurately reflect the real situation.
[0063] This algorithm detects discrepancies between video and sensor data, identifying and addressing outliers for more intelligent fusion. When inconsistencies are detected between video and sensor data, the system dynamically adjusts the weight of each data source based on the degree of inconsistency, or even ignores one data source to ensure the stability and accuracy of the fusion results.
[0064] Algorithm steps: Step 1: Data acquisition and initialization 1. Video data initialization: Extract the 3D coordinates of the object (x video ,y video , z video ) 2. Sensor data initialization: Get the three-dimensional coordinates of the object (x sensor ,y sensor , z sensor ), where z sensor Indicates the depth measured by the sensor.
[0065] Step 2: Inconsistency Detection Calculate the Euclidean distance d between the 3D coordinates of the video data and the sensor data to detect inconsistencies: Set an inconsistency threshold δ. If d>δ, it is determined that there is inconsistency between the video data and the sensor data.
[0066] Step 3: Classify and handle inconsistencies According to the degree of inconsistency, it is divided into three situations and adopts different processing strategies: 1. Slight inconsistency (δ1<d≤δ2): If there is a slight inconsistency, directly adjust the weight ratio for fusion; 2. Moderate inconsistency (δ2<d≤δ3): Use Kalman filtering to correct data sources with large deviations (such as Z coordinates); 3. Severe inconsistency (d>δ3): Directly ignore data sources with large deviations and only use data sources with high consistency.
[0067] Step 4: Adaptive Fusion Based on the inconsistent classification results, the video and sensor 3D coordinates are adaptively fused.
[0068] Slight inconsistency: Adaptive weight fusion For minor inconsistencies, the weights of video and sensor data are dynamically adjusted so that data sources with larger deviations receive lower weights. The weights are calculated based on the distance d between the data sources: 1. Dynamic weight calculation: Where α is the adjustment coefficient.
[0069] 2. Weighted fusion calculation: x fused =w video ·x video +w sensor ·x sensor y fused =w video ·y video +w sensor ·y sensor z fused =wvideo ·z video +u sensor ·z sensor .
[0070] Moderate inconsistency: Kalman filter correction When the inconsistency is large, Kalman filtering is used to correct the data source with large deviation. Assume that the three-dimensional coordinates of the video data are The sensor data is Here ^ represents the predicted value.
[0071] Kalman filter formula: Similarly, the x and y coordinates are corrected.
[0072] Serious inconsistency: Ignore abnormal data When a serious inconsistency is detected, the data source with large deviations is directly ignored and only the data source with high consistency is used: The same process is applied to y fused and z fused .
[0073] Step 5: Smoothing of fusion results After adaptive fusion, the final three-dimensional coordinates are smoothed using Kalman filtering to ensure the continuity of the fused data.
[0074] Smoothed 3D coordinates: in, is the final smoothed 3D coordinate.
[0075] After calculating the three-dimensional coordinates, in order to improve real-time performance, an incremental update mechanism is used to update only the parts of the model that have changed. Assuming that the model at the current time t is M(t), if the video or sensor data changes, the updated model is M(t+1). The updated part can be calculated by calculating the difference: ΔM(t, t+1)=M(t+1)-M(t) Incremental rendering is performed based on this difference, avoiding the need to completely re-render the entire 3D scene each time. After the sensor data and video data are fused, they are sent to the 3D module for mapping.
[0076] 5. Mapping fused data to 3D space After receiving the data, the 3D module refines and optimizes the 3D model by introducing the Laplacian optimization algorithm and multi-resolution technology. Using the Laplacian optimization method, the smoothness and error of the model surface are minimized to reduce the amount of calculation, resulting in an optimized 3D model.
[0077] Among them, v i is the position of the i-th grid point, The optimized target position. When the 3D module finishes processing the data for the first time or there is a subsequent data update, the front end will automatically refresh the display.
[0078] In summary, the present invention combines surveillance video and sensor data, and uses multiple advanced algorithms to process them to generate real-time updated 3D images. Through dynamic data analysis and visual perception fusion algorithms, combined with scene generation and update mechanisms, higher accuracy and lower computing costs are achieved, thereby efficiently realizing the linkage between 3D and video. Through innovative data processing processes, the present invention enables video surveillance and sensor data to be accurately connected and generate a three-dimensional visualization model. In addition, the use of an algorithmic mechanism of "updating the 3D model only when there are changes" ensures efficient use of system resources and reduces unnecessary calculations. The main benefits of the present invention are: 1. Accurate positioning: Combined with three-dimensional spatial location information, the monitoring area can be accurately located, which helps to quickly respond to and handle emergencies; 2. Efficiency improvement: Enhance the efficiency of the monitoring system, reduce human resource investment, and improve management efficiency and accuracy; 3. Improve safety: Improve the safety management level of coal mines, reduce the probability of accidents, and ensure the safety of miners.
[0079] The present invention is not limited to the above-mentioned embodiments. Regardless of any changes in shape or material composition, any structural design provided by the present invention is a variation of the present invention and should be considered within the scope of protection of the present invention.
Claims
1. A method for monitoring coal mines by 3D and video linkage, characterized in that: The following steps are involved: S1: The monitoring module and sensor module collect corresponding data in real time and transmit it to the backend; S2: The video data processing module completes 3D position mapping through feature matching and camera calibration, and then performs depth estimation. The sensor data processing module and the video data processing module process data synchronously; S3: By detecting, classifying and processing inconsistencies between video data and sensor data, adaptive fusion data is mapped into 3D space; S4: Create and dynamically update the status of the 3D model and display it on the front end.
2. The method for monitoring a coal mine by 3D and video linkage according to claim 1, characterized in that: In step S2, the video data processing module includes the following steps: S2.1: Read the video stream through the video decoder, extract each frame in the video according to the timestamp, and perform preprocessing operations on each frame of the surveillance video; S2.2: Use feature detection algorithms to extract significant feature points in each frame. By performing convolution and pooling on each part of the video frame, a feature map is generated to identify key objects and their locations. S2.3: Use stereo vision methods to perform depth estimation. Calculate the disparity between two frames captured by the camera and then calculate the three-dimensional depth of the object. S2.4: The depth information obtained by stereo vision or monocular depth estimation is converted into 3D world coordinates using the camera’s intrinsic parameter matrix based on the pixel coordinates in the image, the optical center of the camera, and the focal length of the camera. The 2D coordinates of each pixel in the image are mapped to 3D space.
3. A method for monitoring a coal mine by 3D and video linkage according to claim 1 or 2, characterized in that: The step S3 includes the following steps: S3.1: Data acquisition and initialization, specifically including extracting the 3D coordinates of objects from video frames and obtaining the 3D coordinates of objects from sensors; S3.2: Calculate the Euclidean distance between the three-dimensional coordinates of the video data and the sensor data, and set an inconsistency threshold; compare the Euclidean distance with the inconsistency threshold. If the Euclidean distance is greater than the inconsistency threshold, it is determined that there is an inconsistency between the video data and the sensor data, and the process proceeds to step S3.
3. If the Euclidean distance is less than or equal to the inconsistency threshold, it is determined that the video data and the sensor data are consistent, and the process proceeds to step S4. S3.3: Set different levels of inconsistency thresholds, and classify them into slight inconsistency, moderate inconsistency, and severe inconsistency based on the Euclidean distance and their size; S3.4: Adaptively fuse the video and sensor 3D coordinates based on the inconsistent classification results; S3.5: Use Kalman filtering to smooth the final 3D coordinates and use an incremental update mechanism to update only the parts of the model that have changed; S3.6: Refine and optimize the 3D model using the Laplacian optimization algorithm and multi-resolution technology. Use the Laplacian optimization method to reduce the amount of calculation by minimizing the smoothness and error of the model surface to obtain the optimized 3D model.
4. The method for monitoring a coal mine by 3D and video linkage according to claim 3, characterized in that: The step S3.4 includes, for cases of slight inconsistency, using adaptive weight fusion to dynamically adjust the weights of video and sensor data so that data sources with larger deviations obtain lower weights. The weight calculation is based on the distance between the data sources, specifically including dynamic weight calculation and weighted fusion calculation; for cases of moderate inconsistency, using Kalman filtering to correct the data sources with larger deviations, that is, the three-dimensional coordinates of the video data and the sensor data; for cases of severe inconsistency, directly ignoring the data sources with larger deviations, and only using data sources with higher consistency.
5. A method for monitoring a coal mine by 3D and video linkage according to claim 1, 2 or 4, characterized in that: In step S4, when the video processing is completed, the data fusion module is notified to integrate the data, and finally sent to the 3D module to generate a 3D model. By collecting the three-dimensional coordinates of each pixel together, point cloud data is formed, and the point cloud is triangulated to convert it into a three-dimensional mesh model. When the sensor data processing is completed, the data fusion module will be notified to integrate the data, and the status of the 3D model will be dynamically updated according to the sensor data.
6. The method for monitoring a coal mine by 3D and video linkage according to claim 2, characterized in that: In step S2.1, the preprocessing operation includes noise removal, image enhancement and color correction; in step S2.3, the method of depth estimation using stereo vision method is that the depth of the object is equal to the product of the focal length of the camera and the baseline distance between the two cameras and the quotient of the parallax in the image.
7. The method for monitoring a coal mine by 3D and video linkage according to claim 5, characterized in that: In step S2, the sensor data processing module performs data preprocessing on the data, specifically including denoising, interpolation and standardization.
8. The method for monitoring a coal mine by 3D and video linkage according to claim 7, characterized in that: The step S1 includes: the monitoring module collects video data through real-time video monitoring, transmits it to the backend via RTSP for real-time frame image processing and object recognition, and the data collected by the sensor module is transmitted to the backend via the Internet of Things.
9. A system for monitoring coal mines by 3D and video linkage, using the method for monitoring coal mines by 3D and video linkage according to any one of claims 1 to 8, characterized in that: include: Front-end and back-end separation architecture; The front end displays monitoring images and 3D images, and the back end includes a processing module and an acquisition module. The acquisition module includes a monitoring module and a sensor module. The processing module includes a video data processing module, a sensor data processing module and a data fusion module. The monitoring module is connected to the video data processing module, and the sensor module is connected to the sensor data processing module. After the video data processing module and the sensor data processing module are completed, they are sent to the data fusion module, and the data fusion module is connected to the 3D generation module.
10. The system for monitoring coal mines by 3D and video linkage according to claim 9, characterized in that: The sensor module includes a gas sensor, a temperature and humidity sensor, and a pressure sensor. The data collected by the sensor module is transmitted to the sensor data processing module through the Internet of Things; the monitoring module performs real-time video monitoring and transmits the video data to the video data processing module through a video transmission protocol; After the backend integrates and processes the data, it is pushed to the front end for display.
Citation Information
Patent Citations
Coal mine underground multi-source sensor data dynamic weight fusion SLAM method
CN117333406A
Cited By
Fully mechanized coal mining face mine pressure real-time monitoring system and method based on multi-sensor fusion
CN120701413A