Coal mine gas occurrence three-dimensional geological modeling method and system based on AI vision

CN122049266BActive Publication Date: 2026-09-29GUIZHOU INST OF COAL SCI +2
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202610157491.3
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2026-02-04
Publication Date
2026-09-29
Estimated Expiration
2046-02-04

AI Technical Summary

Technical Problem

[0008]本发明的核心目的是解决现有煤矿瓦斯赋存三维地质建模中图像质量差、特征识别精度低、模型训练优化不足、建模与实际需求脱节等问题,提出基于AI视觉的煤矿瓦斯赋存三维地质建模方法及系统,具体是融合低光照图像增强、深度学习特征识别与多源数据融合的煤矿瓦斯赋存三维地质建模解决方案,适用于煤矿井下采掘过程中瓦斯赋存状态的高精度建模、动态更新及安全管控场景

Benefits of technology

[0047]1. 图像增强效果显著,适配井下复杂环境:本发明设计“线性化网络+频域分解+光效抑制+动态范围提升”四级架构的低光照图像增强模块,采用半监督学习策略,结合CRF线性化、频域分域处理与多损失协同优化,实现图像线性化转换、光效精准抑制与动态范围提升的协同优化,为后续特征识别提供高质量图像数据,既解决了井下低光照区域提亮问题,又有效抑制了光晕、反光等光效干扰,增强后图像的细节保留度与色彩还原度显著提升,为特征识别提供了高质量数据支撑,相较于传统增强方法,光效抑制率提升40%以上;

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122049266B_ABST
    Figure CN122049266B_ABST
Patent Text Reader

Abstract

The present application relates to the cross field of coal mine geological modeling and artificial intelligence vision application, and provides a coal mine gas occurrence three-dimensional geological modeling method and system based on AI vision, which comprises: acquiring multi-source data of a mining area in a coal mine, including geological basic data, engineering dynamic data, gas monitoring data and underground visual image data; performing low-light enhancement processing on the visual image data to obtain an enhanced image; based on the enhanced image, using a deep learning model to identify coal rock features, outputting semantic segmentation results containing coal rock interfaces, geological structures and fissures, and corresponding geological geometric parameters; taking the semantic segmentation results as spatial geometric constraints, combining the geological data, engineering data and gas monitoring data, and performing data fusion through a spatial interpolation algorithm to construct a three-dimensional geological model reflecting the gas occurrence state. The present application can realize high-precision modeling, dynamic updating and safety management and control of the gas occurrence state in the mining process in a coal mine.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the interdisciplinary field of coal mine geological modeling and artificial intelligence vision applications, specifically to a three-dimensional geological modeling method and system for coal mine gas occurrence based on AI vision. Background Technology

[0002] Coal mine gas disasters are one of the core risks restricting safe production in coal mines. Accurately understanding the three-dimensional geological characteristics of gas occurrence (including gas pressure, content, permeability distribution, and the spatial morphology of geological structures such as coal-rock interfaces, faults, and fissures) is crucial for achieving early warning and efficient management of gas disasters. Current three-dimensional geological modeling of coal mine gas mainly relies on geological exploration data, engineering measurement data, and sensor monitoring data, constructing models through interpolation algorithms. However, it suffers from the following prominent problems:

[0003] 1. Low utilization rate of underground image data: Coal and rock images inside underground boreholes and exposed surfaces of roadways are direct data sources reflecting geological features. However, the underground environment has problems such as low light, strong reflection, and halo interference, resulting in insufficient clarity of the original images. Existing image enhancement methods only focus on brightening low-light areas, which can easily aggravate light effect interference and cannot meet the needs of feature recognition.

[0004] 2. Lack of targeted optimization in model training: Existing image enhancement and feature recognition models related to modeling mostly adopt a single supervised training mode, which is difficult to adapt to the complex data distribution of real-world downhole scenes. Furthermore, they do not design loss functions for core tasks such as CRF estimation, light effect suppression, and dynamic range improvement, resulting in weak model generalization ability and insufficient recognition accuracy.

[0005] 3. The contradiction between accuracy and efficiency in coal and rock feature recognition: Traditional feature recognition relies on manual annotation or simple CNN models. Manual annotation faces bottlenecks such as scarce geological expert resources, high cost and long cycle. Moreover, a single model is difficult to simultaneously achieve semantic segmentation of coal and rock interfaces and microstructures and accurate extraction of geometric parameters such as fault strike and fracture size, and cannot provide quantitative constraints for modeling.

[0006] 4. Limited accuracy of multi-source data fusion modeling: Existing modeling methods do not make full use of the high-precision geological features of visual recognition as hard constraints, and the data density requirements of the interpolation algorithm do not match the actual distribution of downhole data, resulting in the model being unable to truly reflect the relationship between gas occurrence and geological structure, and the dynamic update response is lagging.

[0007] These problems result in low accuracy and poor timeliness of existing 3D geological models, making it difficult to support the actual needs of coal mine gas safety management. Therefore, there is an urgent need to propose an innovative solution that integrates advanced AI vision technology with multi-source data modeling. Summary of the Invention

[0008] The core objective of this invention is to address the problems of poor image quality, low feature recognition accuracy, insufficient model training and optimization, and disconnect between modeling and actual needs in existing 3D geological modeling of coal mine gas occurrence. This invention proposes an AI-based vision-based 3D geological modeling method and system for coal mine gas occurrence. Specifically, it is a 3D geological modeling solution for coal mine gas occurrence that integrates low-light image enhancement, deep learning feature recognition, and multi-source data fusion. This solution is suitable for high-precision modeling, dynamic updating, and safety management of gas occurrence status during underground coal mining.

[0009] To achieve the above objectives, the following technical solution is adopted:

[0010] In a first aspect, embodiments of the present invention provide a three-dimensional geological modeling method for coal mine gas occurrence based on AI vision, comprising the following steps:

[0011] Acquire multi-source data of the underground mining area of ​​the coal mine, wherein the multi-source data includes at least geological basic data, engineering dynamic data, gas monitoring data, and visual image data acquired by underground image acquisition equipment;

[0012] The visual image data is subjected to low-light enhancement processing based on the image enhancement model to obtain an enhanced image; wherein, the low-light enhancement processing includes: estimating the camera response function and linearizing the original image, decomposing the linearized image in the frequency domain to separate light effect interference and detail texture, and performing light effect suppression, noise removal and dynamic range enhancement on the decomposed features respectively, and finally fusing to generate the enhanced image;

[0013] Based on the enhanced image, a deep learning model is used to identify coal and rock features, and the output includes semantic segmentation results containing coal and rock interfaces, geological structures and fractures, as well as the corresponding geological geometric parameters.

[0014] Using the semantic segmentation results as spatial geometric constraints, and combining them with the geological data, engineering data, and gas monitoring data, a three-dimensional geological model reflecting the gas occurrence state is constructed through spatial interpolation algorithms.

[0015] Furthermore, the method also includes: preprocessing the multi-source data, specifically including:

[0016] The geological basic data, engineering dynamic data, and gas monitoring data are processed to unify and standardize units;

[0017] Wavelet thresholding is used to denoise the geophysical data in the geological basic data, and sliding window mean filtering is used to denoise the sensor data in the gas monitoring data.

[0018] The three-dimensional spatial coordinates of each data acquisition point are determined based on GPS and UWB, and the spatiotemporal correspondence between multiple data sources is established with the advance time of the mining face as the time reference. The three-dimensional spatial coordinates of each data acquisition point are uniformly transformed to the preset reference coordinate system through a seven-parameter coordinate transformation model based on multiple control points.

[0019] Furthermore, the image enhancement model is a cascaded processing model built on a convolutional neural network, and its processing flow includes:

[0020] The original image is subjected to camera response function estimation through a first convolutional network, and the output is used to calculate the parameters for the inverse camera response function, so as to transform the image into a linear space;

[0021] The linearized image is decomposed into a low-frequency feature map and a high-frequency feature map in the frequency domain. The low-frequency feature map contains illumination and reflection information, and the high-frequency feature map contains texture and noise information.

[0022] The low-frequency feature map is processed by a second convolutional network to suppress optical interference; the high-frequency feature map is processed by a third convolutional network to remove noise and preserve detailed texture.

[0023] The dynamic range of the low-frequency feature map after light effect suppression is enhanced by using a fourth convolutional network;

[0024] The enhanced image is generated by weighted fusion of the low-frequency feature map after dynamic range enhancement and the high-frequency feature map after denoising.

[0025] Furthermore, the training and optimization of the image enhancement model employs a two-stage semi-supervised learning strategy that includes supervised training and unsupervised fine-tuning, wherein:

[0026] During the supervised training phase, synthetic image data with known camera response function labels are used. The total loss function for supervised training integrates the mean squared error loss for camera response function estimation, the L1 norm loss for image linearization, and the logarithmic transformation loss for dynamic range enhancement.

[0027] In the unsupervised fine-tuning stage, unlabeled real-world downhole image data is used to improve the model's adaptability and generalization ability to real-world scenes by optimizing the total unsupervised fine-tuning loss function. The total unsupervised fine-tuning loss function includes: monotonicity loss for constraining the monotonicity of the predicted camera response function, distribution linearization loss for optimizing the distribution characteristics of image edge pixels, reconstruction loss for ensuring the consistency between the enhanced image and the original image content, smoothing loss for constraining the smoothness of the light effect suppression mask space, and grayscale world loss for correcting the overall color shift of the image.

[0028] Furthermore, the deep learning model is a neural network with a dual-output branch structure, including:

[0029] The semantic segmentation branch, built on an encoder-decoder structure, is used to perform pixel-level classification on the input image and output a semantic segmentation mask containing coal-rock interfaces, geological structures, and fractures.

[0030] The geometric parameter regression branch, connected to the backbone network of the semantic segmentation branch, is used to regress and output multiple geological geometric parameters such as fault strike, dip angle, and fracture size based on the features extracted by the backbone network.

[0031] Furthermore, when training the deep learning model, a pseudo-labeled data augmentation strategy based on semi-supervised learning is adopted, specifically including:

[0032] Obtain a small set of images annotated by experts as a seed annotation set;

[0033] Using an initial model trained on the seed annotation set, predictions are made on a large number of unlabeled downhole images;

[0034] The portion of the prediction results with a confidence level higher than a preset threshold is selected as pseudo-labeled data.

[0035] The seed-labeled data is combined with the verified pseudo-labeled data to form a dataset for model training.

[0036] Furthermore, during the training of the deep learning model, data augmentation processing is applied to the input image; the data augmentation processing includes general geometric transformation enhancement and specific enhancement simulating underground coal mine imaging conditions, the specific enhancement including at least random adjustment of image illumination intensity and noise addition simulating underground dust interference.

[0037] Furthermore, when training the deep learning model, a multi-task joint loss function is used to collaboratively optimize the semantic segmentation branch and the geometric parameter regression branch;

[0038] The joint loss function is composed of a weighted sum of semantic segmentation loss and geometric parameter regression loss; wherein, the semantic segmentation loss is a weighted sum of Dice loss and cross-entropy loss, used to handle the pixel imbalance problem of target categories in the image; the geometric parameter regression loss adopts a smooth L1 loss function, which is robust to outliers in the prediction error.

[0039] Furthermore, the construction of the three-dimensional geological model reflecting the gas occurrence state includes: the spatial interpolation algorithm is the co-kriging interpolation method, which uses gas monitoring parameters as the main variable and geological parameters as auxiliary variables for interpolation calculation;

[0040] An adaptive resolution meshing strategy is adopted, using a first resolution for the core area around the mining face and a second resolution lower than the first resolution for non-core areas far from the working face, in order to balance model accuracy and computational cost.

[0041] Secondly, embodiments of the present invention also provide a three-dimensional geological modeling system for coal mine gas occurrence based on AI vision, used to implement the method described in the first aspect, the system comprising:

[0042] The data acquisition module is used to acquire multi-source data of the underground mining area of ​​the coal mine. The multi-source data includes at least geological basic data, engineering dynamic data, gas monitoring data, and visual image data acquired by underground image acquisition equipment.

[0043] An image enhancement processing module is used to perform low-light enhancement processing on the visual image data based on an image enhancement model to obtain an enhanced image; wherein, the low-light enhancement processing includes: estimating the camera response function and linearizing the original image, decomposing the linearized image in the frequency domain to separate light effect interference and detail texture, and performing light effect suppression, noise removal and dynamic range enhancement on the decomposed features respectively, and finally fusing to generate the enhanced image;

[0044] The coal and rock feature recognition module is used to perform coal and rock feature recognition based on the enhanced image using a deep learning model, and outputs semantic segmentation results including coal and rock interfaces, geological structures and fractures, as well as corresponding geological geometric parameters.

[0045] The data fusion and 3D modeling module is used to take the semantic segmentation results as spatial geometric constraints, combine them with the geological basic data, engineering dynamic data and gas monitoring data, and perform data fusion through spatial interpolation algorithms to construct a 3D geological model that reflects the gas occurrence state.

[0046] Compared with the prior art, the present invention achieves the following beneficial effects:

[0047] 1. Significant image enhancement effect, adaptable to complex downhole environments: This invention designs a low-light image enhancement module with a four-level architecture of "linearization network + frequency domain decomposition + light effect suppression + dynamic range enhancement". It adopts a semi-supervised learning strategy, combining CRF linearization, frequency domain decomposition processing and multi-loss collaborative optimization to achieve synergistic optimization of image linearization transformation, accurate light effect suppression and dynamic range enhancement. This provides high-quality image data for subsequent feature recognition, which not only solves the problem of brightening low-light areas downhole, but also effectively suppresses light effect interference such as halos and reflections. The detail retention and color reproduction of the enhanced image are significantly improved, providing high-quality data support for feature recognition. Compared with traditional enhancement methods, the light effect suppression rate is improved by more than 40%.

[0048] 2. The loss function system is well-developed and the model has strong generalization ability: Through the collaborative design of the total loss of supervised training and the total loss of unsupervised fine-tuning, the full-scene optimization of "labeled training + unlabeled adaptation" is achieved. This not only ensures the training accuracy of the model on standard datasets, but also improves the adaptability to real and complex images in mines. The RMSE of CRF estimation is less than 0.06, and the generalization error is reduced by more than 25% compared with existing fully supervised methods.

[0049] 3. Precise and efficient coal and rock feature identification, breaking through the annotation bottleneck: This invention adopts a feature identification model with a dual-output architecture of "U-Net++ semantic segmentation backbone network + geometric parameter regression branch", which simultaneously achieves semantic segmentation and geometric parameter regression. The segmentation accuracy (mIoU) reaches over 85%, and the geometric parameter prediction error is less than 5%. Furthermore, the annotation bottleneck is alleviated through the strategy of "a small number of expert annotations + semi-supervised pseudo-annotation". At the same time, it achieves semantic segmentation of features such as coal and rock interfaces, faults, folds, and fractures, as well as accurate regression of 12-dimensional geological geometric parameters, providing quantitative constraints for modeling and reducing annotation costs by more than 70%. This solves the core pain points of traditional methods, such as long annotation cycles and high costs, and provides precise quantitative geological feature data for modeling.

[0050] 4. High modeling accuracy and timeliness, adaptable to practical application needs: AI visual recognition results are integrated into multi-source data fusion modeling as hard constraints. Combined with geological, engineering, and gas parameter data, and co-kriging interpolation and adaptive resolution grid, the modeling accuracy and computational resource consumption are balanced. The core area resolution of the model reaches 0.5m×0.5m×0.5m, and the modeling error is less than 5%, providing accurate and efficient technical support for coal mine gas safety management.

[0051] It should be understood that the description in the Summary of the Invention is not intended to limit the key or essential features of the embodiments of the present invention, nor is it intended to restrict the scope of the invention. Other features of the invention will become readily apparent from the following description. Attached Figure Description

[0052] The above and other features, advantages, and aspects of the various embodiments of the present invention will become more apparent from the accompanying drawings and the following detailed description. The drawings are provided for a better understanding of the invention and are not intended to limit the invention. In the drawings, the same or similar reference numerals denote the same or similar elements, wherein:

[0053] Figure 1 This is a flowchart illustrating the three-dimensional geological modeling method for coal mine gas occurrence based on AI vision, according to an embodiment of the present invention.

[0054] Figure 2 This is a flowchart of the downhole low-light image enhancement model processing in an embodiment of the present invention;

[0055] Figure 3 This is a schematic diagram of the deep learning model structure used for coal and rock feature recognition in an embodiment of the present invention;

[0056] Figure 4 This is a three-dimensional geological model simulation image generated in step S4 of this embodiment of the invention;

[0057] Figure 5 This is a schematic diagram of a three-dimensional geological modeling system module for coal mine gas occurrence based on AI vision, according to an embodiment of the present invention.

[0058] Figure 6 This is a comparison chart showing the root mean square error trends of the algorithm of this invention and three existing algorithms under different training iterations. Detailed Implementation

[0059] To make the objectives, technical solutions, and advantages of the embodiments of the present invention clearer, the technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.

[0060] Furthermore, the term "and / or" in this article is merely a description of the relationship between related objects, indicating that three relationships can exist. For example, A and / or B can represent: A existing alone, A and B existing simultaneously, or B existing alone. Additionally, the character " / " in this article generally indicates that the preceding and following related objects have an "or" relationship.

[0061] Figure 1 This is a flowchart illustrating a three-dimensional geological modeling method for coal mine gas occurrence based on AI vision, according to an embodiment of the present invention. Figure 1 As shown in the figure, an AI vision-based three-dimensional geological modeling method 100 for coal mine gas occurrence according to an embodiment of the present invention includes the following steps:

[0062] S1: Acquire multi-source data of the underground mining area of ​​the coal mine, wherein the multi-source data includes at least geological basic data, engineering dynamic data, gas monitoring data, and visual image data acquired by underground image acquisition equipment;

[0063] Step S1 is used to acquire multi-source heterogeneous data, including:

[0064] (1) Geological basic data: The geological exploration system collects borehole columnar sections, seismic wave geophysical data, geological profiles, etc., covering core parameters such as coal seam thickness, burial depth, and lithological combination;

[0065] (2) Engineering dynamic data: real-time extraction of time-series data such as roadway excavation trajectory, coal face advancement progress, and support structure layout from the automated mining platform;

[0066] (3) Gas monitoring data: Using distributed fiber optic sensors and gas concentration sensor arrays, key parameters such as gas pressure (unit: MPa), gas content (unit: m³ / t), and coal seam permeability coefficient (unit: m² / (MPa²·d)) are collected, and the sampling frequency is set to 1 time / 5 minutes.

[0067] (4) Visual image data: Deploy high-definition industrial cameras (resolution ≥ 1920×1080) and infrared supplementary light modules on the front end of the drill rod and the detection equipment on the side wall of the roadway to collect coal and rock images inside the borehole and exposed surface of the roadway in real time. The image format is RGB / JPEG and the acquisition frame rate is set to 10fps.

[0068] Step S1 also includes a preprocessing step for the acquired multi-source heterogeneous data, specifically including:

[0069] Furthermore, the method also includes: preprocessing the multi-source data, specifically including:

[0070] (1) Standardization processing: Units are unified for geological basic data and engineering dynamic data (e.g., length unit is unified to meter, pressure unit is unified to MPa), and the gas monitoring data is normalized using the Z-Score standardization formula.

[0071] (2) Noise Reduction Processing: Wavelet thresholding algorithm is used to remove random noise from geophysical data in geological foundation data, and sliding window (window size) is used for sensor data in gas monitoring data. Mean filtering eliminates fluctuation noise.

[0072] (3) Spatiotemporal registration: Based on GPS and the underground positioning system (UWB), the three-dimensional spatial coordinates (X, Y, Z) of each data acquisition point are determined. The time of advance of the mining face is used as the time reference to establish the spatiotemporal correspondence between the multi-source data. The coal mine independent coordinate system is used as the modeling reference coordinate system, and its transformation relationship with the geodetic coordinate system (CGCS2000) is clarified: The three-dimensional spatial coordinates of each data acquisition point are transformed into the preset reference coordinate system by selecting three or more control points with known CGCS2000 coordinates on the ground (such as the mining area measurement reference point) and using the seven-parameter Bursa-Wolf transformation model. The transformation formula is as follows:

[0073]

[0074] The x-axis coordinate (unit: m) of a data acquisition point in the coal mine independent coordinate system after conversion to the CGCS2000 geodetic coordinate system, which is used as a spatial positioning benchmark for interfacing with regional geological data and emergency rescue systems. The Y-axis coordinates (unit: m) of a data acquisition point in the coal mine's independent coordinate system after transformation to the CGCS2000 geodetic coordinate system, and... Together they form the planar positioning coordinates. The Z-axis coordinates (unit: m) of a data acquisition point in the coal mine independent coordinate system after conversion to the CGCS2000 geodetic coordinate system, corresponding to the geodetic elevation information, supporting the unification of three-dimensional spatial position. : x-axis translation parameter for coordinate transformation (unit: m), used to correct the origin offset in the x-direction between the coal mine independent coordinate system and the CGCS2000 coordinate system. : Y-axis translation parameter for coordinate transformation (unit: m), used to correct the origin offset in the Y direction between the coal mine independent coordinate system and the CGCS2000 coordinate system. Z-axis translation parameter (unit: m) for coordinate transformation, used to correct the origin offset in the elevation direction between the coal mine independent coordinate system and the CGCS2000 coordinate system. The scale parameter (dimensionless) for coordinate transformation is used to correct scale differences between two coordinate systems (such as measurement unit error, projection distortion, etc.). X-axis coordinates of a data collection point in the coal mine's independent coordinate system (unit: ), which is one of the original positioning coordinates for downhole data acquisition. The Y-axis coordinate (unit: m) of a certain data acquisition point in the coal mine independent coordinate system is one of the original positioning coordinates for underground data acquisition. The Z-axis coordinate (unit: m) of a data acquisition point in the coal mine's independent coordinate system represents the original elevation coordinates of the underground data acquisition, reflecting the depth information of the mining face, boreholes, etc. : The rotation parameter (unit: rad) of the coordinate transformation about the X-axis, used to correct the spatial attitude deviation between the two coordinate systems in the X-axis direction. : Rotation parameter (unit: rad) around the Y-axis for coordinate transformation, used to correct spatial attitude deviation between two coordinate systems in the Y-axis direction. : Rotation parameter (unit: rad) around the Z-axis for coordinate transformation, used to correct spatial attitude deviation between two coordinate systems in the Z-axis direction.

[0075] The transformation accuracy is ensured by fitting the control point coordinates. This ensures the compatibility of the model with regional geological data and emergency rescue systems.

[0076] S2: Based on the image enhancement model, perform low-light enhancement processing on the visual image data to obtain an enhanced image; wherein, the low-light enhancement processing includes: estimating the camera response function and linearizing the original image, decomposing the linearized image in the frequency domain to separate light effect interference and detail texture, and performing light effect suppression, noise removal and dynamic range enhancement on the decomposed features respectively, and finally fusing to generate the enhanced image;

[0077] Step S2 is used to construct a downhole low-light image enhancement module to achieve image enhancement processing.

[0078] To address issues such as low lighting, strong reflections (equipment light reflections), and halo interference in underground mining operations, a four-level architecture of "linearized network + frequency domain decomposition + light effect suppression + dynamic range enhancement" was designed. A semi-supervised learning image enhancement model was adopted to achieve synergistic optimization of image linearization transformation, precise light effect suppression, and dynamic range enhancement, providing high-quality image data for subsequent feature recognition.

[0079] S2.1: Constructing an image enhancement model

[0080] like Figure 2 As shown, the image enhancement model is a cascaded processing model based on a convolutional neural network. Its processing flow includes: estimating the camera response function of the original image using a first convolutional network, outputting parameters for calculating the inverse camera response function to transform the image into a linear space; decomposing the linearized image into low-frequency and high-frequency feature maps in the frequency domain, where the low-frequency feature map contains illumination and reflection information, and the high-frequency feature map contains texture and noise information; processing the low-frequency feature map using a second convolutional network to suppress light effect interference; processing the high-frequency feature map using a third convolutional network to remove noise and preserve detail texture; performing dynamic range enhancement on the low-frequency feature map after light effect suppression using a fourth convolutional network; and weighted fusing the dynamically enhanced low-frequency feature map with the denoised high-frequency feature map to generate an enhanced image. The specific implementation is as follows:

[0081] The model structure adopts a four-level architecture of "linearized network + frequency domain decomposition + optical efficiency suppression + dynamic range enhancement":

[0082] 1. Linearized Network: Using ResNet-18 as the backbone network, the input is the original low-light image, and the output is the coefficients of the 11 basis functions of the Camera Response Function (CRF). This is achieved through a basis function model. Calculate the inverse CRF to achieve image linearization transformation, where Based on inverse CRF, As basis functions, These are the network output coefficients.

[0083] The inverse camera response function predicted by the model is used to convert the raw image (nonlinear response) of the downhole low-light image into a linearized image, eliminating the nonlinear distortion of camera imaging. The basic inverse CRF (preset baseline function) provides an initial reference for inverse CRF calculation. It is determined based on the statistical characteristics of 201 CRFs in the DoRF database and is adapted to the imaging characteristics of downhole industrial cameras. Basis function index (value range 1 to 11), corresponding to 11 independent basis function dimensions, used to accurately fit the nonlinear response patterns of different cameras. : No. The basis functions (preset fixed functions) are selected from the CRF modeling-specific basis function set to meet the nonlinear fitting requirements in low-light downhole scenarios. The output of the linearized network is the first... The basis function coefficients are learnable parameters of the model, which are adapted to the specific imaging nonlinear characteristics of downhole images through training.

[0084] 2. Frequency Domain Decomposition: A differentiable decomposition model is used to decompose the linearized image into low-frequency (LF) feature maps (including halos, reflections, and other lighting effects) and high-frequency (HF) feature maps (including noise, texture, and edges). The number of decomposition filters is [not specified]. Ablation experiments revealed that setting the number of frequency domain decomposition filters K to 8 achieves the best balance between model complexity and the accuracy of separating lighting effects / textures. Further increasing the value of K does not significantly improve the enhancement effect, but instead increases the computational burden.

[0085] 3. Feature Processing Network: The LF feature map is input into the LF-DeLight network (encoder-decoder architecture, containing 4 encoder layers and 4 decoder layers, using the ReLU activation function with jumper connections), and outputs a light effect mask through a light effect suppression loss function to achieve halo / reflection removal; the HF feature map is input into the HF-Denoise network (same architecture as LF-DeLight), and texture and edge preservation are achieved through a noise removal loss function. 4. Dynamic Range Enhancement: The light-effect-reduced LF feature map is input into the LF-HDR network (encoder-decoder architecture, 3 output channels), and the dynamic range is enhanced using an HDR loss function. Finally, it is fused with the denoised HF feature map to generate an enhanced image.

[0086]

[0087] The final generated downhole low-light enhanced image (RGB format, resolution consistent with the input) combines dynamic range enhancement, light effect suppression (halo / reflection) and noise removal features, and is used for subsequent coal and rock feature recognition. The number of filters used for frequency domain decomposition (fixed value is 8). By using 8 sets of filters with different frequency responses, the linearized image can be finely decomposed in the frequency domain. Filter index (range of values) ), corresponding to the The processing results of the group frequency domain decomposition filter. : No. The enhanced feature map after processing the low-frequency (LF) feature map by the LF-HDR network has completed dynamic range enhancement and light effect suppression, while preserving large-scale structural information such as coal seams and rock strata. : No. The denoised feature map after processing by the HF-Denoise network has removed downhole dust and sensor noise, while retaining detailed texture information such as coal-rock interface and fractures. The enhanced LF feature map and the denoised HF feature map after K-group frequency domain decomposition are summed to integrate the full frequency domain information. The summed full-frequency feature map is averaged and normalized to avoid brightness overflow or detail distortion caused by feature superposition, thus ensuring enhanced visual consistency of the image.

[0088] As an example and not a limitation, in this embodiment of the invention, a specific implementation of the LF-DeLight network is as follows: The encoder part contains four convolutional layers with output channels of 64, 128, 256, and 512 respectively. Each convolutional layer is followed by a batch normalization layer and a ReLU activation function, and downsampling is performed using convolutions with a stride of 2. The decoder part contains four deconvolutional layers with output channels of 256, 128, 64, and 3 respectively. Each deconvolutional layer is also followed by a batch normalization layer and a ReLU activation function. Feature fusion is performed between corresponding layers of the encoder and decoder through skip connections. Specifically, the feature map output by the encoder is concatenated with the upsampled feature map of the decoder in the channel dimension. The HF-Denoise network can adopt the same or similar architecture as the LF-DeLight network. The LF-HDR network has 3 output channels and its structure is similar to that of the LF-DeLight network, but the final activation function can use the Sigmoid function to map the output to the [0,1] interval.

[0089] S2.2: Constructing the training strategy

[0090] S2.2.1: Training Data

[0091] Synthetic data (based on 201 CRFs from the DoRF database, combined with downhole RAW images to synthesize RGB images with known CRF labels, Gaussian noise and JPEG compression added to simulate real-world scenes) Real data (collected 1000 low-light downhole images, without CRF labels).

[0092] In a preferred embodiment, the method for generating synthetic image data includes: first, acquiring a batch of raw RAW images of an underground scene; then, randomly selecting a known camera response function (CRF) from the DoRF database; next, using the known CRF to convert the selected RAW image into a linear radiosity image; then, applying Gaussian noise simulating the underground environment and a JPEG compression algorithm to the linear image; finally, applying a random tone mapping function conforming to the underground illumination distribution to convert the processed linear image into a low-light RGB image for training. In this process, the known CRF used is the real CRF label corresponding to the synthetic RGB image.

[0093] Furthermore, the training and optimization of the image enhancement model employs a two-stage semi-supervised learning strategy that includes supervised training and unsupervised fine-tuning. Specifically, this includes the following:

[0094] S2.2.2: Supervised Training Phase

[0095] During the supervised training phase, synthetic image data with known camera response function labels are used. The overall supervised training loss function integrates the mean squared error loss for camera response function estimation, the L1 norm loss for image linearization, and the logarithmic transform loss for dynamic range enhancement. Specifically:

[0096] Using MSE loss Training CRF estimation, L1 loss Training image linearization, HDR loss ( , Training dynamic range improvement:

[0097]

[0098] Mean Squared Error Loss (MSE) is used to train the CRF estimation accuracy of a linearized network, measuring the difference between the predicted inverse CRF and the true inverse CRF. The inverse camera response function (InverseCRF) predicted by the model is calculated from the coefficients of 11 basis functions output by the linearization network, which is adapted to the linearization requirements of low-light images in wells. Ground Truth: The ground truth labels of the inverse CRF are generated based on 201 known CRFs in the DoRF database and are preset ground truth values ​​for the synthetic training data. Norm (Euclidean norm), used to calculate and predict the inverse. Contrary to reality The sum of squared errors between the two sides is minimized by gradient descent, which improves the accuracy of CRF estimation. L1 loss (L1Loss) is used to train the linearization effect of images, constrain the deviation between the linearized image and the real linearized image, and reduce the impact of outliers on training. Predicting the inverse CRF applied to synthesized training images The resulting linearized image is the model-linearized transformation of the simulated image of low-light conditions in the well. The training images are synthesized based on downhole RAW images and known CRF tags, and Gaussian noise and JPEG compression are added to simulate the real downhole imaging environment. The real inverse CRF applied to synthetic training images The resulting true linearized image serves as a reference benchmark for the linearization effect. L1 norm (Manhattan norm) is used to calculate the sum of absolute errors between the predicted linearized image and the true linearized image, thus optimizing the stability of the linearization process. High Dynamic Range Loss (HDR Loss) is used to train the dynamic range enhancement capability of the LF-HDR network, adapting to the needs of brightening low-light areas and suppressing excessively bright light effects in wells. : Constant parameter (fixed value of 10), used to adjust the dynamic range compression ratio of the logarithmic transform to adapt to the brightness distribution characteristics of the downhole image. The HDR image predicted by the model is obtained by fusing the enhanced LF feature map and the denoised HF feature map, and contains information about the coal and rock scene after the dynamic range is improved. Ground Truth: The ground truth of the HDR image is the corresponding high-precision HDR image in the training dataset, reflecting the true dynamic range of the coal and rock scene. Logarithmic transformation is performed on the predicted HDR image and the real HDR image to compress the dynamic range of brightness and highlight the differences in details in low-light areas. Normalize the result after logarithmic transformation so that the output value is in the range of [0, 1], which facilitates loss calculation and gradient optimization.

[0099] S2.2.3: Fine-tuning during unsupervised testing

[0100] In the unsupervised fine-tuning stage, unlabeled real-world downhole image data is used to improve the model's adaptability and generalization ability to real-world scenes by optimizing the total unsupervised fine-tuning loss function. The total unsupervised fine-tuning loss function includes: a monotonicity loss to constrain the monotonicity of the predicted camera response function; a distribution linearization loss to optimize the pixel distribution characteristics at image edges; a reconstruction loss to ensure consistency between the enhanced image and the original image content; a smoothing loss to constrain the smoothness of the mask space for suppressing light effects; and a gray-world loss to correct the overall color cast of the image. The specific implementation is as follows:

[0101] Using monotonic loss Constraining the monotonicity of CRF, linearizing the distribution loss ( , Optimize edge pixel distribution and introduce reconstruction loss. Smoothing loss Gray world loss To achieve light effect suppression, the expressions for each loss function are as follows:

[0102] 1. Monotonicity loss (constraining the monotonically increasing property of the CRF function to avoid image color distortion):

[0103]

[0104] Monotonicity Loss constrains the monotonically increasing characteristics of the predicted inverse CRF, avoiding color shift and brightness reversal problems in downhole images caused by nonlinear distortion of CRF. Brightness value variable (values ​​are...) (1024 equally spaced sampling points within the interval), covering the entire brightness range of the downhole image. Predicting inverse CRF in brightness value The output value at that point reflects the camera's linear response to that brightness. Predicting inverse CRF in brightness value The derivative at a point represents the rate of change of the CRF and must remain non-negative to satisfy the monotonicity requirement. The negative value of the derivative is used to trigger the constraint logic of the step function. The Heaviside Step Function (HWF) is a function that, when input... Output 1 if the condition is met, otherwise output 0; this function penalizes cases where the derivative is negative, forcing the CRF to be monotonically increasing. :right The total monotonicity loss is obtained by summing the step function outputs of all sampling points within the interval. The larger the loss value, the more severe the monotonicity violation of the CRF.

[0105] 2. Distributed linearization loss (optimizes the distribution of pixels at image edges, improving detail preservation):

[0106]

[0107] The linearization loss of the single edge region is optimized to improve the linear distribution characteristics of a single group of edge pixels in the RGB space, thereby improving the recognition accuracy of edge details such as coal-rock interfaces and cracks. The number of edge pixels within a single edge region, reflecting the edge detail density of that region. : No. The brightness value of each edge pixel after linearization and normalization (the combined brightness after merging RGB channels, in the range of...) (Interval). The minimum brightness value of an edge pixel within a single edge region, serving as a lower bound reference for the linear distribution. The maximum brightness value of edge pixels within a single edge region, serving as an upper bound reference for the linear distribution. The range of brightness values ​​within the edge region characterizes the dynamic range of brightness in that region; when used as the denominator, it avoids being zero (ensuring this through data preprocessing), and simplifies to the form after the numerator and denominator cancel each other out. That is, to calculate the sum of the absolute errors between each edge pixel and the minimum brightness value, and to constrain the edge pixels to be linearly distributed. The global edge distribution linearization loss integrates the distribution optimization results of all edge regions, improving the linearity of edge details in the entire downhole image. The number of edge regions extracted from the downhole image is determined by the image edge detection algorithm, covering key feature areas such as coal-rock interfaces, faults, and fissures. Edge region index (value range) ), corresponding to the A marginal area. Edge pixel index (range of values) ), corresponding to the The first edge region One edge pixel. The linearization optimization of the global edge pixel distribution is achieved by summing the single-region losses of all edge regions.

[0108] 3. Reconstruction loss (ensuring consistency between the enhanced image and the original image, avoiding over-enhancement):

[0109]

[0110] Reconstruction Loss ensures the consistency of content between the enhanced image and the original downhole image, avoiding distortion of coal and rock features (such as disappearance of fractures and misjudgment of lithology) caused by over-enhancement. The enhanced downhole image output by the model has undergone dynamic range enhancement, light effect suppression, and noise removal. The original low-light image of the well is used as the input image for the model, preserving the original content information of the well scene. L1 norm: Calculates the pixel-level absolute error between the enhanced image and the original image, constraining the enhancement process to not change the core content of the scene.

[0111] 4. Smoothing Loss (Constraining the spatial smoothness of the light effect mask to avoid jagged edges and improve the naturalness of light effect suppression):

[0112]

[0113] The light effect mask smoothing loss constrains the spatial smoothness of the light effect mask, avoids jagged edges in the mask, and ensures that the light effect suppression process such as halo and reflection is natural and does not damage the texture of the coal and rock background. Light effect mask The width (consistent with the width of the input downhole image). Light effect mask The height (consistent with the height of the input downhole image). The total number of pixels in the light effect mask is used to normalize the loss value and avoid the impact of image size differences on training stability. : Light effect mask in coordinates Pixel value at ( (Interval), 1 indicates that the position is a light effect area (to be suppressed), and 0 indicates a non-light effect area (to be retained). : Light effect mask in coordinates The horizontal gradient at a given location represents the rate of change of the mask pixel values ​​in the horizontal direction. : Light effect mask in coordinates The vertical gradient at a given location represents the rate of change of the mask pixel values ​​in the vertical direction. : The absolute values ​​of the horizontal and vertical gradients, to avoid the cancellation of positive and negative gradients. The total gradient loss is obtained by summing the absolute values ​​of the gradients of all pixels in the light effect mask. After normalization, the edges of the optical mask are smoothed by minimizing this loss.

[0114] 5. Gray-world loss (color balance of the image after constraint enhancement, correcting color cast caused by underground lighting):

[0115]

[0116] Gray World Loss: Constrains the RGB three-channel color balance of the enhanced image, corrects the color shift caused by underground lighting (such as equipment spotlights and emergency lights), and ensures accurate color reproduction of coal and rock. Enhanced image in coordinates The red channel pixel value at that location (normalized to the [0,1] range). Enhanced image in coordinates The green channel pixel value at that location (normalized to the [0,1] range). Enhanced image in coordinates The blue channel pixel value at that location (normalized to the [0,1] range). The global mean of the red channel represents the overall brightness level of the red channel. : Global average of the green channel. : Global mean of the blue channel. The target value (after normalization) of the three-channel mean under the grayscale world assumption is that the brightness ratio of each channel is consistent when the color is ideally balanced. Absolute value calculation: Calculates the mean and target value for each channel. The deviations are calculated; the sum of the three deviations yields the total grayscale world loss, and color balance is achieved by minimizing this loss.

[0117] S2.2.4: Complete Loss Function Expression

[0118] 1. Total supervised training loss (combining the various supervised losses according to their weights):

[0119]

[0120] The total loss during supervised training integrates three sub-losses: CRF estimation, image linearization, and dynamic range enhancement. Multi-objective collaborative optimization is achieved through weight allocation, adapting to the training requirements of models for enhancing low-light images in wells. The weight coefficient of MSE loss (fixed value of 10.0) has the highest weight among the three types of sub-losses, prioritizing the accuracy of inverse CRF estimation and laying the foundation for image linearization and dynamic range improvement. Mean squared error loss is used to optimize the accuracy of CRF estimation and measures the difference between the predicted inverse CRF and the true inverse CRF (see the corresponding loss function description above for the specific meaning of the parameters). The weight coefficients of L1 loss (fixed value 1.0) balance the linearization effect of the image with other training objectives, ensuring that the linearization process is stable and does not introduce additional distortion. L1 loss is used to optimize the linearization effect of the image and constrain the deviation between the predicted linearized image and the true linearized image (see the corresponding loss function description above for the specific meaning of the parameters). The weighting coefficient for HDR loss (fixed value of 1.0) is consistent with the weighting of linearization loss, which simultaneously improves the dynamic range enhancement effect and adapts to the brightening needs of low-light areas in underground mines. High dynamic range loss is used to optimize the dynamic range enhancement capability of LF-HDR networks (see the corresponding loss function description above for the specific parameter meanings).

[0121] 2. Unsupervised fine-tuning total loss (integrating various types of unsupervised losses):

[0122]

[0123] During unsupervised testing, the total loss is fine-tuned by integrating five types of sub-losses: monotonicity constraint, edge distribution optimization, content reconstruction, light effect mask smoothing, and color balance. This optimizes the model's adaptability to real downhole images in the absence of real labels. The weighting coefficient for monotonicity loss (fixed value 1.0) ensures the monotonically increasing characteristic of the inverse CRF, avoids color distortion, and provides a basis for other optimization objectives. Monotonic loss constrains the monotonically increasing property of the inverse CRF (see the corresponding loss function explanation above for the specific meaning of the parameters). The weight coefficient of the linearization loss is fixed at 0.1. It has a low weight and moderately optimizes the distribution characteristics of edge pixels while ensuring the core objective. : Global edge distribution linearization loss, which optimizes the linearity of edge details in the entire image (see the corresponding loss function description above for the specific meaning of the parameters). The weighting coefficient of the reconstruction loss (fixed value 1.0) is used to ensure the consistency of content between the enhanced image and the original image, and to avoid over-enhancement that could lead to distortion of coal and rock features. : Reconstruction loss, constraining the consistency of content between the augmented image and the original image (see the corresponding loss function description above for the specific meaning of the parameters). : Weight coefficient for smoothing loss (fixed value 0.5), balances the smoothness of the light effect mask with other targets, and ensures that the light effect suppression is natural and does not destroy the background texture. : Light effect mask smoothing loss, which constrains the spatial smoothness of the light effect mask (see the corresponding loss function description above for the specific meaning of the parameters). The weighting coefficient of grayscale world loss (fixed value 1.0) is consistent with the weighting of monotonicity loss and reconstruction loss. It prioritizes correcting color deviation caused by underground lighting to ensure accurate color restoration of coal and rock. Gray-world loss, constrained enhancement of RGB three-channel color balance of the image.

[0124] In this embodiment of the invention, a complete loss function system for collaborative optimization is constructed through steps S2.2.2-S2.2.4. The total loss for supervised training and the total loss for unsupervised fine-tuning are designed. During the supervised training stage, the accuracy of CRF estimation, image linearization, and dynamic range enhancement is ensured by weighted fusion of MSE loss, L1 loss, and HDR loss. During the unsupervised fine-tuning stage, the model adaptability is optimized and the model's generalization ability is improved in scenarios without real labels by fusing monotonicity loss, distribution linearization loss, reconstruction loss, smoothing loss, and gray-world loss.

[0125] S3: Based on the enhanced image, use a deep learning model to identify coal and rock features, and output semantic segmentation results including coal and rock interfaces, geological structures and fractures, as well as corresponding geological geometric parameters;

[0126] S3: Constructing a deep learning model for coal and rock feature recognition

[0127] S3.1: Model Structure

[0128] Deep learning models are neural networks with a dual-output branch structure, specifically, such as Figure 3 As shown, the model structure adopts a dual-output architecture of "U-Net++ semantic segmentation backbone network + geometric parameter regression branch":

[0129] (1) Semantic segmentation branch: Based on the encoder-decoder structure, it is used to perform pixel-level classification of the input image and output a semantic segmentation mask containing coal-rock interface, geological structure and fracture.

[0130] The input is an enhanced coal and rock image (size 512×512×3). The encoder contains 5 convolutional layers (kernel size 3×3, stride 2, padding=1), each followed by BatchNorm and ReLU activation functions. The decoder contains 5 deconvolutional layers (kernel size 2×2, stride 2), employing an attention mechanism to enhance feature extraction. The output layer is a 5-channel semantic segmentation mask (corresponding to coal and rock interface, fault, fold, fissure, and background), with Softmax activation function.

[0131] (2) Geometric parameter regression branch: The backbone network connected to the semantic segmentation branch is used to regress multiple geological geometric parameters such as fault strike, dip angle, and fracture size based on the features extracted by the backbone network.

[0132] Three fully connected layers (1024, 512, and 12 neurons respectively) are added to the output of the U-Net++ decoder. Using the segmentation mask and the deep feature map of the backbone network as input, the optimal output is a 12-dimensional geological geometric parameter vector. Specifically, the 12-dimensional geological geometric parameters are key morphological parameters that significantly influence gas occurrence and can be effectively regressed from images, selected based on the experience of experts in coal mine geological modeling. These parameters include: fault strike, fault dip, fault length, and fault displacement; fold axis dip, fold wavelength, and fold amplitude; fracture strike, fracture dip, fracture length, fracture width (aperture), and fracture density. These parameters together constitute a set of key geological structural features that can be quantified and extracted from the image. The activation function used is Linear (no activation).

[0133] S3.2: Training optimization strategy (alleviating annotation bottleneck):

[0134] S3.2.1: Annotated Data Expansion

[0135] When training the deep learning model, a pseudo-labeling data augmentation strategy based on semi-supervised learning is adopted, specifically a "small amount of expert annotation + semi-supervised pseudo-labeling" strategy. This includes: acquiring a small number of high-quality images (e.g., 300 images) annotated by geological experts as a seed annotation set; using an initial model trained on the seed annotation set to predict the remaining large number of unannotated downhole images; selecting the portion of the prediction results with a confidence level higher than a preset threshold (e.g., prediction confidence ≥ 0.9) as pseudo-labeling data; merging the seed annotation data with the pseudo-labeling data that has been sampled and reviewed by experts (review rate 20%) to form a dataset for model training, for example, ultimately forming a dataset of 2000 valid annotations (300 seed annotations + 1700 qualified pseudo-labels). The confidence threshold is set to 0.9 based on the evaluation of pseudo-label quality. This threshold effectively filters out most unreliable prediction areas, ensuring the quality of the pseudo-labeling data. Sampling verification shows that its accuracy can reach over 95%.

[0136] S3.2.2: Data Augmentation and Enhancement

[0137] Data augmentation processing is applied to the input image: In addition to the original geometric transformation enhancements such as random flipping, rotation, and scaling, new enhancement methods specific to the coal mine scene are added, including random adjustment of image illumination intensity (simulating different lighting conditions underground), addition of coal dust noise (simulating underground dust interference), and adaptive adjustment of image contrast to improve the model's generalization ability.

[0138] S3.2.3: Training Data Distribution

[0139] The training set (1600 images) and the validation set (400 images) are divided in an 8:2 ratio. The training set contains 240 seed-labeled images and 1360 pseudo-labeled images, while the validation set contains 60 seed-labeled images and 340 pseudo-labeled images to ensure the authenticity of the validation set.

[0140] S3.2.4: Loss Function

[0141] A multi-task joint loss function is employed to collaboratively optimize the semantic segmentation branch and the geometric parameter regression branch. The joint loss function consists of a weighted sum of the semantic segmentation loss and the geometric parameter regression loss; specifically, the semantic segmentation loss is a weighted sum of the Dice loss and the cross-entropy loss, while the geometric parameter regression loss uses a smoothed L1 loss function. The formula for the multi-task joint loss function is as follows:

[0142]

[0143] The coal and rock feature recognition model employs a multi-task joint total loss, which integrates semantic segmentation loss and geometric parameter regression loss to simultaneously optimize the accuracy of coal and rock feature segmentation and the accuracy of geological parameter prediction, thus adapting to the recognition needs of underground coal and rock interfaces, structures, and fractures. Semantic segmentation loss is used to optimize the segmentation mask accuracy of coal and rock features, ensuring accurate spatial positioning of target areas such as coal and rock interfaces, faults, folds, and fissures (see the detailed explanation below for specific parameter meanings). The weighting coefficient of the geometric parameter regression loss (fixed value 0.8) balances the training priority of the segmentation task and the regression task, ensuring both the basic accuracy of segmentation and the reliability of geological geometric parameter prediction. Geometric parameter regression loss, using the SmoothL1 loss function, is used to optimize the prediction accuracy of 12-dimensional geological parameters such as fault strike / dip and fracture size (see the detailed explanation below for the specific meaning of the parameters). The semantic segmentation loss (a weighted sum of Dice loss and cross-entropy loss, with a weight ratio of 1:1) is used to optimize feature segmentation accuracy. The geometric parametric regression loss (using L1 smoothing loss) is calculated as follows:

[0144]

[0145]

[0146] : Semantic segmentation loss, which is a weighted fusion of Dice loss and cross-entropy loss to alleviate the problem of imbalanced feature categories in underground coal and rock (such as the low proportion of crack pixels) and improve the accuracy of the segmentation mask. 0.5: The weight coefficient of the Dice loss (with a weight ratio of 1:1 to the cross-entropy loss), balancing the optimization contributions of the two types of losses to ensure both the overlap of segmented regions and optimize pixel-level classification accuracy. Dice: Dice similarity coefficient, which measures the degree of overlap between the model's predicted segmentation mask and the actual segmentation mask, with a value range of... The closer the value is to 1, the higher the overlap. 1-Dice: The core calculation term of Dice loss, which converts the Dice coefficient into a loss value. The smaller the loss, the higher the segmentation overlap, focusing on optimizing the segmentation effect of small targets (such as small cracks). CrossEntropy: Cross Entropy Loss, which measures the difference between the pixel class probability predicted by the model and the true class label, optimizing the confidence of pixel-level classification and improving the clarity of segmentation boundaries. Geometric parameter regression loss is used to optimize the prediction accuracy of 12-dimensional geological geometric parameters and mitigate the impact of outliers (such as parameter measurement bias under extreme geological conditions) on training. : Geometric parameter index (value range 1-12), corresponding to 12 core geological parameters, including fault strike, fault dip angle, fold axial plane dip angle, fracture length, fracture width, fracture strike, etc. 12: Total number of dimensions of geometric parameters, covering the key geometric information required for coal and rock feature identification, providing quantitative constraints for 3D geological modeling. SmoothL1: Smoothing L1 loss function, combining the advantages of L1 and L2 loss, when the prediction bias is small. The squared loss (L2 property) is used, and the absolute loss (L1 property) is used when the deviation is large, which effectively alleviates the gradient explosion problem caused by outliers. : No. The ground truth of each geometric parameter is determined based on geological expert annotations and measured data (such as fault dip angle measured by boreholes and fracture length annotated in images). : No. The model predictions for each geometric parameter are output by the fully connected regression branch after the U-Net++ decoder. : No. The prediction bias of each geometric parameter reflects the difference between the model's predicted value and the actual value. The total loss of geometric parameter regression is obtained by summing the SmoothL1 loss values ​​of the 12 geometric parameters. The smaller the loss, the higher the prediction accuracy of all parameters.

[0147] The feature recognition output of the deep learning model has two results: one is a semantic segmentation dataset with spatiotemporal labels (including spatial location masks of coal-rock interfaces, geological structures, and fractures); the other is a structured geometric parameter table (including fault strike / dip angle, fold morphology parameters, fracture density, and morphology parameters), which provides precise constraints for subsequent modeling.

[0148] In the coal and rock feature recognition task of this invention, the ratio of targets such as coal and rock interfaces and fractures to the background is often unbalanced (e.g., the pixel proportion of fractures is extremely small). Simple cross-entropy loss will bias towards the dominant class (background) due to class imbalance. Introducing Dice loss can effectively improve the segmentation sensitivity of small targets (such as fractures). The weighted sum of the two (weight ratio 1:1) is used as the semantic segmentation loss, which has been verified to obtain the optimal mIoU on such imbalanced datasets. At the same time, for regression tasks of geometric parameters such as fault dip angle and fracture length, the smooth L1 loss function is more robust to outliers (such as labeling errors) than the mean squared error loss and can provide a more stable training gradient. Therefore, the above-mentioned specific combination of loss functions to construct a multi-task joint loss function is a collaborative solution to the two key challenges of highly imbalanced classes and the possible outliers in geometric parameter labeling in underground coal and rock image feature recognition. Compared with a single loss function or other common combinations, it can effectively improve the recognition accuracy.

[0149] In step S3 of this embodiment, an efficient and accurate coal and rock feature recognition model is constructed. It adopts a dual-output architecture of "U-Net++ semantic segmentation backbone network + geometric parameter regression branch" and combines a training strategy of "a small number of expert annotations + semi-supervised pseudo-annotations" to alleviate the annotation bottleneck. At the same time, it realizes the semantic segmentation of features such as coal and rock interfaces, faults, folds, and fissures and the accurate regression of 12-dimensional geological geometric parameters, providing quantitative constraints for modeling.

[0150] S4: Using the semantic segmentation results as spatial geometric constraints, and combining them with the geological data, engineering data, and gas monitoring data, a three-dimensional geological model reflecting the gas occurrence state is constructed by performing data fusion through a spatial interpolation algorithm.

[0151] Step S4 is used to achieve data fusion and high-precision modeling.

[0152] S4.1: Data Fusion Strategy

[0153] Hard constraint setting: The coal-rock interface and fault location obtained by AI visual recognition are used as hard constraint conditions to constrain the geometric boundaries in 3D modeling.

[0154] The spatial interpolation algorithm is co-kriging interpolation, which uses gas monitoring data (pressure, content, permeability) as the main variables and geological data (coal seam thickness, lithology) as auxiliary variables to construct a co-kriging interpolation model for interpolation calculations:

[0155]

[0156] : Predicted values ​​of gas monitoring parameters at the interpolation point (units determined according to specific parameters, such as gas pressure in MPa, gas content in MPa, etc.). The air permeability coefficient is This corresponds to the gas physical parameters of a node in the 3D modeling mesh and is one of the core outputs of the model. : The three-dimensional spatial coordinates of the point to be interpolated ( Based on the coal mine's independent coordinate system, the node positions of the corresponding modeling grid are determined, covering the entire mining area. The global mean of the main variables of the gas monitoring parameters (with units consistent with the gas parameters) is obtained by statistical calculation of data from all known gas parameter sampling points, providing a benchmark reference for interpolation. The number of known sampling points for the main variables of gas monitoring parameters, i.e., the total number of valid data points for parameters such as gas pressure, content, and permeability coefficient that have been collected. Gas monitoring parameter main variable sampling point index (value range) ), corresponding to the A number of known gas parameter sampling points. : No. The measured values ​​of gas parameters (units consistent with gas parameters) at known sampling points are obtained by a distributed sensor array and after preprocessing. : No. The deviation of the gas parameter value at a known sampling point from the global mean is used to quantify the influence of that sampling point on the interpolation result. : No. The interpolation weights (dimensionless) of the sampling points of the main gas parameter are obtained by fitting the Gaussian variogram. The weights are related to the distance from the sampling point to the interpolation point and the spatial correlation. The closer the distance and the stronger the correlation, the greater the weight. The weighted sum of the main variables of the gas monitoring parameters reflects the comprehensive contribution of all known gas sampling points to the parameter value of the point to be interpolated. The global mean of geological auxiliary variables (the unit is determined according to the auxiliary variables, such as coal seam thickness in meters) is obtained by statistical calculation through sampling points of known geological data and is used to help improve the interpolation accuracy of gas monitoring parameters. The number of known sampling points for geological auxiliary variables, i.e., the total number of valid data points for geological data such as coal seam thickness and lithology that have been collected. Geological auxiliary variable sampling point index (value range) ), corresponding to the A known geological baseline data sampling point. : No. Measured values ​​of geological auxiliary variables (units consistent with auxiliary variables) from known sampling points, derived from preprocessed results such as geological exploration borehole columnar sections and geophysical data. : No. The deviation between the geological auxiliary variable values ​​of a known sampling point and the global mean is used to quantify the auxiliary influence of the geological sampling point on the interpolation of gas monitoring parameters. : No. The interpolation weights (dimensionless) of the sampling points of each geological auxiliary variable are obtained by fitting the spatial correlation between the main variable and the auxiliary variable through a Gaussian variogram function, which reflects the degree of correlation between geological conditions and gas occurrence. The weighted bias sum of geological auxiliary variables, through the spatial correlation between geological data and gas monitoring parameters, helps to correct the interpolation results of gas parameters and improve the modeling accuracy.

[0157] Data density adaptability control: In a preferred embodiment, to support grid resolution requirements, a threshold for the original data spatial distribution is set: borehole data spacing. Sensor deployment density Individual, visual recognition data coverage density Zhang; if the data density in a local area does not meet the standard (such as borehole spacing) The system automatically employs a strategy of "local encrypted interpolation + uncertainty labeling" to add confidence labels (0-1 interval) to the inferred region, with confidence levels... Areas requiring additional data collection should be prompted.

[0158] S4.2: 3D Model Construction

[0159] Geometric modeling: Based on the coal-rock interface and geological structure identification results, a triangulation algorithm is used to construct a three-dimensional geometric model of the coal seam and rock strata to reflect the geological spatial morphology;

[0160] Physical parameter modeling: The interpolated gas pressure, content, and permeability coefficient are mapped to the geometric model mesh nodes to form a multi-field coupled three-dimensional model;

[0161] Mesh resolution optimization: Adaptive resolution mesh is used instead of fixed resolution. The core area (within 50m of the mining face) maintains a high precision of 0.5m×0.5m×0.5m, while the non-core area (far from the mining face) is automatically adjusted to a resolution of 1m×1m×1m, balancing modeling accuracy and computational resource consumption.

[0162] Model format: The output is in an industry standard format (such as VTK, STL), which supports integration with coal mine safety monitoring systems and mining scheduling systems.

[0163] As a specific implementation example of step S4, the data fusion and modeling process is as follows: First, the coal-rock interface point cloud and fault lines extracted from the semantic segmentation results are used as the absolute constraint boundaries for triangulation and geometric modeling. Next, physical parameter interpolation is performed: the gas content measured in the borehole is used as the main variable Z(x), and the coal seam thickness identified and converted visually is used as the auxiliary variable Y(x). The spatial correlation structure of the main and auxiliary variables is fitted using a Gaussian variogram model, and the weighting coefficients are solved based on the co-kriging equations. and For the core area within 50 meters of the mining face, a dense interpolation was performed using a grid resolution of 0.5m × 0.5m × 0.5m; for areas outside this range, the resolution was automatically increased to 1m × 1m × 1m. Finally, the interpolated gas parameter field was mapped onto a constrained geometric grid to generate a 3D geological model that displays gas pressure contour maps and content isosurfaces, and output in VTK format.

[0164] like Figure 4 As shown, this 3D model, based on the identification results of coal-rock interfaces and geological structures, comprehensively employs triangulation geometric modeling and multiphysics parameter mapping methods to construct a high-precision 3D geological model that can realistically reflect the characteristics of coal mine gas occurrence. At the geometric level, the model effectively expresses the control of complex geological structures such as folds and faults on coal seam distribution by meticulously depicting the spatial morphology of the coal seam and the surrounding rocks. At the physical parameter level, key parameters such as gas pressure, gas content, and permeability coefficient, after spatial interpolation, are mapped to 3D mesh nodes, achieving continuous spatial representation and multi-field coupled characterization of the gas occurrence state. Simultaneously, the model introduces an adaptive resolution mesh strategy, using a high-resolution mesh in the mining face and surrounding high-risk areas to enhance the characterization of abnormal gas accumulation and gradient changes, while automatically reducing the mesh resolution in areas far from the mining face, thereby significantly reducing computational and storage costs while ensuring modeling accuracy. The final model can be output in industry standard formats such as VTK and STL, and has good system compatibility and scalability. It can provide a reliable three-dimensional data foundation and decision support for coal mine gas disaster prediction, safety monitoring and intelligent mining scheduling.

[0165] Step S4 achieves high-precision modeling and dynamic updating through multi-source data fusion. It uses AI visual recognition results as hard constraints, combines geological, engineering, and gas parameter data, and balances modeling accuracy and computational resource consumption through collaborative kriging interpolation algorithm and adaptive resolution grid modeling.

[0166] like Figure 5 As shown, this embodiment of the invention also provides a three-dimensional geological modeling system for coal mine gas occurrence based on AI vision, used to implement the aforementioned three-dimensional geological modeling method for coal mine gas occurrence based on AI vision. The system includes:

[0167] The data acquisition module 210 is used to acquire multi-source data of the underground mining area of ​​the coal mine. The multi-source data includes at least geological basic data, engineering dynamic data, gas monitoring data, and visual image data acquired by underground image acquisition equipment.

[0168] The image enhancement processing module 220 is used to perform low-light enhancement processing on the visual image data based on the image enhancement model to obtain an enhanced image; wherein, the low-light enhancement processing includes: estimating the camera response function and linearizing the original image, decomposing the linearized image in the frequency domain to separate light effect interference and detail texture, and performing light effect suppression, noise removal and dynamic range enhancement on the decomposed features respectively, and finally fusing to generate an enhanced image;

[0169] The coal and rock feature recognition module 230 is used to perform coal and rock feature recognition based on the enhanced image using a deep learning model, and output semantic segmentation results including coal and rock interfaces, geological structures and fractures, as well as corresponding geological geometric parameters.

[0170] The data fusion and 3D modeling module 240 is used to use semantic segmentation results as spatial geometric constraints, combine them with geological basic data, engineering dynamic data and gas monitoring data, and perform data fusion through spatial interpolation algorithms to construct a 3D geological model that reflects the gas occurrence state.

[0171] The AI-based vision-based three-dimensional geological modeling system for coal mine gas occurrence provided in this embodiment of the invention can execute the AI-based vision-based three-dimensional geological modeling method for coal mine gas occurrence provided in any of the above embodiments of the invention. It has the corresponding functions and beneficial effects of executing the AI-based vision-based three-dimensional geological modeling method for coal mine gas occurrence. For detailed process, please refer to the relevant operations of the AI-based vision-based three-dimensional geological modeling method for coal mine gas occurrence in the foregoing embodiments.

[0172] According to the above embodiments of the present invention, the AI-based vision-based 3D geological modeling method and system for coal mine gas occurrence firstly overcomes interference from low light and strong reflections by using a cascaded image enhancement model specifically designed for underground environments, achieving high-quality image restoration. Secondly, a dual-branch deep learning model is used to simultaneously extract the semantic contours and precise geometric parameters of geological targets from the images, and a semi-supervised training strategy significantly reduces the reliance on expensive expert annotations. Finally, the visual recognition results are used as hard constraints and fused with multi-source geological, engineering, and monitoring data. Through co-Kriging interpolation and adaptive mesh technology, a high-precision 3D model that accurately reflects the spatial distribution of gas and its correlation with geological structures is constructed, forming a complete technical closed loop from "seeing" to "understanding" to "construction," providing a reliable data foundation for intelligent early warning and precise management of gas disasters.

[0173] like Figure 6 As shown, by comparing the root mean square error (RMSE, unit: %) of the four algorithms under different training iterations, the performance advantages of the proposed technical solution can be intuitively reflected. Overall, as the number of training iterations increases from 100 to 1000, the RMSE of all four algorithms shows a continuous downward trend, reflecting the effectiveness of parameter optimization during model training. However, there is a significant difference between the rate of decrease and the final convergence accuracy: among the two algorithms without image enhancement modules, the RMSE without image enhancement + CNN remains at the highest level, decreasing from the initial 3.5% to 1.6%, while the RMSE without image enhancement + U-Net++, thanks to its superior network architecture design, decreases from 3.2% to 1.1%, with a convergence accuracy improvement of approximately 31.3%, verifying the adaptability of semantic segmentation networks in coal and rock feature recognition tasks; after introducing the underground low-light image enhancement module of this invention, the RMSE of image enhancement module + CNN decreases from 3.0% to 0.8%, with a final convergence accuracy improvement of 50.0% compared to without image enhancement + CNN, fully demonstrating the effectiveness of the underground low-light image enhancement module in improving... The key role of feature recognition quality lies in suppressing interference such as halos and reflections and improving dynamic range, providing high-quality data support for subsequent feature extraction. The algorithm of this invention (underground low-light image enhancement module + coal and rock feature recognition model) performs the best, with its RMSE rapidly decreasing from the initial 2.8% to 0.6%. It not only has the fastest convergence rate, but also significantly outperforms the other three comparison algorithms in terms of final accuracy, improving by 25.0% compared to the enhancement module + CNN and by 45.5% compared to the unenhanced image + U-Net++. This result is attributed to the synergistic effect of "high-quality image input from the enhancement module + U-Net++ semantic segmentation and geometric parameter regression dual-branch architecture + collaboratively optimized loss function system". It not only solves the bottleneck of underground image quality, but also achieves a leapfrog improvement in feature recognition accuracy through dual-output architecture and precise loss function optimization, providing a more reliable quantitative constraint basis for subsequent three-dimensional geological modeling of coal mine gas occurrence.

[0174] It should also be noted that, in the embodiments of this application, relational terms such as "first" and "second" are used only to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any such actual relationship or order between these entities or operations. Furthermore, the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such a process, method, article, or apparatus. Without further limitations, an element defined by the phrase "comprising one..." does not exclude the presence of other identical elements in the process, method, article, or apparatus that includes said element.

[0175] The above description of the disclosed embodiments enables those skilled in the art to make or use this application. Various modifications to these embodiments will be readily apparent to those skilled in the art, and the general principles defined in the embodiments of this application may be implemented in other embodiments without departing from the spirit or scope of this application. Therefore, this application is not to be limited to the embodiments shown in this application, but is to be accorded the widest scope consistent with the principles and novel features disclosed in the embodiments of this application.

Claims

1. A three-dimensional geological modeling method for coal mine gas occurrence based on AI vision, characterized in that, Includes the following steps: Acquire multi-source data of the underground mining area of ​​the coal mine, including geological basic data, engineering dynamic data, gas monitoring data, and visual image data acquired by underground image acquisition equipment; The visual image data is subjected to low-light enhancement processing based on the image enhancement model to obtain an enhanced image; wherein, the low-light enhancement processing includes: estimating the camera response function and linearizing the original image, decomposing the linearized image in the frequency domain to separate light effect interference and detail texture, and performing light effect suppression, noise removal and dynamic range enhancement on the decomposed features respectively, and finally fusing to generate the enhanced image. Based on the enhanced image, a deep learning model is used to identify coal and rock features, and the output includes semantic segmentation results containing coal and rock interfaces, geological structures and fractures, as well as the corresponding geological geometric parameters. The deep learning model is a neural network with a dual-output branch structure, including: a semantic segmentation branch, built on an encoder-decoder structure, used to perform pixel-level classification of the input image and output a semantic segmentation mask containing coal-rock interfaces, geological structures, and fractures; and a geometric parameter regression branch, connected to the backbone network of the semantic segmentation branch, used to regress multiple geological geometric parameters such as fault strike, dip angle, and fracture size based on the features extracted by the backbone network. When training the deep learning model, a pseudo-labeled data augmentation strategy based on semi-supervised learning is adopted, which specifically includes: acquiring a small number of images labeled by experts as a seed label set; using an initial model trained based on the seed label set to predict a large number of unlabeled downhole images to obtain prediction results; selecting the portion of the prediction results with a confidence level higher than a preset threshold as pseudo-labeled data; and merging the seed label data with the verified pseudo-labeled data to form a dataset for model training. Using the semantic segmentation results as spatial geometric constraints, and combining them with the geological basic data, engineering dynamic data, and gas monitoring data, a three-dimensional geological model reflecting the gas occurrence state is constructed by performing data fusion through spatial interpolation algorithms. The construction of a three-dimensional geological model reflecting the state of gas occurrence includes: the spatial interpolation algorithm is the co-kriging interpolation method, which uses gas monitoring data as the main variable and geological basic data as the auxiliary variable for interpolation calculation; an adaptive resolution grid division strategy is adopted, using a first resolution for the core area around the mining face and a second resolution lower than the first resolution for non-core areas far from the working face, in order to balance model accuracy and computational cost.

2. The method according to claim 1, characterized in that, The method further includes: preprocessing the multi-source data, specifically including: The geological basic data, engineering dynamic data, and gas monitoring data are processed to unify and standardize units; Wavelet thresholding is used to denoise the geophysical data in the geological basic data, and sliding window mean filtering is used to denoise the sensor data in the gas monitoring data. The three-dimensional spatial coordinates of each data acquisition point are determined based on GPS and the underground positioning system. The spatiotemporal correspondence between multiple data sources is established with the advance time of the mining face as the time reference. The three-dimensional spatial coordinates of each data acquisition point are uniformly transformed to the preset reference coordinate system through a seven-parameter coordinate transformation model based on multiple control points.

3. The method according to claim 1, characterized in that, The image enhancement model is a cascaded processing model built on a convolutional neural network, and its processing flow includes: The original image is subjected to camera response function estimation through a first convolutional network, and the output is used to calculate the parameters for the inverse camera response function, so as to transform the image into a linear space; The linearized image is decomposed into a low-frequency feature map and a high-frequency feature map in the frequency domain. The low-frequency feature map contains illumination and reflection information, and the high-frequency feature map contains texture and noise information. The low-frequency feature map is processed by a second convolutional network to suppress optical interference; the high-frequency feature map is processed by a third convolutional network to remove noise and preserve detailed texture. The dynamic range of the low-frequency feature map after light effect suppression is enhanced by using a fourth convolutional network; The enhanced image is generated by weighted fusion of the low-frequency feature map after dynamic range enhancement and the high-frequency feature map after denoising.

4. The method according to claim 1, characterized in that, The training and optimization of the image enhancement model employs a two-stage semi-supervised learning strategy that includes supervised training and unsupervised fine-tuning, wherein: During the supervised training phase, synthetic image data with known camera response function labels are used. The total loss function for supervised training integrates the mean squared error loss for camera response function estimation, the L1 norm loss for image linearization, and the logarithmic transformation loss for dynamic range enhancement. In the unsupervised fine-tuning stage, unlabeled real-world downhole image data is used to improve the model's adaptability and generalization ability to real-world scenes by optimizing the total unsupervised fine-tuning loss function. The total unsupervised fine-tuning loss function includes: monotonicity loss for constraining the monotonicity of the predicted camera response function, distribution linearization loss for optimizing the distribution characteristics of image edge pixels, reconstruction loss for ensuring the consistency between the enhanced image and the original image content, smoothing loss for constraining the smoothness of the light effect suppression mask space, and grayscale world loss for correcting the overall color shift of the image.

5. The method according to claim 1, characterized in that, When training the deep learning model, data augmentation processing is applied to the input image; the data augmentation processing includes general geometric transformation enhancement and specific enhancement to simulate underground coal mine imaging conditions, the specific enhancement including at least random adjustment of image illumination intensity and noise addition to simulate underground dust interference.

6. The method according to claim 5, characterized in that, When training the deep learning model, a multi-task joint loss function is used to collaboratively optimize the semantic segmentation branch and the geometric parameter regression branch; The joint loss function is composed of a weighted sum of semantic segmentation loss and geometric parameter regression loss; wherein, the semantic segmentation loss is a weighted sum of Dice loss and cross-entropy loss, used to handle the pixel imbalance problem of target categories in the image; the geometric parameter regression loss adopts a smooth L1 loss function, which is robust to outliers in the prediction error.

7. A three-dimensional geological modeling system for coal mine gas occurrence based on AI vision, characterized in that, The system for implementing the method as described in any one of claims 1 to 6 comprises: The data acquisition module is used to acquire multi-source data of the underground mining area of ​​the coal mine. The multi-source data includes geological basic data, engineering dynamic data, gas monitoring data, and visual image data acquired by underground image acquisition equipment. An image enhancement processing module is used to perform low-light enhancement processing on the visual image data based on an image enhancement model to obtain an enhanced image; wherein, the low-light enhancement processing includes: estimating the camera response function and linearizing the original image, decomposing the linearized image in the frequency domain to separate light effect interference and detail texture, and performing light effect suppression, noise removal and dynamic range enhancement on the decomposed features respectively, and finally fusing to generate the enhanced image; The coal and rock feature recognition module is used to perform coal and rock feature recognition based on the enhanced image using a deep learning model, and outputs semantic segmentation results including coal and rock interfaces, geological structures and fractures, as well as corresponding geological geometric parameters. The data fusion and 3D modeling module is used to take the semantic segmentation results as spatial geometric constraints, combine them with the geological basic data, engineering dynamic data and gas monitoring data, and perform data fusion through spatial interpolation algorithms to construct a 3D geological model that reflects the gas occurrence state.

Citation Information

Patent Citations

  • Complex geological condition-based gas prevention and control robot cluster control method and system

    CN117369254A

  • Deep learning-based weak intercalated layer semantic segmentation method and device, and medium

    CN120807923A