Multi-angle coal gangue stock dynamic monitoring method based on multi-modal spatio-temporal perception
By integrating multimodal spatiotemporal sensing technology with multi-view satellite imagery information and optimizing geometric relationships and feature matching, the problems of low efficiency and insufficient accuracy in traditional coal gangue inventory monitoring have been solved, and high-precision dynamic monitoring of coal gangue inventory has been achieved.
Patent Information
- Application Number
- CN202510748968.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Priority Date
- 2025-06-04
- Filing Date
- 2025-06-06
- Publication Date
- 2025-11-25
- Estimated Expiration
- 2045-06-06
AI Technical Summary
Traditional methods for monitoring coal gangue inventory are inefficient and lack precision, failing to accurately acquire three-dimensional information and making dynamic monitoring difficult.
By constructing a multi-modal spatiotemporal sensing method for monitoring coal gangue inventory, integrating radiometric, spectral, spatial, and temporal information from multi-view satellite images, optimizing geometric relationships using the mapping collinearity equation, performing sub-pixel-level spatial alignment and multispectral feature-guided image compensation matching, constructing the three-dimensional structure of coal gangue, and calculating its volume.
It achieves high-precision and high-efficiency dynamic monitoring of coal gangue inventory, and can accurately monitor changes in coal gangue inventory under complex terrain and variable weather conditions, providing timely and reliable data support for enterprise resource management and environmental protection.
Smart Images

Figure CN120635026B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of coal gangue inventory monitoring, and in particular to a multi-angle dynamic monitoring method for coal gangue inventory based on multimodal spatiotemporal sensing. Background Technology
[0002] Coal, as a crucial global energy source, occupies a vital position in the energy structure. However, the mining and washing of coal inevitably generates a large amount of coal gangue. According to incomplete statistics, approximately 0.1-0.2 tons of coal gangue are produced for every ton of coal mined. Improper disposal of such a large quantity of coal gangue will trigger a series of serious problems. From a resource utilization perspective, coal gangue is not entirely useless; it contains a certain amount of carbon and other valuable elements. If the stock of coal gangue can be accurately assessed and its secondary development and utilization can be achieved through reasonable technical means, it can not only realize resource recycling and reuse but also create additional economic benefits for enterprises. For example, some coal gangue can be used to produce building materials such as bricks and cement, effectively alleviating the demand pressure on the building materials market for resources such as natural sand and gravel. From an environmental protection perspective, the long-term open-air storage of coal gangue allows harmful substances such as sulfur and heavy metals to be gradually released and seep into the surrounding soil and water bodies through natural processes such as rainwater leaching and weathering. This leads to environmental problems such as soil pollution, eutrophication of water bodies, and excessive levels of heavy metals. Furthermore, coal gangue piles also release greenhouse gases such as methane, exacerbating global climate change. Therefore, accurately monitoring the stock and dynamic changes of coal gangue is crucial for developing scientifically sound environmental protection measures and reducing the risk of environmental pollution.
[0003] Traditional methods for monitoring coal gangue inventory mainly include on-site measurement and simple planar remote sensing monitoring. On-site measurement typically involves manual sampling and tape measurement, which requires a large amount of manpower and is extremely inefficient. Because coal gangue dumping areas are often characterized by complex terrain and vast areas, on-site measurement cannot achieve comprehensive coverage, resulting in severely insufficient data representativeness and accuracy. For example, in some large coal gangue dumps, on-site measurement may only cover a small portion of the area, failing to accurately grasp the complex terrain and accumulation conditions within the dump, leading to significant discrepancies between estimated coal gangue inventory and actual conditions. While simple planar remote sensing monitoring improves monitoring efficiency to some extent, it is limited by a single perspective and cannot obtain three-dimensional structural information about the coal gangue pile. Planar remote sensing images only reflect the distribution of coal gangue piles on a two-dimensional plane, failing to accurately obtain key three-dimensional information such as height and slope. This often results in only a rough estimation method when estimating coal gangue inventory, leading to large errors in the estimation results. In addition, traditional remote sensing monitoring methods lack effective integration and utilization of multimodal information during data processing, making it difficult to fully explore the spatiotemporal characteristics in satellite imagery and achieve dynamic monitoring of coal gangue reserves. Summary of the Invention
[0004] The purpose of this invention is to address the technical problems identified in the background section by providing a multi-angle dynamic monitoring method for coal gangue inventory based on multimodal spatiotemporal perception. This method creatively integrates multimodal information from multi-view satellite imagery, including radiometric, spectral, spatial, and temporal data. By constructing a viewpoint-adaptive image registration module, sub-pixel-level spatial alignment and multispectral feature-guided image compensation matching are achieved, significantly improving the accuracy and stability of multi-view image registration. Furthermore, the geometric relationships in the multi-view three-dimensional space are optimized using a mapping collinearity equation, significantly enhancing the accuracy of 3D reconstruction and providing solid technical support for the precise calculation of coal gangue inventory.
[0005] The objective of this invention is achieved through the following technical solution:
[0006] A multi-angle dynamic monitoring method for coal gangue inventory based on multimodal spatiotemporal perception is proposed. The method includes: constructing a multi-angle satellite coal gangue inventory monitoring model incorporating a coal gangue identification model; the coal gangue identification model performs coal gangue identification and detection on the acquired multi-angle satellite image data sets of the study area and constructs a collinearity equation for converting pixel coordinates to three-dimensional geographic coordinates; integrating cross-viewpoint and spatiotemporal features of the multi-angle satellite image data sets to perform location point feature matching and complete feature matching of the multi-angle satellite image data sets; and performing feature extraction, height-sensing encoding processing, spectrally guided cross-attention processing, and sub-pixel spatial encoding on the multi-angle satellite image data sets to obtain a multi-view stereo image. ; Multi-view stereoscopic The coal gangue is identified, detected, and segmented by homography projection onto a multi-view projection perspective plane. Then, the three-dimensional structure of the coal gangue is fused together, the volume of the three-dimensional structure of the coal gangue is calculated, and the volume of the coal gangue is obtained by scaling it proportionally.
[0007] To better implement this invention, the coal gangue identification model is trained using a multi-angle coal gangue sample dataset. This dataset includes coal gangue image samples and coal gangue target annotation information. The target annotation information includes bounding boxes and target categories, with target categories including coal gangue piles and scattered coal gangue. The coal gangue image samples contain satellite images of coal gangue under various weather, lighting, and / or seasonal conditions. Both the coal gangue image samples in the multi-angle coal gangue sample dataset and the satellite image data in the multi-angle satellite image dataset undergo denoising processing. Data augmentation processing is also performed on the coal gangue image samples in the multi-angle coal gangue sample dataset.
[0008] Preferably, the collinearity equation expression for converting the pixel coordinates (u, v) of the satellite image data in the multi-angle satellite image data set into three-dimensional geographic coordinates (X, Y, Z) is as follows:
[0009] ,in( ) is the initial attitude angle of the satellite, ( ( ) represents the satellite position, ( The elements of the rotation matrix are composed of attitude angles; the feature vector f is obtained by averaging the satellite image data in the multi-angle satellite image dataset through a convolutional layer to obtain the feature vector f in the spatial dimension, and thus the attitude angle correction amount. Perform attitude angle correction and update the mapping collinearity equation.
[0010] Preferably, the feature matching method for the multi-angle satellite image data set includes: extracting the band spectral characteristics of the satellite image data in the multi-angle satellite image data set and generating band weights using a multilayer perceptron for band feature enhancement; using a Transformer layer within the multi-angle satellite image data set to perform band feature, spatiotemporal feature, and cross-angle location point feature matching; and selecting high-confidence matching point pairs for geometric constraint compensation.
[0011] Preferably, the multi-view stereoscopic Methods of obtaining include:
[0012] S211. The feature maps of the multi-angle satellite image data set are extracted using convolutional layers and residual blocks to output multi-level feature maps. At the same time, the feature map height information of the multi-angle satellite image data set is extracted through height-aware coding and encoded on the multi-level feature maps. The preliminary multi-view stereo is reconstructed at time t.
[0013] S212, Multi-view stereoscopic view of time t-1 The stereo query input is fed into the subpixel-level spatial alignment module to calculate the subpixel-level offset between time t and time t-1; then the spectral-guided cross-attention module is fed into the module to calculate the spectral similarity and adjust the depth and weight of feature extraction in method S211;
[0014] S213. Using the Feed Forward module, feature fusion and propagation are performed through several fully connected layers to obtain a multi-view stereo image at time t. .
[0015] Preferably, in method S212, multi-view stereoscopic... Subpixel spatial update encoding is performed on the multi-level feature map after time t adjustment with subpixel-level offset.
[0016] Preferably, the method for obtaining the three-dimensional structure of coal gangue includes:
[0017] S221, Multi-view stereoscopic By using homography projection onto a multi-view projection perspective plane, a target detection submodule is used to identify and detect coal gangue. The target detection submodule is trained using coal gangue sample data.
[0018] S222. The target segmentation submodule identifies and extracts boundary information to segment the coal gangue target; then, it is 3D reconstructed and fused into a 3D structure of coal gangue.
[0019] Preferably, the volume of a single coal gangue three-dimensional structure is calculated in the multi-angle satellite coal gangue inventory monitoring model, and then scaled down to the actual scale of the study area to obtain the corresponding coal gangue volume; the volume of all coal gangue three-dimensional structures in the study area is calculated and the corresponding coal gangue volume and the total volume of all coal gangue are obtained respectively.
[0020] Preferably, both the coal gangue identification model and the coal gangue identification and detection using the multi-view projection perspective plane in multi-view stereoscopic model employ the following loss function constraints:
[0021] ;
[0022] ;
[0023] ,in The weights for distance loss and aspect ratio loss are... To predict the bounding box, For the true bounding box, To predict the width of the bounding box, The width of the actual bounding box. To predict the height of the bounding box, The height of the actual bounding box. The scaling parameter is used to scale the distance loss. The scaling parameter is the loss due to scaling the width. The scaling parameter is used to calculate the height loss. The hyperparameters used to control the curvature of the curve, For intersection, union, and comparison, To compare the expected loss with the expected loss, To compare the losses, For distance loss, For aspect ratio loss, The loss result of the loss function;
[0024] The coal gangue identification model and the coal gangue segmentation of the multi-view stereoscopic multi-view projection perspective plane both adopt the following loss function constraints:
[0025] ,in The loss value is used to measure the predicted segmentation result and the actual segmentation result, where N is the number of pixels in the input image and C is the number of categories in the semantic segmentation task. Let be the predicted probability that the i-th pixel belongs to the j-th class. This is the true class label of the i-th pixel, which is 1 if it belongs to the j-th class, and 0 otherwise.
[0026] The gradient descent algorithm is used to reduce the model loss value, while optimizing and updating the model parameters until the minimum number of iterations or the total loss of coal gangue identification, detection and segmentation is reached.
[0027] A multi-angle dynamic monitoring system for coal gangue inventory, comprising a multi-angle satellite coal gangue inventory monitoring model, including a coal gangue identification model, a mapping collinearity equation module, a feature matching module, a multi-view stereo construction module, and a coal gangue volume calculation module. The coal gangue identification model performs coal gangue identification and detection on the acquired multi-angle satellite image data sets of the study area. The mapping collinearity equation module constructs mapping collinearity equations that convert pixel coordinates to three-dimensional geographic coordinates. The feature matching module is used to perform location point feature matching by integrating cross-view features and spatiotemporal features of the multi-angle satellite image data sets, and to complete the feature matching of the multi-angle satellite image data sets. The multi-view stereo construction module is used to perform feature extraction, height-sensory encoding processing, spectrally guided cross-attention processing, and sub-pixel spatial encoding on the multi-angle satellite image data sets to obtain multi-view stereo. The coal gangue volume calculation module will perform multi-view three-dimensional calculations. The coal gangue is identified, detected, and segmented by homography projection onto a multi-view projection perspective plane. Then, the three-dimensional structure of the coal gangue is fused together, the volume of the three-dimensional structure of the coal gangue is calculated, and the volume of the coal gangue is obtained by scaling it proportionally.
[0028] Compared with the prior art, the present invention has the following advantages and beneficial effects:
[0029] (1) This invention performs unified mapping collinearity equation processing on multi-angle satellite images in a multi-angle satellite image data set, integrates cross-view features and spatiotemporal features of the multi-angle satellite image data set to perform location point feature matching and complete feature matching of the multi-angle satellite image data set; performs feature extraction, height-sensory encoding processing, spectral-guided cross-attention processing, and sub-pixel spatial encoding on the multi-angle satellite image data set to obtain multi-view stereo, and gradually fuses multiple satellite image data from the multi-angle satellite image data set to extract multi-modal information such as radiation, spectrum, spatiotemporal, and height and fuses them to construct a multi-view stereo, and then uses homography projection multi-view projection perspective plane to perform coal gangue identification, detection, and segmentation processing, achieving high-precision and high-efficiency acquisition of coal gangue volume. This invention updates the mapping collinearity equation over time and uses the aforementioned time multi-view stereo and stereo query to perform spectral similarity identification between previous and subsequent times, which can dynamically adjust the feature processing depth and weight of the multi-angle satellite image data set to continue accurate extraction and realize dynamic monitoring of multi-angle coal gangue inventory changes.
[0030] (2) This invention creatively integrates multimodal information such as radiometric, spectral, spatial, and temporal data from multi-view satellite imagery. By constructing a view-adaptive image registration module, it achieves sub-pixel-level spatial alignment and multispectral feature-guided image compensation matching, greatly improving the accuracy and stability of multi-view image registration. The geometric relationships of the multi-view three-dimensional space are optimized using a mapping collinearity equation, significantly improving the accuracy of three-dimensional reconstruction. This provides solid technical support for the accurate calculation of coal gangue reserves and timely and accurate data support for coal enterprises' production management and resource planning.
[0031] (3) This invention, through perspective adaptive registration, multi-view feature extraction and fusion mechanism, coal gangue three-dimensional detection and segmentation network and dynamic monitoring and adjustment, can accurately monitor the coal gangue stock and its dynamic changes under complex terrain and variable weather conditions, and provide reliable and timely data support and technical guarantee for practical application scenarios such as coal industry resource management, environmental protection and production planning.
[0032] (4) This invention achieves precise alignment of feature maps by using radiometric correction and multispectral feature-guided image compensation matching, thereby effectively solving the problems of geometric distortion and feature matching in multi-view image registration. It can effectively identify and match complex ground features and improve the robustness and accuracy of coal gangue identification.
[0033] (5) The present invention designs a multi-view feature extraction and fusion mechanism, extracts multi-level feature maps of multi-view images, inputs the multi-view stereo and stereo query calculated in the previous time into the sub-pixel level spatial alignment module to calculate the pixel offset, obtains the actual spatial encoding position of the multi-view feature pixels, and finally obtains the multi-view stereo MPSt through the Feed Forward module. This mechanism effectively integrates the spatiotemporal correlation of features, corrects the multi-view stereo constructed in the previous time period, reduces the time and resource overhead of reconstructing multi-view stereo, and improves the model's ability to describe the spatial distribution and morphological features of coal gangue by combining the positional correlation of objects in the time series.
[0034] (6) The present invention maps the multi-view stereo MPSt to the multi-view projection plane space using the homography transformation module, constructs the multi-view parallax space, and accurately locates and segments the coal gangue target again. Finally, the actual stock of the coal gangue pile is calculated based on the pixel volume. This technical approach effectively eliminates the feature deviation caused by the difference in viewpoint and improves the model's understanding and monitoring ability of the three-dimensional spatial distribution of coal gangue. Attached Figure Description
[0035] Figure 1 This is a flowchart of the multi-angle dynamic monitoring method for coal gangue inventory of the present invention;
[0036] Figure 2 This is a schematic diagram illustrating the principle of the multi-angle dynamic monitoring method for coal gangue inventory in the embodiment.
[0037] Figure 3 This is a schematic diagram illustrating the cross-attention principle of multispectral feature-guided image compensation matching in an embodiment.
[0038] Figure 4 This is a flowchart illustrating the multi-view stereoscopic acquisition method in the embodiment.
[0039] Figure 5 This is a flowchart illustrating the method for obtaining the three-dimensional structure of coal gangue in this embodiment. Detailed Implementation
[0040] The present invention will be further described in detail below with reference to embodiments:
[0041] Example
[0042] like Figure 1 As shown, a multi-angle dynamic monitoring method for coal gangue inventory based on multimodal spatiotemporal sensing is proposed, the method comprising:
[0043] S1. Construct a multi-angle satellite coal gangue inventory monitoring model that includes a coal gangue identification model. The coal gangue identification model performs coal gangue identification and detection on the multi-angle satellite image data sets of the acquired study area and constructs a collinearity equation for converting pixel coordinates into three-dimensional geographic coordinates. For example... Figure 2 As shown, the satellite uses multi-view cameras to collect satellite image data of a certain area from multiple angles (e.g., Figure 2 The collection of satellite image data from four different angles constitutes a multi-angle satellite image data set. Figure 2 Multi-angle imagery (i.e., satellite image data) combines satellite image data from multiple angles of a certain area into a multi-angle satellite image data set.
[0044] The coal gangue identification model is trained using a multi-angle coal gangue sample dataset. This dataset includes coal gangue image samples and target annotation information. The target annotation information includes bounding boxes and target categories, including coal gangue piles and scattered coal gangue. The coal gangue image samples contain satellite images of coal gangue under various weather, lighting, and / or seasonal conditions (the data range of these satellite images covers multiple scenarios such as different mining areas and coal gangue dumping areas). To improve the diversity and representativeness of coal gangue detection, each satellite image may contain different types of coal gangue, and each type of coal gangue differs in morphology, size, and distribution, ensuring the diversity and comprehensiveness of the multi-angle coal gangue sample dataset. The experimental parameters for the coal gangue identification model are shown in Table 1, and the performance of the experimental server is shown in Table 2.
[0045] Table 1 Network Parameter Settings
[0046] method Parameter settings DCGM-CWIM Learning rate: 0.0001; Number of iterations: 500; gamma = 0.9; Optimizer: Adam; Loss function: Focal-EIoU Loss
[0047] Table 2 Server Configuration
[0048] category Configuration CPU Intel(R) Xeon(R) Gold5118 CPU @ 2.30GHz graphics card NVIDIA GeForce RTX 2080 Ti RAM 32 GB operating system Windows 10
[0049] Both the coal gangue image samples in the multi-angle coal gangue sample dataset and the satellite image data in the multi-angle satellite image dataset undergo denoising and other preprocessing. The preprocessing includes: 1. Denoising: The coal gangue image samples (including satellite image sample data) in the multi-angle coal gangue sample dataset or the satellite image data in the multi-angle satellite image dataset are often affected by various noises (including atmospheric interference, sensor noise, etc.). These noises affect image quality and consequently the effectiveness of feature extraction. To eliminate these noises, filtering techniques are used for denoising. This embodiment employs median filtering, mean filtering, and Gaussian filtering, among others. These methods effectively remove random noise from the images while preserving important image features. 2. Image standardization: Since satellite imagery typically has high resolution, the input image data is standardized to ensure consistent processing. All image data is adjusted to the same size, which not only improves training efficiency but also enables the model to process data efficiently and avoids problems caused by inconsistent data sizes. The labeling information for coal gangue targets was generated by labeling each coal gangue in the image using the LabelMe tool. The labeling data included the location of the target detection box (the coordinates of the detection box) and the category of the target detection box (such as coal gangue pile, scattered coal gangue, etc.). The .json file generated by the labeling was converted into the standard COCO dataset format to achieve the standardization and unification of the dataset.
[0050] Data augmentation processing was performed on coal gangue image samples from a multi-angle coal gangue sample dataset. This included rotating the image or rotating it by randomly selecting an angle (e.g., -30° to +30°) to enable the model to learn object features from different perspectives; mirroring the image by horizontally or vertically flipping it to enhance the model's ability to recognize coal gangue in different directions; random cropping, randomly cropping a region from the original image to simulate the object's behavior at different positions and sizes, helping the model better adapt to changes in target scale; scaling the image at different scales to enhance the model's ability to recognize coal gangue of different sizes; and color adjustment, simulating recognition scenarios under different lighting conditions by adjusting the image's brightness, contrast, and saturation.
[0051] In some embodiments, the mapping collinearity equation for converting the pixel coordinates (u, v) of satellite image data in a multi-angle satellite image data set into three-dimensional geographic coordinates (X, Y, Z) is preferred in this invention (such as the residual correction collinearity equation). Figure 2 The expression shown is as follows:
[0052] ,in( ) is the initial attitude angle of the satellite, ( ) represent pitch angle, yaw angle, and roll angle, respectively. ( ) represents the satellite position, ( ) are the elements of the rotation matrix composed of attitude angles. Preferably, this embodiment further includes the following method: extracting features from the satellite image data in the multi-angle satellite image data set through a convolutional layer, averaging the features in the spatial dimension to obtain the feature vector f, and obtaining the attitude angle correction amount (). Attitude angle correction and updating of the collinearity equation are performed. In the multi-angle satellite image dataset, the satellite image data is multispectral. Features are extracted from the satellite image data through a convolutional layer to obtain a feature map F. The feature map F is then averaged spatially to obtain the feature vector f, as shown below:
[0053] H and These represent the height and width of the feature map, respectively. The feature vector f represents the position point at height i and width j. Since the satellite attitude angle changes, the feature vector f of the satellite image data also changes over time. The correction amount is predicted based on the change in feature vector f. Then, the attitude angle is corrected by the correction factor, and the collinearity equation is updated again. When the time series changes are small (such as extreme time periods, such as within a few seconds), the attitude angle correction can be ignored. Since this invention is for dynamic monitoring, the collinearity equation needs to be dynamically corrected over a period of time to accurately achieve the purpose of dynamic monitoring. Although the collinearity equation describes the geometric relationship between image points and object points, in practical applications, factors such as changes in satellite orbit attitude and terrain undulations cause geometric distortion in satellite images. By introducing a correction factor (or using a nonlinear residual term) to correct the collinearity equation, the relationship between image points and object points can be described more accurately, thereby improving monitoring accuracy.
[0054] S2. The cross-viewpoint and spatiotemporal features of the multi-angle satellite image data set are integrated to perform location point feature matching and complete the feature matching of the multi-angle satellite image data set. In some embodiments, the present invention preferably constructs a viewpoint adaptive image registration module. The viewpoint adaptive image registration module uses the aforementioned mapping collinearity equation and a multispectral feature-guided image compensation matching module (preferably using the MSFG-image matching algorithm module) for linked matching. The feature matching method for the multi-angle satellite image data set includes:
[0055] S201. Extract the band spectral characteristics of satellite image data from the multi-angle satellite image data set and combine them with a multilayer perceptron to generate band weights for band feature enhancement. Using an image compensation matching module, based on the spectral characteristics of coal gangue (the spectral characteristics of coal gangue), calculate the global feature vector of the b-th band feature map using band attention. Then, the band weights are generated by combining the multilayer perceptron. In this embodiment, a multilayer perceptron (MLP) is used to generate band weights.
[0056] S202. Within a multi-angle satellite image data set, a Transformer layer is used to perform band feature, spatiotemporal feature, and cross-angle location point feature matching. This invention preferably employs a ResNet50 network (which internally contains a Transformer layer, and the Transformer layer contains a feature matching layer LoFTRLayer). Figure 3 As shown, the feature matching layer performs band feature, spatiotemporal feature, and cross-angle location point feature matching. Band feature matching focuses on band feature matching between satellite image data in a multi-angle satellite image data set. Spatiotemporal feature matching focuses on the matching of the same satellite image data in the time series and / or spatial location within the multi-angle satellite image data set, as well as spatial location matching between satellite image data. Cross-angle location point matching focuses on feature matching between satellite image data in a multi-angle satellite image data set (including key location highlights). High-confidence matching point pairs are selected for geometric constraint compensation, and high-confidence matching point pairs are used to compensate for the geometric constraints of the multi-angle satellite image data set. Preferably, this embodiment introduces the RANSAC algorithm for robust estimation of the fundamental matrix: RANSAC is a robust estimation algorithm used to estimate the fundamental matrix. The reprojection error loss is constructed by dynamically adjusting the error contribution of different bands using spectral weights Wb. ,in For band b, the reprojection error is... The weight of band b is used; through the above radiometric correction and multispectral feature-guided image compensation matching, complex ground features can be effectively identified and matched, improving the robustness and accuracy of coal gangue identification.
[0057] Multi-view stereo is obtained by performing feature extraction, height-sensing encoding, spectrally guided cross-attention processing, and sub-pixel spatial encoding on multi-angle satellite image datasets. In some embodiments, such as Figure 4 As shown, multi-view stereoscopic Methods of obtaining include:
[0058] S211. The feature maps of the multi-angle satellite image data set are used to extract features using convolutional layers and residual blocks to output multi-level feature maps. Simultaneously, height information is extracted from the feature maps of the multi-angle satellite image data set through height-aware encoding and encoded onto the multi-level feature maps, resulting in a preliminary multi-view stereoscopic reconstruction at time t. This embodiment uses a ResNet50 network (which includes several convolutional layers, normalization layers, activation functions, and residual blocks). The ResNet50 network extracts rich image features, including edge, texture, and height information, through convolutional layers, normalization layers, and activation functions. Then, it outputs multi-level feature maps through residual block processing. The feature maps of the satellite image data in the multi-angle satellite image data set are processed by the ResNet50 network to obtain corresponding multi-level feature maps (containing feature information of different scales and levels with rich information). Furthermore, this invention encodes height information onto the multi-level feature maps (ensuring that each pixel in the feature map contains height information, thereby enhancing the expressive ability of the features for the spatial distribution of coal gangue).
[0059] S212, Multi-view stereoscopic view of time t-1 The stereo query input is fed into the subpixel-level spatial alignment module to calculate the subpixel-level offset between time t and time t-1; this invention utilizes multi-view stereo... Alignment iteration and application for multi-view stereo The generation can accurately and quickly obtain multi-view stereoscopic images. This improves model processing speed. Next, the spectral-guided cross-attention module calculates spectral similarity and adjusts the depth and weight of feature extraction in method S211 (the cross-attention module focuses on spectral similarity and adjusts the extraction depth and weight accordingly). The preferred sub-pixel level spatial alignment module of this invention employs the following method:
[0060] Alignment is achieved by constructing a dynamic projection matrix, M. proj It can be decomposed into affine transformations (rotation, translation, scaling) M affine With nonlinear deformation M deform Two parts:
[0061] ,
[0062] , Where a, b, c, and d represent scaling and rotation parameters, respectively. , Let Δa, Δb, Δc, Δd, Δtx, and Δty represent the translation parameters, predicted by the deformation field network with an initial value of 0. The deformation field prediction network takes the multispectral fusion features obtained in the previous step as input and predicts the sub-pixel-level offset of each pixel. First, feature extraction is performed. The multispectral fusion features are input, and the feature Ffeat is extracted through a convolutional layer. Here, Conv2d is a convolutional layer used for feature extraction. Then, the offset O of each pixel is predicted through the convolutional layer: Conv2doffset is a convolutional layer used to predict the offset of each pixel; then feature alignment is performed, which is achieved through deformable convolutional layers based on the predicted offsets. In this layer, DeformCon2d is a deformable convolutional layer used for feature alignment. Finally, based on the dynamic projection matrix M... proj Generate sampling grid G: Where P is the original pixel coordinate. Subpixel-level alignment is achieved through bilinear interpolation. Its BilinearInterpolation represents the bilinear interpolation operation.
[0063] S213. Using the Feed Forward module (employing multiple fully connected neural networks), feature fusion and propagation are performed through several fully connected layers to obtain a multi-view stereo image at time t. (Preferred, such as) Figure 2 As shown, this embodiment also uses the Add & Norm module for normalization processing. The Add & Norm module includes a residual connection module and a layer normalization module. In method S212, a further preferred technical solution is: combining multi-view stereoscopic... Subpixel spatial update encoding is performed on the multi-level feature map after time t adjustment with subpixel-level offset.
[0064] S213. The Feed Forward module utilizes several fully connected layers (connected to contain rich feature information such as multi-view features and height information) to perform feature fusion and propagation to obtain a multi-view stereo image at time t. .
[0065] Multi-view stereo The coal gangue is identified, detected, and segmented separately by homography projection onto a multi-view projection perspective plane. Then, the segments are fused to form a three-dimensional structure of the coal gangue. The volume of the three-dimensional structure is calculated and scaled proportionally to obtain the coal gangue volume. In some preferred embodiments, such as... Figure 5 As shown, the method for obtaining the three-dimensional structure of coal gangue includes:
[0066] S221, Multi-view stereoscopic By using homography projection onto a multi-view projection perspective plane, a target detection submodule is used to identify and detect coal gangue. The target detection submodule is trained using coal gangue sample data.
[0067] S222: The target segmentation submodule identifies and extracts boundary information to segment the coal gangue target. Then, it is 3D reconstructed and fused into a 3D structure of coal gangue.
[0068] In some preferred embodiments, the volume of a single three-dimensional coal gangue structure is calculated in a multi-angle satellite coal gangue inventory monitoring model, and then scaled down to the actual scale of the study area to obtain the corresponding coal gangue volume. The volumes of all three-dimensional coal gangue structures in the study area are calculated, and the corresponding coal gangue volumes and the total volume of all coal gangue are obtained respectively.
[0069] In some preferred embodiments, both the coal gangue identification model and the coal gangue identification and detection using a multi-view stereoscopic projection perspective plane employ the following loss function constraints:
[0070] ;
[0071] ;
[0072] ,in The weights for distance loss and aspect ratio loss are... To predict the bounding box, For the true bounding box, To predict the width of the bounding box, The width of the actual bounding box. To predict the height of the bounding box, The height of the actual bounding box. The scaling parameter is used to scale the distance loss. The scaling parameter is the loss due to scaling the width. The scaling parameter is used to calculate the height loss. The hyperparameters used to control the curvature of the curve, For intersection, union, and comparison, To compare the expected loss with the expected loss, To compare the losses, For distance loss, For aspect ratio loss, The loss result is the loss function. Both the coal gangue identification model and the multi-view projection perspective plane coal gangue segmentation in multi-view stereo model employ the following loss function constraints:
[0073] ,in The loss value is used to measure the predicted segmentation result and the actual segmentation result, where N is the number of pixels in the input image and C is the number of categories in the semantic segmentation task. Let be the predicted probability that the i-th pixel belongs to the j-th class. Let be the true class label of the i-th pixel, set to 1 if it belongs to class j, and 0 otherwise. Gradient descent is used to reduce the model loss value, while simultaneously optimizing and updating the model parameters, until the minimum number of iterations or the total loss for coal gangue identification, detection, and segmentation is reached (or a threshold is set, and iteration ends when the value is less than the threshold).
[0074] The coal gangue volume monitoring and identification results of this invention can be evaluated using evaluation metrics, including: confusion matrix, precision (P), recall (R), F1 score, and coefficient of determination (R²). The confusion matrix displays the correspondence between the model's identification results and the true labels, including the number of true positives, false positives, true negatives, and false negatives. Precision (P) represents the proportion of the volume correctly identified as a coal gangue pile to the predicted volume, used to evaluate the accuracy of the model's identification. Recall (R) represents the proportion of the volume correctly identified as a coal gangue pile to the actual volume, used to evaluate the completeness of the model's identification. The F1 score is the harmonic mean of precision and recall, used to comprehensively evaluate the model's identification performance. The coefficient of determination (R²) is a statistic representing the goodness of fit of the model, used to evaluate the degree of fit between the model's identification results and the true labels. By evaluating the accuracy of the identification results, the performance and effectiveness of the model in the task of identifying coal gangue stockpiles can be comprehensively assessed, providing a basis for subsequent applications and optimizations.
[0075] A multi-angle dynamic monitoring system for coal gangue inventory, comprising a multi-angle satellite coal gangue inventory monitoring model, includes a coal gangue identification model, a mapping collinearity equation module, a feature matching module, a multi-view stereo construction module, and a coal gangue volume calculation module. The coal gangue identification model performs coal gangue identification and detection on the acquired multi-angle satellite image data sets of the study area. The mapping collinearity equation module constructs mapping collinearity equations to convert pixel coordinates into three-dimensional geographic coordinates. The feature matching module is used to perform location point feature matching by integrating cross-view and spatiotemporal features of the multi-angle satellite image data sets and completing feature matching for the multi-angle satellite image data sets. The multi-view stereo construction module is used to extract features, perform height-sensory encoding processing, spectrally guided cross-attention processing, and sub-pixel spatial encoding on the multi-angle satellite image data sets to obtain multi-view stereo data. The coal gangue volume calculation module will use multi-view three-dimensional... The coal gangue is identified, detected, and segmented by homography projection onto a multi-view projection perspective plane. Then, the three-dimensional structure of the coal gangue is fused together, the volume of the three-dimensional structure of the coal gangue is calculated, and the volume of the coal gangue is obtained by scaling it proportionally.
[0076] The above description is only a preferred embodiment of the present invention and is not intended to limit the present invention. Any modifications, equivalent substitutions, and improvements made within the spirit and principles of the present invention should be included within the protection scope of the present invention.
Claims
1. A multi-angle coal gangue stock dynamic monitoring method based on multi-modal spatio-temporal perception, characterized by: The method comprises the following steps: A multi-angle satellite coal gangue stockpile monitoring model comprising a coal gangue recognition model is constructed, the coal gangue recognition model is used for recognizing and detecting coal gangue in a multi-angle satellite image data set of a study area and constructing a mapping collinear equation for converting pixel coordinates into three-dimensional geographic coordinates; The cross-view feature and the spatio-temporal feature of the multi-angle satellite image data set are matched by position point feature matching, and the feature matching of the multi-angle satellite image data set is completed, a view angle adaptive image registration module is constructed, the view angle adaptive image registration module is matched by linkage of the mapping collinear equation and a multi-spectral feature guided image compensation matching module, and the feature matching method of the multi-angle satellite image data set comprises the following steps: S201, the band spectral characteristics of the satellite image data in the multi-angle satellite image data set are extracted, a multi-layer perception machine is used to generate band weights, and band feature enhancement is performed; S202, the band feature, the spatio-temporal feature and the cross-angle position point feature in the multi-angle satellite image data set are matched by using a Transformer layer, high-confidence matching point pairs are screened for geometric constraint compensation, and the geometric constraint compensation of the multi-angle satellite image data set is performed by using the high-confidence matching point pairs. The multi-angle satellite image data set is subjected to feature extraction, height perception coding processing, spectrum-guided cross-attention processing and sub-pixel spatial coding to obtain a multi-view stereo , multi-view stereo The method comprises the following steps: using a convolution layer and a residual block to extract features of a feature map of a multi-angle satellite image data set to output a multi-level feature map, and simultaneously performing height perception coding processing to extract height information of the feature map of the multi-angle satellite image data set and encode the height information on the multi-level feature map to obtain a preliminary multi-view stereo at time t of three-dimensional reconstruction; inputting the multi-view stereo at time t-1 into a sub-pixel level spatial alignment module to calculate a sub-pixel level offset between time t and time t-1, then inputting the sub-pixel level spatial alignment module into a spectrum-guided cross-attention module to calculate a spectrum similarity and adjust a depth and a weight of feature extraction; performing feature fusion and propagation through a Feed Forward module to obtain the multi-view stereo at time t; inputting the multi-view stereo at time t into a multi-view stereo reconstruction module to obtain a multi-view stereo at time t+1 ; projecting the multi-view stereo at time t+1 to a multi-view projection perspective plane through homography projection, and performing coal gangue recognition, detection and segmentation processing, respectively, then fusing to form a coal gangue three-dimensional structure, calculating a volume of the coal gangue three-dimensional structure and obtaining a coal gangue volume amount through scaling according to a proportion. ; projecting the multi-view stereo at time t+1 to a multi-view projection perspective plane through homography projection, and performing coal gangue recognition, detection and segmentation processing, respectively, then fusing to form a coal gangue three-dimensional structure, calculating a volume of the coal gangue three-dimensional structure and obtaining a coal gangue volume amount through scaling according to a proportion. 2. The multi-angle coal gangue stock dynamic monitoring method based on multi-modal spatio-temporal perception according to claim 1, characterized in that: The coal gangue recognition model is trained by using a multi-angle coal gangue sample data set, the multi-angle coal gangue sample data set comprises coal gangue image samples and coal gangue target labeling information, the coal gangue target labeling information comprises a labeling detection frame and a target category, the target category comprises coal gangue piles and scattered coal gangue, and the coal gangue image samples comprise coal gangue satellite images in various weather, illumination or / and seasonal environments; the coal gangue image samples of the multi-angle coal gangue sample data set and the satellite image data of the multi-angle satellite image data set are subjected to denoising processing, and the coal gangue image samples of the multi-angle coal gangue sample data set are subjected to data enhancement processing.
3. The multi-angle coal gangue stock dynamic monitoring method based on multi-modal spatio-temporal perception according to claim 1, characterized in that: The mapping collinear equation expression for converting the pixel coordinates (u, v) of the satellite image data in the multi-angle satellite image data set into three-dimensional geographic coordinates (X, Y, Z) is as follows: where (θ ) is the initial attitude angle of the satellite, (r ) is the position of the satellite, (R ) is the rotation matrix element composed of attitude angles; the satellite image data in the multi-angle satellite image data set is averaged in the spatial dimension to obtain the feature vector f through the convolution layer to obtain the attitude angle correction amount (Δθ ) to correct the attitude angle and update the mapping collinear equation.
4. The multi-angle coal gangue stock dynamic monitoring method based on multi-modal spatio-temporal perception according to claim 1, characterized in that: Combination of multi-view stereos Sub-pixel spatial update coding is performed on the multi-level feature map after adjustment of the sub-pixel level offset with respect to time t.
5. The multi-modal spatio-temporal perception based multi-angle coal gangue stock dynamic monitoring method according to claim 1, characterized in that: The coal gangue three-dimensional structure obtaining method comprises the following steps: S221、the multi-view stereoscopic By homography projection to multi-view projection perspective plane, target detection submodule is carried out respectively to carry out coal gangue recognition detection processing, target detection submodule utilizes coal gangue sample data to learn training. S222, the coal gangue target is segmented by boundary information recognition and extraction by a target segmentation sub-module, and then three-dimensional reconstruction and fusion are performed to obtain a coal gangue three-dimensional structure.
6. The multi-modal spatio-temporal perception based multi-angle coal gangue stock dynamic monitoring method according to claim 1, characterized in that: The volume of a single coal gangue three-dimensional structure is calculated in the multi-angle satellite coal gangue stockpile monitoring model, and then the volume is scaled to the actual scale of the study area to obtain a corresponding coal gangue volume; The volumes of all coal gangue three-dimensional structures in the study area are calculated to obtain corresponding coal gangue volumes and a total amount of all coal gangue volumes.
7. The multi-modal spatio-temporal perception based multi-angle coal gangue stock dynamic monitoring method according to claim 1, characterized in that: The coal gangue recognition model and the coal gangue recognition detection of the multi-view stereoscopic multi-view projection perspective plane are both constrained by the following loss function: ; ; wherein is a weight of distance loss and aspect ratio loss, is a predicted bounding box, is a real bounding box, is a width of the predicted bounding box, is a width of the real bounding box, is a height of the predicted bounding box, is a height of the real bounding box, is a scale parameter of the scaled distance loss, is a scale parameter of the scaled width loss, is a scale parameter of the scaled height loss, is a hyper parameter to control the curvature of the curve, is an intersection over union, is an expected intersection over union loss, is an intersection over union loss, is a distance loss, is an aspect ratio loss, is a loss result of the loss function; The coal gangue segmentation of the coal gangue recognition model and the multi-view stereoscopic multi-view projection perspective plane is constrained by the following loss function: wherein is a loss value for measuring the predicted segmentation result and the real segmentation result, N is the number of pixels divided by the input image, C is the number of classes in the semantic segmentation task, is the probability of the i-th pixel belonging to the j-th class, is the real class label of the i-th pixel, and is 1 when belonging to the j-th class, otherwise 0; The gradient descent algorithm is used to reduce the model loss value, and the model parameters are optimized and updated until the iteration number or the total loss of the coal gangue recognition detection and segmentation is minimized.
8. A multi-angle coal and gangue stock dynamic monitoring system for implementing the multi-angle coal and gangue stock dynamic monitoring method of claim 1, characterized in that: The multi-angle satellite coal gangue stock monitoring model comprises a coal gangue recognition model, a mapping collinear equation module, a feature matching module, a multi-view stereo construction module, and a coal gangue volume calculation module. The coal gangue recognition model performs coal gangue recognition detection on the obtained multi-angle satellite image data set of the research area. The mapping collinear equation module constructs a mapping collinear equation for converting pixel coordinates into three-dimensional geographic coordinates. The feature matching module is used for matching the position point features of the multi-angle satellite image data set by comprehensively matching the cross-view features and the space-time features of the multi-angle satellite image data set, and completing the feature matching of the multi-angle satellite image data set. The multi-view stereo construction module is used for feature extraction, height perception coding processing, spectrum-guided cross-attention processing, and sub-pixel space coding on the multi-angle satellite image data set to obtain a multi-view stereo ; and the coal gangue volume calculation module calculates the volume of the multi-view stereo by performing homographic projection to a multi-view projection perspective plane and performing coal gangue recognition detection and segmentation processing, respectively, and then fusing to form a coal gangue three-dimensional structure, calculating the volume of the coal gangue three-dimensional structure, and scaling the volume to obtain the coal gangue volume.
Citation Information
Patent Citations
Coal gangue detection method based on multi-angle perception and mixed scale Transform feature aggregation
CN119399493A
Generating high-resolution concentration maps for atmospheric gases using geography-informed machine learning
US20240046143A1